Audio apparatus and method of operation therefor

WO2026180487A1PCT designated stage Publication Date: 2026-09-03KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/055080
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-27
Filing Date
2026-02-25
Publication Date
2026-09-03

Smart Images

  • Figure EP2026055080_03092026_PF_FP_ABST
    Figure EP2026055080_03092026_PF_FP_ABST
Patent Text Reader

Abstract

: An audio apparatus has a receiver (301) receiving a data signal including encoded data for a mono downmix audio signal and spatial upmix parameters for upmixing the downmix to a stereo signal. It further includes a directional circuit (305) estimating a directional signal of a signal model for a stereo signal, and a generator 317 generating two estimated signals being residual signals of the signal model. A rendering function (307, 309, 311, 313, 319) renders a stereo signal from the signal model. An adapter (321) controls the generator (317) to adapt a relative property, such as specifically a correlation or relative signal level, between the first estimated signal and the second estimated signal in dependence on the spatial upmix parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AUDIO APPARATUS AND METHOD OF OPERATION THEREFOR

[0002] FIELD OF THE INVENTION

[0003] The invention relates to an audio apparatus and method of operation therefor, and specifically, but not exclusively, to rendering stereo signals using e.g. Parametric Stereo based encoding.

[0004] BACKGROUND OF THE INVENTION

[0005] Spatial audio applications have become numerous and widespread and increasingly form part of many audiovisual experiences. New and improved spatial experiences and applications are continuously being developed which result in increased demands on the audio processing and rendering.

[0006] A lot of research and development effort has focused on providing efficient and high quality audio encoding and audio decoding for spatial audio. A frequently used spatial audio representation is multichannel audio representations, including stereo representation, and efficient encoding of such multichannel audio based on downmixing multichannel audio signals to downmix channels with fewer channels have been developed. One of the main advances in low bit-rate audio coding has been the use of parametric multichannel coding where a downmix signal is generated together with parametric data that can be used to upmix the downmix signal to recreate the multichannel audio signal.

[0007] In particular, instead of traditional mid-side or intensity coding, in parametric multichannel audio coding, a multichannel input signal is downmixed to a lower number of channels (e.g. two to one) and multichannel image (stereo) parameters are extracted. Then the downmix signal is encoded using a more traditional audio coder (e.g. a mono audio encoder). The bitstream of the downmix is multiplexed with the encoded multichannel image parameter bitstream. This bitstream is then transmitted to the decoder, where the process is inverted. First the downmix audio signal is decoded, after which the multichannel audio signal is reconstructed guided by the encoded multichannel image / upmix parameters.

[0008] An example of stereo coding is described in E. Schuijers, W. Oomen, B. den Brinker, J. Breebaart, “Advances in Parametric Coding for High-Quality Audio”, 114th AES Convention, Amsterdam, The Netherlands, 2003, Preprint 5852. In the described approach, the downmixed mono signal is parametrized by exploiting the natural separation of the signal into three components (objects): transients, sinusoids, and noise. In E. Schuijers, J. Breebaart, H. Pumhagen, J. Engdegard, “Low Complexity Parametric Stereo Coding”, 116th AES, Berlin, Germany, 2004, Preprint 6073 more details are provided describing how parametric stereo was realized with a low (decoder) complexity when combining it with Spectral Band Replication (SBR).Parametric Stereo (PS) is a technology which is widely used to efficiently code a stereo signal as a mono downmix and a set of spatial parameters allowing an accurate reconstruction of the stereo image. PS has been used to substantially improve the compression efficiency for AAC and USAC at lower bit-rates, say 32kbps and below.

[0009] In addition to accurately reproducing a stereo signal, it has also been of interest to create high quality binaural rendering of (encoded) stereo signals to emulate a virtual loudspeaker playback.

[0010] Binaural rendering of content authored for multi-channel playback can be achieved by the sum of convolutions of the input channel signals with left and right Head Related Impulse Responses (HRIRs), where each HRIR pair corresponds to a measured / simulated impulse response from a loudspeaker location to the ears. This can be expressed compactly in the z-domain as:

[0011] YL, R(Z) = Xc(z) ■ Hj* (z)

[0012]

[0013] Vc

[0014] where Xc(z) represents the z-transform of the time domain input signal xc[n] with channel c, YL R(z) represents the z-transform of the left and right time domain output signals l[n] and r[n] respectively andR(z) is the z-transform of the HRIR of the left and right channels h| [n] and hr[n] for the angle (and distance) corresponding to loudspeaker position <pc. Approaches for rendering binaural stereo are disclosed in WO2010 / 122455 A 1 and W02007 / 031896 A 1.

[0015] A particular approach for rendering stereo audio using headphones is presented in J. Breebaart and E. Schuijers. " Phantom materialization: A novel method to enhance stereo audio reproduction on headphones.", IEEE transactions on audio, speech, and language processing 16.8 (2008): 1503-1511. The approach seeks to provide accurate rendering of both audio point sources and more ambient and background audio.

[0016] However, whereas current approaches for audio rendering may provide acceptable performance in many applications and scenarios, they tend to not be ideal and may exhibit suboptimal behavior in some scenarios. In particular, it may result in suboptimal perceived quality and / or a reduced user experience with e.g. perceived suboptimal spatial perception / audio scene in some cases. Complexity and / or resource usage may also be higher than desired and may in some case make the approach undesired or impractical for some implementations, such as applications based on small and cheap portable devices.

[0017] Hence, an improved approach would be advantageous. In particular an approach allowing increased flexibility, improved adaptability, improved performance, increased audio quality, improved perceived quality, an improved rendering / generation of a stereo signal, improved spatial perception, reduced complexity and / or resource usage, reduced computational load, facilitated implementation, improved user experience, and / or an improved spatial audio experience would be advantageous.SUMMARY OF THE INVENTION

[0018] Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination.

[0019] According to an aspect of the invention there is provided an audio apparatus comprising: a receiver arranged to receive a data signal comprising encoded data for a mono downmix audio signal (being a downmix of a stereo signal) and a set of spatial upmix parameters for upmixing the mono downmix audio signal to a stereo signal, the set of spatial upmix parameters being indicative of relative signal properties of channels of the stereo signal; a store comprising directional transfer functions for different directions, a directional transfer function for a given direction representing a mapping of a mono audio signal to stereo channels such that the mono audio signal is positioned in the given direction in a stereo image of the stereo channels; a decoder arranged to generate the mono downmix audio signal by decoding the encoded data; a directional circuit arranged to estimate a directional signal, the directional signal representing a point source audio component of the stereo signal; a decorrelator arranged to generate a decorrelated mono downmix audio signal from the mono downmix audio signal; a generator arranged to generate a first estimated signal and a second estimated signal from the mono downmix audio signal, at least the first estimated signal further being generated from the decorrelated mono downmix audio signal and the directional signal, the first estimated signal, and the second estimated signal forming a(n estimated) decomposition of the stereo signal; a direction determining circuit arranged to determine a first direction from the spatial parameters; a first Tenderer arranged to perform a first rendering of the mono downmix audio signal to generate a first intermediate stereo signal, the first rendering being a directional rendering using a first channel transfer function retrieved from the store for the first direction; a second Tenderer arranged to perform a second rendering being a rendering of the first estimated signal and the second estimated signal to generate a second intermediate stereo signal; a combiner arranged to combine at least the first intermediate stereo signal and the second intermediate stereo signal to generate an output stereo signal; and an adapter arranged to control the generator to adapt a first relative property being a relative property between the first estimated signal and the second estimated signal in dependence on the spatial parameters.

[0020] The approach may provide an improved audio experience in many embodiments. For many signals and scenarios, the approach may provide improved rendering of a stereo audio signal allowing improved generation / reconstruction of a stereo audio signal with an improved perceived audio quality. The approach may provide improved representation of an audio scene by a stereo signal. The approach may in many embodiments allow an improved and / or more attractive spatial audio experience from a stereo signal.

[0021] The approach may provide an efficient implementation and may in many embodiments allow reduced complexity and / or resource usage. The approach may in many scenarios allow a reduced computational burden while providing a perceived high quality rendering of a stereo signal, and inparticular based on an efficient representation of a stereo signal using a downmix and spatial upmix parameters, such as specifically a PS encoded stereo signal.

[0022] The processing may be in time frequency segments or tiles. Each time frequency segment / tile may represent a frequency interval in a time interval. In many embodiments, the mono downmix audio signal may be divided into time segments / intervals and a frequency representation of the signal in the time segment / interval may be provided by signal values representing different frequency segments of the signal in the time segment / interval. Some or all of the processing may be performed in the frequency domain / frequency subbands.

[0023] The spatial upmix parameters may comprise sets of upmix parameters, each set of upmix parameters comprising at least one of: a level difference parameter indicative of a level difference between channels of the multichannel audio signal; a correlation parameter indicative of a coherence between channels of the multichannel audio signal; a timing difference parameter indicative of a timing difference between channels of the multichannel audio signal, and a phase difference parameter indicative of a phase difference between channels of the multichannel audio signal.

[0024] The first direction may be a desired / target rendering direction. A direction may be an angle and / or orientation from a listening position / in a stereo image.

[0025] The directional rendering may be a binaural rendering generating the first intermediate stereo signal as a binaural stereo signal comprising a point source positioned in the first direction, the binaural rendering comprising selecting binaural impulse response values as values for a binaural impulse response for a sound source in the first direction. The directional transfer functions may be parameterized transfer functions, and specifically may be represented in the frequency domain as weights for each of a plurality of subbands. A weight may be provided for each stereo channel. The weights may typically be complex valued.

[0026] The binaural impulse response values may be parametric values and may be frequency tile values. The binaural impulse response values may be values representing any suitable binaural impulse response in any suitable way, including HRIR, HRTF, BRIR values etc.

[0027] In many embodiments, the encoded data and the spatial upmix parameters are part of a Parametric Stereo encoding of the stereo signal.

[0028] In some embodiments, the second rendering is arranged to generate the second intermediate stereo signal using a set of directional transfer functions retrieved from the store for a set of predetermined directions.

[0029] In many scenarios and / or embodiments, the first and second estimated signals are estimates of audio components of the mono downmix audio signal not included in the directional signal. The first estimated signal may be a first estimated signal. The second estimated signal may be a second estimated signal.

[0030] In many embodiments, the first and / or second rendering is a binaural rendering and the directional transfer functions are binaural transfer functions.In some embodiments, the second rendering is arranged to generate the second intermediate stereo signal using a set of directional transfer functions retrieved from the store for a set of predetermined directions.

[0031] This may provide an advantageous approach for many scenarios, including e.g. providing an advantageous trade-off between complexity, computational resources, data rate and / or the perceived audio quality of the generated output stereo signal.

[0032] In some embodiments, the set of predetermined directions consists of one predetermined direction.

[0033] This may provide a particularly advantageous trade-off between complexity and user experience / perceived audio quality in many scenarios and embodiments.

[0034] In some embodiments, the set of predetermined directions comprises a plurality of predetermined directions.

[0035] This may provide a particularly advantageous trade-off between complexity and user experience / perceived audio quality in many scenarios and embodiments.

[0036] In some embodiments, the spatial upmix parameters and the directional transfer functions are provided for frequency subbands and the first Tenderer is arranged to generate subband values for subbands of the first intermediate stereo signals from subband values of the modified mono downmix audio signal based on spatial upmix parameters and directional transfer functions for the subbands.

[0037] This may provide improved audio rendering in many embodiments. It may in many scenarios and applications allow an improved and / or more attractive spatial audio experience from a stereo signal.

[0038] In some embodiments, the direction determining circuit is arranged to determine a point source direction in a stereo image of the stereo signal from the spatial upmix parameters, and to determine the first direction by applying a mapping function to the point source direction.

[0039] This may provide an advantageous approach for many scenarios, including e.g. providing an advantageous trade-off between complexity, computational resources, data rate and / or the perceived audio quality of the generated output stereo signal.

[0040] The mapping may be non-uniform. The mapping may be from one range to a different range (e.g. from [0,90°] to [-30°,30°]). The mapping may be non-linear.

[0041] In some embodiments, the direction determining circuit may be arranged to determine an indication of the point source direction as:

[0042] / 1 — IID + v / (IID — I)2+ 4 • ICC2■ IID \

[0043] 7 = arctan - - - - I

[0044] 2 ■ ICC • A / IID

[0045]

[0046] Jwhere IID is an interchannel intensity difference and ICC is an inter-channel cross-correlation. The point source direction may be derived by interpreting y as a relative angle between two loudspeakers (y = 0° is one speaker, 90° is the other). In typical notation a front-left speaker is typically at 30 degrees, and frontright at -30 degrees.

[0047] According to an optional feature of the invention, the first relative property includes a level of the first estimated signal relative to the second estimated signal.

[0048] This may provide improved audio rendering in many embodiments. It may in many scenarios and applications allow an improved and / or more attractive spatial audio experience from a stereo signal. It may provide an advantageous approach for many scenarios, including e.g. providing an advantageous trade-off between complexity, computational resources, data rate and / or the perceived audio quality of the generated output stereo signal.

[0049] According to an optional feature of the invention, the first relative property includes a correlation between the first estimated signal relative to the second estimated signal.

[0050] This may provide improved audio rendering in many embodiments. It may in many scenarios and applications allow an improved and / or more attractive spatial audio experience from a stereo signal. It may provide an advantageous approach for many scenarios, including e.g. providing an advantageous trade-off between complexity, computational resources, data rate and / or the perceived audio quality of the generated output stereo signal.

[0051] According to an optional feature of the invention, the adapter is arranged to adapt a second relative property being a relative property between the directional signal and the first estimated signal in dependence on the spatial parameters.

[0052] This may provide improved audio rendering in many embodiments. It may in many scenarios and applications allow an improved and / or more attractive spatial audio experience from a stereo signal.

[0053] According to an optional feature of the invention, the second relative property includes a level of the directional signal relative to the first estimated signal.

[0054] The approach may allow an improved audio quality to data rate relationship in many scenarios and applications.

[0055] According to an optional feature of the invention, the second relative property includes a correlation between the directional signal relative to the first estimated signal.

[0056] The approach may allow an improved audio quality in many scenarios and applications. According to an optional feature of the invention, at least the directional signal and the first estimated signal are each a linear combination of the mono downmix audio signal and the decorrelated mono downmix audio signal.

[0057] This may provide improved audio rendering in many embodiments. It may in many scenarios and applications allow an improved and / or more attractive spatial audio experience from astereo signal. It may provide an advantageous approach for many scenarios, including e.g. providing an advantageous trade-off between complexity, computational resources, data rate and / or the perceived audio quality of the generated output stereo signal.

[0058] The second estimated signal may also typically be a linear combination of the mono downmix audio signal and the decorrelated mono downmix audio signal.

[0059] According to an optional feature of the invention, the adapter is arranged to determine weights of the linear combination as a function of the spatial parameters.

[0060] This may provide improved audio rendering in many embodiments. It may in many scenarios and applications allow an improved and / or more attractive spatial audio experience from a stereo signal. It may provide an advantageous approach for many scenarios, including e.g. providing an advantageous trade-off between complexity, computational resources, data rate and / or the perceived audio quality of the generated output stereo signal.

[0061] In some embodiments, the directional circuit is arranged to generate the directional signal by applying first frequency band weights to frequency band samples of the mono downmix audio signal, the first frequency band weights being dependent on the spatial parameters.

[0062] In many embodiments no other signals than the mono downmix audio signal are included in the generation of the directional signal. In many embodiments, the first frequency band weights are dependent only on the spatial parameters.

[0063] In some embodiments, the frequency band weights may be determined substantially as:

[0064]

[0065] or

[0066] (IID - I)2+ 4 • ICC2• IID + (1 - IID) • ^4 - ICC2- IID + (IID - I)2

[0067] 9x —

[0068] 1 - IID2+ (IID + 1) ■ ■ ICC2• IID + (IID - l)2

[0069]

[0070] where IID is an interchannel intensity difference and ICC is an inter-channel cross-correlation and make up part of the spatial parameters.

[0071] According to an optional feature of the invention, the directional signal, the first estimated signal, and the second estimated signal are generated to be estimates of a signal model for the stereo signal, the signal model representing the stereo signal as a matrix multiplication of a signal vectorby a matrix having a rank of two, the signal vector comprising a directional signal and two diffuse signals.

[0072] The matrix may be a 2x3 matrix.

[0073] In some embodiments, the directional signal, the first estimated signal, and the second estimated signal are generated to be estimates of a signal model for the stereo signal, the signal model representing the stereo signal as a matrix multiplication of a signal vector comprising a directional signal and two diffuse signals by a matrix having a rank of two.

[0074] This may provide particularly advantageous operation and performance in many embodiments and scenarios.

[0075] According to an optional feature of the invention, the adapter is arranged to adapt the first relative property to match a corresponding relative property between the two diffuse signals of the signal model.

[0076] This may provide particularly advantageous operation and performance in many embodiments and scenarios.

[0077] According to an optional feature of the invention, the generator is arranged to generate the first estimated signal by applying second frequency band weights to frequency band samples of the decorrelated mono downmix audio signal, the second frequency band weights being dependent on the spatial parameters.

[0078] This may provide particularly advantageous operation and performance in many embodiments and scenarios.

[0079] In many embodiments no other signals than the decorrelated mono downmix audio signal are included in the generation of the first estimated signal. In many embodiments, the second frequency band weights are dependent only on the spatial parameters.

[0080] According to an optional feature of the invention, the generator is arranged to generate the second estimated signal from the decorrelated mono downmix audio signal.

[0081] This may provide particularly advantageous operation and performance in many embodiments and scenarios.

[0082] According to an optional feature of the invention, the generator is arranged to generate the second estimated signal by applying third frequency band weights to frequency band samples of the decorrelated mono downmix audio signal, the third frequency band weights being dependent on the spatial parameters.

[0083] This may provide particularly advantageous operation and performance in many embodiments and scenarios.

[0084] In many embodiments no other signals than the decorrelated mono downmix audio signal are included in the generation of the second estimated signal. In many embodiments, the third frequency band weights are dependent only on the spatial parameters.According to another aspect of the invention, there is provided method of operation for an audio apparatus, the method comprising: receiving a data signal comprising encoded data for a mono downmix audio signal and a set of spatial upmix parameters for upmixing the mono downmix audio signal to a stereo signal, the set of spatial upmix parameters being indicative of relative signal properties of channels of the stereo signal; providing directional transfer functions for different directions, a directional transfer function for a given direction representing a mapping of a mono audio signal to stereo channels such that the mono audio signal is positioned in the given direction in a stereo image of the stereo channels; generating the mono downmix audio signal by decoding the encoded data; estimating a directional signal, the directional signal representing a point source audio component of the stereo signal; generating a decorrelated mono downmix audio signal from the mono downmix audio signal; generating a first estimated signal and a second estimated signal from the mono downmix audio signal, at least the first estimated signal further being generated from the decorrelated mono downmix audio signal and the directional signal, the first estimated signal, and the second estimated signal forming a decomposition of the stereo signal; determining a first direction from the spatial parameters; performing a first rendering of the mono downmix audio signal to generate a first intermediate stereo signal, the first rendering being a directional rendering using a first channel transfer function retrieved from the store for the first direction; performing a second rendering being a rendering of the first estimated signal and the second estimated signal to generate a second intermediate stereo signal; combining at least the first intermediate stereo signal and the second intermediate stereo signal to generate an output stereo signal; and controlling the generator to adapt a first relative property being a relative property between the first estimated signal and the second estimated signal in dependence on the spatial parameters.

[0085] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter.

[0086] BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Embodiments of the invention will be described, by way of example only, with reference to the drawings, in which

[0088] FIG. 1 illustrates some elements of an example of an audio distribution system;

[0089] FIG. 2 illustrates some elements of an example of an audio apparatus in accordance with some embodiments of the invention;

[0090] FIG. 3 illustrates some elements of an example of an audio render apparatus in accordance with some embodiments of the invention; and

[0091] FIG. 4 illustrates some elements of a possible arrangement of a processor for implementing elements of an audio apparatus in accordance with some embodiments of the invention.DETAILED DESCRIPTION OF SOME EMBODIMENTS OF THE INVENTION

[0092] FIG. 1 illustrates an example of an audio system wherein a stereo signal may be distributed / communicated for remote rendering. In the system, an audio apparatus referred to as the audio source device 101 generates an audio data signal including a representation of a stereo audio signal. The stereo audio signal may be one captured at the audio source device 101, may be received from another source, or indeed may e.g. be an artificially generated stereo signal (e.g. it may be a virtual audio stereo signal).

[0093] The audio source device 101 may generate the data signal to include encoded data that represents a mono downmix audio signal for the stereo signal. For example, the mono downmix audio signal may be generated as a weighted combination, and specifically as a weighted summation, of the channel signals of an input stereo signal. In many cases, the weights may be fixed and may specifically be the same for the two channel signals.

[0094] The generated mono downmix audio signal is encoded using a suitable mono audio encoding algorithm / standard to generate encoded audio data representing the mono downmix audio signal.

[0095] In addition to the mono downmix audio signal, the audio source device 101 generates spatial upmix parameters for upmixing the mono downmix audio signal to recreate the original stereo signal.

[0096] The spatial upmix parameters are generated to be indicative of / reflect relative properties of the channel signals of the stereo audio signal. In particular, the spatial upmix parameters may be indicated to include parameters that are indicative of at least one of relative intensities / levels of the stereo channels, relative (frequency domain) phases of the stereo channels, a relative time difference between the channels, and / or a correlation between the channels. Specifically, the audio source device 101 may generate spatial upmix parameters including one or more of an inter-channel intensity difference, interchannel level difference, inter-channel time difference, inter-channel phase difference, and / or interchannel correlation.

[0097] The data signal may specifically comprise a Parametric Stereo (PS) encoding of the stereo signal.

[0098] A classical PS downmix is calculated as:

[0099] m = c(l + r)

[0100] where the parameter c is chosen such that the power of the stereo signal is preserved in the downmix, the power being defined using the 2 -norm:

[0101] ||m||2= ||Z||2+ ||r||2

[0102] and thus e.g.:c = √( ||l||2+ ||r||2 / ||l+r||2)

[0103] c =

[0104]

[0105] 4||l+r||

[0106] The PS parameters are specifically an Inter-channel Intensity Difference IID, an Interchannel Correlation ICC, and in some cases an Inter-channel Phase Difference IPD parameter. These may specifically be defined as:

[0107]

[0108] where the complex-valued inner product is defined as:

[0109]

[0110] and:

[0111] | |x| |2=< x,x >

[0112] The spatial upmix parameters are typically generated for specific time frequency tiles, and thus specifically each parameter value is generated / provided for a given frequency subband / interval and for a given time segment / interval.

[0113] The audio source device 101 may accordingly encode a stereo signal as encoded data representing a mono downmix audio signal of the stereo signal and associated spatial upmix parameters that are indicative of relative properties of the channels (the channel signals) of the stereo signal.

[0114] Specifically, the audio source device 101 may be arranged to generate a data signal comprising a conventional PS encoded stereo signal.

[0115] The system of FIG. 1 further comprises an audio apparatus which henceforth will be referred to as the audio render apparatus 103. The audio render apparatus 103 is arranged to receive the data signal generated by the audio source device 101. In the example, the audio render apparatus 103 and audio source device 101 are both coupled to a network 105 through which the data signal can becommunicated and specifically through which it can be communicated from the audio source device 101 to the audio render apparatus 103. The network 105 may specifically be, or include, the Internet.

[0116] Thus, the receiver 101 receives a data signal which comprises encoded data for a mono downmix audio signal and spatial upmix parameters for upmixing the mono downmix audio signal to the stereo signal. The set of spatial upmix parameters comprises one or more parameters indicative of relative signal properties of channels of the stereo signal, and may specifically be indicative of a level / intensity difference between the channels of the stereo signal, a cross-correlation between the channels of the stereo signal, and / or a phase difference or time difference between the channels of the stereo signals. In many cases, the data signal may comprise ICC, IID and / or IPD parameters. The data signal may specifically comprise a PS (Parametric Stereo) encoded stereo signal. The audio source device 101 may be arranged to generate a data signal which comprises a stereo signal encoded in accordance with the Parametric Stereo (PS) specifications / standard. The receiver 101 may accordingly receive a representation of a stereo signal encoded by a mono downmix audio signal and spatial upmix parameters, and specifically a PS encoded stereo signal.

[0117] The audio render apparatus 103 is arranged to process the data signal to render a stereo signal. An audio render apparatus could render the stereo signal of the data signal using a conventional PS rendering approach based on a PS upmixing of the mono downmix audio signal using the PS spatial parameters. However, the audio render apparatus 103 of FIG. 1 uses a specific approach where point source components may be specifically considered, and where in many embodiments different / parallel paths process the received mono downmix audio signal in different ways to generate different stereo signal components which are then combined to generate the output stereo signal.

[0118] The rendering by the audio render apparatus 103 may include a differentiated rendering of different audio components with different properties, and in particular may seek to render some components corresponding to point audio sources with a specific directional property while other components may be rendered less spatially specific including in particular as spatially diffuse components. The approach may be based on considering a signal model for a stereo signal that may allow decomposition of the stereo signal into different signal components.

[0119] A signal model may specifically consider the stereo signal to be a linear combination of three signal components s, di, and dr:

[0120]

[0121] However, one problem with such an approach is that it is in principle not always possible to determine the opposite relationship, i.e. to determine the three signal components from the stereo signals (in principle it is not possible to determine three independent components from two signalcomponents). In other words, it is generally not possible to perform a signal inversion as the inverse matrix G of the signal matrix F cannot be determined.

[0122] In order to perform a decomposition of a stereo signal into three components, it may accordingly be required to include some assumptions or restrictions on the signal model.

[0123] A possible signal model for a stereo signal may be represented by:

[0124] l = cos(γ)ejφs + dl

[0125] r

[0126]

[0127] = sin(γ)ejφs + dr

[0128] This signal model essentially represents a consideration that the stereo signal corresponds to a combination of a directional signal and diffuse background audio. The directional signal s is an audio component that corresponds / reflects audio that is considered to originate from an audio source that has a point source property / characteristic and which thus has a well-defined spatial origin. The audio of the directional signal thus corresponds to audio that can / should be rendered to be perceived to reach the listener from a specific direction. The directional signal may for each time frequency tile represent a single audio point source and thus may for each time frequency tile have a specific source position. The directional signal may in some cases / embodiments correspond to different audio point sources in different time frequency tiles, i.e. it is not required that all time frequency tiles of the directional source represent audio from the same point source. In some cases, the directional signal may correspond to different frequency components being reflected differently in the environment and thus may (e.g. in some frequency bands) correspond to early reflections of the point source within the acoustic environment.

[0129] In the signal model, the directional signal component s is phase shifted using two parameters φland φr. and is further panned / positioned in the stereo image of the original stereo channels I and r. The panning is to an angle represented by the panning angle y. Furthermore, a diffuse signal component is represented by diffuse residual signal components di and drof the respective left and right channels.

[0130] Such a model may be used by a render apparatus to generate estimates of the three signal decomposition signals. Further, these may be generated based on a single received downmix m and two decorrelated signals d1and d2. In particular, a render apparatus may proceed to generate estimates of the decomposition signals as:

[0131] 9 s 0

[0132] 0 9d 0

[0133]

[0134] 0 0 gd] · [d2]The approach thus generates a main / directional signal.s' ’ and two residual signals as decorrelated signals thereby allowing them to be generated from the received mono signal. Further, the gains may be selected to seek to maintain properties corresponding to the stereo signal. In particular, with ||d1||2= ||d2||2= ||m||2, the approach may reconstruct level properties of the underlying signal model, i.e.

[0135] ||s'||2= ||s||2,

[0136] ||dl'||2= ||dl||2,

[0137] ||dr'||2= ||dr||2

[0138]

[0139] KI| = |KI| •

[0140] In addition to the generated decomposition signals being decorrelated, the signal levels of the generated residual signals are the same ||dl'||2= ||dr||2.

[0141] However, whereas such an approach may provide an advantageous decomposition that may allow improved rendering in many cases, it tends to not always provide optimum audio quality.

[0142] A signal model for decomposition of a stereo signal ( / ,r) may be based on considering the linear combination:

[0143]

[0144] but with a matrix F having a rank of two. This allows a signal / matrix inversion allowing the decomposition signals to be determined from the stereo signals:

[0145]

[0146] Based on this consideration, a Tenderer may seek to generate local estimates / replicas of the decomposition signals from the received mono downmix audio signal and a decorrelated signal:

[0147] [s'] = [g_s 0 ] · [m][d_l'] [g_{m,l} g_{d,l}] [d][d_r'] [g_{m,r} g_{d,r}]

[0148]

[0149] -9m,r 9d,r.

[0150] The matrix coefficients may be adapted to seek to maintain the properties of the underlying signal model. Specifically, with | |d| |2= 11 m 112this may allow reconstruction of a range of properties:||s'||2= ||s||2||dl'||2= ||dl||2||dr'||2= ||dr||2

[0151] < s', di >=< s, di >

[0152]

[0153] It should be noted that in this case the correlation between the residual signals (i.e. the last correlation) is typically not 0 (in contrast to the previous signal model where it is an inherent and fundamental property / assumption of the model). It is also noted that the levels of the residual signals are not (necessarily) the same (in contrast to the previous signal model where it is an inherent and fundamental property / assumption of the model).

[0154] Accordingly, whereas both models may reconstruct the properties of the signal components of the underlying model, the Inventors have realized that improved decomposition into specific signal components can be achieved based on the latter model considerations and that this in particular is suitable for a particular and differentiated rendering as will be described in detail later.

[0155] It is noted that the signal model description does not (typically / necessarily) refer to a time-domain signal, but rather can alternatively or additionally refer to individual (potentially relatively small) frequency subbands. For example, the described signal model may individually apply to each of the frequency subbands for which separate spatial upmix parameters are provided.

[0156] The audio render apparatus 103 may be based on a consideration of the signal model as indicated above. In particular, the audio render apparatus 103 is arranged to receive the data signal including the mono downmix audio signal and spatial parameters and to generate local replicas / estimates of the directional signal s’ and the residual signals d’i, d’rfrom the mono downmix audio signal and the spatial parameters. These signal components may then be separately rendered to generate a combined stereo signal that may provide an improved spatial user experience in many embodiments.

[0157] The audio source device 101 of FIG. 1 is arranged to generate a data signal with data representing an audio stereo signal. The audio source device 101 is specifically arranged to generate the data signal to include encoded audio data for a mono downmix audio signal which represents a downmix of the stereo signal. In addition, spatial parameters that indicate relative properties between the channels of the stereo signal are included. Such spatial data may be appropriate for upmixing the mono downmix audio signal to recreate the downmix represented by the mono downmix audio signal.

[0158] As illustrated in FIG. 2, the audio source device 101 comprises a receiver 201 which is arranged to receive audio components from which the mono downmix audio signal is generated. In many embodiments, the receiver 201 may directly receive a stereo signal which is to be represented by the datasignal. In other embodiments, the receiver 201 may additionally or alternatively receive a number of audio components such as audio objects, mono signals, single source signals, multichannel signals etc.

[0159] The audio components are fed to a downmixer 203 which proceeds to generate the mono downmix audio signal. In many embodiments and scenarios, the downmixer 203 may generate a mono downmix audio signal from a received stereo signal, e.g. simply by summing the channels of the stereo signal in accordance with a standard PS downmix approach. In other embodiments, the downmixer 203 may be arranged to generate a stereo signal from received audio components, such as e.g. by generating an intermediate stereo signal for each received audio component followed by a combination of the intermediate stereo signals. For example, a multi-channel signal may be downmixed to an intermediate stereo signal, an audio object may be panned to the stereo image of an intermediate stereo signal based on position data provided for the audio object, etc.

[0160] The audio source device 101 further comprises a spatial parameter circuit 205 which is arranged to determine the spatial parameters for the stereo signal represented by the mono downmix audio signal. Specifically, the received or locally generated stereo signal may be fed to the parameter circuit 205 which may determine the spatial parameters from an analysis / processing of the stereo signal.

[0161] The spatial parameter circuit 205 may provide sets of frequency subband spatial parameters for the stereo signal where the sets of frequency subband spatial parameters are indicative of relative signal properties of the channels of the stereo signal. The frequency subband spatial parameters are provided for individual subbands of the stereo signal.

[0162] The spatial parameters are indicative of / reflect relative properties of the channel signals of the stereo audio signal. In particular, the spatial parameters may be indicated to include parameters that are indicative of at least one of relative intensities / levels of the stereo channels, relative (frequency domain) phases of the stereo channels, a relative time difference between the channels, and / or a correlation between the channels. Specifically, the spatial parameters may include one or more of an inter-channel intensity difference, inter-channel level difference, inter-channel time difference, interchannel phase difference, and / or inter-channel correlation.

[0163] The spatial parameters may specifically be spatial parameters as used for encoding a stereo signal using a Parametric Stereo (PS) encoding of the stereo signal, such as for example by a mono downmix audio signal given as:

[0164] m = c(l + r)

[0165] where the parameter c is chosen such that the power of the stereo signal is preserved in the downmix, the power being defined using the 2-norm:

[0166] ||m||2= ||l||2+ ||r||2and thus e.g.:

[0167] c = √( ||l||2+ ||r||2 / ||l+r||2)

[0168] c =

[0169]

[0170] ||l + r||2

[0171] The PS parameters are specifically an Inter-channel Intensity Difference IID, an Interchannel Correlation ICC, and in some cases an Inter-channel Phase Difference IPD parameter. These may specifically be defined / determined as:

[0172] IID

[0173] ICC

[0174]

[0175] IPD = arg < l,r

[0176] where the complex-valued inner product is defined as:

[0177] < x,y >= ∑∀ixi· yi*

[0178]

[0179] Vi

[0180] and:

[0181] | |x| |2=< x,x >

[0182] The spatial parameters are typically provided for specific time frequency tiles, and thus specifically each parameter value is generated / provided for a given frequency subband and for a given time segment.

[0183] The spatial parameter circuit 205 may determine spatial parameters that are indicative of relative properties of the channels (channel signals) of the stereo signal.

[0184] In many embodiments, the spatial parameter circuit 205 may receive the input stereo signal and process / analyze this to generate the spatial parameters. Specifically, the spatial parameter circuit 205 may calculate the IID, ICC, and IPD values in accordance with the formulas indicated above.

[0185] The spatial parameters may comprise sets of spatial parameters, each set of spatial parameters comprising at least one of: a level difference parameter indicative of a level difference between channels of the multichannel audio signal; a correlation parameter indicative of a coherencebetween channels of the multichannel audio signal; a timing difference parameter indicative of a timing difference between channels of the multichannel audio signal, and a phase difference parameter indicative of a phase difference between channels of the multichannel audio signal.

[0186] The audio source device 101 further comprises a data signal circuit 207 which generates the audio data signal and specifically it generates the audio data signal to include data representing the mono downmix audio signal and data representing the spatial parameters. The data signal circuit 207 may specifically include an audio signal encoder arranged to encode the mono downmix audio signal. It will be appreciated that any suitable audio signal encoding algorithm and approach may be used, such as an Advanced Audio Coding (AAC) mono encoding algorithm. Likewise, the spatial parameters may be encoded using a suitable encoding approach.

[0187] The data signal circuit 207 may generate an output data signal in accordance with any suitable format, and may specifically generate the output data signal to follow a suitable standard for an audio signal.

[0188] FIG. 3 shows examples of elements of the audio render apparatus 103.

[0189] The audio render apparatus 103 comprises a receiver 301 which is arranged to receive the data signal from the audio source device 101. Thus, the receiver 301 receives a data signal comprising encoded data for a mono downmix audio signal of a stereo signal. In addition, the data signal includes spatial upmix parameters for upmixing the mono downmix audio signal to the stereo signal where the spatial upmix parameters are indicative of relative signal properties of channels of the stereo signal. The spatial upmix parameters may as mentioned specifically be inter-channel time, phase, level, intensity differences and / or inter-channel correlation measures.

[0190] The receiver 301 is coupled to a decoder 303 which is arranged to receive the encoded data representing the mono downmix audio signal and to decode this data to generate the mono downmix audio signal. It will be appreciated that any suitable method for encoding and decoding the mono downmix audio signal may be used and in particular that any suitable standardized encoding format and algorithm may be used.

[0191] The receiver 301 may be arranged to receive a time domain audio signal and / or a frequency domain audio signal version / representation of the mono downmix audio signal. In some cases, the received data signal may include the mono downmix audio signal in only one representation, i.e. the data signal may include only one of the frequency domain audio signal and the time domain audio signal. In such cases, the received data signal may be transformed to the other domain as appropriate. Thus, in some cases, a received data signal may include a time domain audio signal being the time domain representation of the mono downmix audio signal, and a time to frequency domain transformer may from this generate the frequency domain audio signal for the mono downmix audio signal. In some cases, a received data signal may include a frequency domain audio signal being the frequency domain representation of the mono downmix audio signal and a frequency to time domain transformer may from this generate the time domain audio signal for the mono downmix audio signal if necessary.In particular, in some embodiments, the receiver may comprise a filter bank which is arranged to generate a frequency subband representation of a received time domain mono downmix audio signal. The receiver 301 may comprise a filter bank that is applied to the mono downmix audio signal such that it is divided into frequency subbands.

[0192] The filter bank may be Quadrature Mirror Filter (QMF) bank or may e.g. be implemented by a Fast Fourier Transform (FFT), but it will be appreciated that many other filter banks and approaches for dividing an audio signal into a plurality of subband signals are known and may be used. The filterbank may specifically be a complex-valued pseudo QMF bank, resulting in e.g. 32 or 64 complex-valued sub-band signals.

[0193] The processing is furthermore typically performed in time segments or time slots. In most embodiments, the audio signal is divided into time intervals / segments with a conversion to the frequency / subband domain by applying e.g. an FFT or QMF filtering to the samples of each signal. For example, each channel of the downmix audio signal may be divided into time segments of e.g. 2048, 1024, or 512 samples. These signals may then be processed to generate samples for e.g. 64, 32 or 16 subbands. Thus, a set of samples may be determined for each subband of the mono downmix audio signal.

[0194] It should be noted that the number of time domain samples is not directly coupled to the number of subbands. Typically, for a so-called critically sampled filterbank of N bands, every N input samples will lead to N sub-band samples (one for every sub-band). An oversampled fdterbank will produce more output samples. E.g. for every N input samples, it would generate k*N output samples, i.e., k consecutive samples for every band.

[0195] In some embodiments, the subbands are generated to have the same bandwidth but in other embodiments subbands are generated to have different bandwidths, e.g. reflecting the sensitivity of human hearing to different frequencies.

[0196] For example, the receiver 301 may employ a hybrid fdterbank with logarithmic fdter band center-frequency spacings that follow that of human perception similar to equivalent rectangular bandwidths (ERBs). In order to compensate for the delay of the fdtering by the small fdter bank, a delay may be introduced for higher frequency subbands.

[0197] As a specific example, a time -domain signal x [n] may be fed through a downsampled complex-exponential modulated QMF bank with K bands. Each frame of 64 time domain samples x[n] results in one slot of QMF samples X[k, I] with k = (0,..., K — 1) at slot I. The lower slots may then be filtered by additional complex-modulated fdterbanks splitting the lower bands further. The higher slots are delayed ensuring that the filtered mono downmix audio signals of the lower bands are in sync with the higher bands as the fdtering introduces a delay. This finally results in a structure where for every 64 timedomain samples x[n], one slot m of hybrid QMF samples K[k, Z] is produced with k = (0,..., L — 1) at slot Z, e.g. with a total number of hybrid bands M = 77.Thus, in many embodiments, the signals and the processing may be performed in subbands and for individual segments. Such blocks of a frequency interval / subband in a given time interval / segment will also be referred to as time frequency segments / tiles.

[0198] The mono downmix audio signal is fed to a directional signal circuit 305 which is arranged to generate a directional signal from the mono downmix audio signal. Thus, the directional signal circuit 305 may generate a modified mono downmix audio signal that seeks to represent / estimate a point source audio component of the stereo signal. The directional signal may be generated to seek to represent parts of the stereo signal which can be considered to have an origin with point source properties. Thus, the directional signal is generated to include audio components that correspond to audio that have a specific position / direction of origin.

[0199] The directional signal may in some cases be generated to represent / estimate audio that has a specific position in the stereo image of the stereo signal, and thus which corresponds to audio reaching the capture position (for the stereo signal) from a single direction.

[0200] The directional signal may accordingly be generated to estimate audio that is linked with audio sources that are spatially limited / small / localized and which specifically are generated by point sources. It may thus seek to extract such audio components / parts from the stereo signal to separate point source audio from audio with a spatial extension, such as e.g. background audio, ambient audio, etc.

[0201] In some cases, the directional signal generator 305 may be arranged to estimate an audio component from a single point source and specifically the directional signal may be generated to represent / estimate a single point source. In such an approach, audio from one specific direction may be identified / estimated and represented by the directional signal.

[0202] In most embodiments, however, the directional signal is generated to represent / estimate audio that has a point source origin but not necessarily from the same single point source. The directional signal may accordingly in many embodiments represent / estimate audio from a plurality of point sources.

[0203] The directional signal may represent a point source audio component of the stereo signal where the point source audio component is a (signal) component / part of the stereo signal representing audio from a point source. A point source may be a source that is spatially limited / small / localized and specifically in the stereo image of the stereo signal. A point source may be a source of audio having a single position in the stereo image of the stereo signal.

[0204] The point source audio component may be a signal component of the stereo signal (representing audio) originating from a point source (in the stereo image of the stereo signal). The point source audio component may be a signal component / part of the stereo signal originating from a single point / position / direction (in the scene / stereo image of the stereo signal).

[0205] The modified mono downmix audio signal / directional signal is fed to a first Tenderer 307 which is arranged to render the modified mono downmix audio signal to generate a first intermediate stereo signal. The rendering by the first Tenderer 307 (also referred to as a first rendering) is a directional rendering which renders the first intermediate stereo signal with a given direction / position in the stereoimage of the first intermediate stereo signal. The first rendering may specifically render the mono downmix audio signal as a point source with a given direction / position in the stereo image.

[0206] The first Tenderer 307 is coupled to a direction determining circuit 309 which is arranged to determine a direction y’ which is fed to the first Tenderer 307 resulting in this rendering the modified mono downmix audio signal from this position / direction. Thus, the first rendering is specifically such that the modified mono downmix audio signal in the first intermediate stereo signal is perceived as a point audio source positioned in the direction corresponding to the direction y’ determined by the direction determining circuit 309. The direction y’ will also be referred to as the rendering direction or rendering angle.

[0207] The direction determining circuit 309 is arranged to determine the direction from the received spatial upmix parameters. The spatial upmix parameters provide information on the relationship between the channels of the stereo signal that is downmixed and as such provide information of the position / orientation of the audio, and specifically of a dominant signal component in the stereo image of the stereo signal. For example, for a PS encoded signal, the spatial upmix parameters provide information of the position of the dominant signal component in the stereo signal, and specifically it provides information of an orientation angle for the dominant signal.

[0208] The direction determining circuit 309 may specifically determine the rendering direction y’ from the spatial upmix parameters. The rendering direction will typically be determined on a frequency tile basis, and specifically in frequency subbands and time segments matching those for which the spatial upmix parameters are provided.

[0209] The first Tenderer 307 may accordingly proceed to render the modified mono downmix audio signal such that is perceived from the given direction and it specifically achieves this directional rendering by applying a channel transfer function to the mono downmix audio signal with the channel transfer function generating the intermediate stereo signal from the mono downmix audio signal. The channel transfer function may specifically include a sub-transfer function for each channel, i.e. it may include one (sub)transfer function for generating a left channel signal and one (sub)transfer function for generating the right channel signal.

[0210] In many cases, the transfer function may be provided as a set of complex weights for the different subbands of a frequency representation of the modified mono downmix audio signal. The audio apparatus may perform many or all of the operations in the frequency domain and thus the transfer function may also be expressed and applied in the frequency domain. For example, for each frequency subband of the representation of the mono downmix audio signal, the transfer function may provide a complex weight for each of the output channels and a frequency representation of the first intermediate stereo signal may be generated by applying / multiplying the subband samples of the mono downmix audio signal by these weights to generate the subband samples of the first intermediate stereo signal.

[0211] The first transfer function is determined to correspond to the desired direction, i.e. it reflects the mapping from the mono downmix audio signal to the channels of the first intermediate stereosignal such that it is perceived as / corresponds to an audio source at a position in the stereo image corresponding the rendering direction / angle.

[0212] For example, in some cases, the transfer function for a given direction may correspond to a panning of the mono downmix audio signal to the given direction in the stereo image.

[0213] In many embodiments, the first rendering may be a binaural rendering and the first intermediate stereo signal may be a binaural stereo signal providing an enhanced spatial experience / perception when heard through headphones. Thus, the first Tenderer 307 may specifically be a binaural audio Tenderer which generates binaural audio signals for the left and right ear of a user. Binaural audio signals are generated to provide a desired spatial experience and are typically reproduced by headphones or earphones that specifically may be part of a headset worn by a user (the headset typically also comprises left and right eye displays).

[0214] Thus, in many embodiments, the audio rendering by the first Tenderer 307 is a binaural render process using suitable binaural transfer functions to provide the desired spatial effect for a user wearing a headphone. For example, the first Tenderer 307 may be arranged to generate an audio component to be perceived to arrive from a specific position using binaural processing.

[0215] Binaural processing is known to be used to provide a spatial experience by virtual positioning of sound sources using individual signals for the listener’s ears. With an appropriate binaural rendering processing, the signals required at the eardrums in order for the listener to perceive sound from any desired direction can be calculated, and the signals can be rendered such that they provide the desired effect. These signals are then recreated at the eardrum using either headphones or a crosstalk cancelation method (suitable for rendering over closely spaced speakers). Binaural rendering can be considered to be an approach for generating signals for the ears of a listener resulting in tricking the human auditory system into perceiving that a sound is coming from the desired positions.

[0216] The binaural rendering is based on binaural transfer functions which vary from person to person due to the acoustic properties of the head, ears and reflective surfaces, such as the shoulders. Binaural transfer functions may therefore be personalized for an optimal binaural experience. For example, binaural filters can be used to create a binaural recording simulating multiple sources at various locations. This can be realized by convolving each sound source with the pair of e.g., Head Related Impulse Responses (HRIRs) that correspond to the position of the sound source.

[0217] A well-known method to determine binaural transfer functions is binaural recording. It is a method of recording sound that uses a dedicated microphone arrangement and is intended for replay using headphones. The recording is made by either placing microphones in the ear canal of a subject or using a dummy head with built-in microphones, a bust that includes pinnae (outer ears). The use of such dummy head including pinnae provides a very similar spatial impression as if the person listening to the recordings was physically present during the recording.

[0218] By measuring e.g., the responses from a sound source at a specific location in 2D or 3D space to microphones placed in or near the human ears, the appropriate binaural filters can be determined.Based on such measurements, binaural filters reflecting the acoustic transfer functions to the user’s ears can be generated. The binaural filters can be used to create a binaural recording simulating multiple sources at various locations. This can be realized e.g., by convolving each sound source with the pair of measured impulse responses for a desired position of the sound source. In order to create the illusion that a sound source is moving around the listener, a large number of binaural filters is typically required with a certain spatial resolution, e.g., 10 degrees.

[0219] The head related binaural transfer functions may be represented e.g., as Head Related Impulse Responses (HRIR), or equivalently as Head Related Transfer Functions (HRTFs) or, Binaural Room Impulse Responses (BRIRs). The (e.g., estimated or assumed) transfer function from a given position to the listener’s ears (or eardrums) may for example be represented in the frequency domain in which case it is typically referred to as an HRTF or BRTF, or in the time domain in which case it is typically referred to as a HRIR or BRIR. In some scenarios, the head related binaural transfer functions are determined to include aspects or properties of the acoustic environment and specifically of the environment in which the measurements are made, whereas in other examples only the user characteristics are considered. Examples of the first type of functions are the BRIRs and BRTFs.

[0220] The audio render apparatus 103 comprises a store 311 which stores directional transfer functions for different directions. The directional transfer function for a given direction represents the mapping of a mono audio signal to stereo channels such that the mono audio signal is positioned in the given direction in a stereo image of the stereo channels. Thus, applying the directional transfer function for a given direction to the mono downmix audio signal may generate a stereo signal representing the mono downmix audio signal as an audio source positioned in the given direction. The mapping may in some cases be a time domain mapping (such as a gain, filter or other transfer function) or may in many cases be a frequency domain mapping, such as a set of parameter values / scale values (typically complex values) for different subbands. In the latter case, a frequency domain intermediate stereo signal may be generated by for each subband multiplying the subband sample of the mono downmix audio signal with respectively a complex value for that subband for a first channel of the intermediate stereo signal and with a complex value forthat subband for a second channel of the intermediate stereo signal.

[0221] For example, in examples where a panning is performed in the horizontal 2D plane, the store 311 may comprise panning parameters for different directions. For example, panning parameters for azimuth angles in a 0-360° interval may be provided for each 1° angle increment. The first Tenderer 307 may be coupled to the store 311 and be arranged to extract the directional transfer function for the rendering direction and then proceed to perform the rendering using the extracted directional transfer function. The rendering of the modified mono downmix audio signal may accordingly be rendered such that it is positioned / perceived in the stereo image to arrive from the rendering position.

[0222] It will be appreciated that the store 311 may not have directional transfer function stored for the desired rendering direction. In such cases, the first Tenderer 307 may be arranged to retrieve the nearest directional transfer function from the store 311 and use this for rendering. In such cases, therendering direction may be considered to correspond to the direction for the retrieved directional transfer function, i.e. the rendered direction may be a quantized value y’ of the desired rendering direction determined by the direction determining circuit 309.

[0223] In other embodiments, the first Tenderer 307 may be arranged to estimate a desired directional transfer function for a desired rendering direction by interpolating between two directional transfer functions from the store 311 corresponding to the two rendering angles nearest to the desired rendering direction determined by the direction determining circuit 309.

[0224] In most embodiments, the first Tenderer 307 is as mentioned arranged to perform a binaural rendering and the directional transfer functions stored in the store 311 are binaural transfer functions. Thus, the store may store data describing binaural transfer functions for different directions. The binaural transfer functions may for example be HRTFs, BRIRs, or HRIRs. The store 311 may specifically store frequency subband complex values for each channel for each frequency subband for a range of different frequencies. The first Tenderer 307 may thus perform the binaural rendering by multiplying the subband samples of the mono downmix audio signal with the corresponding subband coefficients / complex values of the selected binaural transfer function to generate subband sample values of the intermediate binaural stereo signal.

[0225] It will be appreciated that in many embodiments, the directional transfer functions may be stored as a plurality of functions linked with different directions. For example, the store 311 may be a look-up table which can receive the rendering direction as an index and provide a set of values of the directional transfer function forthat direction. The directional transfer function may for example be represented by individual subband values / coefficients, or may e.g. in other embodiments be represented by e.g. parameter values defining the directional transfer function operation (e.g. coefficients for the transfer function), a mathematical description / function from which suitable values of the transfer function can be generated etc.

[0226] Thus, the audio render apparatus 103 comprises a processing path which generates an intermediate stereo signal comprising the modified mono downmix audio signal represented as an audio source at a specific position in the spatial image of the first intermediate stereo signal. The mono downmix audio signal may typically be represented as a point audio source at the given direction. The rendering is adaptive with the direction being given by the spatial parameters and thus is dynamically adapted to reflect the characteristics of the stereo signal.

[0227] Further, the spatially definite / well defined and directional rendering is specifically of a directional signal representing point source audio. Thus, an improved spatial rendering of point source audio can be achieved with point source audio typically being rendered with a higher audio quality.

[0228] The generated first intermediate stereo signal is fed to an output circuit 313 which is arranged to generate an output stereo signal that includes the first intermediate stereo signal. As will be described more in the following, the output stereo signal may be generated to include other signalcomponents, including audio components representing non-point source audio, such as background or ambient sounds.

[0229] In many embodiments, the rendering of the mono downmix audio signal includes extracting a modified mono downmix audio signal / directional signal s’ which may specifically represent a point source audio component of the stereo signal.

[0230] It will be appreciated that different approaches, algorithms, and functions may be used to estimate the directional signal from the mono downmix audio signal using the spatial parameters.

[0231] For example, a rudimentary estimate of the directional signal from the mono signal may follow from: x' = ICC • m. This example may follow the consideration that for fully decorrelated signals all of the signal power of the original stereo signal is captured by the directional signal, whereas for fully decorrelated signals none of the signal power of the original stereo signal is captured by the directional signal.

[0232] In a particular approach, the directional signal is estimated as a spectro-temporally shaped version of the received mono downmix audio signal, where the shaping is done in such a way that it approximates the signal power of the directional signal of the signal model as previously described, i.e. where the original stereo signal is considered to represent a main component s and residual / more diffuse components di, d2. The directional signal can specifically be generated by a spectral shaping wherein frequency band samples of the mono downmix audio signal are multiplied by subband weights determined from the spatial parameters.

[0233] A particular approach can be determined from the signal model of:

[0234] I = cos^e^x + ni

[0235] r = sin (7) e-”6’' a? + nr

[0236] where x can be considered to correspond to a directional / point source signal.

[0237] The audio render apparatus 103 / first Tenderer 307 may then seek to generate an estimate of the directional signal component x by scaling of the mono downmix audio signal (with the scaling typically being separate for different frequency band). The estimated direct signal component may be represented by:

[0238]

[0239] where the resulting approximation x' is the estimate of the directional signal in a given subband and gxis determined from the spatial parameters, and specifically withgx= f(IID, ICC, IPD)

[0240] The spatial parameters are indicative of the relative properties of the channel signals for the stereo signal and, as has been realized by the inventors, this may also provide information on the relative levels for respectively a directional component (and for a non-directional component). The first renderer 307 may specifically determine a scaling factor that scales the effective gain of subbands of the mono downmix audio signal to generate the directional signal estimate.

[0241] In many embodiments, the first renderer 307 is accordingly arranged to generate the gains dependent on the spatial parameters. In particular, the gain for the directional rendering may be set in dependence on a relative power level of a directional component of the stereo signal relative to a power level of the mono downmix audio signal. The spatial upmix parameters provide information on the relative properties of the channel signals of the stereo signal, and specifically may provide information on both the interchannel levels / intensity differences as well as on the interchannel correlation. Accordingly, the spatial upmix parameters can be considered to provide information on the directional signal component and the diffiise / residual components. Accordingly, the gains gxmay be estimated / calculated from the received spatial parameters.

[0242] The gains may in many embodiments be determined to ensure power preservation. For the indicated signal model of the stereo signal, the gains for an actual point source with no other audio being present may be 1 and for a completely diffuse signal, the gains would be 0.

[0243] The gains for estimating the directional signal may in particular in many embodiments advantageously be determined in line with one or more of the following:

[0244] (IID - I)2+ 4ICC2IID

[0245] IID + 1

[0246] (IID - I)2+ 4 • ICC2• IID + (1 - IID) • ■ ICC2• IID + (IID - I)2

[0247] 1 - IID2+ (IID + 1) • ^4 • ICC2■ IID + (IID - I)2

[0248]

[0249] where IID is an interchannel intensity difference and ICC is an inter-channel cross-correlation, and specificallyIID

[0250] ||r||

[0251] ICC

[0252]

[0253] and

[0254]

[0255] and:

[0256] I |x| |2=< x,x >

[0257] Thus, in many embodiments, one or more of the gains may be determined based on received interchannel intensity differences, and interchannel correlations (being part of the spatial upmix parameters).

[0258] Different approaches for determining the rendering direction from the spatial upmix parameters may be used in different embodiments. In particular, the signal model as indicated above is based on directional component x being at a direction y in the stereo image of the stereo signal, henceforth also referred to as the orientation direction. In many embodiments, the direction determining circuit 309 may determine the orientation direction y and then determine the (desired) rendering direction y’ from the orientation direction y. Indeed, in some embodiments or scenarios, the rendering direction y’ may simply be set equal to the orientation direction y.

[0259] The determination of the orientation direction y may be based on the signal model indicated above. The spatial upmix parameters provide information on the relative properties of the channel signals of the stereo signal and specifically they may provide information on both the interchannel levels / intensity differences as well as on the interchannel correlation. Accordingly, the spatial upmix parameters can be considered to provide information on the directional signal component x and on the position of this in the stereo image of the stereo signal, i.e. the spatial upmix parameters provide information on the orientation direction y allowing this to be determined from the provided parameter values.

[0260] The direction determining circuit 309 may determine the orientation direction y as a direction to a directional signal component in a stereo image of the stereo signal from the spatial upmix parameters, and to map this to a direction in a stereo image of the output stereo signal. The directional signal component may be a dominant signal component. The direction determining circuit 309 may bearranged to determine the orientation direction y as a direction of a dominant sound source in the stereo signal where the direction of the dominant sound source is represented by the spatial upmix parameters.

[0261] The directional signal component may specifically be a signal component (estimated / determined) to originate from a point source. Specifically, the direction determining circuit 309 may be arranged to determine the orientation direction y as a direction for which a single point source audio source will result in spatial upmix parameter values matching the spatial upmix parameters of the data signal.

[0262] In some embodiments, the direction determining circuit may be arranged to determine the first direction in line with:

[0263] / 1 - IID + (IID - I)2+ 4 - ICC2IID

[0264] 7 — arctan I - - - - I 2 • icc • VnD

[0265]

[0266]

[0267] where IID is an interchannel intensity difference and ICC is an inter-channel cross-correlation, and specifically with these given by the equations provided above in connection with the equations for determining gains. The determination of the orientation direction y above results from an assumption that the signals x. ni and nrof the signal model are mutually decorrelated, and that the power of the left and right diffuse signals and nrare equal.

[0268] The direction determining circuit 309 may, as previously mentioned, in some embodiments be used directly as the rendering direction y’, i.e. y = y’. However, in many embodiments, a mapping may be included which for at least some values of the orientation direction y may result in a different rendering direction y’.

[0269] Thus, in many embodiments, the direction determining circuit 309 may be arranged to apply a mapping function to the orientation direction y to determine the rendering direction y’.

[0270] For example, the mapping may map the position in the stereo image of the original stereo signal as represented by the orientation direction y to a desired position in the stereo image of the output stereo signal as represented by the rendering direction y’. In many cases, where the output stereo signal is a binaural signal, the mapping may include a consideration / determination of a distance to the audio sources. For example, a range of the orientation direction y in the interval of [0,180°] may be mapped to a location between two virtual stereo speakers in the audio scene created by the binaural rendering. Such speakers may for example be positioned at angles of -30° and +30° relative to a center direction for the binaural signal. Thus, in such situations, the direction determining circuit 309 may include a mapping between an orientation direction y in the range of [0, 180°] to a rendering direction y’ in the range of [-30°, +30°].Thus, in some embodiments, the directional component (the mono downmix audio signal) may be rendered to a virtual angle in the range of a virtual loudspeaker angle range generated by a binaural rendering. The rendered directional component may be combined with a diffuse rendering of the residual signal(s).

[0271] In many embodiments, the direction determining circuit 309 may be arranged to map an orientation direction y representing an angle in one interval / range to a rendering direction y’ representing an angle in a different interval / range.

[0272] The audio render apparatus 103 comprises a second processing path which generates a second intermediate stereo signal. The mono downmix audio signal is fed to a decorrelator 315 which is arranged to apply a decorrelation to the mono downmix audio signal to generate a decorrelated mono downmix audio signal. It will be appreciated that a large number of different algorithms and functions for decorrelating an audio signal is known to the skilled person, and that any suitable approach or algorithm may be used without detracting from the invention.

[0273] The decorrelated mono downmix audio signal is fed to a signal generator 317 which is typically also fed the mono downmix audio signal and which is arranged to generate a first and second estimated signals based on the decorrelated mono downmix audio signal and the mono downmix audio signal. The first and second estimated signals are generated / estimated to, together with the directional signal, correspond to a decomposition of the stereo signal. The first and second estimated signals may be estimated signals representing signal components of the stereo signal that are not represented by the directional signal. The first and second estimated signal may be estimated to represent audio that is more diffuse such as ambient or background audio for the scene. The directional signal and the first and second estimated signals may be generated / estimated to represent respectively point source audio and non-point source audio / diffuse / ambient audio.

[0274] In some cases, each of the two signals may be generated by a filtering of the decorrelated mono downmix audio signal. The two filters may be different and generate different estimated signals from the decorrelated mono downmix audio signal.

[0275] The signal generator 317 is coupled to a second Tenderer 319 which is arranged to perform a second rendering being a rendering of the first estimated signal and the second estimated signal to generate a second intermediate stereo signal. However, in contrast to the first rendering process, the second rendering process is typically a predetermined rendering which is not dependent on the spatial parameters, and which typically is not depending on properties of the stereo signal. The second rendering may typically be a diffuse rendering seeking to generate the second intermediate stereo signal to provide a perception of a more diffuse and spatially less definite audio source. The second rendering is specifically a predetermined rendering employing a predetermined mapping of the first and second estimated signals to channel signals of the second intermediate stereo signal.

[0276] As a specific example, the second rendering may typically generate the second intermediate stereo signal by simply mapping the first estimated signals and the second estimated signalto two phase-inverse signals, i.e. the second intermediate stereo signal may be generated with the first and second estimated signals being mapped to the two stereo channels. For example, in some embodiments, the first estimated signal may be mapped to the right signal of the second intermediate stereo signal and the second estimated signal may be mapped to the left signals of the second intermediate stereo signal.

[0277] The first Tenderer 307 and the second Tenderer 319 are coupled to a combiner 313 which is arranged to combine at least the first intermediate stereo signal and the second intermediate stereo signal to generate an output stereo signal. In many embodiments, the combiner 313 may be arranged to combine / sum the samples / values of the individual channels of the first and second intermediate stereo signals to generate the samples / values of the output stereo signal. In many cases, the combination may be performed by combining / summing subband values of the intermediate stereo signals. In other embodiments, the combination may be performed in the time domain by combining / summing time domain values of the intermediate stereo signals.

[0278] In many embodiments, the combination of the intermediate stereo signals may be by a (possibly weighted) combination / summation of corresponding channel signals for the first intermediate stereo signal and the second intermediate stereo signal.

[0279] In some embodiments where binaural processing is used, each of the estimated (residual) signals may be rendered from a specific position, such as each estimated (residual) signal being rendered from a different virtual position, such as for example from different virtual positions.

[0280] In some embodiments, the rendering for the first and / or second estimated signal, may position the signal at a specific position.

[0281] In many embodiments, the rendering of the estimated signals may be performed by the Tenderer 319 retrieving a set of directional transfer functions from the store 311 and rendering the estimated signals using the retrieved transfer functions.

[0282] In many embodiments, the audio render apparatus 103 may be arranged to extract a directional transfer function for a single predetermined direction and render the first and / or second estimated signal using this directional transfer function. Accordingly, (each of) the estimated signals may be rendered from one predetermined direction / position, such as a direction / position corresponding to a virtual speaker position.

[0283] An example of subband parametric rendering may e.g. result in left and right signals:

[0284] 1 = 9s - md- Gi [f(r)] ■ eiMW]+gn.Hi[my.Gl[p^.ei<t>ilPi\+gn. H2{m} ■ Gt[pr] ■ r = gs- md- Gr[f(yy] ‘ ei<l>rlf(r)]+gn■ H^m} ■ Gr[^] ■ + gn■ H2{m] ■ Gr[(3r]

[0285]

[0286] .ej<t> APr}

[0287] where Gj, Gr. cf>t, (prform the parametric HRIRs, f (y) is a mapping function converting the estimated angles (orientation direction y) to HRIR direction angles, Pi and are two pre-determined angles and} and H2{. } represent the processing generating the first and second estimated signals respectively, m is the received mono downmix audio signal and n is the directional signal generated by the directional signal generator 305, and gsand gnare suitable scaling factors.

[0288] In some embodiments, the Tenderer 319 may retrieve directional transfer functions for a plurality of predetermined directions and it may use multiple directional transfer functions in performing the predetermined rendering. For example, different directional transfer functions may be used for different frequency subbands. This may provide a more diffuse perception with the audio being generated such that it is perceived from different directions for different subbands thereby resulting in a perception of a more distributed and spread audio source.

[0289] In the latter case, the sets of predetermined directions for the different estimated signals are different in order to enhance the perceived diffuseness of the non-directional audio components.

[0290] The Tenderer 319 may generate the second intermediate stereo signal using a first set of directional transfer functions retrieved from the store 311 for a first set of predetermined directions, and may generate the third intermediate stereo signal using a second set of directional transfer functions retrieved from the store 311 for a second set of predetermined directions where the first set of set of predetermined directions is different from the second set of predetermined directions.

[0291] In particular, the directional transfer functions may be binaural transfer functions and the Tenderer 319 may be arranged to perform binaural rendering to generate the first intermediate stereo signal using binaural impulse response values for a first set of predetermined directions and may be arranged to perform binaural rendering to generate the first intermediate stereo signal using binaural impulse response values for a second set of predetermined directions where the first set of predetermined directions are different from the second set of predetermined directions.

[0292] In many cases, the use of multiple directional transfer functions may be achieved by using directional transfer functions for different directions in different frequency subbands.

[0293] Thus, instead of rendering the diffuse / non-directional estimated signals using fixed angles, e.g. mimicking a virtual stereo speaker setup, the diffuse signals may also be rendered using composite, e.g. pre-calculated HRIRs for many sources / directions, e.g. spread over a (part of a) circle, or (part of) a sphere.

[0294] 1 = gs- ™d- Gt[f(y)] ■ + gn■ H^m] ■ G l, comp ' eJ<^l-comV + 9n ■ H2{m} ■ G I, comp ■ ej$l>comp r = gs- md- Gr[f(y)] ■ + gn■ ■ GT: Comp■eJ<pr-comP + gn■ H2{m} ■ GT: Comp

[0295]

[0296] . eJ4>r,comp

[0297] where e.g.:Gl’Comp 9norm GM] ■

[0298] PEBi { '

[0299] ‘ e7<^] ■

[0300] r _

[0301] ur,comp 9 norm ^ Gr\J3] - e7^^

[0302] / ?eBr{ '

[0303] Gr[j3] ■ e7<w / ?]■

[0304]

[0305] / ?eBr

[0306] with B(being a set of angles at which the left diffuse signal is to be rendered, Bra set of angles at which the right diffuse signal is to be rendered, and gnormanormalisation factor.

[0307] In some embodiments, the diffuse signal component may be directly rendered onto left and right channels without any HRIR processing.

[0308] The audio apparatus may accordingly be arranged to generate an output stereo signal, and often an output binaural stereo signal from the received mono downmix audio signal and spatial upmix parameters. The audio apparatus specifically implements two different rendering paths with one being a directional (binaural) rendering of a directional (e.g. a dominant) signal component which is generated from the mono downmix audio signal and the difference signal. The other rendering path may employ a predetermined rendering / mapping of estimated signals generated from the mono downmix audio signal and a decorrelation thereof. The rendering of the output stereo signal is typically not a conventional adaptive upmixing of the received and decorrelated mono signals, and is specifically not a conventional 2x2 matrix upmixing of the mono signal and a decorrelated signal, but rather is a direct generation of a stereo signal by parallel processing of respectively the mono downmix audio signal and estimated signals generated from the mono downmix audio signal and decorrelated versions of the mono downmix audio signal, with the former rendering being directional dependent on the spatial upmix parameters and the latter rendering being a predetermined rendering.

[0309] The processing seeks to render direct / dominant / directional point source components using a direct rendering with a direction that is given by the spatial parameters. The rendering employs a directionally dependent transfer function to the left and right stereo output signal for that purpose. The approach further seeks to render a residual / remaining signal component as more diffuse audio, andspecifically it may use a predetermined rendering where a decorrelated signal is mapped directly to the channels of the output binaural signal using a transfer function. The mapping may be predetermined and may specifically be such that it allows a more diffuse and non-directional perception of this signal component. The rendering process thus uses fundamentally different approaches to provide different signal components in the output binaural signal, but does so without specifically decomposing the mono downmix audio signal into a dominant and diffiise / residual signal component with these subsequently being individually rendered. Rather a direct rendering of respectively the mono downmix audio signal and estimated signals generated using a decorrelated version of the mono downmix audio signal is performed to generate the binaural output signal.

[0310] The approach allows low complexity and computationally efficient rendering of a stereo signal encoded as a downmix and spatial upmix parameters, such as a PS encoded signal. It may further allow high performance rendering with a perceived improved audio quality. In many cases, a substantially improved spatial perception and user experience may be achieved, and indeed, can be provided using headphones. The approach may in many cases provide a user perception of an audio scene where individual / dominant audio sources are well defined at specific directions / positions in a stereo image whereas other sources (e.g. ambient orbackground audio sources) are perceived more diffuse.

[0311] The approach may typically allow a very efficient operation and rendering with reduced complexity. A particular advantage of the approach is that it does not require a decomposition of the received mono downmix audio signal into different components with different specific properties.

[0312] The audio render apparatus 103 accordingly generates an output stereo signal which is the combination of a directional rendering putting point source audio at a specific and spatially definite position determined from the received spatial upmix parameters, and of a typically predetermined rendering providing a more diffuse and decorrelated perception of the corresponding audio source. The approach provides two parallel rendering processes / paths for the mono downmix audio signal with the rendered results being combined to generate the output stereo signal.

[0313] However, further, the audio render apparatus 103 of FIG. 3 is arranged to control the generation of the signals such that relative properties between the two estimated signals are controlled based on the spatial parameters. Thus, rather than merely generating the two signal components representing non-point source audio (typically ambient / diffuse sound) by two signal components that are decorrelated with respect to each other and with the same signal level, the audio render apparatus 103 of FIG. 3 includes an adapter 321 which is specifically arranged to dynamically control the relative property, and specifically the correlation between and / or relative level of the estimated signals.

[0314] The adapter 321 receives the spatial parameters and proceeds to control the signal generator 317 to adapt the generation of the estimated signals such that they have appropriate relative properties, and specifically such that they have appropriate (cross)correlation and signal levels as indicated by the spatial properties. The adapter 321 and signal generator 317 accordingly operate to generate estimated signals that have varying and signal dependent (via the spatial parameters) correlationand signal level properties. The estimated signals may typically be residual signal estimated to represent audio components of the stereo signal that are not represented / included in the directional signal.

[0315] As mentioned, the directional signal circuit 305 may generate the directional signal as

[0316] s’ = 9s - m

[0317] Further the signal generator 317 may generate the two estimated signals as linear combinations of the mono downmix audio signal and the decorrelated mono downmix audio signal, e.g.:

[0318] dt'l _ rgm,i 9a,i~l rmi

[0319] dr'_ L9m,r 9d,rJ Id J

[0320] Accordingly the overall decomposition can be represented by:

[0321] s' 9s 0

[0322] rmi

[0323] dt' 9m, l 9d,l Ld J

[0324]

[0325] -9m,r 9d,r-

[0326] Thus, the audio render apparatus 103 may generate the directional signal and the estimated signals to correspond to a signal composition of the stereo signal in accordance with the previously introduced signal model. The signals may be generated as linear combinations of the mono downmix audio signal and the decorrelated mono downmix audio signal.

[0327] The adapter 321 may proceed to adapt the (matrix) coefficients to seek to maintain the properties of the underlying signal model. Specifically, by imposing that the decorrelator signal power equals the mono downmix power, | |d| |2= 11 m 112this may allow reconstruction of a range of properties:

[0328] Ill'll2= INI2

[0329] Ill'll2= INII2

[0330] ||d / ||2= IKII2

[0331] < s', di >=< s, di >

[0332]

[0333] Accordingly, the adapter 321 may by determining suitable weights / scale factors / gains / (matrix) coefficients proceed to generate the estimated signals such that they have a relative signal leveland a cross correlation that match those of a decomposition of the original stereo signal according to the signal model. The estimated signals are thus not merely generated as decorrelated signals with the same signal level but are generated to be signals having relative / cross properties that correspond more closely to a decomposition according to a suitable signal model for the stereo signals. In addition, the relative properties, and specifically the relative levels and the correlations, between the difference signal and respectively the first and second estimated signals may be adapted to correspond to those of the underlying signal model.

[0334] The directional signal and the estimated signals are each a linear combination of the mono downmix audio signal and the decorrelated mono downmix audio signal and the weights of the linear combination may be determined as a function of the spatial parameters. The adapter 321 may specifically apply functions that result in one or more of the properties indicated above being the same for the generated signals as for a decomposition in accordance with the signal model.

[0335] The linear combinations and the determination of suitable weights based on the spatial parameters are typically performed in the frequency domain and may be individual and separate for each frequency subband / time frequency tile.

[0336] In some embodiments, as in particular in the example above, no other signal than the mono downmix audio signal is included in the generation of the directional signal. Similarly, no other signals than the mono downmix audio signal and the decorrelated mono downmix audio signal are included in the generation of the estimated signals.

[0337] It will be appreciated that different detailed approaches can be used to determine suitable weights and coefficients from the spatial parameters, and that different approaches may be used in different embodiments.

[0338] For example, using the overall decomposition matrix:

[0339] s' 9 s 0

[0340] rmi

[0341] dt' 9m, l 9d,l

[0342] Ld J

[0343]

[0344] -9m,r 9d,r-

[0345] If the parameter gmis chosen to be set to 0, the correlation between the generated signal pair (s', d'i) becomes 0, under the condition that the correlation between the mono signal m and the decorrelated signal d was 0. The same holds for the for the parameter gm rand the generated signal pair (s', d'r). Similarly, if the parameter gd iis chosen to be set to 0, the (absolute) correlation between the generated signal pair (s', d'i) becomes 1. Likewise, if the parameter gd ris chosen to be set to 0, the (absolute) correlation between the generated signal pair (s', d

[0346]

[0347] 'r) becomes 1. By balancing gmand gdi an arbitrary correlation between the signal pair (s', d'i) can be generated. By balancing gm rand gd ran arbitrary correlation between the signal pair (s', d'r) can be generated. The relative levels of the signalpair (s', d'i) can be realized by weighting gsto the combined weights of gmi and g^ i (typically as 9m,i2+ 9d,i2Y F°rexample, if the weights of gmi and g^ i are relatively close to 0, compared to the weight gsa large relative level is realized. If the weight of gsis relatively close to 0, compared to the weights of gmi and g^ i, a small relative level is realized. Similarly, the relative levels of the signal (s', d'r) can be realized by weighting gsto the combined weights of gm rand gci r(typically as

[0348] 9m,r " 9d,r )•

[0349] It can be shown that for two signals a, b that are a linear combination of the left and right signals according to:

[0350] pi _ pia2i rn

[0351] Lb J - b2’ LrJ

[0352] that:

[0353] |< a, b >|

[0354] if* / '"1|< ai - l + a2• r, bi • 1 + b2• r >| i / l |ai - l + a2- r||2• ||bi • / + &2• ’'ll2

[0355] | < ai • l + a2■ r, &i ■ I + fe2• r > | ICCO,6 vftlail2< 1,1 > +|a2|2< r,r > +2R{ax • a2- < l,r >}) • (|bi|2< 1,1 > +I&2I2< r,r > +2R{&i • b^- < l,r >})

[0356] |ai • bj ■ IIP + a2• b2* + oi • b*2■ ICC • e>1PD■ VHP + a2• bj ■ ICC • e^'IPD• Z5P|

[0357]

[0358] This means that if there are given relationships between the coefficients a, a2, b, b2and the PS parameters, i.e., a = f^IID, ICC, IPD), a2= f2(JID, ICC, IPD b =

[0359] / 3( / / £), ICC, IPD), b2= f4(HD, ICC, IPD), it is possible to calculate the cross-correlation between the resulting two signals a and b.

[0360] Similarly, it can be shown that the intensity ratio can be determined as:

[0361] i t i >||a||2|ai|2- IID + |a2|2+ 2 - ICC - VnD - |ai| - |a2| - cos(IPD + Zai - Za2)

[0362]

[0363] Ill’ll2H2• IID + |b2|2+ 2 • ICC ■ yilD ■ 1^1 • |62| • cos(IPD + Z^ - Zb2)This approach may be generalized as following. Suppose a set of n signals

[0364]

[0365] are generated as a linear combination of a left and right signal using a matrix G, where each element is a function of the PS parameters (gtj = ftj(IID, ICC, IPD)

[0366]

[0367] then, the resulting normalized covariance matrix between each pair of the linearly combined signals Xt, Xj is a function of the PS parameters and the underlying matrix functions gt j:

[0368] _ < XiL, X Ji _ >A / < xi,xi>< Xj, Xj > _ UlIDij ICCij y!<xj'xj > ICCj.i JHD~

[0369] ICCij

[0370]

[0371] ICCij*

[0372] In many embodiments, the adapter 321 may accordingly adapt the generation of the estimated signals to have desired relative properties, and specifically to have a desired correlation and signal level properties.

[0373] In many embodiments, the adapter 321 may further be arranged to adapt a relative property between the directional signal and the first and / or second estimated signal in dependence on the spatial parameters. The relative property may specifically be the level difference, and specifically the IID, between the directional signal and one of the estimated signals, and / or a correlation, and specifically an ICC, between the directional signal and one of the estimated signals.

[0374] For example, using the overall decomposition matrix:

[0375] s' 9 s 0

[0376] mi

[0377] 9m, i 9d,l

[0378]

[0379] -9m,r 9d,r-the IID between the directional signal and the left and right estimated signals are determined as:

[0380] HDs>dl= Igsl2

[0381] ^9m,l | + ^9d,l |

[0382] HDs>d

[0383]

[0384] If we want to reinstate these IID values, without affecting the correlation, we can set 9m, i=9m, r=0- 9s is equal to gxas defined before. Then the desired IIDs dand IIDs dr, which can be calculated from the linear equations of the decomposition model, can be realized by setting the values of 9d,l and gd raccordingly.

[0385] The ICC between the direct and the left estimated signal using the overall decomposition matrix is given by:

[0386] lrir, _ 9s ' 9m, l

[0387] ILLs,d

[0388] 1, ~ |

[0389] I / 2 2 \

[0390] + \9d,l | J

[0391]

[0392] where gsis again equal to gxdefined before, resulting in two equations, one for the IIDs dand one for ICCs,dtand two unknowns, gm iand gdd.

[0393] As a particular example, the processing may be based on a Principal Component Analysis (PCA) and the directional signal 5 may be considered as the principal component determined by a PCA process.

[0394] The general decomposition of a stereo signal into a directional signal s, and a left and right residual signal di and dr, can be determined as:

[0395] s

[0396] di = H ■ Pl

[0397] [djLrJ

[0398] where the matrix H is a 3x2 matrix, where each element is a function of the PS parameters:

[0399] hx y= f{HD, ICC, IPD}

[0400] An example of such matrix is given by a PCA decomposition, for which:

[0401] wzwr

[0402] l - |w;|2-W*L- Wr

[0403]

[0404] -Wr* - Wi 1 - |wr|2

[0405] where:

[0406]

[0407] and:

[0408] IPD = <pt— <pr

[0409] The adapter 321 may in many embodiments seek to maintain the decomposition interchannel properties:

[0410] HDSidl=

[0411] HDs>d

[0412] iccs,dl—

[0413] iccsd.

[0414] iccd d

[0415]

[0416] It is noted that IIDdl drfollows from the other IID values.

[0417] For the PCA decomposition it can be shown that:

[0418] < s,di >

[0419] iccs,dl

[0420] / < s,s >■< di, di >

[0421] < Wi ■ I + wr■ r, \l — Wi\2■ I — Wi* - wr- r >

[0422]

[0423] y / < wt• I + wr• r,wt• I + wr• r >■< |1 — wj2■ Z — wt* ■ wr■ r, |1 — wt\2■ I — wf ■ wr■ r >By fdling in the expressions for WLand wras provided above, it can be shown that:

[0424] iccs dl= 0

[0425] which is well-known property of the PCA decomposition. It realizes an orthogonalization of the resulting signals.

[0426] In a similar way, it can be shown that

[0427] ICCS'dr=o

[0428] Finally, using the same method as described above, it is possible to show that:

[0429] ICCdbdr= -eJ 'PD

[0430] Using the overall decomposition matrix, it becomes obvious that gm iand gm rcan be set to 0, as no correlation is required between the directional signal and one of the estimated signals:

[0431] s' 9 s 0

[0432] rmi

[0433] dt’ 9m, l 9d,l

[0434] Ld J

[0435]

[0436] -9m,r

[0437] As a result, the decomposition inter-channel properties can be reconstructed by applying the following matrix on the mono downmix audio signal and (single) decorrelated mono downmix audio signal:

[0438] 9s 0

[0439] G = 0 9d,i ■

[0440]

[0441] 0 -9d,r ■ e-r1PD / \

[0442] where:

[0443] IID + si + 2 ca- sa- ICC JUD

[0444]

[0445] II D + 1

[0446] 9 d,l a ' 9 d

[0447] 9 d,r -a ' 9 dcl + si ■ IID — 2 ■ ca■ sa■ ICC ■ VTTD

[0448] IID + 1

[0449]

[0450] It is noted that in the matrix G, the phase angle IPD can be spread differently.

[0451] Thus, in the specific case where the decomposition is based on a PCA approach, the directional signal can be generated by scaling the mono downmix audio signal by suitable frequency subband weights and the first and second estimated signals can be generated by scaling the decorrelated mono downmix audio signal by suitable frequency subband weights that are different for the two estimated signals.

[0452] In many embodiments, the signal generator 317 may be arranged to generate the first estimated signal by applying frequency band weights to frequency band samples of the decorrelated mono downmix audio signal where the frequency band weights are dependent on the spatial parameters.

[0453] Specifically, it may be generated as / from:

[0454] 9

[0455]

[0456] d,i ■

[0457] Similarly, the signal generator 317 may be arranged to generate the second estimated signal by applying frequency band weights to frequency band samples of the decorrelated mono downmix audio signal where the frequency band weights being dependent on the spatial parameters. Specifically, it may be generated as / from:

[0458] ~9d,r ■ e~i IPD / 2

[0459] The processing may be performed in subbands and may be performed in time segments. The processing in each subband may for some (any) or all steps be performed separately / independently ineach subband (with respect to the processing in other subbands). The processing in each time segment may for some (any) or all steps be performed separately / independently in each time segment (with respect to the processing in other time segments).

[0460] The processing may be time interval / segment based with all processing being performed for each time segment. Equivalently, the signal(s) for each segment may be considered a signal (and in particular signals of different time segments, may be considered different signals).

[0461] The audio apparatus(s) may specifically be implemented in one or more suitably programmed processors. An example of a suitable processor is provided in the following.

[0462] FIG. 4 is a block diagram illustrating an example processor 400 according to embodiments of the disclosure. Processor 400 may be used to implement one or more processors implementing an apparatus as previously described or elements thereof (including in particular one more artificial neural network). Processor 400 may be any suitable processor type including, but not limited to, a microprocessor, a microcontroller, a Digital Signal Processor (DSP), a Field ProGrammable Array (FPGA) where the FPGA has been programmed to form a processor, a Graphical Processing Unit (GPU), an Application Specific Integrated Circuit (ASIC) where the ASIC has been designed to form a processor, or a combination thereof.

[0463] The processor 400 may include one or more cores 402. The core 402 may include one or more Arithmetic Eogic Units (AEU) 404. In some embodiments, the core 402 may include a Floating Point Logic Unit (FPLU) 406 and / or a Digital Signal Processing Unit (DSPU) 408 in addition to or instead of the ALU 404.

[0464] The processor 400 may include one or more registers 412 communicatively coupled to the core 402. The registers 412 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments the registers 412 may be implemented using static memory. The register may provide data, instructions and addresses to the core 402.

[0465] In some embodiments, processor 400 may include one or more levels of cache memory 410 communicatively coupled to the core 402. The cache memory 410 may provide computer-readable instructions to the core 402 for execution. The cache memory 410 may provide data for processing by the core 402. In some embodiments, the computer-readable instructions may have been provided to the cache memory 410 by a local memory, for example, local memory attached to the external bus 416. The cache memory 410 may be implemented with any suitable cache memory type, for example, Metal-Oxide Semiconductor (MOS) memory such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), and / or any other suitable memory technology.

[0466] The processor 400 may include a controller 414, which may control input to the processor 400 from other processors and / or components included in a system and / or outputs from the processor 400 to other processors and / or components included in the system. Controller 414 may control the data paths in the ALU 404, FPLU 406 and / or DSPU 408. Controller 414 may be implemented as one or more statemachines, data paths and / or dedicated control logic. The gates of controller 414 may be implemented as standalone gates, FPGA, ASIC or any other suitable technology.

[0467] The registers 412 and the cache 410 may communicate with controller 414 and core 402 via internal connections 420A, 420B, 420C and 420D. Internal connections may be implemented as a bus, multiplexer, crossbar switch, and / or any other suitable connection technology.

[0468] Inputs and outputs for the processor 400 may be provided via a bus 416, which may include one or more conductive lines. The bus 416 may be communicatively coupled to one or more components of processor 400, for example the controller 414, cache 410, and / or register 412. The bus 416 may be coupled to one or more components of the system.

[0469] The bus 416 may be coupled to one or more external memories. The external memories may include Read Only Memory (ROM) 432. ROM 432 may be a masked ROM, Electronically Programmable Read Only Memory (EPROM) or any other suitable technology. The external memory may include Random Access Memory (RAM) 433. RAM 433 may be a static RAM, battery backed up static RAM, Dynamic RAM (DRAM) or any other suitable technology. The external memory may include Electrically Erasable Programmable Read Only Memory (EEPROM) 435. The external memory may include Flash memory 434. The External memory may include a magnetic storage device such as disc 436. In some embodiments, the external memories may be included in a system.

[0470] The invention can be implemented in any suitable form including hardware, software, firmware, or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of an embodiment of the invention may be physically, functionally and logically implemented in any suitable way. Indeed, the functionality may be implemented in a single unit, in a plurality of units or as part of other functional units. As such, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.

[0471] Although the present invention has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the accompanying claims. Additionally, although a feature may appear to be described in connection with particular embodiments, one skilled in the art would recognize that various features of the described embodiments may be combined in accordance with the invention. In the claims, the term comprising does not exclude the presence of other elements or steps.

[0472] Furthermore, although individually listed, a plurality of means, elements, circuits or method steps may be implemented by e.g. a single circuit, unit or processor. Additionally, although individual features may be included in different claims, these may possibly be advantageously combined, and the inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. Also, the inclusion of a feature in one category of claims does not imply a limitation to this category but rather indicates that the feature is equally applicable to other claim categories as appropriate.Furthermore, the order of features in the claims do not imply any specific order in which the features must be worked and in particular the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. In addition, singular references do not exclude a plurality. Thus references to "a", "an", "first", "second" etc. do not preclude a plurality. Reference signs in the claims are provided merely as a clarifying example shall not be construed as limiting the scope of the claims in any way.

[0473] Generally, examples of an audio apparatus and a method therefor are indicated by below embodiments.

[0474] EMBODIMENTS:

[0475] Embodiment 1. An audio apparatus comprising:

[0476] a receiver (301) arranged to receive a data signal comprising encoded data for a mono downmix audio signal and a set of spatial upmix parameters for upmixing the mono downmix audio signal to a stereo signal, the set of spatial upmix parameters being indicative of relative signal properties of channels of the stereo signal;

[0477] a store (311) comprising directional transfer functions for different directions, a directional transfer function for a given direction representing a mapping of a mono audio signal to stereo channels such that the mono audio signal is positioned in the given direction in a stereo image of the stereo channels;

[0478] a decoder (303) arranged to generate the mono downmix audio signal by decoding the encoded data;

[0479] a directional circuit (305) arranged to estimate a directional signal, the directional signal representing a point source audio component of the stereo signal;

[0480] a decorrelator (315) arranged to generate a decorrelated mono downmix audio signal from the mono downmix audio signal;

[0481] a generator (317) arranged to generate a first estimated signal and a second estimated signal from the mono downmix audio signal, at least the first estimated (residual) signal further being generated from the (first) decorrelated mono downmix audio signal and the directional signal, the first estimated signal, and the second estimated signal forming a decomposition of the stereo signal;

[0482] a direction determining circuit (309) arranged to determine a first direction from the spatial upmix parameters;

[0483] a first Tenderer (307) arranged to perform a first rendering of the mono downmix audio signal to generate a first intermediate stereo signal, the first rendering being a directional rendering using a first channel transfer function retrieved from the store for the first direction;a second Tenderer (319) arranged to perform a second rendering being a rendering of the first estimated signal and the second estimated signal to generate a second intermediate stereo signal;

[0484] a combiner (313) arranged to combine at least the first intermediate stereo signal and the second intermediate stereo signal to generate an output stereo signal; and

[0485] an adapter (321) arranged to control the generator (317) to adapt a first relative property being a relative property between the first estimated signal and the second estimated signal in dependence on the spatial upmix parameters.

[0486] Embodiment 2. The audio apparatus of embodiment 1 wherein the first relative property includes a level of the first estimated (residual) signal relative to the second estimated (residual) signal.

[0487] Embodiment 3. The audio apparatus of embodiment 1 or 2 wherein the first relative property includes a correlation between the first estimated (residual) signal relative to the second estimated (residual) signal.

[0488] Embodiment 4. The audio apparatus of any previous embodiment wherein the adapter (321) is arranged to adapt a second relative property being a relative property between the directional signal and the first estimated signal in dependence on the spatial upmix parameters.

[0489] Embodiment 5. The audio apparatus of embodiment 4 wherein the second relative property includes a level of the directional signal relative to the first estimated (residual) signal.

[0490] Embodiment 6. The audio apparatus of embodiment 4 or 5 wherein the second relative property includes a correlation between the directional signal relative and the first estimated (residual) signal.

[0491] Embodiment 7. The audio apparatus of any previous embodiment wherein at least the directional signal and the first estimated signal are each a linear combination of the mono downmix audio signal and the decorrelated mono downmix audio signal.

[0492] Embodiment 8. The audio apparatus of embodiment 7 wherein the adapter (321) is arranged to determine weights of the linear combination as a function of the spatial upmix parameters.

[0493] Embodiment 9. The audio apparatus of any previous embodiment wherein the directional signal, the first estimated signal, and the second estimated signal are generated to be estimates of a signal model for the stereo signal, the signal model representing the stereo signal as a matrix multiplication of a signal vector by a matrix having a rank of two, the signal vector comprising a directional signal and two diffuse signals.Embodiment 10. The audio apparatus of embodiment 9 wherein the adapter (321) is arranged to adapt the first relative property to match a corresponding relative property between the two diffuse signals of the signal model.

[0494] Embodiment 11. The audio apparatus of any previous embodiment wherein the generator (317) is arranged to generate the first estimated signal by applying second frequency band weights to frequency band samples of the decorrelated mono downmix audio signal, the second frequency band weights being dependent on the spatial upmix parameters.

[0495] Embodiment 12. The audio apparatus of any previous embodiment wherein the generator (317) is arranged to generate the second estimated signal from the decorrelated mono downmix audio signal.

[0496] Embodiment 13. The audio apparatus of embodiment 12 wherein the generator (317) is arranged to generate the second estimated signal by applying third frequency band weights to frequency band samples of the decorrelated mono downmix audio signal, the third frequency band weights being dependent on the spatial upmix parameters.

[0497] Embodiment 14. A method of operation for an audio apparatus, the method comprising:

[0498] receiving a data signal comprising encoded data for a mono downmix audio signal and a set of spatial upmix parameters for upmixing the mono downmix audio signal to a stereo signal, the set of spatial upmix parameters being indicative of relative signal properties of channels of the stereo signal;

[0499] providing directional transfer functions for different directions, a directional transfer function for a given direction representing a mapping of a mono audio signal to stereo channels such that the mono audio signal is positioned in the given direction in a stereo image of the stereo channels;

[0500] generating the mono downmix audio signal by decoding the encoded data; estimating a directional signal, the directional signal representing a point source audio component of the stereo signal;

[0501] generating a decorrelated mono downmix audio signal from the mono downmix audio signal;

[0502] generating a first estimated signal and a second estimated signal from the mono downmix audio signal, at least the first estimated (residual) signal further being generated from the (first) decorrelated (mono downmix audio) signal and the directional signal, the first estimated signal, and the second estimated signal forming a decomposition of the stereo signal;

[0503] determining a first direction from the spatial upmix parameters;

[0504] performing a first rendering of the mono downmix audio signal to generate a first intermediate stereo signal, the first rendering being a directional rendering using a first channel transfer function of the directional transfer functions;performing a second rendering being a rendering of the first estimated signal and the second estimated signal to generate a second intermediate stereo signal;

[0505] combining at least the first intermediate stereo signal and the second intermediate stereo signal to generate an output stereo signal; and

[0506] controlling the generator (317) to adapt a first relative property being a relative property between the first estimated signal and the second estimated signal in dependence on the spatial upmix parameters.

[0507] More specifically, the invention is defined by the appended CLAIMS.

Claims

CLAIMS:Claim 1. An audio apparatus comprising:a receiver (301) arranged to receive a data signal comprising encoded data for a mono downmix audio signal and a set of spatial upmix parameters for upmixing the mono downmix audio signal to a stereo signal, the set of spatial upmix parameters being indicative of relative signal properties of channels of the stereo signal;a store (311) comprising directional transfer functions for different directions, a directional transfer function for a given direction representing a mapping of a mono audio signal to stereo channels such that the mono audio signal is positioned in the given direction in a stereo image of the stereo channels;a decoder (303) arranged to generate the mono downmix audio signal by decoding the encoded data;a directional circuit (305) arranged to estimate a directional signal, the directional signal representing a point source audio component of the stereo signal by representing audio that has a point source origin;a decorrelator (315) arranged to generate a decorrelated mono downmix audio signal from the mono downmix audio signal;a generator (317) arranged to generate a first estimated signal and a second estimated signal from the mono downmix audio signal, at least the first estimated signal further being generated from the decorrelated mono downmix audio signal and the directional signal, the first estimated signal, and the second estimated signal forming a decomposition of the stereo signal;a direction determining circuit (309) arranged to determine a first direction from the spatial upmix parameters;a first Tenderer (307) arranged to perform a first rendering of the mono downmix audio signal to generate a first intermediate stereo signal, the first rendering being a directional rendering using a first channel transfer function retrieved from the store for the first direction;a second Tenderer (319) arranged to perform a second rendering being a rendering of the first estimated signal and the second estimated signal to generate a second intermediate stereo signal;a combiner (313) arranged to combine at least the first intermediate stereo signal and the second intermediate stereo signal to generate an output stereo signal; andan adapter (321) arranged to control the generator (317) to adapt a first relative property being a relative property between the first estimated signal and the second estimated signal in dependence on the spatial upmix parameters.Claim 2. The audio apparatus of claim 1 wherein the first relative property includes a level of the first estimated signal relative to the second estimated signal.Claim 3. The audio apparatus of claim 1 or 2 wherein the first relative property includes a correlation between the first estimated signal relative to the second estimated signal.Claim 4. The audio apparatus of any previous claim wherein the adapter (321) is arranged to adapt a second relative property being a relative property between the directional signal and the first estimated signal in dependence on the spatial upmix parameters.Claim 5. The audio apparatus of claim 4 wherein the second relative property includes a level of the directional signal relative to the first estimated signal.Claim 6. The audio apparatus of claim 4 or 5 wherein the second relative property includes a correlation between the directional signal relative and the first estimated signal.Claim 7. The audio apparatus of any previous claim wherein at least the directional signal and the first estimated signal are each a linear combination of the mono downmix audio signal and the decorrelated mono downmix audio signal.Claim 8. The audio apparatus of claim 7 wherein the adapter (321) is arranged to determine weights of the linear combination as a function of the spatial upmix parameters.Claim 9. The audio apparatus of any previous claim wherein the directional signal, the first estimated signal, and the second estimated signal are generated to be estimates of a signal model for the stereo signal, the signal model representing the stereo signal as a matrix multiplication of a signal vector by a matrix having a rank of two, the signal vector comprising a directional signal and two diffuse signals.Claim 10. The audio apparatus of claim 9 wherein the adapter (321) is arranged to adapt the first relative property to match a corresponding relative property between the two diffuse signals of the signal model.Claim 11. The audio apparatus of any previous claim wherein the generator (317) is arranged to generate the first estimated signal by applying second frequency band weights to frequency band samplesof the decorrelated mono downmix audio signal, the second frequency band weights being dependent on the spatial upmix parameters.Claim 12. The audio apparatus of any previous claim wherein the generator (317) is arranged to generate the second estimated signal from the decorrelated mono downmix audio signal.Claim 13. The audio apparatus of claim 12 wherein the generator (317) is arranged to generate the second estimated signal by applying third frequency band weights to frequency band samples of the decorrelated mono downmix audio signal, the third frequency band weights being dependent on the spatial upmix parameters.Claim 14. A method of operation for an audio apparatus, the method comprising:receiving a data signal comprising encoded data for a mono downmix audio signal and a set of spatial upmix parameters for upmixing the mono downmix audio signal to a stereo signal, the set of spatial upmix parameters being indicative of relative signal properties of channels of the stereo signal;providing directional transfer functions for different directions, a directional transfer function for a given direction representing a mapping of a mono audio signal to stereo channels such that the mono audio signal is positioned in the given direction in a stereo image of the stereo channels;generating the mono downmix audio signal by decoding the encoded data; estimating a directional signal, the directional signal representing a point source audio component of the stereo signal by representing audio that has a point source origin;generating a decorrelated mono downmix audio signal from the mono downmix audio signal;generating a first estimated signal and a second estimated signal from the mono downmix audio signal, at least the first estimated signal further being generated from the decorrelated mono downmix audio signal and the directional signal, the first estimated signal, and the second estimated signal forming a decomposition of the stereo signal;determining a first direction from the spatial upmix parameters;performing a first rendering of the mono downmix audio signal to generate a first intermediate stereo signal, the first rendering being a directional rendering using a first channel transfer function of the directional transfer functions;performing a second rendering being a rendering of the first estimated signal and the second estimated signal to generate a second intermediate stereo signal;combining at least the first intermediate stereo signal and the second intermediate stereo signal to generate an output stereo signal; andcontrolling the generator (317) to adapt a first relative property being a relative property between the first estimated signal and the second estimated signal in dependence on the spatial upmix parameters.Claim 15. A computer program product comprising computer program code means adapted to perform all the steps of claim 14 when said program is run on a computer.