Apparatus, method, or computer program for processing audio scenes encoded using parameter transformations

By transforming listener-specific parameters into channel-specific parameters using STFT and DFT-stereo processing, the latency and complexity issues in DirAC upmixing are addressed, enabling low-latency stereo output in immersive audio systems.

JP7787169B2Active Publication Date: 2025-12-16FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023521514
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-22
Filing Date
2021-10-08
Publication Date
2025-12-16
Estimated Expiration
2041-10-08

AI Technical Summary

Technical Problem

Existing audio processing systems, particularly in immersive VR/AR technologies, face challenges in achieving low-latency stereo output without additional delay due to the use of Directional Audio Coding (DirAC) upmixing, which is suboptimal in terms of delay and complexity, especially when only one transport channel is available.

Method used

The proposed solution involves transforming listener-specific parameters into channel-specific parameters using a short-time Fourier transform (STFT) filter bank and DFT-stereo processing, allowing low-latency upmixing of encoded audio scenes to stereo output within the latency requirements of communication codecs like IVAS, while reducing algorithmic complexity.

Benefits of technology

This approach reduces overall latency to 32 milliseconds, aligns stereo output configurations, and improves audio quality by avoiding additional delays and complexity, maintaining flexibility in processing and rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007787169000191
    Figure 0007787169000191
  • Figure 0007787169000192
    Figure 0007787169000192
  • Figure 0007787169000193
    Figure 0007787169000193
Patent Text Reader

Abstract

An apparatus for processing an encoded audio scene (130) representing a sound field associated with a virtual listener position, the encoded audio scene including information about a transport signal (122) and a first parameter set (112) associated with the virtual listener position, the apparatus comprising: a parameter converter (110) for converting the first parameter set (112) into a second parameter set (114) related to a channel representation including two or more channels for playback at a predetermined spatial position for the two or more channels; and an output interface (120) for generating a processed audio scene (124) using the second parameter set and information about the transport signal (122).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to audio processing, and in particular to the processing of encoded audio scenes for the purpose of generating processed audio scenes for rendering, storage or transmission. [Background technology]

[0002] Traditionally, audio applications providing a means for user communication, such as telephone or videoconferencing, have been primarily limited to mono recording and playback. However, in recent years, the emergence of new immersive VR / AR technologies has also increased interest in spatial rendering of communication scenarios. To meet this interest, a new 3GPP (registered trademark) audio standard called Immersive Voice and Audio Services (IVAS) is currently under development. Based on the recently released Enhanced Voice Services (EVS) standard, IVAS offers multi-channel and VR enhancements that can render immersive audio scenes, such as for spatial videoconferencing, while still meeting the low-latency requirements for smooth audio communication. This ongoing need to keep the overall latency of a codec to a minimum without sacrificing playback quality provides the motivation for the work described below.

[0003] Coding scene-based audio (SBA) material, such as third-order Ambisonics content, at low bitrates (e.g., 32 kbps or less) using parametric audio coding systems like Directional Audio Coding (DirAC) [1][2] allows for direct coding of only a single (transport) channel while restoring spatial information via side parameters in the filterbank-domain decoder. If the speaker configuration in the decoder allows only stereo playback, full reconstruction of the 3D audio scene is not required. Because higher bitrate coding of two or more transport channels is possible, in those cases, the stereophonic reproduction of the scene can be extracted and reproduced directly (e.g., due to additional filterbank analysis / synthesis, such as a complex-valued low-delay filterbank (CLDFB)) without any parametric spatial upmixing (which completely skips the spatial renderer) and the associated extra delay. However, for low-rate systems with only one transport channel, this is not possible. Therefore, for DirAC, a FOA (first-order Ambisonics) upmixing with the following L / R conversion was previously required for stereo output: This is problematic because in this case the overall delay is larger than other possible stereo output configurations in the system, and it is desirable to have all stereo output configurations aligned.

[0004] High-latency DirAC stereo rendering example FIG. 12 shows an example block diagram of a conventional decoder process for high-delay DirAC stereo upmix.

[0005] For example, in an encoder not shown, a single downmix channel is derived via spatial downmix in a DirAC encoder process and then encoded by a core coder such as Enhanced Voice Services (EVS) [3].

[0006] At the decoder, for example using the conventional DirAC upmix process depicted in Figure 12, one available transport channel is first decoded from the bitstream 1212 by using a mono or IVAS mono decoder 1210, resulting in a time-domain signal that can be viewed as a decoded mono downmix 1214 of the original audio scene.

[0007] The decoded mono signal 1214 is input to a CLDFB 1220 to analyze the delay-inducing signal 1214 (transform the signal into the frequency domain). The significantly delayed output signal 1222 is input to a DirAC renderer 1230. The DirAC renderer 1230 processes the delayed output signal 1222 and the transmitted side information, i.e., DirAC side parameters 1213, is used to convert the signal 1222 into a FOA representation, i.e., a FOA upmix 1232 of the original scene with spatial information recovered from the DirAC side parameters 1213.

[0008] The transmitted parameters 1213 may include directivity angles, e.g., one azimuth angle value for the horizontal plane and one elevation angle for the vertical plane, as well as one diffuseness value per frequency band to perceptually describe the entire 3D audio scene. Due to the per-band processing of the DirAC stereo upmix, the parameters 1213 are transmitted multiple times per frame, i.e., one set per frequency band. Furthermore, each set comprises multiple directivity parameters for individual subframes within the entire frame (e.g., 20 ms long) to increase the temporal resolution.

[0009] The result of the DirAC renderer 1230 can be, for example, a complete 3D scene in FOA format, i.e., FOA upmix 1232, which can be converted using a matrix transform 1240 into an L / R signal 1242 suitable for playback on a stereo speaker setup. In other words, the L / R signal 1242 can be input to stereo speakers or can be input to a CLDFB synthesis 1250 using predetermined channel weights. The CLDFB synthesis 1250 converts the two input frequency-domain output channels (L / R signal 1242) into the time domain, resulting in a stereo-playable output signal 1252.

[0010] Alternatively, the same DirAC stereo upmix can be used to directly generate a rendering of the stereo output configuration, avoiding the intermediate step of generating the FOA signal. This reduces the algorithmic complexity of the framework, potentially complicating it. Nevertheless, both approaches require the use of an additional filter bank after the core encoding, resulting in an additional delay of 5 ms. Further examples of DirAC rendering can be found in [2].

[0011] The DirAC stereo upmix approach is rather suboptimal in terms of both delay and complexity. By using the CLDFB filter bank, the output is significantly delayed (an additional 5 ms in the DirAC example) and therefore has the same overall delay as a full SBA upmix (compared to the delay of a stereo output configuration where no additional step of rendering is required). It is also a reasonable assumption that performing a full SBA upmix to generate a stereo signal is not ideal in terms of system complexity. Summary of the Invention [Problem to be solved by the invention]

[0012] The object of the present invention is to provide an improved concept for processing coded audio scenes. [Means for solving the problem]

[0013] This object is achieved by an apparatus for processing encoded audio scenes according to claim 1, a method for processing encoded audio scenes according to claim 32 or a computer program according to claim 33.

[0014] According to a first aspect relating to parameter transformation, the present invention is based on the discovery that an improved concept for processing an encoded audio scene can be obtained by transforming given parameters in the encoded audio scene associated with a virtual listener position into transformed parameters associated with a channel representation of a given output format. This procedure offers high flexibility in processing and ultimately rendering the processed audio scene in a channel-based environment.

[0015] An embodiment according to a first aspect of the present invention comprises an apparatus for processing an encoded audio scene representing a sound field associated with a virtual listener position, the encoded audio scene including information about a transport signal, e.g., a core-coded audio signal, and a first parameter set associated with the virtual listener position. The apparatus comprises a parameter converter for converting the first parameter set, e.g., Directional Audio Coding (DirAC) side parameters in B-Format or First Order Ambisonics (FOA) format, into a second parameter set, e.g., stereo parameters associated with a channel representation including two or more channels for reproduction at a predetermined spatial position of the two or more channels, and an output interface for generating a processed audio scene using the second parameter set and the information about the transport signal.

[0016] In an embodiment, a short-time Fourier transform (STFT) filter bank is used for the upmix rather than a directional audio coding (DirAC) renderer. This allows one downmix channel (contained in the bitstream) to be upmixed to a stereo output without additional overall delay. Using a window with a very short overlap for analysis in the decoder allows the upmix to stay within the overall delay required for communication codecs or upcoming immersive voice and audio services (IVAS). This value can be, for example, 32 milliseconds. In such an embodiment, any post-processing for bandwidth extension purposes can be avoided, since such processing can be performed in parallel with parameter conversion or parameter mapping.

[0017] By mapping listener-specific parameters of the low-band (LB) signal to a channel-specific stereo parameter set for the low-band, low-latency upmixing of the low-band in the DFT domain can be achieved. For the high-band, a single stereo parameter set allows for upmixing in the high-band in the time domain, preferably in parallel with low-band spectral analysis, spectral upmixing, and spectral synthesis.

[0018] Illustratively, the parameter transformer is configured to use a single-sided gain parameter for panning and a residual prediction parameter that is closely related to stereo width and also to the diffuseness parameter used in directional audio coding (DirAC).

[0019] This "DFT-stereo" approach, in embodiments, allows the IVAS codec to stay within the same overall latency as EVS, specifically 32 milliseconds, when processing encoded audio scenes (scene-based audio) to obtain a stereo output. By implementing simple processing via DFT-stereo instead of spatial DirAC rendering, the complexity of the parametric stereo upmix is ​​reduced.

[0020] The invention is based on the discovery that a second aspect relating to bandwidth extension provides an improved concept for processing coded audio scenes.

[0021] An embodiment according to a second aspect of the present invention comprises an apparatus for processing an audio scene representing a sound field, the audio scene including information about a transport signal and a parameter set, the apparatus further comprising an output interface for generating a processed audio scene using the information about the parameter set and the transport signal, the output interface being configured to generate a raw representation of two or more channels using the parameter set and the transport signal, a multi-channel enhancer for generating an augmented representation of the two or more channels using the transport signal, and a signal combiner for combining the raw representations of the two or more channels and the augmented representation of the two or more channels to obtain the processed audio scene.

[0022] Generating a raw representation of two or more channels on the one hand and an augmented representation of two or more channels separately on the other hand allows for great flexibility in choosing the algorithms for the raw and augmented representations. The final combination is already performed for each of the output channels, i.e., in the multi-channel output domain, rather than in the lower channel input or coded scene domain. Thus, following the combination, the two or more channels can be combined and used for further procedures such as rendering, transmission, or storage.

[0023] In an embodiment, part of the core processing, such as the bandwidth extension (BWE) of the Algebraic Code Excited Linear Prediction (ACELP) audio coder for the extended representation, can be performed in parallel with the DFT-stereo processing for the raw representation. Therefore, the delays caused by both algorithms are not accumulated, and only the given delay caused by one algorithm becomes the final delay. In an embodiment, only a transport signal, e.g., a low-band (LB) signal (channel), is input to an output interface, e.g., the DFT-stereo processing, while the high-band (HB) is separately upmixed in the time domain, e.g., using a multi-channel enhancer, so that the stereo decoding can be processed within a target time window of 32 milliseconds. For example, based on the mapped side gain from the parameter converter, e.g., by using wideband panning, a linear time-domain upmix of the entire high-band can be obtained without significant delay.

[0024] In an embodiment, the delay reduction in DFT-stereo may not result entirely from the difference in the overlap of the two transforms—for example, the 5 ms transform delay caused by CLDFB and the 3,125 ms transform delay caused by STFT. Instead, DFT-stereo exploits the fact that the last 3,25 ms from the 32 ms EVS coder target delay essentially comes from the ACELP BWE. Everything else (the remaining milliseconds until the EVS coder target delay is reached) is simply artificially delayed to finally achieve alignment of the two transformed signals (the HB stereo upmix signal and the HB filling signal with the LB stereo core signal). Therefore, to avoid additional delay in DFT-stereo, all other components of the encoder are transformed only within, for example, a very short DFT window overlap, while the ACELP BWE, for example, using a multi-channel enhancer, is mixed in the time domain with almost no delay.

[0025] According to a third aspect relating to parameter smoothing, the present invention is based on the discovery that an improved concept for processing coded audio scenes can be obtained by performing parameter smoothing over time according to a smoothing rule. Thus, processed audio scenes obtained by applying smoothed parameters rather than raw parameters to transport channels have improved audio quality. This is particularly true when the smoothed parameters are upmix parameters, but for any other parameters, such as envelope parameters, LPC parameters, noise parameters, or scale factor parameters, the use or smoothed parameters obtained by the smoothing rule result in an improved subjective audio quality of the resulting processed audio scenes.

[0026] An embodiment according to a third aspect of the present invention comprises an apparatus for processing an audio scene representing a sound field, the audio scene including information relating to a transport signal and a first parameter set, the apparatus further comprising: a parameter processor for processing the first parameter set to obtain a second parameter set, the parameter processor being configured to calculate at least one raw parameter for each output time frame using at least one parameter of the first parameter set for an input time frame, to calculate smoothing information, such as a coefficient, for each raw parameter according to a smoothing rule, and to apply the corresponding smoothing information to the corresponding raw parameter to derive parameters of the second parameter set for the output time frame, and an output interface for generating a processed audio scene using the second parameter set and the information relating to the transport signal.

[0027] Smoothing the raw parameters over time avoids strong fluctuations in gain or parameters from one frame to the next. The smoothing coefficient determines the strength of the smoothing, which is adaptively calculated in a preferred embodiment by a parameter processor, which also functions as a parameter converter for converting listener position-related parameters into channel-related parameters. The adaptive calculation allows for a faster response whenever the audio scene changes suddenly. The adaptive smoothing coefficient is calculated for each band from the energy change in the current band. The energy for each band is calculated for all subframes included in the frame. Furthermore, energy changes over time, characterized by two averages (short-term and long-term averages), do not affect smoothing in extreme cases, while small sudden increases in energy do not reduce smoothing. Therefore, a smoothing coefficient is calculated for each DTF-stereo subframe in the current frame from the quotient of the averages.

[0028] It should be mentioned herein that all the above and below mentioned alternatives or aspects can be used individually, i.e. without any aspect, however in other embodiments two or more aspects are combined with each other, and in other embodiments all aspects are combined with each other to obtain an improved compromise between overall delay, achievable audio quality and required implementation effort.

[0029] Preferred embodiments of the present invention are described below with reference to the accompanying drawings. [Brief explanation of the drawings]

[0030] [Figure 1] 1 is a block diagram of an apparatus for processing an audio scene encoded using a parameter transformer according to an embodiment; [Figure 2a] 1 shows a schematic diagram of a first set of parameters and a second set of parameters according to an embodiment; [Figure 2b]1 is an embodiment of a parameter transformer or parameter processor for calculating raw parameters. [Figure 2c] 1 is an embodiment of a parameter transformer or parameter processor for combining raw parameters. [Figure 3] 1 is an embodiment of a parameter transformer or parameter processor for performing a weighted combination of raw parameters. [Figure 4] 1 is an embodiment of a parameter transformer for generating side gain parameters and residual prediction parameters. [Figure 5a] 1 is an embodiment of a parameter transformer or parameter processor for calculating smoothing coefficients for raw parameters. [Figure 5b] 1 is an embodiment of a parameter transformer or parameter processor for calculating smoothing coefficients for frequency bands. [Figure 6] 1 shows a schematic diagram of averaging a transport signal of a smoothing factor according to an embodiment; [Figure 7] 1 is an embodiment of a parameter transformer parameter processor for computing recursive smoothing. [Figure 8] 1 is an embodiment of an apparatus for decoding a transport signal; [Figure 9] 1 is an embodiment of an apparatus for processing an audio scene coded using bandwidth extension; [Figure 10] 1 is an embodiment of an apparatus for obtaining a processed audio scene; [Figure 11] FIG. 1 is a block diagram of an embodiment of a multi-channel enhancer. [Figure 12] FIG. 1 is a block diagram of a conventional DirAC stereo upmix process. [Figure 13] 1 is an embodiment of an apparatus for obtaining a processed audio scene using parameter mapping; [Figure 14] 1 is an embodiment of an apparatus for obtaining an audio scene processed using bandwidth extension; DETAILED DESCRIPTION OF THE INVENTION

[0031] FIG. 1 illustrates an apparatus for processing an encoded audio scene 130, e.g., representing a sound field associated with a virtual listener position. The encoded audio scene 130 includes information about a transport signal 122, e.g., a bitstream, and a first parameter set 112, e.g., a plurality of DirAC parameters also included in the bitstream, associated with the virtual listener position. The first parameter set 112 is input to a parameter converter 110 or parameter processor, which converts the first parameter set 112 into a second parameter set 114 associated with a channel representation including at least two or more channels. The apparatus can support different audio formats. The audio signals may be acoustic in nature, picked up by a microphone, or electrical in nature and intended to be transmitted to a speaker. Supported audio formats can include mono signals, low-band signals, high-band signals, multi-channel signals, first-order and higher-order Ambisonics components, and audio objects. An audio scene can also be described by combining different input formats.

[0032] The parameter converter 110 is configured to calculate the second parameter set 114 as parametric stereo or multi-channel parameters, e.g., two or more channels, which are input to the output interface 120. The output interface 120 is configured to generate the processed audio scene 124 by combining the transport signal 122 or information about the transport signal with the second parameter set 114 to obtain the transcoded audio scene as the processed audio scene 124. Another embodiment includes upmixing the transport signal 122 using the second parameter set 114 to an upmix signal including two or more channels. In other words, the parameter converter 120 maps the first parameter set 112, e.g., used for DirAC rendering, to the second parameter set 114. The second parameter set may include side gain parameters used for panning and residual prediction parameters that, when applied in the upmix, result in an improved spatial image of the audio scene. For example, the parameters of the first parameter set 112 may include at least one of a direction of arrival parameter, a diffuseness parameter, a directional information parameter related to a sphere with the virtual listening position as the origin of the sphere, and a distance parameter. For example, the parameters of the second parameter set 114 may include at least one of a side gain parameter, a residual prediction gain parameter, an inter-channel level difference parameter, an inter-channel time difference parameter, an inter-channel phase difference parameter, and an inter-channel coherence parameter.

[0033] FIG. 2a shows a schematic diagram of a first parameter set 112 and a second parameter set 114 according to an embodiment. In particular, the parameter resolution of both parameters (first and second) is depicted. Each horizontal axis in FIG. 2a represents time, and each vertical axis in FIG. 2a represents frequency. As shown in FIG. 2a, an input time frame 210 associated with the first parameter set 112 includes two or more input time subframes 212 and 213. Directly below, an output time frame 220 associated with the second parameter set 114 is shown in a corresponding diagram related to the diagram above. This indicates that the output time frame 220 is smaller than the input time frame 210 and longer than the input time subframes 212 or 213. It should be noted that the input time subframes 212 or 213 and the output time frame 220 can include multiple frequencies as frequency bands. The input frequency band 230 can include the same frequencies as the output frequency band 240. According to an embodiment, the frequency bands of the input frequency band 230 and the output frequency band 240 may not be connected or correlated with each other.

[0034] 4 are typically calculated per frame, such that a single side gain and a single residual gain are calculated per input frame 210. However, in other embodiments, rather than just a single side gain and a single residual gain being calculated for each frame, a group of side gains and a group of residual gains are calculated for the input time frame 210, where each side gain and each residual gain is associated with a particular input time subframe 212 or 213 of a frequency band, for example. Thus, in an embodiment, the parameter converter 110 calculates a group of side gains and a group of residual gains for each frame of the first parameter set 112 and the second parameter set 114, and the number of side and residual gains for the input time frame 210 is typically equal to the number of input frequency bands 230.

[0035] 2b shows an embodiment of the parameter converter 110 for calculating 250 raw parameters 252 of the second parameter set 114. The parameter converter 110 calculates the raw parameters 252 for each of two or more input time subframes 212 and 213 in a temporally sequential manner. For example, the calculation 250 derives, for each input frequency band 230 and time point (input time subframes 212, 213), a dominant direction of arrival (DOA) in azimuth angle θ and a dominant direction of arrival in elevation angle φ and divergence parameter ψ.

[0036] For directional components such as X, Y, and Z, the first order spherical harmonic at the center position is given by the omnidirectional components w(b,n) and DirAC parameters, which can be derived using the following equation: The W channel represents the omnidirectional mono component of the signal, corresponding to the output of an omnidirectional microphone. The X, Y, and Z channels are three-dimensional directional components. From these four FOA channels, a stereo signal (stereo version, stereo output) can be obtained by decoding the W and Y channels using the parameter converter 110, which results in two cardioids pointing at azimuth angles of +90 and -90 degrees. Therefore, the following equation shows the left-right relationship of the stereo signal, where the left channel L is represented by adding the Y channel to the W channel and the right channel R is represented by subtracting the Y channel from the W channel: TIFF0007787169000005.tif1325 In other words, this decoding corresponds to first-order beamforming pointing in two directions, which can be expressed using the following equations: TIFF0007787169000006.tif7109As a result, there is a direct link between the stereo output (left and right channels) and the first parameter set 112, ie the DirAC parameters.

[0037] However, on the other hand, the second parameter set 114, i.e., the DFT parameters, depends on a model of the left L and right R channels based on the intermediate signal M and the side signal S, which can be expressed using the following equations: TIFF0007787169000007.tif1323 where M is transmitted as a mono signal (channel) corresponding to the omnidirectional channel W in scene-based audio (SBA) mode. Furthermore, in the DFT, the stereo S is predicted from M using the side gain parameters described below.

[0038] 4 shows an embodiment of the parameter transformer 110 for generating side gain parameters 455 and residual prediction parameters 456, for example using calculation process 450. The parameter transformer 110 preferably processes calculations 250 and 450 to calculate raw parameters 252, for example side parameters 455 for output frequency band 241, using the following equations: According to the equation, b is the output frequency band, sidegain is the side gain parameter 455, azimuth is the azimuth component of the direction of arrival parameter, and elevation is the elevation component of the direction of arrival parameter. As shown in FIG. 4 , the first parameter set 112 includes the direction of arrival (DOA) parameters 456 for the input frequency bands 231 as described above, and the second parameter set 114 includes the side gain parameters 455 for each input frequency band 230. However, if the first parameter set 112 further includes a spread parameter ψ 453 for the input frequency bands 231, then the parameter converter 110 is configured to calculate 250 the side gain parameters 455 for the output frequency bands 241 using the following equation: TIFF0007787169000009.tif13157 According to the formula, diff(b) is the diffuseness parameter ψ 453 for the input frequency band b 230. Note that the directivity parameters 456 of the first parameter set 112 may include different value ranges, e.g., the azimuth parameter 451 is [0;360], the elevation parameter 452 is [0;180], and the resulting side gain parameter 455 is [-1;1]. As shown in FIG. 2c, the parameter converter 110 uses a combiner 260 to combine at least two raw parameters 252, resulting in the parameters of the second parameter set 114 associated with the output time frame 220.

[0039] According to an embodiment, the second parameter set 114 further includes residual prediction parameters 456 for the output frequency band 241 of the output frequency band 240 shown in Fig. 4. The parameter converter 110 can use the spreadness parameter ψ 453 from the input frequency band 231, as indicated by the residual selector 410, as the residual prediction parameters 456 for the output frequency band 241. If the input frequency band 231 and the output frequency band 241 are equal to each other, the parameter converter 110 uses the spreadness parameter ψ 453 from the input frequency band 231. The spreadness parameter ψ 453 for the output frequency band 241 is derived from the spreadness parameter ψ 453 for the input frequency band 231, and the spreadness parameter ψ 453 is used for the output frequency band 241 as the residual prediction parameters 456 for the output frequency band 241. The parameter converter 110 can then use the spreadness parameter ψ 453 from the input frequency band 231.

[0040] In DFT stereo processing, the residual of the prediction using the residual selector 410 is assumed and expected to be incoherent, modeled by its energy decorrelating the residual signals towards the left L and right R. The residual of the prediction of the side signal S with the intermediate signal M as a mono signal (channel) can be expressed as: TIFF0007787169000010.tif766 That energy is modeled in DFT stereo processing using the residual prediction gain using the following formula: The residual gain represents the inter-channel incoherence and spatial width of the stereo signal, and is therefore directly linked to the diffuse part modeled by DirAC. Therefore, the residual energy can be rewritten as a function of the DirAC diffuseness parameter: TIFF0007787169000012.tif748

[0041] 3 illustrates a parameter converter 110 for performing a weighted combination 310 of raw parameters 252 according to an embodiment. At least two raw parameters 252 are input to the weighted combination 310, and weighting factors 324 for the weighted combination 310 are derived based on an amplitude-related measure 320 of the transport signal 122 in the corresponding input time subframe 212. Furthermore, the parameter converter 110 is configured to use the energy or power value of the transport signal 122 in the corresponding input time subframe 212 or 213 as the amplitude-related measure 320. The amplitude-related measure 320 may, for example, measure the energy or power of the transport signal 122 in the corresponding input time subframe 212, such that the weighting factor 324 for that input subframe 212 is larger if the transport signal 122 in the corresponding input time subframe 212 has a higher energy or power than the weighting factor 324 for an input subframe 212 with a lower energy or power of the transport signal 122 in the corresponding input time subframe 212.

[0042] As mentioned above, the directivity parameters, azimuth parameters, and elevation parameters have corresponding ranges of values. However, the direction parameters of the first parameter set 112 typically have a higher time resolution than the second parameter set 114, which means that more than one azimuth and elevation value must be used to calculate one side gain value. According to an embodiment, the calculation is based on energy-dependent weights that can be obtained as the output of the amplitude-related measure 320. For example, all For TIFF0007787169000013.tif74 input temporal subframes 212 and 213, the energy nrg of the subframe is calculated using the following formula: TIFF0007787169000014.tif2582 where, TIFF0007787169000015.tif73 is the time-domain input signal, TIFF0007787169000016.tif74 is the number of samples in each subframe, and TIFF0007787169000017.tif73 is the sample index. Furthermore, each output time frame For TIFF0007787169000018.tif72230, the weights 324 are then Each input temporal subframe in TIFF0007787169000019.tif72 The contribution of TIFF0007787169000020.tif73212, 213 can be calculated as follows: TIFF0007787169000021.tif1352The side gain parameter 455 is then finally calculated using the following formula: Due to the similarity between the TIFF0007787169000022.tif13116 parameters, the diffuseness parameters 453 per band are directly mapped to the residual prediction parameters 456 of all subframes in the same band. The similarity can be expressed by the following equation: TIFF0007787169000023.tif786

[0043] 5a shows an embodiment of the parameter converter 110 or parameter processor for calculating a smoothing coefficient 512 for each raw parameter 252 according to a smoothing rule 514. Furthermore, the parameter converter 110 is configured to apply the smoothing coefficient 512 (corresponding smoothing coefficient for one raw parameter) to the raw parameters 252 (one raw parameter corresponding to the smoothing coefficient) to derive the parameters of the second parameter set 114 for the output time frame 220, i.e., the parameters of the output time frame.

[0044] 5b shows an embodiment of the parameter converter 110 or parameter processor for calculating smoothing coefficients 522 for frequency bands using a compression function 540. The compression function 540 may be different for different frequency bands, such that the compression strength of the compression function 540 is stronger for lower frequency bands than for higher frequency bands. The parameter converter 110 is further configured to calculate the smoothing coefficients 512, 522 using maximum boundary selection 550. In other words, the parameter converter 110 may obtain the smoothing coefficients 512, 522 by using different maximum boundaries for different frequency bands, such that the maximum boundary for the lower frequency band is higher than the maximum boundary for the higher frequency band.

[0045] Both the compression function 540 and the maximum boundary selection 550 are input to a calculation 520 that obtains a smoothing coefficient 522 for the frequency band 522. For example, the parameter converter 110 is not limited to using two calculations 510 and 520 to calculate the smoothing coefficients 512 and 522; as a result, the parameter converter 110 is configured to calculate the smoothing coefficients 512 and 522 using only one calculation block that can output the smoothing coefficients 512 and 522. In other words, the smoothing coefficients are calculated for each band (for each raw parameter 252) from the energy changes in the current frequency band. For example, by using a parameter smoothing process, the side gain parameters 455 and the residual prediction parameters 456 are smoothed over time to avoid large gain fluctuations. This requires relatively strong smoothing in most cases, but a faster response is required whenever the audio scene 130 changes suddenly, so the smoothing coefficients 512 and 522, which determine the strength of the smoothing, are adaptively calculated.

[0046] Therefore, the energy per band nrg is calculated for all subframes using the following formula: Calculated for TIFF0007787169000024.tif73: TIFF0007787169000025.tif1981 where, TIFF0007787169000026.tif73 is the frequency bins (real and imaginary) of the DFT transformed signal, TIFF0007787169000027.tif72 is the current frequency band This is the bin index across all bins in TIFF0007787169000028.tif73.

[0047] To capture the change in energy across the two averages, one short-term average 331 and one long-term average 332 are calculated using the amplitude-related measure 320 of the transport signal 122, as shown in FIG.

[0048] FIG. 6 shows a schematic diagram of an amplitude-related measure 320 averaging a transport signal 122 over a smoothing factor 512, according to an embodiment. The x-axis represents time and the y-axis represents energy (of the transport signal 122). The transport signal 122 shows a schematic portion of a sine function 122. As shown in FIG. 6, the second time portion 631 is shorter than the first time portion 632. The change in energy over the averages 331 and 332 is calculated for each band according to the following equation: Calculated for TIFF0007787169000029.tif73: TIFF0007787169000030.tif1361 and TIFF0007787169000031.tif1958 where, TIFF0007787169000032.tif712 and TIFF0007787169000033.tif711 is the number of time subframes before which the individual averages are calculated TIFF0007787169000034.tif73. For example, in this particular embodiment, TIFF0007787169000035.tif712 is set to a value of 3, TIFF0007787169000036.tif711 is set to the value 10.

[0049] Furthermore, the parameter converter or parameter processor 110 is configured to use calculation 510 to calculate smoothing factors 512, 522 based on the ratio between the long-term average 332 and the short-term average 331. In other words, the quotient of the two averages 331 and 332 is calculated, so that a higher short-term average, indicating a recent increase in energy, leads to less smoothing. The following equation shows the correlation between the smoothing factor 512 and the two averages 331 and 312: Due to the fact that a higher long-term average 332, which indicates a decrease in energy, does not lead to a decrease in smoothing, the smoothing factor 512 is (currently) set to a maximum of 1. As a result, the above equation becomes: TIFF0007787169000038.tif725 minimum TIFF0007787169000039.tif1310 (0.3 in this embodiment). However, in extreme cases the coefficient needs to be close to 0, which can be achieved by using the following formula to find the value in the range Range from TIFF0007787169000040.tif1317 This is why it is converted to [TIFF0007787169000041.tif79]. TIFF0007787169000042.tif19116

[0050] In an embodiment, the smoothing is excessively reduced compared to the smoothing shown previously, so that the coefficients are compressed by a root function that tends towards the value 1. Since stability is particularly important in the lowest bands, the fourth root is used in the frequency band TIFF0007787169000043.tif711 and Used in TIFF0007787169000044.tif711. The formula for minimum bandwidth is: TIFF0007787169000045.tif1359 All other bands The formula for TIFF0007787169000046.tif711 performs compression with a square root function using the following formula: TIFF0007787169000047.tif1359 All other bands By applying a square root function to TIFF0007787169000048.tif711, extreme cases where the energy may increase exponentially are made smaller, and less rapid increases in energy do not reduce the smoothing as much.

[0051] Furthermore, the maximum smoothing is set according to the frequency band for the following equation: Note that a factor of 1 simply repeats the previous value without the contribution of the current gain. TIFF0007787169000049.tif788 where, TIFF0007787169000050.tif721 represents a given implementation with five bands configured according to the following table:

[0052] [Table 1] The smoothing coefficients are the sum of the DFT stereo subframes in the current frame. Calculated for each of TIFF0007787169000052.tif73.

[0053] Figure 7 shows the side gain parameter TIFF0007787169000053.tif722455 and residual prediction gain parameters TIFF0007787169000054.tif723456 show a parameter transformer 110 according to an embodiment using recursive smoothing 710, where both are recursively smoothed: TIFF0007787169000055.tif7161 and A recursive smoothing 710 over temporally subsequent output time frames of the current output time frame is calculated by combining the parameters of the preceding output time frame 532 weighted by the first weight value and the raw parameters 252 for the current output time frame 220 weighted by the second weight value. In other words, the smoothed parameters for the current output time frame are calculated such that the first weight value and the second weight value are derived from the smoothing coefficients for the current time frame.

[0054] These mapped and smoothed parameters (g side ,g pred ) is input to the output interface 120 for DFT stereo processing, i.e., the stereo signal ( TIFF0007787169000057.tif710 is downmix TIFF0007787169000058.tif710, Residual prediction signal TIFF0007787169000059.tif712, and the mapped parameters TIFF0007787169000060.tif710 and Generated from TIFF0007787169000061.tif711. For example, downmix TIFF0007787169000062.tif710 is obtained from the downmix either by enhanced stereo filling using an all-pass filter, or by stereo filling using a delay.

[0055] The upmix is ​​described by the following equation: TIFF0007787169000063.tif7148 and TIFF0007787169000064.tif7156 Upmixing is performed in the frequency bands listed in the table above. All bins in TIFF0007787169000065.tif73 Subframe in TIFF0007787169000066.tif72 Each side gain is processed. TIFF0007787169000068.tif710 is the downmix as above Energy and residual prediction gain parameters in TIFF0007787169000069.tif710 TIFF0007787169000070.tif712 or Energy normalization factor calculated from TIFF0007787169000071.tif722 Weighted by TIFF0007787169000072.tif712.

[0056] The mapped and smoothed side gains 755 and the mapped and smoothed residual gains 756 are input to the output interface 120 to obtain a smoothed audio scene. Based on the above description, processing the encoded audio scene using smoothed parameters therefore provides an improved compromise between achievable audio quality and implementation effort.

[0057] 8 shows an apparatus for decoding a transport signal 122 according to an embodiment. An (encoded) audio signal 816 is input to a transport signal core decoder 810 for core-decoding the (core-encoded) audio signal 816 to obtain a (decoded raw) transport signal 812, which is input to an output interface 120. For example, the transport signal 122 may be the encoded transport signal 812 output from the transport signal core encoder 810. The (decoded) transport signal 812 is input to the output interface 120, which is configured to generate a raw representation 818 of two or more channels, e.g., a left channel and a right channel, using a parameter set 814 including the second parameter set 114. For example, the transport signal core decoder 810 for decoding the core-encoded audio signal to obtain the transport signal 122 is an ACELP decoder. Furthermore, the core decoder 810 is configured to provide the decoded raw transport signal 812 to two parallel branches, a first of which comprises the output interface 120 and a second of which comprises the transport signal enhancer 820 or the multi-channel enhancer 990, or both. The signal combiner 940 is configured to receive a first input to be combined from the first branch and a second input to be combined from the second branch.

[0058] 9, the apparatus for processing the encoded audio scene 130 can use a bandwidth extension processor 910. The low-band transport signal 901 is input to the output interface 120 to obtain a two-channel low-band representation of the transport signal 972. It should be noted that the output interface 120 processes the transport signal 901 in the frequency domain 955, for example during the upmixing process 960, and transforms the two-channel transport signal 901 in the time domain 966. This is done by a transformer 970, which transforms the upmixed spectral representation 962 representing the frequency domain 955 into the time domain to obtain the two-channel low-band representation of the transport signal 972.

[0059] 8, the single-channel low-band transport signal 901 is input to a transformer 950, which performs a transformation of, for example, a time portion of the transport signal 901 corresponding to the output time frame 220 into a spectral representation 952 of the transport signal 901, i.e., a transformation from the time domain 966 to the frequency domain 955. For example, as described in FIG. 2, the portion (of the output time frame) is shorter than the input time frame 210 in which the parameters 252 of the first parameter set 112 are organized.

[0060] The spectral representation 952 is input to the upmixer 960, which upmixes the spectral representation 952, for example, using the second parameter set 114, to obtain an upmixed spectral representation 962 that is (still) processed in the frequency domain 955. As described above, the upmixed spectral representation 962 is input to the transformer 970 to transform the upmixed spectral representation 962, i.e., each channel of the two or more channels, from the frequency domain 955 to the time domain 966 (time representation) to obtain a low-band representation 972. Thus, the two or more channels in the upmixed spectral representation 962 are calculated. Preferably, the output interface 120 is configured to operate in the complex discrete Fourier transform domain, and the upmix operation is performed in the complex discrete Fourier transform domain. The transformation from the complex discrete Fourier transform domain to a real-valued time-domain representation is performed using the transformer 970. In other words, the output interface 120 is configured to generate a raw representation of two or more channels using an upmixer 960 in a second domain, i.e., the frequency domain 955, while the first domain represents the time domain 966.

[0061] In an embodiment, the upmix operation of the upmixer 960 is based on the following equation: TIFF0007787169000073.tif77= TIFF0007787169000074.tif1345 and TIFF0007787169000075.tif1310= TIFF0007787169000076.tif1359, where: TIFF0007787169000077.tif78 is the transport signal 901 for frame t and frequency bin k, TIFF0007787169000078.tif77 is the side gain parameter 455 for frame t and subband b, TIFF0007787169000079.tif76 is the residual prediction gain parameter 456 for frame t and subband b, and g norm is an energy adjustment factor that may or may not be present, TIFF0007787169000080.tif77 is the raw residual signal for frame t and frequency bin k.

[0062] The transport signals 902, 122, in contrast to the low-band transport signal 901, are processed in the time domain 966. The transport signal 902 is input to a bandwidth extension processor (BWE processor) 910 to generate a high-band signal 912, and to a multi-channel filter 930 to apply a multi-channel filling operation. The high-band signal 912 is input to an upmixer 920 to upmix the high-band signal 912 into an upmixed high-band signal 922 using parameters from a second parameter set 144, i.e., an output time frame 262, 532. For example, the upmixer 920 may apply a wideband panning process to the high-band signal 912 in the time domain 966 using at least one parameter from the second parameter set 114.

[0063] The low-band representation 972, the upmixed high-band signal 922, and the multi-channel filling transport signal 932 are input to a signal combiner 940, which combines, in the time domain 966, the results of the wideband panning 922, the results of the stereo filling 932, and the low-band representation of the two or more channels 972. This combination results in a full-band multi-channel signal 942 in the time domain 966 as a channel representation. As outlined above, the converter 970 converts each channel of the two or more channels in the spectral representation 962 into a time representation to obtain a raw time representation of the two or more channels 972. Thus, the signal combiner 940 combines the raw time representation of the two or more channels with an extended time representation of the two or more channels.

[0064] In an embodiment, only the low-band (LB) transport signal 901 is input to the output interface 120 (DFT stereo) processing, and the high-band (HB) transport signal 912 is separately upmixed in the time domain (using an upmixer 920). Such a process is implemented for a panning operation using the BWE processor 910 and time-domain stereo filling using a multi-channel filler 930 to generate ambience contributions. The panning process includes wideband panning based on mapped side gains, e.g., frame-wise mapped and smoothed side gains 755. Here, there is only one gain per frame covering the complete high-band frequency range, which simplifies the calculation of the left and right high-band channels from the downmix channel based on the following equation: Each subframe Sample in TIFF0007787169000081.tif73 For each TIFF0007787169000082.tif72, TIFF0007787169000083.tif7105 and TIFF0007787169000084.tif7121.

[0065] High-bandwidth stereo-filling signal TIFF0007787169000085.tif716, i.e., the multi-channel filling transport signal 932, is expressed as follows: Delay TIFF0007787169000086.tif715, TIFF0007787169000087.tif714 Weight it by the energy normalization factor TIFF0007787169000088.tif712 is obtained by further using: All samples in the current time frame For TIFF0007787169000089.tif72 (done for the whole time frame 210, not the time subframes 213 and 213), TIFF0007787169000090.tif799 and TIFF0007787169000091.tif13101. TIFF0007787169000092.tif73 is the number of samples by which the highband downmix is ​​delayed to generate the filling signal 932 obtained by the multi-channel filler 930. Other methods for generating the filling signal apart from the delay can be implemented, such as a more advanced decorrelation process or the use of a noise signal or any other signal derived from the transport signal in a different way compared to the delay.

[0066] Both the panned stereo signals 972 and 922 and the generated stereo filling signal 932 are combined (mixed back) into the core signal after DFT synthesis using a signal combiner 940 .

[0067] This described process of the ACELP highband is also in contrast to the high-delay DirAC processing, where the ACELP core and TCX frames are artificially delayed to be aligned with the ACELP highband, so the CLDFB (analysis) is performed on the complete signal, which means that the upmix of the ACELP highband is also done in the CLDFB domain (frequency domain).

[0068] 10 shows an embodiment of an apparatus for obtaining a processed audio scene 124. The transport signal 122 is input to the output interface 120 to generate a raw representation of two or more channels 972 using the second parameter set 114 and a multi-channel enhancer 990 to generate an enhanced representation 992 of the two or more channels. For example, the multi-channel enhancer 990 is configured to perform at least one operation from a group of operations including a bandwidth extension operation, a gap filling operation, a quality enhancement operation, or an interpolation operation. To obtain the processed audio scene 124, both the raw representation of the two or more channels 972 and the enhanced representation 992 of the two or more channels are input to a signal combiner 940.

[0069] 11 shows a block diagram of an embodiment of a multi-channel enhancer 990 for generating an enhanced representation 992 of two or more channels, including a transport signal enhancer 820, an upmixer 830, and a multi-channel filler 930. The transport signal 122 and / or the decoded raw transport signal 812 are input to the transport signal enhancer 820, which generates an enhanced transport signal 822, which is input to the upmixer 830 and the multi-channel filler 930. For example, the transport signal enhancer 820 is configured to perform at least one operation from a group of operations including a bandwidth expansion operation, a gap filling operation, a quality enhancement operation, or an interpolation operation.

[0070] 9 , the multi-channel filler 930 generates a multi-channel filling transport signal 932 using the transport signal 902 and at least one parameter 532. In other words, the multi-channel enhancer 990 is configured to generate an extended representation of two or more channels 992 using the extended transport signal 822 and the second parameter set 114, or using the extended transport signal 822 and the upmixed extended transport signal 832. For example, the multi-channel enhancer 990 includes either the upmixer 830 or the multi-channel filler 930, or both, to generate the extended representation 992 of two or more channels using the transport signal 122 or the extended transport signal 933 and at least one parameter of the second parameter set 532. In an embodiment, the transport signal enhancer 820 or the multi-channel enhancer 990 is configured to operate in parallel with the output interface 120 when generating the raw representation 972, or the parameter converter 110 is configured to operate in parallel with the transport signal enhancer 820.

[0071] In Figure 13, the bitstream 1312 transmitted from the encoder to the decoder may be the same as the DirAC-based upmixing scheme shown in Figure 12. The single transport channel 1312 derived from the DirAC-based spatial downmixing process is input to the core decoder 1310, decoded by the core decoder, e.g., an EVS or IVAS mono decoder, and transmitted together with the corresponding DirAC side parameters 1313.

[0072] In this DFT stereo approach to processing audio scenes without extra delay, the initial decoding in the mono core decoder of the transport channel (IVAS mono decoder) also remains unchanged. Instead of passing through the CLDFB filter bank 1220 from Figure 12, the decoded downmix signal 1314 is input to a DFT analysis 1320 to transform the decoded mono signal 1314 into the STFT domain (frequency domain), such as by using windows with very short overlap. Thus, the DFT analysis 1320 does not introduce any additional delay to the target system delay of 32 ms, using only the remaining headroom between the overall delay and that already introduced by the MDCT analysis / synthesis of the core decoder.

[0073] The DirAC side parameters 1313 or first parameter set 112 are input to a parameter mapping 1360, which may include, for example, a parameter transformer 110 or a parameter processor to obtain DFT stereo side parameters, i.e., the second parameter set 114. The frequency-domain signal 1322 and the DFT side parameters 1362 are input to a DFT stereo decoder 1330, which generates a stereo upmix signal 1332, for example by using the upmixer 960 described in FIG. 9. The two channels of the stereo upmix 1332 are input to a DFT synthesis, which transforms the stereo upmix 1332 from the frequency domain to the time domain, for example using the transformer 970 described in FIG. 9, resulting in an output signal 1342 that can represent the processed audio scene 124.

[0074] FIG. 14 illustrates an embodiment for processing an encoded audio scene using bandwidth extension 1470. A bitstream 1412 is input to an ACELP core or low-band decoder 1410, instead of an IVAS mono decoder as described in FIG. 13, to generate a decoded low-band signal 1414. The decoded low-band signal 1414 is input to a DFT analysis 1420 to convert the signal 1414 into a frequency-domain signal 1422, e.g., the spectral representation 952 of the transport signal 901 from FIG. 9. The DFT stereo decoder 1430 may represent an upmixer 960 that generates a LB stereo upmix 1432 using a decoded low-band signal 1442 in the frequency domain and DFT stereo side parameters 1462 from a parameter mapping 1460. The generated LB stereo upmix 1432 is input to a DFT synthesis block 1440, which performs a transformation to the time domain, e.g., using the transformer 970 of FIG. 9. The low-band representation 972 of the transport signal 122, i.e., the output signal 1442 of the DFT synthesis stage 1440, is input to a signal combiner 940 which combines the upmixed high-band stereo signal 922 and the multi-channel filling high-band transport signal 932 with the low-band representation of the transport signal 972 to result in a full-band multi-channel signal 942.

[0075] The decoded LB signal 1414 and parameters 1415 for the BWE 1470 are input to the ACELP BWE decoder 910 to generate the decoded highband signal 912. A mapped side gain 1462, e.g., the mapped and smoothed side gain 755 for the lowband spectral region, is input to the DFT stereo block 1430, and the mapped and smoothed single side gain for the entire highband is forwarded to the highband upmix block 920 and the stereo filling block 930. The HB upmix block 920 for upmixing the decoded HB signal 912 using the highband side gain 1472, e.g., the parameters 532 of the output time frame 262 from the second parameter set 114, generates the upmixed highband signal 922. The stereo filling block 930 for filling the decoded highband transport signal 912, 902 uses the parameters 532, 456 of the output time frame 262 from the second parameter set 114 and generates the highband filling transport signal 932.

[0076] In conclusion, embodiments according to the present invention create a concept for processing coded audio scenes using parameter transformation and / or using bandwidth extension and / or using parameter smoothing, resulting in an improved compromise between overall delay, achievable audio quality and implementation effort.

[0077] Further embodiments of aspects of the present invention, particularly combinations of aspects of the present invention, are presented below. The proposed solution for achieving low-delay upmixing relies on the use of parametric stereo techniques, such as those described in [4], using short-time Fourier transform (STFT) filter banks rather than the DirAC renderer. This "DFT-stereo" technique describes the upmixing of one downmix channel to a stereo output. The advantage of this method is that windows with very short overlap are used for the DFT analysis in the decoder, allowing it to stay within the much lower overall delay required for communication codecs such as EVS [3] or the upcoming IVAS codec (32 ms). Also, unlike DirAC CLDFB, DFT stereo processing is not a post-processing step relative to the core coder, but is performed in parallel with part of the core processing, namely the bandwidth extension (BWE) of the Algebraic Code Exit Excitation Prediction (ACELP) speech coder, without exceeding this already-given delay. Therefore, with respect to the 32 ms delay of EVS, DFT stereo processing can be referred to as delay-free, since it operates with the same overall coder delay. DirAC, on the other hand, can be seen as a post-processor that introduces an additional 5ms of delay to extend the overall delay to 37ms.

[0078] Generally, latency gains are achieved: the low latency comes from processing steps that occur in parallel with the core processing, whereas in the exemplary CLDFB version, post-processing steps to perform the necessary rendering occur after the core encoding.

[0079] Unlike DirAC, DFT stereo utilizes an artificial delay of 3.25 ms for all components except ACELP BWE, by only transforming those components into the DFT domain using windows with a very short overlap of 3.125 ms that fit into the available headroom without introducing more delay. Thus, only TCX and ACELP without BWE are upmixed in the frequency domain, while ACELP BWE is upmixed in the time domain by a separate delay-free processing step called inter-channel bandwidth extension (ICBWE) [5]. For the special stereo output of a given embodiment, this time-domain BWE processing is slightly modified, as will be explained towards the end of the embodiment.

[0080] The transmitted DirAC parameters cannot be directly used for DFT stereo upmixing. Therefore, it is necessary to map the given DirAC parameters to the corresponding DFT stereo parameters. While DirAC uses azimuth and elevation angles for spatial positioning along with the diffuseness parameter, DFT stereo has a single side gain parameter used for panning and a residual prediction parameter that is closely related to the stereo width and thus the diffuseness parameter of DirAC. From the perspective of parameter resolution, each frame is divided into two subframes and several frequency bands per subframe. The side gains and residual gains used in DFT stereo are described in [6].

[0081] The DirAC parameters are derived from a band-by-band analysis of the audio scene, originally in B format or FOA. Then, for each band k and time instant n, the azimuth angle TIFF0007787169000093.tif714 and elevation angle TIFF0007787169000094.tif715 and diffusion coefficient The main direction of arrival of TIFF0007787169000095.tif714 is derived. For the directional component, the first order spherical harmonic function at the center position is TIFF0007787169000096.tif714 and can be derived using DirAC parameters. TIFF0007787169000097.tif1366TIFF0007787169000098.tif13118TIFF0007787169000099.tif13118TIFF0007787169000100.tif1393

[0082] Furthermore, from the FOA channel, a stereo version can be obtained by decoding with W and Y, which results in two cardioids pointing at azimuth angles of +90 and -90 degrees. TIFF0007787169000101.tif1323This decoding corresponds to first-order beamforming pointing in two directions. As a result, there is a direct link between the stereo output and the DirAC parameters, while the DFT parameters depend on a model of the L and R channels based on the mid-signal M and side-signal S. TIFF0007787169000103.tif1322M is transmitted as a mono channel and corresponds to the omnidirectional channel W in case of SBA mode. In the DFT, the stereo S is predicted from M using the side gain, which can be expressed using the DirAC parameters as follows: TIFF0007787169000104.tif13152

[0083] In DFT stereo, the residual of the prediction is assumed and expected to be incoherent, modeled by its energy, which decorrelates the residual signal towards left and right. The residual of the prediction of S by M can be expressed as: TIFF0007787169000105.tif1362And that energy is modeled in DFT stereo using prediction gain as follows: The residual gain represents the inter-channel incoherence and spatial width of the stereo signal, and is therefore directly linked to the diffuse part modeled by DirAC. Therefore, the residual energy can be rewritten as a function of the DirAC diffuseness parameter: TIFF0007787169000107.tif748

[0084] The band configuration of commonly used DFT stereo is not the same as that of DirAC, so it must be adapted to cover the same frequency range as the DirAC bands. For these bands, the directivity angle of DirAC is: can be mapped to the DFT stereo side gain parameters by TIFF0007787169000108.tif13149, where TIFF0007787169000109.tif73 is the current band, and the parameter range is TIFF0007787169000110.tif715, About elevation angle TIFF0007787169000111.tif715 and the resulting side gain values TIFF0007787169000112.tif714. However, the directivity parameters in DirAC usually have a higher time resolution than DFT stereo, which means that more than one azimuth and elevation value must be used to calculate one side gain value. One way to do this is to do averaging across subframes, but in this implementation the calculation is based on energy-dependent weights. All For TIFF0007787169000113.tif74DirAC subframes, the energy of the subframe is Calculated as TIFF0007787169000114.tif2582, where: TIFF0007787169000115.tif73 is the time-domain input signal, TIFF0007787169000116.tif74 is the number of samples in each subframe, and TIFF0007787169000117.tif73 is the sample index for each DFT stereo subframe Regarding TIFF0007787169000118.tif72, Internal as TIFF0007787169000119.tif1352 Each DirAC subframe of TIFF0007787169000120.tif72 A weight can be calculated for the contribution of TIFF0007787169000121.tif73.

[0085] The side gain is then The final result is calculated as TIFF0007787169000122.tif19166.

[0086] Due to the similarity between the parameters, one diffuseness value per band is directly mapped to the residual prediction parameters of all subframes in the same band. TIFF0007787169000123.tif763 Furthermore, the parameters are smoothed over time to avoid strong fluctuations in the gain. This requires relatively strong smoothing in most cases, but a faster response whenever the scene changes suddenly, so the smoothing coefficient that determines the strength of the smoothing is calculated adaptively. This adaptive smoothing coefficient is calculated for each band from the change in energy in the current band. Thus, initially all subframes The bandwidth energy needs to be calculated for TIFF0007787169000124.tif73: TIFF0007787169000125.tif1981 where, TIFF0007787169000126.tif73 is the frequency bins (real and imaginary) of the DFT transformed signal, TIFF0007787169000127.tif72 is the current bandwidth These are the bin indices of all bins in TIFF0007787169000128.tif73.

[0087] To capture the energy change over time, one short-term and one long-term time period are then calculated for each band. Regarding TIFF0007787169000129.tif73, TIFF0007787169000130.tif1361 and Calculated according to TIFF0007787169000131.tif1958.

[0088] where: TIFF0007787169000132.tif712 and TIFF0007787169000133.tif711 is the number of subframes before which the individual averages are calculated TIFF0007787169000134.tif73. In this particular implementation, TIFF0007787169000135.tif712 is set to 3, TIFF0007787169000136.tif711 is set to 10. A smoothing factor is then calculated from the quotient of the means, so that a higher short-term average, indicating a recent increase in energy, leads to less smoothing. TIFF0007787169000137.tif1371The smoothing factor is set here to a maximum of 1, since a higher long-term average, which indicates a decrease in energy, does not lead to a decrease in smoothing.

[0089] The above formula is TIFF0007787169000138.tif725 minimum TIFF0007787169000139.tif1310 (0.3 in this implementation). However, in extreme cases the coefficient needs to be close to 0, which means Values ​​range through TIFF0007787169000140.tif19116 Range from TIFF0007787169000141.tif1317 This is why it is converted to [TIFF0007787169000142.tif79].

[0090] In the less extreme cases, the smoothing is reduced too much, so the coefficients are compressed by the root function towards the value 1. Since stability is especially important in the lowest bands, the fourth order roots are TIFF0007787169000143.tif711 and Used in TIFF0007787169000144.tif711: TIFF0007787169000145.tif1359 Meanwhile, all other bands TIFF0007787169000146.tif711 is the square root Compressed by TIFF0007787169000147.tif1359.

[0091] In this way, extreme cases remain close to 0, but sudden increases in energy do not reduce smoothing too much.

[0092] Finally, the maximum smoothing is set according to the band (a factor of 1 simply repeats the previous value without the contribution of the current gain): TIFF0007787169000148.tif788 where, in the given implementation, it has five bands TIFF0007787169000149.tif721 is set according to the table below.

[0093] [Table 2] The smoothing coefficients are calculated for each DFT stereo subframe in the current frame. Calculated for TIFF0007787169000151.tif73.

[0094] In the final step, both the side gains and the residual prediction gains are recursively smoothed according to: TIFF0007787169000152.tif13153 and TIFF0007787169000153.tif13156 These mapped and smoothed parameters are now fed into the DFT stereo processing, where the stereo signal TIFF0007787169000154.tif78 is downmixed Residual prediction signal generated from TIFF0007787169000155.tif710 TIFF0007787169000156.tif712 (obtained from the downmix either by "enhanced stereo filling" using an all-pass filter [7] or by regular stereo filling using delay) as well as the mapped parameters TIFF0007787169000157.tif710 and TIFF0007787169000158.tif711 is generated. The upmix is ​​generally described by the following equation [6]: TIFF0007787169000159.tif7148 and Bandwidth All bins in TIFF0007787169000160.tif73 Each subframe of TIFF0007787169000161.tif72 Regarding TIFF0007787169000162.tif73, TIFF0007787169000163.tif13149Furthermore, each side gain TIFF0007787169000164.tif710 is TIFF0007787169000165.tif710 and Energy normalization factor calculated from the energy of TIFF0007787169000166.tif712 Weighted by TIFF0007787169000167.tif712.

[0095] Finally, the upmix signal is transformed back to the time domain via IDFT and played back in a given stereo setting.

[0096] The "time-domain bandwidth extension" (TBE) [8] used in ACELP generates its own delay (in the implementation, this embodiment is based on exactly 2.3125 ms), so it cannot be transformed into the DFT domain while the overall delay remains within 32 ms (3.25 ms remains for a stereo decoder whose STFT already uses 3.125 ms). Therefore, only the low band (LB) is put into the DFT stereo processing shown by 1450 in Figure 14, while the high band (HB) must be upmixed separately in the time domain as shown in block 920 in Figure 14. In regular DFT stereo, this is done via inter-channel bandwidth extension (ICBWE) [5] for panning for ambience and time-domain stereo filling. In a given case, the stereo filling in block 930 is calculated in the same way as regular DFT stereo. However, the ICBWE processing is skipped entirely due to missing parameters and replaced by a low-resource wideband panning required in block 920 based on the mapped side gain 1472. In the given embodiment, there is only a single gain that covers the complete HB region, which simplifies the calculation of the left and right HB channels in block 920 from the downmix channel to: TIFF0007787169000168.tif7105 and Each subframe Sample in TIFF0007787169000169.tif73 About TIFF0007787169000170.tif72 TIFF0007787169000171.tif13115HB stereo filling signal TIFF0007787169000172.tif716 is delayed in block 930. TIFF0007787169000173.tif714 and Weighted by TIFF0007787169000174.tif714, energy normalization factor as follows Obtained by TIFF0007787169000175.tif712. TIFF0007787169000176.tif1393 and All samples in the current frame (not subframes, but full frames) About TIFF0007787169000177.tif72 TIFF0007787169000178.tif13101, where TIFF0007787169000179.tif73 is the number of samples the HB downmix is ​​delayed relative to the filling signal.

[0097] Both the panned stereo signal and the generated stereo filling signal are finally mixed back into the core signal after DFT synthesis in combiner 940 .

[0098] This special processing of ACELP HB is also in contrast to the high-delay DirAC processing, where the ACELP Core and TCX frames are artificially delayed to be aligned with the ACELP HB, so CLDFB is performed on the complete signal, i.e. the upmix of ACELP HB is also done in the CLDFB domain.

[0099] Advantages of the proposed method The lack of additional delay allows the IVAS codec to stay within the same overall delay as in EVS (32 ms) for this particular case of SBA input to stereo output.

[0100] Due to the overall simplicity and easier processing, the complexity of the DFT-based parametric stereo upmix is ​​much lower than that of spatial DirAC rendering.

[0101] Further Preferred Embodiments 1. An apparatus, method or computer program for encoding or decoding as described above.

[0102] 2. An apparatus or method for encoding or decoding, or related computer program, a system in which the input is encoded by a model based on a spatial audio representation of an acoustic scene having a first set of parameters and is decoded at the output using a stereo model for two output channels or a multi-channel model for more than two output channels having a second set of parameters; and / or Mapping spatial parameters to stereo parameters, and / or Transformation of one frequency domain based input representation / parameters into another frequency domain based output representation / parameters, and / or Transformation of parameters with higher time resolution into parameters with lower time resolution, and / or Lower output delay due to shorter window overlap of the second frequency translation, and / or Mapping DirAC parameters (direction angle, diffusion) to DFT stereo parameters (side gain, residual prediction gain) to output SBA DirAC encoded content as stereo, and / or Transformation from CLDFB-based input representations / parameters to DFT-based output representations / parameters, and / or Conversion of 5ms resolution parameters to 10ms resolution parameters, and / or Advantage: An apparatus or method, or related computer program, for encoding or decoding, including lower output delay due to shorter DFT window overlap compared to CLDFB.

[0103] It should be noted that, herein, all alternatives or aspects described above, and all aspects defined by the following independent claims, can be used individually, i.e., without any alternatives or purposes other than those contemplated in the alternatives, purposes or independent claims. However, in other embodiments, two or more alternatives or aspects or independent claims can be combined with each other, and in other embodiments, all aspects or alternatives and all independent claims can be combined with each other.

[0104] It should be outlined that different aspects of the present invention relate to a parameter transformation aspect, a smoothing aspect, and a bandwidth extension aspect, which can be implemented separately or independently of each other, or any two of at least the three aspects can be combined, or all three aspects can be combined in the above-mentioned embodiments.

[0105] The coded signals of the present invention can be stored on a digital or non-transitory storage medium or can be transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0106] While some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or apparatus corresponds to a method step or feature of a method step, and similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or function of a corresponding apparatus.

[0107] Depending on particular implementation requirements, embodiments of the present invention can be implemented in hardware or software, using a digital storage medium such as, for example, a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM or flash memory, on which electronically readable control signals are stored and which cooperate (or can cooperate) with a programmable computer system to perform the respective method.

[0108] Some embodiments of the present invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0109] Generally, embodiments of the present invention can be implemented as a computer program product comprising program code that operates to perform one of the methods when the computer program product is run on a computer, and the program code may for example be stored on a machine-readable carrier.

[0110] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier or a non-transitory storage medium.

[0111] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0112] A further embodiment of the inventive method is, therefore, a data carrier (or digital storage medium, or computer readable medium) having recorded thereon the computer program for performing one of the methods described herein.

[0113] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, the data stream or sequence of signals being adapted to be transferred via a data communication connection, such as, for example, the Internet.

[0114] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0115] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0116] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0117] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented by way of illustration and description of the embodiments herein.

[0118] References [1] V. Pulkki, M.-VVJ Laitinen, J. Ahonen, T. Lokki and T. Pihlajamaeki, "Directional audio coding-perception - based reproduction of spatial sound," in INTERNATIONAL WORKSHOP ON THE PRINCIPLES AND APPLICATION ON SPATIAL HEARING, 2009.

[0119] [2] G. Fuchs, O. Thiergart, S. Korse, S. Doehla, M. Multrus, F. Kuech, Boutheon, A. Eichenseer and S. Bayer, "Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to dirac based spatial audio coding using low-order, mid-order and high-order components generators". WO Patent 2020115311A1, 11 06 2020.

[0120] [3] 3GPP TS 26.445, Codec for Enhanced Voice Services (EVS); Detailed algorithmic description.

[0121] [4] S. Bayer, M. Dietz, S. Doehla, E. Fotopoulou, G. Fuchs, W. Jaegers, G. Markovic, M. Multrus, E. Ravelli and M. Schnell, " APPARATUS AND METHOD FOR ESTIMATING AN INTER-CHANNEL TIME DIFFERENCE". Patent WO17125563, 27 07 2017.

[0122] [5] V. S. C. S. Chebiyyam and V. Atti, "Inter-channel bandwidth extension". WO Patent 2018187082A1, 11 10 2018.

[0123] [6] J. Buethe, G. Fuchs, W. Jaegers, F. Reutelhuber, J. Herre, E. Fotopoulou, M. Multrus and S. Korse, "Apparatus and method for encoding or decoding a multichannel signal using a side gain and a residual gain". WO Patent WO2018086947A1, 17 05 2018.

[0124] [7] J. Buethe, F. Reutelhuber, S. Disch, G. Fuchs, M. Multrus and R. Geiger, "Apparatus for Encoding or Decoding an Encoded Multichannel Signal Using a Filling Signal Generated by a Broad Band Filter". WO Patent WO2019020757A2, 31 01 2019.

[0125] [8] V. A. e. al., "Super-wideband bandwidth extension for speech in the 3GPP EVS codec," in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, 2015.

Claims

1. 1. An apparatus for processing an encoded audio scene (130) representing a sound field associated with a virtual listener position, the encoded audio scene (130) comprising information about a transport signal (122) and a first set of parameters (112) associated with the virtual listener position; a parameter converter (110) for converting the first parameter set (112), which is related to the virtual listener position and is listener-specific, into a second parameter set (114), which is related to and channel-specific channel representations, which channel representations include two or more channels for playback at predetermined spatial positions for the two or more channels, and which includes at least one DirAC (Directional Audio Coding) parameter for each input time frame (210) of a plurality of input time frames and each input frequency band (231) of a plurality of input frequency bands (230), and which is configured to calculate the second parameter set (114) as parametric stereo or multi-channel parameters; an output interface (120) for generating a processed audio scene (124) using the second parameter set (114) and the information about the transport signal (122).

2. the output interface (120) is configured to upmix the transport signal (122) using the second parameter set (114) into an upmix signal including the two or more channels.

10. The apparatus of claim 1.

3. 2. The apparatus of claim 1, wherein the output interface is configured to generate the processed audio scene by combining the transport signal or the information about the transport signal with the second parameter set to obtain a transcoded audio scene as the processed audio scene.

4. the at least one DirAC parameter includes at least one of a direction of arrival parameter, a diffuseness parameter, a directional information parameter related to a sphere with the virtual listening position as the sphere's origin, and a distance parameter; 2. The apparatus of claim 1, wherein the parametric stereo or multi-channel parameters include at least one of a side gain parameter (455), a residual prediction gain parameter (456), an inter-channel level difference parameter, an inter-channel time difference parameter, an inter-channel phase difference parameter, and an inter-channel coherence parameter.

5. the input time frame (210) to which the first parameter set (112) is associated includes two or more input time subframes (212, 213), and the output time frame (220) to which the second parameter set (114) is associated is smaller than the input time frame (210) and longer than the input time subframes (212, 213) of the two or more input time subframes (212, 213); 5. The apparatus of claim 1, wherein the parameter converter is configured to calculate channel-specific raw parameters of the second parameter set for each of the two or more temporally succeeding input time subframes from listener-specific parameters for each of the two or more temporally succeeding input time subframes, and to combine at least two channel-specific raw parameters for the two or more temporally succeeding input time subframes to derive channel-specific parameters of the second parameter set associated with the output time frame.

6. 6. The apparatus of claim 5, wherein the parameter converter is configured to perform a weighted combination of the at least two channel-specific raw parameters for the two or more temporally subsequent input time subframes, the weighting coefficients of the weighted combination being derived based on amplitude-related measures of the transport signal in the corresponding input time subframes.

7. 7. The apparatus of claim 6, wherein the amplitude-related measure is the energy or power of the transport signal in the corresponding input time subframe, and wherein a weighting factor for an input time subframe is greater if the energy or power of the transport signal in the corresponding input time subframe is higher compared to a weighting factor for an input time subframe in which the energy or power of the transport signal in the corresponding input time subframe is lower.

8. the parameter transformer (110) is configured to calculate at least one channel-specific raw parameter (252) for each output time frame (220) using at least one parameter of the first parameter set (112) for the input time frame (210); the parameter converter (110) is configured to calculate a smoothing coefficient (512; 522) for each channel-specific raw parameter (252) according to a smoothing rule; 5. The apparatus of claim 1, wherein the parameter converter is configured to apply corresponding smoothing coefficients to the corresponding channel-specific raw parameters to derive the parameters of the second parameter set for the output time frame.

9. The parameter converter (110) calculating a long-term average (332) over the amplitude-related measure (320) of the first time portion of the transport signal (122); calculating a short-term average (331) over the amplitude-related measure (320) of a second time portion of the transport signal (122), the second time portion being shorter than the first time portion; 9. The device according to claim 8, configured to calculate a smoothing factor (512; 522) based on the ratio of the long-term average (332) to the short-term average (331).

10. 10. The apparatus according to claim 8 or 9, wherein the parameter converter (110) is configured to calculate the smoothing coefficients (512; 522) for the bands using a compression function (540), the compression function being different for different frequency bands, the compression strength of the compression function being stronger for lower frequency bands than for higher frequency bands.

11. 11. The apparatus according to claim 8, wherein the parameter converter (110) is configured to calculate the smoothing coefficients (512; 522) using different maximum bounds for different bands, the maximum bound for a low band being higher than the maximum bound for a high band.

12. 12. The apparatus of claim 8, wherein the parameter converter is configured to apply, as the smoothing rule, a recursive smoothing rule over temporally subsequent output time frames, such that smoothed parameters for a current output time frame are calculated by combining the parameters for a previous output time frame weighted by a first weight value and the channel-specific raw parameters for the current output time frame weighted by a second weight value, the first weight value and the second weight value being derived from the smoothing coefficients for the current time frame.

13. the input time frame (210) to which the first parameter set (112) is associated includes two or more input time subframes (212, 213), and the output time frame (220) to which the second parameter set (114) is associated is smaller than the input time frame (210) and longer than the input time subframes (212, 213) of the two or more input time subframes (212, 213); the parameter converter (110) is configured to calculate channel-specific raw parameters (252) of the second parameter set (114) for each of the two or more temporally succeeding input time subframes (212, 213) from listener-specific parameters (251) for each of the two or more temporally succeeding input time subframes (212, 213), and to combine at least two channel-specific raw parameters for the two or more temporally succeeding input time subframes (212, 213) to derive channel-specific parameters of the second parameter set (114) associated with the output time frame (220); The output interface (120) performing a transformation into a spectral representation of a time portion of the transport signal (122) corresponding to the output time frame (220), the time portion of the transport signal (122) being shorter than the input time frame (210) in which the parameters of the first parameter set (112) are organized; performing an upmix operation on the spectral representation using the second set of channel-specific parameters (114) to obtain the two or more channels in the spectral representation; 5. The apparatus of claim 1, configured to convert each of the two or more channels in the spectral representation to a temporal representation.

14. The output interface (120) Transform into the complex discrete Fourier transform domain, performing the upmix operation in the complex discrete Fourier transform domain; 14. The apparatus of claim 13, configured to perform the transformation from the complex discrete Fourier transform domain to a real-valued time domain representation.

15. The output interface (120) is configured to perform the upmix operation based on the following equation: = and where: is the transport signal (122) for frame t and frequency bin k, is a first channel of the two or more channels in the spectral representation for frame t and frequency bin k; is a second channel of the two or more channels in the spectral representation for frame t and frequency bin k; is the side gain parameter (455) for frame t and subband b, is the residual prediction gain parameter (456) for frame t and subband b, and g norm is an energy adjustment factor that may or may not be present, 15. The apparatus of claim 13 or 14, wherein {right arrow over (x)} is the raw residual signal for frame t and frequency bin k.

16. the first set of listener-specific parameters (112) includes direction-of-arrival parameters for the input frequency bands (231), and the second set of channel-specific parameters (114) includes side gain parameters (455) for each input frequency band (231); The parameter converter (110) is configured to calculate the side gain parameter (455) for an output frequency band (241) using the following formula: where b is the output frequency band (241), sidegain is the side gain parameter (455), azimuth is the azimuth component of the direction of arrival parameter, and elevation is the elevation component of the direction of arrival parameter.

16. An apparatus according to any one of claims 1 to 15.

17. the first listener-specific parameter set (112) further comprises a diffuseness parameter for the input frequency band (231), and the parameter converter (110) is configured to calculate the side gain parameter (455) for the output frequency band (241) using the following formula: where diff(b) is the spreading parameter for the input frequency band (231)b.

17. The apparatus of claim 16.

18. the first parameter set (112) includes a spreading parameter for an input frequency band (231); the second parameter set (114) includes a residual prediction gain parameter (456) for an output frequency band (241); the parameter converter (110) is configured to use the spread parameter from the input frequency band (231) as the residual prediction gain parameter (456) for the output frequency band (241) when the input frequency band (231) and the output frequency band (241) are equal to each other, or to derive the spread parameter for the output frequency band (241) from the spread parameter for the input frequency band (231) and then use the spread parameter for the output frequency band (241) as the residual prediction gain parameter (456) for the output frequency band (241).

18. Apparatus according to any one of claims 1 to 17.

19. The information about the transport signal (122) includes a core encoded audio signal, and the device:

19. The apparatus of any one of claims 13 to 18, further comprising a core decoder (810) for core decoding the core encoded audio signal to obtain the transport signal (122).

20. said core decoder (810) being within an ACELP decoder; or the output interface (120) is configured to convert the transport signal (122), which is a low-band signal, into a spectral representation, upmix the spectral representation, and convert the upmixed spectral representation in the time domain to obtain a low-band representation of the two or more channels; the apparatus comprising a bandwidth extension processor (910) for generating a high-bandwidth signal from the transport signal (122) in the time domain; the apparatus comprising a multi-channel filler (930) for applying a multi-channel filling operation to the transport signal (122) in the time domain; the apparatus comprising an upmixer (920) for applying wideband panning in the time domain to the highband signal using at least one parameter from the second parameter set (114); 20. The apparatus of claim 19, comprising a signal combiner (940) for combining, in the time domain, a result of the wideband panning, a result of the multi-channel filling operation, and the low-band representations of the two or more channels to obtain a full-band multi-channel signal in the time domain as the channel representation.

21. the output interface (120) is configured to generate a raw representation of the two or more channels using the second parameter set (114) and the transport signal (122); the apparatus further comprising a multi-channel enhancer (990) for generating an enhanced representation of the two or more channels using the transport signal (122); the apparatus further comprising a signal combiner (940) for combining the raw representation of the two or more channels and the augmented representation of the two or more channels to obtain the processed audio scene (124).

19. Apparatus according to any one of claims 1 to 18.

22. the multi-channel enhancer (990) is configured to generate an enhanced representation (992) of the two or more channels using the enhanced transport signal (822) and the second set of parameters (114); or 22. The apparatus of claim 21, wherein the multi-channel enhancer comprises a transport signal enhancer for generating an enhanced transport signal and an upmixer for upmixing the enhanced transport signal.

23. The transport signal (122) is an encoded transport signal, and the device a core decoder (810) for generating a decoded raw transport signal; the transport signal enhancer (820) is configured to generate the enhanced transport signal using the decoded raw transport signal; 23. The apparatus of claim 22, wherein the output interface (120) is configured to generate the raw representation of the two or more channels using the second parameter set (114) and the decoded raw transport signal.

24. 24. The apparatus of claim 22 or 23, wherein the multi-channel enhancer (990) comprises either the upmixer or the multi-channel filler (930), or both the upmixer and the multi-channel filler (930), for generating the extended representation of the two or more channels using the transport signal (122) or the extended transport signal (822) and at least one parameter of the second parameter set (114).

25. the output interface (120) is configured to generate a raw representation of the two or more channels using an upmix in a second domain; the transport signal enhancer (820) is configured to generate the extended transport signal (822) in a first domain different from the second domain, or the multi-channel enhancer (990) is configured to generate the extended representation of the two or more channels using the extended transport signal (822) in the first domain; 25. The apparatus of claim 22, 23, or 24, wherein the signal combiner (940) is configured to combine the raw representation of the two or more channels and the augmented representation of the two or more channels in the first region.

26. 26. The apparatus of claim 25, wherein the first domain is the time domain and the second domain is the spectral domain.

27. 27. The apparatus of claim 22, wherein the transport signal enhancer (820) or the multi-channel enhancer (990) is configured to perform at least one operation from a group of operations comprising a bandwidth extension operation, a gap filling operation, a quality enhancement operation, or an interpolation operation.

28. the transport signal enhancer (820) or the multi-channel enhancer (990) is configured to operate in parallel with the output interface (120) when generating the raw representation, or the parameter converter (110) is configured to operate in parallel with the transport signal enhancer (820); 28. Apparatus according to any one of claims 22 to 27.

29. 24. The apparatus of claim 23, wherein the core decoder (810) is configured to supply the decoded raw transport signal to two parallel branches, a first branch of the two parallel branches comprising the output interface (120), a second branch of the two parallel branches comprising the transport signal enhancer (820) or the multi-channel enhancer (990) or both, and the signal combiner (940) is configured to receive a first input to be combined from the first branch and a second input to be combined from the second branch.

30. The output interface (120) performing a transformation into a spectral representation of a time portion of said transport signal (122) corresponding to an output time frame (220); performing an upmix operation on the spectral representation using the second parameter set (114) to obtain the two or more channels in the spectral representation; converting each of the two or more channels in the spectral representation to a time representation to obtain a raw time representation of the two or more channels; 30. The apparatus of claim 20, 21, 25 or 29, wherein the signal combiner (940) is configured to combine the raw time representations of the two or more channels and the extended representations of the two or more channels.

31. 1. A method for processing an encoded audio scene representing a sound field associated with a virtual listener position, the encoded audio scene comprising information about a transport signal and a first set of parameters (112) associated with the virtual listener position, converting the first parameter set (112) associated with the virtual listener position and listener-specific into a second parameter set (114), the second parameter set (114) associated with and channel-specific to a channel representation, the channel representation including two or more channels for playback at a predetermined spatial position for the two or more channels, the first parameter set (112) including at least one DirAC (Directional Audio Coding) parameter for each input time frame (210) of a plurality of input time frames and each input frequency band (231) of a plurality of input frequency bands (230), the converting including calculating the second parameter set (114) as parametric stereo or multi-channel parameters; generating a processed audio scene using the second parameter set (114) and the information about the transport signal.

32. 32. A computer program for performing the method of claim 31 when the computer program is run on a computer or processor.

Citation Information

Patent Citations

  • Apparatus and method for multi-channel parameter conversion

    JP2010507114A

  • Apparatus, method, and computer program for upmixing a downmixed audio signal using phase smoothing.

    JP2012512438A

  • Apparatus and method for converting a first parametric spatial audio signal to a second parametric spatial audio signal.

    JP2013514696A

  • Temporal Offset Estimation

    JP2019504344A

  • Apparatus and method for encoding or decoding a multi-channel signal using side gains and residual gains - Patents.com

    JP2019536112A