Methods for parametric multi-channel encoding

By generating bitstreams of downmix signals and spatial metadata, the parameter multi-channel audio encoding system is optimized using the Dolby Digital Plus encoder and parameter processing unit, bandwidth efficiency and computational efficiency problems are solved, the system robustness is enhanced and the auditory quality is maintained.

JP2025122080APending Publication Date: 2025-08-20DOLBY INTERNATIONAL AB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025083802
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-02-21
Filing Date
2025-05-20
Publication Date
2025-08-20

AI Technical Summary

Technical Problem

The existing parameter multi-channel audio encoding system has the need for improvement in bandwidth efficiency, computing efficiency and robustness.

Method used

By generating a bit stream representing the downmix signal and spatial metadata, the downmix signal is encoded using the Dolby Digital Plus encoder, and the quantized spatial metadata is inserted into the bit stream, the spatial parameters are determined in combination with the parameter processing unit, the number of bits is reduced by quantization and entropy coding technology, and the metadata transmission is optimized using window functions and frequency differential coding.

Benefits of technology

Improve bandwidth efficiency and computing efficiency, enhance system robustness while maintaining auditory quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025122080000001_ABST
    Figure 2025122080000001_ABST
Patent Text Reader

Abstract

To provide an efficient method and a system for parametric multi-channel audio encoding.SOLUTION: An audio encoding system 500 that generates a bitstream 564 indicating spatial metadata includes a downmix processing unit 510 that generates a downmix signal from a multi-channel input signal 561. The downmix signal has m channels, and the multi-channel input signal 561 has n channels (where n and m are integers and m<n). The system also includes a parameter processing unit 520 that determines spatial metadata from the multi-channel input signal, and a configuration unit 540 that determines control settings for the parameter processing unit on the basis of external settings. The external settings include a target data-rate for the bitstream 564, and the control settings include a maximum data-rate for the spatial metadata.SELECTED DRAWING: Figure 5a
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 61 / 767,673, filed on February 21, 2013. The content of the application is hereby incorporated by reference in its entirety.

[0002] Technical Field This document relates to audio encoding systems. Specifically, this document relates to efficient methods and systems for parametric multi - channel audio encoding.

Background Art

[0003] Parametric multi - channel audio encoding systems can be used to provide improved listening quality, especially at particularly low data rates. Nevertheless, there is a need to further improve such parametric multi - channel audio encoding systems, especially with respect to bandwidth efficiency, computational efficiency, and / or robustness.

Summary of the Invention

Means for Solving the Problems

[0004] According to one aspect, an audio encoding system is described that is configured to generate a bitstream indicative of a downmix signal and spatial metadata. The spatial metadata may be used by a corresponding decoding system to generate a multi - channel upmix signal from the downmix signal. The downmix signal may have m channels, and the multi - channel upmix signal may have n channels, where n and m are integers and m < n. In one example, n = 6 and m = 2. The spatial metadata may allow a corresponding decoding system to generate n channels of the multi - channel upmix signal from m channels of the downmix signal.

[0005] The audio encoding system may be configured to quantize and / or encode the downmix signal and the spatial metadata and insert the quantized / encoded data into a bitstream. In particular, the downmix signal may be encoded using a Dolby Digital Plus encoder and the bitstream may correspond to a Dolby Digital Plus bitstream. The quantized / encoded spatial metadata may be inserted into a data field of the Dolby Digital Plus bitstream.

[0006] The audio encoding system may include a downmix processing unit configured to generate a downmix signal from a multi-channel input signal. The downmix processing unit is also referred to herein as a downmix coding unit. The multi-channel input signal may have n channels, as may the multi-channel upmix signal regenerated based on the downmix signal. In particular, the multi-channel upmix signal may provide an approximation of the multi-channel input signal. The downmix unit may include a Dolby Digital Plus encoder as described above. The multi-channel upmix signal and the multi-channel input signal may be 5.1 or 7.1 signals, and the downmix signal may be a stereo signal.

[0007] The audio encoding system may have a parameter processing unit configured to determine spatial metadata from the multi-channel input signal. In particular, the parameter processing unit (also referred to herein as a parameter encoding unit) may be configured to determine one or more spatial parameters, e.g., a set of spatial parameters. The parameters may be determined based on various combinations of channels of the multi-channel input signal. A spatial parameter of the set of spatial parameters may indicate a cross-correlation between different channels of the multi-channel input signal. The parameter processing unit may be configured to determine spatial metadata for frames of the multi-channel input signal, referred to as spatial metadata frames. A frame of the multi-channel input signal typically includes a predetermined number (e.g., 1536) of samples of the multi-channel input signal. Each spatial metadata frame may include one or more sets of spatial parameters.

[0008] The audio encoding system may further include a configuration setting unit configured to determine one or more control settings for the parameter processing unit based on one or more external settings. The one or more external settings may include a target data rate for the bitstream. Alternatively or additionally, the one or more external settings may include one or more of: a sampling rate of the multi-channel input signal, a number m of channels of the downmix signal, a number n of channels of the multi-channel input signal, and / or an update period indicating a time period during which a corresponding decoding system is required to synchronize to the bitstream. The one or more control settings may include a maximum data rate for spatial metadata. In the case of a spatial metadata frame, the maximum data rate for spatial metadata may indicate a maximum number of metadata bits for a spatial metadata frame. Alternatively or additionally, the one or more control settings may include one or more of: a temporal resolution setting indicating the number of sets of spatial parameters per spatial metadata frame to be determined; a frequency resolution setting indicating the number of frequency bands over which spatial parameters are to be determined; a quantizer setting indicating the type of quantizer to be used to quantize spatial metadata; and an indication of whether the current frame of the multi-channel input signal should be encoded as an independent frame.

[0009] The parameter processing unit may be configured to determine whether the number of bits for the spatial metadata frame determined according to the one or more control settings exceeds a maximum number of metadata bits. Further, the parameter processing unit may be configured to reduce the number of bits for a particular spatial metadata frame if it is determined that the number of bits for the particular spatial metadata frame exceeds the maximum number of metadata bits. This reduction in the number of bits may be performed in a resource (processing power) efficient manner. In particular, this reduction in the number of bits may be performed without having to recalculate the complete spatial metadata frame.

[0010] As indicated above, a spatial metadata frame may include one or more sets of spatial parameters. The one or more control settings may include a temporal resolution setting indicating the number of sets of spatial parameters to be determined by the parameter processing unit per spatial metadata frame. The parameter processing unit may be configured to determine, for the current spatial metadata frame, the number of sets of spatial parameters indicated by the temporal resolution setting. Typically, the temporal resolution setting takes a value of 1 or 2. Furthermore, the parameter processing unit may be configured to discard a set of spatial parameters from the current spatial metadata frame if the current spatial metadata frame has multiple sets of spatial parameters and if the number of bits of the current spatial metadata frame exceeds the maximum number of metadata bits. The parameter processing unit may be configured to retain at least one set of spatial parameters per spatial metadata frame. By discarding a set of spatial parameters from the spatial metadata frame, the number of bits of the spatial metadata frame can be reduced with little computational effort and without significantly affecting the perceived listening quality of the multi-channel upmix signal.

[0011] The one or more sets of spatial parameters are typically associated with one or more corresponding sampling points, which may indicate one or more corresponding time points. In particular, a sampling point may indicate a time point at which the decoding system should fully apply the corresponding set of spatial parameters. In other words, a sampling point may indicate a time point for which the corresponding set of spatial parameters was determined.

[0012] The parameter processing unit may be configured to discard a first set of spatial parameters from the current spatial metadata frame if the sampling points of the current metadata frame are not associated with a transient component of the multi-channel input signal, where the first set of spatial parameters is associated with a first sampling point that precedes a second sampling point. On the other hand, the parameter processing unit may be configured to discard a second set of spatial parameters (typically the last set) from the current spatial metadata frame if the sampling points of the current metadata frame are associated with a transient component of the multi-channel input signal. In this way, the parameter processing unit may be configured to reduce the impact of discarding a set of spatial parameters on the listening quality of the multi-channel upmix signal.

[0013] The one or more control settings may include a quantizer setting indicating a first type quantizer from a plurality of predetermined types of quantizers. The plurality of predetermined types of quantizers may each provide a different quantizer resolution. In particular, the plurality of predetermined types of quantizers may include fine quantization and coarse quantization. The parameter processing unit may be configured to quantize the one or more sets of spatial parameters of the current spatial metadata frame according to the first type quantizer. Furthermore, the parameter processing unit may be configured to requantize one, some, or all of the spatial parameters of the one or more sets of spatial parameters according to a second type quantizer having a lower resolution than the first type quantizer if it is determined that the number of bits of the current spatial metadata frame exceeds the maximum number of metadata bits. In this way, the number of bits of the current spatial metadata frame can be reduced with only a limited impact on the quality of the upmix signal and without significantly increasing the computational complexity of the audio encoding system.

[0014] The parameter processing unit may be configured to determine a set of temporal difference parameters based on a difference between a current set of spatial parameters and a previous set of spatial parameters. In particular, the temporal difference parameters may be determined by determining a difference between a parameter of the current set of spatial parameters and a corresponding parameter of the previous set of spatial parameters. The set of spatial parameters may include, for example, the parameters α1, α2, α3, β1, β2, β3, g, k1, and k2 described herein. Typically, only one of the parameters k1 and k2 may need to be transmitted. Both parameters are related by the relationship k1. 2 +k2 2 = 1. As an example, only parameter k1 may be transmitted and parameter k2 may be calculated at the receiving side. The temporal difference parameter may relate to the difference between corresponding ones of the above mentioned parameters.

[0015] The parameter processing unit may be configured to encode the set of temporal difference parameters using entropy encoding, for example using a Huffman code. The parameter processing unit may further be configured to insert the encoded set of temporal difference parameters into the current spatial metadata frame. The parameter processing unit may further be configured to reduce the entropy of the set of temporal difference parameters when it is determined that the number of bits of the current spatial metadata frame exceeds a maximum number of metadata bits. As a result, the number of bits required to entropy encode the temporal difference parameters may be reduced, thereby reducing the number of bits used for the current spatial metadata frame. For example, the parameter processing unit may be configured to set one, some, or all of the temporal difference parameters of the set of temporal difference parameters equal to a value with an increased (e.g., highest) probability of the possible values of the temporal difference parameters, in order to reduce the entropy of the set of temporal difference parameters. In particular, the probability may be increased compared to the probability of the temporal difference parameters prior to the setting operation. Typically, the most probable possible value of the temporal difference parameter corresponds to zero.

[0016] It should be noted that temporal differential encoding of said set of spatial parameters may typically not be used for independent frames. Thus, the parameter processing unit may be configured to verify whether the current spatial metadata frame is an independent frame and to apply temporal differential encoding only if the current spatial metadata frame is not an independent frame. On the other hand, frequency differential encoding, as described below, may also be used for independent frames.

[0017] The one or more control settings may include a frequency resolution setting, which indicates the number of different frequency bands for which each spatial parameter, referred to as a band parameter, is to be determined. The parameter processing unit may be configured to determine different corresponding spatial parameters (band parameters) for different frequency bands. In particular, different parameters α1, α2, α3, β1, β2, β3, g, k1, k2 may be determined for different frequency bands. Thus, the set of spatial parameters may include corresponding band parameters for the different frequency bands. For example, the set of spatial parameters may include T corresponding band parameters for T frequency bands, where T is an integer, e.g., T=7, 9, 12, or 15.

[0018] The parameter processing unit may be configured to determine a set of frequency difference parameters based on a difference between one or more band parameters in a first frequency band and one or more corresponding band parameters in a second, adjacent frequency band. The parameter processing unit may further be configured to encode the set of frequency difference parameters using entropy encoding, e.g., based on a Huffman code. The parameter processing unit may further be configured to insert the encoded set of frequency difference parameters into the current spatial metadata frame. The parameter processing unit may further be configured to reduce the entropy of the set of frequency difference parameters when it is determined that the number of bits of the current spatial metadata frame exceeds a maximum number of metadata bits. In particular, the parameter processing unit may be configured to set one, some, or all of the frequency difference parameters of the set of frequency difference parameters equal to a value (e.g., 0) that increases the probability of a possible value of the frequency difference parameter, in order to reduce the entropy of the set of frequency difference parameters. In particular, the probability may be increased compared to the probability of the frequency difference parameter before the setting operation.

[0019] Alternatively or additionally, the parameter processing unit may be configured to reduce the number of frequency bands when it is determined that the number of bits of the current spatial metadata frame exceeds the maximum number of metadata bits. Furthermore, the parameter processing unit may be configured to redetermine some or all of the one or more sets of spatial parameters for the current spatial metadata frame using a reduced number of frequency bands. Typically, a change in the number of frequency bands primarily affects high-frequency bands. As a result, band parameters for one or more frequencies may not be affected, and thus the parameter processing unit may not need to recalculate all band parameters.

[0020] As indicated above, the one or more external settings may include an update period indicating a time period during which a corresponding decoding system is required to synchronize to the bitstream. Furthermore, the one or more control settings may include an indication of whether a current spatial metadata frame should be encoded as an independent frame. The parameter processing unit may be configured to determine a sequence of spatial metadata frames for a corresponding sequence of frames of the multi-channel input signal. The configuration setting unit may be configured to determine, based on the update period, the one or more spatial metadata frames to be encoded as independent frames from the sequence of spatial metadata frames.

[0021] In particular, the one or more independent spatial metadata frames may be determined such that the update period is (on average) satisfied. For this purpose, the configuration unit may be configured to determine whether the current frame of the sequence of frames of the multi-channel input signal includes samples at a point in time that is an integer multiple of the update period (with respect to the start point of the multi-channel input signal). Further, the configuration unit may be configured to determine that the current spatial metadata frame corresponding to the current frame is an independent frame (since it includes samples at a point in time that is an integer multiple of the update period). The parameter processing unit may be configured to independently encode one or more sets of spatial parameters of the current spatial metadata frame from data included in previous (and / or future) spatial metadata frames if the current spatial metadata frame is to be encoded as an independent frame. Typically, if the current spatial metadata frame is to be encoded as an independent frame, all sets of spatial parameters of the current spatial metadata are independently encoded from data included in previous (and / or future) spatial metadata frames.

[0022] According to another aspect, a parameter processing unit is described that is configured to determine a spatial metadata frame for generating a frame of a multi-channel upmix signal from a corresponding frame of a downmix signal. The downmix signal may have m channels, the multi-channel upmix signal may have n channels, where n and m are integers and m < n. As outlined above, the spatial metadata frame may include one or more sets of spatial parameters.

[0023] The parameter processing unit may include a transform unit configured to determine a plurality of spectra from a current frame and a subsequent frame (referred to as a look-ahead frame) of a channel of the multi-channel input signal. The transform unit may utilize a filter bank, e.g., a QMF filter bank. A spectrum of the plurality of spectra may include a predetermined number of transform coefficients in a corresponding predetermined number of frequency bins. The plurality of spectra may be associated with a corresponding plurality of time bins (or time points). Thus, the transform unit may be configured to provide a time / frequency representation of the current frame and the look-ahead frame. By way of example, the current frame and the look-ahead frame may each have K samples. The transform unit may be configured to determine 2×K / Q spectra, each including Q transform coefficients.

[0024] The parameter processing unit may comprise a parameter determination unit configured to determine a spatial metadata frame for a current frame of a channel of the multi-channel input signal by weighting the plurality of spectra with a window function. The window function may be used to adjust the influence of a spectrum of the plurality of spectra on a particular spatial parameter or on a particular set of spatial parameters. For example, the window function may take a value between 0 and 1.

[0025] The window function may depend on one or more of: the number of sets of spatial parameters contained in a spatial metadata frame, the presence of one or more transient components in the current frame or a immediately following frame of the multi-channel input signal and / or the time points of said transient components. In other words, the window function may be adapted according to attributes of the current frame and / or the look-ahead frame. In particular, the window function used to determine the set of spatial parameters (referred to as a set-dependent window function) may depend on one or more attributes of the current frame and / or the look-ahead frame.

[0026] Thus, the window function may comprise a set-dependent window function. In particular, the window function for determining the spatial parameters of the spatial metadata frame may comprise (or consist of) one or more set-dependent window functions, respectively for the one or more sets of spatial parameters. The parameter determination unit may be configured to determine a set of spatial parameters for a current frame of the channel of the multi-channel input signal (i.e. for the current spatial metadata frame) by weighting the plurality of spectra with a set-dependent window function. As outlined above, the set-dependent window function may depend on one or more attributes of the current frame. In particular, the set-dependent window function may depend on whether the set of spatial parameters is associated with a transient component.

[0027] By way of example, if the set of spatial parameters is not associated with a transient component, the set-dependent window function may be configured to provide a phase-in of the plurality of spectra starting from a sampling point of a preceding set of spatial parameters to a sampling point of the present set of spatial parameters. The phase-in may be provided by a window function that transitions from 0 to 1. Alternatively or additionally, if the set of spatial parameters is not associated with a transient component, the set-dependent window function may include (or fully consider or leave unaffected) the plurality of spectra starting from the sampling point of the present set of spatial parameters to a sampling point of the subsequent set of spatial parameters if the subsequent set of spatial parameters is associated with a transient component. This may be achieved by a window function having a value of 1. Alternatively or additionally, if the set of spatial parameters is not associated with a transient component, the set-dependent window function may cancel (or eliminate or attenuate) the plurality of spectra starting from the sampling point of the subsequent set of spatial parameters if a subsequent set of spatial parameters is associated with a transient component. This may be achieved by a window function having a value of 0. Alternatively or additionally, if the set of spatial parameters is not associated with a transient component, the set-dependent window function may phase out the plurality of spectra starting from the sampling point of the set of spatial parameters to a spectrum of the plurality of spectra prior to the sampling point of the subsequent set of spatial parameters if the subsequent set of spatial parameters is not associated with a transient component. The phase-out may be provided by a window function transitioning from 1 to 0.

[0028] If the set of spatial parameters is associated with transient components, the set-dependent window function may cancel (or alternatively, exclude or attenuate) the spectra from the plurality of spectra before the sampling points of the set of spatial parameters. Alternatively or additionally, if the set of spatial parameters is associated with transient components and the sampling points of a subsequent set of spatial parameters are associated with transient components, the set-dependent window function may include (i.e., leave unaffected) the spectra from the plurality of spectra starting from the sampling points of the set of spatial parameters up to the spectra of the plurality of spectra before the sampling points of the subsequent set of spatial parameters, and may cancel (i.e., exclude or attenuate) the spectra from the plurality of spectra starting from the sampling points of the subsequent set of spatial parameters. Alternatively or additionally, if the set of spatial parameters is associated with transient components and the subsequent set of spatial parameters is not associated with transient components, the set-dependent window function may include (i.e., leave unaffected) the spectra from the plurality of spectra from the sampling points of the set of spatial parameters up to the end of the current frame, and may provide a fade-out of the spectra from the plurality of spectra from the start of the immediately subsequent frame up to the sampling points of the subsequent set of spatial parameters (i.e., gradually attenuate).

[0029] According to a further aspect, a parameter processing unit is described that is configured to determine a spatial metadata frame for generating a frame of a multi-channel upmix signal from a corresponding frame of a downmix signal. The downmix signal may have m channels, the multi-channel upmix signal may have n channels, where n and m are integers and m < n. As discussed above, the spatial metadata frame may include a set of spatial parameters.

[0030] As outlined above, the parameter processing unit may include a transform unit. The transform unit may be configured to determine a first plurality of transform coefficients from a frame of a first channel of the multi-channel input signal. Further, the transform unit may be configured to determine a second plurality of transform coefficients from a corresponding frame of a second channel of the multi-channel input signal. The first and second channels may be different. Thus, the first and second plurality of transform coefficients provide first and second time / frequency representations of the corresponding frames of the first and second channels, respectively. As outlined above, the first and second time / frequency representations may include a plurality of frequency bins and a plurality of time bins.

[0031] Furthermore, the parameter processing unit may include a parameter determination unit configured to determine a set of spatial parameters based on the first and second plurality of transform coefficients using fixed-point arithmetic. As indicated above, the set of spatial parameters typically includes corresponding band parameters for various frequency bands, where different frequency bands may include different numbers of frequency bins. A specific band parameter for a specific frequency band may be determined based on transform coefficients from the first and second plurality of transform coefficients of the specific frequency band (typically without considering transform coefficients of other frequency bands). The parameter determination unit may be configured to determine a shift used by the fixed-point arithmetic to determine the specific band parameter depending on the specific frequency band. In particular, the shift used by the fixed-point arithmetic to determine the specific band parameter for the specific frequency band may depend on the number of frequency bins included in the specific frequency band. Alternatively or additionally, the shift used by the fixed-point arithmetic to determine the specific band parameter for the specific frequency band may depend on the number of time bins to be considered to determine the specific band parameter.

[0032] The parameter determination unit may be configured to determine a shift for the particular frequency band so as to maximize accuracy of the particular band parameters, which may be achieved by determining the shift required for each multiply-accumulate operation of the particular band parameter determination process.

[0033] The parameter determination unit determines the specific band parameters for the specific frequency band p by calculating a first energy (or energy estimate) E based on transform coefficients that fall within the specific frequency band p from the first plurality of transform coefficients. 1,1 (p) based on the transform coefficients from the second plurality of transform coefficients that fall within the particular frequency band p. 2,2 (p) may be determined based on the transform coefficients from the first and second plurality of transform coefficients that fall within the particular frequency band p. 1,2 (p) may be determined. The parameter determination unit may determine the first energy estimate E 1,1 (p), the second energy estimate E 2,2 (p) and the covariance E 1,2 (p) based on the maximum of the absolute values of the shift z for the specific band parameter p p may be configured to determine

[0034] According to another aspect, an audio encoding system is described that is configured to generate a bitstream indicating a sequence of frames of a downmix signal and a corresponding sequence of frames of spatial metadata for generating a corresponding sequence of frames of a multi-channel upmix signal from the sequence of frames of the downmix signal. The system may have a downmix processing unit configured to generate the sequence of frames of the downmix signal from a corresponding sequence of frames of a multi-channel input signal. As indicated above, the downmix signal may have m channels, the multi-channel input signal may have n channels, where n and m are integers and m < n. Further, the audio encoding system may have a parameter processing unit configured to determine the sequence of frames of the spatial metadata from the sequence of frames of the multi-channel input signal.

[0035] The audio encoding system may further include a bitstream generation unit configured to generate the bitstream, the bitstream including a sequence of bitstream frames, where the bitstream frames represent frames of the downmix signal corresponding to first frames of the multi-channel input signal and spatial metadata frames corresponding to second frames of the multi-channel input signal. The second frames may be different from the first frames. In particular, the first frames may precede the second frames. In this way, the spatial metadata frame for a current frame may be transmitted together with the corresponding frame of a subsequent frame. This ensures that the spatial metadata frame arrives at the corresponding decoding system only when needed. The decoding system typically decodes the current frame of the downmix signal and generates a decorrelated frame based on the current frame of the downmix signal. This process introduces an algorithmic delay, delaying the spatial metadata frame for the current frame, ensuring that the spatial metadata frame only arrives at the decoding system once the decoded current frame and the decorrelated frame are available. As a result, the processing power and memory requirements of the decoding system can be reduced.

[0036] In other words, an audio encoding system configured to generate a bitstream based on a multi-channel input signal is described. As outlined above, the system may have a downmix processing unit configured to generate a sequence of frames of a downmix signal from corresponding sequences of first frames of the multi-channel input signal. The downmix signal may have m channels, the multi-channel input signal may have n channels, where n and m are integers and m < n. Further, the audio encoding system may have a parameter processing unit configured to determine a sequence of spatial metadata frames from a sequence of second frames of the multi-channel input signal. The sequence of frames of the downmix signal and the sequence of spatial metadata frames may be used by a corresponding decoding system to generate a multi-channel upmix signal including n channels.

[0037] The audio encoding system may further have a bitstream generation unit configured to generate the bitstream including a sequence of bitstream frames. Here, a bitstream frame indicates a frame of the downmix signal corresponding to a first frame of a sequence of first frames of the multi-channel input signal and a spatial metadata frame corresponding to a second frame of a sequence of second frames of the multi-channel input signal. The second frame may be different from the first frame. In other words, the frame configuration used to determine the spatial metadata frame and the frame configuration used to determine the frame of the downmix signal may be different. As outlined above, different frame configurations may be used to ensure that data is aligned in a corresponding decoding system.

[0038] The first frame and the second frame may typically include the same number of samples (e.g., 1536 samples). Some of the samples of the first frame may precede the samples of the second frame. In particular, the first frame may precede the second frame by a predetermined number of samples. The predetermined number of samples may, for example, correspond to a percentage of the number of samples of the frame. By way of example, the predetermined number of samples may correspond to 50% or more of the number of samples of the frame. In a specific example, the predetermined number of samples corresponds to 928 samples. As shown herein, this particular number of samples provides the minimum overall delay and optimal alignment for a particular implementation of an audio encoding and decoding system.

[0039] According to a further aspect, an audio encoding system configured to generate a bitstream based on a multi-channel input signal is described. The system may include a downmix processing unit configured to determine a sequence of clipping protection gains (also referred to herein as clip gains and / or DRC2 parameters) for a corresponding sequence of frames of the multi-channel input signal. The current clipping protection gain may indicate an attenuation to be applied to a current frame of the multi-channel input signal to prevent clipping of the corresponding current frame of the downmix signal. Similarly, the sequence of clipping protection gains may indicate respective attenuations to be applied to frames of the sequence of frames of the multi-channel input signal to prevent clipping of corresponding frames of the sequence of frames of the downmix signal.

[0040] The downmix processing unit may be configured to interpolate a current clipping protection gain and a previous clipping protection gain of a previous frame of the multi-channel input signal to provide a clipping protection gain curve. This may be performed in a similar manner for the sequence of clipping protection gains. Furthermore, the downmix processing unit may be configured to apply the clipping protection gain curve to a current frame of the multi-channel input signal to provide an attenuated current frame of the multi-channel input signal. Again, this may be performed in a similar manner for the sequence of frames of the multi-channel input signal. Furthermore, the downmix processing unit may be configured to generate a current frame of the sequence of frames of the downmix signal from the attenuated current frame of the multi-channel input signal. The sequence of frames of the downmix signal may be generated in a similar manner.

[0041] The audio processing system may further comprise a parameter processing unit configured to determine a sequence of spatial metadata frames from the multi-channel input signal. The sequence of frames of the downmix signal and the sequence of spatial metadata frames may be used to generate a multi-channel up-mix signal comprising n channels, the multi-channel up-mix signal being an approximation of the multi-channel input signal. The audio processing system may further comprise a bitstream generation unit configured to generate a bitstream indicating the sequence of clipping protection gains, the sequence of frames of the downmix signal and the sequence of spatial metadata frames, such that a corresponding decoding system can generate the multi-channel up-mix signal.

[0042] The clipping protection gain curve may include a transition segment that provides a smooth transition from a previous clipping protection gain to a current clipping protection gain and a flat segment that remains flat at the current clipping protection gain. The transition segment may extend through a predetermined number of samples of a current frame of the multi-channel input signal. The predetermined number of samples may be greater than one and less than the total number of samples of the current frame of the multi-channel input signal. In particular, the predetermined number of samples may correspond to a block of samples (where a frame may include multiple blocks) or to a frame. In a specific example, a frame may have 1536 samples and a block may have 256 samples.

[0043] According to a further aspect, an audio encoding system is described that is configured to generate a bitstream indicative of a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal. The system may include a downmix processing unit configured to generate the downmix signal from a multi-channel input signal. The system may further include a parameter processing unit configured to determine a sequence of frames of spatial metadata for a corresponding sequence of frames of the multi-channel input signal.

[0044] The audio encoding system may further include a configuration unit configured to determine one or more control settings for the parameter processing unit based on one or more external settings, which may include an update period indicating a time period during which a corresponding decoding system is required to synchronize to the bitstream. The configuration unit may be configured to determine, based on the update period, one or more independent frames of spatial metadata to be encoded independently from a sequence of frames of spatial metadata.

[0045] According to another aspect, a method for generating a bitstream indicating a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal is described. The method may include generating the downmix signal from a multi-channel input signal. Furthermore, the method may include determining one or more control settings based on one or more external settings. The one or more external settings include a target data rate for the bitstream, and the one or more control settings include a maximum data rate for spatial metadata. Furthermore, the method may include determining the spatial metadata from the multi-channel input signal in accordance with the control settings.

[0046] According to a further aspect, a method for determining a spatial metadata frame for generating a frame of a multi-channel up-mix signal from a corresponding frame of a down-mix signal is described. The method includes determining a plurality of spectra from a current frame and a subsequent frame of a channel of a multi-channel input signal. The method may further include weighting the plurality of spectra using a window function to obtain a plurality of weighted spectra. The method may further include determining the spatial metadata frame for the current frame of the channel of the multi-channel input signal based on the plurality of weighted spectra. The window function may depend on one or more of: the number of sets of spatial parameters included in the spatial metadata frame; the presence of transient components in the current frame or a subsequent frame of the multi-channel input signal; and / or the time points of the transient components.

[0047] According to a further aspect, a method for determining spatial metadata frames for generating frames of a multi-channel up-mix signal from corresponding frames of a down-mix signal is described. The method may include determining a first plurality of transform coefficients from frames of a first channel of the multi-channel input signal and determining a second plurality of transform coefficients from corresponding frames of a second channel of the multi-channel input signal. As outlined above, the first and second plurality of transform coefficients typically provide first and second time / frequency representations of corresponding frames of the first and second channels, respectively. The first and second time / frequency representations may include multiple frequency bins and multiple time bins. The set of spatial parameters may include corresponding band parameters for different frequency bands, each including a different number of frequency bins. The method may further include determining a shift to be applied when determining specific band parameters for a specific frequency band using fixed-point arithmetic. The shift may be determined based on the specific frequency band. Furthermore, the shift may be determined based on the number of time bins to be considered for determining the specific band parameters. Further, the method may include determining the specific band parameters based on the first and second plurality of transform coefficients falling within the specific frequency band using fixed-point arithmetic and a determined shift.

[0048] A method for generating a bitstream based on a multi-channel input signal is described. The method may include generating a sequence of frames of a downmix signal from a corresponding sequence of frames of a first multi-channel input signal. Furthermore, the method may include determining a sequence of spatial metadata frames from a second sequence of frames of the multi-channel input signal. The sequence of frames of the downmix signal and the sequence of spatial metadata frames may be for generating a multi-channel upmix signal. Furthermore, the method may include generating the bitstream including a sequence of bitstream frames. The bitstream frames may indicate frames of the downmix signal corresponding to first frames of the first sequence of frames of the multi-channel input signal and spatial metadata frames corresponding to second frames of the second sequence of frames of the multi-channel input signal. The second frames may be different from the first frames.

[0049] According to a further aspect, a method for generating a bitstream based on a multi-channel input signal is described. The method may include determining a sequence of clipping protection gains for a corresponding sequence of frames of the multi-channel input signal. The current clipping protection gain may indicate an attenuation to be applied to a current frame of the multi-channel input signal to prevent clipping of the corresponding current frame of the downmix signal. The method may proceed with interpolating the current clipping protection gain and a previous clipping protection gain of a previous frame of the multi-channel input signal to provide a clipping protection gain curve. The method may further include applying the clipping protection gain curve to the current frame of the multi-channel input signal to provide an attenuated current frame of the multi-channel input signal. A current frame of the sequence of frames of the downmix signal may be generated from the attenuated current frame of the multi-channel input signal. The method may further include determining a sequence of spatial metadata frames from the multi-channel input signal. The sequence of frames of the downmix signal and the sequence of spatial metadata frames may be used to generate a multi-channel upmix signal. The bitstream may be generated such that it indicates a sequence of clipping protection gains, a sequence of frames of a downmix signal and a sequence of spatial metadata frames to enable generation of the multi-channel upmix signal based on the bitstream.

[0050] According to a further aspect, a method for generating a bitstream indicating a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal is described. The method may include generating the downmix signal from a multi-channel input signal. Furthermore, the method may include determining one or more control settings based on one or more external settings. The one or more external settings may include an update period indicating a time period during which a corresponding decoding system is required to synchronize to the bitstream. The method may further include determining a sequence of frames of spatial metadata for a corresponding sequence of frames of the multi-channel input signal in accordance with the control settings. Furthermore, the method may include encoding one or more frames of spatial metadata from the sequence of frames of spatial metadata as independent frames in accordance with the update period.

[0051] According to a further aspect, a software program is described, which may be adapted for execution on a processor so as to perform the method steps outlined herein when executed on the processor.

[0052] According to another aspect, a storage medium is described, which may have a software program for execution on a processor adapted to perform the method steps outlined herein when executed on the processor.

[0053] According to a further aspect, a computer program product is described, which may include executable instructions for performing the method steps outlined herein when executed on a computer.

[0054] It should be noted that the methods and systems, including the preferred embodiments, outlined in this patent application may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this patent application may be combined in any manner. In particular, the features of the claims may be combined with each other in any manner. [Brief explanation of the drawings]

[0055] The invention is described below, by way of example, with reference to the accompanying drawings, in which: [Figure 1] FIG. 1 is a generalized block diagram of an exemplary audio processing system for performing spatial synthesis. [Figure 2] FIG. 2 illustrates exemplary details of the system of FIG. 1. [Figure 3] 1 and 2. FIG. 3 illustrates an exemplary audio processing system for performing spatial synthesis, similar to FIG. [Figure 4] FIG. 1 illustrates an exemplary audio processing system for performing spatial decomposition. [Figure 5a] FIG. 1 is a block diagram of an exemplary parametric multi-channel audio encoding system. [Figure 5b] FIG. 1 is a block diagram of an exemplary spatial decomposition and encoding system. [Figure 5c] 1 illustrates an exemplary time-frequency representation of a frame of a multi-channel audio signal. [Figure 5d] FIG. 1 illustrates an exemplary time-frequency representation of multiple channels of a multi-channel audio signal. [Figure 5e] FIG. 5b illustrates an exemplary windowing applied by the transform unit of the spatial decomposition and encoding system shown in Figure 5b. [Figure 6] 1 is a flow diagram of an exemplary method for reducing the data rate of spatial metadata. [Figure 7a]FIG. 1 illustrates an exemplary transition scheme for spatial metadata implemented in a decoding system. [Figure 7b] FIG. 10 illustrates an exemplary window function applied for determining spatial metadata. [Figure 7c] FIG. 10 illustrates an exemplary window function applied for determining spatial metadata. [Figure 7d] FIG. 10 illustrates an exemplary window function applied for determining spatial metadata. [Figure 8] FIG. 1 is a block diagram of an exemplary processing path of a parametric multi-channel codec system. [Figure 9a] FIG. 1 is a block diagram of an example parametric multi-channel audio encoding system configured to perform clipping protection and / or dynamic range control. [Figure 9b] FIG. 1 is a block diagram of an example parametric multi-channel audio encoding system configured to perform clipping protection and / or dynamic range control. [Figure 10] FIG. 1 illustrates an exemplary method for compensating DRC parameters. [Figure 11] FIG. 10 illustrates an exemplary interpolation curve for clipping protection. DETAILED DESCRIPTION OF THE INVENTION

[0056] As outlined in the introduction, this paper relates to a multi-channel audio coding system that utilizes parametric multi-channel representations. In the following, an exemplary multi-channel audio coding and decoding (codec) system is described. In the context of Figures 1 to 3, it is described how a decoder of the audio codec system uses a received parametric multi-channel representation to generate an n-channel upmix signal Y (typically n>2) from a received m-channel downmix signal X (e.g., m=2). Then, the encoder-related processing of the multi-channel audio codec system is described. In particular, it is described how the parametric multi-channel representation and the m-channel downmix signal can be generated from an n-channel input signal.

[0057] 1 shows a block diagram of an exemplary audio processing system 100 configured to generate an upmix signal Y from a downmix signal X and a set of mixing parameters. In particular, the audio processing system 100 is configured to generate an upmix signal based solely on the downmix signal X and the set of mixing parameters. From a bitstream P, an audio decoder 140 generates a downmix signal X=[l0r0] Tand extract a set of mixing parameters. In the illustrated example, the set of mixing parameters includes parameters α1, α2, α3, β1, β2, β3, g, k1, and k2. The mixing parameters may be included in quantized and / or entropy-coded form in respective mixing parameter data fields in the bitstream P. These mixing parameters may be referred to as metadata (or spatial metadata), which are transmitted together with the encoded downmix signal X. In some instances of the present disclosure, it is explicitly shown that some connecting lines are adapted to transmit multi-channel signals, where these lines are given crossing lines adjacent to the respective channel numbers. In the system 100 shown in FIG. 1, the downmix signal X includes m=2 channels, and the upmix signal Y defined below includes n=6 channels (e.g., 5.1 channels).

[0058] An upmix stage 110, whose operation depends parametrically on the mixing parameters, receives the downmix signal. A downmix modification processor 120 modifies the downmix signal by non-linear processing and by forming a linear combination of the downmix channels, thereby producing a modified downmix signal D=[d1d2]. T The first mixing matrix 130 receives the downmix signal X and the modified downmix signal D and obtains the upmix signal Y=[l f l s r f r s c lfe] T Output.

[0059]

number

[0060] The contributions from the modified downmix signal to the spatially left and right channels in the upmix signal may be controlled separately by parameters β1 (contribution of the first modified channel to the left channel) and β2 (contribution of the second modified channel to the right channel). Furthermore, the contribution from each channel in the downmix signal to its spatially corresponding channel in the upmix signal may be individually controllable by varying an independent mixing parameter g. Preferably, the gain parameter g is non-uniformly quantized to avoid large quantization errors.

[0061] Referring now still to FIG. 2, the downmix modification processor 120 may perform the following linear combination (which is a cross-mix) of the downmix channels in a second mixing matrix 121:

[0062]

number

[0063] FIG. 3 illustrates a first mixing matrix 130 of a type similar to that shown in FIG. 1, and its associated transform stages 301, 302 and inverse transform stages 311, 312, 313, 314, 315, and 316. These transform stages may include, for example, a filter bank, such as a Quadrature Mirror Filterbank (QMF). Thus, signals upstream of transform stages 301, 302 are time-domain representations, as are signals downstream of inverse transform stages 311, 312, 313, 314, 315, and 316. Other signals are frequency-domain representations. The time dependence of other signals may be represented, for example, as discrete values or blocks of values related to the time blocks into which the signals are segmented. Note that FIG. 3 uses an alternative notation compared to the matrix formulas above. For example, X L0 ~l0,X R0 ~r0, Y L ~l f , Y Ls ~l s Furthermore, the notation in Figure 3 allows us to express the time domain representation of a signal X L0 (t) is the frequency domain representation of the same signal as L0 (f) emphasizes the distinction between the frequency domain representation and the time domain representation, which is understood to be segmented into time frames and thus a function of both time and frequency variables.

[0064] FIG. 4 illustrates an audio processing system 400 for generating a downmix signal X and mixing parameters α1, α2, α3, β1, β2, β3, g, k1, and k2 that control the gains applied by the upmix stage 110. This audio processing system 400 is typically located on the encoder side, e.g., in a broadcasting or recording facility. Meanwhile, the system 100 of FIG. 1 is typically deployed on the decoder side, e.g., in a playback facility. The downmix stage 410 generates an m-channel signal X based on an n-channel signal Y. Preferably, the downmix stage 410 operates on a time-domain representation of these signals. A parameter extractor 420 may analyze the n-channel signal Y and generate values for the mixing parameters α1, α2, α3, β1, β2, β3, g, k1, and k2 by taking into account quantitative and qualitative attributes of the downmix stage 410. The mixing parameters may be a vector of values for frequency blocks, as suggested by the notation in FIG. 4, or may be further segmented into time blocks. In one exemplary implementation, the downmix stage 410 is time-invariant and / or frequency-invariant. Due to the time-invariant and / or frequency-invariant nature, there is typically no need for a communication connection between the downmix stage 410 and the parameter extractor 420; parameter extraction may proceed independently. This provides significant flexibility for implementation. It also offers the potential to reduce the overall latency of the system, since several processing stages may be performed in parallel. As an example, the Dolby Digital Plus format (or enhanced AC-3) may be used to encode the downmix signal X.

[0065] The parameter extractor 420 may have knowledge of quantitative and / or qualitative attributes of the downmix stage 410 by accessing a downmix specification. The downmix specification may specify one of a set of gain values, an index identifying a predefined downmix mode in which the gain is predefined, etc. The downmix specification may be a data record preloaded in memory in each of the downmix stage 410 and the parameter extractor 420. Alternatively or additionally, the downmix specification may be transmitted from the downmix stage 410 to the parameter extractor 420 over a communication line connecting these units. As a further alternative, each of the downmix stages 410 to the parameter extractor 420 may access the downmix specification from a common data source, such as a memory within the audio processing system (e.g., in the configuration unit 520 shown in FIG. 5a), or in a metadata stream associated with the input signal Y.

[0066] Figure 5a shows an exemplary multi-channel encoding system 500 that encodes a multi-channel audio input signal Y (including n channels) using a downmix signal X (including m channels, m < n) and a parametric representation. The system 500 has a downmix encoding unit 510 that has, for example, the downmix stage 410 of FIG. 4. The downmix encoding unit 510 may be configured to provide an encoded version of the downmix signal X. The downmix encoding unit 510 may utilize, for example, a Dolby Digital Plus encoder to encode the downmix signal X. Further, the system 500 has a parameter encoding unit 520 that may have the parameter extractor 420 of FIG. 4. The parameter encoding unit 520 may be configured to quantize and encode a set of mixing parameters α1, α2, α3, β1, β2, β3, g, k1 (also referred to as spatial parameters) to provide an encoded spatial parameter 562. As shown above, the parameter k2 may be determined from the parameter k1. Further, the system 500 may have a bitstream generation unit 530 that is configured to generate a bitstream P 564 from the encoded downmix signal 563 and from the encoded spatial parameter 562. The bitstream 564 may be encoded according to a predetermined bitstream syntax. In particular, the bitstream 564 may be encoded in a format compliant with Dolby Digital Plus (DD+ or E-AC-3, Enhanced AC-3).

[0067] The system 500 may comprise a configuration setting unit 540 configured to determine one or more control settings 552, 554 for the parameter coding unit 520 and / or the downmix coding unit 510. The one or more control settings 552, 554 may be determined based on one or more external settings 551 of the system 500. By way of example, the one or more external settings may include an overall (maximum or fixed) data rate of the bitstream 564. The configuration setting unit 540 may be configured to determine the one or more control settings 552 depending on the one or more external settings 551. The one or more control settings 552 for the parameter coding unit 520 may include one or more of the following:

[0068] · Maximum data rate for encoded spatial metadata 562. This control setting is referred to herein as the metadata data rate setting. The maximum and / or specific number of parameter sets to be determined by the parameter coding unit 520 per frame of the audio signal 561. This control setting is referred to herein as the temporal resolution setting, as it allows to influence the temporal resolution of the spatial parameters. The number of frequency bands for which the spatial parameters should be determined by the parameter coding unit 520. This control setting is called the frequency resolution setting, as it allows to influence the frequency resolution of the spatial parameters. The resolution of the quantizer that should be used to quantize the spatial parameters. This control setting is referred to as the quantizer setting in this paper.

[0069] The parameter coding unit 520 may use one or more of the control settings 552 described above to determine and / or encode the spatial parameters to be included in the bitstream 564. Typically, the input audio signal Y 561 is segmented into a sequence of frames, where each frame contains a predetermined number of samples of the input audio signal Y 561. The metadata data rate setting may indicate the maximum number of bits available for encoding the spatial parameters of a frame of the input audio signal 561. The actual number of bits used for encoding the spatial parameters 562 of a frame may be less than the number of bits allocated by the metadata data rate setting. The parameter coding unit 520 may be configured to inform the configuration setting unit 540 about the actual number of bits used 553, thereby enabling the configuration setting unit 540 to determine the number of bits available for encoding the downmix signal X. This number of bits may be communicated to the downmix encoding unit 510 as a control setting 554. The downmix encoding unit 510 may be configured to encode the downmix signal X (e.g., using a multi-channel encoder such as Dolby Digital Plus) based on the control settings 554. Thus, bits not used for encoding the spatial parameters may be used for encoding the downmix signal.

[0070] FIG. 5b shows a block diagram of an exemplary parameter coding unit 520. The parameter coding unit 520 may include a transform unit 521 configured to determine a frequency representation of the input signal 561. In particular, the transform unit 521 may be configured to transform a frame of the input signal 561 into one or more spectra, each spectrum including a plurality of frequency bins. By way of example, the transform unit 521 may be configured to apply a filter bank, e.g., a QMF filter bank, to the input signal 561. The filter bank may be a critically sampled filter bank. The filter bank may have a predetermined number Q of filters (e.g., Q=64 filters). Thus, the transform unit 521 may be configured to determine Q subband signals from the input signal 561, where each subband signal is associated with a corresponding frequency bin 571. By way of example, a frame of K samples of the input signal 561 may be transformed into Q subband signals with K / Q frequency coefficients per subband signal. In other words, a frame of K samples of the input signal 561 may be transformed into K / Q spectra. Here, each spectrum has Q frequency bins. In one particular example, the frame length is K=1536, the number of frequency bins is Q=64, and the number of spectra is K / Q=24.

[0071] The parameter coding unit 520 may include a banding unit 522 configured to group one or more frequency bins 571 into frequency bands 572. The grouping of the frequency bins 571 into frequency bands 572 may depend on the frequency resolution setting 552. Table 1 shows an example mapping of frequency bins 571 to frequency bands 572, where the mapping may be applied by the banding unit 522 based on the frequency resolution setting 552. In the illustrated example, the frequency resolution setting 552 may indicate banding of the frequency bins 571 into 7, 9, 12, or 15 frequency bands. The banding typically models the psychoacoustic behavior of the human ear. As a result, the number of frequency bins 571 per frequency band 572 typically increases with increasing frequency.

[0072] [Table 1] The parameter determination unit 523 (particularly the parameter extractor 420) of the parameter coding unit 520 may be configured to determine one or more sets of mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for each of the frequency bands 572. For this reason, the frequency bands 572 may also be referred to as parameter bands. The mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for the frequency bands 572 may also be referred to as band parameters. Thus, a complete set of mixing parameters typically includes a band parameter for each frequency band 572. The band parameters may be applied to determine subband versions of the decoded upmix signal in the mixing matrix 130 of FIG. 3 .

[0073] The number of sets of mixing parameters per frame to be determined by parameter determination unit 523 may be indicated by temporal resolution setting 552. As an example, temporal resolution setting 552 may indicate that one or more sets of mixing parameters are to be determined for each frame.

[0074] Determining a set of blending parameters including band parameters for multiple frequency bands 572 is shown in Figure 5c. Figure 5c shows an example set of transform coefficients 580 derived from a frame of input signal 561. The transform coefficients 580 correspond to specific time points 582 and specific frequency bins 571. A frequency band 572 may include multiple transform coefficients 580 from one or more frequency bins 571. As can be seen from Figure 5c, transforming the time-domain samples of input signal 561 provides a time-frequency representation of a frame of input signal 561.

[0075] It should be noted that the set of blending parameters for the current frame may be determined based on the transform coefficients 580 of the current frame and also based on the transform coefficients 580 of the immediately following frame (also referred to as the look-ahead frame).

[0076] The parameter determination unit 523 may be configured to determine mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for each frequency band 572. When the temporal resolution setting is set to 1, all transform coefficients 580 (of the current frame and the look-ahead frame) of a particular frequency band 572 may be considered to determine mixing parameters for that particular frequency band 572. On the other hand, the parameter determination unit 523 may be configured to determine two sets of mixing parameters per frequency band 572 (e.g., when the temporal resolution setting is set to 2). In this case, the first half in time of the transform coefficients 580 of that particular frequency band 572 (e.g., corresponding to the transform coefficients 580 of the current frame) may be used to determine the first set of mixing parameters, and the second half in time of the transform coefficients 580 of that particular frequency band 572 (e.g., corresponding to the transform coefficients 580 of the look-ahead frame) may be used to determine the second set of mixing parameters.

[0077] In general terms, the parameter determination unit 523 may be configured to determine one or more sets of mixing parameters based on the transform coefficients 580 of the current frame and the look-ahead frame. A window function may be used to define the influence of the transform coefficients 580 on the one or more sets of mixing parameters. The shape of the window function may depend on the number of sets of mixing parameters per frequency band 572 and / or attributes of the current frame and / or the look-ahead frame (e.g., the presence of one or more transient components). Exemplary window functions are described in the context of Figure 5e and Figures 7b-7d.

[0078] It should be noted that the above may apply when a frame of the input signal 561 does not contain a transient signal portion. The system 500 (e.g., the parameter determination unit 523) may be configured to perform transient detection based on the input signal 561. If one or more transient components are detected, one or more transient indicators 583, 584 may be set, where the transient indicators 583, 584 may identify the time points 582 of the corresponding transient components. The transient indicators 583, 584 may be referred to as sampling points of the respective sets of mixing parameters. In the case of a transient component, the parameter determination unit 523 may be configured to determine the set of mixing parameters based on the transform coefficients 580 starting from the time point of the transient component (this is indicated by the different shaded areas in FIG. 5c). On the other hand, the transform coefficients 580 before the time point of the transient component are ignored, thereby ensuring that the set of mixing parameters reflects the multi-channel situation after the transient component.

[0079] Figure 5c shows transform coefficients 580 for a channel of a multi-channel input signal Y 561. The parameter coding unit 520 is typically configured to determine transform coefficients 580 for multiple channels of the multi-channel input signal 561. Figure 5d shows exemplary transform coefficients for a first 561-1 and a second 561-2 channel of the input signal 561. A frequency band p 572 includes frequency bins 571 ranging from frequency index i to j. The transform coefficients 580 for the first channel 561-1 in frequency bin i at time instant (or spectrum) q are a q,i In a similar manner, the transform coefficients 580 of the second channel 561-2 in frequency bin i at time (or spectrum) q may be referred to as b q,i The transform coefficients 580 may be complex numbers. Determining the mixing parameters for the frequency band p may involve determining the energy and / or covariance of the first and second channels 561-1, 561-2 based on the transform coefficients 580. As an example, the covariance of the transform coefficients 580 of the first and second channels 561-1, 561-2 for the time interval [q,v] in the frequency band p may be

number

number

[0080] Thus, the parameter determination unit 523 may be configured to determine one or more sets 573 of band parameters for different frequency bands 572. The number of frequency bands 572 typically depends on the frequency resolution setting 552, and the number of sets of mixing parameters per frame typically depends on the temporal resolution setting 552. As an example, the frequency resolution setting 552 may indicate the use of 15 frequency bands 572, and the temporal resolution setting 552 may indicate the use of two sets of mixing parameters. In this case, the parameter determination unit 523 may be configured to determine two temporally distinct sets of mixing parameters, where each set of mixing parameters includes 15 sets 573 of band parameters (i.e., mixing parameters for different frequency bands 572).

[0081] As indicated above, the mixing parameters for the current frame may be determined based on the transform coefficients 580 of the current frame and based on the transform coefficients 580 of the subsequent look-ahead frame. The parameter determination unit 523 may apply a window to the transform coefficients 580 to ensure a smooth transition between the mixing parameters of successive frames in the sequence of frames and / or to take into account sudden portions (e.g., transient components) in the input signal 561. This is illustrated in FIG. 5e, which shows K / Q spectra 589 of a current frame 585 and an immediately following frame 590 of the input audio signal 561 at corresponding K / Q consecutive points in time 582. Furthermore, FIG. 5e shows an exemplary window 586 used by the parameter determination unit 523. The window 586 reflects the influence of the K / Q spectra 589 of the current frame 585 and the immediately following frame 590 (referred to as look-ahead frames) on the mixing parameters. As will be outlined in more detail below, window 586 reflects the case where current frame 585 and look-ahead frame 590 do not contain any transient components. In this case, window 586 ensures smooth phase-in and phase-out of spectra 589 of current frame 585 and look-ahead frame 590, respectively, thereby allowing for smooth evolution of spatial parameters. Furthermore, FIG. 5e shows exemplary windows 587 and 588. Dashed window 587 reflects the influence of K / Q spectra 589 of current frame 585 on the blending parameters of the immediately preceding frame. Dashed window 588 reflects the influence of K / Q spectra 589 of immediately succeeding frame 590 on the blending parameters of immediately succeeding frame 590 (in the case of smooth interpolation).

[0082] The one or more sets of mixing parameters may then be quantized and encoded using an encoding unit 524 of the parameter coding unit 520. The encoding unit 524 may apply various encoding schemes. By way of example, the encoding unit 524 may be configured to perform differential encoding of the mixing parameters. The differential encoding may be based on a temporal difference (between a current mixing parameter and a preceding corresponding mixing parameter for the same frequency band 572) or a frequency difference (between a current mixing parameter of a first frequency band 572 and a corresponding current mixing parameter of an adjacent second frequency band 572).

[0083] Further, the encoding unit 524 may be configured to quantize the set of mixing parameters and / or the temporal or frequency differences of the mixing parameters. The quantization of the mixing parameters may depend on a quantizer setting 552. For example, the quantizer setting 552 may have two values, a first value indicating fine quantization and a second value indicating coarse quantization. Thus, the encoding unit 524 may be configured to perform fine quantization (with a relatively low quantization error) or coarse quantization (with a relatively increased quantization error) based on the quantization type indicated by the quantizer setting 552. The quantized parameters or parameter differences may then be encoded using an entropy-based code, such as a Huffman code. The result is encoded spatial parameters 562. The number of bits 553 used for the encoded spatial parameters 562 may be communicated to the configuration unit 540.

[0084] In one embodiment, the encoding unit 524 may be configured to first quantize the various mixing parameters (taking into account the quantizer settings 552) to provide quantized mixing parameters. The quantized mixing parameters may then be entropy coded (e.g., using a Huffman code). The entropy coding may encode the quantized mixing parameters of a frame (without taking into account previous frames), frequency differences of the quantized mixing parameters, or temporal differences of the quantized mixing parameters. Temporal difference encoding may not be used in the case of so-called independent frames, which are encoded independently of previous frames.

[0085] Thus, the parameter encoding unit 520 may utilize a combination of differential encoding and Huffman coding for determining the encoded spatial parameters 562. As outlined above, the encoded spatial parameters 562 may be included in the bitstream 564 as metadata (also referred to as spatial metadata) together with the encoded downmix signal 563. To reduce redundancy and thus increase the spare bitrate available for encoding the downmix signal 563, differential encoding and Huffman coding may be used for transmitting the spatial metadata. Because Huffman codes are variable-length codes, the size of the spatial metadata can vary significantly depending on the statistics of the encoded spatial parameters 562 to be transmitted. The data rate required for transmitting the spatial metadata deducts from the data rate available to the core codec (e.g., Dolby Digital Plus) for encoding the stereo downmix signal. To avoid compromising the audio quality of the downmix signal, the number of bytes that may be spent on transmitting spatial metadata per frame is typically limited. This limit may be subject to encoder tuning considerations, which may be taken into account by configuration unit 540. However, due to the variable-length nature of the underlying differential / Huffman coding of spatial parameters, it is typically not possible to guarantee without further measures that the data rate ceiling (e.g., as reflected in metadata data rate setting 552) will not be exceeded.

[0086] This document describes a method for post-processing of encoded spatial parameters 562 and / or spatial metadata including the encoded spatial parameters 562. A method 600 for post-processing spatial metadata is described in the context of FIG. 6. The method 600 may be applied when it is determined that the total size of one frame of spatial metadata exceeds a predefined limit, e.g., dictated by a metadata data rate setting 552. The method 600 is directed to gradually reducing the amount of metadata. Reducing the size of the spatial metadata typically also reduces the accuracy of the spatial metadata, thereby impairing the quality of the spatial image of the reproduced audio signal. However, the method 600 typically ensures that the total amount of spatial metadata does not exceed the predefined limit, thus allowing for determining an improved trade-off between spatial metadata (for regenerating an m-channel multi-channel signal) and audio codec metadata (for decoding an encoded downmix signal 563) in terms of overall audio quality. Furthermore, the method 600 for post-processing of spatial metadata can be implemented with relatively low computational complexity (compared to a complete recalculation of the encoded spatial parameters with modified control settings 552).

[0087] The method 600 for spatial metadata post-processing includes one or more of the following steps. As outlined above, a spatial metadata frame may include multiple (e.g., one or two) parameter sets per frame, and the use of additional parameter sets allows for increased temporal resolution of the mixing parameters. The use of multiple parameter sets per frame can improve audio quality, especially for attack-rich (i.e., transient) signals. Even for audio signals with a fairly slowly changing spatial image, spatial parameter updates using a grid with twice the density of sampling points can improve audio quality. However, transmitting multiple parameter sets per frame leads to approximately a two-fold increase in the data rate. Thus, if it is determined (step 601) that the data rate for spatial metadata exceeds the metadata data rate setting 552, it may be checked whether the spatial metadata frame includes more than one set of mixing parameters. In particular, it may be checked (step 602) whether the metadata frame includes two sets of mixing parameters that are supposed to be transmitted. If it is determined that the spatial metadata includes multiple sets of mixing parameters, one or more of the sets in excess of a single set of mixing parameters may be discarded (step 603). As a result, the data rate for the spatial metadata can be significantly reduced (typically by a factor of two, in the case of two sets of mixing parameters), while still impairing audio quality to a relatively low extent.

[0088] The decision as to which of the two (or more) sets of mixing parameters to discard may depend on whether the encoding system 500 detects a transient location ("attack") in the portion of the input signal 561 covered by the current frame. If multiple transients are present in the current frame, earlier transients are more important than later transients due to the psychoacoustic post-masking effect of any single attack. Thus, if a transient is present, it may be advisable to discard a later set of mixing parameters (e.g., the second of the two). On the other hand, if there is no attack, an earlier set of mixing parameters (e.g., the first of the two) may be discarded. This may be due to the windowing (shown in FIG. 5e) used when calculating the spatial parameters. The window 586 used to window out the portion of the input signal 561 used to calculate the spatial parameters for the second set of mixing parameters typically has the greatest impact at the time when the upmix stage 130 places the sampling points for parameter reconstruction (i.e., at the end of the current frame). On the other hand, the first set of mixing parameters typically has a half-frame offset relative to this point in time. As a result, the error introduced by dropping the first set of mixing parameters is very likely to be lower than the error introduced by dropping the second set of mixing parameters. This is shown in Figure 5e, where it can be seen that the second half of spectrum 589 of current frame 585, which is used to determine the second set of mixing parameters, is influenced to a greater extent by samples of current frame 585 than the first half of spectrum 589 of current frame 585 (window function 586 has a lower value for the first half of spectrum 589 than for the second half).

[0089] The spatial cues (i.e., mixing parameters) calculated in the encoding system 500 are transmitted to the corresponding decoder 100 via a bitstream 562 (which may be part of a bitstream 564 carrying an encoded stereo downmix signal 563). Between the calculation of the spatial cues and their representation in the bitstream 562, the encoding unit 524 typically applies a two-stage encoding approach: a first stage, quantization, is a lossy stage because it adds error to the spatial cues; and a second stage, differential / Huffman coding, is a lossless stage. As outlined above, the encoder 500 can choose between different types of quantization (e.g., two types of quantization): a high-resolution quantization scheme that adds a relatively small error but provides a larger number of potential quantization indexes, and a low-resolution quantization scheme that adds a relatively large error but provides a smaller number of quantization indexes and therefore does not require as large a Huffman codeword. It should be noted that different types of quantization may be applicable to some or all of the mixing parameters. For example, different types of quantization may be applicable to the mixing parameters α1, α2, α3, β1, β2, β3, k1, while the gain g may be quantized with a fixed type of quantization.

[0090] The method 600 may include a step 604 of verifying which type of quantization was used to quantize the spatial parameters. If it is determined that a relatively fine quantization resolution was used, the encoding unit 524 may be configured to reduce 605 the quantization resolution to a lower type of quantization. As a result, the spatial parameters are quantized again. However, this does not add significant computational overhead (compared to redetermining the spatial parameters using different control settings 552). It should be noted that different types of quantization may be used for different spatial parameters α1, α2, α3, β1, β2, β3, g, and k1. Thus, the encoding unit 524 may be configured to select a quantization resolution individually for each type of spatial parameter and thereby adjust the data rate of the spatial metadata.

[0091] The method 600 may also include a step (not shown in FIG. 6 ) of reducing the frequency resolution of the spatial parameters. As outlined above, the set of mixing parameters for a frame is typically clustered into frequency bands or parameter bands 572. Each parameter band represents a frequency range, and for each band, a separate set of spatial cues is determined. Depending on the data rate available for transmitting the spatial metadata, the number of parameter bands 572 may be varied in stages (e.g., 7, 9, 12, or 15 bands). The number of parameter bands 572 has an approximately linear relationship to the data rate, so reducing the frequency resolution can significantly reduce the data rate of the spatial metadata, while only moderately affecting the audio quality. However, such a reduction in frequency resolution typically requires recalculating the set of mixing parameters using the changed frequency resolution, thereby increasing the amount of computation.

[0092] As outlined above, the encoding unit 524 may utilize differential encoding of (quantized) spatial parameters. The configuration unit 551 may be configured to impose direct encoding of spatial parameters of frames of the input audio signal 561 to ensure that transmission errors do not propagate over an unlimited number of frames and to allow the decoder to synchronize to the received bitstream 562 at intermediate points in time. Thus, a certain percentage of frames may not utilize differential encoding along the timeline. Such frames that do not utilize differential encoding may be referred to as independent frames. The method 600 may include a step 606 of verifying whether the current frame is an independent frame and / or whether the independent frame is a forced independent frame. The encoding of the spatial parameters may depend on the result of step 606.

[0093] As outlined above, differential encoding is typically designed to calculate differences between successive temporal or neighboring frequency bands of quantized spatial cues. In either case, the statistics of the spatial cues are such that small differences appear more frequently than large differences and are therefore represented by shorter Huffman codewords. This paper proposes performing smoothing (over time or over frequency) of the quantized spatial parameters. Smoothing spatial parameters over time or over frequency typically results in smaller differences and thus reduces the data rate. Due to psychoacoustic considerations, temporal smoothing is usually preferred over smoothing in the frequency direction. If it is determined that the current frame is not a constrained independent frame, method 600 may proceed to perform temporal differential encoding (step 607), possibly in combination with temporal smoothing. On the other hand, if it is determined that the current frame is an independent frame, method 600 may proceed to perform frequency differential encoding (step 608) and possibly smoothing along frequency.

[0094] The differential encoding in step 607 may be subjected to a smoothing process over time to reduce the data rate. The degree of smoothing may vary depending on the amount by which the data rate is to be reduced. The most severe kind of temporal "smoothing" corresponds to keeping the previous set of mixing parameters unchanged, which corresponds to transmitting only delta values equal to 0. Temporal smoothing of the differential encoding may be performed on one or more (e.g., all) of the spatial parameters.

[0095] Similar to temporal smoothing, frequency smoothing may also be performed. In its most extreme form, frequency smoothing corresponds to transmitting the same quantized spatial parameters for the complete frequency range of the input signal 561. While ensuring that the limits set by the metadata data rate setting are not exceeded, frequency smoothing may have a relatively large impact on the quality of the spatial image that can be reproduced using the spatial metadata. Therefore, it may be preferable to apply frequency smoothing only when temporal smoothing is not allowed (for example, when the current frame is a forced independent frame where temporal differential encoding with respect to the previous frame must not be used).

[0096] As outlined above, the system 500 may be operated according to one or more external settings, such as an overall target data rate of the bitstream 564 or a sampling rate of the input audio signal 561. Typically, there is no single optimal operating point for all combinations of external settings. The configuration unit 540 may be configured to map valid combinations of external settings 551 to combinations of control settings 552, 554. By way of example, the configuration unit 540 may rely on the results of psychoacoustic listening tests. In particular, the configuration unit 540 may be configured to determine a combination of control settings 552, 554 that ensures (on average) an optimal psychoacoustic encoding result for a particular combination of external settings 551.

[0097] As outlined above, the decoding system 100 must be able to synchronize to the received bitstream 564 within a given time period. To ensure this, the encoding system 500 may periodically encode so-called independent frames, i.e., frames that do not rely on knowledge of previous frames. The average distance in frames between two independent frames may be given by the ratio of a given maximum time delay for synchronization to the duration of one frame. This ratio does not necessarily have to be an integer; the distance between two independent frames is always an integer number of frames.

[0098] The encoding system 500 (e.g., the configuration unit 540) may be configured to receive a maximum time delay for synchronization or a desired update time period as an external setting 551. Additionally, the encoding system 500 (e.g., the configuration unit 540) may include a timer module configured to track the absolute amount of time that has elapsed since the first encoded frame of the bitstream 564. The first encoded frame of the bitstream 564 is, by definition, an independent frame. The encoding system 500 (e.g., the configuration unit 540) may be configured to determine whether the next frame to be encoded has samples corresponding to times that are integer multiples of the desired update period. Whenever the next frame to be encoded has samples at times that are integer multiples of the desired update period, the encoding system 500 (e.g., the configuration unit 540) may be configured to ensure that the next frame to be encoded is encoded as an independent frame. This ensures that the desired update time period is maintained even if the ratio between the desired update time period and the frame length is not an integer.

[0099] As outlined above, the parameter determination unit 523 is configured to calculate spatial cues based on a time / frequency representation of the multi-channel input signal 561. A frame of spatial metadata may be determined based on K / Q (e.g., 24) spectra 589 (QMF spectra) of the current frame and / or based on K / Q (e.g., 24) spectra 589 (QMF spectra) of a previous frame, where each spectrum 589 may have a frequency resolution of Q (e.g., 64) frequency bins 571. Depending on whether the encoding system 500 detects transient components in the input signal 561, the time length of the signal portion used to calculate a single set of spatial cues may have a different number of spectra 589 (e.g., from 1 spectrum to 2 × K / Q spectra). As shown in FIG. 5c, each spectrum 589 is divided into a number of frequency bands 572 (e.g., 7, 9, 12, or 15 frequency bands). These frequency bands contain different numbers of frequency bins 571 (e.g., from one frequency bin to 41 frequencies) due to psychoacoustic considerations. Different frequency bands p 572 and different temporal segments [q,v] define a grid on the time / frequency representation of the current and future frames of the input signal 561. For different squares in this grid, different sets of spatial cues may be calculated based on energy and / or covariance estimates of at least some of the input channels within each square. As outlined above, the energy estimates and / or covariances may be calculated by summing the squares of the transform coefficients 580 of one channel and / or by summing the products of the transform coefficients 580 of different channels (as indicated by the formulas given above). The different transform coefficients 580 may be weighted according to a window function 586 used to determine the spatial parameters.

[0100] Energy estimate E 1,1 (p), E 2,2 (p) and / or covariance E 1,2The calculation of (p) may be performed in fixed-point arithmetic. In this case, different sizes of the squares of the time / frequency grid may have an impact on the arithmetic precision of the values determined for the spatial parameters. As outlined above, the number (j-i+1) of frequency bins 571 per frequency band 572 and / or the length of the time interval [q,v] of the squares of the time / frequency grid may vary significantly (e.g., between 1x1x2 and 48x41x2 transform coefficients 580 (e.g., real and imaginary parts of complex QMF coefficients)). As a result, the energy E 1,1 (p) / covariance E 1,2 The products Re{a that need to be summed to determine (p) t,f}Re{b t,f} and Im{a t,f}Im{b t,f} can vary significantly. To prevent the result of the above calculation from exceeding the range of numbers that can be represented in fixed-point arithmetic, the signal is scaled by the maximum number of bits (e.g., 2 6 2 6 41 2). However, this approach leads to a significant loss of arithmetic precision for smaller mas and / or for mas that have only relatively low signal energy.

[0101] This paper proposes using individual scaling for each square of the time / frequency lattice. The individual scaling may depend on the number of transform coefficients 580 included in the square of the time / frequency lattice. Typically, the spatial parameters for a particular square of the time / frequency lattice (i.e., for a particular frequency band 572 and a particular time interval [q,v]) are determined solely based on the transform coefficients 580 from that particular square (and not on the transform coefficients 580 from other squares). Furthermore, the spatial parameters are typically determined solely based on the energy estimates and / or covariance ratios (and are typically not affected by absolute energy estimates and / or covariances). In other words, a single spatial cue typically uses only the energy estimates and / or cross-channel products from a single time / frequency square. Furthermore, the spatial cue is typically not affected by absolute energy estimates / covariances, but only by the energy estimate / covariance ratios. Therefore, it is possible to use individual scaling for every single square. This scaling should be consistent for channels that contribute to a particular spatial cue.

[0102] Energy estimates E of the first and second channels 561-1, 561-2 for frequency band p 572 and time interval [q,v] 1,1 (p), E 2,2 (p) and the covariance E between the first and second channels 561-1, 561-2 1,2 (p) may be determined, for example, as shown by the formula above. The energy estimates and covariances are scaled by a scaling factor s p Scaled by the scaled energy and covariance s p E 1,1 (p), s p E 2,2 (p) and s p E 1,2 (p) may be given. The energy estimate E 1,1 (p), E 2,2 (p) and covariance E 1,2The spatial parameter P(p) derived based on (p) typically depends on the ratio of energies and / or covariances, and therefore the value of the spatial parameter P(p) depends on the scaling factor s p As a result, different scaling factors s for different frequency bands p, p+1, and p+2 are used. p , s p+1 , s p+2 may be used.

[0103] It should be noted that one or more of the spatial parameters may depend on more than two different input channels (e.g., three different channels), in which case the one or more spatial parameters are determined by the energy estimates E 1,1 (p), E 2,2 (p) Based on the covariances between different pairs of channels, i.e., E 1,2 (p), E 1,3 (p), E 2,3 (p), etc., where the values of the one or more spatial parameters are independent of scaling factors applied to the energy estimates and / or covariances.

[0104] In particular, z p is a positive integer that indicates the shift in fixed-point arithmetic, then for a particular frequency band p, the scaling factor s p =2 -zp but 0.5 p max{|E 1,1 (p)|,|E 2,2 (p)|,|E 1,2 (p)|}≦1.0 and shift z p may be determined such that is minimized. By ensuring this individually for each frequency band p and / or each time interval [q,v] for which the mixing parameters are determined, increased (e.g., maximum) precision in fixed-point arithmetic may be achieved while still guaranteeing a valid range of values.

[0105] ​As an example, individual scaling can be implemented by checking whether for every single MAC (multiply-accumulate) operation the result of the MAC operation can exceed ±1. If and only if so, the individual scaling for that square may be increased by one bit. Once this is done for all channels, the maximum scaling for each square may be determined and all deviating scalings for the squares may be adapted accordingly.

[0106] As outlined above, the spatial metadata may include one or more (e.g., two) sets of spatial parameters per frame. Accordingly, the encoding system 500 may transmit one or more sets of spatial parameters per frame to the corresponding decoding system 100. Each of these sets of spatial parameters corresponds to a specific spectrum among the K / Q temporally consecutive spectra 289 of the frame of spatial metadata. This specific spectrum corresponds to a specific time point, which may be referred to as a sampling point. FIG. 5c shows two exemplary sampling points 583, 584 for each of the two sets of spatial parameters. The sampling points 583, 584 may be associated with specific events contained within the input audio signal 561. Alternatively, the sampling points may be predetermined.

[0107] The sampling points 583 and 584 indicate the time points at which the corresponding spatial parameters should be fully applied in the decoding system 100. In other words, the decoding system 100 may be configured to update the spatial parameters according to the transmitted sets of spatial parameters at the sampling points 583 and 584. Furthermore, the decoding system 100 may be configured to interpolate the spatial parameters between two successive sampling points. The spatial parameters may indicate the type of transition performed between successive sets of spatial parameters. Examples of the type of transition are a "smooth" transition and an "abrupt" transition between the spatial parameters, which respectively mean that the spatial parameters may be interpolated in a smooth (e.g., linear) manner or updated abruptly.

[0108] In the case of a "smooth" transition, the sampling points may be fixed (i.e., predetermined) and therefore do not need to be signaled in the bitstream 564. If a frame of spatial metadata conveys a single set of spatial parameters, the predetermined sampling point may be a position at the very end of the frame; i.e., the sampling point may correspond to the K / Qth spectrum 589. If the spatial metadata conveys two sets of spatial parameters, the first sampling point may correspond to the K / 2Qth spectrum 589 and the second sampling point may correspond to the K / Qth spectrum 589.

[0109] In the case of "sharp" transitions, the sampling points 583, 584 may be variable and may be signaled in the bitstream 562. The portion of the bitstream 562 that carries information about the number of sets of spatial parameters used in a frame, the selection between "smooth" and "sharp" transitions, and the location of the sampling points in the case of "sharp" transitions may be referred to as the "framing" portion of the bitstream 562. Figure 7a shows an exemplary transition scheme that may be applied by the decoding system 100 depending on the framing information contained in the received bitstream 562.

[0110] As an example, the frame configuration information for a particular frame may indicate a "smooth" transition and a single set of spatial parameters 711. In this case, the decoding system 100 (e.g., the first mixing matrix 130) may assume that the sampling points for the set of spatial parameters 711 correspond to the last spectrum of the particular frame. Further, the decoding system 100 may be configured to interpolate 701 (e.g., linearly) between the last received set of spatial parameters 710 for the immediately preceding frame and the set of spatial parameters 711 for the particular frame. In another example, the frame configuration information for a particular frame may indicate a "smooth" transition and two sets of spatial parameters 711, 712. In this case, the decoding system 100 (e.g., the first mixing matrix 130) may assume that the sampling points for the first set of spatial parameters 711 correspond to the last spectrum of the first half of the particular frame, and the sampling points for the second set of spatial parameters 712 correspond to the last spectrum of the second half of the particular frame. Furthermore, the decoding system 100 may be configured to interpolate 702 (e.g., linearly) between the last received set of spatial parameters 710 for the immediately preceding frame and said set of spatial parameters 711, and between the first set of spatial parameters 711 and the second set of spatial parameters 712.

[0111] In a further example, the frame configuration information for a particular frame may indicate a "sharp" transition, a single set of spatial parameters 711, and a sampling point 583 for the single set of spatial parameters 711. In this case, the decoding system 100 (e.g., the first mixing matrix 130) may be configured to apply the last received set of spatial parameters 710 for the immediately preceding frame up to the sampling point 583, and apply the set of spatial parameters 711 starting from the sampling point 583 (as shown by the curve 703). In another example, the frame configuration information for a particular frame may indicate a "sharp" transition, two sets of spatial parameters 711, 712, and two corresponding sampling points 583, 584 for the two sets of spatial parameters 711, 712. In this case, the decoding system 100 (e.g., the first mixing matrix 130) may be configured to apply the last received set of spatial parameters 710 for the immediately preceding frame up to the first sampling point 583, apply the first set of spatial parameters 711 starting from the first sampling point 583 up to the second sampling point 584, and apply the second set of spatial parameters 712 starting from the second sampling point 584 at least until the end of that particular frame (as shown by curve 704).

[0112] The encoding system 500 should ensure that the frame configuration information matches the signal characteristics and that appropriate portions of the input signal 561 are selected to calculate one or more sets of spatial parameters 711, 712. To this end, the encoding system 500 may have a detector configured to detect signal locations where the signal energy in one or more channels increases abruptly. If at least one such signal location is found, the encoding system 500 may be configured to switch from a "smooth" transition to a "sharp" transition; otherwise, the encoding system 500 may continue with the "smooth" transition.

[0113] As outlined above, the encoding system 500 (e.g., the parameter determination unit 523) may be configured to calculate spatial parameters for a current frame based on multiple frames 585, 590 of the input audio signal 561 (e.g., based on the current frame 585 and based on the immediately following frame 590, i.e., the so-called look-ahead frame). Thus, the parameter determination unit 523 may be configured to determine spatial parameters based on 2 × K / Q spectra 589 (as shown in FIG. 5e). The spectra 589 may be windowed by a window 586, as shown in FIG. 5e. It is proposed herein to adapt the window 586 based on the number of sets 711, 712 of spatial parameters to be determined, based on the type of transition, and / or based on the positions of the sampling points 583, 584. This ensures that the frame structure information matches the signal characteristics and that appropriate portions of the input signal 561 are selected for calculating the one or more sets 711, 712 of spatial parameters.

[0114] Below, exemplary window functions are described for various encoder / signal situations.

[0115] a) Situation: Single set of spatial parameters 711, smooth transition, no transients in the look-ahead frame 590 Window function 586: Between the last spectrum of the previous frame and the K / Qth spectrum 589, the window function 586 may rise linearly from 0 to 1. Between the K / Qth spectrum and the 48th spectrum 589, the window function 586 may fall linearly from 1 to 0 (see Figure 5e).

[0116] b) Situation: single set of spatial parameters 711, smooth transition, transient component in Nth spectrum (N>K / Q), i.e., transient component in look-ahead frame 590 Window function 721 as shown in Figure 7b: Between the last spectrum of the previous frame and the K / Qth spectrum, window function 721 rises linearly from 0 to 1. Between the K / Qth spectrum and the (N-1)th spectrum, window function 721 remains constant at 1. Between the Nth spectrum and the 2*K / Qth spectrum, window function 586 remains constant at 0. The transient component in the Nth spectrum is represented by a transient point 724 (which corresponds to a sampling point for the set of spatial parameters of the immediately following frame 590). Also shown in Figure 7b are complementary window function 722 (which is applied to the spectrum of the current frame 585 when determining the one or more sets of spatial parameters for the immediately preceding frame) and window function 723 (which is applied to the spectrum of the immediately following frame 590 when determining the one or more sets of spatial parameters for the immediately following frame). Overall, the window function 721 ensures that, in the case of one or more transient components in the look-ahead frame 590, the spectrum of the look-ahead frame before the first transient point 724 is fully taken into account to determine the set of spatial parameters 711 for the current frame 585, while the spectrum of the look-ahead frame 590 after the transient point 724 is ignored.

[0117] c) Situation: Single set of spatial parameters 711, sharp transition, transient component in Nth spectrum (N≦K / Q), no transient component in the immediately following frame 590 Window function 731 as shown in Figure 7c: Between the first spectrum and the (N-1)th spectrum, window function 731 remains constant at 0. Between the Nth spectrum and the K / Qth spectrum, window function 731 remains constant at 1. Between the K / Qth spectrum and the 2*K / Qth spectrum, window function 731 drops linearly from 1 to 0. Figure 7c shows a transition point 734 in the Nth spectrum, which corresponds to a sampling point for a single set of spatial parameters 711. Furthermore, Figure 7c shows window function 732 applied to the spectrum of current frame 585 when determining the one or more sets of spatial parameters for the immediately preceding frame, and window function 733 applied to the spectrum of immediately succeeding frame 590 when determining the one or more sets of spatial parameters for the immediately succeeding frame.

[0118] d) Situation: single set of spatial parameters, sharp transitions, transient components in the Nth and Mth spectra (N≦K / Q, M>K / Q) Window function 741 in Figure 7d: Between the first spectrum and the (N-1)th spectrum, window function 741 remains constant at 0. Between the Nth spectrum and the (M-1)th spectrum, window function 741 remains constant at 1. Between the Mth spectrum and the 48th spectrum, window function remains constant at 0. Figure 7d shows a transition point 744 (i.e., a sampling point of the set of spatial parameters) in the Nth spectrum and a transition point 745 in the Mth spectrum. Furthermore, Figure 7d shows a window function 742 applied to the spectrum of current frame 585 when determining the one or more sets of spatial parameters for the immediately preceding frame, and a window function 743 applied to the spectrum of immediately succeeding frame 590 when determining the one or more sets of spatial parameters for the immediately succeeding frame.

[0119] e) Situation: two sets of spatial parameters, smooth transition, no transient components in subsequent frames Window function: i) The first set of spatial parameters: Between the last spectrum of the previous frame and the K / 2Q-th spectrum, the window function linearly rises from 0 to 1. Between the K / 2Q-th spectrum and the K / Q-th spectrum, the window linearly descends from 1 to 0. Between the K / Q-th spectrum and the 2*K / Q-th spectrum, the window remains constant at 0.

[0120] ii) The second set of spatial parameters: Between the first spectrum and the K / 2Q-th spectrum, the window remains constant at 0. Between the K / 2Q-th spectrum and the K / Q-th spectrum, the window linearly rises from 0 to 1. Between the K / Q-th spectrum and the 3*K / 2Q-th spectrum, the window linearly descends from 1 to 0. Between the 3*K / 2Q-th spectrum and the 2*K / Q-th spectrum, the window remains constant at 0.

[0121] f) Situation: Two sets of spatial parameters, smooth transition, transient component at the N-th spectrum (N > K / Q) Window function: i) The first set of spatial parameters: Between the last spectrum of the previous frame and the K / 2Q-th spectrum, the window linearly rises from 0 to 1. Between the K / 2Q-th spectrum and the K / Q-th spectrum, the window linearly descends from 1 to 0. Between the K / Q-th spectrum and the 2*K / Q-th spectrum, the window remains constant at 0.

[0122] ii) The second set of spatial parameters: Between the first spectrum and the K / 2Q-th spectrum, the window remains constant at 0. Between the K / 2Q-th spectrum and the K / Q-th spectrum, the window linearly rises from 0 to 1. Between the K / Q-th spectrum and the (N - 1)-th spectrum, the window remains constant at 1. Between the N-th spectrum and the 2*K / Q-th spectrum, the window remains constant at 0.

[0123] g) Situation: Two sets of parameters, sharp transition, transient components at the N-th spectrum and the M-th spectrum (N < M ≤ K / Q), no transient components in subsequent frames Window function: i) A first set of spatial parameters: Between the first spectrum and the (N-1)th spectrum, the window remains constant at 0. Between the Nth spectrum and the (M-1)th spectrum, the window remains constant at 1. Between the Mth spectrum and the 2*K / Qth spectrum, the window remains constant at 0.

[0124] ii) Second set of spatial parameters: Between the first spectrum and the (M-1)th spectrum, the window remains constant at 0. Between the Mth spectrum and the K / Qth spectrum, the window remains constant at 1. Between the K / Qth spectrum and the 2*K / Qth spectrum, the window drops linearly from 1 to 0.

[0125] h) Situation: Two sets of spatial parameters, sharp transition, transient components (N<M≦K / Q、O> K / Q) Window function: i) A first set of spatial parameters: Between the first spectrum and the (N-1)th spectrum, the window remains constant at 0. Between the Nth spectrum and the (M-1)th spectrum, the window remains constant at 1. Between the Mth spectrum and the 2*K / Qth spectrum, the window remains constant at 0.

[0126] ii) A second set of spatial parameters: Between the first spectrum and the (M-1)th spectrum, the window remains constant at 0. Between the Mth spectrum and the (O-1)th spectrum, the window remains constant at 1. Between the Oth spectrum and the 2*K / Qth spectrum, the window remains constant at 0.

[0127] Overall, the following exemplary rules for the window function for determining the current set of spatial parameters may be defined:

[0128] If the current set of spatial parameters is not associated with a transient component The window function provides a smooth phasing of the spectra from the sampling point of the previous set of spatial parameters to the sampling point of the current set of spatial parameters; If the subsequent set of spatial parameters is not associated with a transient component, the window function provides a smooth phase-out of the spectra from the sampling points of the current set of spatial parameters to the sampling points of the subsequent set of spatial parameters; If a subsequent set of spatial parameters is associated with a transient component, the window function fully considers the spectra from the sampling point of the current set of spatial parameters to the spectrum before the sampling point of the subsequent set of spatial parameters and cancels the spectra starting from the sampling point of the subsequent set of spatial parameters.

[0129] If the current set of spatial parameters is associated with a transient component The window function cancels out spectra preceding the sampling point of the current set of spatial parameters; If the sampling point of the subsequent set of spatial parameters is associated with a transient component, the window function fully considers the spectra from the sampling point of the current set of spatial parameters to the spectrum before the sampling point of the subsequent set of spatial parameters and cancels the spectra starting from the sampling point of the subsequent set of spatial parameters; If the subsequent set of spatial parameters is not associated with a transient component, the window function takes into full account the spectra from the sampling point of the current set of spatial parameters to the spectrum at the end of the current frame and provides a smooth phase-out of the spectra from the beginning of the look-ahead frame to the sampling point of said subsequent set of spatial parameters.

[0130] The following describes a method for reducing delay in a parametric multi-channel codec system having an encoding system 500 and a decoding system 100. As outlined above, the encoding system 500 has several processing paths, such as generating and encoding a downmix signal and determining and encoding parameters. The decoding system 100 typically performs decoding of the encoded downmix signal and generating a decorrelated downmix signal. Furthermore, the decoding system 100 performs decoding of the encoded spatial metadata. The decoded spatial metadata is then applied to the decoded downmix signal and the decorrelated downmix signal in a first upmix matrix 130 to generate an upmix signal.

[0131] It is desirable to provide an encoding system 500 configured to provide a bitstream 564 that enables the decoding system 100 to generate the upmix signal Y with reduced delay and / or reduced buffer memory. As outlined above, the encoding system 500 has several different paths by which the encoded data provided to the decoding system 100 in the bitstream 564 can be aligned to properly match upon decoding. As outlined above, the encoding system 500 performs downmixing and encoding of the PCM signal 561. Furthermore, the encoding system 500 determines spatial metadata from the PCM signal 561. Furthermore, the encoding system 500 may be configured to determine one or more clip gains (typically one clip gain per frame). The clip gain indicates the clipping prevention gain applied to the downmix signal X to ensure that the downmix signal X is not clipped. The one or more clip gains may be transmitted in the bitstream 564 (typically in a spatial metadata frame) to enable the decoding system 100 to regenerate the upmix signal Y. Additionally, encoding system 500 may be configured to determine one or more dynamic range control (DRC) values (e.g., one or more DRC values per frame). The one or more DRC values may be used by decoding system 100 to perform dynamic range control of upmixed signal Y. In particular, the one or more DRC values may ensure that the DRC performance of the parametric multi-channel codec system described herein is similar to (or equal to) the DRC performance of a legacy multi-channel codec system, such as Dolby Digital Plus. The one or more DRC values may be transmitted within the downmix audio frame (e.g., within an appropriate field of the Dolby Digital Plus bitstream).

[0132] Thus, encoding system 500 may have at least four signal processing paths. To align these four paths, encoding system 500 may also take into account delays introduced into the system by various processing components not directly related to encoding system 500, such as core encoder delay, core decoder delay, spatial metadata decoder delay, LFE filter delay (for filtering the LFE channel), and / or QMF decomposition delay.

[0133] To align the various paths, the delay of the DRC processing path may be taken into account. The DRC processing delay may typically only be frame-aligned, and not aligned per time sample. Therefore, the DRC processing delay typically only depends on the core encoder delay, which may be rounded up to the next frame alignment. That is, DRC processing delay = round up (core encoder delay / frame size). Based on this, the downmix processing delay for generating the downmix signal may be determined. This is because the downmix processing delay can be delayed per time sample. That is, downmix processing delay = DRC delay × frame size - core encoder delay. The remaining delays can be calculated by summing the individual delay lines and ensuring that the delays match at the decoder stage. This is shown in Figure 8.

[0134] By taking into account the various processing delays, when writing the bitstream 564, processing power (number of input channels - 1 x 1536 fewer copy operations) and memory in the decoding system can be reduced when delaying the resulting spatial metadata by one frame (number of input channels x 1536 x 4 bytes - 245 bytes less memory) instead of delaying the encoded PCM data by 1536 samples. As a result of the delays, all signal paths are more closely aligned in time samples, not just roughly matched.

[0135] As outlined above, FIG. 8 illustrates various delays incurred by the exemplary encoding system 500. The bracketed numbers in FIG. 8 indicate exemplary delays in terms of the number of samples of the input signal 561. The encoding system 500 typically includes a delay 801 caused by filtering the LFE channel of the multi-channel input signal 561. Additionally, a delay 802 (referred to as the "clipgainpcmdelayline") may be introduced by determining a clip gain (i.e., the DRC2 parameter, described below) to be applied to the input signal 561 to prevent clipping of the downmix signal. In particular, this delay 802 may be introduced to synchronize the clip gain application in the encoding system 500 with the clip gain application in the decoding system 100. For this purpose, the input to the downmix calculation (performed by the downmix processing unit 510) may be delayed by an amount equal to the delay 811 (referred to as the "coredecdelay") of the decoder 140 of the downmix signal. This means that in the example shown, clipgainpcmdelayline=coredecdelay=288 samples.

[0136] The downmix processing unit 510 (e.g., having a Dolby Digital Plus encoder) delays the processing path of the audio data, i.e., the downmix signal, but the downmix processing unit 510 does not delay the processing path of the spatial metadata or the processing path for the DRC / clip gain data. As a result, the downmix processing unit 510 should delay the calculated DRC gain, clip gain, and spatial metadata. For DRC gain, this delay typically needs to be an integer multiple of one frame. The delay 807 of the DRC delay line (referred to as "drcdelayline") can be calculated as drcdelayline = ceil((corencdelay + clipgainpcmdelayline) / frame_size) = 2 frames. Here, "coreencdelay" refers to the delay 810 of the encoder of the downmix signal.

[0137] The DRC gain delay can typically only be an integer multiple of the frame size. Therefore, an additional delay may need to be added in the downmix processing path to compensate for this and round up to the next integer multiple of the frame size. The additional downmix delay 806 (referred to as "dmxdelayline") may be determined by dmxdelayline+coreencdelay+clipgainpcmdelayline=drcdelayline*frame_size, where dmxdelayline=drcdelayline*frame_size-coreencdelay-clipgainpcmdelayline, resulting in dmxdelayline=100.

[0138] When spatial parameters are applied in the frequency domain (e.g., in the QMF domain) at the decoder side, they should be synchronized with the downmix signal. To compensate for the fact that the encoder of the downmix signal does not delay the spatial metadata frames but rather the downmix processing path, the input to the parameter extractor 420 should be delayed so that the following condition holds: dmxdelayline + coreencdelay + coredecdelay + aspdecanadelay = aspdelayline + qmfanadelay + framingdelay. In the above formula, "qmfanadelay" (QMF decomposition delay) specifies the delay 804 introduced by the transform unit 521, and "framingdelay" (framing delay) specifies the delay 805 introduced by the windowing of the transform coefficients 580 and the determination of the spatial parameters. As outlined above, the framing calculation uses two frames as input: the current frame and the look-ahead frame. Due to the look-ahead, the framing introduces a delay 805 of exactly one frame length. Furthermore, since delay 804 is known, the additional delay to be applied to the processing path to determine the spatial metadata is aspdelayline=dmxdelayline+coreencdelay+coredecdelay+aspdecanadelay-qmfanadelay-framingdelay=1856. Since this delay is greater than one frame, the memory size of the delay line can be reduced by delaying the calculated bitstream instead of delaying the input PCM data, so that aspbsdelayline=floor(aspdelayline / frame_size)=1 frame (delay 809) and asppcmdelayline=aspdelayline-aspbsdelayline*frame_size=320 (delay 803).

[0139] After calculating the one or more clip gains, the one or more clip gains are provided to the bitstream generation unit 530. Thus, the one or more clip gains undergo a delay that is applied to the final bitstream by aspbsdelayline 809. Thus, the additional delay 808 for a clip gain should be: clipgainbsdelayline+aspbsdelayline=dmxdelayline+coreencdelay+coredecdelay, which gives clipgainbsdelayline=dmxdelayline+coreencdelay+coredecdelay-aspbsdelayline=1 frame. In other words, it should be ensured that the one or more clip gains are provided to the decoding system 500 immediately after the decoding of the corresponding frame of the downmix signal. Thereby, the one or more clip gains can be applied to the downmix signal before performing the upmix in the upmix stage 130.

[0140] 8 illustrates additional delays incurred in the decoding system 100, such as a delay 812 (referred to as "aspdecanadelay") caused by the time-domain to frequency-domain transformations 301, 302 of the decoding system 100, a delay 813 (referred to as "aspdecsyndelay") caused by the frequency-domain to time-domain transformations 311-316, and a further delay 814.

[0141] 8, the various processing paths of the codec system have processing-related delays and alignment delays that ensure that the various output data from the various processing paths is available when needed in the decoding system 100. The alignment delays (e.g., delays 803, 809, 807, 808, 806) are provided within the encoding system 500, thereby reducing the processing power and memory required in the decoding system 100. The total delays for the various processing paths (excluding the LFE filter delay 801, which is applicable to all processing paths) are as follows:

[0142] Downmix processing path: sum of delays 802, 806, 810 = 3072, i.e. 2 frames; ·DRC processing path: delay 807 = 3072, i.e. 2 frames; Clip gain processing path: sum of delays 808, 809, 802 = 3360, which corresponds to the delay 811 of the decoder of the downmix signal plus the delay of the downmix processing path; Spatial metadata processing path: sum of delays 802, 803, 804, 805, 809 = 4000. This corresponds to the delay of the downmix processing path plus the delay 811 of the decoder of the downmix signal and the delay 812 caused by the time domain to frequency domain conversion stages 301, 302.

[0143] Thus, it is ensured that the DRC data is available to the decoding system 100 at time 821 , the clip gain data is available at time 822 , and the spatial metadata is available at time 823 .

[0144] 8, it can be seen that the bitstream generation unit 530 may combine encoded audio data and spatial metadata that may relate to different excerpts of the input audio signal 561. In particular, it can be seen that the downmix processing path, the DRC processing path, and the clip gain processing path have a delay of exactly two frames (3072 samples) (neglecting delay 801) by the output of the encoding system 500 (indicated by interfaces 831, 832, and 833). The encoded downmix signal is provided by interface 831, the DRC gain data is provided by interface 832, and the spatial metadata and clip gain data are provided by interface 833. Typically, the encoded downmix signal and DRC gain data are provided in regular Dolby Digital Plus frames, and the clip gain data and spatial metadata may be provided in spatial metadata frames (e.g., in auxiliary fields of the Dolby Digital Plus frames).

[0145] It can be seen that the spatial metadata processing path in the interface 833 has a delay of 4000 samples (neglecting the delay 801), which differs from the delay of the other processing paths (3072 samples). This means that a spatial metadata frame may relate to a different excerpt of the input signal 561 than a frame of the downmix signal. In particular, to ensure alignment in the decoding system 100, the bitstream generation unit 530 should be configured to generate a bitstream 564 including a sequence of bitstream frames, where a bitstream frame indicates a frame of the downmix signal corresponding to a first frame of the multi-channel input signal 561 and a spatial metadata frame corresponding to a second frame of the multi-channel input signal 561. The first and second frames of the multi-channel input signal 561 may contain the same number of samples. Nevertheless, the first and second frames of the multi-channel input signal 561 may differ from each other. In particular, the first and second frames may correspond to different excerpts of the multi-channel input signal 561. More particularly, the first frame may include samples that precede the samples of the second frame. For example, the first frame may include samples of the multi-channel input signal 561 that precede the samples of the second frame of the multi-channel input signal 561 by a predetermined number of samples, for example 928 samples.

[0146] As outlined above, encoding system 500 may be configured to determine dynamic range control (DRC) and / or clip gain data. In particular, encoding system 500 may be configured to ensure that downmix signal X is not clipped. Furthermore, encoding system 500 may be configured to provide dynamic range control (DRC) parameters that ensure that the DRC behavior of multi-channel signal Y, encoded using the above-mentioned parametric encoding scheme, is similar to or equal to the DRC behavior of multi-channel signal Y encoded using a reference multi-channel encoding system (such as Dolby Digital Plus).

[0147] FIG. 9a is a block diagram of an exemplary dual-mode encoding system 900. It should be noted that the portions 930, 931 of the dual-mode encoding system 900 are typically separate. An n-channel input signal Y 561 is provided to each of an upper portion 930, which is active in at least the multi-channel encoding mode of the encoding system 900, and a lower portion 931, which is active in at least the parametric encoding mode of the encoding system 900. The lower portion 931 of the encoding system 900 may correspond to or include, for example, the encoding system 500. The upper portion 930 may correspond to a reference multi-channel encoder (such as a Dolby Digital Plus encoder). The upper portion 930 typically includes a discrete-mode DRC analyzer 910 arranged in parallel with an encoder 911, both of which receive the audio signal Y 561 as an input. Based on this input signal 561, the encoder 911 outputs an encoded n-channel signal (Y). Meanwhile, the DRC analyzer 910 outputs one or more post-processing DRC parameters DRC1 that quantify the decoder-side DRC to be applied. The DRC parameters DRC1 may be "compr" gain (compressor gain) and / or "dynrng" gain (dynamic range gain) parameters. The parallel outputs from both units 910, 911 are collected by a discrete-mode multiplexer 912, which outputs a bitstream P. The bitstream P may have a predetermined syntax, for example the syntax of Dolby Digital Plus.

[0148] The lower portion 931 under the encoding system 900 has a parametric analysis stage 922 arranged in parallel with the parametric mode DRC analyzer 921. The parametric mode DRC analyzer 921, like the parametric analysis stage 922, receives an n-channel input signal Y. The parametric analysis stage 922 may have a parameter extractor 420. Based on the n-channel audio signal Y, the parametric analysis stage 922 outputs one or more mixing parameters, collectively represented by α in FIGS. 9a and 9b (as outlined above), and an m-channel (1 < m < n) downmix signal X. The downmix signal X is then processed by a core signal encoder 923 (e.g., a Dolby Digital Plus encoder), which outputs an encoded downmix signal (X with a hat) based on it. The parametric analysis stage 922 applies dynamic range limiting in the time blocks or frames of the input signal when it may be necessary. A possible condition for controlling when to apply dynamic range limiting can be the "non-clipping condition" or the "in-range condition". This implies that in time blocks or frame segments where the downmix signal has a large amplitude, the signal is processed to fit within a defined range. This condition may be implemented based on one time block or a one-hour frame containing several time blocks. As an example, a frame of the input signal 561 may contain a predetermined number (e.g., 6) of blocks. Preferably, the above condition is implemented by applying a wide-spectrum gain reduction rather than just clipping the peak value or using a similar approach.

[0149] 9b shows a possible implementation of the parametric decomposition stage 922, which includes a preprocessor 927 and a parametric decomposition processor 928. The preprocessor 927 is responsible for performing dynamic range limitation on the n-channel input signal 561, thereby outputting a dynamically range-limited n-channel signal, which is fed to the parametric decomposition processor 928. The preprocessor 527 also outputs a block- or frame-wise value of a preprocessing DRC parameter DRC2. The parameter DRC2, together with the mixing parameter α and the m-channel downmix signal X from the parametric decomposition processor 928, are included in the output from the parametric decomposition stage 922.

[0150] The parameter DRC2 may also be referred to as clip gain. The parameter DRC2 may indicate a gain applied to the multi-channel input signal 561 to ensure that the downmix signal X is not clipped. The one or more channels of the downmix signal X may be determined from the channels of the input signal Y by determining a linear combination of some or all of the channels of the input signal Y. As an example, the input signal Y may be a 5.1 multi-channel signal and the downmix signal may be a stereo signal. Samples of the left and right channels of the downmix signal may be generated based on different linear combinations of samples of the 5.1 multi-channel input signal.

[0151] The DRC2 parameters may be determined so that the maximum amplitude of the channels of the downmix signal does not exceed a predetermined threshold. This may be ensured on a block-by-block or frame-by-frame basis. A single gain (clip gain) may be applied to the channels of the multi-channel input signal Y to ensure that the above-mentioned conditions are met. The DRC2 parameters may indicate this gain (e.g., the inverse of this gain).

[0152] Referring to FIG. 9a, it should be noted that the discrete-mode DRC analyzer 910 functions similarly to the parametric-mode DRC analyzer 921 in that it outputs one or more post-processing DRC parameters DRC1 that quantify the decoder-side DRC to be applied. Thus, the parametric-mode DRC analyzer 921 may be configured to simulate the DRC processing performed by the reference multi-channel encoder 930. The parameters DRC1 provided by the parametric-mode DRC analyzer 921 are typically not included in the bitstream P in parametric coding mode; instead, they are compensated to account for the dynamic range limitation performed by the parametric decomposition stage 922. To this end, the DRC up-compensator 924 receives the post-processing DRC parameters DRC1 and the pre-processing DRC parameters DRC2. For each block or frame, the DRC up-compensator 924 derives values for one or more compensated post-processing DRC parameters DRC3. These post-processing DRC parameters are such that the combined effect of the compensated post-processing DRC parameter DRC3 and the pre-processing DRC parameter DRC2 is quantitatively equivalent to the DRC quantified by the post-processing DRC parameter DRC1. In other words, the DRC up-compensator 924 is configured to reduce the post-processing DRC parameters output by the DRC analyzer 921 by the portion, if any, already performed by the parametric decomposition stage 922. Included in the bitstream P is the compensated post-processing DRC parameter DRC3.

[0153] Referring to the lower portion 931 of the system 900, the parametric mode multiplexer 925 collects the compensated post-processing DRC parameters DRC3, the pre-processing DRC parameters DRC2, the mixing parameter α, and the encoded downmix signal X, and forms a bitstream P based thereon. Thus, the parametric mode multiplexer 925 may include or correspond to the bitstream generation unit 530. In one possible implementation, the compensated post-processing DRC parameters DRC3 and the pre-processing DRC parameters DRC2 may be encoded in logarithmic form as dB values that affect decoder-side amplitude upscaling or downscaling. The compensated post-processing DRC parameter DRC3 may have any sign. However, the post-processing DRC parameter DRC2, resulting from implementations such as the "no-clip condition," is typically represented by a non-negative dB value at all times.

[0154] FIG. 10 illustrates exemplary processing that may be performed, for example, in parametric mode DRC analyzer 921 and DRC up compensator 924 to determine modified DRC parameters DRC3 (e.g., modified “dynrng gain” and “compr gain” parameters).

[0155] The DRC2 and DRC3 parameters may be used to ensure that a decoding system reproduces different audio bitstreams at consistent loudness levels. Furthermore, it may be ensured that bitstreams generated by the parametric encoding system 500 have consistent loudness levels relative to bitstreams generated by legacy and / or reference encoding systems (such as Dolby Digital Plus). As outlined above, this may be ensured by generating an unclipped downmix signal by the encoding system 500 (using the DRC2 parameters) and by providing DRC2 parameters (e.g., the inverse of the attenuation applied to prevent clipping of the downmix signal) in the bitstream to enable the decoding system 100 (when generating the upmix signal) to recreate the original loudness.

[0156] As outlined above, the downmix signal is typically generated based on a linear combination of some or all of the channels of the multi-channel input signal 561. Thus, the scaling factor (or attenuation) applied to a channel of the multi-channel input signal 561 may depend on all channels of the multi-channel input signal 561 that contributed to the downmix signal. In particular, the one or more channels of the downmix signal may be determined based on the LFE channel of the multi-channel input signal 561. As a result, the scaling factor (or attenuation) applied for clipping protection should also take the LFE channel into account. This differs from other multi-channel encoding systems (such as Dolby Digital Plus), in which the LFE channel is typically not taken into account for clipping protection. By taking into account the LFE channel and / or all channels that contributed to the downmix signal, the quality of the clipping protection may be improved.

[0157] Thus, the one or more DRC2 parameters provided to the corresponding decoding system 100 may depend on all channels of the input signal 561 that contributed to the downmix signal. In particular, the DRC2 parameters may depend on the LFE channel. This may improve the quality of the clipping protection.

[0158] It should be noted that the dialnorm parameter may not be taken into account for the calculation of the scaling factor and / or the DRC2 parameter (as shown in FIG. 10).

[0159] As outlined above, the encoding system 500 may be configured to write so-called "clip gains" (i.e., DRC2 parameters) into the spatial metadata frame, indicating what gains have been applied to the input signal 561 to prevent clipping in the downmix signal. The corresponding decoding system 100 may be configured to exactly undo the clip gains applied in the encoding system 500. However, only the clip gain sampling points are transmitted in the bitstream. In other words, the clip gain parameters are typically determined only on a frame-by-frame or block-by-block basis. Between these sampling points, the decoding system 100 may be configured to interpolate clip gain values (e.g., received DRC2 parameters) between neighboring sampling points.

[0160] An exemplary interpolation curve for interpolating DRC2 parameters for adjacent frames is shown in FIG. 11. In particular, FIG. 11 shows first DRC2 parameters 953 for a first frame and second DRC2 parameters 954 for a subsequent second frame 950. Decoding system 100 may be configured to interpolate between first DRC2 parameters 953 and second DRC2 parameters 954. Interpolation may be performed within a subset 951 of samples of second frame 950, for example, within first block 951 of second frame 950 (as shown by interpolation curve 952). Interpolating DRC2 parameters ensures a smooth transition between adjacent audio frames, thereby avoiding audible artifacts that may be caused by differences between successive DRC2 parameters 953, 954.

[0161] The encoding system 500 (in particular the downmix processing unit 510) may be configured to apply a clip gain interpolation corresponding to the DRC2 interpolation 952 performed by the decoding system 500 when generating the downmix signal. This ensures that the clip gain protection of the downmix signal is consistently removed when generating the upmix signal. In other words, the encoding system 500 may be configured to simulate a curve of DRC2 values resulting from the DRC2 interpolation 952 applied by the decoding system 100. Furthermore, the encoding system 500 may be configured to apply the exact (sample-by-sample) inverse of this curve of DRC2 values to the multi-channel input signal 561 when generating the downmix signal.

[0162] The methods and systems described herein may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software running on, for example, a digital signal processor or microprocessor. Other components may be implemented as hardware and / or application-specific integrated circuits. Signals encountered in the described methods and systems may be stored on media such as random access memory or optical storage media. The signals may be transmitted over networks such as radio, satellite, wireless, or wired networks, e.g., the Internet. Typical devices utilizing the methods and systems described herein are portable electronic devices or other consumer equipment that store and / or render audio signals.

[0163] Several aspects will be described. [Aspect 1] 1. An audio encoding system configured to generate a bitstream indicative of a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal, the bitstream comprising: a downmix processing unit (510) configured to generate the downmix signal from a multi-channel input signal, the downmix signal having m channels and the multi-channel input signal having n channels, n and m being integers, and m <nである、ダウンミックス処理ユニットと;a parameter processing unit (520) configured to determine the spatial metadata from the multi-channel input signal; a configuration unit (540) configured to determine one or more control settings for the parameter processing unit based on one or more external settings, the one or more external settings comprising a target data rate for the bitstream, and the one or more control settings comprising a maximum data rate for the spatial metadata; Audio encoding system. [Aspect 2] the parameter processing unit is configured to determine spatial metadata for a frame of the multi-channel input signal, referred to as a spatial metadata frame; a frame of the multi-channel input signal comprising a predetermined number of samples of the multi-channel input signal; the maximum data rate for the spatial metadata indicates a maximum number of metadata bits for a spatial metadata frame; 2. The audio encoding system of embodiment 1. Aspect 3 3. The audio encoding system of claim 2, wherein the parameter processing unit is configured to determine whether a number of bits for a spatial metadata frame determined based on the one or more control settings exceeds the maximum number of metadata bits. Aspect 4 The spatial metadata frame contains one or more sets of spatial parameters; the one or more control settings include a temporal resolution setting indicating the number of sets of spatial parameters per spatial metadata frame to be determined by the parameter processing unit; the parameter processing unit is configured to discard a set of spatial parameters (711) from the current spatial metadata frame if the current spatial metadata frame has multiple sets of spatial parameters (711, 712) and if the number of bits of the current spatial metadata frame exceeds the maximum number of metadata bits; 4. The audio encoding system of embodiment 3. Aspect 5 the one or more sets of spatial parameters are associated with a corresponding one or more sampling points; the one or more sampling points indicate one or more corresponding time points; the parameter processing unit is configured to discard a first set of spatial parameters (711) from the current spatial metadata frame if the plurality of sampling points (583, 584) of the current metadata frame are not associated with a transient component of the multi-channel input signal, and the first set of spatial parameters is associated with a first sampling point (583) that precedes a second sampling point (584); the parameter processing unit is configured to discard a second set of spatial parameters (712) from the current spatial metadata frame if the plurality of sampling points of the current metadata frame are associated with a transient component of the multi-channel input signal; 5. The audio encoding system of embodiment 4. Aspect 6 the one or more control settings include a quantizer setting indicating a first type of quantizer from a plurality of predetermined types of quantizers; the parameter processing unit is configured to quantize the one or more sets of spatial parameters according to a quantizer of the first type; the plurality of predetermined types of quantizers each providing a different quantizer resolution; the parameter processing unit is configured to requantize one, some or all of the spatial parameters of the one or more sets of spatial parameters according to a second type of quantizer having a lower resolution than the first type of quantizer if it is determined that the number of bits of the current spatial metadata frame exceeds the maximum number of metadata bits; 6. The audio encoding system according to embodiment 4 or 5. Aspect 7 7. The audio encoding system of claim 6, wherein the plurality of predetermined types of quantizers include fine quantization and coarse quantization. Aspect 8 The parameter processing unit: determining a set of temporal difference parameters based on the difference of a current set of spatial parameters (712) relative to a previous set of spatial parameters (711); encoding the set of temporal difference parameters using entropy coding; Inserting the encoded set of temporal difference parameters into the current spatial metadata frame; If it is determined that the number of bits of the current spatial metadata frame exceeds the maximum number of metadata bits, reducing the entropy of the set of temporal difference parameters. 8. The audio encoding system of any one of aspects 4 to 7, configured to: Aspect 9 9. The audio encoding system of claim 8, wherein the parameter processing unit is configured to set one, some, or all of the temporal difference parameters of the set of temporal difference parameters equal to a value that has an increased probability of being one of the possible values of the temporal difference parameters in order to reduce the entropy of the set of temporal difference parameters. Aspect 10 the one or more control settings include a frequency resolution setting; The frequency resolution setting indicates the number of different frequency bands; the parameter processing unit is configured to determine different spatial parameters, called band parameters, for different frequency bands; The set of spatial parameters includes corresponding band parameters for the different frequency bands; 10. The audio encoding system according to any one of aspects 4 to 9. Aspect 11 The parameter processing unit determining a set of frequency difference parameters based on differences of one or more band parameters in a first frequency band relative to corresponding one or more band parameters in a second, adjacent frequency band; encoding the set of frequency difference parameters using entropy coding; Inserting the encoded set of frequency difference parameters into the current spatial metadata frame; reducing the entropy of the set of frequency difference parameters when it is determined that the number of bits of the current spatial metadata frame exceeds the maximum number of metadata bits. 11. The audio encoding system of claim 10, configured to: Aspect 12 12. The audio encoding system of claim 11, wherein the parameter processing unit is configured to set one, some, or all of the frequency difference parameters of the set of frequency difference parameters equal to a value that has an increased probability of a possible value of the frequency difference parameter in order to reduce the entropy of the set of frequency difference parameters. Aspect 13 The parameter processing unit: reducing the number of frequency bands when it is determined that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits; redetermining the one or more sets of spatial parameters for the current spatial metadata frame using a reduced number of frequency bands; 13. The audio encoding system of any one of aspects 10 to 12, configured to: Aspect 14 the one or more external settings further include one or more of: a sampling rate of the multi-channel input signal, a number m of channels of the downmix signal, a number n of channels of the multi-channel input signal, and an update period indicating a time period during which a corresponding decoding system is required to synchronize to the bitstream; the one or more control settings further comprise one or more of: a temporal resolution setting indicating the number of sets of spatial parameters per frame of spatial metadata to be determined; a frequency resolution setting indicating the number of frequency bands over which spatial parameters are to be determined; a quantizer setting indicating the type of quantizer to be used to quantize the spatial metadata; and an indication of whether a current frame of the multi-channel input signal should be encoded as an independent frame; 14. The audio encoding system of any one of aspects 1 to 13. Aspect 15 the one or more external settings further include an update period indicating a time period during which a corresponding decoding system is required to synchronize to the bitstream; the one or more control settings further include an indication of whether the current spatial metadata frame should be encoded as an independent frame; the parameter processing unit is configured to determine a sequence of spatial metadata frames for a corresponding sequence of frames of the multi-channel input signal; the configuration unit is configured to determine, from the sequence of spatial metadata frames, the one or more spatial metadata frames to be encoded as independent frames based on the update period; 15. The audio encoding system of any one of aspects 2 to 14. Aspect 16 The configuration unit: determining whether a current frame of the sequence of frames of the multi-channel input signal includes a sample at a time that is an integer multiple of the update period; Determine whether the current spatial metadata frame corresponding to the current frame is an independent frame. 16. The audio encoding system of claim 15, configured to: Aspect 17 16. The audio encoding system of claim 15, wherein the parameter processing unit is configured to encode one or more sets of spatial parameters of a current spatial metadata frame independently of data included in a previous spatial metadata frame if the current spatial metadata frame is to be encoded as an independent frame. Aspect 18 n=6 and m=2; and / or the multi-channel upmix signal is a 5.1 signal; and / or the downmix signal is a stereo signal; and / or the multi-channel input signal is a 5.1 signal; 18. The audio encoding system of any one of aspects 1 to 17. Aspect 19 the downmix processing unit is configured to encode the downmix signal using a Dolby Digital Plus encoder; The bitstream corresponds to a Dolby Digital Plus bitstream; the spatial metadata is contained within a data field of the Dolby Digital Plus bitstream; 19. The audio encoding system of any one of aspects 1 to 18. Aspect 20 the spatial metadata includes one or more sets of spatial parameters; a spatial parameter of the set of spatial parameters is indicative of a cross-correlation between different channels of the multi-channel input signal; 20. The audio encoding system of any one of aspects 1 to 19. Aspect 21 a parameter processing unit (520) configured to determine spatial metadata frames for generating frames of a multi-channel up-mix signal from corresponding frames of a down-mix signal, the down-mix signal having m channels and the multi-channel up-mix signal having n channels, n, m being integers, and m <nであり、前記空間的メタデータ·フレームは、空間的パラメータの一つまたは複数の集合を含み、当該パラメータ処理ユニットは、a transformation unit (521) configured to determine a plurality of spectra from a current frame and a subsequent frame of a channel of the multi-channel input signal; a parameter determination unit (523) configured to determine the spatial metadata frame for a current frame of the channels of the multi-channel input signal by weighting the plurality of spectra using a window function; the window function depends on one or more of: the number of sets of spatial parameters included in the spatial metadata frame, the presence of one or more transient components in the current frame or a immediately following frame of the multi-channel input signal and / or the time points of the transient components; Parameter Processing Unit. Aspect 22 The window function comprises a set-dependent window function; the parameter determination unit is configured to determine a set of spatial parameters for a current frame of the channels of the multi-channel input signal by weighting the plurality of spectra with the set-dependent window function; The set-dependent window function depends on whether the set of spatial parameters is associated with a transient component or not. 22. The parameter processing unit according to claim 21. Aspect 23 If the set of spatial parameters (711) is not associated with a transient component, the set-dependent window function provides a phasing of the plurality of spectra from a sampling point of a preceding set of spatial parameters (710) to a sampling point of the set of spatial parameters (711); and / or if the subsequent set of spatial parameters (712) is associated with a transient component, the set-dependent window function cancels out the spectra starting from the sampling point of the subsequent set of spatial parameters (712), including the spectra from the sampling point of the set of spatial parameters (711) to the spectrum of the plurality of spectra preceding the sampling point of the subsequent set of spatial parameters (712); 23. The parameter processing unit according to claim 22. Aspect 24 If said set of spatial parameters (711) is associated with a transient component, the set-dependent window function cancels spectra from the plurality of spectra before a sampling point of the set of spatial parameters (711); and / or if a sampling point of a subsequent set of spatial parameters (712) is associated with a transient component, the set-dependent window function cancels out spectra from the plurality of spectra starting from the sampling point of the subsequent set of spatial parameters (712), including spectra from the plurality of spectra from the sampling point of the set of spatial parameters (711) to the spectrum of the plurality of spectra preceding the sampling point of the subsequent set of spatial parameters (712); and / or if the subsequent set (712) of spatial parameters is not associated with a transient component, the set-dependent window function includes a spectrum of the plurality of spectra from a sampling point of the set (711) of spatial parameters to a spectrum of the plurality of spectra at the end of the current frame (585) and provides a phase-out of a spectrum of the plurality of spectra from the beginning of the immediately following frame (590) to a sampling point of the subsequent set (712) of spatial parameters; 23. The parameter processing unit according to claim 22. Aspect 25 a parameter processing unit (520) configured to determine spatial metadata frames for generating frames of a multi-channel up-mix signal from corresponding frames of a down-mix signal, the down-mix signal having m channels and the multi-channel up-mix signal having n channels, n, m being integers, and m <nであり、前記空間的メタデータ·フレームは空間的パラメータの集合を含み、当該パラメータ処理ユニットは:a transform unit (561) configured to determine a first plurality of transform coefficients from a frame of a first channel of a multi-channel input signal and to determine a second plurality of transform coefficients from a corresponding frame of a second channel of the multi-channel input signal, the first and second plurality of transform coefficients providing first and second time / frequency representations of the frames of the first and second channels, respectively, the first and second time / frequency representations including a plurality of frequency bins and a plurality of time bins; a parameter determination unit (523) configured to determine the set of spatial parameters based on the first and second plurality of transform coefficients using fixed-point arithmetic, the set of spatial parameters including corresponding band parameters for different frequency bands including different numbers of frequency bins, a specific band parameter for a specific frequency band being determined based on a transform coefficient from the first and second plurality of transform coefficients of the specific frequency band, and a shift used by the fixed-point arithmetic to determine the specific band parameter depends on the specific frequency band; Parameter Processing Unit. Aspect 26 26. The parameter processing unit of claim 25, wherein the shift used by the fixed-point arithmetic to determine the specific band parameter for the specific frequency band depends on the number of frequency bins included in the specific frequency band. Aspect 27 27. The parameter processing unit of claim 25 or 26, wherein the shift used by the fixed-point arithmetic to determine the specific band parameter for the specific frequency band depends on the number of time bins used to determine the specific band parameter. Aspect 28 28. The parameter processing unit of any one of aspects 25 to 27, wherein the parameter determination unit is configured to determine, for the particular frequency band, a corresponding shift that maximizes accuracy of the particular band parameter. Aspect 29 The parameter determination unit determines the specific band parameters for the specific frequency band by: determining a first energy estimate based on transform coefficients from the first plurality of transform coefficients that fall within the particular frequency band; determining a second energy estimate based on transform coefficients from the second plurality of transform coefficients that fall within the particular frequency band; determining a covariance based on transform coefficients from the first and second plurality of transform coefficients that fall within the particular frequency band; determining the shift for the particular band parameter based on a maximum of the first energy estimate, the second energy estimate, and the covariance; 29. The parameter processing unit of any one of aspects 25 to 28, configured to perform by: Aspect 30 1. An audio encoding system configured to generate a bitstream based on a multi-channel input signal, comprising: a downmix processing unit (510) configured to generate a sequence of frames of a downmix signal from a corresponding sequence of first frames of the multi-channel input signal, the downmix signal having m channels and the multi-channel input signal having n channels, n, m being integers, and m <nである、ダウンミックス処理ユニットと;a parameter processing unit (520) configured to determine a sequence of spatial metadata frames from a second sequence of frames of the multi-channel input signal, wherein the sequence of frames of the downmix signal and the sequence of spatial metadata frames are for generating a multi-channel upmix signal comprising n channels; a bitstream generation unit (503) configured to generate the bitstream comprising a sequence of bitstream frames, the bitstream frames indicating a frame of the downmix signal corresponding to a first frame of the sequence of first frames of the multi-channel input signal and a spatial metadata frame corresponding to a second frame of the sequence of second frames of the multi-channel input signal, the second frame being different from the first frame, Audio encoding system. Aspect 31 the first frame and the second frame have the same number of samples; and / or the samples of the first frame precede the samples of the second frame; 31. The audio encoding system of embodiment 30. Aspect 32 32. The audio encoding system of embodiment 30 or 31, wherein the first frame precedes the second frame by a predetermined number of samples. Aspect 33 33. The audio encoding system of embodiment 32, wherein the predetermined number of samples is 928 samples. Aspect 34 1. An audio encoding system configured to generate a bitstream based on a multi-channel input signal, comprising: A downmix processing unit (510), determining a sequence of clipping protection gains for a corresponding sequence of frames of the multi-channel input signal, a current clipping protection gain indicating an attenuation to be applied to a corresponding current frame of the multi-channel input signal to prevent clipping of the corresponding current frame of the downmix signal; interpolating a current clipping protection gain and a previous clipping protection gain of a previous frame of the multi-channel input signal to provide a clipping protection gain curve; applying the clipping protection gain curve to a current frame of the multi-channel input signal to provide an attenuated current frame of the multi-channel input signal; generating a current frame of the sequence of frames of the downmix signal from an attenuated current frame of the multi-channel input signal, the downmix signal having m channels and the multi-channel input signal having n channels, n and m being integers, and m <nである、段階とを実行するよう構成されているa downmix processing unit; a parameter processing unit (520) configured to determine a sequence of spatial metadata frames from the multi-channel input signal, the sequence of frames of the downmix signal and the sequence of spatial metadata frames being for generating a multi-channel upmix signal comprising n channels; a bitstream generation unit (503) configured to generate the bitstream indicative of the sequence of clipping protection gains, the sequence of frames of the downmix signal and the sequence of spatial metadata frames so as to enable a corresponding decoding system to generate the multi-channel upmix signal, Audio encoding system. Aspect 35 The clipping protection gain curve is: a transition segment that provides a smooth transition from the previous clipping protection gain to the current clipping protection gain; a flat segment that remains flat at the current clipping protection gain; 35. The audio encoding system of embodiment 34. Aspect 36 the transition segment extends through a predetermined number of samples of a current frame of the multi-channel input signal, the predetermined number of samples is greater than 1 and less than the total number of samples of the current frame of the multi-channel input signal; 36. The audio encoding system of embodiment 35. Aspect 37 1. An audio encoding system configured to generate a bitstream indicative of a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal, the bitstream comprising: a downmix processing unit (510) configured to generate the downmix signal from a multi-channel input signal, the downmix signal having m channels and the multi-channel input signal having n channels, n and m being integers, and m <nである、ダウンミックス処理ユニットと;a parameter processing unit configured to determine a sequence of frames of spatial metadata for a corresponding sequence of frames of the multi-channel input signal; a configuration unit (540) configured to determine one or more control settings for the parameter processing unit based on one or more external settings, the one or more external settings include an update period indicating a time period during which a corresponding decoding system is required to synchronize to the bitstream, and the configuration unit is configured to determine, based on the update period, one or more frames of spatial metadata from the sequence of frames of spatial metadata to be encoded as independent frames. Audio encoding system. Aspect 38 1. A method for generating a bitstream indicative of a downmix signal and spatial metadata for generating a multi-channel upmix signal from said downmix signal, comprising: In a step of generating the downmix signal from a multi-channel input signal, the downmix signal has m channels, the multi-channel input signal has n channels, n and m are integers, and m <nである、段階と;determining one or more control settings based on one or more external settings, the one or more external settings including a target data rate for the bitstream, and the one or more control settings including a maximum data rate for the spatial metadata; determining the spatial metadata from the multi-channel input signal in accordance with the one or more control settings. method. Aspect 39 1. A method for determining spatial metadata frames for generating frames of a multi-channel up-mix signal from corresponding frames of a down-mix signal, wherein the down-mix signal has m channels and the multi-channel up-mix signal has n channels, n and m being integers, and m <nであり、前記空間的メタデータ·フレームは、空間的パラメータの一つまたは複数の集合を含み、当該方法は、determining a plurality of spectra from a current frame and a immediately following frame of a channel of the multi-channel input signal; weighting the plurality of spectra using a window function to provide a plurality of weighted spectra; determining the spatial metadata frame for a current frame of the channels of the multi-channel input signal based on the plurality of weighted spectra, wherein the window function depends on one or more of: the number of sets of spatial parameters included in the spatial metadata frame, the presence of one or more transient components and / or time points of the transient components in the current frame or the immediately following frame of the multi-channel input signal, method. Aspect 40 1. A method for determining spatial metadata frames for generating frames of a multi-channel up-mix signal from corresponding frames of a down-mix signal, wherein the down-mix signal has m channels and the multi-channel up-mix signal has n channels, n and m being integers, and m <nであり、前記空間的メタデータ·フレームは、空間的パラメータの集合を含み、当該方法は、determining a first plurality of transform coefficients from frames of a first channel of the multi-channel input signal; determining a second plurality of transform coefficients from corresponding frames of a second channel of the multi-channel input signal, the first and second plurality of transform coefficients providing first and second time / frequency representations of frames of the first and second channels, respectively, the first and second time / frequency representations including a plurality of frequency bins and a plurality of time bins, and the set of spatial parameters including corresponding band parameters for different frequency bands including different numbers of frequency bins; determining a shift to be applied when determining specific band parameters for a specific frequency band using fixed-point arithmetic, the shift being determined based on the specific frequency band; determining the specific band parameters based on the first and second plurality of transform coefficients falling within the specific frequency band using fixed-point arithmetic and the determined shift; method. Aspect 41 1. A method for generating a bitstream based on a multi-channel input signal, comprising: generating a sequence of frames of a downmix signal from a corresponding sequence of first frames of the multi-channel input signal, the downmix signal having m channels and the multi-channel input signal having n channels, n, m being integers, and m <nである、段階と;determining a sequence of spatial metadata frames from a second sequence of frames of the multi-channel input signal, wherein the sequence of frames of the downmix signal and the sequence of spatial metadata frames are for generating a multi-channel upmix signal having n channels; generating the bitstream comprising a sequence of bitstream frames, the bitstream frames indicating a frame of the downmix signal corresponding to a first frame of the sequence of first frames of the multi-channel input signal and a spatial metadata frame corresponding to a second frame of the sequence of second frames of the multi-channel input signal, the second frame being different from the first frame, method. Aspect 42 1. A method for generating a bitstream based on a multi-channel input signal, comprising: determining a sequence of clipping protection gains for a corresponding sequence of frames of the multi-channel input signal, a current clipping protection gain indicating an attenuation to be applied to a corresponding current frame of the multi-channel input signal to prevent clipping of the corresponding current frame of the downmix signal; interpolating a current clipping protection gain and a previous clipping protection gain of a previous frame of the multi-channel input signal to provide a clipping protection gain curve; applying the clipping protection gain curve to a current frame of the multi-channel input signal to provide an attenuated current frame of the multi-channel input signal; generating a current frame of the sequence of frames of the downmix signal from an attenuated current frame of the multi-channel input signal, the downmix signal having m channels and the multi-channel input signal having n channels, n and m being integers, and m <nである、段階と;determining a sequence of spatial metadata frames from the multi-channel input signal, wherein the sequence of frames of the downmix signal and the sequence of spatial metadata frames are for generating a multi-channel upmix signal having n channels; generating the bitstream indicating the sequence of clipping protection gains, the sequence of frames of the downmix signal and the sequence of spatial metadata frames to enable generation of the multi-channel upmix signal based on the bitstream, method. Aspect 43 1. A method for generating a bitstream indicative of a downmix signal and spatial metadata for generating a multi-channel upmix signal from said downmix signal, comprising: In a step of generating the downmix signal from a multi-channel input signal, the downmix signal has m channels, the multi-channel input signal has n channels, n and m are integers, and m <nである、段階と;determining one or more control settings based on one or more external settings, the one or more external settings including an update period indicating a time period during which a decoding system is required to synchronize to the bitstream; determining a sequence of frames of spatial metadata for a corresponding sequence of frames of the multi-channel input signal in accordance with the one or more control settings; encoding one or more frames of spatial metadata from the sequence of frames of spatial metadata as independent frames based on the update period; method. Aspect 44 An audio decoder (140) configured to decode a bitstream generated according to any one of aspects 38, 41 to 43.

Claims

1. receiving, by an audio processor, a multi-channel input audio signal; determining a first set of dynamic range control (DRC) values configured to control the dynamic range of the output audio signal; determining a second set of DRC values configured to prevent the multi-channel input audio signal from being clipped during downmixing by the audio processor, the second set of DRC values being expressed in logarithmic form as dB values; applying the second set of DRC values to the multi-channel input audio signal to obtain an attenuated multi-channel input audio signal; downmixing the attenuated multi-channel input audio signal to obtain a downmix signal; generating the output audio signal from the first set of DRC values and the downmix audio signal. method.

2. one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: receiving a multi-channel input audio signal; determining a first set of dynamic range control (DRC) values configured to control the dynamic range of the output audio signal; determining a second set of DRC values configured to prevent the multi-channel input audio signal from being clipped during downmixing by the device, the second set of DRC values being expressed in logarithmic form as dB values; applying the second set of DRC values to the multi-channel input audio signal to obtain an attenuated multi-channel input audio signal; downmixing the attenuated multi-channel input audio signal to obtain a downmix signal; generating the output audio signal from the first set of DRC values and the downmix audio signal. Device.

3. 1. A non-transitory computer-readable storage medium having a sequence of instructions, which when executed by an audio signal processing apparatus, causes the audio signal processing apparatus to perform a method, the method comprising: receiving, by an audio processor, a multi-channel input audio signal; determining a first set of dynamic range control (DRC) values configured to control the dynamic range of the output audio signal; determining a second set of DRC values configured to prevent the multi-channel input audio signal from being clipped during downmixing by the audio processor, the second set of DRC values being expressed in logarithmic form as dB values; applying the second set of DRC values to the multi-channel input audio signal to obtain an attenuated multi-channel input audio signal; downmixing the attenuated multi-channel input audio signal to obtain a downmix signal; generating the output audio signal from the first set of DRC values and the downmix audio signal. A non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Apparatus and method for encoding and decoding audio signals

    JP2009500658A

  • Audio device

    JP2011035459A

  • Protecting signal clipping using existing audio gain metadata

    JP2012507059A

  • Advanced stereo coding based on adaptively selectable left / right or mid / side stereo coding and parametric stereo coding combinations.

    JP2012521012A

  • Manufacturing method of microstructure

    JP2019009146A