Methods for parametric multichannel encoding
Patent Information
- Application Number
- JP2024110637
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-02-21
- Filing Date
- 2024-07-10
- Publication Date
- 2026-09-14
- Estimated Expiration
- 2034-02-21
Smart Images

Figure 0007920238000006 
Figure 0007920238000007 
Figure 0007920238000008
Abstract
Description
Technical Field
[0001] Cross-Reference to Related Application The present application claims priority to U.S. Provisional Patent Application No. 61 / 767,673 filed on February 21, 2013. The content of this application is incorporated herein by reference in its entirety.
[0002] Technical Field The present document relates to audio encoding systems. In particular, the present document relates to an efficient method and system for parametric multi-channel audio encoding.
Background Art
[0003] Parametric multi-channel audio encoding systems can be used to provide improved listening quality, particularly at low data rates. Nevertheless, there is a need for further improvement of such parametric multi-channel audio encoding systems, particularly in terms of bandwidth efficiency, computational efficiency and / or robustness.
Summary of the Invention
Means for Solving the Problems
[0004] According to one aspect, there is described an audio encoding system configured to generate a bitstream indicative of a downmixed signal and spatial metadata. The spatial metadata may be used by a corresponding decoding system to generate a multi-channel upmixed signal from the downmixed signal. The downmixed signal may have m channels, and the multi-channel upmixed signal may have n channels, where n and m are integers, and m < n. In one example, n=6 and m=2. The spatial metadata may allow a corresponding decoding system to generate the n channels of the multi-channel upmixed signal from the m channels of the downmixed signal.
[0005] An audio encoding system may be configured to quantize and / or encode a downmix signal and spatial metadata, and insert the quantized / encoded data into a bitstream. In particular, the downmix signal may be encoded using a Dolby Digital Plus encoder, and the bitstream may correspond to a Dolby Digital Plus bitstream. The quantized / encoded spatial metadata may be inserted into a data field of the Dolby Digital Plus bitstream.
[0006] An audio encoding system may comprise a downmix processing unit configured to generate a downmix signal from a multi-channel input signal. The downmix processing unit is also referred to as a downmix encoding unit herein. The multi-channel input signal may have n channels, similarly to the multi-channel upmix signal regenerated based on the downmix signal. In particular, the multi-channel upmix signal may provide an approximation of the multi-channel input signal. The downmix unit may comprise the above-mentioned Dolby Digital Plus encoder. The multi-channel upmix signal and the multi-channel input signal may be 5.1 or 7.1 signals, and the downmix signal may be a stereo signal.
[0007] An audio encoding system may have a parameter processing unit configured to determine spatial metadata from a multichannel input signal. In particular, the parameter processing unit (also referred to hereby as a parameter encoding unit) may be configured to determine one or more spatial parameters, for example, a set of spatial parameters. These parameters may be determined based on various combinations of channels in the multichannel input signal. The spatial parameters of the set of spatial parameters may represent the cross-correlations between different channels of the multichannel input signal. The parameter processing unit may be configured to determine spatial metadata for frames of the multichannel input signal, which are referred to as spatial metadata frames. A frame of a multichannel input signal typically contains a predetermined number (e.g., 1536) of samples of the multichannel input signal. Each spatial metadata frame may contain one or more sets of spatial parameters.
[0008] The audio encoding system may further have a configuration setting unit configured to determine one or more control settings for a parameter processing unit based on one or more external settings. The one or more external settings may include a target data rate for a bitstream. Alternatively or additionally, the one or more external settings may include one or more update cycles indicating the sampling rate of the multichannel input signal, the number of channels m of the downmix signal, the number of channels n of the multichannel input signal, and / or the time period for which the corresponding decoding system is required to synchronize with the bitstream. The one or more control settings may include a maximum data rate for spatial metadata. In the case of a spatial metadata frame, the maximum data rate for spatial metadata may indicate the maximum number of metadata bits for the spatial metadata frame. Alternatively or additionally, one or more of the control settings may include: a temporal resolution setting indicating the number of sets of spatial parameters per spatial metadata frame to be determined; a frequency resolution setting indicating the number of frequency bands for which the spatial parameters should be determined; a quantizer setting indicating the type of quantizer to be used to quantize the spatial metadata; and an indication of whether the current frame of the multi-channel input signal should be encoded as independent frames.
[0009] The parameter processing unit may be configured to determine whether the number of bits in a spatial metadata frame determined according to one or more control settings exceeds the maximum number of metadata bits. Furthermore, if the parameter processing unit determines that the number of bits in a particular spatial metadata frame exceeds the maximum number of metadata bits, it may be configured to reduce the number of bits in that particular spatial metadata frame. This reduction in the number of bits may be performed in a resource (processing power) efficient manner. In particular, this reduction in the number of bits may be performed without requiring the complete spatial metadata frame to be recalculated.
[0010] As described above, a spatial metadata frame may contain one or more sets of spatial parameters. The one or more control settings may include a temporal resolution setting indicating the number of sets of spatial parameters per spatial metadata frame to be determined by a parameter processing unit. The parameter processing unit may be configured to determine the number of spatial parameters for the current spatial metadata frame, as indicated by the temporal resolution setting. Typically, the temporal resolution setting takes a value of 1 or 2. Furthermore, the parameter processing unit may be configured to discard sets of spatial parameters from the current spatial metadata frame if the current spatial metadata frame has multiple sets of spatial parameters and if the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits. The parameter processing unit may be configured to retain at least one set of spatial parameters per spatial metadata frame. By discarding sets of spatial parameters from a spatial metadata frame, the number of bits in the spatial metadata frame can be reduced with little computational effort and without significantly affecting the perceived listening quality of the multi-channel upmix signal.
[0011] The set of spatial parameters described above is typically associated with a corresponding set of sampling points, which may indicate a corresponding set of time points. In particular, the sampling points may indicate a time when the decoding system should fully apply the corresponding set of spatial parameters. In other words, the sampling points may indicate a time when the corresponding set of spatial parameters was determined for that set.
[0012] The parameter processing unit may be configured to discard a first set of spatial parameters from the current spatial metadata frame if the plurality of sampling points in the current metadata frame are not associated with transient components of the multichannel input signal, where the first set of spatial parameters is associated with the first sampling point prior to the second sampling point. On the other hand, the parameter processing unit may be configured to discard a second set (typically the last set) of spatial parameters from the current spatial metadata frame if the plurality of sampling points in the current metadata frame are associated with transient components of the multichannel input signal. By doing so, the parameter processing unit may be configured to reduce the impact of discarding a set of spatial parameters on the listening quality of the multichannel upmix signal.
[0013] The one or more control settings may have a quantizer setting that indicates a first type of quantizer from a plurality of predetermined types of quantizers. The plurality of predetermined types of quantizers may each provide a different quantizer resolution. In particular, the plurality of predetermined types of quantizers may include fine quantization and coarse quantization. The parameter processing unit may be configured to quantize the one or more sets of spatial parameters of the current spatial metadata frame according to the first type of quantizer. Furthermore, if the parameter processing unit determines that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits, it may be configured to requantize one, some, or all of the spatial parameters of the one or more sets of spatial parameters according to a second type of quantizer having a lower resolution than the first type of quantizer. In this way, the number of bits in the current spatial metadata frame can be reduced without significantly increasing the computational complexity of the audio encoding system and with only a limited impact on the quality of the upmix signal.
[0014] The parameter processing unit may be configured to determine a set of temporal difference parameters based on the difference between the current set of spatial parameters and the previous set of spatial parameters. In particular, the temporal difference parameters may be determined by determining the difference between a parameter in the current set of spatial parameters and the corresponding parameter in the previous set of spatial parameters. The set of spatial parameters may include, for example, the parameters α1, α2, α3, β1, β2, β3, g, k1, k2 described in this paper. Typically, only one of the parameters k1 or k2 needs to be transmitted. Both parameters are related by relation k1. 2 +k2 2 This is because they can be related by =1. For example, only parameter k1 may be sent, and parameter k2 may be calculated on the receiving end. The time difference parameter may relate to the difference between the corresponding parameters mentioned above.
[0015] The parameter processing unit may be configured to encode a set of time-difference parameters using entropy encoding, for example, using Huffman coding. Furthermore, the parameter processing unit may be configured to insert the encoded set of time-difference parameters into the current spatial metadata frame. Furthermore, the parameter processing unit may be configured to reduce the entropy of the set of time-difference parameters if it is determined that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits. As a result, the number of bits required to entropy encode the time-difference parameters may be reduced. This may reduce the number of bits used for the current spatial metadata frame. As an example, the parameter processing unit may be configured to set one, some, or all of the time-difference parameters in the set of time-difference parameters to a value that has an increased (e.g., best) probability of being one of the possible values of the time-difference parameter. In particular, the probability may be increased compared to the probability of the time-difference parameter prior to the setting operation. Typically, the value with the highest probability of being a time-difference parameter corresponds to 0.
[0016] It should be noted that the temporal difference encoding of the set of spatial parameters mentioned above is typically not required for independent frames. Therefore, the parameter processing unit may be configured to verify whether the current spatial metadata frame is an independent frame and to apply temporal difference encoding only if the current spatial metadata frame is not an independent frame. On the other hand, the frequency difference encoding described later may also be used for independent frames.
[0017] The one or more control settings may include a frequency resolution setting, where the frequency resolution setting indicates the number of different frequency bands for which each spatial parameter, referred to as a band parameter, should be determined. The parameter processing unit may be configured to determine different corresponding spatial parameters (band parameters) for different frequency bands. In particular, different parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for different frequency bands may be determined. Thus, the set of spatial parameters may include the corresponding band parameters for the different frequency bands. For example, the set of spatial parameters may include T corresponding band parameters for T frequency bands, where T is an integer, for example, T = 7, 9, 12, or 15.
[0018] The parameter processing unit may be configured to determine a set of frequency difference parameters based on the difference between one or more band parameters in a first frequency band and one or more corresponding band parameters in a second, adjacent frequency band. Furthermore, the parameter processing unit may be configured to encode the set of frequency difference parameters using entropy encoding, for example, based on Huffman coding. Furthermore, the parameter processing unit may be configured to insert the encoded set of frequency difference parameters into the current spatial metadata frame. Furthermore, the parameter processing unit may be configured to reduce the entropy of the set of frequency difference parameters if it is determined that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits. In particular, the parameter processing unit may be configured to set one, some, or all of the frequency difference parameters in the set of frequency difference parameters to a value (e.g., 0) that has an increased probability of being a possible value for the frequency difference parameter. In particular, the probability may be increased compared to the probability of the frequency difference parameter before the setting operation.
[0019] Alternatively or additionally, the parameter processing unit may be configured to reduce the number of frequency bands if it determines that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits. Furthermore, the parameter processing unit may be configured to redetermine some or all of the aforementioned sets of spatial parameters for the current spatial metadata frame using the reduced number of frequency bands. Typically, changes in the number of frequency bands primarily affect the higher frequency bands. As a result, one or more frequency band parameters may remain unaffected, and therefore the parameter processing unit may not need to recalculate all band parameters.
[0020] As described above, one or more external settings may include an update cycle indicating a time period for which the corresponding decoding system is required to synchronize with the bitstream. Furthermore, one or more control settings may include an indicator of whether the current spatial metadata frame should be encoded as an independent frame. A parameter processing unit may be configured to determine a sequence of spatial metadata frames for a corresponding sequence of frames in the multichannel input signal. A configuration setting unit may be configured to determine, from the sequence of spatial metadata frames, one or more spatial metadata frames that should be encoded as independent frames, based on the update cycle.
[0021] In particular, the one or more independent spatial metadata frames may be determined such that the update period is satisfied (on average). For this purpose, the configuration unit may be configured to determine whether a current frame of the sequence of frames of the multi-channel input signal comprises a sample at a time point (relative to the start point of the multi-channel input signal) that is an integer multiple of the update period. Furthermore, the configuration unit may be configured to determine that a current spatial metadata frame corresponding to the current frame is an independent frame (because it comprises a sample at the time point that is an integer multiple of the update period). The parameter processing unit may be configured to, if the current spatial metadata frame is to be encoded as an independent frame, encode one or more sets of spatial parameters of the current spatial metadata frame independently from data included in previous (and / or future) spatial metadata frames. Typically, if the current spatial metadata frame is to be encoded as an independent frame, all sets of spatial parameters of the current spatial metadata are encoded independently from data included in previous (and / or future) spatial metadata frames.
[0022] According to another aspect, a parameter processing unit configured to determine a spatial metadata frame for generating a frame of a multi-channel upmix signal from a corresponding frame of a downmix signal is described. The downmix signal may have m channels, the multi-channel upmix signal may have n channels, where n and m are integers and m < n. As outlined above, the spatial metadata frame may comprise one or more sets of spatial parameters.
[0023] The parameter processing unit may have a transformation unit configured to determine multiple spectra from the current frame and the immediately following frame (referred to as a look-ahead frame) of a channel of the multi-channel input signal. The transformation unit may utilize a filter bank, such as a QMF filter bank. The spectra of the multiple spectra may include a predetermined number of transformation coefficients within a corresponding predetermined number of frequency bins. The multiple spectra may be associated with a corresponding number of time bins (or time points). Thus, the transformation unit may be configured to provide a time / frequency representation of the current frame and the look-ahead frame. For example, the current frame and the look-ahead frame may each have K samples. The transformation unit may be configured to determine 2 x K / Q spectra, each containing Q transformation coefficients.
[0024] The parameter processing unit may have a parameter determination unit configured to determine a spatial metadata frame for the current frame of a channel of the multi-channel input signal by weighting the multiple spectra using a window function. The window function may be used to adjust the influence of one of the multiple spectra on a particular spatial parameter or on a particular set of spatial parameters. For example, the window function may take a value between 0 and 1.
[0025] The window function may depend on: the number of sets of spatial parameters contained within the spatial metadata frame, the presence of one or more transient components in the current frame or the immediately following frame of the multi-channel input signal, and / or the time points of the transient components. In other words, the window function may be adapted according to the attributes of the current frame and / or the look-ahead frame. In particular, the window function used to determine the set of spatial parameters (referred to as a set-dependent window function) may depend on one or more attributes of the current frame and / or the look-ahead frame.
[0026] Therefore, the window function may include set-dependent window functions. In particular, a window function for determining the spatial parameters of a spatial metadata frame may include (or be composed of) one or more set-dependent window functions for each of the one or more sets of spatial parameters. The parameter determination unit may be configured to determine the set of spatial parameters for the current frame of the channel of the multichannel input signal (i.e., for the current spatial metadata frame) by weighting the multiple spectra using set-dependent window functions. As outlined above, the set-dependent window function may depend on one or more attributes of the current frame. In particular, the set-dependent window function may depend on whether the set of spatial parameters is associated with a transient component.
[0027] For example, if the set of spatial parameters is not associated with transient components, the set-dependent window function may be configured to provide a phase-in of the plurality of spectra starting from the sampling point of the preceding set of spatial parameters to the sampling point of the set of spatial parameters. The phase-in may be provided by a window function that transitions from 0 to 1. Alternatively or additionally, if the set of spatial parameters is not associated with transient components, and the subsequent set of spatial parameters is associated with transient components, the set-dependent window function may include (or fully consider or leave unaffected) the plurality of spectra starting from the sampling point of the set of spatial parameters to the sampling point of the subsequent set of spatial parameters. This may be achieved by a window function with a value of 1. Alternatively or additionally, if the set of spatial parameters is not associated with transient components, and the subsequent set of spatial parameters is associated with transient components, the set-dependent window function may cancel out (or eliminate or attenuate) the plurality of spectra starting from the sampling point of the subsequent set of spatial parameters. This may be achieved by a window function with a value of 0. Alternatively or additionally, if the set of spatial parameters is not associated with transient components, and the subsequent set of spatial parameters is not associated with transient components, the set-dependent window function may phase out the plurality of spectra starting from the sampling point of the set of spatial parameters and up to the spectra of the plurality of spectra prior to the sampling point of the subsequent set of spatial parameters. Phase-out may be provided by a window function that transitions from 1 to 0.
[0028] On the other hand, if the set of spatial parameters is associated with a transient component, the set-dependent window function may cancel the spectra from the plurality of spectra before the sampling points of said set of spatial parameters (or alternatively, may exclude said spectra or attenuate said spectra). Alternatively or additionally, when said set of spatial parameters is associated with a transient component, if a sampling point of a subsequent set of spatial parameters is associated with a transient component, the set-dependent window function may include the spectra from said plurality of spectra starting from the sampling point of said set of spatial parameters up to the spectra of said plurality of spectra before the sampling point of said subsequent set of spatial parameters (that is, may leave said spectra unaffected), and may cancel the spectra from said plurality of spectra starting from the sampling point of said subsequent set of spatial parameters (that is, may exclude said spectra or attenuate said spectra). Alternatively or additionally, when said set of spatial parameters is associated with a transient component, if a subsequent set of spatial parameters is not associated with a transient component, the set-dependent window function may include the spectra of said plurality of spectra starting from the sampling point of said set of spatial parameters up to the spectra of said plurality of spectra at the end of the current frame (that is, may leave said spectra unaffected), and may provide a phase-out of the spectra of said plurality of spectra from the start of the immediately following frame up to the sampling point of said subsequent set of spatial parameters (that is, may gradually attenuate said spectra).
[0029] According to a further aspect, there is described a parameter processing unit configured to determine a spatial metadata frame for generating a frame of a multi-channel upmix signal from a corresponding frame of a downmix signal. The downmix signal may have m channels, and the multi-channel upmix signal may have n channels, where n and m are integers and m<n. As discussed above, the spatial metadata frame may comprise a set of spatial parameters.
[0030] As outlined above, the parameter processing unit may have a conversion unit. The conversion unit may be configured to determine a first set of conversion coefficients from the frames of the first channel of a multi-channel input signal. Furthermore, the conversion unit may be configured to determine a second set of conversion coefficients from the corresponding frames of the second channel of the multi-channel input signal. The first and second channels may be different. Thus, the first and second sets of conversion coefficients provide the first and second time / frequency representations of the corresponding frames of the first and second channels, respectively. As outlined above, the first and second time / frequency representations may include multiple frequency bins and multiple time bins.
[0031] Furthermore, the parameter processing unit may have a parameter determination unit configured to determine a set of spatial parameters based on a plurality of first and second transformation coefficients using fixed-point arithmetic. As shown above, the set of spatial parameters typically includes corresponding band parameters for various frequency bands, where different frequency bands may contain a different number of frequency bins. A specific band parameter for a particular frequency band may be determined based on transformation coefficients from a plurality of first and second transformation coefficients for the particular frequency band (typically without considering transformation coefficients for other frequency bands). The parameter determination unit may be configured to determine the shift used by the fixed-point arithmetic to determine the particular band parameter, depending on the particular frequency band. In particular, the shift used by the fixed-point arithmetic to determine the particular band parameter for the particular frequency band may depend on the number of frequency bins contained within the particular frequency band. Alternatively or additionally, the shift used by the fixed-point arithmetic to determine the particular band parameter for the particular frequency band may depend on the number of time bins to be considered in determining the particular band parameter.
[0032] The parameter determining unit may be configured to determine the shift for the specific frequency band such that the accuracy of the specific band parameter is maximized. This may be achieved by determining the shift required for each product-sum operation in the process of determining said specific band parameter.
[0033] The parameter determining unit may be configured to determine the specific band parameter for the specific frequency band p by determining a first energy (or first energy estimation value) E based on transform coefficients falling within the specific frequency band p among the first plurality of transform coefficients 1,1 (p). Furthermore, a second energy (or second energy estimation value) E may be determined based on transform coefficients falling within the specific frequency band p among the second plurality of transform coefficients 2,2 (p). Furthermore, a cross product or covariance E may be determined based on transform coefficients falling within the specific frequency band p among said first and second plurality of transform coefficients 1,2 (p). The parameter determining unit is configured to obtain the first energy estimation value E 1,1 (p), the second energy estimation value E 2,2 (p) and the covariance E 1,2 (p), and may be configured to determine the shift z for the specific band parameter p based on the maximum among the absolute values of p the foregoing values.
[0034] According to another aspect, described is an audio encoding system configured to generate a bitstream representing a sequence of frames of a downmix signal and a corresponding sequence of spatial metadata frames for generating a corresponding sequence of frames of a multichannel upmix signal from said sequence of frames of the downmix signal. The system may comprise a downmix processing unit configured to generate said sequence of frames of the downmix signal from a corresponding sequence of frames of a multichannel input signal. As indicated above, the downmix signal may have m channels, the multichannel input signal may have n channels, n and m are integers, and m < n. Furthermore, the audio encoding system may comprise a parameter processing unit configured to determine said sequence of spatial metadata frames from said sequence of frames of the multichannel input signal.
[0035] Furthermore, the audio encoding system may have a bitstream generation unit configured to generate the bitstream, which includes a sequence of bitstream frames. Here, the bitstream frames represent a frame of the downmix signal corresponding to a first frame of the multichannel input signal and a spatial metadata frame corresponding to a second frame of the multichannel input signal. The second frame may be different from the first frame. In particular, the first frame may precede the second frame. This allows the spatial metadata frame for the current frame to be transmitted together with the corresponding frame of subsequent frames. This ensures that the spatial metadata frame arrives at the corresponding decoding system only when needed. The decoding system typically decodes the current frame of the downmix signal and generates a decorrelated frame based on the current frame of the downmix signal. This process introduces an algorithmic delay, delaying the spatial metadata frame for the current frame, so that the spatial metadata frame arrives at the decoding system only after the decoded current frame and decorrelated frame have been provided. As a result, the processing power and memory requirements of the decoding system can be reduced.
[0036] In other words, an audio encoding system configured to generate a bitstream based on a multi-channel input signal is described. As outlined above, the system may comprise a downmix processing unit configured to generate a sequence of frames of a downmix signal from a corresponding sequence of first frames of the multi-channel input signal. The downmix signal may have m channels, the multi-channel input signal may have n channels, where n and m are integers and m<n. Furthermore, the audio encoding system may comprise a parameter processing unit configured to determine a sequence of spatial metadata frames from a sequence of second frames of the multi-channel input signal. The sequence of frames of the downmix signal and the sequence of spatial metadata frames may be used by a corresponding decoding system to generate a multi-channel upmix signal comprising n channels.
[0037] The audio encoding system may further comprise a bitstream generation unit configured to generate the bitstream comprising a sequence of bitstream frames. Here, a bitstream frame indicates the frame of the downmix signal corresponding to a first frame of the sequence of first frames of the multi-channel input signal, and a spatial metadata frame corresponding to a second frame of the second frames of the multi-channel input signal. The second frame may be different from the first frame. In other words, the frame configuration used to determine the spatial metadata frames and the frame configuration used to determine the frames of the downmix signal may be different. As outlined above, different frame configurations may be used to ensure that data is aligned in the corresponding decoding system.
[0038] The first and second frames may typically contain the same number of samples (e.g., 1536 samples). Some of the samples in the first frame may precede those in the second frame. In particular, the first frame may precede the second frame by a predetermined number of samples. This predetermined number of samples may correspond, for example, to a certain percentage of the frame's sample count. For example, the predetermined number of samples may correspond to 50% or more of the frame's sample count. In a specific example, the predetermined number of samples corresponds to 928 samples. As shown in this paper, this particular number of samples provides the minimum overall delay and optimal alignment for a particular implementation of the audio encoding and decoding system.
[0039] In a further aspect, an audio encoding system configured to generate a bitstream based on a multi-channel input signal is described. The system may have a downmix processing unit configured to determine a sequence of clipping protection gains (also referred to here as clip gains and / or DRC2 parameters) for a corresponding sequence of frames in the multi-channel input signal. The current clipping protection gain may indicate the attenuation to be applied to the current frame of the multi-channel input signal to prevent clipping of the corresponding current frame of the downmix signal. Similarly, the sequence of clipping protection gains may indicate the respective attenuations to be applied to the frames in the sequence of frames of the multi-channel input signal to prevent clipping of the corresponding frames in the sequence of frames of the downmix signal.
[0040] The downmix processing unit may be configured to interpolate the current clipping protection gain with the preceding clipping protection gain of the preceding frame of the multichannel input signal to obtain a clipping protection gain curve. This may be done in a similar manner for the sequence of clipping protection gains. Furthermore, the downmix processing unit may be configured to apply the clipping protection gain curve to the current frame of the multichannel input signal to obtain an attenuated current frame of the multichannel input signal. Again, this may be done in a similar manner for the sequence of frames of the multichannel input signal. Furthermore, the downmix processing unit may be configured to generate the current frame of the downmix signal's frame sequence from the attenuated current frame of the multichannel input signal. The downmix signal's frame sequence may be generated in a similar manner.
[0041] The audio processing system may further include a parameter processing unit configured to determine a sequence of spatial metadata frames from a multi-channel input signal. The sequence of frames of the downmix signal and the sequence of spatial metadata frames may be used to generate a multi-channel upmix signal containing n channels, the multi-channel upmix signal being an approximation of the multi-channel input signal. Furthermore, the audio processing system may include a bitstream generation unit configured to generate a bitstream showing a sequence of clipping protection gains, a sequence of frames of the downmix signal, and a sequence of spatial metadata frames, so that a corresponding decoding system can generate a multi-channel upmix signal.
[0042] The clipping protection gain curve may include a transition segment that provides a smooth transition from a preceding clipping protection gain to the current clipping protection gain, and a flat segment that remains flat at the current clipping protection gain. The transition segment may extend through a predetermined number of samples of the current frame of the multichannel input signal. The predetermined number of samples may be greater than 1 and less than the total number of samples in the current frame of the multichannel input signal. In particular, the predetermined number of samples may correspond to blocks of samples (where a frame may contain multiple blocks) or to frames. In specific examples, a frame may have 1536 samples, and a block may have 256 samples.
[0043] In a further aspect, an audio encoding system is described configured to generate a bitstream showing a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal. The system may have a downmix processing unit configured to generate the downmix signal from a multi-channel input signal. Furthermore, the system may have a parameter processing unit configured to determine a sequence of frames of spatial metadata for a corresponding sequence of frames of the multi-channel input signal.
[0044] Furthermore, the audio encoding system may have a configuration unit configured to determine one or more control settings for a parameter processing unit based on one or more external settings. The one or more external settings may include an update cycle indicating a time period for which the corresponding decoding system is required to synchronize with the bitstream. The configuration unit may be configured to determine, based on the update cycle, one or more independent frames of spatial metadata from a sequence of frames of spatial metadata to be encoded independently.
[0045] Another aspect describes a method for generating a bitstream that includes a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal. The method may include a step of generating the downmix signal from a multi-channel input signal. Furthermore, the method may include a step of determining one or more control settings based on one or more external settings, the one or more external settings including a target data rate for the bitstream, and the one or more control settings including a maximum data rate for the spatial metadata. Furthermore, the method may include a step of determining spatial metadata from a multi-channel input signal according to the control settings.
[0046] In a further aspect, a method is described for determining a spatial metadata frame for generating a frame of a multi-channel upmix signal from corresponding frames of a downmix signal. The method includes the step of determining a plurality of spectra from the current frame and the immediately following frame of a channel of a multi-channel input signal. Furthermore, the method may include the step of weighting the plurality of spectra using a window function to obtain a plurality of weighted spectra. Furthermore, the method may include the step of determining the spatial metadata frame for the current frame of the channel of the multi-channel input signal based on the plurality of weighted spectra. The window function may depend on one or more of the number of sets of spatial parameters contained in the spatial metadata frame, the presence of transient components in the current frame or the immediately following frame of the multi-channel input signal, and / or the time points of the transient components.
[0047] In a further aspect, a method is described for determining spatial metadata frames for generating frames of a multi-channel upmix signal from corresponding frames of a downmix signal. The method may include determining a first set of transformation coefficients from frames of a first channel of a multi-channel input signal and determining a second set of transformation coefficients from corresponding frames of a second channel of the multi-channel input signal. As outlined above, the first and second sets of transformation coefficients typically provide first and second time / frequency representations of the corresponding frames of the first and second channels, respectively. The first and second time / frequency representations may include multiple frequency bins and multiple time bins. The set of spatial parameters may include corresponding band parameters for different frequency bands, each containing a different number of frequency bins. The method may further include determining a shift to be applied when determining a particular band parameter for a particular frequency band using fixed-point arithmetic. The shift may be determined based on the particular frequency band. Furthermore, the shift may be determined based on the number of time bins to be considered in determining the particular band parameter. Furthermore, the method may include determining the specific band parameters using fixed-point arithmetic and determined shifts based on the first and second conversion coefficients that fall within the specific frequency band.
[0048] A method for generating a bitstream based on a multichannel input signal is described. The method may include a step of generating a sequence of frames of a downmix signal from a corresponding sequence of frames of a first multichannel input signal. Furthermore, the method may include a step of determining a sequence of spatial metadata frames from a sequence of frames of a second multichannel input signal. The sequence of frames of the downmix signal and the sequence of spatial metadata frames may be for generating a multichannel upmix signal. Furthermore, the method may include a step of generating the bitstream, which includes a sequence of bitstream frames. The bitstream frames may represent the downmix signal frames corresponding to the first frames of the sequence of frames of the first multichannel input signal, and the spatial metadata frames corresponding to the second frames of the sequence of frames of the second multichannel input signal. The second frames may be different from the first frames.
[0049] In a further aspect, a method for generating a bitstream based on a multichannel input signal is described. The method may include a step of determining a sequence of clipping protection gains for a corresponding sequence of frames in the multichannel input signal. The current clipping protection gain may indicate the attenuation to be applied to the current frame of the multichannel input signal to prevent clipping of the corresponding current frame of the downmix signal. The method may proceed to interpolate the current clipping protection gain with the preceding clipping protection gain of the preceding frame of the multichannel input signal to obtain a clipping protection gain curve. Furthermore, the method may include a step of applying the clipping protection gain curve to the current frame of the multichannel input signal to obtain an attenuated current frame of the multichannel input signal. A current frame of a sequence of frames in the downmix signal may be generated from the attenuated current frame of the multichannel input signal. Furthermore, the method may include a step of determining a sequence of spatial metadata frames from the multichannel input signal. The sequence of frames in the downmix signal and the sequence of spatial metadata frames may be used to generate a multichannel upmix signal. To enable the generation of the multi-channel upmix signal based on the bitstream, the bitstream may be generated such that it contains a sequence of clipping protection gains, a sequence of downmix signal frames, and a sequence of spatial metadata frames.
[0050] In a further aspect, a method for generating a bitstream showing a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal is described. The method may include the step of generating the downmix signal from a multi-channel input signal. Furthermore, the method may include the step of determining one or more control settings based on one or more external settings. The one or more external settings may include an update cycle indicating a time period for which a corresponding decoding system is required to synchronize with the bitstream. The method may further include the step of determining a sequence of frames of spatial metadata for a corresponding sequence of frames of the multi-channel input signal according to the control settings. Furthermore, the method may include encoding one or more frames of spatial metadata from the sequence of frames of spatial metadata as independent frames according to the update cycle.
[0051] In a further aspect, a software program is described. The software program may be adapted for execution on a processor to perform the method steps outlined in this paper when executed on said processor.
[0052] In another aspect, a storage medium is described. The storage medium may have a software program adapted for execution on a processor to perform the method steps outlined in this paper when executed on the processor.
[0053] In a further respect, a computer program product is described. A computer program may contain executable instructions for performing the method steps outlined in this paper when executed on a computer.
[0054] It should be noted that the methods and systems, including preferred embodiments outlined in this patent application, can be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this patent application can be combined in any way. In particular, the features of the claims may be combined with each other in any way. [Brief explanation of the drawing]
[0055] The present invention is described below in an illustrative manner with reference to the accompanying drawings. [Figure 1] This is a generalized block diagram of an exemplary audio processing system for performing spatial synthesis. [Figure 2] This diagram shows an example of the system in Figure 1. [Figure 3] Similar to Figure 1, this figure shows an exemplary audio processing system for performing spatial synthesis. [Figure 4] This figure shows an exemplary audio processing system for performing spatial decomposition. [Figure 5a] This is a block diagram of an exemplary parametric multichannel audio encoding system. [Figure 5b] This is a block diagram of an exemplary spatial decomposition and encoding system. [Figure 5c] This figure shows an exemplary time-frequency representation of a frame in a multi-channel audio signal. [Figure 5d] This figure shows an exemplary time-frequency representation of multiple channels in a multi-channel audio signal. [Figure 5e] This figure shows an example of windowing applied by the conversion unit of the spatial decomposition and encoding system shown in Figure 5b. [Figure 6] This is a flowchart illustrating an exemplary method for reducing the data rate of spatial metadata. [Figure 7a]This figure shows an exemplary transition scheme for spatial metadata performed in a decoding system. [Figure 7b] This figure shows an exemplary window function applied for determining spatial metadata. [Figure 7c] This figure shows an exemplary window function applied for determining spatial metadata. [Figure 7d] This figure shows an exemplary window function applied for determining spatial metadata. [Figure 8] This is a block diagram of an exemplary processing path for a parametric multichannel codec system. [Figure 9a] This is a block diagram of an exemplary parametric multi-channel audio encoding system configured to perform clipping protection and / or dynamic range control. [Figure 9b] This is a block diagram of an exemplary parametric multi-channel audio encoding system configured to perform clipping protection and / or dynamic range control. [Figure 10] This figure shows an example method for compensating for DRC parameters. [Figure 11] This figure shows an exemplary interpolation curve for clipping protection. [Modes for carrying out the invention]
[0056] As outlined in the introduction, this paper concerns a multi-channel audio coding system that utilizes a parametric multi-channel representation. Below, an exemplary multi-channel audio coding and decoding (codec) system is described. In the context of Figures 1 to 3, the decoder of the audio codec system is described as using the received parametric multi-channel representation to generate an n-channel upmix signal Y (typically n>2) from an received m-channel downmix signal X (e.g., m=2). Subsequently, the encoder-related processing of the multi-channel audio codec system is described. In particular, how the parametric multi-channel representation and the m-channel downmix signal can be generated from an n-channel input signal is described.
[0057] Figure 1 shows a block diagram of an exemplary audio processing system 100 configured to generate an upmix signal Y from a downmix signal X and a set of mixing parameters. In particular, the audio processing system 100 is configured to generate the upmix signal based solely on the downmix signal X and the set of mixing parameters. From the bitstream P, the audio decoder 140 generates the downmix signal X = [l0r0] Tand extract a set of mixing parameters. In the illustrated example, the set of mixing parameters includes parameters α1, α2, α3, β1, β2, β3, g, k1, k2. The mixing parameters may be contained in a quantized and / or entropy-encoded form within each mixing parameter data field in the bitstream P. These mixing parameters may also be referred to as metadata (or spatial metadata), which is transmitted together with the encoded downmix signal X. In some examples of this disclosure, it is explicitly shown that several connecting lines are adapted to transmit multichannel signals, where these lines are given adjacent cross lines for each number of channels. In system 100 shown in Figure 1, the downmix signal X contains m=2 channels, and the upmix signal Y, defined below, contains n=6 channels (e.g., 5.1 channels).
[0058] An upmix stage 110, which has an action that is parametrically dependent on the mixing parameters, receives the downmix signal. A downmix correction processor 120 corrects the downmix signal by nonlinear processing and by forming a linear combination of downmix channels, thereby correcting the downmix signal D=[d1d2] T The first mixing matrix 130 receives the downmix signal X and the modified downmix signal D, and by forming the following linear combination, the upmix signal Y = [l f l s r f r s [c lfe] T Outputs.
[0059]
number
[0060] The contributions of the modified downmix signal to the spatially left and right channels in the upmix signal may be controlled separately by parameters β1 (the contribution of the first modified channel to the left channel) and β2 (the contribution of the second modified channel to the right channel). Furthermore, the contributions of each channel in the downmix signal to its spatially corresponding channel in the upmix signal may be individually controlled by changing an independent mixing parameter g. Preferably, the gain parameter g is non-uniformly quantized to avoid large quantization errors.
[0061] Referring further to Figure 2, the downmix correction processor 120 may perform the following linear combination of downmix channels (which is cross-mixing) in the second mixing matrix 121.
[0062]
number
[0063] Figure 3 shows a first mixing matrix 130 of a similar type to that shown in Figure 1, along with its associated transform stages 301, 302, and inverse transform stages 311, 312, 313, 314, 315, and 316. These transform stages may have filter banks, for example, a quadrature mirror filter bank (QMF). Thus, signals located upstream of transform stages 301 and 302 are time-domain representations, as are signals located downstream of inverse transform stages 311, 312, 313, 314, 315, and 316. Other signals are frequency-domain representations. The time dependence of other signals may be represented, for example, as discrete values or blocks of values relating to the time blocks into which the signal is segmented. Note that Figure 3 uses an alternative notation compared to the matrix equations above. For example, X L0 ~l0, X R0 ~r0, Y L ~l f , Y Ls ~l s It can have correspondences such as the above. Furthermore, the notation in Figure 3 is the time-domain representation X of the signal. L0 (t) Frequency domain representation of the same signal X L0 This emphasizes the distinction between (f) and (f). The frequency domain representation is segmented into time frames and is therefore understood to be a function of both time and frequency variables.
[0064] Figure 4 shows an audio processing system 400 for generating a downmix signal X and mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 that control the gain applied by the upmix stage 110. This audio processing system 400 is typically located on the encoder side, for example, in a broadcast or recording facility. System 100 in Figure 1, on the other hand, is typically deployed on the decoder side, for example, in a playback facility. The downmix stage 410 generates an m-channel signal X based on an n-channel signal Y. Preferably, the downmix stage 410 acts on the time-domain representation of these signals. A parameter extractor 420 may generate the values of the mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 by analyzing the n-channel signal Y and taking into account the quantitative and qualitative attributes of the downmix stage 410. The mixing parameters may be a vector of values in frequency blocks, as suggested by the notation in Figure 4, and may further be segmented into time blocks. In one exemplary implementation, the downmix stage 410 is time-invariant and / or frequency-invariant. Thanks to this time and / or frequency invariance, typically there is no need for a communication connection between the downmix stage 410 and the parameter extractor 420, and parameter extraction may proceed independently. This provides a significant degree of flexibility for implementation. It also offers the possibility of reducing the overall latency of the system, as several processing stages may be executed in parallel. As an example, the Dolby Digital Plus format (or Enhanced AC-3) may be used to encode the downmix signal X.
[0065] The parameter extractor 420 may have knowledge of the quantitative and / or qualitative attributes of the downmix stage 410 by accessing the downmix specification. The downmix specification may specify one of the following: a set of gain values, an index that identifies a predefined downmix mode in which the gain is predefined. The downmix specification may also be a data record preloaded into memory in each of the downmix stage 410 and the parameter extractor 420. Alternatively or additionally, the downmix specification may be transmitted from the downmix stage 410 to the parameter extractor 420 through a communication line connecting these units. As a further alternative, each of the downmix stage 410 and the parameter extractor 420 may access the downmix specification in a common data source, such as memory in the audio processing system (for example, in the configuration setting unit 520 shown in Figure 5a), or in a metadata stream associated with the input signal Y.
[0066] FIG. 5a shows an exemplary multi-channel encoding system 500 that encodes a multi-channel audio input signal Y 561 (including n channels) using a downmix signal X (including m channels, where m<n) and a parametric representation. System 500 includes a downmix encoding unit 510 having the downmix stage 410 of FIG. 4, for example. The downmix encoding unit 510 may be configured to provide an encoded version of the downmix signal X. The downmix encoding unit 510 may use, for example, a Dolby Digital Plus encoder to encode the downmix signal X. Furthermore, system 500 includes a parameter encoding unit 520, which may include the parameter extractor 420 of FIG. 4. The parameter encoding unit 520 may be configured to quantize and encode a set of mixing parameters α1, α2, α3, β1, β2, β3, g, k1 (also referred to as spatial parameters) to provide encoded spatial parameters 562. As indicated above, parameter k2 may be determined from parameter k1. Furthermore, system 500 may include a bitstream generation unit 530 configured to generate a bitstream P 564 from the encoded downmix signal 563 and from the encoded spatial parameters 562. The bitstream 564 may be encoded according to a predetermined bitstream syntax. In particular, the bitstream 564 may be encoded in a format compliant with Dolby Digital Plus (DD+ or E-AC-3, Enhanced AC-3).
[0067] The system 500 may have a configuration setting unit 540 configured to determine one or more control settings 552, 554 for the parameter coding unit 520 and / or the downmix coding unit 510. The one or more control settings 552, 554 may be determined based on one or more external settings 551 of the system 500. For example, the one or more external settings may include the overall (maximum or fixed) data rate of the bitstream 564. The configuration setting unit 540 may be configured to determine one or more control settings 552 depending on the one or more external settings 551. The one or more control settings 552 for the parameter coding unit 520 may include one or more of the following:
[0068] • The maximum data rate for encoded spatial metadata 562. This control setting is referred to as the metadata data rate setting in this paper. The maximum and / or specific number of parameter sets to be determined by the parameter coding unit 520 per frame of the audio signal 561. This control setting is allowed to affect the temporal resolution of the spatial parameters and is therefore referred to in this paper as the temporal resolution setting. The number of frequency bands for which the spatial parameters should be determined by the parameter coding unit 520. This control setting is referred to as the frequency resolution setting because it allows for an effect on the frequency resolution of the spatial parameters. • The resolution of the quantizer that should be used to quantize spatial parameters. This control setting is referred to as the quantizer setting in this paper.
[0069] The parameter coding unit 520 may use one or more of the control settings 552 described above to determine and / or encode the spatial parameters to be included in the bitstream 564. Typically, the input audio signal Y 561 is segmented into a sequence of frames, where each frame contains a predetermined number of samples of the input audio signal Y 561. The metadata data rate setting may indicate the maximum number of bits available to encode the spatial parameters of the frames of the input audio signal 561. The actual number of bits used to encode the spatial parameters 562 of the frames may be less than the number of bits allocated by the metadata data rate setting. The parameter coding unit 520 may be configured to inform the configuration setting unit 540 of the number of bits actually used 553, so that the configuration setting unit 540 can determine the number of bits available to encode the downmix signal X. This number of bits may be communicated to the downmix encoding unit 510 as a control setting 554. The downmix encoding unit 510 may be configured to encode the downmix signal X based on the control setting 554 (for example, using a multi-channel encoder such as Dolby Digital Plus). Thus, bits that were not used to encode the spatial parameters may be used to encode the downmix signal.
[0070] Figure 5b shows a block diagram of an exemplary parameter coding unit 520. The parameter coding unit 520 may have a transform unit 521 configured to determine the frequency representation of an input signal 561. In particular, the transform unit 521 may be configured to transform frames of the input signal 561 into one or more spectra. Each spectrum contains multiple frequency bins. For example, the transform unit 521 may be configured to apply a filter bank, for example, a QMF filter bank, to the input signal 561. The filter bank may be a critically sampled filter bank. The filter bank may have a predetermined number of Q filters (for example, Q=64 filters). Thus, the transform unit 521 may be configured to determine Q subband signals from the input signal 561, where each subband signal is associated with a corresponding frequency bin 571. For example, frames of K samples of the input signal 561 may be transformed into Q subband signals, each having K / Q frequency coefficients per subband signal. In other words, frames of K samples of the input signal 561 may be transformed into K / Q spectra. Here, each spectrum has Q frequency bins. In a particular example, the frame length is K=1536, the number of frequency bins is Q=64, and the number of spectra is K / Q=24.
[0071] The parameter coding unit 520 may have a banding unit 522 configured to group one or more frequency bins 571 into frequency bands 572. The grouping of frequency bins 571 into frequency bands 572 may depend on a frequency resolution setting 552. Table 1 shows an exemplary mapping of frequency bins 571 into frequency bands 572, where the mapping may be applied by the banding unit 522 based on the frequency resolution setting 552. In the illustrated example, the frequency resolution setting 552 may result in banding of frequency bins 571 into 7, 9, 12, or 15 frequency bands. Banding typically models the psychoacoustic behavior of the human ear. As a result, the number of frequency bins 571 per frequency band 572 typically increases with increasing frequency.
[0072] [Table 1] The parameter determination unit 523 (particularly the parameter extractor 420) of the parameter coding unit 520 may be configured to determine one or more sets of mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for each of the frequency bands 572. For this reason, the frequency bands 572 are sometimes also referred to as parameter bands. The mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for the frequency bands 572 are sometimes referred to as band parameters. Thus, the complete set of mixing parameters typically includes the band parameters for each frequency band 572. The band parameters may be applied in the mixing matrix 130 in Figure 3 to determine the subband versions of the decoded upmixed signal.
[0073] The number of sets of mixed parameters per frame to be determined by the parameter determination unit 523 may be indicated by the time resolution setting 552. For example, the time resolution setting 552 may indicate that one or more sets of mixed parameters are determined per frame.
[0074] The determination of a set of mixed parameters, including bandwidth parameters for multiple frequency bands 572, is shown in Figure 5c. Figure 5c shows an exemplary set of transformation coefficients 580 derived from frames of the input signal 561. The transformation coefficients 580 correspond to a particular time point 582 and a particular frequency bin 571. The frequency band 572 may include multiple transformation coefficients 580 from one or more frequency bins 571. As can be seen from Figure 5c, the transformation of time-domain samples of the input signal 561 provides a time-frequency representation of frames of the input signal 561.
[0075] It should be noted that the set of mixed parameters for the current frame may be determined based on the transformation coefficient of 580 for the current frame, as well as the transformation coefficient of 580 for the immediately following frame (also known as the look-ahead frame).
[0076] The parameter determination unit 523 may be configured to determine the mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for each frequency band 572. When the temporal resolution setting is set to 1, all the conversion coefficients 580 (of the current frame and look-ahead frame) for a particular frequency band 572 may be considered to determine the mixing parameters for that particular frequency band 572. On the other hand, the parameter determination unit 523 may be configured to determine two sets of mixing parameters per frequency band 572 (for example, when the temporal resolution setting is set to 2). In this case, the first half of the time frame of the conversion coefficients 580 for that particular frequency band 572 (for example, corresponding to the conversion coefficients 580 for the current frame) may be used to determine the first set of mixing parameters, and the second half of the time frame of the conversion coefficients 580 for that particular frequency band 572 (for example, corresponding to the conversion coefficients 580 for the look-ahead frame) may be used to determine the second set of mixing parameters.
[0077] In general terms, the parameter determination unit 523 may be configured to determine one or more sets of mixed parameters based on conversion coefficients 580 of the current frame and look-ahead frame. A window function may be used to define the effect of the conversion coefficients 580 on the one or more sets of mixed parameters. The shape of the window function may depend on the number of sets of mixed parameters per frequency bandwidth 572 and / or attributes of the current frame and / or look-ahead frame (e.g., the presence of one or more transient components). Exemplary window functions are described in the context of Figures 5e and 7b through 7d.
[0078] It should be noted that the above may apply when the frame of the input signal 561 does not contain transient signal portions. The system 500 (for example, the parameter determination unit 523) may be configured to perform transient detection based on the input signal 561. If one or more transient components are detected, one or more transient indicators 583, 584 may be set, where the transient indicators 583, 584 may identify the time point 582 of the corresponding transient component. The transient indicators 583, 584 may be referred to as sampling points for each set of mixed parameters. In the case of transient components, the parameter determination unit 523 may be configured to determine the set of mixed parameters based on a conversion coefficient 580 starting from the time point of the transient component (this is shown by different shaded areas in Figure 5c). On the other hand, conversion coefficients 580 prior to the time point of the transient component are ignored, thereby ensuring that the set of mixed parameters reflects the multi-channel situation after the transient component.
[0079] Figure 5c shows the transformation coefficient 580 for one channel of a multichannel input signal Y 561. A parameter coding unit 520 is typically configured to determine the transformation coefficient 580 for multiple channels of the multichannel input signal 561. Figure 5d shows exemplary transformation coefficients for the first 561-1 and second 561-2 channels of the input signal 561. The frequency band p 572 contains frequency bins 571 in the range of frequency indices i to j. The transformation coefficient 580 for the first channel 561-1 in frequency bin i at time (or spectrum) q is a q,i It may also be called. Similarly, the transformation coefficient 580 of the second channel 561-2 in frequency bin i at time (or spectrum) q is b q,i It may also be called the conversion coefficient 580. The conversion coefficient 580 may be a complex number. The determination of the mixing parameters for the frequency band p may involve the determination of the energy and / or covariance of the first and second channels 561-1, 561-2 based on the conversion coefficient 580. As an example, the covariance of the conversion coefficient 580 for the first and second channels 561-1, 561-2 for the time interval [q,v] in the frequency band p is:
number
number
[0080] Therefore, the parameter determination unit 523 may be configured to determine one or more sets 573 of band parameters for various frequency bands 572. The number of frequency bands 572 typically depends on the frequency resolution setting 552, and the number of sets of mixed parameters per frame typically depends on the time resolution setting 552. For example, the frequency resolution setting 552 may instruct the use of 15 frequency bands 572, and the time resolution setting 552 may instruct the use of 2 sets of mixed parameters. In this case, the parameter determination unit 523 may be configured to determine two temporally distinct sets of mixed parameters, where each set of mixed parameters includes 15 sets 573 of band parameters (i.e., mixed parameters for various frequency bands 572).
[0081] As shown above, the mixing parameter for the current frame may be determined based on the conversion coefficient 580 of the current frame and the conversion coefficient 580 of the subsequent look-ahead frame. The parameter determination unit 523 may apply a window to the conversion coefficient 580 to ensure a smooth transition between the mixing parameters of successive frames in the frame sequence and / or to take into account abrupt portions (e.g., transient components) in the input signal 561. This is shown in Figure 5e, which shows K / Q spectra 589 of the current frame 585 and the immediately following frame 590 of the input audio signal 561 at the corresponding K / Q successive time points 582. Furthermore, Figure 5e shows an exemplary window 586 used by the parameter determination unit 523. The window 586 reflects the influence of K / Q spectra 589 of the current frame 585 and the immediately following frame 590 (referred to as the look-ahead frame) on the mixing parameter. As will be outlined in more detail later, window 586 reflects the case where the current frame 585 and the look-ahead frame 590 contain no transient components. In this case, window 586 ensures smooth phase-in and phase-out of the spectra 589 of the current frame 585 and the look-ahead frame 590, respectively, thereby allowing for a smooth evolution of the spatial parameters. Furthermore, Figure 5e shows exemplary windows 587 and 588. The dashed window 587 reflects the influence of the K / Q spectra 589 of the current frame 585 on the mixture parameters of the preceding frame. Furthermore, the dashed window 588 reflects the influence of the K / Q spectra 589 of the following frame 590 on the mixture parameters of the following frame 590 (in the case of smooth interpolation).
[0082] One or more sets of mixed parameters may then be quantized and encoded using the encoding unit 524 of the parameter coding unit 520. The encoding unit 524 may apply various encoding schemes. For example, the encoding unit 524 may be configured to perform differential encoding of mixed parameters. Differential encoding may be based on a temporal difference (between the current mixed parameter and a preceding corresponding mixed parameter for the same frequency band 572) or on a frequency difference (between the current mixed parameter in a first frequency band 572 and the corresponding current mixed parameter in an adjacent second frequency band 572).
[0083] Furthermore, the encoding unit 524 may be configured to quantize a set of mixed parameters and / or a temporal or frequency difference of mixed parameters. The quantization of the mixed parameters may depend on the quantizer setting 552. For example, the quantizer setting 552 may take two values: a first value indicating fine quantization and a second value indicating coarse quantization. Thus, the encoding unit 524 may be configured to perform fine quantization (with relatively low quantization error) or coarse quantization (with relatively increased quantization error) based on the quantization type indicated by the quantizer setting 552. The quantized parameters or parameter differences may then be encoded using an entropy-based code such as a Huffman code. As a result, an encoded spatial parameter 562 is obtained. The number of bits 553 used for the encoded spatial parameter 562 may be communicated to the configuration setting unit 540.
[0084] In one embodiment, the encoding unit 524 may be configured to first quantize various mixing parameters (taking into account the quantizer setting 552) and provide quantized mixing parameters. The quantized mixing parameters may then be entropy-coded (for example, using Huffman coding). Entropy coding may encode the quantized mixing parameters of a frame (without considering preceding frames), the frequency difference of the quantized mixing parameters, or the temporal difference of the quantized mixing parameters. Encoding of the temporal difference may not be used in the case of so-called independent frames that are encoded independently of preceding frames.
[0085] Therefore, the parameter encoding unit 520 may utilize a combination of differential coding and Huffman coding to determine the encoded spatial parameters 562. As outlined above, the encoded spatial parameters 562 may be included in the bitstream 564 as metadata (also referred to as spatial metadata) together with the encoded downmix signal 563. Differential coding and Huffman coding may be used for transmitting the spatial metadata to reduce redundancy and thus increase the spare bitrate available for encoding the downmix signal 563. Since Huffman coding is a variable-length coding, the size of the spatial metadata can vary considerably depending on the statistics of the encoded spatial parameters 562 to be transmitted. The data rate required to transmit the spatial metadata is deducted from the data rate available to the core codec (e.g., Dolby Digital Plus) for encoding the stereo downmix signal. To avoid compromising the audio quality of the downmix signal, the number of bytes that may be spent per frame for transmitting spatial metadata is typically limited. This limitation may be subject to encoder tuning considerations, which may be taken into account by the configuration setting unit 540. However, due to the variable-length nature of the underlying differential / Huffman coding of the spatial parameters, it is typically not possible to guarantee, without further measures, that the data rate limit (e.g., reflected in the metadata data rate setting 552) will not be exceeded.
[0086] This paper describes a method for post-processing encoded spatial parameters 562 and / or spatial metadata containing encoded spatial parameters 562. Method 600 for post-processing spatial metadata is described in the context of Figure 6. Method 600 may be applied when it is determined that the total size of one frame of spatial metadata exceeds a predefined limit, for example, indicated by the metadata data rate setting 552. Method 600 is directed toward reducing the amount of metadata step by step. Reducing the size of spatial metadata typically also reduces the precision of the spatial metadata, thus impairing the quality of the spatial image of the audio signal being reproduced. However, Method 600 typically ensures that the total amount of spatial metadata does not exceed a predefined limit, and thus allows for a determined improved trade-off between spatial metadata (for regenerating the m-channel multi-channel signal) and audio codec metadata (for decoding the encoded downmix signal 563) in terms of overall audio quality. Furthermore, method 600 for post-processing spatial metadata can be implemented with relatively low computational complexity (compared to a complete recalculation of the encoded spatial parameters using the modified control settings 552).
[0087] Method 600 for post-processing spatial metadata includes one or more of the following steps: As outlined above, a spatial metadata frame may contain multiple (e.g., one or two) parameter sets per frame, and the use of additional parameter sets allows for increased temporal resolution of the mixed parameters. The use of multiple parameter sets per frame can improve audio quality, especially for signals with a lot of attack (i.e., transients). Even for audio signals with a fairly slowly changing spatial image, spatial parameter updates using a grid with twice the density of sampling points can improve audio quality. However, transmitting multiple parameter sets per frame leads to approximately a doubling of the data rate. Therefore, if it is determined that the data rate for spatial metadata exceeds the metadata data rate setting 552 (step 601), it may be checked whether the spatial metadata frame contains two or more sets of mixed parameters. In particular, it may be checked whether the metadata frame contains two sets of mixed parameters that are assumed to be transmitted (step 602). If it is determined that the spatial metadata contains multiple sets of mixed parameters, one or more sets exceeding a single set of mixed parameters may be discarded (step 603). As a result, the data rate for spatial metadata can be significantly reduced (typically by half in the case of two sets of mixed parameters) while maintaining a relatively low degree of audio quality degradation.
[0088] The decision of which of the two (or more) sets of mixed parameters to discard may depend on whether the encoding system 500 has detected a transient position ("attack") in the portion of the input signal 561 currently covered by the frame. If multiple transient components exist in the current frame, earlier transient components are more important than later ones due to the psychoacoustic post-masking effect of all individual attacks. Therefore, if transient components are present, it may be prudent to discard the later set of mixed parameters (e.g., the second of the two sets). On the other hand, if there is no attack, the earlier set of mixed parameters (e.g., the first of the two sets) may be discarded. This may be due to the windowing used when calculating the spatial parameters (shown in Figure 5e). The window 586 used to window out the portion of the input signal 561 used to calculate the spatial parameters for the second set of mixed parameters typically has the greatest impact at the point when the upmix stage 130 places the sampling point for parameter reconstruction (i.e., at the end of the current frame). On the other hand, the first set of mixing parameters typically has a half-frame offset relative to this point in time. As a result, the error that can be made by dropping the first set of mixing parameters is very likely to be lower than the error that can be made by dropping the second set of mixing parameters. This is shown in Figure 5e, where it can be seen that the latter half of the spectrum 589 of current frame 585, used to determine the second set of mixing parameters, is influenced more to a greater extent by the samples of current frame 585 than the first half of the spectrum 589 of current frame 585 (the window function 586 has a lower value for the first half of the spectrum 589 than for the second half).
[0089] The spatial cue (i.e., mixed parameter) computed in the encoding system 500 is transmitted to the corresponding decoder 100 via bitstream 562 (which may be part of bitstream 564 on which the encoded stereo downmix signal 563 is carried). Between the computation of the spatial cue and its representation in bitstream 562, the encoding unit 524 typically applies a two-stage encoding approach: the first stage of quantization is a lossy stage as it adds an error to the spatial cue; the second stage of difference / Huffman coding is a lossless stage. As outlined above, the encoder 500 can choose between various types of quantization (e.g., two types of quantization): a high-resolution quantization scheme that adds a relatively small error but gives a larger number of potential quantization indices, and a low-resolution quantization scheme that adds a relatively large error but gives a smaller number of quantization indices and thus does not require a large Huffman codeword. It should be noted that different types of quantization may be applicable to some or all of the mixed parameters. For example, different types of quantization may be applicable to the mixed parameters α1, α2, α3, β1, β2, β3, and k1. On the other hand, the gain g may be quantized using a fixed type of quantization.
[0090] Method 600 may include a step 604 to verify which type of quantization was used to quantize the spatial parameters. If it is determined that a relatively fine quantization resolution was used, the encoding unit 524 may be configured to reduce the quantization resolution to a lower type of quantization 605. As a result, the spatial parameters are quantized once again. However, this does not add significant computational overhead (compared to re-determining the spatial parameters using different control settings 552). It should be noted that different types of quantization may be used for different spatial parameters α1, α2, α3, β1, β2, β3, g, k1. Therefore, the encoding unit 524 may be configured to individually select a quantization resolution for each type of spatial parameter, thereby adjusting the data rate of the spatial metadata.
[0091] Method 600 may include a step (not shown in Figure 6) of reducing the frequency resolution of the spatial parameters. As outlined above, the set of mixed parameters in a frame is typically clustered into frequency bands or parameter bands 572. Each parameter band represents a certain frequency range, and for each band, a separate set of spatial cues is determined. Depending on the data rate available for transmitting the spatial metadata, the number of parameter bands 572 may be varied in steps (e.g., 7, 9, 12, or 15 bands). The number of parameter bands 572 has a nearly linear relationship with the data rate, so a reduction in frequency resolution can significantly reduce the data rate of the spatial metadata. Audio quality, on the other hand, is only moderately affected. However, such a reduction in frequency resolution typically requires recalculating the set of mixed parameters using the modified frequency resolution, thus increasing the computational cost.
[0092] As outlined above, the encoding unit 524 may utilize differential encoding of (quantized) spatial parameters. The configuration unit 551 may be configured to impose direct encoding of the spatial parameters of frames of the input audio signal 561 in order to ensure that transmission errors do not propagate across an unlimited number of frames and to enable the decoder to synchronize with the bitstream 562 received at intermediate points in time. Thus, a certain percentage of the frames may not utilize differential encoding along the timeline. Such frames that do not utilize differential encoding may be referred to as independent frames. Method 600 may include a step 606 for verifying whether a current frame is an independent frame and / or whether an independent frame is a forced independent frame. The encoding of spatial parameters may depend on the result of step 606.
[0093] As outlined above, differential coding is typically designed to compute differences between temporally consecutive elements or between neighboring frequency bands of quantized spatial cues. In either case, the statistics of the spatial cues are such that small differences occur more frequently than large differences, and thus small differences are represented by shorter Huffman codewords than large differences. In this paper, it is proposed to perform (time- or frequency-based) smoothing of the quantized spatial parameters. Smoothing spatial parameters over time or frequency typically yields smaller differences, thus leading to a reduction in data rate. Due to psychoacoustic considerations, temporal smoothing is usually preferred over frequency smoothing. If it is determined that the current frames are not forced independent frames, method 600 may proceed to perform temporal differential encoding (step 607), possibly in combination with temporal smoothing. On the other hand, if it is determined that the current frames are independent frames, method 600 may proceed to perform frequency differential encoding (step 608) and possibly frequency-based smoothing.
[0094] The differential encoding in step 607 may be subjected to a smoothing process over time to reduce the data rate. The degree of smoothing may vary depending on the amount by which the data rate should be reduced. The most stringent type of temporal "smoothing" corresponds to preserving the original, unchanged set of mixed parameters, which corresponds to transmitting only delta values equal to 0. Temporal smoothing of the differential encoding may be performed on one or more (e.g., all) spatial parameters.
[0095] Similar to temporal smoothing, frequency smoothing may also be performed. In its most extreme form, frequency smoothing corresponds to transmitting the same quantized spatial parameters over the full frequency range of the input signal 561. While ensuring that the limits set by the metadata data rate setting are not exceeded, frequency smoothing can have a relatively large impact on the quality of the spatial image that can be reconstructed using the spatial metadata. Therefore, it may be preferable to apply frequency smoothing only when temporal smoothing is not permitted (for example, when the current frame is a forced independent frame for which time-difference coding for the preceding frame should not be used).
[0096] As outlined above, the system 500 may be operated according to one or more external settings, such as the overall target data rate of the bitstream 564 or the sampling rate of the input audio signal 561. Typically, there is no single optimal operating point for all combinations of external settings. The configuration setting unit 540 may be configured to map valid combinations of external settings 551 to combinations of control settings 552, 554. For example, the configuration setting unit 540 may rely on the results of a psychoacoustic listening test. In particular, the configuration setting unit 540 may be configured to determine a combination of control settings 552, 554 that guarantees (on average) the best psychoacoustic coding result for a particular combination of external settings 551.
[0097] As outlined above, the decoding system 100 needs to be able to synchronize with the received bitstream 564 within a given time period. To ensure this, the encoding system 500 may periodically encode so-called independent frames, i.e., frames that do not depend on knowledge of preceding frames. The average distance in frames between two independent frames may be given by the ratio of a given maximum time delay for synchronization to the duration of one frame. This ratio does not necessarily have to be an integer. The distance between two independent frames is always an integer number of frames.
[0098] The encoding system 500 (e.g., configuration setting unit 540) may be configured to receive a maximum time delay for synchronization or a desired update time period as an external setting 551. Furthermore, the encoding system 500 (e.g., configuration setting unit 540) may have a timer module configured to track the absolute amount of time elapsed since the first encoded frame of the bitstream 564. The first encoded frame of the bitstream 564 is, by definition, an independent frame. The encoding system 500 (e.g., configuration setting unit 540) may be configured to determine whether the next frame to be encoded has a sample corresponding to a point in time that is an integer multiple of the desired update cycle. Whenever the next frame to be encoded has a sample at a point in time that is an integer multiple of the desired update cycle, the encoding system 500 (e.g., configuration setting unit 540) may be configured to ensure that the next frame to be encoded is encoded as an independent frame. This ensures that the desired update time period is maintained, even if the ratio of the desired update time period to the frame length is not an integer.
[0099] As outlined above, the parameter determination unit 523 is configured to compute spatial cues based on the time / frequency representation of the multi-channel input signal 561. The frames of spatial metadata may be determined based on K / Q (e.g., 24) spectra 589 (QMF spectra) of the current frame and / or based on K / Q (e.g., 24) spectra 589 (QMF spectra) of the look-ahead frame, where each spectrum 589 may have a frequency resolution of Q (e.g., 64) frequency bins 571. Depending on whether the encoding system 500 detects transient components in the input signal 561, the temporal length of the signal portion used to compute a single set of spatial cues may have a different number of spectra 589 (e.g., from 1 spectrum to 2 x K / Q spectra). As shown in Figure 5c, each spectrum 589 is divided into a number of frequency bands 572 (e.g., 7, 9, 12, or 15 frequency bands). These frequency bands, for psychoacoustic reasons, contain a different number of frequency bins 571 (e.g., from one frequency bin to 41 frequencies). Different frequency bands p 572 and different temporal segments [q,v] define a grid on the time / frequency representation of the current and look-ahead frames of the input signal 561. For different "squares" in this grid, different sets of spatial cues may be calculated based on estimates of the energy and / or covariance of at least some of the input channels within each of those different "squares". As outlined above, the energy estimates and / or covariances may be calculated by summing the squares of the transformation coefficients 580 for one channel and / or by summing the products of the transformation coefficients 580 for different channels (as shown by the formulas given above). The different transformation coefficients 580 may be weighted according to a window function 586 used to determine the spatial parameters.
[0100] Energy estimate E 1,1 (p), E 2,2 (p) and / or covariance E 1,2The calculation of (p) may be performed using fixed-point arithmetic. In this case, different sizes of the time / frequency grid "mesh" may affect the arithmetic precision of the values determined for the spatial parameters. As outlined above, the number of frequency bins 571 per frequency band 572 (j-i+1) and / or the length of the time interval [q,v] of the time / frequency grid "mesh" can vary considerably (e.g., between 1×1×2 and 48×41×2 conversion coefficients 580 (e.g., the real and imaginary parts of the complex QMF coefficients)). As a result, the energy E 1,1 (p) / covariance E 1,2 The product Re{a} needs to be summed to determine (p) t,f Re{b t,f} and Im{a t,f}Im{b t,f The number of} can vary significantly. To prevent the result of the above calculation from exceeding the range of numbers that can be represented by fixed-point arithmetic, the signal is defined by the maximum number of bits (for example, 2). 6 ·2 6 It may be scaled down by 6 bits (because = 4096 ≥ 48 · 41 · 2). However, this approach leads to a significant decrease in arithmetic precision for smaller "mass" and / or "mass" that have only relatively low signal energy.
[0101] This paper proposes using individual scaling for each "mesh" of the time / frequency grid. Each scaling may depend on the number of transformation coefficients 580 contained within the "mesh" of the time / frequency grid. Typically, the spatial parameters for a particular "mesh" of the time / frequency grid (i.e., for a particular frequency band 572 and a particular time interval [q,v]) are determined solely on the transformation coefficients 580 from that particular "mesh" (and not on the transformation coefficients 580 from other "mesh"). Furthermore, the spatial parameters are typically determined solely on the ratio of energy estimates and / or covariances (and not typically influenced by absolute energy estimates and / or covariances). In other words, a single spatial cue typically uses only the energy estimate and / or cross-channel product from a single time / frequency "mesh." Furthermore, the spatial cue is typically not influenced by absolute energy estimates / covariances, but only by the ratio of energy estimates / covariances. Therefore, it is possible to use individual scaling for every single "mesh." This scaling should be consistent for channels that contribute to specific spatial cues.
[0102] Energy estimates E for the first and second channels 561-1 and 561-2 for the frequency band p 572 and time interval [q,v] 1,1 (p), E 2,2 (p) and the covariance E between the first and second channels 561-1, 561-2 1,2 (p) may be determined, for example, by the formula shown above. The energy estimate and covariance are given by the scaling factor s p The scaled energy and covariances are scaled by s p ·E 1,1 (p), s p ·E 2,2 (p) and s p ·E 1,2 (p) may be given. Energy estimate E 1,1 (p), E 2,2 (p) and covariance E 1,2The spatial parameter P(p) derived based on (p) typically depends on the ratio of energy and / or covariance, and thus the value of the spatial parameter P(p) is determined by the scaling factor s p This is independent of the above. As a result, different scaling factors s are obtained for different frequency bands p, p+1, and p+2. p , s p+1 , s p+2 It may be used.
[0103] It should be noted that one or more spatial parameters may depend on more than two different input channels (e.g., three different channels). In this case, the one or more spatial parameters depend on the energy estimates E of those different channels. 1,1 (p), E 2,2 (p) ...based on the respective covariances between different pairs of those channels, i.e., E 1,2 (p), E 1,3 (p), E 2,3 (p) may be derived based on the above. In this case, the values of one or more spatial parameters are independent of the scaling factors applied to the energy estimate and / or covariance.
[0104] In particular, z p Let be a positive integer indicating a shift in fixed-point arithmetic, and for a specific frequency band p, the scaling factor s p =2 -zp but 0.5 p ·max{|E 1,1 (p)|,|E 2,2 (p)|,|E 1,2 (p)|}≦1.0 And shift z p It may be determined such that it is minimized. By individually ensuring this for each frequency band p and / or each time interval [q,v] in which the mixing parameters are determined, increased (e.g., maximum) precision in fixed-point arithmetic can be achieved while guaranteeing a valid range of values.
[0105] For example, individual scaling can be implemented by checking whether the result of any single MAC (multiply-accumulate) operation can exceed ±1. Only if so, the individual scaling for that "mass" may be increased by one bit. Once this has been done for all channels, the maximum scaling for each "mass" may be determined, and all deviant scalings for "mass" may be applied accordingly.
[0106] As outlined above, spatial metadata may include one or more (e.g., two) sets of spatial parameters per frame. Thus, the encoding system 500 may transmit one or more sets of spatial parameters per frame to the corresponding decoding system 100. Each of these sets of spatial parameters corresponds to one specific spectrum among K / Q temporally consecutive spectra 289 of the frame of spatial metadata. This specific spectrum corresponds to a specific time point, which may be referred to as a sampling point. Figure 5c shows two exemplary sampling points 583, 584 for each of two sets of spatial parameters. The sampling points 583, 584 may be associated with specific events contained within the input audio signal 561. Alternatively, the sampling points may be predetermined.
[0107] Sampling points 583 and 584 indicate the point in time when the corresponding spatial parameters should be fully applied in the decoding system 100. In other words, the decoding system 100 may be configured to update the spatial parameters at sampling points 583 and 584 according to the transmitted set of spatial parameters. Furthermore, the decoding system 100 may be configured to interpolate the spatial parameters between two consecutive sampling points. The spatial parameters may indicate the type of transition performed between consecutive sets of spatial parameters. Examples of transition types are “smooth” and “steep” transitions between spatial parameters. These mean that the spatial parameters may be interpolated in a smooth (e.g., linear) manner or updated abruptly, respectively.
[0108] In the case of a "smooth" transition, the sampling point may be fixed (i.e., predetermined) and therefore does not need to be transmitted in bitstream 564. If the spatial metadata frame transmits a single set of spatial parameters, the predetermined sampling point may be located at the very end of the frame. That is, the sampling point may correspond to the K / Q-th spectral 589. If the spatial metadata transmits two sets of spatial parameters, the first sampling point may correspond to the K / 2Q-th spectral 589, and the second sampling point may correspond to the K / Q-th spectral 589.
[0109] In the case of a “steep” transition, sampling points 583 and 584 may be variable and may be signaled in bitstream 562. The location of bitstream 562 that carries information about the number of spatial parameters used in a given frame, information about the choice between a “smooth” transition and a “steep” transition, and information about the location of the sampling points in the case of a “steep” transition may be referred to as the “framing” portion of bitstream 562. Figure 7a shows an exemplary transition scheme that may be applied by the decoding system 100 depending on the framing information contained in the received bitstream 562.
[0110] For example, frame configuration information for a particular frame may indicate a single set 711 of “smooth” transitions and spatial parameters. In this case, the decoding system 100 (e.g., the first mixture matrix 130) may assume that the sampling points for the set of spatial parameters 711 correspond to the last spectrum of that particular frame. Furthermore, the decoding system 100 may be configured to interpolate (e.g., linearly) 701 between the last received set 710 of spatial parameters for the previous frame and the set 711 of spatial parameters for that particular frame. In another example, frame configuration information for a particular frame may indicate two sets 711, 712 of “smooth” transitions and spatial parameters. In this case, the decoding system 100 (e.g., the first mixture matrix 130) may assume that the sampling points for the first set of spatial parameters 711 correspond to the last spectrum of the first half of that particular frame, and the sampling points for the second set of spatial parameters 712 correspond to the last spectrum of the second half of that particular frame. Furthermore, the decoding system 100 may be configured to interpolate (for example linearly) 702 between the last received set 710 of spatial parameters for the previous frame and the set 711 of spatial parameters, and between the first set 711 of spatial parameters and the second set 712 of spatial parameters.
[0111] In one further example, the frame configuration information for a particular frame may indicate a “steep” transition, a single set of spatial parameters 711, and a sampling point 583 for the single set of spatial parameters 711. In this case, the decoding system 100 (e.g., the first mixture matrix 130) may be configured to apply the last received set of spatial parameters 710 for the previous frame up to the sampling point 583, and then apply the set of spatial parameters 711 starting from the sampling point 583 (as shown in curve 703). In another example, the frame configuration information for a particular frame may indicate a “steep” transition, two sets of spatial parameters 711, 712, and two corresponding sampling points 583, 584 for the two sets of spatial parameters 711, 712. In this case, the decoding system 100 (for example, the first mixture matrix 130) may be configured to apply the last received set 710 of spatial parameters for the previous frame up to the first sampling point 583, to apply the first set 711 of spatial parameters from the first sampling point 583 to the second sampling point 584, and to apply the second set 712 of spatial parameters from the second sampling point 584 to at least the end of that particular frame (as shown in curve 704).
[0112] The encoding system 500 should ensure that the frame configuration information matches the signal characteristics and that appropriate portions of the input signal 561 are selected to compute one or more sets of spatial parameters 711, 712. For this purpose, the encoding system 500 may have detectors configured to detect signal locations in one or more channels where the signal energy increases sharply. If at least one such signal location is found, the encoding system 500 may be configured to switch from a “smooth” transition to a “steep” transition; otherwise, the encoding system 500 may continue with a “smooth” transition.
[0113] As outlined above, the encoding system 500 (e.g., the parameter determination unit 523) may be configured to calculate spatial parameters for the current frame based on multiple frames 585, 590 of the input audio signal 561 (e.g., based on the current frame 585 and the immediately following frame 590, i.e., the so-called look-ahead frame). Thus, the parameter determination unit 523 may be configured to determine spatial parameters based on 2 x K / Q spectra 589 (as shown in Figure 5e). The spectra 589 may be windowed by a window 586 as shown in Figure 5e. In this paper, it is proposed to adapt the window 586 based on the number of sets of spatial parameters 711, 712 to be determined, based on the type of transition, and / or based on the location of the sampling points 583, 584. In this way, it can be ensured that the frame configuration information matches the signal characteristics and that appropriate portions of the input signal 561 are selected to compute the one or more sets of spatial parameters 711, 712.
[0114] The following describes exemplary window functions for various encoder / signal conditions.
[0115] a) Situation: Single set of spatial parameters 711, smooth transitions, no transient components within lookahead frames 590 Window function 586: Between the last spectrum of the previous frame and the K / Qth spectrum 589, the window function 586 may rise linearly from 0 to 1. Between the K / Qth spectrum and the 48th spectrum 589, the window function 586 may fall linearly from 1 to 0 (see Figure 5e).
[0116] b) Situation: A single set of spatial parameters 711, smooth transition, transient component in the Nth spectrum (N>K / Q), i.e., transient component within look-ahead frame 590 Window function 721, as shown in Figure 7b: Between the last spectrum of the previous frame and the K / Qth spectrum, window function 721 increases linearly from 0 to 1. Between the K / Qth spectrum and the (N-1)th spectrum, window function 721 remains constant at 1. Between the Nth spectrum and the 2*K / Qth spectrum, window function 586 remains constant at 0. The transient component in the Nth spectrum is represented by the transient point 724 (which corresponds to the sampling point for the set of spatial parameters of the immediately following frame 590). Furthermore, complementary window functions 722 (which are applied to the spectrum of the current frame 585 when determining the aforementioned set of spatial parameters for the previous frame) and window function 723 (which are applied to the spectrum of the immediately following frame 590 when determining the aforementioned set of spatial parameters for the immediately following frame) are shown in Figure 7b. Overall, the window function 721 ensures that, in the case of one or more transient components in the look-ahead frame 590, the spectrum of the look-ahead frame prior to the first transient point 724 is fully taken into account in determining the set of spatial parameters 711 for the current frame 585. On the other hand, the spectrum of the look-ahead frame 590 after the transient point 724 is ignored.
[0117] c) Situation: Single set of spatial parameters 711, steep transition, transient component in the Nth spectrum (N ≤ K / Q), no transient component in the immediately following frame 590. The window function 731 is as shown in Figure 7c: Between the first spectrum and the (N-1)th spectrum, the window function 731 remains constant at 0. Between the Nth spectrum and the K / Qth spectrum, the window function 731 remains constant at 1. Between the K / Qth spectrum and the 2*K / Qth spectrum, the window function 731 drops linearly from 1 to 0. Figure 7c shows the transient point 734 in the Nth spectrum (which corresponds to the sampling point for a single set 711 of spatial parameters). Furthermore, Figure 7c shows the window function 732 applied to the spectrum of the current frame 585 when determining the one or more sets of spatial parameters for the preceding frame, and the window function 733 applied to the spectrum of the following frame 590 when determining the one or more sets of spatial parameters for the following frame.
[0118] d) Situation: A single set of spatial parameters, steep transitions, transient components in the Nth and Mth spectra (N ≤ K / Q, M > K / Q) Window function 741 in Figure 7d: Between the first spectrum and the (N-1)th spectrum, the window function 741 remains constant at 0. Between the Nth spectrum and the (M-1)th spectrum, the window function 741 remains constant at 1. Between the Mth spectrum and the 48th spectrum, the window function remains constant at 0. Figure 7d shows the transient point 744 in the Nth spectrum (i.e., the sampling point of the set of spatial parameters) and the transient point 745 in the Mth spectrum. Furthermore, Figure 7d shows the window function 742 applied to the spectrum of the current frame 585 when determining the one or more sets of spatial parameters for the immediately preceding frame, and the window function 743 applied to the spectrum of the immediately following frame 590 when determining the one or more sets of spatial parameters for the immediately following frame.
[0119] e) Situation: Two sets of spatial parameters, smooth transition, no transient components in subsequent frames. Window function: i) First set of spatial parameters: between the last spectrum of the immediately preceding frame and the K / 2Q-th spectrum, the window function rises linearly from 0 to 1. between the K / 2Q-th spectrum and the K / Q-th spectrum, the window falls linearly from 1 to 0. between the K / Q-th spectrum and the 2*K / Q-th spectrum, the window remains constant at 0.
[0120] ii) Second set of spatial parameters: between the first spectrum and the K / 2Q-th spectrum, the window remains constant at 0. between the K / 2Q-th spectrum and the K / Q-th spectrum, the window rises linearly from 0 to 1. between the K / Q-th spectrum and the 3*K / 2Q-th spectrum, the window falls linearly from 1 to 0. between the 3*K / 2Q-th spectrum and the 2*K / Q-th spectrum, the window remains constant at 0.
[0121] f) Situation: two sets of spatial parameters, smooth transition, transient component at the N-th spectrum (N>K / Q) Window function: i) First set of spatial parameters: between the last spectrum of the immediately preceding frame and the K / 2Q-th spectrum, the window rises linearly from 0 to 1. between the K / 2Q-th spectrum and the K / Q-th spectrum, the window falls linearly from 1 to 0. between the K / Q-th spectrum and the 2*K / Q-th spectrum, the window remains constant at 0.
[0122] ii) Second set of spatial parameters: between the first spectrum and the K / 2Q-th spectrum, the window remains constant at 0. between the K / 2Q-th spectrum and the K / Q-th spectrum, the window rises linearly from 0 to 1. between the K / Q-th spectrum and the (N-1)-th spectrum, the window remains constant at 1. between the N-th spectrum and the 2*K / Q-th spectrum, the window remains constant at 0.
[0123] g) Situation: two sets of parameters, abrupt transition, transient components at the N-th spectrum and the M-th spectrum (N<M≦K / Q), no transient component in the subsequent frame Window function: i) The first set of spatial parameters: The window remains constant at 0 between the first spectrum and the (N-1)th spectrum. The window remains constant at 1 between the Nth spectrum and the (M-1)th spectrum. The window remains constant at 0 between the Mth spectrum and the 2*K / Qth spectrum.
[0124] ii) A second set of spatial parameters: The window remains constant at 0 between the first spectrum and the (M-1)th spectrum. Between the Mth spectrum and the K / Qth spectrum, the window remains constant at 1. Between the K / Qth spectrum and the 2*K / Qth spectrum, the window drops linearly from 1 to 0.
[0125] h) Situation: Two sets of spatial parameters, steep transition, transient components (N) in the Nth, Mth and Oth spectra.<M≦K / Q、O> K / Q) Window function: i) The first set of spatial parameters: The window remains constant at 0 between the first spectrum and the (N-1)th spectrum. The window remains constant at 1 between the Nth spectrum and the (M-1)th spectrum. The window remains constant at 0 between the Mth spectrum and the 2*K / Qth spectrum.
[0126] ii) A second set of spatial parameters: The window remains constant at 0 between the first spectrum and the (M-1)th spectrum. The window remains constant at 1 between the Mth spectrum and the (O-1)th spectrum. The window remains constant at 0 between the Oth spectrum and the 2*K / Qth spectrum.
[0127] Overall, the following exemplary rules may be established for window functions to determine the current set of spatial parameters.
[0128] ● When the current set of spatial parameters is not associated with the transient component The window function provides a smooth phase-in of the spectra from the sampling point of the immediately preceding set of spatial parameters to the sampling point of the current set of spatial parameters; • If the subsequent set of spatial parameters is not associated with transient components, the window function provides a smooth phase-out of the spectra from the sampling point of the current set of spatial parameters to the sampling point of the subsequent set of spatial parameters; If the subsequent set of spatial parameters is associated with transient components, the window function fully considers the spectra from the sampling point of the current set of spatial parameters to the spectrum before the sampling point of the subsequent set of spatial parameters, and cancels out the spectra starting from the sampling point of the subsequent set of spatial parameters.
[0129] ●When the current set of spatial parameters is associated with transient components The window function cancels out the spectra preceding the sampling point of the current set of spatial parameters; If the sampling point of the subsequent set of spatial parameters is associated with transient components, the window function fully considers the spectra from the sampling point of the current set of spatial parameters to the spectrum before the sampling point of the subsequent set of spatial parameters, and cancels out the spectra starting from the sampling point of the subsequent set of spatial parameters; If the subsequent set of spatial parameters is not associated with transient components, the window function fully considers the spectra from the sampling point of the current set of spatial parameters to the spectrum at the end of the current frame, providing a smooth phase-out of the spectra from the beginning of the look-ahead frame to the sampling point of the said subsequent set of spatial parameters.
[0130] The following describes a method for reducing latency in a parametric multichannel codec system having an encoding system 500 and a decoding system 100. As outlined above, the encoding system 500 has several processing paths, such as generating and encoding the downmix signal and determining and encoding the parameters. The decoding system 100 typically performs decoding of the encoded downmix signal and generating a decorrelated downmix signal. Furthermore, the decoding system 100 performs decoding of the encoded spatial metadata. Subsequently, in the first upmix matrix 130, the decoded spatial metadata is applied to the decoded downmix signal and the decorrelated downmix signal to generate an upmix signal.
[0131] It is desirable to provide an encoding system 500 configured to provide a bitstream 564 that enables the decoding system 100 to generate an upmix signal Y with reduced delay and / or reduced buffer memory. As outlined above, the encoding system 500 has several different paths through which the encoded data provided to the decoding system 100 in the bitstream 564 can be aligned so that it matches correctly at the time of decoding. As outlined above, the encoding system 500 performs downmixing and encoding of the PCM signal 561. Furthermore, the encoding system 500 determines spatial metadata from the PCM signal 561. Furthermore, the encoding system 500 may be configured to determine one or more clip gains (typically one clip gain per frame). The clip gains represent the clipping prevention gains applied to the downmix signal X to ensure that the downmix signal X is not clipped. The one or more clip gains may be transmitted in the bitstream 564 (typically in the spatial metadata frame) so that the decoding system 100 can regenerate the upmix signal Y. Furthermore, the encoding system 500 may be configured to determine one or more dynamic range control (DRC) values (for example, one or more DRC values per frame). The one or more DRC values may be used by the decoding system 100 to perform dynamic range control of the upmixed signal Y. In particular, the one or more DRC values may ensure that the DRC performance of the parametric multichannel codec system described herein is similar to (or equal to) the DRC performance of legacy multichannel codec systems such as Dolby Digital Plus. The one or more DRC values may be transmitted within the downmixed audio frame (for example, within the appropriate field of the Dolby Digital Plus bitstream).
[0132] Therefore, the encoding system 500 may have at least four signal processing paths. In order to align these four paths, the encoding system 500 may also take into account delays introduced into the system by various processing components not directly related to the encoding system 500, such as core encoder delay, core decoder delay, spatial metadata decoder delay, LFE filter delay (for filtering LFE channels), and / or QMF decomposition delay.
[0133] To align the various paths described above, the delay of the DRC processing path may be considered. The DRC processing delay may typically only be aligned to the frame, and not necessarily to each time sample. Therefore, the DRC processing delay typically depends only on the core encoder delay, which may be rounded up to the next frame alignment. That is, DRC processing delay = round up (core encoder delay / frame size). Based on this, the downmix processing delay for generating the downmix signal may be determined, since the downmix processing delay can be delayed to each time sample. That is, downmix processing delay = DRC delay × frame size - core encoder delay. The remaining delays can be calculated by summing the individual delay lines and ensuring that the delays match at the decoder stage. This is shown in Figure 8.
[0134] By considering various processing delays, when writing bitstream 564, delaying the resulting spatial metadata by one frame instead of delaying the encoded PCM data by 1536 samples (number of input channels × 1536 × 4 bytes - 245 bytes less memory) can reduce the processing power (number of input channels - 1 × 1536 fewer copy operations) and memory in the decoding system. As a result of the delay, all signal paths are not only precisely aligned by time samples but also roughly matched.
[0135] As outlined above, Figure 8 shows the various delays experienced by the exemplary encoding system 500. The numbers in parentheses in Figure 8 indicate exemplary delays in the number of samples of the input signal 561. The encoding system 500 typically has a delay 801 caused by filtering the LFE channel of the multi-channel input signal 561. Furthermore, a delay 802 (referred to as the "clipgainpcmdelayline") may be caused by determining the clipping gain (i.e., the DRC2 parameter described later) applied to the input signal 561 to prevent the downmix signal from being clipped. In particular, this delay 802 may be introduced to synchronize the application of clipping gain in the encoding system 500 with the application of clipping gain in the decoding system 100. For this purpose, the input to the downmix calculation (performed by the downmix processing unit 510) may be delayed by an amount equal to the delay 811 (referred to as the "coredecdelay") of the decoder 140 of the downmix signal. This means that, in the illustrated example, clipgainpcmdelayline=coredecdelay=288 samples.
[0136] The downmix processing unit 510 (for example, having a Dolby Digital Plus encoder) delays the processing path of audio data, i.e., the downmix signal, but does not delay the processing path for spatial metadata and DRC / clip gain data. As a result, the downmix processing unit 510 should delay the calculated DRC gain, clip gain, and spatial metadata. For DRC gain, this delay typically needs to be an integer multiple of one frame. The delay of the DRC delay line 807 (referred to as "drcdelayline") can be calculated as drcdelayline = ceil((corencdelay + clipgainpcmdelayline) / frame_size) = 2 frames, where "coreencdelay" refers to the encoder delay 810 of the downmix signal.
[0137] The delay of the DRC gain can typically only be an integer multiple of the frame size. Therefore, an additional delay may need to be added in the downmix processing path to compensate for this and round it to the next integer multiple of the frame size. The additional downmix delay 806 (referred to as "dmxdelayline") may be determined by dmxdelayline + coreencdelay + clipgainpcmdelayline = drcdelayline * frame_size, where dmxdelayline = drcdelayline * frame_size - coreencdelay - clipgainpcmdelayline, resulting in dmxdelayline = 100.
[0138] When spatial parameters are applied in the frequency domain (e.g., in the QMF domain) on the decoder side, the spatial parameters should be synchronized with the downmix signal. To compensate for the fact that the encoder of the downmix signal does not delay the spatial metadata frame but delays the downmix processing path, the input to the parameter extractor 420 should be delayed such that the following condition holds: dmxdelayline + coreencdelay + coredecdelay + aspdecanadelay = aspdelayline + qmfanadelay + framingdelay. In the above formula, "qmfanadelay" [QMF decomposition delay] specifies a delay 804 caused by the conversion unit 521, and "framingdelay" [frame structuring delay] specifies a delay 805 caused by windowing of the conversion coefficient 580 and the determination of the spatial parameters. As outlined above, the frame structuring calculation uses two frames as input: the current frame and the look-ahead frame. For the look-ahead, the frame structuring introduces a delay 805 of exactly one frame length. Furthermore, since delay 804 is known, the additional delay to be applied to the processing path to determine spatial metadata is aspdelayline=dmxdelayline+coreencdelay+coredecdelay+aspdecanadelay-qmfanadelay-framingdelay=1856. Since this delay is greater than one frame, the memory size of the delay line can be reduced by delaying the computed bitstream instead of delaying the input PCM data. Thus, aspbsdelayline=floor(aspdelayline / frame_size)=1 frame (delay 809) and asppcmdelayline=aspdelayline-aspbsdelayline*frame_size=320 (delay 803).
[0139] After calculating the one or more clipping gains, the one or more clipping gains are provided to the bitstream generation unit 530. Thus, the one or more clipping gains experience a delay applied to the final bitstream by aspbsdelayline 809. Therefore, the additional delay 808 for the clipping gains should be: clipgainbsdelayline + aspbsdelayline = dmxdelayline + coreencdelay + coredecdelay, which gives clipgainbsdelayline = dmxdelayline + coreencdelay + coredecdelay - aspbsdelayline = 1 frame. In other words, it should be ensured that the one or more clipping gains are provided to the decoding system 500 immediately after decoding the corresponding frame of the downmix signal. This allows the one or more clipping gains to be applied to the downmix signal before the upmix is performed in the upmix stage 130.
[0140] Figure 8 shows further delays experienced in the decoding system 100. These include, for example, delays 812 (referred to as "aspdecanadelay") caused by the time-domain to frequency-domain conversion 301, 302 of the decoding system 100, delays 813 (referred to as "aspdecsyndelay") caused by the frequency-domain to time-domain conversion 311 to 316, and further delays 814.
[0141] As can be seen from Figure 8, the various processing paths of the codec system have processing-related delays and alignment delays that ensure that the various output data from the various processing paths are available in the decoding system 100 when needed. Alignment delays (e.g., delays 803, 809, 807, 808, 806) are provided within the encoding system 500, thereby reducing the processing power and memory required in the decoding system 100. The total delays for the various processing paths (excluding the LFE filter delay 801 applicable to all processing paths) are as follows:
[0142] • Downmix processing path: delays 802, 806, 810 sum = 3072, i.e., 2 frames; • DRC processing path: delay 807 = 3072, i.e., 2 frames; • Clipping gain processing path: The sum of delays 808, 809, and 802 = 3360. This corresponds to the delay of the downmix signal decoder (811) plus the delay of the downmix processing path; • Spatial metadata processing path: The sum of delays 802, 803, 804, 805, and 809 is 4000. This corresponds to the delay 811 of the downmix signal decoder and the delay 812 caused by the time-domain to frequency-domain conversion stages 301 and 302, plus the delay of the downmix processing path.
[0143] Therefore, it is guaranteed that the DRC data is available in the decoding system 100 at time 821, the clip gain data is available at time 822, and the spatial metadata is available at time 823.
[0144] Furthermore, Figure 8 shows that the bitstream generation unit 530 may combine encoded audio data and spatial metadata, which may relate to different excerpts of the input audio signal 561. In particular, it can be seen that the downmix processing path, DRC processing path, and clip gain processing path have a delay of just 2 frames (3072 samples) (ignoring delay 801) to the output of the encoding system 500 (indicated by interfaces 831, 832, and 833). The encoded downmix signal is provided by interface 831, the DRC gain data is provided by interface 832, and the spatial metadata and clip gain data are provided by interface 833. Typically, the encoded downmix signal and DRC gain data are provided in a normal Dolby Digital Plus frame, while the clip gain data and spatial metadata may be provided in a spatial metadata frame (for example, in an auxiliary field of the Dolby Digital Plus frame).
[0145] The spatial metadata processing path in interface 833 has a delay of 4000 samples (ignoring delay 801), which can be seen as different from the delays of the other processing paths (3072 samples). This means that the spatial metadata frames may relate to different excerpts of the input signal 561 than to the frames of the downmix signal. In particular, to ensure alignment in the decoding system 100, the bitstream generation unit 530 should be configured to generate a bitstream 564 containing a sequence of bitstream frames. Here, the bitstream frames represent the downmix signal frames corresponding to the first frame of the multichannel input signal 561 and the spatial metadata frames corresponding to the second frame of the multichannel input signal 561. The first and second frames of the multichannel input signal 561 may contain the same number of samples. Nevertheless, the first and second frames of the multichannel input signal 561 may be different from each other. In particular, the first and second frames may correspond to different excerpts of the multichannel input signal 561. More specifically, the first frame may include samples that precede the samples in the second frame. For example, the first frame may include samples of the multi-channel input signal 561 that precede the samples in the second frame of the multi-channel input signal 561 by a predetermined number of samples, for example, 928 samples.
[0146] As outlined above, the encoding system 500 may be configured to determine dynamic range control (DRC) and / or clip gain data. In particular, the encoding system 500 may be configured to ensure that the downmix signal X is not clipped. Furthermore, the encoding system 500 may be configured to provide dynamic range control (DRC) parameters that ensure the DRC behavior of a multichannel signal Y encoded using the parametric encoding scheme described above is similar to or equal to the DRC behavior of a multichannel signal Y encoded using a reference multichannel encoding system (such as Dolby Digital Plus).
[0147] Figure 9a is a block diagram of an exemplary dual-mode encoding system 900. It should be noted that parts 930 and 931 of the dual-mode encoding system 900 are typically provided separately. An n-channel input signal Y 561 is provided to each of the upper part 930, which is active in at least the multi-channel encoding mode of the encoding system 900, and the lower part 931, which is active in at least the parametric encoding mode of the encoding system 900. The lower part 931 of the encoding system 900 may correspond to, or include, an encoding system 500, for example. The upper part 930 may correspond to a reference multi-channel encoder (such as a Dolby Digital Plus encoder). The upper part 930 generally has a discrete-mode DRC analyzer 910 arranged in parallel with an encoder 911, both of which receive the audio signal Y 561 as input. Based on this input signal 561, the encoder 911 outputs an encoded n-channel signal (Y with ^). Meanwhile, the DRC analyzer 910 outputs one or more post-processing DRC parameters DRC1 that quantify the decoder-side DRC to be applied. The DRC parameters DRC1 may be "compr" gain (compressor gain) and / or "dynrng" gain (dynamic range gain) parameters. The parallel outputs from both units 910 and 911 are collected by a discrete-mode multiplexer 912, which outputs a bitstream P. The bitstream P may have a predetermined syntax, for example, the Dolby Digital Plus syntax.
[0148] A lower portion 931 of encoding system 900 includes a parametric analysis stage 922 arranged in parallel with a parametric mode DRC analyzer 921. The parametric mode DRC analyzer 921, similarly to the parametric analysis stage 922, receives an n-channel input signal Y. The parametric analysis stage 922 may include a parameter extractor 420. Based on the n-channel audio signal Y, the parametric analysis stage 922 outputs one or more mixing parameters collectively represented by α in FIGS. 9a and 9b (as outlined above) and an m-channel (1<m<n) downmix signal X. The downmix signal X is then processed by a core signal encoder 923 (e.g., a Dolby Digital Plus encoder), which outputs an encoded downmix signal (X with a caret symbol) based thereon. The parametric analysis stage 922 applies dynamic range limitation to time blocks or frames of the input signal when this may be necessary. A possible condition for controlling when to apply dynamic range limitation can be a "non-clipping condition", that is, an "in-range condition". This implies that in time blocks or frame segments where the downmix signal has a large amplitude, the signal is processed so as to fall within a defined range. This condition may be implemented based on one time block or one time frame including several time blocks. By way of example, a frame of the input signal 561 may include a predetermined number (e.g., six) of blocks. Preferably, the above condition is implemented by applying a wide-spectrum gain reduction, rather than clipping only peak values or using a similar approach.
[0149] Figure 9b shows a possible implementation of the parametric decomposition stage 922, which includes a preprocessor 927 and a parametric decomposition processor 928. The preprocessor 927 is responsible for performing dynamic range limiting on the n-channel input signal 561, thereby outputting a dynamic range-limited n-channel signal, which is fed to the parametric decomposition processor 928. The preprocessor 527 also outputs block-by-block or frame-by-frame values of the preprocessing DRC parameter DRC2. The parameter DRC2, along with the mixed parameter α and the m-channel downmix signal X from the parametric decomposition processor 928, is included in the output from the parametric decomposition stage 922.
[0150] The parameter DRC2 may also be referred to as the clipping gain. The parameter DRC2 may represent the gain applied to the multi-channel input signal 561 to ensure that the downmix signal X is not clipped. One or more channels of the downmix signal X may be determined from the channels of the input signal Y by determining a linear combination of some or all of the channels of the input signal Y. For example, the input signal Y may be a 5.1 multi-channel signal, and the downmix signal may be a stereo signal. Samples of the left and right channels of the downmix signal may be generated based on different linear combinations of samples from the 5.1 multi-channel input signal.
[0151] The DRC2 parameter may be determined so that the maximum amplitude of the channels in the downmix signal does not exceed a predetermined threshold. This may be guaranteed on a block-by-block or frame-by-frame basis. A single gain (clip gain) on a block-by-block or frame-by-frame basis may be applied to the channels of the multi-channel input signal Y to ensure that the above conditions are met. The DRC2 parameter may represent this gain (for example, its reciprocal).
[0152] Referring to Figure 9a, it should be noted that the discrete-mode DRC analyzer 910 functions similarly to the parametric-mode DRC analyzer 921 in that it outputs one or more post-processing DRC parameters DRC1 that quantify the decoder-side DRC to be applied. Thus, the parametric-mode DRC analyzer 921 may be configured to simulate the DRC processing performed by the reference multi-channel encoder 930. The parameter DRC1 provided by the parametric-mode DRC analyzer 921 is typically not included in the bitstream P in parametric coding mode, but instead is compensated to take into account the dynamic range limiting performed by the parametric decomposition stage 922. For this purpose, the DRC up-compensator 924 receives the post-processing DRC parameter DRC1 and the pre-processing DRC parameter DRC2. For each block or frame, the DRC up-compensator 924 derives one or more compensated post-processing DRC parameter DRC3 values. These post-processing DRC parameters are such that the combined effect of the compensated post-processing DRC parameter DRC3 and the pre-processing DRC parameter DRC2 is quantitatively equivalent to the DRC quantified by the post-processing DRC parameter DRC1. In other words, the DRC up-compensator 924 is configured to reduce the post-processing DRC parameters output by the DRC analyzer 921 by the portion already performed by the parametric decomposition stage 922. The compensated post-processing DRC parameter DRC3 may be included in the bitstream P.
[0153] Referring to the lower part 931 of system 900, the parametric mode multiplexer 925 collects the compensated post-processing DRC parameter DRC3, the pre-processing DRC parameter DRC2, the mixing parameter α, and the encoded downmix signal X, and forms a bitstream P based on them. Thus, the parametric mode multiplexer 925 may include or correspond to the bitstream generation unit 530. In one possible implementation, the compensated post-processing DRC parameter DRC3 and the pre-processing DRC parameter DRC2 may be encoded in logarithmic form as dB values that affect amplitude upscaling or downscaling on the decoder side. The compensated post-processing DRC parameter DRC3 may have any sign. However, the post-processing DRC parameter DRC2 resulting from implementations such as “non-clipping conditions” is typically represented by a non-negative dB value at all time points.
[0154] Figure 10 shows an exemplary process that may be performed, for example, in a parametric mode DRC analyzer 921 and a DRC up compensator 924 to determine the modified DRC parameters DRC3 (for example, modified “dynrng gain” and “compr gain” parameters).
[0155] The DRC2 and DRC3 parameters may be used to ensure that the decoding system reproduces different audio bitstreams at a consistent loudness level. Furthermore, it may be ensured that the bitstream generated by the parametric encoding system 500 has a consistent loudness level with respect to bitstreams generated by legacy and / or reference encoding systems (such as Dolby Digital Plus). As outlined above, this can be ensured by the encoding system 500 generating a clipped downmix signal (using the DRC2 parameter) and by providing the DRC2 parameter (e.g., the reciprocal of the attenuation applied to prevent clipping of the downmix signal) within the bitstream so that the decoding system 100 can regenerate the original loudness (when generating the upmix signal).
[0156] As outlined above, the downmix signal is typically generated based on a linear combination of some or all of the channels of the multichannel input signal 561. Therefore, the scaling factor (or attenuation) applied to the channels of the multichannel input signal 561 may depend on all the channels of the multichannel input signal 561 that contributed to the downmix signal. In particular, one or more of the channels of the downmix signal may be determined based on the LFE channel of the multichannel input signal 561. Consequently, the scaling factor (or attenuation) applied for clipping protection should also take the LFE channel into consideration. This differs from other multichannel encoding systems (such as Dolby Digital Plus) where the LFE channel is typically not considered for clipping protection. By taking into account the LFE channel and / or all channels that contributed to the downmix signal, the quality of clipping protection can be improved.
[0157] Therefore, one or more DRC2 parameters provided to the corresponding decoding system 100 may depend on all channels of the input signal 561 that contributed to the downmix signal. In particular, the DRC2 parameters may depend on the LFE channel. Doing so may improve the quality of clipping protection.
[0158] It should be noted that the dialnorm parameter does not necessarily have to be taken into account for the calculation of the scaling factor and / or DRC2 parameter (as shown in Figure 10).
[0159] As outlined above, the encoding system 500 may be configured to write a so-called "clip gain" (i.e., a DRC2 parameter) in a spatial metadata frame, indicating what gain was applied to the input signal 561 to prevent clipping in the downmix signal. The corresponding decoding system 100 may be configured to precisely cancel out the clip gain applied in the encoding system 500. However, only the sampling points of the clip gain are transmitted in the bitstream. In other words, the clip gain parameter is typically determined only on a frame-by-frame or block-by-block basis. The decoding system 100 may be configured to interpolate the clip gain value (e.g., the received DRC2 parameter) between neighboring sampling points.
[0160] An exemplary interpolation curve for interpolating DRC2 parameters for adjacent frames is shown in Figure 11. In particular, Figure 11 shows a first DRC2 parameter 953 for the first frame and a second DRC2 parameter 954 for the subsequent second frame 950. The decoding system 100 may be configured to interpolate between the first DRC2 parameter 953 and the second DRC2 parameter 954. The interpolation may be performed within a subset 951 of the samples of the second frame 950, for example, within the first block 951 of the second frame 950 (as shown by the interpolation curve 952). The interpolation of DRC2 parameters ensures a smooth transition between adjacent audio frames, thereby avoiding audible artifacts that may be caused by differences between successive DRC2 parameters 953 and 954.
[0161] The encoding system 500 (in particular, the downmix processing unit 510) may be configured to apply a corresponding clip gain interpolation to the DRC2 interpolation 952 performed by the decoding system 500 when generating the downmix signal. This ensures that clip gain protection for the downmix signal is consistently removed when generating the upmix signal. In other words, the encoding system 500 may be configured to simulate the curve of DRC2 values resulting from the DRC2 interpolation 952 applied by the decoding system 100. Furthermore, the encoding system 500 may be configured to apply the exact (sample-by-sample) reciprocal of this curve of DRC2 values to the multichannel input signal 561 when generating the downmix signal.
[0162] The methods and systems described in this paper may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software running on, for example, a digital signal processor or microprocessor. Other components may be implemented as hardware and / or as application-specific integrated circuits. Signals encountered in the methods and systems described may be stored on a medium such as random-access memory or optical storage media. These signals may be transmitted over a network such as a radio network, satellite network, wireless network, or wired network, such as the Internet. Typical devices utilizing the methods and systems described in this paper are portable electronic devices or other consumer equipment that store and / or render audio signals.
[0163] Several aspects are described below. [Aspect 1] An audio encoding system configured to generate a bitstream showing a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal: - A downmix processing unit (510) configured to generate the downmix signal from a multichannel input signal, wherein the downmix signal has m channels, and the multichannel input signal has n channels, where n and m are integers, and m <nである、ダウンミックス処理ユニットと;A parameter processing unit (520) configured to determine the spatial metadata from the multi-channel input signal; A configuration setting unit (540) configured to determine one or more control settings for the parameter processing unit based on one or more external settings, wherein the one or more external settings include a target data rate for the bitstream, and the one or more control settings include a maximum data rate for the spatial metadata, Audio encoding system. [Aspect 2] The parameter processing unit is configured to determine spatial metadata for the frames of the multi-channel input signal, referred to as spatial metadata frames; The frame of the multi-channel input signal includes a predetermined number of samples of the multi-channel input signal; The maximum data rate for the spatial metadata indicates the maximum number of metadata bits for the spatial metadata frame. The audio encoding system according to Embodiment 1. [Aspect 3] The audio encoding system according to embodiment 2, wherein the parameter processing unit is configured to determine whether the number of bits in a spatial metadata frame determined based on one or more control settings exceeds the maximum number of metadata bits. [Aspect 4] • A spatial metadata frame contains one or more sets of spatial parameters; The one or more control settings include a temporal resolution setting that indicates the number of sets of spatial parameters per spatial metadata frame to be determined by the parameter processing unit; The parameter processing unit is configured to discard the set of spatial parameters (711) from the current spatial metadata frame if the current spatial metadata frame has a set of multiple spatial parameters (711, 712) and the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits. The audio encoding system according to embodiment 3. [Aspect 5] The set of spatial parameters is associated with one or more corresponding sampling points; The one or more sampling points indicate the corresponding one or more time points; The parameter processing unit is configured to discard a first set of spatial parameters (711) from the current spatial metadata frame if the plurality of sampling points (583, 584) of the current metadata frame are not associated with transient components of the multi-channel input signal, wherein the first set of spatial parameters is associated with a first sampling point (583) prior to a second sampling point (584); The parameter processing unit is configured to discard a second set (712) of spatial parameters from the current spatial metadata frame if the plurality of sampling points in the current metadata frame are associated with transient components of the multi-channel input signal. The audio encoding system described in Embodiment 4. [Aspect 6] The one or more control settings include a quantizer setting that indicates a first type of quantizer from a plurality of predetermined types of quantizers; The parameter processing unit is configured to quantize one or more sets of spatial parameters according to the first type of quantizer; The aforementioned plurality of predetermined types of quantizers each provide a different quantizer resolution; The parameter processing unit is configured to requantize one, some, or all of the spatial parameters of one or more sets of spatial parameters according to a second type quantizer having a lower resolution than the first type quantizer, if it is determined that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits. The audio encoding system according to embodiment 4 or 5. [Aspect 7] The audio encoding system according to embodiment 6, wherein the plurality of predetermined types of quantizers include fine quantization and coarse quantization. [Aspect 8] The parameter processing unit is: The set of temporal difference parameters is determined based on the difference between the current set of spatial parameters (712) and the previous set of spatial parameters (711); Encode the aforementioned set of time difference parameters using entropy coding; - Insert the encoded set of temporal difference parameters into the current spatial metadata frame; If it is determined that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits, the entropy of the set of temporal difference parameters is reduced. An audio encoding system according to any one of embodiments 4 to 7, configured as described above. [Aspect 9] The audio encoding system according to embodiment 8, wherein the parameter processing unit is configured to set one, some, or all of the time difference parameters in the set of time difference parameters to a value that has an increased probability of being a possible value of the time difference parameter, in order to reduce the entropy of the set of time difference parameters. [Aspect 10] The one or more control settings mentioned above include frequency resolution settings; The frequency resolution setting indicates the number of different frequency bands; The parameter processing unit is configured to determine different spatial parameters, referred to as band parameters, for different frequency bands; The set of spatial parameters includes the corresponding band parameters for the different frequency bands. An audio encoding system according to any one of the descriptions in paragraphs 4 through 9. [Aspect 11] The parameter processing unit is A set of frequency difference parameters is determined based on the difference between one or more band parameters in a first frequency band and one or more corresponding band parameters in a second, adjacent frequency band; Encode the set of frequency difference parameters using entropy coding; - Insert the encoded set of frequency difference parameters into the current spatial metadata frame; - If it is determined that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits, the entropy of the set of frequency difference parameters is reduced. An audio encoding system according to embodiment 10, configured as described above. [Aspect 12] The audio encoding system according to embodiment 11, wherein the parameter processing unit is configured to set one, some, or all of the frequency difference parameters in the set of frequency difference parameters to a value that has an increased probability of being a possible value of the frequency difference parameter, in order to reduce the entropy of the set of frequency difference parameters. [Aspect 13] The parameter processing unit, If it is determined that the number of bits in the current spatial metadata frame exceeds the aforementioned maximum number of metadata bits, the number of frequency bands shall be reduced; - Redetermine one or more sets of spatial parameters for the current spatial metadata frame using a reduced number of frequency bands. An audio encoding system according to any one of embodiments 10 to 12, configured as described above. [Aspect 14] The one or more external settings further include: the sampling rate of the multi-channel input signal, the number of channels m of the downmix signal, the number of channels n of the multi-channel input signal, and an update cycle indicating a time period for which the corresponding decoding system is required to synchronize with the bitstream; The one or more control settings further include: a temporal resolution setting indicating the number of sets of spatial parameters per frame of spatial metadata to be determined; a frequency resolution setting indicating the number of frequency bands for which spatial parameters should be determined; a quantizer setting indicating the type of quantizer to be used to quantize the spatial metadata; and an indication of whether the current frame of the multi-channel input signal should be encoded as an independent frame. An audio encoding system according to any one of the descriptions in 1 to 13. [Aspect 15] The one or more external settings further include an update cycle indicating a time period for which the corresponding decoding system is required to synchronize with the bitstream; The one or more control settings further include an indicator of whether the current spatial metadata frame should be encoded as an independent frame; The parameter processing unit is configured to determine the sequence of spatial metadata frames for the corresponding sequence of frames in the multi-channel input signal; The configuration unit is configured to determine, based on the update cycle, one or more spatial metadata frames to be encoded as independent frames from the sequence of spatial metadata frames. An audio encoding system according to any one of the descriptions in Actuals 2 to 14. [Aspect 16] The aforementioned configuration setting unit is - Determine whether the current frame of the sequence of frames of the multi-channel input signal contains a sample at a point in time that is an integer multiple of the update cycle; • Determine if the current spatial metadata frame corresponding to the current frame is an independent frame. An audio encoding system according to embodiment 15, configured as described above. [Aspect 17] The audio encoding system according to embodiment 15, wherein the parameter processing unit is configured to encode one or more sets of spatial parameters of the current spatial metadata frame independently of the data contained in previous spatial metadata frames, when the current spatial metadata frame should be encoded as an independent frame. [Aspect 18] n=6 and m=2; and / or • The multi-channel upmix signal is a 5.1 signal; and / or • The downmix signal is a stereo signal; and / or The aforementioned multi-channel input signal is a 5.1 signal. An audio encoding system according to any one of the descriptions in Actuals 1 to 17. [Aspect 19] The downmix processing unit is configured to encode the downmix signal using a Dolby Digital Plus encoder; The bitstream is compatible with Dolby Digital Plus bitstream; The spatial metadata is contained within the data field of the Dolby Digital Plus bitstream. An audio encoding system according to any one of the descriptions in Actuals 1 to 18. [Aspect 20] The spatial metadata includes one or more sets of spatial parameters; • A spatial parameter in the set of spatial parameters indicates the cross-correlation between different channels of the multi-channel input signal. An audio encoding system according to any one of the descriptions in 1 to 19. [Aspect 21] A parameter processing unit (520) is configured to determine a spatial metadata frame for generating a frame of a multi-channel upmix signal from corresponding frames of a downmix signal, wherein the downmix signal has m channels and the multi-channel upmix signal has n channels, where n and m are integers, and m <nであり、前記空間的メタデータ·フレームは、空間的パラメータの一つまたは複数の集合を含み、当該パラメータ処理ユニットは、A conversion unit (521) configured to determine multiple spectra from the current frame and the immediately following frame of a channel in a multi-channel input signal; The parameter determination unit (523) is configured to determine the spatial metadata frame for the current frame of the channel of the multichannel input signal by weighting the plurality of spectra using a window function; The window function depends on: the number of spatial parameters included in the spatial metadata frame, the presence of one or more transient components in the current frame or the immediately following frame of the multi-channel input signal, and / or one or more time points of the transient components. Parameter processing unit. [Aspect 22] The aforementioned window function includes set-dependent window functions; The parameter determination unit is configured to determine a set of spatial parameters for the current frame of the channel of the multichannel input signal by weighting the multiple spectra using the set-dependent window function; The set-dependent window function depends on whether the set of spatial parameters is associated with transient components. A parameter processing unit according to embodiment 21. [Aspect 23] If the set of spatial parameters (711) is not associated with transient components, The set-dependent window function provides a phase-in of the plurality of spectra from the sampling point of the preceding set (710) of spatial parameters to the sampling point of the set (711) of spatial parameters; and / or If the subsequent set of spatial parameters (712) is associated with the transient component, the set-dependent window function includes the multiple spectra from the sampling point of the set of spatial parameters (711) to the spectrum of the multiple spectra prior to the sampling point of the subsequent set of spatial parameters (712), and cancels out the multiple spectra starting from the sampling point of the subsequent set of spatial parameters (712). A parameter processing unit according to embodiment 22. [Aspect 24] If the set of spatial parameters (711) is associated with transient components, The set-dependent window function cancels out spectra from the plurality of spectra prior to the sampling point of the set (711) of spatial parameters; and / or If the sampling point of the subsequent set (712) of spatial parameters is associated with a transient component, the set-dependent window function includes the spectrum from the plurality of spectra up to the spectrum of the plurality of spectra prior to the sampling point of the subsequent set (711) of spatial parameters, and cancels out the spectrum from the plurality of spectra starting from the sampling point of the subsequent set (712) of spatial parameters; and / or If the subsequent set (712) of spatial parameters is not associated with transient components, the set-dependent window function includes the spectrum of the plurality of spectra from the sampling point of the set (711) of spatial parameters up to the spectrum of the plurality of spectra at the end of the current frame (585), and provides a phase-out of the spectrum of the plurality of spectra from the beginning of the immediately following frame (590) up to the sampling point of the subsequent set (712) of spatial parameters. A parameter processing unit according to embodiment 22. [Aspect 25] A parameter processing unit (520) configured to determine a spatial metadata frame for generating a multi-channel upmix signal frame from corresponding frames of a downmix signal, wherein the downmix signal has m channels and the multi-channel upmix signal has n channels, where n and m are integers, and m <nであり、前記空間的メタデータ·フレームは空間的パラメータの集合を含み、当該パラメータ処理ユニットは:A conversion unit (561) configured to determine a first plurality of conversion coefficients from frames of a first channel of a multi-channel input signal, and to determine a second plurality of conversion coefficients from corresponding frames of a second channel of the multi-channel input signal, wherein the first and second plurality of conversion coefficients provide first and second time / frequency representations of frames of the first and second channels, respectively, and the first and second time / frequency representations include a plurality of frequency bins and a plurality of time bins; A parameter determination unit (523) configured to determine the set of spatial parameters based on a plurality of first and second transformation coefficients using fixed-point arithmetic, wherein the set of spatial parameters includes corresponding band parameters for different frequency bands containing a different number of frequency bins, a particular band parameter for a particular frequency band is determined based on transformation coefficients from the plurality of first and second transformation coefficients for the particular frequency band, and the shift used by the fixed-point arithmetic to determine the particular band parameter depends on the particular frequency band, the parameter determination unit comprises Parameter processing unit. [Aspect 26] The parameter processing unit according to embodiment 25, wherein the shift used by fixed-point arithmetic to determine the specific band parameter for the specific frequency band depends on the number of frequency bins contained within the specific frequency band. [Aspect 27] A parameter processing unit according to embodiment 25 or 26, wherein the shift used by fixed-point arithmetic to determine the specific bandwidth parameter for the specific frequency band depends on the number of time bins used to determine the specific bandwidth parameter. [Aspect 28] The parameter processing unit according to any one of embodiments 25 to 27, wherein the parameter determination unit is configured to determine a corresponding shift that maximizes the accuracy of the specific band parameter for the specific frequency band. [Aspect 29] The parameter determination unit determines the specific band parameters for the specific frequency band, A first energy estimate is determined based on the conversion coefficients from the first plurality of conversion coefficients that fall into the specific frequency band; A second energy estimate is determined based on the conversion coefficients from the second set of conversion coefficients that fall into the specific frequency band; The covariance is determined based on the conversion coefficients that fall into the specific frequency band from the first and second plurality of conversion coefficients; Based on the first energy estimate, the second energy estimate, and the maximum of the covariances, the shift for the specific band parameter is determined. A parameter processing unit according to any one of embodiments 25 to 28, configured to perform the processing by the means described above. [Aspect 30] An audio encoding system configured to generate a bitstream based on a multi-channel input signal, wherein: - A downmix processing unit (510) configured to generate a sequence of frames of a downmix signal from a corresponding sequence of first frames of the multichannel input signal, wherein the downmix signal has m channels and the multichannel input signal has n channels, where n and m are integers, and m <nである、ダウンミックス処理ユニットと;A parameter processing unit (520) configured to determine a sequence of spatial metadata frames from a sequence of second frames of the multi-channel input signal, wherein the sequence of frames of the downmix signal and the sequence of spatial metadata frames are for generating a multi-channel upmix signal containing n channels; A bitstream generation unit (503) configured to generate the bitstream including a sequence of bitstream frames, wherein the bitstream frames represent frames of the downmix signal corresponding to the first frames of the sequence of the first frames of the multichannel input signal, and spatial metadata frames corresponding to the second frames of the sequence of the second frames of the multichannel input signal, the second frames being different from the first frames, the bitstream generation unit comprising Audio encoding system. [Aspect 31] The first frame and the second frame have the same number of samples; and / or The sample of the first frame precedes the sample of the second frame. The audio encoding system according to embodiment 30. [Aspect 32] The audio encoding system according to embodiment 30 or 31, wherein the first frame precedes the second frame by a predetermined number of samples. [Aspect 33] The audio encoding system according to embodiment 32, wherein the predetermined number of samples is 928 samples. [Aspect 34] An audio encoding system configured to generate a bitstream based on a multi-channel input signal, • Downmix processing unit (510), - A step in which a sequence of clipping protection gains is determined for a corresponding sequence of frames of the multi-channel input signal, wherein the current clipping protection gain indicates the attenuation to be applied to the current frame of the multi-channel input signal in order to prevent clipping of the corresponding current frame of the downmix signal; - A step of interpolating the current clipping protection gain with the preceding clipping protection gain of the preceding frame of the multi-channel input signal to obtain a clipping protection gain curve; The step of applying the clipping protection gain curve to the current frame of the multi-channel input signal to obtain an attenuated current frame of the multi-channel input signal; - A step of generating the current frame of a sequence of frames of the downmix signal from the attenuated current frame of the multichannel input signal, wherein the downmix signal has m channels and the multichannel input signal has n channels, where n and m are integers, and m <nである、段階とを実行するよう構成されているDownmix processing unit and; A parameter processing unit (520) configured to determine a sequence of spatial metadata frames from the multi-channel input signal, wherein the sequence of frames of the downmix signal and the sequence of spatial metadata frames are for generating a multi-channel upmix signal including n channels; The bitstream generation unit (503) is configured to generate the bitstream showing the sequence of clipping protection gain, the sequence of frames of the downmix signal, and the sequence of spatial metadata frames, so that a corresponding decoding system can generate the multi-channel upmix signal. Audio encoding system. [Aspect 35] The clipping protection gain curve is, A transition segment that provides a smooth transition from the preceding clipping protection gain to the current clipping protection gain; - Including a flat segment that remains flat in the current clipping protection gain, The audio encoding system described in embodiment 34. [Aspect 36] The transition segment extends through a predetermined number of samples of the current frame of the multi-channel input signal. The predetermined number of samples is greater than 1 and less than the total number of samples in the current frame of the multi-channel input signal. The audio encoding system according to embodiment 35. [Aspect 37] An audio encoding system configured to generate a bitstream showing a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal: - A downmix processing unit (510) configured to generate the downmix signal from a multichannel input signal, wherein the downmix signal has m channels, and the multichannel input signal has n channels, where n and m are integers, and m <nである、ダウンミックス処理ユニットと;A parameter processing unit configured to determine the sequence of frames of spatial metadata for the corresponding sequence of frames of the multi-channel input signal; The system includes a configuration setting unit (540) configured to determine one or more control settings for the parameter processing unit based on one or more external settings, The one or more external settings include an update cycle indicating a time period for which the corresponding decoding system is required to synchronize with the bitstream, and the configuration setting unit is configured to determine, based on the update cycle, one or more frames of spatial metadata from the sequence of frames of spatial metadata to be encoded as independent frames. Audio encoding system. [Aspect 38] A method for generating a bitstream showing a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal, - A step of generating the downmix signal from a multichannel input signal, wherein the downmix signal has m channels, and the multichannel input signal has n channels, where n and m are integers, and m <nである、段階と;A step of determining one or more control settings based on one or more external settings, wherein the one or more external settings include a target data rate for the bitstream, and the one or more control settings include a maximum data rate for the spatial metadata; The step of determining the spatial metadata from the multi-channel input signal according to one or more control settings, method. [Aspect 39] A method for determining a spatial metadata frame for generating a frame of a multi-channel upmix signal from corresponding frames of a downmix signal, wherein the downmix signal has m channels and the multi-channel upmix signal has n channels, where n and m are integers, and m <nであり、前記空間的メタデータ·フレームは、空間的パラメータの一つまたは複数の集合を含み、当該方法は、The step of determining multiple spectra from the current frame and the immediately following frame of a channel in a multi-channel input signal; The step of weighting the multiple spectra using a window function to obtain multiple weighted spectra; - A step of determining the spatial metadata frame for the current frame of the channel of the multichannel input signal based on the plurality of weighted spectra, wherein the window function depends on one or more of the following: the number of sets of spatial parameters contained within the spatial metadata frame, the presence of one or more transient components in the current frame or the immediately following frame of the multichannel input signal, and / or the time of the transient components. method. [Aspect 40] A method for determining a spatial metadata frame for generating a frame of a multi-channel upmix signal from corresponding frames of a downmix signal, wherein the downmix signal has m channels and the multi-channel upmix signal has n channels, where n and m are integers, and m <nであり、前記空間的メタデータ·フレームは、空間的パラメータの集合を含み、当該方法は、The step of determining a first set of conversion coefficients from the frame of the first channel of a multi-channel input signal; The step of determining a second set of conversion coefficients from corresponding frames of the second channel of the multi-channel input signal, wherein the first and second sets of conversion coefficients provide first and second time / frequency representations of the frames of the first and second channels, respectively, and the first and second time / frequency representations include a set of frequency bins and a set of time bins, and the set of spatial parameters includes corresponding band parameters for different frequency bands, each including a different number of frequency bins; The step of determining a shift to be applied when determining a specific bandwidth parameter for a particular frequency band using fixed-point arithmetic, wherein the shift is determined based on the particular frequency band; The step includes determining the specific band parameter using fixed-point arithmetic and the determined shift, based on the first and second plurality of conversion coefficients that fall within the specific frequency band, method. [Aspect 41] A method for generating a bitstream based on a multi-channel input signal, - A step of generating a sequence of frames of a downmix signal from a corresponding sequence of first frames of the multichannel input signal, wherein the downmix signal has m channels and the multichannel input signal has n channels, where n and m are integers, and m <nである、段階と;The step of determining a sequence of spatial metadata frames from a second sequence of frames of the multi-channel input signal, wherein the sequence of frames of the downmix signal and the sequence of spatial metadata frames are for generating a multi-channel upmix signal having n channels; - A step of generating the bitstream, which includes a sequence of bitstream frames, wherein the bitstream frames represent frames of the downmix signal corresponding to the first frames of the sequence of the first frames of the multichannel input signal, and spatial metadata frames corresponding to the second frames of the sequence of the second frames of the multichannel input signal, wherein the second frames are different from the first frames, method. [Aspect 42] A method for generating a bitstream based on a multi-channel input signal, - A step in which a sequence of clipping protection gains is determined for a corresponding sequence of frames of the multi-channel input signal, wherein the current clipping protection gain indicates the attenuation to be applied to the current frame of the multi-channel input signal in order to prevent clipping of the corresponding current frame of the downmix signal; - A step of interpolating the current clipping protection gain with the preceding clipping protection gain of the preceding frame of the multi-channel input signal to obtain a clipping protection gain curve; The step of applying the clipping protection gain curve to the current frame of the multi-channel input signal to obtain an attenuated current frame of the multi-channel input signal; - A step of generating the current frame of a sequence of frames of the downmix signal from the attenuated current frame of the multichannel input signal, wherein the downmix signal has m channels and the multichannel input signal has n channels, where n and m are integers, and m <nである、段階と;The step of determining a sequence of spatial metadata frames from the multi-channel input signal, wherein the sequence of frames of the downmix signal and the sequence of spatial metadata frames are for generating a multi-channel upmix signal having n channels; The steps include generating the bitstream that shows the sequence of clipping protection gains, the sequence of frames of the downmix signal, and the sequence of spatial metadata frames, in order to enable the generation of the multi-channel upmix signal based on the bitstream, method. [Aspect 43] A method for generating a bitstream showing a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal, - A step of generating the downmix signal from a multichannel input signal, wherein the downmix signal has m channels, and the multichannel input signal has n channels, where n and m are integers, and m <nである、段階と;A step of determining one or more control settings based on one or more external settings, wherein the one or more external settings include an update cycle that indicates a time period for which the decoding system is required to synchronize with the bitstream; The steps include: determining the sequence of frames of spatial metadata for the corresponding sequence of frames of the multi-channel input signal according to one or more control settings; The step of encoding one or more frames of spatial metadata from the sequence of spatial metadata frames as independent frames, based on the update cycle, method. [Aspect 44] An audio decoder (140) configured to decode a bitstream generated by any one of embodiments 38, 41 to 43.
Claims
1. The audio processor receives the multi-channel input audio signal; A step of determining a first set of dynamic range control (DRC) values configured to control the dynamic range of an output audio signal, wherein the first set of DRC values is expressed logarithmically as dB values, and the dB values may be positive or negative; The steps include determining a second set of DRC values configured to prevent the multi-channel input audio signal from being clipped during downmixing by the audio processor; A step of obtaining an attenuated multi-channel input audio signal by applying the DRC values of the second set to the multi-channel input audio signal, wherein applying the DRC values of the second set includes interpolating between successive values of the DRC values of the second set; The steps include: downmixing the attenuated multi-channel input audio signal to obtain a downmixed signal; The step of generating the output audio signal from the DRC values of the first set and the downmix signal, method.
2. The method according to claim 1, wherein one or more channels of the downmix signal are obtained from a linear combination of some or all of the channels of the multi-channel input audio signal, and the linear combination is controlled by a gain value.
3. One or more processors; A device having a memory that stores instructions for causing one or more processors to perform an operation when executed by the one or more processors, wherein the operation is: The stage of receiving a multi-channel input audio signal; A step of determining a first set of dynamic range control (DRC) values configured to control the dynamic range of an output audio signal, wherein the first set of DRC values is expressed logarithmically as dB values, and the dB values may be positive or negative; The steps include determining a second set of DRC values configured to prevent the multi-channel input audio signal from being clipped during downmixing by the device; A step of obtaining an attenuated multi-channel input audio signal by applying the DRC values of the second set to the multi-channel input audio signal, wherein applying the DRC values of the second set includes interpolating between successive values of the DRC values of the second set; The steps include: downmixing the attenuated multi-channel input audio signal to obtain a downmixed signal; The step of generating the output audio signal from the DRC values of the first set and the downmix signal, Device.
4. The apparatus according to claim 3, wherein one or more channels of the downmix signal are obtained from a linear combination of some or all of the channels of the multi-channel input audio signal, and the linear combination is controlled by a gain value.
5. A non-temporary computer-readable storage medium having a sequence of instructions that, when executed by an audio signal processing device, cause the audio signal processing device to perform the method according to claim 1 or 2.
Citation Information
Patent Citations
Apparatus and method for encoding and decoding audio signals
JP2009500658A
Audio device
JP2011035459A
Protecting signal clipping using existing audio gain metadata
JP2012507059A
Advanced stereo coding based on adaptively selectable left / right or mid / side stereo coding and parametric stereo coding combinations.
JP2012521012A