Methods for parametric multi-channel encoding

HK40137700APending Publication Date: 2026-09-18DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
HK42026126933
Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-02-21
Filing Date
2026-07-31
Publication Date
2026-09-18
Estimated Expiration
2034-02-20

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a method for parametric multichannel coding. Specifically, this document relates to an efficient method and system for parametric multichannel audio coding. An audio coding system (500) is described, which is configured to generate a bitstream (564) indicating a downmixed signal and spatial metadata for generating a multichannel upmixed signal from the downmixed signal. The system (500) includes a downmix processing unit (510) configured to generate a downmixed signal from a multichannel input signal (561); wherein the downmixed signal includes m channels, and wherein the multichannel input signal (561) includes n channels; n, m are integers, where m < n. In addition, the system (500) includes a parameter processing unit (520) configured to determine spatial metadata from the multichannel input signal (561). Additionally, the system (500) includes a configuration unit (540) configured to determine one or more control settings for the parameter processing unit (520) based on one or more external settings; wherein the one or more external settings include the target data rate of the bitstream (564), and wherein the one or more control settings include the maximum data rate of the spatial metadata.
Need to check novelty before this filing date? Find Prior Art

Description

(19) State Intellectual Property Office of the People's Republic of China (12) Invention Patent Application (10) Application Publication No. (43) Application Publication Date (21) Application No. 202610667699.X (22) Application Date February 21, 2014 (30) Priority Data 61 / 767,673 February 21, 2013 US (62) Divisional Original Application Data 201480010021.X February 21, 2014 (71) Applicant Dolby International BV Address Netherlands (72) Inventors T. Vreugdenhil, A. Müller, K. Linzmeier, C-C. Springer, T. R. Wangeblås (74) Patent Agency China Council for the Promotion of International Trade Patent and Trademark Office Co., Ltd. 11038 Patent Agent Liu Qianhong (51) Int.Cl. G10L 19 / 008(2013.01) G10L 19 / 16(2013.01) (54) Title of the Invention Method for Parametric Multichannel Coding (57) Abstract The present invention relates to a method for parametric multichannel coding. In particular, this document relates to an efficient method and system for parametric multichannel audio coding. An audio coding system (500) is described, which is configured to generate a bitstream (564) indicating a downmix signal and spatial metadata, where the spatial metadata is used to generate a multichannel upmixed signal from the downmix signal. The system (500) comprises a downmix processing unit (510) configured to generate a downmix signal from a multichannel input signal (561); wherein the downmix signal comprises m channels, and the multichannel input signal (561) comprises n channels; n and m are integers, where m < n. Furthermore, the system (500) comprises a parameter processing unit (520) configured to determine the spatial metadata from the multichannel input signal (561). In addition, the system (500) comprises a configuration unit (540) configured to determine one or more control settings for the parameter processing unit (520) based on one or more external settings; wherein the one or more external settings comprise a target data rate of the bitstream (564), and wherein the one or more control settings comprise a maximum data rate of the spatial metadata.Claims 4 pages, Description 30 pages, Drawings 13 pages, CN 122369474 A 2026.07.10 CN 1 22 36 94 74 A 1. A method comprising: receiving a multi-channel input audio signal and a downmixing instruction including a set of gain values ​​via an audio processor; determining a first dynamic range control (DRC) value set configured to control the dynamic range of an output audio signal; determining a second DRC value set configured to prevent the multi-channel input audio signal from being trimmed during downmixing by the audio processor; applying the second DRC value set to the multi-channel input audio signal to obtain a attenuated multi-channel input audio signal; downmixing the attenuated multi-channel input audio signal in response to the downmixing instruction to obtain a downmixed signal; and generating the output audio signal from the first DRC value set and the downmixed audio signal. 2. An apparatus comprising: one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including: receiving a multichannel input audio signal and a downmixing instruction including a set of gain values; determining a first dynamic range control (DRC) value set configured to control the dynamic range of an output audio signal; determining a second DRC value set configured to prevent the multichannel input audio signal from being trimmed during downmixing by the apparatus; applying the second DRC value set to the multichannel input audio signal to obtain a attenuated multichannel input audio signal; downmixing the attenuated multichannel input audio signal in response to the downmixing instruction to obtain a downmixed signal; and generating the output audio signal from the first DRC value set and the downmixed audio signal. 3. The apparatus of claim 2, wherein generating the output audio signal includes applying the first DRC value set to the downmixed audio signal. 4. The apparatus of claim 2, wherein the first DRC value set and / or the second DRC value set are represented in logarithmic form as dB values. 5. The apparatus of claim 2, wherein the multichannel input audio signal is divided into a frame sequence of samples of the multichannel audio signal, and determining the first DRC value set and / or the second DRC value set comprises determining the DRC value of each sample of each frame in the frame sequence. 6. The apparatus of claim 5, wherein determining the DRC value of each sample of a frame comprises interpolating between the DRC value of the frame and the DRC value of the previous frame. 7. The apparatus of claim 6, wherein the interpolation is spline interpolation. 8. The apparatus of claim 2, wherein the downmix signal is a stereo signal.9. The apparatus according to claim 2, wherein the left channel and right channel of the downmix signal are generated based on different linear combinations of channels of the multi-channel input audio signal. 10. A non-transitory computer-readable storage medium comprising an instruction sequence which, when executed by an audio signal processing apparatus, causes the audio signal processing apparatus to perform a method, the method comprising: Receiving, by an audio processor, a multi-channel input audio signal and a downmix instruction comprising a set of gain values; Determining a first set of dynamic range control (DRC) values, wherein the first set of DRC values is configured to control the dynamic range of an output audio signal; Determining a second set of DRC values, wherein the second set of DRC values is configured to prevent the multi-channel input audio signal from being clipped during downmixing by the audio processor; Applying the second set of DRC values to the multi-channel input audio signal to obtain an attenuated multi-channel input audio signal; In response to the downmix instruction, downmixing the attenuated multi-channel input audio signal to obtain a downmix signal; and Generating the output audio signal from the first set of DRC values and the downmixed audio signal. 11. A parameter processing unit (520), wherein the parameter processing unit (520) is configured to determine a spatial metadata frame for a frame of a multi-channel upmixed signal generated from a corresponding frame of a downmix signal; wherein the downmix signal comprises m channels, and wherein the multi-channel upmixed signal comprises n channels; n and m are integers, wherein m < n; wherein the spatial metadata frame comprises one or more sets of spatial parameters (711, 712); the parameter processing unit (520) comprises: - a transform unit (521), wherein the transform unit (521) is configured to determine a plurality of frequency spectrums (589) from a current frame (585) and an immediate next frame (590) of channels of a multi-channel input signal (561); and - a parameter determining unit (523), wherein the parameter determining unit (523) is configured to determine the spatial metadata frame for the current frame of the channels of the multi-channel input signal (561) by weighting the plurality of frequency spectrums (589) using a window function (586); wherein the window function (586) depends on one or more of the following: the number of sets of spatial parameters (711, 712) included in the spatial metadata frame, the presence of one or more transients in the current frame or the immediate next frame of the multi-channel input signal (561), and / or the time instant of the transients.12. A parameter processing unit (520), wherein the parameter processing unit (520) is configured to determine a spatial metadata frame for generating a frame of a multi-channel upmixed signal from a corresponding frame of a downmixed signal; wherein the downmixed signal comprises m channels, and wherein the multi-channel upmixed signal comprises n channels; n and m are integers, where m < n; wherein the spatial metadata frame comprises a spatial parameter set (711); the parameter processing unit (520) comprises: - a transform unit (521), wherein the transform unit (521) is configured to: determine a first plurality of transform coefficients (580) from a frame (585) of a first channel (561-1) of a multi-channel input signal (561), and determine a second plurality of transform coefficients (580) from a corresponding frame of a second channel (561-2) of the multi-channel input signal (561); wherein the first channel (561-1) and the second channel (561-2) are different; wherein the first plurality of transform coefficients (580) and the second plurality of transform coefficients (580) respectively provide a first time / frequency representation and a second time / frequency representation of the frame (585) of the first channel and the second channel; wherein the first time / frequency representation and the second time / frequency representation comprise a plurality of frequency bins (571) and a plurality of time intervals (582); and - a parameter determination unit (523), wherein the parameter determination unit (523) is configured to determine the spatial parameter set (711) based on the first plurality of transform coefficients (580) and the second plurality of transform coefficients (580) using fixed-point arithmetic; wherein the spatial parameter set (711) comprises corresponding band parameters for different frequency bands (572) comprising different numbers of frequency bins (571); wherein a specific band parameter for a specific frequency band (572) is determined based on transform coefficients (580) from the first plurality of transform coefficients (580) and the second plurality of transform coefficients (580) of the specific frequency band (572); and wherein, Claims Page 2 of 4 3 CN 122369474 A a shift used by the fixed-point arithmetic for determining the specific band parameter depends on the specific frequency band (572).13. An audio encoding system (500), wherein the audio encoding system (500) is configured to generate a bitstream (564) based on a multi-channel input signal (561); the system (500) comprises: - a downmix processing unit (510), wherein the downmix processing unit (510) is configured to generate a frame sequence of a downmix signal from a corresponding first frame sequence of the multi-channel input signal (561); wherein the downmix signal comprises m channels, and wherein the multi-channel input signal (561) comprises n channels; n and m are integers, where m < n; - a parameter processing unit (520), wherein the parameter processing unit (520) is configured to determine a spatial metadata frame sequence from a second frame sequence of the multi-channel input signal (561); wherein the frame sequence of the downmix signal and the spatial metadata frame sequence are used to generate a multi-channel upmixed signal comprising n channels; and - a bitstream generating unit (503), wherein the bitstream generating unit (503) is configured to generate the bitstream (564) comprising a bitstream frame sequence, wherein a bitstream frame indicates a frame of the downmix signal corresponding to a first frame of the first frame sequence of the multi-channel input signal (561) and a spatial metadata frame corresponding to a second frame of the second frame sequence of the multi-channel input signal (561); wherein the second frame is different from the first frame. 14. A method for determining a spatial metadata frame, wherein the spatial metadata frame is used to generate a frame of a multi-channel upmixed signal from a corresponding frame of a downmix signal; wherein the downmix signal comprises m channels, and wherein the multi-channel upmixed signal comprises n channels; n and m are integers, where m < n; wherein the spatial metadata frame comprises one or more spatial parameter sets (711, 712); the method comprises: - determining a plurality of spectrums (589) from a current frame (585) and an immediately following frame (590) of channels of a multi-channel input signal (561); - weighting the plurality of spectrums (589) using a window function (586) to obtain a plurality of weighted spectrums; and - determining, based on the plurality of weighted spectrums, the spatial metadata frame for the current frame of the channels of the multi-channel input signal (561); wherein the window function (586) depends on one or more of: the number of spatial parameter sets (711, 712) included in the spatial metadata frame, the presence of one or more transients in the current frame or the immediately following frame of the multi-channel input signal (561), and / or a time position of the transient.15. A method for determining a spatial metadata frame, wherein the spatial metadata frame is configured to generate a frame of a multi-channel upmixed signal from a corresponding frame of a downmixed signal; wherein the downmixed signal comprises m channels, and the multi-channel upmixed signal comprises n channels; n and m are integers, where m < n; wherein the spatial metadata frame comprises a spatial parameter set (711); the method comprises: - determining a first plurality of transform coefficients (580) from a frame (585) of a first channel (561-1) of a multi-channel input signal (561); - determining a second plurality of transform coefficients (580) from a corresponding frame of a second channel (561-2) of the multi-channel input signal (561); wherein the first channel (561-1) and the second channel (561-2) are different; wherein the first plurality of transform coefficients (580) and the second plurality of transform coefficients (580) respectively provide a first time / frequency representation and a second time / frequency representation of the frames (585) of the first channel and the second channel; wherein the first time / frequency representation and the second time / frequency representation comprise a plurality of frequency intervals (571) and a plurality of time intervals (582); wherein the spatial parameter set (711) comprises corresponding band parameters for different frequency bands (572) comprising different numbers of frequency intervals (571); - determining a shift to be applied when a specific band parameter for a specific frequency band (572) is determined using fixed-point arithmetic; wherein the shift is determined based on the specific frequency band (572); and - determining the specific band parameter using fixed-point arithmetic and the determined shift based on the first plurality of transform coefficients (580) and the second plurality of transform coefficients (580) falling within the specific frequency band (572). 16. A method for generating a bitstream (564) based on a multi-channel input signal (561); the method comprises: - generating a frame sequence of a downmixed signal from a corresponding first frame sequence of the multi-channel input signal (561); wherein the downmixed signal comprises m channels, and the multi-channel input signal (561) comprises n channels; n and m are integers, where m < n; - determining a spatial metadata frame sequence from a second frame sequence of the multi-channel input signal (561); wherein the frame sequence of the downmixed signal and the spatial metadata frame sequence are configured to generate a multi-channel upmixed signal comprising n channels; and - generating a bitstream (564) comprising a bitstream frame sequence; wherein a bitstream frame indicates a frame of the downmixed signal corresponding to a first frame of the first frame sequence of the multi-channel input signal (561) and a spatial metadata frame corresponding to a second frame of the second frame sequence of the multi-channel input signal (561); wherein the second frame is different from the first frame.CLAIMS Page 4 of 4 5 CN 122369474 A A Method for Parametric Multi-channel Encoding

[0001] This application is a divisional application of the invention patent application with application number 202310791753.8, filing date February 21, 2014, and the invention title "Method for Parametric Multi-channel Encoding".

[0002] Cross-Reference to Related Applications

[0003] This application claims the priority of U.S. Provisional Patent Application No. 61 / 767,673 filed on February 21, 2013, the entire content of which is hereby incorporated by reference. Technical Field

[0004] This document relates to audio encoding systems. Specifically, this document relates to efficient methods and systems for parametric multi-channel audio encoding. Background Art

[0005] Parametric multi-channel audio encoding systems can be used to provide improved listening quality at a particularly low data rate. Nevertheless, there is still a need for further improvements to such parametric multi-channel audio encoding systems, particularly in terms of bandwidth efficiency, computational efficiency and / or robustness. Summary of the Invention

[0006] According to one aspect, an audio encoding system configured to generate a bitstream indicating a downmix signal and spatial metadata is described. The spatial metadata can be used by a corresponding decoding system to generate a multi-channel upmixed signal from the downmix signal. The downmix signal may comprise m channels, and the multi-channel upmixed signal may comprise n channels, wherein n and m are integers, and m < n. In an example, n=6, m=2. The spatial metadata enables the corresponding decoding system to generate the n channels of the multi-channel upmixed signal from the m channels of the downmix signal.

[0007] The audio encoding system may be configured to quantize and / or encode the downmix signal and the spatial metadata, and insert the quantized / encoded data into the bitstream. Specifically, the downmix signal may be encoded using a Dolby Digital Plus encoder, and the bitstream may correspond to a Dolby Digital Plus bitstream. The quantized / encoded spatial metadata may be inserted into a data field of the Dolby Digital Plus bitstream.

[0008] The audio encoding system may include a downmix processing unit configured to generate a downmix signal from a multi-channel input signal. The downmix processing unit is also referred to herein as a downmix encoding unit. The multi-channel input signal may comprise n channels, such as the multi-channel upmixed signal regenerated based on the downmix signal. Specifically, the multi-channel upmixed signal may provide an approximation of the multi-channel input signal. The downmix unit may comprise the aforementioned Dolby Digital Plus encoder. The multi-channel upmixed signal and the multi-channel input signal may be 5.1 or 7.1 signals, and the downmix signal may be a stereo signal.

[0009] The audio encoding system may include a parameter processing unit configured to determine spatial metadata from the multi-channel input signal.Specifically, the parameter processing unit (also referred to herein as the parameter encoding unit) can be configured to determine one or more spatial parameters, such as a set of spatial parameters, which can be determined based on different combinations of channels of the multichannel input signal. The spatial parameters of the set of spatial parameters can indicate the cross-correlation between different channels of the multichannel input signal. The parameter processing unit can be configured to determine spatial metadata of a frame of the multichannel input signal, called a spatial metadata frame. A frame of the multichannel input signal typically includes a predetermined number (e.g., 1536) samples of the multichannel input signal. Each spatial metadata frame may include one or more sets of spatial parameters.

[0010] The audio encoding system may also include a configuration unit configured to determine one or more control settings for the parameter processing unit based on one or more external settings. The one or more external settings may include a target data rate for the bitstream. Alternatively or additionally, the one or more external settings may include one or more of the following: the sampling rate of the multichannel input signal, the number of channels m of the downmix signal, the number of channels n of the multichannel input signal, and / or an update period indicating the time period required for the corresponding decoding system to synchronize with the bitstream. The one or more control settings may include the maximum data rate of the spatial metadata. In the case of a spatial metadata frame, the maximum data rate of the spatial metadata may indicate the maximum number of metadata bits of the spatial metadata frame. Alternatively or additionally, the one or more control settings may include one or more of the following: a time resolution setting indicating the number of spatial parameter sets for each spatial metadata frame to be determined; a frequency resolution setting indicating the number of frequency bands for which spatial parameters will be determined; a quantizer setting indicating the type of quantizer to be used to quantize the spatial metadata; and an indication of whether the current frame of the multichannel input signal will be encoded as an independent frame.

[0011] The parameter processing unit may be configured to determine whether the number of bits of the spatial metadata frame determined according to the one or more control settings exceeds the maximum number of metadata bits. Furthermore, the parameter processing unit can be configured to reduce the number of bits in a particular spatial metadata frame if it is determined that the number of bits in that particular spatial metadata frame exceeds the maximum number of metadata bits. This bit reduction can be performed in a resource-efficient (processing power) manner. Specifically, this bit reduction can be performed without recalculating the entire spatial metadata frame.

[0012] As indicated above, a spatial metadata frame may include one or more sets of spatial parameters. The one or more control settings may include a time resolution setting that indicates the number of sets of spatial parameters for each spatial metadata frame determined by the parameter processing unit.The parameter processing unit can be configured to determine a plurality of spatial parameter sets for the current spatial metadata frame, as indicated by the time resolution setting. Typically, the time resolution setting is a value of 1 or 2. Furthermore, the parameter processing unit can be configured to discard spatial parameter sets from the current spatial metadata frame if the current spatial metadata frame includes multiple spatial parameter sets, and if the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits. The parameter processing unit can be configured to retain at least one spatial parameter set for each spatial metadata frame. By discarding spatial parameter sets from spatial metadata frames, the number of bits in the spatial metadata frame can be reduced with minimal computational effort and without significantly affecting the perceived listening quality of the multi-channel upmixed signal.

[0013] The one or more spatial parameter sets are typically associated with one or more corresponding sample points. The one or more sample points can indicate one or more corresponding moments. Specifically, a sample point can indicate the moment when the decoding system should adequately apply the corresponding spatial parameter set. In other words, a sample point can indicate the moment when the corresponding spatial parameter set has been determined.

[0014] The parameter processing unit can be configured to discard a first set of spatial parameters from the current spatial metadata frame if multiple sampling points of the current metadata frame are not associated with transients of the multichannel input signal, wherein the first set of spatial parameters is associated with a first sampling point preceding a second sampling point. Alternatively, the parameter processing unit can be configured to discard a second (typically the last) set of spatial parameters from the current spatial metadata frame if multiple sampling points of the current metadata frame are associated with transients of the multichannel input signal. By doing so, the parameter processing unit can be configured to reduce the impact of discarding spatial parameter sets on the listening quality of the multichannel upmixed signal.

[0015] The one or more control settings may include quantizer settings indicating a first type of quantizer among a plurality of predetermined types. The plurality of predetermined types of quantizers may each provide different quantizer resolutions. Specifically, the plurality of predetermined types of quantizers may include fine quantization and coarse quantization. The parameter processing unit can be configured, as per specification page 2 / 30 7 CN 122369474 A, to quantize one or more sets of spatial parameters of the current spatial metadata frame according to the first type of quantizer. Furthermore, the parameter processing unit can be configured to, if it is determined that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits, requantize one, some, or all of the spatial parameters of the one or more sets of spatial parameters according to a second type of quantizer with a resolution lower than that of the first type of quantizer.By doing so, the number of bits in the current spatial metadata frame can be reduced, while only affecting the quality of the upmixed signal to a limited extent, and without significantly increasing the computational complexity of the audio coding system.

[0016] The parameter processing unit can be configured to determine the time difference parameter set based on the difference between the current spatial parameter set and the immediately preceding spatial parameter set. Specifically, the time difference parameter can be determined by determining the difference between the parameters of the current spatial parameter set and the corresponding parameters of the immediately preceding spatial parameter set. The spatial parameter set may include, for example, the parameters α1, α2, α3, β1, β2, β3, g, k1, k2 described in this document. Typically, only one of the parameters k1 and k2 may need to be sent, as these parameters can be correlated using a relation. For example only, only parameter k1 can be sent, and parameter k2 can be computed at the receiver. The time difference parameter can be correlated with the difference between the corresponding parameters among the parameters mentioned above.

[0017] The parameter processing unit can be configured to encode the time difference parameter set using entropy coding (e.g., using Huffman codes). Furthermore, the parameter processing unit can be configured to insert the encoded time difference parameter set into the current spatial metadata frame. Additionally, the parameter processing unit can be configured to reduce the entropy of the time difference parameter set if it is determined that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits. As a result, the number of bits required for entropy encoding of the time difference parameters can be reduced, thereby reducing the number of bits used for the current spatial metadata frame. For example, the parameter processing unit can be configured to set one, some, or all of the time difference parameters in the time difference parameter set to a value with an increased (e.g., highest) probability among the possible values ​​of the time difference parameter, in order to reduce the entropy of the time difference parameter set. Specifically, the probability can be increased compared to the probability of the time difference parameter before the setting operation. Typically, the value with the highest probability among the possible values ​​of the time difference parameter corresponds to zero.

[0018] It should be noted that time difference encoding of the spatial parameter set generally cannot be used for independent frames. Thus, the parameter processing unit can be configured to verify whether the current spatial metadata frame is an independent frame, and only apply time difference encoding if the current spatial metadata frame is not an independent frame. On the other hand, the frequency difference encoding described below can also be used for independent frames.

[0019] The one or more control settings may include a frequency resolution setting, wherein the frequency resolution setting indicates the number of different frequency bands for which respective spatial parameters (referred to as band parameters) will be determined. The parameter processing unit may be configured to determine different corresponding spatial parameters (band parameters) for different frequency bands. Specifically, different parameters α1, α2, α3, β1, β2, β3, g, k1, k2 may be determined for different frequency bands. The set of spatial parameters may therefore include corresponding band parameters for different frequency bands.For example, the spatial parameter set may include T corresponding band parameters for T frequency bands, where T is an integer, such as T = 7, 9, 12, or 15.

[0020] The parameter processing unit may be configured to determine the frequency difference parameter set based on the difference between one or more band parameters in the first frequency band and the corresponding one or more band parameters in the adjacent second frequency band. Furthermore, the parameter processing unit may be configured to encode the frequency difference parameter set using entropy coding (e.g., based on Huffman codes). Additionally, the parameter processing unit may be configured to insert the encoded frequency difference parameter set into the current spatial metadata frame. Furthermore, the parameter processing unit may be configured to reduce the entropy of the frequency difference parameter set if it is determined that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits. Specifically, the parameter processing unit may be configured to set one, some, or all of the frequency difference parameters in the frequency difference parameter set to a value with an increased probability (e.g., zero) among the possible values ​​of the frequency difference parameter, in order to reduce the entropy of the frequency difference parameter set. Specifically, the probability may be increased compared to the probability of the frequency difference parameter before the setting operation.

[0021] Alternatively or additionally, the parameter processing unit may be configured to reduce the number of frequency bands if it is determined that the number of bits in the current spatial metadata frame exceeds the maximum number of metadata bits. Additionally, the parameter processing unit may be configured to use the reduced number of frequency bands to redetermine some or all of the one or more sets of spatial parameters for the current spatial metadata frame. Typically, the change in the number of frequency bands primarily affects higher frequency bands. As a result, the band parameters of one of several frequencies may be unaffected, making it possible that the parameter processing unit may not need to recalculate all the band parameters.

[0022] As indicated above, the one or more external settings may include an update period indicating the time period required for the corresponding decoding system to synchronize with the bitstream. Furthermore, the one or more control settings may include an indication of whether the current spatial metadata frame will be encoded as an independent frame. The parameter processing unit may be configured to determine a spatial metadata frame sequence for a corresponding frame sequence of multichannel input signals. The configuration unit may be configured to determine one or more spatial metadata frames to be encoded as independent frames from the spatial metadata frame sequence based on the update period.

[0023] Specifically, the one or more independent spatial metadata frames can be determined such that an update period (on average) is satisfied. For this purpose, the configuration unit can be configured to determine whether the current frame of the frame sequence of the multichannel input signal includes samples of times that are integer multiples of the update period (relative to the start point of the multichannel input signal). Furthermore, the configuration unit can be configured to determine that the current spatial metadata frame corresponding to the current frame is an independent frame (because it includes samples of times that are integer multiples of the update period).The parameter processing unit may be configured to encode one or more sets of spatial parameters of the current spatial metadata frame independently of data included in a previous (and / or future) spatial metadata frame, if the current spatial metadata frame is to be encoded as an independent frame. Generally, if the current spatial metadata frame is to be encoded as an independent frame, all sets of spatial parameters of current spatial metadata are encoded independently of data included in the previous (and / or future) spatial metadata frame.

[0024] According to another aspect, a parameter processing unit is described, which is configured to determine a spatial metadata frame for a frame of a multichannel upmixed signal generated from a corresponding frame of a downmixed signal. The downmixed signal may include m channels, and the multichannel upmixed signal may include n channels; n and m are integers, where m < n. As outlined above, the spatial metadata frame may include one or more sets of spatial parameters.

[0025] The parameter processing unit may include a transform unit configured to determine a plurality of spectra from a current frame and a succeeding frame (referred to as a look-ahead frame) of channels of a multichannel input signal. The transform unit may use a filter bank, for example, a QMF filter bank. A spectrum among the plurality of spectra includes a predetermined number of transform coefficients in a corresponding predetermined number of frequency bins. The plurality of spectra may be associated with a corresponding plurality of time intervals (or time instants). In this way, the transform unit may be configured to provide a time / frequency representation of the current frame and the look-ahead frame. For example, both the current frame and the look-ahead frame may include K samples. The transform unit may be configured to determine 2×K / Q spectra, each of which includes Q transform coefficients.

[0026] The parameter processing unit may include a parameter determination unit configured to determine a spatial metadata frame for a current frame of channels of a multichannel input signal by weighting the plurality of spectra using a window function. The window function may be used to adjust the influence of a spectrum among the plurality of spectra on a specific spatial parameter or a specific set of spatial parameters. For example, the window function may take values between 0 and 1.

[0027] The window function may depend on one or more of the following: the number of sets of spatial parameters included in the spatial metadata frame, the presence of one or more transients in the current frame or a succeeding frame of the multichannel input signal, and / or the time instant of a transient. In other words, the window function may be changed according to the properties of the current frame and / or the look-ahead frame. Specifically, the window function for determining a set of spatial parameters (which is referred to as a set-dependent window function) may depend on the properties of the current frame and / or the look-ahead frame.

[0028] In this way, the window function may include a set-dependent window function. Specifically, the window function for determining spatial parameters of a spatial metadata frame may include (or may consist of) one or more set-dependent window functions respectively for one or more sets of spatial parameters. Description 4 / 30 page 9 CN 122369474 AThe parameter determination unit can be configured to determine the set of spatial parameters for the current frame (i.e., for the current spatial metadata frame) of the channels of the multi-channel input signal by weighting the plurality of spectra using a set-dependent window function. As outlined above, the set-dependent window function may depend on one or more properties of the current frame. Specifically, the set-dependent window function may depend on whether the set of spatial parameters is associated with a transient.

[0029] For example, if the set of spatial parameters is not associated with a transient, the set-dependent window function can be configured to provide a phase-in of the plurality of spectra from the sampling point of the previous set of spatial parameters to the sampling point of the next set of spatial parameters. The phase-in can be provided by a window function that transitions from 0 to 1. Alternatively or additionally, if the set of spatial parameters is not associated with a transient, the set-dependent window function may include a plurality of spectra from the sampling point of the spatial parameter set to the spectra preceding the sampling point of the next set of spatial parameters (or these spectra may be adequately considered, or these spectra may be unaffected), if the next set of spatial parameters is associated with a transient. This can be achieved by a window function with a value of 1. Alternatively or additionally, if the spatial parameter set is not associated with the transient, a set-dependent window function can cancel out (or exclude, or attenuate) the plurality of spectra starting from the sampling points of the subsequent spatial parameter set, if the subsequent spatial parameter set is associated with the transient. This can be achieved using a window function with a value of 0. Alternatively or additionally, if the spatial parameter set is not associated with the transient, a set-dependent window function can phase-out the plurality of spectra from the sampling points of the spatial parameter set up to the spectra preceding the sampling points of the subsequent spatial parameter set, if the subsequent spatial parameter set is not associated with the transient. Phase-out can be provided by a window function transitioning from 1 to 0. On the other hand, if the spatial parameter set is associated with the transient, a set-dependent window function can cancel out (or exclude, or attenuate) the spectra preceding the sampling points of the spatial parameter set. Alternatively or additionally, if the set of spatial parameters is associated with a transient, the set-related window function may include the spectrum of the plurality of spectra from the sampling point of the spatial parameter set up to the spectrum of the plurality of spectra preceding the sampling point of the latter spatial parameter set (or may make these spectra unaffected), and may eliminate the spectrum of the plurality of spectra starting from the sampling point of the latter spatial parameter set (or may exclude these spectra, or may attenuate these spectra), if the sampling point of the latter spatial parameter set is associated with a transient.Alternatively or additionally, if the set of spatial parameters is associated with a transient, the set-related window function may comprise the spectra of the plurality of spectra from the sampling point of the set of spatial parameters up to the spectrum at the end of the current frame (or may leave these spectra unaffected), and may provide fading of the spectra of the plurality of spectra from the start of the immediately following frame up to the sampling point of the subsequent set of spatial parameters (or may gradually attenuate these spectra), if said subsequent set of spatial parameters is not associated with a transient.

[0030] According to another aspect, a parameter processing unit is described, which is configured to determine a spatial metadata frame for generating a frame of a multi-channel upmixed signal from a corresponding frame of a downmixed signal. The downmixed signal may comprise m channels, and the multi-channel upmixed signal may comprise n channels; n and m are integers, wherein m < n. As discussed above, the spatial metadata frame may comprise a set of spatial parameters.

[0031] As outlined above, the parameter processing unit may comprise a transform unit. The transform unit may be configured to determine a first plurality of transform coefficients from a frame of a first channel of a multi-channel input signal. Furthermore, the transform unit may be configured to determine a second plurality of transform coefficients from a corresponding frame of a second channel of the multi-channel input signal. The first channel and the second channel may be different. In this way, the first plurality of transform coefficients and the second plurality of transform coefficients respectively provide a first time / frequency representation and a second time / frequency representation of the corresponding frames of the first channel and the second channel. As outlined above, the first time / frequency representation and the second time / frequency representation comprise a plurality of frequency intervals and a plurality of time intervals.

[0032] Furthermore, the parameter processing unit may comprise a parameter determination unit, which is configured to determine the set of spatial parameters based on the first plurality of transform coefficients and the second plurality of transform coefficients using fixed-point arithmetic. As indicated above, the set of spatial parameters generally comprises respective band parameters for different frequency bands, wherein said different frequency bands may comprise different numbers of frequency intervals. A specific band parameter for a specific frequency band may be determined based on the transform coefficients among the first plurality of transform coefficients and the second plurality of transform coefficients of the specific frequency band (generally, transform coefficients of other frequency bands are not considered). The parameter determination unit may be configured to determine the shift used by fixed-point arithmetic for determining a specific band parameter dependent on a specific frequency band. In particular, the shift used by fixed-point arithmetic for determining a specific band parameter for a specific frequency band may depend on the number of frequency intervals included within the specific frequency band. Alternatively or additionally, the shift used by fixed-point arithmetic for determining a specific band parameter for a specific frequency band may depend on the number of time intervals to be considered for determining the specific band parameter.

[0033] The parameter determination unit may be configured to determine the shift for a specific frequency band such that the precision of the specific band parameter is maximized.This can be achieved by determining the shifts required for each multiplication and addition operation of the definite processing with specific band parameters.

[0034] The parameter determining unit may be configured to determine a specific band parameter for a specific frequency band p by determining a first energy (or energy estimate) E1,1(p) based on transform coefficients falling into the specific frequency band p among the first plurality of transform coefficients. Furthermore, a second energy (or energy estimate) E2,2(p) may be determined based on transform coefficients falling into the specific frequency band p among the second plurality of transform coefficients. In addition, a cross product or covariance E1,2(p) may be determined based on transform coefficients falling into the specific frequency band p among the first plurality of transform coefficients and the second plurality of transform coefficients. The parameter determining unit may be configured to determine the shift zp for the specific frequency band parameter p based on the maximum value among the absolute values of the first energy estimate E1,1(p), the second energy estimate E2,2(p) and the covariance E1,2(p).

[0035] According to another aspect, an audio encoding system is described, which is configured to generate a bitstream, the bitstream indicating a frame sequence of a downmix signal and a corresponding sequence of spatial metadata frames, the corresponding sequence of spatial metadata frames being configured to generate a corresponding frame sequence of a multi-channel upmix signal from the frame sequence of the downmix signal. The system may comprise a downmix processing unit configured to generate the frame sequence of the downmix signal from a corresponding frame sequence of a multi-channel input signal. As indicated above, the downmix signal may comprise m channels, and the multi-channel input signal may comprise n channels; n and m are integers, where m<n. Furthermore, the audio encoding system may comprise a parameter processing unit configured to determine the sequence of spatial metadata frames from the frame sequence of the multi-channel input signal.

[0036] In addition, the audio encoding system may comprise a bitstream generating unit configured to generate a bitstream comprising a sequence of bitstream frames, wherein a bitstream frame indicates a frame of the downmix signal corresponding to a first frame of the multi-channel input signal and a spatial metadata frame corresponding to a second frame of the multi-channel input signal. The second frame may be different from the first frame. Specifically, the first frame may precede the second frame. By doing so, the spatial metadata frame for a current frame can be transmitted together with the frame of a subsequent frame. This ensures that the spatial metadata frame only arrives at a corresponding decoding system when it is needed. A decoding system typically decodes a current frame of the downmix signal and generates a decorrelated frame based on the current frame of the downmix signal. This processing introduces algorithmic delay, and by delaying the spatial metadata frame for the current frame, it is ensured that the spatial metadata frame arrives at the decoding system only after the decoded current frame and the decorrelated frame are provided. As a result, processing power and memory requirements of the decoding system can be reduced.

[0037] In other words, an audio encoding system is described, which is configured to generate a bitstream based on a multi-channel input signal.As outlined above, the system may comprise a downmix processing unit configured to generate a frame sequence of a downmix signal from a respective first frame sequence of a multi-channel input signal. The downmix signal may comprise m channels, and the multi-channel input signal may comprise n channels, wherein n and m are integers and m < n. Furthermore, the audio encoding system may comprise a parameter processing unit configured to generate a spatial metadata frame sequence from a second frame sequence of the multi-channel input signal. The frame sequence of the downmix signal and the spatial metadata frame sequence may be used by a corresponding decoding system to generate a multi-channel upmix signal comprising n channels.

[0038] The audio encoding system may further comprise a bitstream generating unit configured to generate a bitstream including a bitstream frame sequence, wherein a bitstream frame may indicate a frame of the downmix signal corresponding to a first frame of the first frame sequence of the multi-channel input signal and a spatial metadata frame corresponding to a second frame of the second frame sequence of the multi-channel input signal. The second frame may be different from the first frame. In other words, the framing used to determine the spatial metadata frame and the framing used to determine the frame of the downmix signal may be different. As outlined above, different framing may be used to ensure alignment of data at the corresponding decoding system.

[0039] The first frame and the second frame typically comprise the same number of samples (e.g., 1536 samples). Some of the samples of the first frame may lead the samples of the second frame. Specifically, the first frame may lead the second frame by a predetermined number of samples. The predetermined number of samples may, for example, correspond to a fraction of the number of samples of one frame. For example, the predetermined number of samples may correspond to 50% or more of the number of samples of one frame. In a specific example, the predetermined number of samples corresponds to 928 samples. As shown in this document, this specific number of samples provides the minimum total delay and optimal alignment for a specific implementation of audio encoding and decoding systems.

[0040] According to another aspect, an audio encoding system is described, which is configured to generate a bitstream based on a multi-channel input signal. The system may comprise a downmix processing unit configured to determine a sequence of clip protection gains (also referred to herein as clip-gain and / or DRC2 parameters) for respective frame sequences of the multi-channel input signal. A current clip protection gain may indicate an attenuation to be applied to a current frame of the multi-channel input signal to prevent clipping of a corresponding current frame of the downmix signal. In a similar manner, the sequence of clip protection gains may indicate respective attenuations to be applied to frames of the frame sequence of the multi-channel input signal to prevent clipping of corresponding frames of the frame sequence of the downmix signal.

[0041] The downmix processing unit may be configured to interpolate between the current clip protection gain and a previous clip protection gain of a previous frame of the multi-channel input signal to obtain a clip protection gain curve. This may be performed in a manner similar to that for the sequence of clip protection gains.Furthermore, the downmixing unit can be configured to apply a trim protection gain curve to the current frame of the multichannel input signal to obtain the current frame of the attenuated multichannel input signal. Again, this can be performed in a manner similar to the frame sequence of the multichannel input signal. Furthermore, the downmixing unit can be configured to generate the current frame of the frame sequence of the downmixed signal from the current frame of the attenuated multichannel input signal. In a similar manner, the frame sequence of the downmixed signal can be generated.

[0042] The audio processing system may also include a parameter processing unit configured to determine a spatial metadata frame sequence from the multichannel input signal. The frame sequence of the downmixed signal and the spatial metadata frame sequence can be used to generate a multichannel upmixed signal including n channels, such that the multichannel upmixed signal is an approximation of the multichannel input signal. In addition, the audio processing system may include a bitstream generation unit configured to generate a bitstream indicating the trim protection gain sequence, the frame sequence of the downmixed signal, and the spatial metadata frame sequence, such that a corresponding decoding system is able to generate a multichannel upmixed signal.

[0043] The trim protection gain curve may include a transition segment and a flat segment, the transition segment providing a smooth transition from the previous trim protection gain to the current trim protection gain, and the flat segment remaining flat at the current trim protection gain. The transition segment may extend across a predetermined number of samples of the current frame of the multichannel input signal. The predetermined number of samples may be more than one but less than the total number of samples of the current frame of the multichannel input signal. Specifically, the predetermined number of samples may correspond to a sample block (wherein a frame may include multiple blocks) or a frame. In a particular example, a frame may include 1536 samples, and a block may include 256 samples.

[0044] According to another aspect, an audio coding system is described that is configured to generate a bitstream indicating a downmixed signal and spatial metadata for generating a multichannel upmixed signal from the downmixed signal. The system may include a downmixing processing unit configured to generate the downmixed signal from the multichannel input signal. Furthermore, the system may include a parameter processing unit configured to determine a spatial metadata frame sequence for a corresponding frame sequence of multichannel input signals.

[0045] Furthermore, the audio encoding system may include a configuration unit configured to determine one or more control settings for the parameter processing unit based on one or more external settings. The one or more external settings may include an update period indicating a time period required for the corresponding decoding system to synchronize with the bitstream. The configuration unit may be configured to determine one or more independent spatial metadata frames to be independently encoded from the spatial metadata frame sequence based on the update period.

[0046] According to another aspect, a method for generating a bitstream is described, the bitstream indicating a downmixing signal and spatial metadata for generating a multichannel upmixing signal from the downmixing signal.The method can generate a downmixed signal from a multichannel input signal. Furthermore, the method may include determining one or more control settings based on one or more external settings; wherein the one or more external settings include a target data rate for the bitstream, and wherein the one or more control settings include a maximum data rate for spatial metadata. Additionally, the method may include determining spatial metadata from the multichannel input signal according to the one or more control settings.

[0047] According to another aspect, a method for determining a spatial metadata frame for generating a multichannel upmixed signal from a corresponding frame of the downmixed signal is described. The method may include determining multiple spectra from the current frame and the immediately following frame of a channel of the multichannel input signal. Furthermore, the method may include weighting the multiple spectra using a window function to obtain multiple weighted spectra. Additionally, the method may include determining a spatial metadata frame for the current frame of the channel of the multichannel input signal based on the multiple weighted spectra. The window function may depend on one or more of the following: the number of spatial parameter sets included within the spatial metadata frame, the presence of a transient in the current frame or the immediately following frame of the multichannel input signal, and / or the timing of that transient.

[0048] According to another aspect, a method for determining a spatial metadata frame for generating a frame of a multichannel upmix signal from a corresponding frame of an downmix signal is described. The method may include: determining a first plurality of transform coefficients from a frame of a first channel of the multichannel input signal, and determining a second plurality of transform coefficients from a corresponding frame of a second channel of the multichannel input signal. As outlined above, the first plurality of transform coefficients and the second plurality of transform coefficients typically provide a first time / frequency representation and a second time / frequency representation of the corresponding frames of the first and second channels, respectively. The first time / frequency representation and the second time / frequency representation may include a plurality of frequency intervals and a plurality of time intervals. The set of spatial parameters may include corresponding band parameters for different frequency bands comprising different numbers of frequency intervals. The method may further include determining a shift to be applied when determining a specific band parameter for a specific frequency band using fixed-point arithmetic. Furthermore, the shift may be determined based on the number of time intervals to be considered when determining the specific band parameter. Additionally, the method may include determining the specific band parameter using fixed-point arithmetic and the determined shift, based on the first plurality of transform coefficients and the second plurality of transform coefficients falling within the specific frequency band.

[0049] A method for generating a bitstream based on a multichannel input signal is described. The method may include generating a frame sequence of a downmixed signal from a corresponding first frame sequence of the multichannel input signal. Furthermore, the method may include determining a spatial metadata frame sequence from a second frame sequence of the multichannel input signal. The frame sequence of the downmixed signal and the spatial metadata frame sequence can be used to generate a multichannel upmixed signal.Additionally, the method may include generating a bitstream comprising a sequence of bitstream frames. The bitstream frames may indicate a frame corresponding to the first frame of a first frame sequence of the multichannel input signal and a spatial metadata frame corresponding to the second frame of a second frame sequence of the multichannel input signal. The second frame may be different from the first frame.

[0050] According to another aspect, a method for generating a bitstream based on a multichannel input signal is described. The method may include determining a trim protection gain sequence for a corresponding frame sequence of the multichannel input signal. The current trim protection gain may indicate that the current frame of the multichannel input signal will be applied to prevent attenuation of the corresponding current frame of the undermixed signal. The method may continue to interpolate the current trim protection gain and the previous trim protection gain of the previous frame of the multichannel input signal to obtain a trim protection gain curve. Furthermore, the method may include applying the trim protection gain curve to the current frame of the multichannel input signal to obtain a current frame of attenuation of the multichannel input signal. The current frame of the frame sequence of the undermixed signal may be generated from the current frame of attenuation of the multichannel input signal. Additionally, the method may include determining a spatial metadata frame sequence from a multi-channel input signal. The frame sequence of the downmixing signal and the spatial metadata frame sequence may be used to generate a multi-channel upmixing signal. A bitstream may be generated such that the bitstream indicates a trimmed guard gain sequence, a frame sequence of the downmixing signal, and a spatial metadata frame sequence, so that a multi-channel upmixing signal can be generated based on the bitstream.

[0051] According to another aspect, a method for generating a bitstream that indicates a downmixing signal and spatial metadata, the spatial metadata being used to generate a multi-channel upmixing signal from the downmixing signal, is described. The method may include generating a downmixing signal from a multi-channel input signal. Furthermore, the method may include determining one or more control settings based on one or more external settings, wherein the one or more external settings include an update period indicating a time period required for the decoding system to synchronize with the bitstream. The method may also include determining a spatial metadata frame sequence for a corresponding frame sequence of the multi-channel input signal based on one or more control settings. Additionally, the method may include encoding one or more spatial metadata frames in the spatial metadata frame sequence as independent frames according to the update period.

[0052] According to another aspect, a software program is described. The software program may be adapted to execute on a processor, and is adapted to perform the method steps outlined in this document when executed on a processor.

[0053] According to another aspect, a storage medium is described. The storage medium may include a software program that may be adapted to execute on a processor, and is adapted to perform the method steps outlined in this document when executed on a processor.

[0054] According to another aspect, a computer program product is described.The computer program product may include executable instructions for performing the method steps outlined in this document when executed on a computer.

[0055] It should be noted that the methods and systems including the preferred embodiments outlined herein can be used independently or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined herein can be combined arbitrarily. Specifically, the features of the claims can be combined with each other in any manner.

[0056] The invention will now be described by way of example with reference to the accompanying drawings, in which: FIG1 shows a generalized block diagram of an example audio processing system for performing spatial synthesis; FIG2 shows example details of the system of FIG1; FIG3 shows an example audio processing system for performing spatial synthesis similar to FIG1; FIG4 shows an example audio processing system for performing spatial analysis; FIG5a shows a block diagram of an example parametric multichannel audio coding system; FIG5b shows a block diagram of an example spatial analysis and coding system; FIG5c illustrates an example time-frequency representation of the frames of the channels of a multichannel audio signal; FIG5d illustrates an example time-frequency representation of multiple channels of a multichannel audio signal; FIG5e shows an example windowing applied to the transform unit of the spatial analysis and coding system shown in FIG5b; FIG6 shows a flowchart of an example method for reducing the data rate of spatial metadata; FIG7a illustrates an example transition scheme for spatial metadata performed at the decoding system; FIG7b to 7d illustrate example window functions applied to determine spatial metadata; Figure 8 shows a block diagram of an example processing path for a parametric multichannel codec system; Figures 9a and 9b show block diagrams of an example parametric multichannel audio coding system configured to perform trim protection and / or dynamic range control; Figure 10 illustrates an example method for compensating DRC parameters; and Figure 11 shows an example interpolation curve for trim protection. Detailed Description

[0057] As outlined in the introduction, this document relates to a multichannel audio coding system using a parametric multichannel representation. Hereinafter, an example multichannel audio coding and decoding (encoding / decoding) system is described. In the context of Figures 1 to 3, the decoder of the audio codec system is described as being able to use the received parametric multichannel representation to generate an n-channel upmix signal Y (typically, n>2) from a received m-channel downmix signal X (e.g., m=2). Subsequently, the encoder-related processing of the multichannel audio codec system is described. Specifically, how the parametric multichannel representation and the m-channel downmix signal can be generated from the n-channel input signal is described.

[0058] FIG1 illustrates a block diagram of an example audio processing system 100 configured to generate an overmix signal Y from a downmix signal X and a set of mixing parameters. Specifically, the audio processing system 100 is configured to generate the overmix signal based solely on the downmix signal X and the set of mixing parameters.From the bitstream P, the audio decoder 140 extracts the downmix signal X = [l 0 r0] T and a set of mixing parameters. In the illustrated example, the set of mixing parameters includes parameters α1, α2, α3, β1, β2, β3, g, k1, k2. The mixing parameters may be included in each mixing parameter data field in the bitstream P in quantized and / or entropy encoded form. The mixing parameters may be referred to as metadata (or spatial metadata), which is transmitted together with the encoded downmix signal X. In some instances of this disclosure, it has been explicitly indicated that some connection lines are adapted to transmit multichannel signals, wherein these lines are provided with cross lines adjacent to each number of channels. In the system 100 shown in FIG. 1, the downmix signal X includes m = 2 channels, and the upmix signal Y, as defined below, includes n = 6 channels (e.g., 5.1 channels).

[0059] The upmix stage 110, whose operation is parameterized to the mixing parameters, receives the downmix signal. The downmixing modification processor 120 modifies the downmixing signal through nonlinear processing and by forming a linear combination of downmixing channels to obtain a modified downmixing signal D=[d1 d2]T. The first mixing matrix 130 receives the downmixing signal X and the modified downmixing signal D, and outputs the upmixing signal Y=[lf ls rf rs c lfe]T by forming the following linear combination:

[0060] In the above linear combination, the mixing parameter α3 controls the contribution of the intermediate type signal (proportional to l0+r0) formed from the downmixing signal to all channels in the upmixing signal. The mixing parameter β3 controls the contribution of the side type signal (proportional to l0-r0) to all channels in the upmixing signal. Therefore, in practical applications, it is reasonable to expect that the mixing parameters α3 and β3 will have different statistical properties, which enables more efficient encoding. (For comparison, consider the reference parameterization (where the independent mixing parameters control the contribution of the downmix to the spatial left and right channels of the upmix, as per specification page 10 / 30, 15 CN 122369474 A), note that the statistical observables of such mixing parameters may not be significantly different.)

[0061] Returning to the linear combination shown in the above equations, further note that the gain parameters k1, k2 can depend on a single, shared mixing parameter in the bitstream P. Furthermore, the gain parameters can be normalized such that...

[0062] The contribution of the modified downmix to the spatial left and right channels of the upmix can be controlled by parameters β1 (the contribution of the first modified channel to the left channel) and β2 (the contribution of the second modified channel to the right channel), respectively. Furthermore, the contribution of each channel in the downmix to its spatially corresponding channel in the upmix can be individually controlled by changing the independent mixing parameter g. Preferably, the gain parameter g is non-uniformly quantized to avoid large quantization errors.

[0063] Referring now further to Figure 2, the downmixing modification processor 120 can perform the following linear combination (which is cross-mixing) of the downmixing channels in the second mixing matrix 121:

[0064] As indicated by the formula, the gain filling the second mixing matrix can be parameterized to depend on some of the mixing parameters encoded in the bitstream P. The processing performed by the second mixing matrix 121 yields an intermediate signal Z = [z1 z2]T, which is supplied to the decorrelector 122. Figure 1 shows an example of the decorrelector 122 comprising two sub-decorrelectors 123, 124, which can be configured identically (i.e., providing the same output in response to the same input) or differently. As an alternative, Figure 2 shows an example of all decorrelectable operations performed by a single unit 122, which outputs a preliminarily modified downmixing signal D'. The downmixing modification processor 120 in Figure 2 may also include an artifact attenuator 125. In an example embodiment, as outlined above, the artifact attenuator 125 is configured to detect tails in the intermediate signal Z and take corrective action by attenuating unwanted artifacts in the signal based on the position of the detected tail. This attenuation generates a modified downmixed signal D, which is output from the downmixing modification processor 120.

[0065] FIG3 illustrates a first mixing matrix 130 of a similar type to that shown in FIG1 and its associated transform stages 301, 302 and inverse transform stages 311, 312, 313, 314, 315, 316. The transform stages may include, for example, filter banks, such as quadrature mirror filter banks (QMF). Thus, signals upstream of transform stages 301, 302 are representations in the time domain, as are signals downstream of inverse transform stages 311, 312, 313, 314, 315, 316. Other signals are representations in the frequency domain. The time dependence of other signals may, for example, be expressed as block values ​​or discrete values ​​related to the time blocks to which the signal is segmented. Note that Figure 3 uses alternative notation compared to the matrix equations above; one could, for example, have correspondences such as XL0~l0, XR0~r0, YL~lf, YLS~lS, etc. Furthermore, the notation in Figure 3 emphasizes the distinction between the time-domain representation of the signal XL0(t) and the frequency-domain representation of the same signal XL0(f). It is understood that the frequency-domain representation is divided into time frames; therefore, it is a function of both time and frequency variables.

[0066] Figure 4 shows an audio processing system 400 for generating a downmixed signal X and mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for controlling the gain applied by the upmixing stage 110. This audio processing system 400 is typically located on the encoder side, for example, in a broadcast or recording device, while the system 100 shown in Figure 1 is typically deployed on the decoder side, for example, in a playback device.The downmixing stage 410 generates an m-channel signal X based on an n-channel signal Y. Preferably, the downmixing stage 410 operates on time domain representations of these signals. The parameter extractor 420 can generate values of mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 by analyzing the n-channel signal Y and taking into account the quantitative and qualitative properties of the downmixing stage 410. The mixing parameters may be vectors of frequency block values as indicated by the notation in Fig. 4, and may be further segmented into time blocks. In an example implementation, the downmixing stage 410 is time-invariant and / or frequency-invariant. Due to the time invariance and / or frequency invariance, there is typically no need for a communication connection between the downmixing stage 410 and the parameter extractor 420, and parameter extraction can be performed independently. This provides great flexibility for implementation. It also offers the possibility of reducing the overall delay of the system, because several processing steps can be executed in parallel. By way of example, Dolby Digital Plus format (or Enhanced AC‑3) can be used to encode the downmix signal X.

[0067] The parameter extractor 420 can learn the quantitative and / or qualitative properties of the downmixing stage 410 by accessing a downmix specification, which may specify one of the following: a set of gain values, an index of a predefined downmix mode for which predefined gains are available, etc. The downmix specification may be data in a memory preloaded into each of the downmixing stage 410 and the parameter extractor 420. Alternatively or additionally, the downmix specification may be transmitted from the downmixing stage 410 to the parameter extractor 420 via a communication line connecting these units. As a further alternative, each of the downmixing stage 410 and the parameter extractor 420 can access the downmix specification from a common data source, such as a memory (e.g., of the configuration unit 540 shown in Fig. 5a) in an audio processing system or in a metadata stream associated with the input signal Y.

[0068] Fig. 5a shows an example multi-channel encoding system 500 for encoding a multi-channel audio input signal Y 561 (comprising n channels) using a downmix signal X (comprising m channels, where m < n) and a parametric representation. The system 500 comprises a downmix encoding unit 510, which includes, for example, the downmixing stage 410 of Fig. 4. The downmix encoding unit 510 may be configured to provide an encoded version of the downmix signal X. The downmix encoding unit 510 may encode the downmix signal X using, for example, a Dolby Digital Plus encoder. Furthermore, the system 500 comprises a parameter encoding unit 510, which may include the parameter extractor 420 of Fig. 4.The parameter encoding unit 510 can be configured to quantize and encode the mixed parameter set α1, α2, α3, β1, β2, β3, g, k1 (also referred to as spatial parameters) to obtain encoded spatial parameters 562. As indicated above, parameter k2 can be determined from parameter k1. In addition, the system 500 may include a bitstream generation unit 530, which is configured to generate a bitstream P 564 from the encoded downmixed signal 563 and the encoded spatial parameters 562. The bitstream 564 can be encoded according to a predetermined bitstream syntax. Specifically, the bitstream 564 can be encoded in a format conforming to Dolby Digital Plus (DD+ or E-AC-3, Enhanced AC-3).

[0069] The system 500 may include a configuration unit 540, which is configured to determine one or more control settings 552, 554 for the parameter encoding unit 520 and / or the downmixing encoding unit 510. The one or more control settings 552, 554 can be determined based on one or more external settings 551 of system 500. For example, the one or more external settings 551 may include the total (maximum or fixed) data rate of bitstream 564. Configuration unit 540 can be configured to determine one or more control settings 552 based on the one or more external settings 551. The one or more control settings 552 for parameter encoding unit 520 may include one or more of the following: the maximum data rate of encoded spatial parameters 562. This control setting is referred to herein as metadata data rate setting.

[0070] The maximum number and / or a specific number of parameter sets to be determined by parameter encoding unit 520 for each frame of audio signal 561. This control setting is referred to herein as temporal resolution setting because it allows the temporal resolution to affect spatial parameters.

[0071] The number of parameter bands that parameter encoding unit 520 will determine for spatial parameters. This control setting is referred to herein as frequency resolution setting because it allows the frequency resolution to affect spatial parameters.

[0072] The resolution of the quantizer used to quantize spatial parameters. This control setting is referred to herein as quantizer setting.

[0073] The parameter encoding unit 520 may use one or more of the control settings 552 mentioned above for determining and / or encoding spatial parameters to be included in the bitstream 564. Typically, the input audio signal Y 561 is divided into a sequence of frames, wherein each frame includes a predetermined number of samples of the input audio signal Y 561. The metadata data rate setting may indicate the maximum number of bits available for encoding the spatial parameters of the frames of the input audio signal 561. The actual number of bits used to encode the spatial parameters 562 of the frames may be less than the number of bits allocated by the metadata data rate setting.The parameter encoding unit 520 can be configured to notify the configuration unit 540 about the actual number of bits 553 used, thereby enabling the configuration unit 540 to determine the number of bits available for encoding the downmix signal X. This number of bits can be transmitted to the downmix encoding unit 510 as a control setting 554. The downmix encoding unit 510 can be configured (e.g., using a multi-channel encoder such as Dolby Digital Plus) to encode the downmix signal X based on the control setting 554. In this way, bits that have not yet been used to encode spatial parameters can be used to encode the downmix signal.

[0074] FIG5b shows a block diagram of an example parameter encoding unit 520. The parameter encoding unit 520 may include a transformation unit 521 configured to determine the frequency representation of the input signal 561. Specifically, the transformation unit 521 may be configured to transform frames of the input signal 561 into one or more spectra, each spectrum including multiple frequency intervals. For example, the transformation unit 521 can be configured to apply a filter bank (e.g., a QMF filter bank) to the input signal 561. The filter bank can be a critical sampling filter bank. The filter bank can include a predetermined number of Q filters (e.g., Q = 64 filters). Thus, the transformation unit 521 can be configured to determine Q sub-band signals from the input signal 561, wherein each sub-band signal is associated with a corresponding frequency interval 571. For example, K sampled frames of the input signal 561 can be transformed into Q sub-band signals, wherein each sub-band signal has K / Q frequency coefficients. In other words, K sampled frames of the input signal 561 are transformed into K / Q spectra, wherein each spectrum includes Q frequency intervals. In a particular example, the frame length is K = 1536, the number of frequency intervals is Q = 64, and the number of spectra is K / Q = 24.

[0075] The parameter encoding unit 520 can include a banding unit 522, which is configured to group one or more frequency intervals 571 into frequency bands 572. The grouping of frequency intervals 571 to frequency bands 572 can depend on the frequency resolution setting 552. Table 1 illustrates example mappings of frequency intervals 571 to frequency bands 572, which can be applied by the banding unit 522 based on the frequency resolution setting 552. In the illustrated example, the frequency resolution setting 552 can indicate banding of frequency intervals 571 to 7, 9, 12, or 15 frequency bands. Banding typically models the psychoacoustic behavior of the human ear. As a result, the number of frequency intervals 571 for each frequency band 572 typically increases with increasing frequency.Instruction manual 13 / 30 pages 18 CN 122369474 A

[0076]

[0077] Table 1

[0078] The parameter determination unit 523 of the parameter encoding unit 520 (and specifically, the parameter extractor 420) can be configured to determine one or more sets of mixed parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for each frequency band 572. Therefore, the frequency band 572 can also be referred to as the parameter band. The mixed parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for the frequency band 572 can be referred to as the band parameters. Thus, the entire set of mixed parameters generally includes the band parameters for each frequency band 572. The band parameters can be applied to the mixing matrix 130 of FIG3 to determine the sub-band version of the decoded upmixed signal.

[0079] The number of mixed parameter sets for each frame determined by the parameter determination unit 523 can be indicated by the time resolution setting 552. For example, time resolution setting 552 can instruct one or two sets of hybrid parameters to be determined for each frame.

[0080] Figure 5c illustrates the determination of a set of hybrid parameters including parameters for multiple frequency bands 572. Figure 5c illustrates an example set of transform coefficients 580 derived from a frame of input signal 561. Transform coefficients 580 correspond to a specific time 582 and a specific frequency range 571. Frequency band 572 may include multiple transform coefficients 580 from one or more frequency ranges 571. As can be seen from Figure 5c, the transform of time-domain sampling of input signal 561 provides a time-frequency representation of a frame of input signal 561.

[0081] It should be noted that the set of hybrid parameters for the current frame can be determined based on the transform coefficients 580 of the current frame and possibly also based on the transform coefficients 580 of the immediately following frame (which is also referred to as the look-ahead frame).

[0082] The parameter determination unit 523 can be configured to determine the mixing parameters α1, α2, α3, β1, β2, β3, g, k1, k2 for each frequency band 572. If the time resolution setting is set to 1, all transform coefficients 580 of a specific frequency band 572 (current frame and forward-looking frame) can be considered for determining the mixing parameters for the specific frequency band 572. On the other hand, the parameter determination unit specification 14 / 30 page 19 CN 122369474 A 523 can be configured to determine two sets of mixing parameters for each frequency band 572 (e.g., when the time resolution setting is set to 2). In this case, the first half of the time of the transform coefficients 580 of the specific frequency band 572 (corresponding to, for example, the transform coefficients 580 of the current frame) can be used to determine the first set of mixing parameters, while the second half of the time of the transform coefficients 580 of the specific frequency band 572 (corresponding to, for example, the transform coefficients 580 of the forward-looking frame) can be considered for determining the second set of mixing parameters.

[0083] Generally, the parameter determination unit 523 can be configured to determine one or more sets of mixed parameters based on the transform coefficients 580 of the current frame and the forward-looking frame. A window function can be used to limit the effect of the transform coefficients 580 on the one or more sets of mixed parameters. The shape of the window function can depend on the number of sets of mixed parameters for each frequency band 572 and / or the nature of the current frame and / or the forward-looking frame (e.g., the presence of one or more transients). An example window function will be described in the context of Figures 5e and 7b to 7d.

[0084] It should be noted that the above can be applied to cases where the frame of the input signal 561 does not include a transient signal portion. The system 500 (e.g., the parameter determination unit 523) can be configured to perform transient detection based on the input signal 561. In the event that one or more transients are detected, one or more transient indicators 583, 584 can be set, wherein the transient indicators 583, 584 can identify the moment 582 of the corresponding transient. The transient indicators 583, 584 can also be referred to as sampling points of each set of mixed parameters. In the case of a transient, the parameter determination unit 523 can be configured to determine the set of mixed parameters based on the transform coefficients 580 starting from the moment of the transient (this is illustrated by the areas with different shaded lines in FIG. 5c). On the other hand, the transform coefficients 580 before the moment of the transient can be ignored, thereby ensuring that the set of mixed parameters reflects the multi-channel situation after the transient.

[0085] FIG. 5c illustrates the transform coefficients 580 of the channels of the multi-channel input signal Y 561. The parameter encoding unit 520 is generally configured to determine the transform coefficients 580 for the multiple channels of the multi-channel input signal 561. FIG. 5d shows example transform coefficients for the first 561-1 channel and the second 561-2 channel of the input signal 561. The frequency band p 572 includes frequency intervals 571 ranging from frequency index i to j. The transform coefficients 580 of the first channel 561-1 at time (or in the spectrum) q in frequency interval i can be referred to as aq,i. Similarly, the transform coefficient 580 of the second channel 561-2 at time (or in the spectrum) q, in the frequency interval i, can be referred to as bq,i. The transform coefficient 580 can be a complex number. The determination of the mixing parameters for the frequency band p may involve determining the energy and / or covariance of the first channel 561-1 and the second channel 561-2 based on the transform coefficient 580.For example, the covariance of the transform coefficients 580 of the first channel 561-1 and the second channel 561-2 in frequency band p with respect to the time interval [q, v] can be determined as follows:

[0086] The energy estimate of the transform coefficients 580 of the first channel 561-1 in frequency band p with respect to the time interval [q, v] can be determined as follows:

[0087] The energy estimate E2,2(p) of the transform coefficients 580 of the second channel 561-2 in frequency band p with respect to the time interval [q, v] can be determined in a similar manner.

[0088] Thus, the parameter determination unit 523 can be configured to determine one or more band parameter sets 573 for different frequency bands 572. The number of frequency bands 572 typically depends on the frequency resolution setting 552, while the number of mixed parameter sets for each frame typically depends on the time resolution setting 552. For example, the frequency resolution setting 552 may indicate the use of 15 frequency bands 572, while the time resolution setting 552 may indicate the use of 2 mixed parameter sets. In this case, the parameter determination unit specification (page 15 / 30, 20 CN 122369474 A 523) can be configured to determine two temporally distinct sets of mixed parameters, wherein each set of mixed parameters includes 15 band parameter sets 573 (i.e., mixed parameters for different frequency bands 572).

[0089] As indicated above, the mixed parameters for the current frame can be determined based on the transform coefficients 580 of the current frame and based on the transform coefficients 580 of the following preceding frame. The parameter determination unit 523 can apply a window to the transform coefficients 580 to ensure a smooth transition between the mixed parameters of consecutive frames in the frame sequence, and / or to account for destructive portions (e.g., transients) within the input signal 561. This is illustrated in FIG5e, which shows the K / Q spectra 589 of the current frame 585 and the immediately following frame 590 of the input audio signal 561 at corresponding K / Q consecutive moments 582. Furthermore, FIG5e shows an example window 586 used by the parameter determination unit 523. Window 586 reflects the effect of the K / Q spectra 589 of the current frame 585 and the immediately following frame 590 (referred to as the foreground frame) on the mixing parameters. As will be outlined in more detail below, window 586 reflects the case where the current frame 585 and the foreground frame 590 do not include any transients. In this case, window 586 ensures smooth inflation and deflation of the spectra 589 of the current frame 585 and the foreground frame 590, respectively, thus allowing for smooth evolution of the spatial parameters. Furthermore, Figure 5e shows example windows 587 and 588. The dashed window 587 reflects the effect of the K / Q spectra 589 of the current frame 585 on the mixing parameters of the previous frame. Additionally, the dashed window 588 reflects the effect of the K / Q spectra 589 of the immediately following frame 590 on the mixing parameters of the immediately following frame 590 (in the case of smooth interpolation).

[0090] Subsequently, the encoding unit 524 of the parameter encoding unit 520 can be used to quantize and encode the one or more sets of mixed parameters. The encoding unit 524 can apply various encoding schemes. For example, the encoding unit 524 can be configured to perform differential encoding of the mixed parameters. Differential encoding can be based on a time difference (for the same frequency band 572, the time difference between the current mixed parameter and the corresponding previous mixed parameter) or a frequency difference (the frequency difference between the current mixed parameter of the first frequency band 572 and the corresponding current mixed parameter of the adjacent second frequency band 572).

[0091] Furthermore, the encoding unit 524 can be configured to quantize the set of mixed parameters and / or the time difference or frequency difference of the mixed parameters. The quantization of the mixed parameters can depend on the quantizer setting 552. For example, the quantizer setting 552 can take two values, a first value indicating fine quantization and a second value indicating coarse quantization. Thus, the encoding unit 524 can be configured to perform fine quantization (with relatively low quantization error) or coarse quantization (with relatively increased quantization error) based on the quantization type indicated by the quantizer setting 552. The quantized parameters or parameter differences can then be encoded using entropy-based codes (such as Huffman codes). As a result, an encoded spatial parameter 562 is obtained. The number of bits 553 used for encoding the spatial parameter 562 can be transmitted to the configuration unit 540.

[0092] In an embodiment, the encoding unit 524 can be configured to first quantize different mixing parameters (within consideration of the quantizer setting 552) to obtain quantized mixing parameters. The quantized mixing parameters can then be entropy-coded (using, for example, Huffman codes). Entropy coding can then encode the quantized mixing parameters of a frame (regardless of previous frames), the frequency difference of the quantized mixing parameters, or the time difference of the quantized mixing parameters. Encoding of the time difference may not be used in the case of so-called independent frames, which are encoded independently of previous frames.

[0093] Therefore, the parameter encoding unit 520 can use a combination of differential coding and Huffman coding to determine the encoded spatial parameter 562. As outlined above, the encoded spatial parameters 562 can be included as metadata (also referred to as spatial metadata) in the bitstream 564 along with the encoded downmix signal 563. Differential coding and Huffman coding can be used for transmitting the spatial metadata to reduce redundancy and thus increase the available bit rate for encoding the downmix signal 563. Because Huffman codes are variable-length codes, the size of the spatial metadata can vary considerably depending on the statistics of the encoded spatial parameters 562 to be transmitted. The data rate required to transmit the spatial metadata is subtracted from the data rate available to the core codec (e.g., Dolby Digital Plus) for encoding the stereo downmix signal.To avoid compromising the audio quality of the downmixed signal, the number of bytes that may be spent sending spatial metadata for each frame is typically limited. This limit may be subject to encoder tuning considerations, which may be taken into account by configuration unit 540. However, due to the variable-length nature of the basic differential / Huffman coding of the spatial parameters, it is generally not guaranteed that the data rate limit (e.g., reflected in the metadata data rate setting 552) will not be exceeded without any further means.

[0094] In this document, a method for post-processing encoded spatial parameters 562 and / or spatial metadata including encoded spatial parameters 562 is described. The method 600 for post-processing spatial metadata is described in the context of FIG. 6. Method 600 can be applied when it is determined that the total size of a frame of spatial metadata exceeds, for example, a predefined limit indicated by metadata data rate setting 552. Method 600 aims to progressively reduce the amount of metadata. Reducing the size of spatial metadata typically also decreases the accuracy of the spatial metadata, and thus compromises the quality of the spatial image of the reproduced audio signal. However, method 600 generally guarantees that the total amount of spatial metadata does not exceed a predefined limit, and thus allows determining the trade-off between the improvement in overall audio quality between spatial metadata (used to regenerate the m-channel multichannel signal) and audio codec metadata (used to decode the encoded downmix signal 563). Furthermore, method 600 for post-processing spatial metadata can be implemented with relatively low computational complexity (compared to completely recalculating the encoded spatial parameters with modified control settings 552).

[0095] Method 600 for post-processing spatial metadata may include one or more of the following steps. As outlined above, each frame of spatial metadata may include multiple (e.g., one or two) sets of parameters, wherein the use of additional sets of parameters allows for increased temporal resolution of the mixing parameters. The use of multiple sets of parameters per frame can improve audio quality, especially in the case of attack-rich (i.e., transient) signals. Even with audio signals exhibiting relatively slow changes in spatial imagery, spatial parameter updates that are twice as large as a dense grid of sampling points can improve audio quality. However, transmitting multiple parameter sets per frame results in an approximately two-fold increase in data rate. Therefore, if it is determined that the data rate of the spatial metadata exceeds the metadata data rate setting 552 (step 601), it can be checked whether the spatial metadata frame includes more than one mixed parameter set. Specifically, it can be checked whether the metadata frame includes two mixed parameter sets that should be transmitted (step 602).If it is determined that the spatial metadata comprises multiple sets of mixed parameters, one or more of the sets exceeding a single set of mixed parameters can be discarded (step 603). As a result, the data rate of the spatial metadata can be significantly reduced (typically by half in the case of two sets of mixed parameters) while only impairing audio quality to a relatively low degree.

[0096] The decision of which of the two (or more) sets of mixed parameters to discard can depend on whether the encoding system 500 detects a transient location (“attack”) in the portion of the input signal 561 covered by the current frame: if multiple transients exist in the current frame, earlier transients are generally more important than later transients due to the psychoacoustic backmasking effect of each individual attack. Therefore, if a transient exists, it can be recommended to discard the later set of mixed parameters (e.g., the second of two). On the other hand, in the absence of an attack, the earlier set of mixed parameters (e.g., the first of two) can be discarded. This may be due to the windowing used when calculating the spatial parameters (as shown in Figure 5e). The window 586 used to window out the input signal 561 for calculating the spatial parameters for the second mixed parameter set typically has its greatest influence at the time point when the sampling points for parameter reconstruction are placed in the upper mixing stage 130 (i.e., at the end of the current frame). On the other hand, the first mixed parameter set typically receives a half-shift of the frame at that time point. Therefore, the error resulting from discarding the first mixed parameter set is most likely to be lower than the error resulting from discarding the second mixed parameter set. This is illustrated in Figure 5e, where it can be seen that the second half of the spectrum 589 of the current frame 585 used to determine the second mixed parameter set is more influenced by the sampling of the current frame 585 than the first half of the spectrum 589 of the current frame 585 (for the first half, the value of the window function 586 is lower than for the second half of the spectrum 589).

[0097] The spatial cue (i.e., mixing parameters) calculated in the encoding system 500 is sent to the corresponding decoder via a bitstream 562 (which may be part of a bitstream 564 in which the encoded stereo downmix signal 563 is delivered). Between the calculation of the spatial cue and its representation in the bitstream 562, the encoding unit 524 typically applies a two-step encoding method: the first step, quantization, is a lossy step because it introduces error into the spatial cue; the second step, differential / Huffman coding, is a lossless step.As outlined above, encoder 500 can choose between different types of quantization (e.g., two types of quantization): a high-resolution quantization scheme, which adds relatively little error but results in a larger potential quantization index, thus requiring a larger Huffman codeword; and a low-resolution quantization scheme, which adds relatively more error but results in a lower amount of quantization index, thus not requiring such a large Huffman codeword. It should be noted that different types of quantization can be applied to some or all of the mixed parameters. For example, different types of quantization can be applied to the mixed parameters α1, α2, α3, β1, β2, β3, k1. On the other hand, the gain g can be quantized using a fixed type of quantization.

[0098] Method 600 may include a step 604 of verifying which type of quantization has been used to quantize the spatial parameters. If it is determined that a relatively fine quantization resolution has been used, encoding unit 524 can be configured to reduce the quantization resolution to a lower type of quantization 605. As a result, the spatial parameters are quantized again. However, this does not increase significant computational overhead (compared to re-determining the spatial parameters using different control settings 552). It should be noted that different types of quantization can be used for different spatial parameters α1, α2, α3, β1, β2, β3, g, k1. Therefore, the coding unit 524 can be configured to select the quantizer resolution individually for each type of spatial parameter, thereby adjusting the data rate of the spatial metadata.

[0099] Method 600 may include a step of reducing the frequency resolution of the spatial parameters (not shown in FIG. 6). As outlined above, the set of mixed parameters of a frame is typically clustered into frequency bands or parameter bands 572. Each parameter band represents a frequency range, and for each band, a separate set of spatial clues is determined. The number of parameter bands 572 (e.g., 7, 9, 12, or 15 bands) can be gradually changed depending on the data rate available for transmitting the spatial metadata. The number of parameter bands 572 is approximately linearly related to the data rate, and therefore a reduction in frequency resolution can significantly reduce the data rate of the spatial metadata while only moderately affecting audio quality. However, such a reduction in frequency resolution typically requires recalculating the set of mixed parameters using the changed frequency resolution, and thus increases computational complexity.

[0100] As outlined above, encoding unit 524 may use differential coding of (quantized) spatial parameters. Configuration unit 551 may be configured to apply direct coding of the spatial parameters of the frames of the input audio signal 561 to ensure that transmission errors do not propagate over an infinite number of frames and to allow the decoder to synchronize with the received bitstream 562 at intermediate moments. Thus, a small portion of a frame may not use differential coding along the timeline. Such frames that do not use differential coding may be referred to as independent frames. Method 600 may include step 606 of verifying whether the current frame is an independent frame and / or whether the independent frame is a forced independent frame.The encoding of the spatial parameters may depend on the result of step 606.

[0101] As outlined above, differential coding is typically designed to compute differences between temporal successors or between adjacent frequency bands of quantized spatial cues. In both cases, the statistics of spatial cues make smaller differences occur more frequently than larger differences, and therefore, smaller differences are represented by shorter Huffman codewords compared to larger differences. In this document, smoothing of the quantized spatial parameters (temporally or frequency-wise) is proposed. Smoothing spatial parameters temporally or frequency-wise generally results in smaller differences and thus a reduction in data rate. Due to psychoacoustic considerations, temporal smoothing is generally preferred over frequency-wise smoothing. If it is determined that the current frame is not a forced independent frame, method 600 may proceed with temporal differential coding (step 607), possibly in conjunction with temporal smoothing. On the other hand, if the current frame is determined to be an independent frame, method 600 may proceed with frequency differential coding (step 608), and possibly with frequency smoothing.

[0102] The differential coding in step 607 can be submitted to temporal smoothing to reduce the data rate. The degree of smoothing can be varied depending on the amount by which the data rate will be reduced. The most severe type of temporal "smoothing" corresponds to keeping the previous set of mixed parameters unchanged, which corresponds to sending only incremental values ​​equal to zero. Temporal smoothing of differential coding can be performed on one or more (e.g., all) of the spatial parameters. Specification 18 / 30 pages 23 CN 122369474 A

[0103] In a similar manner to temporal smoothing, frequency smoothing can be performed. In its most extreme form, frequency smoothing corresponds to sending the same quantized spatial parameters over the entire frequency range of the input signal 561. While ensuring that the limits set by the metadata data rate settings are not exceeded, frequency smoothing can have a relatively high impact on the quality of the spatial image that can be reproduced using spatial metadata. Therefore, it may be preferable to apply frequency smoothing only when temporal smoothing is not permitted (e.g., if the current frame is a forced independent frame for which temporal differential coding of the previous frame is not available).

[0104] As outlined above, system 500 may operate subject to one or more external settings 551, such as the overall target data rate of bitstream 564 or the sampling rate of input audio signal 561. Typically, there is no single optimal operating point for all combinations of external settings. Configuration unit 540 may be configured to map effective combinations of external settings 551 to combinations of control settings 552, 554. For example, configuration unit 540 may rely on the results of psychoacoustic listening tests. Specifically, configuration unit 540 may be configured to determine combinations of control settings 552, 554 that ensure (on average) optimal psychoacoustic coding results for a particular combination of external settings 551.

[0105] As outlined above, the decoding system 100 should be able to synchronize with the received bitstream 564 within a given time period. To ensure this, the encoding system 500 may periodically encode so-called independent frames (i.e., frames that do not depend on knowledge of their predecessors). The average distance between two independent frames can be given by a ratio between the maximum delay for synchronization and the duration of a frame. This ratio does not necessarily have to be an integer, where the distance between two independent frames is always an integer of the frame.

[0106] The encoding system 500 (e.g., configuration unit 540) may be configured to receive the maximum delay or desired update time period for synchronization as an external setting 551. Furthermore, the encoding system 500 (e.g., configuration unit 540) may include a timer module configured to track the absolute amount of time elapsed since the first encoded frame of the bitstream 564. The first encoded frame of the bitstream 564 is, by definition, an independent frame. The encoding system 500 (e.g., configuration unit 540) may be configured to determine whether the next encoded frame includes a sample corresponding to a time that is an integer multiple of the desired update time period. Whenever the next encoded frame includes a sample at a time point that is an integer multiple of the desired update period, the encoding system 500 (e.g., configuration unit 540) can be configured to ensure that the next encoded frame is encoded as an independent frame. By doing so, it can be ensured that the desired update period is maintained even if the ratio of the desired update period to the frame length is not an integer.

[0107] As outlined above, the parameter determination unit 523 is configured to calculate spatial cues based on the time / frequency representation of the multichannel input signal 561. Spatial metadata frames can be determined based on the K / Q (e.g., 24) spectra 589 (e.g., QMF spectra) of the current frame and / or based on the K / Q (e.g., 24) spectra 589 (e.g., QMF spectra) of the forward-looking frame, wherein each spectrum 589 can have a frequency resolution of Q (e.g., 64) frequency intervals 571. Depending on whether the encoding system 500 detects a transient in the input signal 561, the time length of the signal portion used to calculate a single spatial cue set can include a different number of spectra 589 (e.g., 1 spectrum up to twice the number of K / Q spectra). As shown in Figure 5c, each spectrum 589 is divided into a certain number of frequency bands 572 (e.g., 7, 9, 12, or 15 bands), which, due to psychoacoustic considerations, include a different number of frequency intervals 571 (e.g., 1 frequency interval up to 41 frequencies). Different frequency bands p 572 and different time segments [q, v] define a grid on the time / frequency representation of the current frame and the forward-looking frame of the input signal 561.For different “boxes” in the grid, different sets of spatial cues can be calculated based on estimates of the energy and / or covariance of at least some of the input channels within each “box”. As outlined above, the energy estimates and / or covariance can be calculated by summing the squares of the transform coefficients 580 of a channel and / or by summing the products of the transform coefficients 580 of different channels (as indicated by the formulas provided above). The different transform coefficients 580 can be weighted according to the window function 586 used to determine the spatial parameters. Specification 19 / 30 pages 24 CN 122369474 A

[0108] The calculation of the energy estimates E1,1(p), E2,2(p) and / or covariance E1,2(p) can be implemented using fixed-point arithmetic. In this case, the different sizes of the “boxes” in the time / frequency grid may affect the arithmetic accuracy of the values ​​determined for the spatial parameters. As outlined above, the number of frequency intervals (j-i+1) 571 for each frequency band 572 and / or the length of the time interval [q,v] of the “boxes” of the time / frequency grid can be significantly varied (e.g., between 1×1×2 and 48×41×2 transform coefficients 580 (e.g., the real and complex parts of complex QMF coefficients)). Consequently, the number of products Re{at,f}Re{bt,f} and Im{at,f} Im{bt,f} that need to be summed to determine the energy E1,1(p) / covariance E1,2(p) can be significantly varied. To prevent the calculated results from exceeding a range that can be expressed in fixed-point arithmetic, the signal can be scaled down by a maximum number of bits (e.g., scaled down by 6 bits since 26·26=4096≥48·41·2). However, for smaller “boxes” and / or for “boxes” that only include relatively low signal energy, this method results in a significant reduction in arithmetic accuracy.

[0109] In this document, it is proposed that each “box” of the time / frequency grid uses a separate scale. The separate scale may depend on the number of transform coefficients 580 included within the “box” of the time / frequency grid. Typically, the spatial parameters for a particular “box” of the time / frequency grid (i.e., for a particular frequency band 572 and for a particular time interval [q, v]) are determined solely based on the transform coefficients 580 from that particular “box” (and not depending on the transform coefficients 580 from other “boxes”). Furthermore, the spatial parameters are typically determined solely based on the energy estimate and / or covariance ratio (and are generally unaffected by the absolute energy estimate and / or covariance). In other words, a single spatial line typically does not use the energy estimate and / or cross-channel product from a single time / frequency “box.” Furthermore, the spatial line is generally unaffected by the absolute energy estimate / covariance, but only by the energy estimate / covariance ratio. Therefore, a separate scale can be used for each individual “box.”The scaling should be matched for channels that contribute to a particular spatial cue.

[0110] For frequency band p 572 and for time interval [q, v], the energy estimates E1,1(p), E2,2(p) of the first channel 561-1 and the second channel 561-2, and the covariance E1,2(p) between the first channel 561-1 and the second channel 561-2, can be determined, for example, as indicated by the formula above. The energy estimates and covariance can be scaled by a scaling factor sp to provide scaled energy and covariance: sp·E1,1(p), sp·E2,2(p) and sp·E1,2(p). The spatial parameter P(p) derived from the energy estimates E1,1(p), E2,2(p) and covariance E1,2(p) generally depends on the ratio of energy and / or covariance such that the value of the spatial parameter P(p) is independent of the scaling factor sp. As a result, different scaling factors sp, sp+1, sp+2 can be used for different frequency bands p, p+1, p+2.

[0111] It should be noted that one or more of the spatial parameters may depend on more than two different input channels (e.g., three different channels). In this case, the one or more spatial parameters can be derived based on the energy estimates E1,1(p), E2,2(p)... of different channels, and based on the covariances between different pairs of channels (i.e., E1,2(p), E1,3(p), E2,3(p) etc.). And, in this case, the values ​​of the one or more spatial parameters are independent of the scaling factors applied to the energy estimates and / or covariances.

[0112] Specifically, the scaling factor sp = 2 - zp (where zp is a positive integer indicating the shift in fixed-point arithmetic) for a particular frequency band p can be determined such that

[0113] and such that the shift zp is minimized. By ensuring this individually for each frequency band p and / or for each time interval [q, v] for which the mixing parameters are determined, increased (e.g., maximum) precision in fixed-point arithmetic can be achieved while ensuring a valid range of values.

[0114] For example, individual scaling can be achieved by checking whether the result of each individual MAC (multiplicative-accumulator) operation exceeds + / - 1. Only if this is the case can the individual scaling for the “box” be increased by one bit. Once this has been done for all channels, the maximum scaling for each “box” can be determined, and all deviations in scaling for the “box” can be adjusted accordingly.

[0115] As outlined above, spatial metadata may include one or more (e.g., two) sets of spatial parameters per frame. Thus, encoding system 500 may send one or more sets of spatial parameters per frame to the corresponding decoding system 100.Each of these spatial parameter sets corresponds to a specific spectrum in the K / Q time-series of the spectrum 289 of the spatial metadata frame. This specific spectrum corresponds to a specific moment, and this specific moment may be referred to as a sampling point. Figure 5c shows two example sampling points 583 and 584 for the two spatial parameter sets, respectively. Sampling points 583 and 584 may be associated with specific events included within the input audio signal 561. Alternatively, the sampling points may be predetermined.

[0116] Sampling points 583 and 584 indicate the moment when the corresponding spatial parameters should be fully utilized by the decoding system 100. In other words, the decoding system 100 may be configured to update the spatial parameters at sampling points 583 and 584 according to the transmitted spatial parameter sets. Furthermore, the decoding system 100 may be configured to interpolate spatial parameters between two subsequent sampling points. The spatial metadata may indicate the type of transition to be performed between successive sets of spatial parameters. Examples of transition types are “smooth” and “steep” transitions between spatial parameters, meaning that the spatial parameters may be interpolated in a smooth (e.g., linear) manner or may be updated abruptly, respectively.

[0117] In the case of a “smooth” transition, the sampling point can be fixed (i.e., predetermined) and therefore does not need to be signaled in bitstream 564. If the spatial metadata frame delivers a single set of spatial parameters, the predetermined sampling point can be the position at the very end of the frame, i.e., the sampling point can correspond to the (K / Q)th spectrum 589. If the spatial metadata frame delivers two sets of spatial parameters, the first sampling point can correspond to the (K / 2Q)th spectrum 589, and the second sampling point can correspond to the (K / Q)th spectrum 589.

[0118] In the case of a “steep” transition, sampling points 583, 584 can be variable and can be signaled in bitstream 562. The portion of bitstream 562 carrying the following information can be referred to as the “framing” portion of bitstream 562: information about the number of spatial parameter sets used in a frame, information about the choice between “smooth” and “steep” transitions, and information about the position of the sampling point in the case of a “steep” transition. Figure 7a illustrates an example transition scheme that can be applied by the decoding system 100 based on the framing information included in the received bitstream 562.

[0119] For example, the framing information for a particular frame may indicate a “smooth” transition and a single set of spatial parameters 711. In this case, the decoding system 100 (e.g., the first mixing matrix 130) may assume that the sampling points of the spatial parameter set 711 correspond to the last spectrum of the particular frame. Furthermore, the decoding system 100 may be configured to perform (e.g., linear) interpolation 701 between the last received spatial parameter set 710 for the immediately preceding frame and the spatial parameter set 711 for the particular frame.In another example, the framing information for a particular frame may indicate a “smooth” transition and two spatial parameter sets 711, 712. In this case, the decoding system 100 (e.g., a first mixing matrix 130) may assume that the sampling points of the first spatial parameter set 711 correspond to the last spectrum of the first half of the particular frame, and the sampling points of the second spatial parameter set 712 correspond to the last spectrum of the second half of the particular frame. Furthermore, the decoding system 100 may be configured to perform (e.g., linear) interpolation 702 between the last received spatial parameter set 710 for the immediately preceding frame and the first spatial parameter set 711, and between the first spatial parameter set 711 and the second spatial parameter set 712.

[0120] In another example, the framing information for a particular frame may indicate a “steep” transition, a single spatial parameter set 711, and the sampling points 583 of that single spatial parameter set 711. In this scenario, decoding system 100 (e.g., first hybrid matrix 130) can be configured to apply the last received spatial parameter set 710 to the immediately preceding frame up to sample point 583, and apply spatial parameter set 711 from sample point 583 onwards (as shown by curve 703). In another example, the framing information for a particular frame may indicate a “steep” transition, two spatial parameter sets 711 and 712, and two corresponding sample points 583 and 584 for each of the two spatial parameter sets 711 and 712, respectively. In this case, decoding system 100 (e.g., hybrid matrix 130, page 21 / 30 of specification, CN 122369474 A) can be configured to apply the last received spatial parameter set 710 to the immediately preceding frame up to the first sample point 583, apply the first spatial parameter set 711 from the first sample point 583 up to the second sample point 584, and apply the second spatial parameter set 712 from the second sample point 584 at least until the end of the particular frame (as shown by curve 704).

[0121] The encoding system 500 shall ensure that the framing information matches the signal characteristics and that a suitable portion of the input signal 561 is selected to compute the one or more sets of spatial parameters 711, 712. For this purpose, the encoding system 500 may include a detector configured to detect signal locations in one or more channels where the signal energy suddenly increases. If at least one such signal location is found, the encoding system 500 may be configured to switch from a “smooth” transition to a “steep” transition; otherwise, the encoding system 500 may continue with a “smooth” transition.

[0122] As outlined above, the encoding system 500 (e.g., parameter determination unit 523) can be configured to calculate spatial parameters for the current frame based on multiple frames 585, 590 of the input audio signal 561 (e.g., based on the current frame 585 and the immediately following frame 590 (i.e., the so-called forward frame)). Thus, the parameter determination unit 523 can be configured to determine the spatial parameters based on twice the number of K / Q spectra 589 (as shown in FIG. 5e). As shown in FIG. 5e, the spectra 589 can be windowed using window 586. In this document, it is proposed to adjust the window 586 based on the number of spatial parameter sets 711, 712 to be determined, based on the transition type, and / or based on the position of sampling points 583, 584. By doing so, it can be ensured that the framing information matches the signal characteristics, and that appropriate portions of the input signal 561 are selected to calculate the one or more spatial parameter sets 711, 712.

[0123] Hereinafter, example window functions for different encoder / signal conditions are described: a) Case: Single spatial parameter set 711, smooth transition, no transient in the forward-looking frame 590; Window function 586: Between the last spectrum of the previous frame and the (K / Q)th spectrum 589, window function 586 can linearly increase from 0 to 1. Between the (K / Q)th spectrum 589 and the 48th spectrum 589, window function 586 can linearly decrease from 1 to 0 (see Figure 5e).

[0124] b) Case: Single spatial parameter set 711, smooth transition, transient exists in the Nth spectrum (N>K / Q), that is, transient exists in the forward-looking frame 590; Window function 721 as shown in Figure 7b: Between the last spectrum of the previous frame and the (K / Q)th spectrum, window function 721 linearly increases from 0 to 1. Between the (K / Q)th and (N-1)th spectra, window function 721 remains constant at 1. Between the Nth and (2K / Q)th spectra, window function remains constant at 0. The transient at the Nth spectra is represented by transient point 724 (which corresponds to the sampling point for the set of spatial parameters immediately following frame 590). Furthermore, Figure 7b shows complementary window function 722 (applied to the spectrum of the current frame 585 when determining the one or more sets of spatial parameters for the previous frame) and window function 723 (applied to the spectrum of the next frame 590 when determining the one or more sets of spatial parameters for the next frame). In summary, window function 721 ensures that, in the case of one or more transients in the preceding frame 590, the spectrum of the preceding frame before the first transient point 724 is adequately considered for determining the set of spatial parameters 711 for the current frame 585. On the other hand, the spectrum of the forward frame 590 following the transient point 724 is ignored.

[0125] c) Case: Single spatial parameter set 711, steep transition, transient in the Nth spectrum (N<=K / Q), no transient in subsequent frame 590.

[0126] Window function 731 as shown in FIG7c: Between the 1st spectrum and the (N-1)th spectrum, window function 731 remains constant at 0. Between the Nth spectrum and the (K / Q)th spectrum, window function 731 remains constant at 1. Between the (K / Q)th spectrum and the (2K / Q)th spectrum, window function 731 linearly decreases from 1 to 0. FIG7c indicates the transient point 734 at the Nth spectrum (which corresponds to the sampling point of the single spatial parameter set 711). Furthermore, Figure 7c shows window functions 732 and 733. Window function 732 is applied to the spectrum of the current frame 585 when determining the set of one or more spatial parameters for the previous frame, and window function 733 is applied to the spectrum of the next frame 590 when determining the set of one or more spatial parameters for the next frame.

[0127] d) Cases: Single set of spatial parameters, steep transition, transients in the Nth and Mth spectra (N<= K / Q, M>K / Q); Window function 741 in Figure 7d: Between the 1st and (N-1)th spectra, window function 741 remains constant at 0. Between the Nth and (M-1)th spectra, window function 741 remains constant at 1. Between the Mth and 48th spectra, window function 741 remains constant at 0. Figure 7d indicates the transient point 744 (i.e., the sampling point of the spatial parameter set) at the Nth spectrum and the transient point 745 at the Mth spectrum. Furthermore, Figure 7d shows window functions 742 and 743, where window function 742 is applied to the spectrum of the current frame 585 when determining the one or more spatial parameter sets for the previous frame, and window function 743 is applied to the spectrum of the next frame 590 when determining the one or more spatial parameter sets for the next frame.

[0128] e) Case: Two spatial parameter sets, smooth transition, no transient in subsequent frames; Window functions: i.) First spatial parameter set: Between the last spectrum of the previous frame and the (K / 2Q)th spectrum, the window linearly increases from 0 to 1. Between the (K / 2Q)th spectrum and the (K / Q)th spectrum, the window linearly decreases from 1 to 0. Between the (K / Q)th spectrum and the (2K / Q)th spectrum, the window remains constant at 0.

[0129] ii.) The second set of spatial parameters: Between the first spectrum and the (K / 2Q)th spectrum, the window remains constant at 0. Between the (K / 2Q)th spectrum and the (K / Q)th spectrum, the window linearly increases from 0 to 1. Between the (K / Q)th spectrum and the (3K / 2Q)th spectrum, the window linearly decreases from 1 to 0.Between the (3K / 2Q)-th spectrum and the (2K / Q)-th spectrum, the window remains constantly 0.

[0130] f) Case: two sets of spatial parameters, smooth transition, with a transient existing in the N-th spectrum (N>K / Q); window function: i.) First set of spatial parameters: between the last spectrum of the previous frame and the (K / 2Q)-th spectrum, the window rises linearly from 0 to 1. Between the (K / 2Q)-th spectrum and the (K / Q)-th spectrum, the window drops linearly from 1 to 0. Between the (K / Q)-th spectrum and the (2K / Q)-th spectrum, the window remains constantly 0.

[0131] ii.) Second set of spatial parameters: between the first spectrum and the (K / 2Q)-th spectrum, the window remains constantly 0. Between the (K / 2Q)-th spectrum and the (K / Q)-th spectrum, the window rises linearly from 0 to 1. Between the (K / Q)-th spectrum and the (N‑1)-th spectrum, the window remains constantly 1. Between the N-th spectrum and the (2K / Q)-th spectrum, the window remains constantly 0.

[0132] g) Case: two sets of spatial parameters, steep transition, with transients existing in the N-th spectrum and the M-th spectrum (N<M ≤K / Q), and no transient existing in the subsequent frame; window function: i.) First set of spatial parameters: between the first spectrum and the (N‑1)-th spectrum, the window remains constantly 0. Between the N-th spectrum and the (M‑1)-th spectrum, the window remains constantly 1. Between the M-th spectrum and the (2K / Q)-th spectrum, the window remains constantly 0.

[0133] ii.) Second set of spatial parameters: between the first spectrum and the (M‑1)-th spectrum, the window remains constantly 0. Between the M-th spectrum and the (K / Q)-th spectrum, the window remains constantly 1. Between the (K / Q)-th spectrum and the (2K / Q)-th spectrum, the window drops linearly from 1 to 0.

[0134] h) Case: two sets of spatial parameters, steep transition, with transients existing in the N-th, M-th and O-th spectra (N<M≤K / Q, O>K / Q); Description Page 23 of 30 28 CN 122369474 A window function: i.) First set of spatial parameters: between the first spectrum and the (N‑1)-th spectrum, the window remains constantly 0. Between the N-th spectrum and the (M‑1)-th spectrum, the window remains constantly 1. Between the M-th spectrum and the (2K / Q)-th spectrum, the window remains constantly 0.

[0135] ii.) Second set of spatial parameters: between the first spectrum and the (M‑1)-th spectrum, the window remains constantly 0. Between the M-th spectrum and the (O‑1)-th spectrum, the window remains constantly 1. Between the O-th spectrum and the (2K / Q)-th spectrum, the window remains constantly 0.

[0136] In general, the following example rules can be specified for the window function used to determine the current spatial parameter set: If the current spatial parameter set is not associated with a transient, - the window function provides a smooth rise in the spectrum from the sampling points of the previous spatial parameter set to the sampling points of the current spatial parameter set; - the window function provides a smooth fall in the spectrum from the sampling points of the current spatial parameter set to the sampling points of the next spatial parameter set, if the next spatial parameter set is not associated with a transient; - the window function fully considers the spectrum preceding the sampling points of the current spatial parameter set to the sampling points of the next spatial parameter set, and eliminates the spectrum starting from the sampling points of the next spatial parameter set, if the next spatial parameter set is associated with a transient; If the current spatial parameter set is associated with a transient, - the window function eliminates the spectrum preceding the sampling points of the current spatial parameter set; - The window function fully considers the spectrum from the sampling point of the current spatial parameter set to the spectrum preceding the sampling point of the next spatial parameter set, and eliminates the spectrum starting from the sampling point of the next spatial parameter set if the sampling point of the next spatial parameter set is associated with a transient; - The window function fully considers the spectrum from the sampling point of the current spatial parameter set to the spectrum at the end of the current frame, and provides a smooth fading of the spectrum from the beginning of the previous frame to the sampling point of the next spatial parameter set if the next spatial parameter set is not associated with a transient.

[0137] Hereinafter, a method for reducing latency in a parameterized multichannel codec system including an encoding system 500 and a decoding system 100 is described. As outlined above, the encoding system 500 includes several processing paths, such as downmixing signal generation and encoding, and parameter determination and encoding. The decoding system 100 typically performs decoding of the encoded downmixing signal and generation of the decorrelated downmixing signal. In addition, the decoding system 100 performs decoding of the encoded spatial metadata. Subsequently, the decoded spatial metadata is applied to the decoded downmixed signal and the decorrelated downmixed signal to generate an upmixed signal in the first upmixing matrix 130.

[0138] It is desirable to provide an encoding system 500 configured to provide a bitstream 564 that enables the decoding system 100 to generate the upmixed signal Y with reduced latency and / or reduced buffer memory. As outlined above, the encoding system 500 includes several different paths that can be aligned such that the encoded data provided to the decoding system 100 within the bitstream 564 matches correctly during decoding. As outlined above, the encoding system 500 performs downmixing encoding of the PCM signal 561. Furthermore, the encoding system 500 determines spatial metadata from the PCM signal 561. Additionally, the encoding system 500 can be configured to determine one or more trimming gains (typically one trimming gain per frame). The trimming gain indicates a trimming prevention gain that has been applied to the downmixed signal X to ensure that the downmixed signal X is not trimmed.The one or more trimming gains may be transmitted within the bitstream 564 (typically within a spatial metadata frame) to enable the decoding system 100 to regenerate the upmix signal Y. Additionally, the encoding system 500 may be configured to determine one or more Dynamic Range Control (DRC) values ​​(e.g., one or more DRC values ​​per frame). These one or more DRC values ​​may be used by the decoding system 100 to perform dynamic range control on the upmix signal Y. Specifically, the one or more DRC values ​​may ensure that the DRC performance of the parameterized multichannel codec system described herein is similar to (or equal to) the DRC performance of older multichannel codec systems such as Dolby Digital Plus. These one or more DRC values ​​may be transmitted within the downmix audio frame (e.g., within a suitable field of the Dolby Digital Plus bitstream).

[0139] Thus, the encoding system 500 may include at least four signal processing paths. To align these four paths, the encoding system 500 may also consider the delays introduced into the system by different processing components not directly related to the encoding system 500, such as core encoder delay, core decoder delay, spatial metadata decoder delay, LFE filter delay (used for filtering LFE channels), and / or QMF analysis delay.

[0140] To align the different paths, the delay of the DRC processing path can be considered. The DRC processing delay can typically only be aligned to the frame, rather than based on time-samples. Thus, the DRC processing delay typically depends only on the core encoder delay that can be rounded up to the next frame for alignment, i.e., DRC processing delay = round up(core encoder delay / frame size). Based on this, the downmixing processing delay used to generate the downmixed signal can be determined, since the downmixing processing delay can be delayed based on time samples, i.e., downmixing processing delay = DRC delay frame size - core encoder delay. As shown in Figure 8, the remaining delays can be calculated by summing the individual delay lines and by ensuring that the delays match at the decoder level.

[0141] By considering the different processing delays when writing bitstream 564, the processing power at decoding system 100 and memory can be reduced by delaying the resulting spatial metadata by one frame (reducing the number of input channels by 1536 4 bytes - 245 bytes) instead of delaying the encoded PCM data by 1536 samples (reducing the number of input channels by 11536) when copying is performed. As a result of the delay, all signal paths are accurately aligned by time sampling, not just roughly matched.

[0142] As outlined above, Figure 8 illustrates the different delays caused by example encoding system 500.The numbers in parentheses in Figure 8 indicate example delays based on the number of samples of the input signal 561. The encoding system 500 typically includes a delay 801 caused by filtering the LFE channels of the multi-channel input signal 561. Furthermore, a delay 802 (referred to as “clipgainpcmdelayline”) can be caused by determining a clipping gain (i.e., the DRC2 parameter below) that will be applied to the input signal 561 to prevent clipping of the downmixed signal. Specifically, this delay 802 can be introduced to synchronize the application of the clipping gain in the encoding system 500 with the application of the clipping gain in the decoding system 100. For this purpose, the input delay of the downmixing calculation (performed by the downmixing processing unit 510) can be made equal to the delay 811 (referred to as “coredecdelay”) of the decoder 140 of the downmixed signal. This means that, in the illustrated example, clipgainpcmdelayline = coredecdelay = 288 samples.

[0143] The downmixing processing unit 510 (which includes, for example, a Dolby Digital Plus encoder) delays the processing path of audio data (e.g., the downmix signal), but the downmixing processing unit 510 does not delay the processing path of spatial metadata and the processing path for DRC / trimmed gain data. Therefore, the downmixing processing unit 510 should delay the calculated DRC gain, trimmed gain, and spatial metadata. For DRC gain, this delay typically needs to be a multiple of one frame. The delay 807 of the DRC delay line (which is referred to as "drcdela yline") can be calculated as drcdela yline = ceil ((coreencela y + clipgainpcmdelayline) / frame_size) = 2 frames; where "coreencdelay" refers to the delay 810 of the encoder of the downmix signal.

[0144] The delay of the DRC gain can typically only be a multiple of the frame size. As a result, additional delays may need to be added in the downmixing processing path to compensate for this and round up to the next multiple of the frame size.The additional downmixing delay 806 (which is referred to as "dmxdelayline") can be determined by dmxdelayline + coreencdelay + clipgainpcmdelayline = drcdelayline frame_size; and dmxdelayline = drcdelayline frame_size - coreencdelay - clipgainpcmdelayline, such that dmxdelayline = 100.

[0145] When the spatial parameters are applied in the frequency domain (e.g., in the QMF domain) on the decoder side, the spatial parameters should be synchronized with the downmixing signal. In order to compensate for the fact that the encoder of the downmixing signal does not delay the spatial metadata frame, but delays the downmixing processing path, the input of the parameter extractor 420 should be delayed such that the following condition applies: dmxdelayline + coreencdelay + coredecdelay + aspdecanadelay = aspdelayline + qmfanadelay + framingdelay. In the above formula, "qmfanadelay" specifies the delay 804 caused by the transform unit 521, and "framingdelay" specifies the delay 805 caused by the windowing of the transform coefficients 580 and the determination of spatial parameters. As outlined above, the framing calculation uses two frames (the current frame and the looking-forward frame) as input. Due to the looking-forward, framing introduces a delay of exactly one frame length 805. Furthermore, the delay 804 is known such that the additional delay to be applied to the processing path used to determine spatial metadata is aspdelayline = dmxdelayline + coreencdelay + coredecdelay + aspdecanadelay - qmfanadelay - framingdelay = 1856. Because the delay is greater than one frame, the memory size of the delay line can be reduced by delaying the calculated bitstream instead of delaying the input PCM data, thereby providing aspbsdelayline = floor (aspdelayline / frame_size) = 1 frame (delay 809) and asppcmdelayline = aspdelayline - aspbsdelayline frame_size = 320 (delay 803).

[0146] After calculating the one or more trimming gains, the one or more trimming gains are provided to the bitstream generation unit 530.Therefore, the one or more trimming gains undergo a delay applied to the final bitstream by aspbsdelayline 809. Thus, the additional delay 808 for the trimming gains should be: clipgainbsdelayline + as pbsd ela yline = dm xd ela yline + coreencd ela y + cored ecd ela y, which provides: clipgainbsdelayline = dmxdelayline + coreencdelay + coredecdelay - aspbsdelayline = 1 frame. In other words, it should be ensured that the one or more trimming gains are provided to the decoding system 100 immediately after the corresponding frame of the downmixed signal is decoded, so that the one or more trimming gains can be applied to the downmixed signal before upmixing is performed in the upmixing stage 130.

[0147] Figure 8 illustrates further delays induced at the decoding system 100, such as delay 812 (referred to as "aspdecanadelay") caused by time-domain to frequency-domain transformations 301, 302 of the decoding system 100, delay 813 (referred to as "aspdecsyndelay") caused by frequency-domain to time-domain transformations 311 to 316, and further delay 814.

[0148] As can be seen from Figure 8, the different processing paths of the encoding / decoding system include processing-related delays or alignment delays, which ensure that different output data from different processing paths are available at the decoding system 100 when needed. Alignment delays (e.g., delays 803, 809, 807, 808, 806) are provided within the encoding system 500, thereby reducing the processing power and memory required at the decoding system 100. The total delays for different processing paths (excluding the LFE filter delay 801 applicable to all processing paths) are as follows: Downmixing processing path: the sum of delays 802, 806, and 810 = 3072, i.e., two frames; DRC processing path: delay 807 = 3072, i.e., two frames; Trim gain processing path: the sum of delays 808, 809, and 802 = 3360, which corresponds to the delay of the downmixing processing path in addition to the delay 811 of the decoder of the downmixed signal; Spatial metadata processing path: the sum of delays 802, 803, 804, 805, and 809 = 4000, which corresponds to the delay of the downmixing processing path in addition to the delay 811 of the decoder of the downmixed signal, and in addition to the delay 812 caused by the time-to-frequency transformation stages 301 and 302.

[0149] Therefore, it is ensured that DRC data is available at the decoding system 100 at time 821, trimmed gain data is available at time 822, and spatial metadata is available at time 823.

[0150] Furthermore, as can be seen from FIG8, the bitstream generation unit 530 can combine encoded audio data and spatial metadata that may be associated with different segments of the input audio signal 561. Specifically, it can be seen that the downmixing processing path, the DRC processing path, and the trimmed gain processing path have a delay of exactly two frames (3072 samples) until the output of the encoding system 500 (indicated by interfaces 831, 832, 833) (when delay 801 is ignored). The encoded downmixed signal is provided by interface 831, the DRC gain data is provided by interface 832, and the spatial metadata and trimmed gain data are provided by interface 833. Typically, the encoded downmixed signal and DRC gain data are provided in a conventional Dolby Digital Plus frame, while the trimmed gain data and spatial metadata can be provided in a spatial metadata frame (e.g., in an auxiliary field of the Dolby Digital Plus frame).

[0151] It can be seen that the spatial metadata processing path at interface 833 has a delay of 4000 samples (when delay 801 is ignored), which is different from the delay of other processing paths (3072 samples). This means that the spatial metadata frame may be associated with a segment of the input signal 561 that is different from the downmixed signal. Specifically, it can be seen that in order to ensure alignment at the decoding system 100, the bitstream generation unit 530 should be configured to generate a bitstream 564 comprising a sequence of bitstream frames, wherein the bitstream frames indicate the frame of the downmixed signal corresponding to the first frame of the multi-channel input signal 561 and the spatial metadata frame corresponding to the second frame of the multi-channel input signal 561. The first and second frames of the multi-channel input signal 561 may include the same number of samples. Nevertheless, the first and second frames of the multi-channel input signal 561 may be different from each other. Specifically, the first frame and the second frame may correspond to different sections of the multi-channel input signal 561. More specifically, the first frame may include samples preceding the samples of the second frame. For example, the first frame may include samples of the multi-channel input signal 561 that precede the samples of the second frame of the multi-channel input signal 561 by a predetermined number of samples (e.g., 928 samples).

[0152] As outlined above, the encoding system 500 may be configured to determine dynamic range control (DRC) and / or trim gain data. Specifically, the encoding system 500 may be configured to ensure that the downmixed signal X is not trimmed.Furthermore, the encoding system 500 can be configured to provide dynamic range control (DRC) parameters that ensure that the DRC behavior of a multichannel signal Y encoded using the parametric encoding scheme mentioned above is similar to or equal to the DRC behavior of a multichannel signal Y encoded using a reference multichannel encoding system (such as Dolby Digital Plus).

[0153] FIG9a shows a block diagram of an example dual-mode encoding system 900. It should be noted that portions 930, 931 of the dual-mode encoding system 900 are typically provided separately. An n-channel input signal Y 561 is provided to each of the upper portion 930 and the lower portion 931, the upper portion 930 being valid at least in the multichannel encoding mode of the encoding system 900, and the lower portion 931 being valid at least in the parametric encoding mode of the system 900. The lower portion 931 of the encoding system 900 may correspond to or may include, for example, the encoding system 500. The upper portion 930 may correspond to a reference multichannel encoder (such as a Dolby Digital Plus encoder). The upper part 930 generally includes a discrete-mode DRC analyzer 910 arranged in parallel with the encoder 911. Both the encoder 911 and the discrete-mode DRC analyzer 910 receive the audio signal Y 561 as input. Based on this input signal 561, the encoder 911 outputs the encoded n-channel signal, while the DRC analyzer 910 outputs one or more post-processing DRC parameters DRC1 to be applied to the decoder-side DRC. The DRC parameter DRC1 can be a "compr" gain (compressor gain) and / or a "dynrng" gain (dynamic range gain) parameter. The parallel outputs of the two units 910 and 911 are acquired by a discrete-mode multiplexer 912, which outputs a bitstream P. The bitstream P can have a predetermined syntax, such as Dolby Digital Plus syntax.

[0154] The lower portion 931 of the encoding system 900 includes a parametric analysis stage 922 arranged in parallel with a parametric mode DRC analyzer 921, which, like the parametric analysis stage 922, receives an n-channel input signal Y. The parametric analysis stage 922 may include a parameter extractor 420. Based on the n-channel audio signal Y, the parametric analysis stage 922 outputs one or more mixing parameters (as outlined above) (represented by a in both Figures 9a and 9b) and m-channel (1).

Claims

1. A method comprising: The audio processor receives multi-channel input audio signals and downmixing instructions, including a set of gain values. A first dynamic range control (DRC) value set is determined, the first DRC value set being configured to control the dynamic range of the output audio signal; A second set of DRC values ​​is determined, the second set of DRC values ​​being configured to prevent the multi-channel input audio signal from being trimmed during downmixing by the audio processor; The second set of DRC values ​​is applied to the multi-channel input audio signal to obtain a decayed multi-channel input audio signal; In response to the downmixing command, the attenuated multi-channel input audio signal is downmixed to obtain a downmixed signal; as well as The output audio signal is generated from the first set of DRC values ​​and the downmixed audio signal.

2. An apparatus comprising: One or more processors; A memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including the following: Receives multi-channel input audio signals and downmixing commands including a set of gain values; A first dynamic range control (DRC) value set is determined, the first DRC value set being configured to control the dynamic range of the output audio signal; A second set of DRC values ​​is determined, the second set of DRC values ​​being configured to prevent the multi-channel input audio signal from being trimmed during downmixing by the device; The second set of DRC values ​​is applied to the multi-channel input audio signal to obtain a decayed multi-channel input audio signal; In response to the downmixing command, the attenuated multi-channel input audio signal is downmixed to obtain a downmixed signal; as well as The output audio signal is generated from the first set of DRC values ​​and the downmixed audio signal.

3. The apparatus according to claim 2, wherein, Generating the output audio signal includes applying the first set of DRC values ​​to the downmixed audio signal.

4. The apparatus according to claim 2, wherein, The first set of DRC values ​​and / or the second set of DRC values ​​are represented as dB values ​​in logarithmic form.

5. The apparatus according to claim 2, wherein, The multi-channel input audio signal is divided into a frame sequence of samples of the multi-channel audio signal, and determining the first DRC value set and / or the second DRC value set includes determining the DRC value of each sample of each frame in the frame sequence.

6. The apparatus according to claim 5, wherein, Determining the DRC value for each sample of a frame involves interpolation between the DRC value of the frame and the DRC value of the previous frame.

7. The apparatus according to claim 6, wherein, The interpolation is spline interpolation.

8. The apparatus according to claim 2, wherein, The downmixed signal is a stereo signal.

9. The apparatus according to claim 2, wherein, The left and right channels of the downmix signal are generated based on different linear combinations of the channels of the multichannel input audio signal.

10. A non-transitory computer-readable storage medium comprising a sequence of instructions, the sequence of instructions causing the audio signal processing device to perform a method when executed by the audio signal processing device, the method comprising: The audio processor receives multi-channel input audio signals and downmixing instructions, including a set of gain values. A first dynamic range control (DRC) value set is determined, the first DRC value set being configured to control the dynamic range of the output audio signal; Determine a second set of DRC values, the second set of DRC values being configured to prevent the multichannel input audio signal from being clipped during downmixing by the audio processor; Apply the second set of DRC values to the multichannel input audio signal to obtain a attenuated multichannel input audio signal; In response to the downmixing instruction, downmix the attenuated multichannel input audio signal to obtain a downmixed signal; And Generate the output audio signal from the first set of DRC values and the downmixed audio signal.

11. A parameter processing unit (520) configured to determine a spatial metadata frame for generating a multichannel upmix signal from a corresponding frame of an undermix signal; wherein, The downmixed signal includes m channels, and wherein, the multichannel upmixed signal includes n channels; n, m are integers, where m < n; wherein, the spatial metadata frame includes one or more sets of spatial parameters (711, 712); the parameter processing unit (520) includes: - A transformation unit (521), the transformation unit (521) being configured to determine a plurality of spectra (589) from a current frame (585) and a following frame (590) of channels of a multichannel input signal (561); and - A parameter determination unit (523), the parameter determination unit (523) being configured to determine a spatial metadata frame for the current frame of the channels of the multichannel input signal (561) by weighting the plurality of spectra (589) using a window function (586); Wherein, the window function (586) depends on one or more of the following: the number of sets of spatial parameters (711, 712) included within the spatial metadata frame, the presence of one or more transients in the current frame or the following frame of the multichannel input signal (561), and / or the moment of the transients.

12. A parameter processing unit (520) configured to determine a spatial metadata frame for generating a multichannel upmix signal from a corresponding frame of an undermix signal; wherein, The downmixed signal includes m channels, and wherein, the multichannel upmixed signal includes n channels; n, m are integers, where m < n; wherein, the spatial metadata frame includes a set of spatial parameters (711); the parameter processing unit (520) includes: - A transformation unit (521), the transformation unit (521) being configured to: determine a first plurality of transformation coefficients (580) from a frame (585) of a first channel (561-1) of a multichannel input signal (561), and determine a second plurality of transformation coefficients (580) from a corresponding frame of a second channel (561-2) of the multichannel input signal (561); wherein, the first channel (561-1) and the second channel (561-2) are different; wherein, the first plurality of transformation coefficients (580) and the second plurality of transformation coefficients (580) respectively provide a first time / frequency representation and a second time / frequency representation of the frames (585) of the first channel and the second channel; wherein, the first time / frequency representation and the second time / frequency representation include a plurality of frequency intervals (571) and a plurality of time intervals (582); and - A parameter determination unit (523) configured to determine the set of spatial parameters (711) using fixed-point arithmetic based on the first plurality of transform coefficients (580) and the second plurality of transform coefficients (580); wherein the set of spatial parameters (711) includes corresponding band parameters for different frequency bands (572) including different numbers of frequency intervals (571); wherein a specific band parameter for the specific frequency band (572) is determined based on the transform coefficients (580) of the first plurality of transform coefficients (580) and the second plurality of transform coefficients (580) from the specific frequency band (572); and wherein the shift used in the fixed-point arithmetic for determining the specific band parameter depends on the specific frequency band (572).

13. An audio coding system (500) configured to generate a bitstream (564) based on a multichannel input signal (561); the system (500) comprising: - A downmix processing unit (510) configured to generate a sequence of frames of a downmix signal from a corresponding first sequence of frames of the multichannel input signal (561); wherein the downmix signal includes m channels, and wherein the multichannel input signal (561) includes n channels; n, m are integers, where m < n; - A parameter processing unit (520) configured to determine a sequence of spatial metadata frames from a second sequence of frames of the multichannel input signal (561); wherein the sequence of frames of the downmix signal and the sequence of spatial metadata frames are used to generate a multichannel upmix signal including n channels; and - A bitstream generation unit (503) configured to generate a bitstream (564) including a sequence of bitstream frames, wherein the bitstream frames indicate the frame of the downmix signal corresponding to the first frame of the first sequence of frames of the multichannel input signal (561) and the spatial metadata frame corresponding to the second frame of the second sequence of frames of the multichannel input signal (561); wherein the second frame is different from the first frame.

14. A method for determining a spatial metadata frame, the spatial metadata frame being used to generate a frame of a multichannel upmix signal from a corresponding frame of an downmix signal; wherein, The downmix signal includes m channels, and wherein the multichannel upmix signal includes n channels; n, m are integers, where m < n; wherein the spatial metadata frame includes one or more sets of spatial parameters (711, 712); the method comprising: - Determining a plurality of spectra (589) from the current frame (585) and the following frame (590) of the channels of the multichannel input signal (561); - Weighting the plurality of spectra (589) using a window function (586) to obtain a plurality of weighted spectra; and - Determining a spatial metadata frame for the current frame of the channels of the multichannel input signal (561) based on the plurality of weighted spectra. The window function (586) depends on one or more of the following: the number of sets of spatial parameters (711, 712) included in the spatial metadata frame, the presence of one or more transients in the current frame or the immediately following frame of the multichannel input signal (561), and / or the moments of the transients.

15. A method for determining a spatial metadata frame, the spatial metadata frame being used to generate a frame of a multichannel upmix signal from a corresponding frame of an downmix signal; wherein, The downmixed signal includes m channels, and wherein, the multichannel upmixed signal includes n channels; n and m are integers, where m < n; wherein, the spatial metadata frame includes a set of spatial parameters (711); the method includes: - Determining a first plurality of transform coefficients (580) from a frame (585) of a first channel (561-1) of the multichannel input signal (561); - Determining a second plurality of transform coefficients (580) from a corresponding frame of a second (561-2) channel of the multichannel input signal (561); wherein, the first channel (561-1) and the second channel (561-2) are different; Wherein, the first plurality of transform coefficients (580) and the second plurality of transform coefficients (580) respectively provide a first time / frequency representation and a second time / frequency representation of the frames (585) of the first channel and the second channel; wherein, the first time / frequency representation and the second time / frequency representation include a plurality of frequency intervals (571) and a plurality of time intervals (582); wherein, the set of spatial parameters (711) includes corresponding band parameters for different frequency bands (572) including different numbers of frequency intervals (571); - Determining a shift to be applied when using fixed-point arithmetic to determine a specific band parameter for a specific frequency band (572); wherein, the shift is determined based on the specific frequency band (572); and - Using fixed-point arithmetic and the determined shift, determining the specific band parameter based on the first plurality of transform coefficients (580) and the second plurality of transform coefficients (580) falling within the specific frequency band (572).

16. A method for generating a bitstream (564) based on a multichannel input signal (561); the method includes: - Generating a sequence of frames of a downmixed signal from a corresponding first sequence of frames of the multichannel input signal (561); wherein, the downmixed signal includes m channels, and wherein, the multichannel input signal (561) includes n channels; n and m are integers, where m < n; - Determining a sequence of spatial metadata frames from a second sequence of frames of the multichannel input signal (561); wherein, the sequence of frames of the downmixed signal and the sequence of spatial metadata frames are used to generate a multichannel upmixed signal including n channels; and - Generating a bitstream (564) including a sequence of bitstream frames; Wherein, the bitstream frames indicate the frame of the downmixed signal corresponding to the first frame of the first sequence of frames of the multichannel input signal (561) and the spatial metadata frame corresponding to the second frame of the second sequence of frames of the multichannel input signal (561); wherein, the second frame is different from the first frame.