Apparatus and method for stereo filling in multichannel coding
The proposed device and method address quantization-induced spectral holes in multichannel audio encoding by using a noise filling module to adaptively fill spectral lines, enhancing encoding quality and reducing noise artifacts in immersive audio scenarios.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
- Filing Date
- 2024-07-24
- Publication Date
- 2026-07-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing audio encoding methods struggle with quantization-induced spectral holes and inefficient stereo filling, particularly in immersive audio scenarios with dynamic channel setups, leading to noise artifacts and limited encoding quality at low bitrates.
A device and method for decoding and encoding multichannel audio signals that utilize a noise filling module to generate a mixing channel by selecting suitable pre-audio output channels and filling spectral lines with noise, based on inter-channel correlation and side information, to address quantization issues and improve encoding quality.
Enhances encoding quality by effectively filling spectral holes and adapting to time-varying inter-channel dependencies, reducing noise artifacts and improving audio fidelity in immersive audio scenarios.
Smart Images

Figure 0007885285000045 
Figure 0007885285000046 
Figure 0007885285000047
Abstract
Description
Technical Field
[0001] The present invention relates to audio signal encoding, and more particularly, to an apparatus and method for stereo filling in multi-channel encoding.
Background Art
[0002] Audio encoding is an area of compression that exploits the redundancy and irrelevance of audio signals.
[0003] In MPEG USAC (see, for example, [3]), joint stereo encoding of two channels is performed using complex prediction with band-limited or full-band residual signals, MPS 2-1-2, or integrated stereo. MPEG Surround (see, for example, [4]) hierarchically combines 1to2 (OTT) and 2to3 (TTT) boxes for joint encoding of multi-channel audio, with or without transmission of the residual signal.
[0004] In MPEG-H, the quad-channel element hierarchically applies the MPS 2-1-2 stereo box, followed by applying a complex prediction / MS stereo box that constructs a fixed 4×4 remix tree (see, for example, [1]).
[0005] AC4 (see, for example, [6]) introduces new 3-, 4-, and 5-channel elements, which enable remixing of the transmitted channels via the transmitted mix matrix and subsequent joint stereo encoding information. Furthermore, conventional publications have proposed using an orthogonal transform such as the Karhunen-Loeve transform (KLT) for enhanced multi-channel audio encoding (see, for example, [7]).
Summary of the Invention
Problems to be Solved by the Invention
[0006] For example, in the context of 3D audio, loudspeaker channels are distributed across several height layers, resulting in horizontal and vertical channel pairs. As defined in USAC, combined coding of only two channels is insufficient to account for the spatial and perceptual relationships between channels. MPEG surround is applied with additional pre- and post-processing steps, and residual signals are transmitted separately without the possibility of combined stereo coding that utilizes, for example, the dependency between left and right vertical residual signals. AC-4 dedicated N-channel elements are introduced to enable efficient coding of combined coding parameters, but fail for common speaker setups with more channels proposed for new immersive playback scenarios (7.1+4, 22.2). MPEG-H quad-channel elements are also limited to only four channels and cannot be dynamically applied to arbitrary channels, but can only be applied to a pre-configured fixed number of channels.
[0007] The MPEG-H multichannel coding tool enables the creation of discretely coded stereo boxes, i.e., arbitrary trees of combined coded channel pairs, see [2].
[0008] A common problem in audio signal encoding is caused by quantization, such as spectral quantization. Quantization can create spectral holes. For example, all spectral values within a particular frequency band may be set to zero on the encoder side as a result of quantization. For instance, the exact values of such spectral lines before quantization may be relatively low, and quantization can result in a situation where the spectral values of all spectral lines within a particular frequency band are set to zero. On the decoder side, this can create undesirable spectral holes during decoding.
[0009] Modern frequency-domain speech / audio coding systems such as the IETF[9] Opus / Celt codec, MPEG-4(HE-)AAC
[10] , or in particular MPEG-D xHE-AAC(USAC)
[11] , present means of encoding audio frames using either a long block, which is one long transform, or a short block, which is eight consecutive short transforms, depending on the temporal stationarity of the signal. Furthermore, for low-bitrate coding, these schemes provide tools for reconstructing the frequency coefficients of a channel using pseudo-random noise or low-frequency coefficients of the same channel. In xHE-AAC, these tools are called noise filling and spectral band duplication, respectively.
[0010] However, in the case of highly tonal or transient stereo inputs, the encoding quality achievable at very low bitrates is limited primarily by noise filling and / or spectral band duplication, because there are too many spectral coefficients in both channels that need to be transmitted clearly.
[0011] MPEG-H stereo filling is a parametric tool that relies on the use of a downmix of the previous frame to improve the filling of spectral holes by quantization in the frequency domain. Like noise filling, stereo filling operates directly in the MDCT domain of the MPEG-H corecoder, see [1], [5], [8].
[0012] However, the use of MPEG surround and stereo filling in MPEG-H is limited to fixed channel pair elements and therefore cannot utilize time-varying inter-channel dependencies.
[0013] While the Multi-Channel Coding Tool (MCT) in MPEG-H allows adaptation to changing inter-channel dependencies, it does not enable stereo filling because it uses single-channel elements in its normal operating configuration. Prior art has not disclosed a perceptually optimal method for generating a downmix of the previous frame for time-varying and arbitrary combined coded channel pairs. Using noise filling instead of stereo filling in combination with MCT to fill spectral holes can lead to noise artifacts, particularly in tonal signals.
[0014] The object of the present invention is to provide an improved concept of audio coding. The object of the present invention is solved by the decoding device described in claim 1, the coding device described in claim 15, the decoding method described in claim 18, the coding method described in claim 19, the computer program described in claim 20, and the coded multichannel signal described in claim 21. [Means for solving the problem]
[0015] A device is provided for decoding an encoded multichannel signal of the current frame in order to obtain three or more current audio output channels. The multichannel processing unit is adapted to select two decoded channels from three or more decoded channels, depending on a first multichannel parameter. Furthermore, the multichannel processing unit is adapted to generate a first group of two or more processed channels based on the selected channels. A noise filling module identifies one or more frequency bands in which all spectral lines are quantized to zero for at least one of the selected channels, and, depending on the side information, generates a suitable subset of the three or more decoded pre-audio output channels, and is adapted to fill the spectral lines of the frequency bands in which all spectral lines are quantized to zero with noise generated using the spectral lines of the mixing channels.
[0016] According to one embodiment, a device is provided for decoding a pre-encoded multichannel signal from a previous frame to obtain three or more previous audio output channels, and for decoding a currently encoded multichannel signal from the current frame to obtain three or more current audio output channels.
[0017] The device comprises an interface, a channel decoder, a multi-channel processing unit for generating three or more current audio output channels, and a noise filling module. The interface is adapted to receive the current encoded multichannel signal and to receive side information including the first multichannel parameters. A channel decoder is adapted to decode the currently encoded multichannel signal of the current frame and obtain a set of three or more decoded channels of the current frame. The multichannel processing unit is adapted to select a first selected pair of two decoded channels from a set of three or more decoded channels, depending on the first multichannel parameter.
[0018] Furthermore, the multi-channel processing unit is adapted to generate a first group of two or more processed channels based on the first selected pair of two decoded channels, and to obtain an updated set of three or more decoded channels.
[0019] Before the multichannel processing unit generates a first pair of two or more processed channels based on a first selected pair of two decoded channels, the noise filling module identifies one or more frequency bands in at least one of the two channels of the first selected pair of two decoded channels where all spectral lines are quantized to zero, and adapts to generate a mixing channel using two or more of the three or more pre-audio output channels, rather than all of them, and to fill the spectral lines of one or more frequency bands where all spectral lines are quantized to zero with noise generated using the spectral lines of the mixing channel, and the noise filling module adapts to select two or more pre-audio output channels to be used to generate the mixing channel from the three or more pre-audio output channels according to the side information.
[0020] A particular concept of an embodiment that may be used by a noise-filling module that specifies how to generate and fill noise is called stereo filling.
[0021] Furthermore, a device for encoding a multi-channel signal having at least three channels is provided.
[0022] The device includes an iterative processing unit adapted to calculate inter-channel correlation values between each pair of at least three channels in a first iteration step, in order to select a pair having the highest value or a pair having a value above a threshold, process the selected pair using a multi-channel processing operation to derive initial multi-channel parameters for the selected pair, and derive a first processed channel.
[0023] The iterative processing unit is adapted to use at least one of the processed channels to perform calculations, selections, and processing in a second iteration step to derive further multichannel parameters and a second processed channel.
[0024] Furthermore, the apparatus includes a channel encoder adapted to encode a channel resulting from an iterative process executed by an iterative processing unit in order to obtain an encoded channel.
[0025] Furthermore, the apparatus has an encoded channel, initial multi-channel parameters and further multi-channel parameters, and uses noise generated based on previously decoded audio output channels that have been previously decoded by a decoding device, to generate an encoded multi-channel signal having information indicating whether a decoding device should fill spectral lines in one or more frequency bands where all spectral lines are quantized to zero. The apparatus includes an output interface adapted to do so.
[0026] Furthermore, a method is provided for decoding a previous encoded multi-channel signal of a previous frame to obtain three or more previous audio output channels, and decoding a current encoded multi-channel signal of a current frame to obtain three or more current audio output channels. The method includes the following. - Receiving a current encoded multi-channel signal and receiving side information including first multi-channel parameters. - Decoding the current encoded multi-channel signal of the current frame to obtain a set of three or more decoded channels of the current frame. - Selecting a first selected pair of two decoded channels from the set of three or more decoded channels according to the first multi-channel parameters. - Generating a first group of two or more processed channels based on the first selected pair of two decoded channels, and obtaining an updated set of three or more decoded channels.
[0027] Before a first pair of two or more processed channels is generated based on the first selected pair of two decoded channels, the following steps are executed. - For at least one of the two channels of the first selected pair of two decoded channels, one or more frequency bands in which all spectral lines are quantized to zero are identified; a mixing channel is generated using two or more of the three or more pre-audio output channels, but not all of them; the spectral lines of one or more frequency bands in which all spectral lines are quantized to zero are filled using noise generated using the spectral lines of the mixing channel; and two or more pre-audio output channels are selected from the three or more pre-audio output channels to be used to generate the mixing channel, depending on the side information.
[0028] Furthermore, a method for encoding a multichannel signal having at least three channels is provided. This method includes the following: -In the first iteration step, in order to select the pair having the highest value or a pair having a value above a threshold, the inter-channel correlation value between each pair of at least three channels is calculated, the selected pair is processed using a multi-channel processing operation to derive initial multi-channel parameters for the selected pair, and a first processed channel is derived. - Using at least one of the processed channels, perform calculation, selection, and processing in a second iteration step to derive further multichannel parameters and a second processed channel. - Encoding the channels resulting from iterative processing performed by the iterative processing unit in order to obtain an encoded channel. - To generate an encoded multichannel signal having encoded channels, initial multichannel parameters and further multichannel parameters, and using noise generated based on a previously decoded audio output channel that was previously decoded by a decoder, the decoder has information indicating whether or not spectral lines in one or more frequency bands in which all spectral lines are quantized to zero should be filled by the decoder.
[0029] Furthermore, computer programs are provided, each computer program configured to perform one of the above methods when executed on a computer or signal processing unit, and each of the above methods is performed by one of the computer programs.
[0030] Furthermore, an encoded multichannel signal is provided. The encoded multichannel signal includes the encoded channel, multichannel parameters, and information indicating whether the decoder should fill in spectral lines in one or more frequency bands where all spectral lines are quantized to zero, using spectral data generated based on previously decoded audio output channels, which have been previously decoded by the decoder. Embodiments of the present invention will be described in more detail below with reference to the drawings. [Brief explanation of the drawing]
[0031] [Figure 1a] A decoding device according to one embodiment is shown. [Figure 1b] Another embodiment of a decoding device is shown. [Figure 2] A block diagram of a parametric frequency-domain decoder according to one embodiment of the present invention is shown. [Figure 3] To facilitate understanding of the decoder in Figure 2, a schematic diagram is shown illustrating the spectral sequence that forms the channel spectrogram of a multi-channel audio signal. [Figure 4] To facilitate understanding of the explanation in Figure 2, a schematic diagram showing the current spectrum from the spectrogram shown in Figure 3 is provided. [Figure 5a] A block diagram of a parametric frequency-domain audio decoder in another embodiment, where the downmix of the previous frame is used as the basis for inter-channel noise filling, is shown. [Figure 5b] A block diagram of a parametric frequency-domain audio decoder in another embodiment, where the downmix of the previous frame is used as the basis for inter-channel noise filling, is shown. [Figure 6] A block diagram of a parametric frequency-domain audio encoder according to one embodiment is shown. [Figure 7] This is a schematic block diagram of a device for encoding a multichannel signal having at least three channels according to one embodiment. [Figure 8] This is a schematic block diagram of a device for encoding a multichannel signal having at least three channels according to one embodiment. [Figure 9] A schematic block diagram of a stereo box according to one embodiment is shown. [Figure 10] This is a schematic block diagram of an apparatus for decoding an encoded multichannel signal having an encoded channel and at least two multichannel parameters, according to one embodiment. [Figure 11] A flowchart of a method for encoding a multichannel signal having at least three channels, according to one embodiment, is shown. [Figure 12] A flowchart of a method for decoding an encoded multichannel signal having an encoded channel and at least two multichannel parameters, according to one embodiment, is shown. [Figure 13] This shows a system according to one embodiment. [Figure 14] Scenario (a) demonstrates the generation of a composite channel for the first frame of the scenario, and scenario (b) demonstrates the generation of a composite channel for a second frame following the first frame according to one embodiment. [Figure 15] The indexing scheme for multichannel parameters according to the embodiment is shown. [Modes for carrying out the invention]
[0032] Elements that are equal or equivalent, or elements that have equal or equivalent functions, are indicated in the following description by their equal or equivalent reference numbers.
[0033] The following description provides several details to give a more complete description of embodiments of the present invention. However, it will be obvious to those skilled in the art that embodiments of the present invention can be carried out without these specific details. In other examples, well-known structures and apparatus are shown in block diagrams rather than details, in order to avoid obscuring embodiments of the present invention. Also, features of the different embodiments described below can be combined with each other unless otherwise noted.
[0034] Before describing the decoding apparatus 201 in Figure 1a, we will first describe noise filling for multi-channel audio coding. In embodiments, the noise filing module 220 in Figure 1a may be configured to perform, for example, one or more of the following techniques described with respect to noise filling for multi-channel audio coding.
[0035] Figure 2 shows a frequency-domain audio decoder according to one embodiment of the present invention. The decoder is generally denoted by reference numeral 10 and includes a scale factor band identification unit 12, an inverse quantization unit 14, a noise filling unit 16 and an inverse transform unit 18, as well as a spectral line extraction unit 20 and a scale factor extraction unit 22. Optional further elements that may be included in the decoder 10 include a complex stereo prediction unit 24, an MS (intermediate) decoder 26 and an inverse TNS (time noise shaping) filter tool, two examples 28a and 28b are shown in Figure 2. Furthermore, a downmix provider is shown and outlined below in more detail using reference numeral 30.
[0036] The frequency-domain audio decoder 10 in Figure 2 is a parametric decoder that supports noise filling by having a zero-quantized scale factor band filled with noise using the scale factor of that scale factor band as a means of controlling the level of noise that fills that scale factor band. Beyond this, the decoder 10 in Figure 2 represents a multi-channel audio decoder configured to reconstruct a multi-channel audio signal from an inbound data stream 30. However, Figure 2 focuses on the elements of the decoder 10 involved in the reconstruction of one of the multi-channel audio signals encoded in the data stream 30, outputting this (output) channel at output 32. Reference numeral 34 indicates that the decoder 10 may include further elements or may include several pipeline operation controls that play a role in reconstructing other channels of the multi-channel audio signal, and what is described below shows how the reconstruction of the target channel at output 32 of the decoder 10 interacts with the decoding of other channels.
[0037] The multichannel audio signal represented by the data stream 30 may include two or more channels. In the following description of embodiments of the present application, the case in which the multichannel audio signal is stereo and includes only two channels is focused, but in principle, the embodiments described below can be easily moved to alternative embodiments relating to multichannel audio signals and their encoding including three or more channels.
[0038] As will become clearer from the description of Figure 2 below, decoder 10 in Figure 2 is a transform decoder. That is, according to the encoding technique underlying decoder 10, channels are encoded in a transform domain, such as by using channel-wrapped transforms. Furthermore, depending on the creator of the audio signal, there exists a temporal phase in which the channels of the audio signal represent roughly the same audio content, shifted by small or decisive changes such as different amplitudes and / or phases, and the difference between channels represents an audio scene that allows for the virtual positioning of the audio source of the audio scene with respect to the virtual speaker position associated with the output channels of the multi-channel audio signal. However, in some other temporal phases, different channels of the audio signal may be more or less uncorrelated with each other, and may even represent completely different audio sources, for example.
[0039] To illustrate the potentially time-varying relationships between channels in an audio signal, the underlying audio codec of decoder 10 in Figure 2 allows for the time-varying use of different measurements to leverage channel redundancy. For example, MS coding allows switching between representing the left and right channels of a stereo audio signal directly and representing them as a pair of M (mid) and S (side) channels, representing the downmix of the left and right channels and their halved differences, respectively. That is, while the spectrograms of the two channels transmitted by the data stream 30 exist continuously in terms of spectral time, the meaning of these (transmitted) channels can change over time and with respect to the output channels, respectively.
[0040] Another inter-channel redundancy tool, complex stereo prediction, predicts frequency-domain coefficients or spectral lines of a given channel using lines at the spectrally identical position of another channel in the spectral domain. Further details will be discussed later.
[0041] To facilitate the following explanation of Figure 2 and its illustrated components, Figure 3 illustrates a possible way in which sample values for the spectral lines of two channels can be encoded into the data stream 30, as processed by the decoder 10 in Figure 2, for an exemplary case of a stereo audio signal represented by the data stream 30. In particular, the upper half of Figure 3 shows the spectrogram 40 of the first channel of the stereo audio signal, while the lower half of Figure 3 shows the spectrogram 42 of the other channel of the stereo audio signal. Again, it is worth noting that the “meaning” of spectrograms 40 and 42 may change over time, for example, due to time-varying switching between MS-encoded and un-MS-encoded regions. In the first example, spectrograms 40 and 42 relate to the M channel and S channel, respectively, and later, spectrograms 40 and 42 relate to the left and right channels. Switching between MS-encoded and un-encoded MS-encoded regions may be signaled in the data stream 30.
[0042] Figure 3 shows that spectrograms 40 and 42 can be encoded into data stream 30 with a time-varying spectral resolution. For example, both (transmitted) channels may be subdivided into a sequence of frames, indicated by curly braces 44, that are of equal length, non-overlapping, and adjacent to each other in a time-matched manner. As described above, the spectral resolution of spectrograms 40 and 42 as represented in data stream 30 can change over time. We assume beforehand that the spectral resolution of spectrograms 40 and 42 changes equally over time, but as will become clear from the following explanation, this simplification can be extended. The change in spectral resolution is signaled in units of frames 44 in data stream 30, for example. That is, the spectral resolution changes in units of frames 44. The change in spectral resolution of spectrograms 40 and 42 is achieved by switching the transformation length and the number of transformations used to describe spectrograms 40 and 42 within each frame 44. In the example in Figure 3, frames 44a and 44b illustrate frames in which one long transform was used to sample the channels of the audio signal, thereby resulting in the highest spectral resolution with one spectral line sample value for each spectral line for each such frame per channel. In Figure 3, the spectral line sample values are indicated using small crosses in boxes, which may be arranged in rows and columns to represent a spectral time grid, with each row corresponding to one spectral line and each column corresponding to a sub-interval of frame 44 corresponding to the shortest transform involved in the formation of spectrograms 40 and 42. In particular, Figure 3 shows that, for example, frame 44d, the frame may instead undergo a short-length continuous transform, resulting in a spectrum with reduced spectral resolution for frames like frame 44d over several subsequent times.Eight short transformations are used exemplarily on frame 44d, resulting in spectral time sampling of spectrograms 40 and 42 within frame 42d with spaced-out spectral lines, so that only every eight spectral lines are populated, but to transform frame 44d, each of the eight transformation windows' sample values or shorter transformation lengths are used. For illustrative purposes, other number of transformations for a frame, such as the use of two transformation lengths, may also be feasible, for example, half the transformation length of the long transformation for frames 44a and 44b, thereby resulting in a spectral time grid or sampling of spectrograms 40 and 42 where two spectral line sample values are obtained for every two spectral lines, one related to the preceding transformation and the other to the subsequent transformation.
[0043] In Figure 3, the transformation windows for frame-refined transformations are shown below each spectrogram using overlapping window-like lines. Temporal overlap is useful, for example, for the purpose of TDAC (Time-Domain Aliasing Cancellation).
[0044] Furthermore, although this can be carried out in other ways in the embodiments described below, Figure 3 shows a case where the switching between different spectral time resolutions for individual frames 44 is performed in such a way that for each frame 44, the same number of spectral line values, indicated by the small cross in Figure 3, result in spectrograms 40 and 42, where the difference lies simply in how the lines spectrally sample each spectral time tile corresponding to each frame 44, spanning temporally over the time of each frame 44, from zero frequency to the maximum frequency f max It spans spectrally up to that point.
[0045] Using the arrows in Figure 3, Figure 3 shows that, with respect to frame 44d, a similar spectrum may be obtained for all frames 44 by appropriately distributing spectral line sample values belonging to short transformation windows within one frame of one channel, up to the next occupied spectral line of the same frame, onto the unoccupied (empty) spectral lines within that frame. The spectrum thus obtained is hereafter referred to as the "interleaved spectrum". For example, in the interleaving of n transformations in one frame of one channel, the spectrally identical spectral line values of n short transformations follow each other before the set of spectrally identical spectral line values of n short transformations of spectrally succeeding spectral lines. An intermediate form of interleaving may also be feasible; instead of interleaving all spectral line coefficients of a frame, it may be possible to interleave only a suitable subset of spectral line coefficients of the short transformations of frame 44d. In any case, whenever the spectra of the two channel frames corresponding to spectrograms 40 and 42 are discussed, these spectra may refer to interleaved spectra or non-interleaved spectra.
[0046] To efficiently encode the spectral line coefficients representing spectrograms 40 and 42 via the data stream 30 sent to the decoder 10, the spectral line coefficients are quantized. To control quantization noise spectrally in time, the quantization step size is controlled via a scale factor set to a specific spectral time grid. In particular, in each sequence of spectra of each spectrogram, the spectral lines are grouped into spectrally continuous, non-overlapping scale factor groups. Figure 4 shows the spectrum 46 of spectrogram 40 and the same-time spectrum 48 from spectrogram 42 in its upper half. As shown, spectra 46 and 48 are subdivided into scale factor bands along the spectral axis f, grouping the spectral lines into non-overlapping groups. The scale factor bands are shown in Figure 4 using curly braces 50. For simplicity, it is assumed that the boundary between scale factor bands coincides between spectra 46 and 48, but this is not necessarily required.
[0047] That is, by encoding the data stream 30, the spectrograms 40 and 42 are subdivided into temporal sequences of spectra, each of which is spectrally subdivided into scale factor bands, and for each scale factor band, the data stream 30 encodes or transmits information about the scale factor corresponding to each scale factor band. The spectral line coefficients entering each scale factor band 50 can be quantized using their respective scale factors, or, as far as the decoder 10 is concerned, can be inversely quantized using the scale factor of the corresponding scale factor band.
[0048] Before returning to Figure 2 and its description, it is assumed below that the specially processed channel, which is one of the decodes containing specific elements of the decoder in Figure 2 except for 34, is the transmitted channel of the spectrogram 40, which, as mentioned above, can represent one of the left and right channels, the M channel, or the S channel, assuming that the multi-channel audio signal encoded in the data stream 30 is a stereo audio signal.
[0049] The spectral line extraction unit 20 is configured to extract spectral line coefficients of frame 44 from spectral line data, i.e., data stream 30, while the scale factor extraction unit 22 is configured to extract scale factors corresponding to each frame 44. For this purpose, the extraction units 20 and 22 can use entropy decoding. According to one embodiment, the scale factor extraction unit 22 is configured to sequentially extract scale factors of, for example, the spectrum 46 in Figure 4, i.e., scale factors of the scale factor band 50, from the data stream 30 using context-adaptive entropy decoding. The order of sequential decoding can follow, for example, a spectral order defined within the scale factor band from low frequency to high frequency. The scale factor extraction unit 22 may also use context-adaptive entropy decoding and may determine the context for each scale factor by depending on already extracted scale factors in the spectral neighborhood of the currently extracted scale factor, such as depending on the scale factor of the immediately preceding scale factor band. Alternatively, the scale factor extraction unit 22 can predict and decode the scale factor from the data stream 30 by using differential decoding, for example, while predicting the current decoded scale factor based on one of the previously decoded scale factors, such as the immediately preceding scale factor. Notably, this scale factor extraction process is agnostic with respect to scale factors belonging to scale factor bands that are exclusively populated by zero-quantized spectral lines, or populated by spectral lines where at least one is quantized to a non-zero value. Scale factors belonging to scale factor bands populated only by zero-quantized spectral lines can serve as a basis for predictions of subsequent decoded scale factors that may belong to scale factor bands populated by at least one non-zero spectral line, or they may be predicted based on previously decoded scale factors that may belong to scale factor bands populated by at least one non-zero spectral line.
[0050] For completeness only, it should be noted that the spectral line extraction unit 20 extracts spectral line coefficients such that the scale factor bandwidth 50 is similarly populated, for example, by using entropy coding and / or predictive coding. Entropy coding may use context adaptability based on spectral line coefficients near the spectral time of the currently decoded spectral line coefficient, and similarly, prediction may be spectral prediction, time prediction, or spectral time prediction that predicts the currently decoded spectral line coefficient based on previously decoded spectral line coefficients near that spectral time. To improve coding efficiency, the spectral line extraction unit 20 may be configured to perform decoding of spectral lines or line coefficients in tuples that collect or group spectral lines along the frequency axis.
[0051] Therefore, the output of the spectral line extraction unit 20 provides spectral line coefficients, for example, in spectral units, such as spectrum 46, which either collects all spectral line coefficients for the corresponding frame or, instead, collects all spectral line coefficients for a specific short transformation of the corresponding frame. The output of the scale factor extraction unit 22 outputs the corresponding scale factor for each spectrum.
[0052] The scale factor band identification unit 12 and the inverse quantization unit 14 have spectral line inputs coupled to the output of the spectral line extraction unit 20, and the inverse quantization unit 14 and the noise filling unit 16 have scale factor inputs coupled to the output of the scale factor extraction unit 22. The scale factor band identification unit 12 is configured to identify so-called zero-quantized scale factor bands in the current spectrum 46, i.e., scale factor bands in which all spectral lines are quantized to zero, such as the scale factor band 50c in Figure 4, and the remaining scale factor bands of the spectrum in which at least one spectral line is quantized to a non-zero value. In particular, in Figure 4, the spectral line coefficients are shown using the shaded region of Figure 4. In the spectrum 46, it can be seen that all scale factor bands except for the scale factor band 50b have at least one spectral line and the spectral line coefficients are quantized to a non-zero value. It will become clear later that zero-quantized scale factor bands such as 50d form the target of inter-channel noise filling, which will be further described below. Before proceeding, it should be noted that the scale factor band identification unit 12 may limit its identification to an appropriate subset of the scale factor band 50, such as a scale factor band above a specific starting frequency 52. In Figure 4, this may limit the identification procedure to scale factor bands 50d, 50e, and 50f.
[0053] The scale factor band identification unit 12 notifies the noise filling unit 16 on these scale factor bands, which are zero-quantized scale factor bands. The inverse quantization unit 14 uses the scale factor associated with the inbound spectrum 46 to inversely quantize or scale the spectral line coefficients of the spectral lines of the spectrum 46 according to the associated scale factor, i.e., the scale factor associated with the scale factor band 50. In particular, the inverse quantization unit 14 uses the scale factor associated with each scale factor band to inversely quantize and scale the spectral line coefficients that fall into each scale factor band. Figure 4 is to be interpreted as showing the results of the inverse quantization of spectral lines.
[0054] The noise filling unit 16 acquires information regarding the zero-quantized scale factor bands that form the target of subsequent noise filling, the inverse quantized spectrum, and the scale factors of at least these scale factor bands identified as zero-quantized scale factor bands, as well as information regarding signal transmission obtained from the data stream 30 for the current frame to determine whether inter-channel noise filling should be performed on the current frame.
[0055] The inter-channel noise filling process described in the following embodiments actually includes two types of noise filling: namely, the insertion of a noise floor 54 for all spectral lines quantized to zero regardless of their potential membership in any zero-quantized scale factor band, and the actual inter-channel noise filling procedure. This combination will be discussed later, but it should be emphasized that according to another embodiment, the noise floor insertion can be omitted. Furthermore, the signaling for noise filling on and off obtained from the current frame and from the data stream 30 can relate only to inter-channel noise filling, or a combination of both noise filling types can be controlled together.
[0056] With respect to noise floor insertion, the noise filling unit 16 can operate as follows. In particular, the noise filling unit 16 can use artificial noise generation, such as a pseudo-random number generator or other random number source, to fill spectral lines where the spectral line coefficients are zero. The level of the noise floor 54 inserted into the thus zero-quantized spectral lines can be set according to explicit signal transmission within the data stream 30 to the current frame or current spectrum 46. The "level" of the noise floor 54 can be determined, for example, using root mean square (RMS) or energy measurement.
[0057] Therefore, the insertion of the noise floor represents a kind of pre-filling of scale factor bands identified as zero-quantized, such as the scale factor band 50d in Figure 4. It also affects other scale factor bands that are not zero-quantized, the latter of which are further subject to inter-channel noise filling. As will be described later, the inter-channel noise filling process is to fill the zero-quantized scale factor bands to a level controlled by the scale factor of each zero-quantized scale factor band. The latter can be used directly for this purpose because all spectral lines of each zero-quantized scale factor band are quantized to zero. Nevertheless, the data stream 30 may include additional signaling of parameters for each frame or each spectrum 46, which is applied in common to the scale factor of all zero-quantized scale factor bands in the corresponding frame or spectrum 46, and when applied on the scale factor of the zero-quantized scale factor band by the noise filling unit 16, results in each zero-quantized scale factor band being individually filled to its respective level. That is, the noise filling unit 16 may use the same correction function to correct the scale factor of each individual scale factor band for each zero-quantized scale factor band of the spectrum 46, using the above-described parameters for the spectrum 46 of the current frame included in the data stream 30, thereby obtaining a target filling level for each zero-quantized scale factor band, which serves as a measure of the extent to which, in terms of energy or RMS, the inter-channel noise filling process should fill each zero-quantized scale factor band with (optionally selected) additional noise (in addition to the noise floor 54).
[0058] In particular, to perform inter-channel noise filling 56, the noise filling unit 16 takes a spectrally identical portion of the spectrum 48 of another channel, which is already largely or completely decoded, copies the obtained portion of the spectrum 48 to a zero-quantized scale factor band in which this portion is spectrally identical, and scales the resulting overall noise level in the zero-quantized scale factor band obtained by integration across the spectral lines of each scale factor band to be equal to the aforementioned filling target level obtained from the scale factor of the zero-quantized scale factor band. By this means, the tonality of the noise filled into each zero-quantized scale factor band is improved compared to artificially generated noise that forms the basis of the noise floor 54, and is also better than an uncontrolled spectral copy / duplicate from very low frequency lines in the same spectrum 46.
[0059] More precisely, the noise filling unit 16 places spectrally equivalent portions within the spectrum 48 of other channels for the current band such as 50d, and scales its spectral lines depending on the scale factor of the zero-quantized scale factor band 50d, the method which may optionally include any additional offset or noise factor parameter included in the data stream 30 for the current frame or spectrum 46, so that each zero-quantized scale factor band 50d is filled to a desired level as defined by the scale factor of the zero-quantized scale factor band 50d. In this embodiment, this means that the filling is performed in an additional manner to the noise floor 54.
[0060] According to a simplified embodiment, the resulting noise-filled spectrum 46 may be directly input to the inverse transformer 18, thereby obtaining the time-domain portion of the respective channel audio time signal for each transform window to which the spectral line coefficients of the spectrum 46 belong, and then combining these time-domain portions by an overlap addition process (not shown in Figure 2). That is, if the spectrum 46 is a non-interleaved spectrum and the spectral line coefficients belong to only one transform, the inverse transformer 18 performs the transform to result in one time-domain portion, and the leading and trailing ends of the time-domain portion may undergo an overlap addition process with the leading and trailing time-domain portions obtained by inverse transforming the leading and trailing transforms, for example, to achieve time-domain aliasing cancellation. However, if the spectrum 46 interleaves the spectral line coefficients of two or more consecutive transformations, the inverse transform unit 18 may perform separate inverse transforms on them to obtain one time-domain portion for each inverse transform, and these time-domain portions may be subjected to overlap addition with respect to preceding and succeeding time-domain portions of other spectra or frames, according to the temporal order defined between them.
[0061] However, for the sake of completeness, it should be noted that further processing can be performed on the noise-filled spectrum. As shown in Figure 2, the inverse TNS filter can perform inverse TNS filtering on the noise-filled spectrum. That is, controlled via TNS filter coefficients for the current frame or spectrum 46, the spectrum obtained so far undergoes linear filtering along the spectral direction.
[0062] With or without inverse TNS filtering, the complex stereo prediction unit 24 can treat the spectrum as a prediction residual of inter-channel prediction. More specifically, the inter-channel prediction unit 24 can use spectrally identical portions of other channels to predict the spectrum 46 or at least a subset of its scale factor band 50. The complex prediction process is shown in Figure 4 with a dashed box 58 in relation to the scale factor band 50b. That is, the data stream 30 can include, for example, inter-channel prediction parameters that control which of the scale factor bands 50 are to be inter-channel predicted and which are not. Furthermore, the inter-channel prediction parameters in the data stream 30 may further include complex inter-channel prediction factors applied by the inter-channel prediction unit 24 to obtain inter-channel prediction results. These factors may be included in the data stream 30 individually for each scale factor band in which inter-channel prediction is activated or signaled, or alternatively, individually for each group of one or more scale factor bands.
[0063] The source of the inter-channel prediction may be the spectrum 48 of another channel, as shown in Figure 4. More precisely, the source of the inter-channel prediction may be a spectrally identical portion of the spectrum 48 that is in the same position as the inter-channel predicted scale factor band 50b, extended by the estimation of its imaginary part. The estimation of the imaginary part may be performed based on a spectrally identical portion 60 of the spectrum 48 itself, and / or using a downmix of the already decoded channel of the previous frame, i.e., the frame immediately preceding the current decoded frame to which spectrum 46 belongs. In short, the inter-channel prediction unit 24 adds the prediction signal obtained as described above to the inter-channel predicted scale factor band, such as the scale factor band 50b in Figure 4.
[0064] As already stated in the above explanation, the channel to which spectrum 46 belongs may be an MS-coded channel, or it may be a speaker-related channel such as the left or right channel of a stereo audio signal. Therefore, optionally, the MS decoder 26 may optionally perform MS decoding on the inter-channel predicted spectrum 46, and in that MS decoding, it may perform addition or subtraction for each spectral line or spectrum 46 with the spectrally corresponding spectral line of another channel corresponding to spectrum 48. For example, spectrum 48, as shown in Figure 4 (though not shown in Figure 2), is obtained by part 34 of the decoder 10 in a similar manner to that previously described with respect to the channel to which spectrum 46 belongs, and when the MS decoding module 26 performs MS decoding, it adds or subtracts spectral lines on spectra 46 and 48, meaning that both spectra 46 and 48 are at the same stage in the processing line, for example, both have just been obtained by inter-channel prediction, or both have just been obtained by noise filling or inverse TNS filtering.
[0065] It should be noted that, optionally, MS decoding may be performed comprehensively over the entire spectrum 46, or it may be activated individually by the data stream 30, for example, in units of scale factor bandwidth 50. In other words, MS decoding may be switched on or off in the data stream 30 using each signal transmission, for example, in units of frames or individually for scale factor bandwidths of spectra 46 and / or 48 of spectrograms 40 and / or 42, assuming that the same boundary of scale factor bandwidth is defined for both channels.
[0066] As shown in Figure 2, inverse TNS filtering by the inverse TNS filter 28 can also be performed after any inter-channel processing, such as inter-channel prediction 58 or MS decoding by the MS decoder 26. The performance before or downstream of inter-channel processing may be fixed or controlled for each frame in the data stream 30, or at some other granularity, via their respective signal transmissions. Whenever inverse TNS filtering is performed, each TNS filter coefficient present in the data stream of the current spectrum 46 controls the TNS filter, i.e., a linear prediction filter operating along the spectral direction, to linearly filter the inbound spectrum to each inverse TNS filter module 28a and / or 28b.
[0067] Therefore, the spectrum 46 arriving at the input of the inverse transform unit 18 may undergo further processing as described above. Again, the above explanation is not intended to be understood as meaning that all of these optional tools must be present simultaneously or not simultaneously. These tools may be present in the decoder 10 partially or collectively.
[0068] In any case, the resulting spectrum at the input of the inverse transform represents the final reconstruction of the channel's output signal and forms the basis for the aforementioned downmix to the current frame, serving as the basis for the estimation of the potential imaginary part of the next frame to be decoded, as explained with respect to complex prediction 58. It can further serve as the final inter-channel reconstruction for predicting another channel that is not the channel to which the elements other than 34 in Figure 2 are related.
[0069] Each downmix is formed by the downmix provider 31 by combining this final spectrum 46 with each final version of spectrum 48. The latter entities, i.e., each final version of spectrum 48, formed the basis for the complex inter-channel prediction in the prediction unit 24.
[0070] As long as the basis for inter-channel noise filling is represented by a downmix of spectral lines at the same spectral location in the previous frame, Figure 5 shows an alternative to Figure 2, and in the optional case of using complex inter-channel prediction, the source of this complex inter-channel prediction is used twice: once as the source for inter-channel noise filling and once as the source for imaginary part estimation in the complex inter-channel prediction. Figure 5 shows the decoder 10 including the internal structure of part 70 related to decoding the first channel to which spectrum 46 belongs, and the aforementioned other part 34 involved in decoding the other channel, which includes spectrum 48. The same reference numerals are used for the internal elements of part 70 on the one hand and part 34 on the other hand. As can be understood, the configuration is the same. At output 32, one channel of the stereo audio signal is output, and at the output of the inverse transformer 18 of the second decoder part 34, the other (output) channel of the stereo audio signal is obtained, and this output is indicated by reference numeral 74. Again, the embodiments described above can be easily adapted when using three or more channels.
[0071] The downmix provider 31 is shared by both sections 70 and 34 and receives spectra 48 and 46 that are at the same temporal position in spectrograms 40 and 42. It forms a downmix based on these spectra by summing them line by line, and optionally forms an average by dividing the sum at each spectral line by the number of channels being downmixed, i.e., 2 in the case of Figure 5. The output of the downmix provider 31 is obtained by this measurement of the downmix of the previous frame. In this regard, it should be noted that if the previous frame contains two or more spectra in either spectrogram 40 or 42, there are different possibilities regarding how the downmix provider 31 operates in that case. For example, in this case, the downmix provider 31 may use the spectrum of the subsequent transformation of the current frame, or it may use the interleaved result of interleaving all spectral line coefficients of the current frame in spectrograms 40 and 42. The delay element 74 shown in Figure 5, connected to the output of the downmix provider 31, indicates that the downmix provided at the output of the downmix provider 31 forms the downmix of the previous frame 76 (see Figure 4 for inter-channel noise filling 56 and complex prediction 58, respectively). Therefore, one output of the delay element 74 is connected to the input of the inter-channel prediction unit 24 of the decoder sections 34 and 70, and the other is connected to the input of the noise filling unit 16 of the decoder sections 70 and 34.
[0072] That is, in Figure 2, the noise filling unit 16 receives the spectrally identically positioned, ultimately reconstructed spectrum 48 of the other channel in the same current frame as the basis for inter-channel noise filling, whereas in Figure 5, inter-channel noise filling is instead performed based on a downmix of the previous frame, such as that provided by the downmix provider 31. The method by which inter-channel noise filling is performed is the same. That is, in the case of Figure 2, the inter-channel noise filling unit 16 takes spectrally identical portions from each spectrum of the other channel in the current frame, and in the case of Figure 5, it takes the almost or completely decoded final spectrum obtained from the previous frame representing the downmix of the previous frame, and further adds the same “source” portion to the spectral lines in the scale factor band to be noise-filled, such as 50d in Figure 4, scaled according to the target noise level determined by the scale factor of each scale factor band.
[0073] Concluding the above discussion of embodiments illustrating inter-channel noise filling in audio decoders, it will be apparent to readers of the art that certain preprocessing can be applied to the “source” spectral lines without deviating from the general concept of inter-channel filling, before adding the captured spectrally or temporally identical portions of the “source” spectrum to the spectral lines of the “target” scale factor band. In particular, to improve the audio quality of the inter-channel noise filling process, it may be beneficial to apply filtering operations, such as spectral flattening or slope removal, to the spectral lines of the “source” region added to the “target” scale factor band, as in 50d in Figure 4. Similarly, and also as an example of a nearly (instead of completely) decoded spectrum, the aforementioned “source” portion can be obtained from a spectrum that has not yet been filtered by an available inverse (i.e., composite) TNS filter.
[0074] Thus, the above embodiments concerned the concept of inter-channel noise filling. Below, we will explain how the above concept of inter-channel noise filling can be incorporated into an existing codec, namely xHE-AAC, in a semi-backward compatible manner. In particular, a preferred implementation of the above embodiments in which the stereo filling tool is incorporated into an xHE-AAC-based audio codec in a semi-backward compatible signaling scheme is described below. By using the embodiments further described below, stereo filling of the conversion coefficients of either one of the two channels in an MPEG-D xHE-AAC (USAC) based audio codec is possible, thereby improving the encoding quality of certain audio signals, especially at low bitrates. The stereo filling tool is signaled in a semi-backward compatible manner so that a legacy xHE-AAC decoder can analyze and decode the bitstream without obvious audio errors or omissions. As already mentioned above, better overall quality can be obtained if the audio coder can use a combination of previously decoded / quantized coefficients of the two stereo channels to reconstruct the zero-quantized (not transmitted) coefficient of either one of the currently decoded channels. In audio coders, particularly xHE-AAC or coders based thereon, it is desirable to enable stereo filling (from previous channel coefficients to current channel coefficients), in addition to spectral band duplication (from low-frequency channel coefficients to high-frequency channel coefficients) and noise filling (from uncorrelated pseudo-random sources).
[0075] To enable encoded bitstreams using stereo filling to be read and parsed by legacy xHE-AAC decoders, the desired stereo filling tool should be used in a semi-backward compatible manner, and its presence should not cause—or even initiate—decoding by legacy decoders. The readability of the bitstream by the xHE-AAC infrastructure also facilitates market introduction.
[0076] To achieve the requirements for quasi-backward compatibility with respect to the stereo-filling tool, as described above in the context of xHE-AAC or its potential derivatives, the following embodiments include the functionality of stereo-filling and the ability to signal the functionality of stereo-filling via syntax in the data stream that is actually related to noise-filling. The stereo-filling tool operates in accordance with the above description. When the stereo-filling tool is activated as an alternative to noise-filling (or in addition to noise-filling as described above) in a channel pair having a common window configuration, the coefficients of the zero-quantized scale factor band are reconstructed by the sum or difference of the coefficients of the previous frame in either of the two channels, preferably in the right channel. Stereo-filling is performed in the same way as noise-filling. Signal transmission is performed via the noise-filling signal transmission of xHE-AAC. Stereo-filling is transmitted by 8 bits of noise-filling side information. This is achievable as described in the MPEG-D USAC standard [3], where all 8 bits are transmitted even if the applied noise level is zero. In such situations, some of the noise-filling bits can be reused for the stereo-filling tool.
[0077] Near backward compatibility with bitstream analysis and playback by the legacy xHE-AAC decoder is guaranteed as follows: Stereo filling is signaled via a zero noise level (i.e., the first three noise filling bits, all with zero values) followed by five non-zero bits containing side information and loss noise levels for the stereo filling tool (traditionally representing the noise offset). Since the legacy xHE-AAC decoder ignores the values of the five noise offset bits if the three noise level bits are zero, the presence of signaling for the stereo filling tool only affects noise filling in the legacy decoder, and since the first three bits are zero, noise filling is turned off, and the remaining decoding operation works as intended. In particular, stereo filling is not performed due to the fact that it is operated similarly to the deactivated noise filling process. Therefore, when a frame with stereo filling on is reached, the legacy decoder does not need to mute the output signal, or even interrupt decoding, and thus the legacy decoder still performs a "graceful" decoding of the enhanced bitstream 30. Naturally, it is impossible to reconstruct stereo-filled linear coefficients exactly as intended, resulting in a degradation of quality in affected frames compared to decoding with a suitable decoder that can adequately handle the new stereo-filling tool. Nevertheless, assuming that the stereo-filling tool is used as intended, i.e., only for low-bitrate stereo inputs, the quality from the xHE-AAC decoder should be better than when affected frames are dropped due to muting or result in other obvious playback errors.
[0078] Below, we will explain in detail how, as an extension, a stereo filler tool can be integrated into the xHE-AAC codec.
[0079] If incorporated into the standard, the stereo filler tool can be described as follows. In particular, such a stereo filler (SF) tool would represent a new tool in the frequency domain (FD) portion of MPEG-H 3D audio. Following the above description, the purpose of such a stereo filler tool would be the parametric reconstruction of MDCT spectral coefficients at low bitrates, similar to what can already be achieved by noise filler according to section 7.2 of the standard described in [3]. However, unlike noise filler, which utilizes a pseudo-random noise source to generate MDCT spectral values for any FD channel, SF could also be used to reconstruct the MDCT value of the right channel of a combined-encoded stereo pair of channels by using a downmix of the left and right MDCT spectra of the previous frame. SF is transmitted semi-backward compatible with noise filler side information that can be accurately analyzed by a legacy MPEG-D USAC decoder, according to the embodiments described below.
[0080] The tool may also be described as follows: When SF is activated in a coupled stereo FD frame, the MDCT coefficients of the empty (i.e., completely zero-quantized) scale factor band of the right (second) channel, such as 50d, are replaced with the sum or difference of the corresponding decoded left and right channel MDCT coefficients of the previous frame (in the case of FD). If legacy noise filling is activated for the second channel, pseudo-random values are also added to each coefficient. The resulting coefficients of each scale factor band are then scaled so that the RMS (root mean square of the coefficients) of each band matches the value transmitted by the scale factor of that band. See section 7.3 of the standard in [3].
[0081] Using new SF tools in the MPEG-D USAC standard may impose several operational constraints. For example, an SF tool may only be available for use in the right FD channel of a common FD channel pair, i.e., a channel pair element that transmits StereoCoreToolInfo() using common_window==1. In addition, due to quasi-backward compatible signaling, an SF tool may only be available for use in the syntax container UsacCoreConfig() when noiseFilling==1. If either channel in that pair is in LPD core_mode, the SF tool may not be used, even if the right channel is in FD mode.
[0082] As described in [3], the following terms and definitions are used to more clearly describe the extension of the standard.
[0083] In particular, as far as data elements are concerned, the following data elements will be newly introduced: stereo_filling: A binary flag indicating whether stereofill (SF) is used in the current frame and channel. Furthermore, new auxiliary elements will be introduced. noise_offset Noise-filling offset to correct the scale factor of zero-quantized bandwidth (Section 7.2) noise_level: The noise filling level that represents the amplitude of the added spectral noise (Section 7.2). downmix_prev[] Downmix (i.e., sum or difference) of the left and right channels of the previous frame. sf_index[g][sfb] Scale factor index (i.e., the integer transmitted) for window group g and bandwidth sfb.
[0084] This standard decoding process can be extended as follows. In particular, decoding of a combined stereo-encoded FD channel with the SF tool activated is performed in the following three sequential steps:
[0085] First, the stereo_filling flag may be decrypted. stereo_filling does not represent an independent bitstream element, but is derived from the noise-filling elements, noise_offset and noise_level, within UsacChannelPairElement(), and the common_window flag in StereoCoreToolInfo(). If noiseFilling==0, common_window==0, or the current channel is the left (first) channel in that element, stereo_filling is 0 and the stereo filling process is complete. Otherwise, if ((noiseFilling != 0) && (common_window != 0) && (noise_level == 0)) { stereo_filling = (noise_offset & 16) / 16; noise_level = (noise_offset & 14) / 2; noise_offset = (noise_offset & 1) * 16; } else { stereo_filling = 0; }
[0086] In other words, if noise_level == 0, noise_offset includes the stereo_filling flag and the following 4 bits of noise-filling data, which are then rearranged. This operation must be performed before the noise-filling process in Section 7.2 because it modifies the values of noise_level and noise_offset. Furthermore, the above pseudocode is not performed on the left (first) channel of UsacChannelPairElement() or any other element.
[0087] Next, the calculation of downmix_prev will likely take place. The spectral downmix `downmix_prev[]`, which should be used for stereo filling, is identical to `dmx_re_prev[]`, which is used for MDST spectral estimation in complex stereo prediction (Section 7.7.2.3). This means the following:
[0088] If any of the frames and elements on which downmixing is performed, i.e., any of the channels of the frame preceding the currently decoded frame, use core_mode==1 (LPD), or if the channels have non-uniform transformation lengths (split_transform==1 or a block switch to window_sequence==EIGHT_SHORT_SEQUENCE in a single channel) or use usacIndependencyFlag==1, then all coefficients in downmix_prev[] must be zero.
[0089] If the channel transformation length in the current element has changed from the last frame to the current frame (i.e., split_transform==1 before split_transform==0, or window_sequence==EIGHT_SHORT_SEQUENCE before window_sequence !=EIGHT_SHORT_SEQUENCE, or vice versa), all coefficients in downmix_prev[] must be zero throughout the stereo filling process.
[0090] • If transformation decomposition is applied to the channels of the previous or current frame, downmix_prev[] represents the spectral downmix interleaved per line. See the transformation decomposition tool for details.
[0091] • If complex stereo prediction is not used in the current frame and element, pred_dir is equal to 0.
[0092] As a result, the pre-downmix only needs to be calculated once for both tools, saving computation. The only difference between downmix_prev[] and dmx_re_prev[] in Section 7.7.2 is the behavior when complex stereo prediction is not currently used, or when complex stereo prediction is activated but use_prev_frame==0. In that case, downmix_prev[] is calculated for stereo fill decoding according to Section 7.7.2.3, even if dmx_re_prev[] is not required for complex stereo prediction decoding and is therefore undefined / zero.
[0093] Subsequently, stereo filling of the empty scale factor bandwidth will be performed.
[0094] If stereo_filling==1, after noise filling in all initially empty scale factor bands sfb[] below max_sfb_ste, i.e., all bands where all MDCT lines were quantized to zero, the following procedure is performed: First, the energies of the corresponding lines in this given sfb[] and downmix_prev[] are calculated by the sum of the squares of the lines. Then, for the spectrum of each group window, an sfbWidth is given, which contains the above number of lines per sfb[].
[0095] if (energy[sfb] < sfbWidth[sfb]) { / * noise level isn't maximum, or band starts below noise-fill region * / facDmx = sqrt((sfbWidth[sfb] - energy[sfb]) / energy_dmx[sfb]); factor = 0.0; / * if the previous downmix isn't empty, add the scaled downmix lines such that band reaches unity energy * / for (index = swb_offset[sfb]; index < swb_offset[sfb+1]; index++) { spectrum[window][index] += downmix_prev[window][index] * facDmx; factor += spectrum[window][index] * spectrum[window][index]; } if ((factor != sfbWidth[sfb]) && (factor > 0)) { / * unity energy isn't reached, so modify band * / factor = sqrt(sfbWidth[sfb] / (factor + 1e-8)); for (index = swb_offset[sfb]; index < swb_offset[sfb+1]; index++) { spectrum[window][index] *= factor; } } }
[0096] Next, a scale factor is applied to the resulting spectrum as in Section 7.3, and the scale factor for empty bandwidths is handled in the same way as a normal scale factor.
[0097] An alternative form of the above extension to the xHE-AAC standard would use an implicit, semi-backward compatible signaling method.
[0098] The above embodiment within the framework of the xHE-AAC code describes a method of signaling the use of a new stereo filling tool to the decoder shown in Figure 2 using a single bit in the bitstream contained in stereo_filling. More precisely, such signaling (referred to as explicit quasi-backward compatible signaling) allows subsequent legacy bitstream data—in this case noise-filling side information—to be used independently of SF signaling, and in embodiments of the present invention, the noise-filling data does not depend on the stereo-filling information, and vice versa. For example, noise-filling data consisting entirely of zeros (noise_level=noise_offset=0) may be transmitted, while stereo_filling may signal any possible value (a binary flag of either 0 or 1).
[0099] Strict independence between legacy bitstream data and the bitstream data of the present invention is not required, and if the signal of the present invention is a binary determination, explicit transmission of signal transmission bits can be avoided, and the binary determination can also be transmitted by the presence or absence of a signal which may be called implicit quasi-backward compatible signal transmission. Taking the above embodiment again as an example, the use of stereo filling can be transmitted by simply utilizing a new signal transmission, and if noise_level is zero and noise_offset is non-zero, the stereo_filling flag is set to equal 1. If both noise_level and noise_offset are non-zero, stereo_filling is equal to 0. This implicit signal dependency on legacy noise-filled signals occurs when both noise_level and noise_offset are zero. In this case, it is not clear whether legacy or new SF implicit signal transmission is being used. To avoid such ambiguity, the value of stereo_filling must be predefined. In this example, it is appropriate to define stereo_filling=0 when all noise-filling data consists of zeros, because this is how legacy encoders without stereo-filling capabilities transmit signals when noise-filling should not be applied to the frame.
[0100] An unresolved issue in the case of implicit quasi-backward compatible signaling is how to signal that stereo_filling==1 and at the same time there is no noise filling. As mentioned above, the noise-filling data cannot be "all zeros", and if a noise magnitude of zero is required, noise_level (as mentioned above, (noise_offset&14) / 2) must be equal to 0. This leaves only noise_offset (as mentioned above, (noise_offset&1)*16) greater than 0 as a solution. However, even if noise_level is zero, noise_offset is taken into account when applying the scale factor in the case of stereo filling. Conveniently, when writing the bitstream, the encoder can compensate for the fact that a zero noise_offset may not be transmitted by modifying the affected scale factor so that the affected scale factor includes an offset that is not performed in the decoder via noise_offset. This makes the implicit signaling in the above embodiment possible at the cost of the potential increase in the data rate of the scale factor. Therefore, the stereo filling signal transmission in the pseudocode described above can be modified as follows by using the saved SF signal transmission bits to transmit noise_offset with 2 bits (4 values) instead of 1 bit:
[0101] if ((noiseFilling) && (common_window) && (noise_level == 0) && (noise_offset > 0)) { stereo_filling = 1; noise_level = (noise_offset & 28) / 4; noise_offset = (noise_offset & 3) * 8; } else { stereo_filling = 0; }
[0102] For the sake of completeness, Figure 6 shows a parametric audio encoder according to one embodiment of the present invention. First, the encoder of Figure 6, generally indicated by reference numeral 90, comprises a conversion unit 92 for performing a distortion-free conversion of the audio signal reconstructed in output 32 of Figure 2 to its original version. Wrapped conversion may be used, with a plurality of different conversion lengths having corresponding conversion windows, switched in units of frames 44, as described in relation to Figure 3. The different conversion lengths and corresponding conversion windows are indicated in Figure 3 by reference numeral 104. Similar to Figure 2, Figure 6 focuses on a portion of the encoder 90 that is responsible for encoding one channel of a multi-channel audio signal, while another channel region portion of the encoder 90 is generally indicated in Figure 6 by reference numeral 96.
[0103] At the output of the transformer 92, the spectral lines and scale factor are not quantized, and virtually no coding loss has occurred yet. The spectrogram output by the transformer 92 enters the quantization unit 98, which is configured to quantize the spectral lines of the spectrogram output by the transformer 92, spectrum by spectrum, by setting and using a preliminary scale factor of the scale factor band. That is, at the output of the quantization unit 98, the preliminary scale factor and the corresponding spectral line coefficients are obtained, and the sequence of noise filling unit 16', an optional inverse TNS filter 28a', inter-channel prediction unit 24', MS decoder 26', and inverse TNS filter 28b' is connected sequentially, thereby giving the encoder 90 in Figure 6 the ability to obtain a reconstructed final version of the current spectrum, which can be obtained at the input of the decoder-side downmix provider (see Figure 2). When using the inter-channel prediction unit 24' and / or the inter-channel noise filling in the version that uses a downmix of the previous frame to form inter-channel noise, the encoder 90 also includes a downmix provider 31' that forms a downmix of the final reconfigured version of the spectrum of the channels of the multi-channel audio signal. Of course, to save computational resources, instead of the final version, the unquantized original version of the spectrum of the channels may be used by the downmix provider 31' in forming the downmix.
[0104] The encoder 90 may use information about the available reconstructed final version of the spectrum to perform inter-frame spectral prediction, such as the aforementioned possible version which performs inter-channel prediction using imaginary part estimation, and / or perform rate control, i.e., within a rate control loop, the encoder 90 may determine that the possible parameters ultimately encoded into the data stream 30 are optimally set in terms of rate / distortion.
[0105] For example, one parameter set within such a prediction loop and / or rate control loop of encoder 90 is the scale factor of each scale factor band, which is simply pre-set by the quantization unit 98 for each zero-quantized scale factor band identified by the identification unit 12'. Within the prediction and / or rate control loop of encoder 90, the scale factor of the zero-quantized scale factor band is set to be psychoacoustically or rate / distortion optimized, thereby determining the aforementioned optional correction parameter, along with the aforementioned target noise level, which is carried to the decoder side by the data stream for the corresponding frame. Note that this scale factor may be calculated using only the spectral lines of the spectrum and the channel to which the spectrum belongs (i.e., the aforementioned "target" spectrum), or alternatively, it may be determined using both the spectral lines of the "target" channel spectrum and, additionally, the spectral lines of other channel spectra, or the downmix spectrum from the previous frame obtained from the downmix provider 31' (i.e., the aforementioned "source" spectrum). In particular, to stabilize the target noise level and reduce temporal level fluctuations in the decoded audio channel to which inter-channel noise filling is applied, the target scale factor may be calculated using the relationship between the energy scale of a spectral line in the “target” scale factor band and the energy scale of a spectral line at the same location in the corresponding “source” region. Finally, as described above, this “source” region may originate from the reconfigured final version of another channel or a downmix of the previous frame, or, if the encoder computation should be reduced, from the unquantized original version of the other channel or a downmix of the unquantized original version of the spectrum of the previous frame.
[0106] The following describes multi-channel coding and multi-channel decoding according to embodiments. In embodiments, the multi-channel processing unit 204 of the decoding apparatus 201 in Figure 1a may be configured to perform one or more of the following techniques, for example, with respect to noise multi-channel decoding.
[0107] However, before explaining multichannel decoding, we will first describe multichannel coding according to the embodiment with reference to Figures 7 to 9, and then explain multichannel decoding with reference to Figures 10 and 12.
[0108] Here, with reference to Figures 7 to 9 and Figure 11, a multi-channel coding embodiment will be described.
[0109] Figure 7 shows a schematic block diagram of a device (encoder) 100 that encodes a multichannel signal 101 having at least three channels CH1 to CH3.
[0110] The device 100 comprises an iteration processing unit 102, a channel encoder 104, and an output interface 106.
[0111] The iterative processing unit 102 is configured to calculate inter-channel correlation values between at least three pairs of channels CH1 to CH3 in the first iteration step in order to select the pair having the highest value or a pair having a value above a threshold, and to process the selected pair using a multi-channel processing operation to derive a multi-channel parameter MCH_PAR1 for the selected pair, and to derive first processed channels P1 and P2. Hereinafter, such processed channels P1 and P2 will also be referred to as composite channel P1 and composite channel P2, respectively. Furthermore, the iterative processing unit 102 is configured to use at least one of the processed channels P1 or P2 to perform calculation, selection and processing in the second iteration step to derive a multi-channel parameter MCH_PAR2 and second processed channels P3 and P4.
[0112] For example, as shown in Figure 7, the iterative processing unit 102 may, in the first iteration step, calculate the inter-channel correlation value between at least three first pairs of channels CH1 to CH3, where the first pair consists of the first channel CH1 and the second channel CH2; the inter-channel correlation value between at least three second pairs of channels CH1 to CH3, where the second pair consists of the second channel CH2 and the third channel CH3; and the inter-channel correlation value between at least three third pairs of channels CH1 to CH3, where the third pair consists of the first channel CH1 and the third channel CH3.
[0113] In Figure 7, it is assumed that in the first iteration step, a third pair consisting of the first channel CH1 and the third channel CH3 contains the highest inter-channel correlation value, and that the iteration processing unit 102 selects the third pair having the highest inter-channel correlation value in the first iteration step, derives the multi-channel parameter MCH_PAR1 for the selected pair using a multi-channel processing operation, and processes the selected pair, i.e., the third pair, to derive the first processed channels P1 and P2.
[0114] Furthermore, the iteration processing unit 102 can be configured to calculate inter-channel correlation values between at least three channels CH1 to CH3 and each pair of processed channels P1 and P2 in the second iteration step in order to select the pair having the highest value or the pair having a value above a threshold. This allows the iteration processing unit 102 to be configured not to select the pair selected in the first iteration step in the second iteration step (or any further iteration step).
[0115] Referring to the example shown in Figure 7, the iterative processing unit 102 may further calculate the inter-channel correlation value between a fourth channel pair consisting of a first channel CH1 and a first processed channel P1, the inter-channel correlation value between a fifth pair consisting of a first channel CH1 and a second processed channel P2, the inter-channel correlation value between a sixth pair consisting of a second channel CH2 and a first processed channel P1, the inter-channel correlation value between a seventh pair consisting of a second channel CH2 and a second processed channel P2, the inter-channel correlation value between an eighth pair consisting of a third channel CH3 and a first processed channel P1, the inter-channel correlation value between a ninth pair consisting of a third channel CH3 and a second processed channel P2, and the inter-channel correlation value between a tenth pair consisting of a first processed channel P1 and a second processed channel P2.
[0116] In Figure 7, we assume that in the second iteration step, the sixth pair consisting of the second channel CH2 and the first processed channel P1 contains the highest inter-channel correlation value, and that the iteration processing unit 102 selects the sixth pair in the second iteration step, uses a multi-channel processing operation to derive the multi-channel parameter MCH_PAR2 for the selected pair, and processes the selected pair, i.e., the sixth pair, to derive the second processed channels P3 and P4.
[0117] The iterative processing unit 102 can be configured to select a pair only when the level difference between the pairs is less than a threshold, and the threshold is less than 40 dB, 25 dB, 12 dB, or less than 6 dB. Thus, a threshold of 25 or 40 dB corresponds to a rotation angle of 3 or 0.5 degrees.
[0118] The iterative processing unit 102 can be configured to calculate a normalized integer correlation value, and the iterative processing unit 102 can be configured to select a pair when the integer correlation value is greater than, for example, 0.2, preferably 0.3.
[0119] Furthermore, the iterative processing unit 102 may provide the channels obtained as a result of the multi-channel processing to the channel encoder 104. For example, referring to Figure 7, the iterative processing unit 102 may provide the channel encoder 104 with the third processed channel P3 and the fourth processed channel P4, which are the results of the multi-channel processing performed in the second iteration step, as well as the second processed channel P2, which is the result of the multi-channel processing performed in the first iteration step. In this way, the iterative processing unit 102 can provide the channel encoder 104 with only these processed channels that are not processed (further) in subsequent iteration steps. As shown in Figure 7, the first processed channel P1 is not provided to the channel encoder 104 because it is processed further in the second iteration step.
[0120] The channel encoder 104 can be configured to encode channels P2 to P4, which are the results of iterative processing (or multi-channel processing) performed by the iteration processing unit 102, to obtain encoded channels E1 to E3.
[0121] For example, the channel encoder 104 can be configured to use mono encoders (or monoboxes or monotools) 120_1 to 120_3 to encode channels P2 to P4, which are the result of iterative processing (or multi-channel processing). The monoboxes may be configured to encode channels such that fewer bits are required to encode channels with less energy (or smaller amplitude) than to encode channels with more energy (or higher amplitude). The monoboxes 120_1 to 120_3 may be, for example, conversion-based audio encoders. Furthermore, the channel encoder 104 can be configured to use stereo encoders (e.g., parametric stereo encoders or lossy stereo encoders) to encode channels P2 to P4, which are the result of iterative processing (or multi-channel processing).
[0122] The output interface 106 can be configured to generate an encoded multichannel signal 107 having encoded channels E1 to E3 and multichannel parameters MCH_PAR1 and MCH_PAR2.
[0123] For example, the output interface 106 can generate the encoded multichannel signal 107 as a serial signal or serial bitstream, and can be configured such that the multichannel parameter MCH_PAR2 precedes the multichannel parameter MCH_PAR1 in the encoded signal 107. Thus, the decoder of the embodiment described later with respect to Figure 10 receives the multichannel parameter MCH_PAR2 before the multichannel parameter MCH_PAR1.
[0124] In Figure 7, the iteration processing unit 102 exemplary performs two multi-channel processing operations, namely a multi-channel processing operation in the first iteration step and a multi-channel processing operation in the second iteration step. Of course, the iteration processing unit 102 can also perform further multi-channel processing operations in subsequent iteration steps. This allows the iteration processing unit 102 to be configured to execute iteration steps until an iteration termination criterion is reached. The iteration termination criterion may be that the maximum number of iteration steps is equal to or two or more greater than the total number of channels in the multi-channel signal 101, or the iteration termination criterion may be that the inter-channel correlation value is not greater than a threshold, where the threshold is preferably greater than 0.2, or preferably 0.3. In a further embodiment, the iteration termination criterion may be that the maximum number of iteration steps is greater than or equal to the total number of channels in the multi-channel signal 101, or the iteration termination criterion may be that the inter-channel correlation value is not greater than a threshold, where the threshold is preferably greater than 0.2, or preferably 0.3.
[0125] For illustrative purposes, the multi-channel processing operations performed by the iteration processing unit 102 in the first and second iteration steps are illustrated in Figure 7 by processing boxes 110 and 112. Processing boxes 110 and 112 can be implemented in hardware or software. Processing boxes 110 and 112 can be, for example, stereo boxes.
[0126] This allows us to leverage inter-channel signal dependencies by hierarchically applying known combined stereo coding tools. In contrast to previous MPEG methods, the signal pairs to be processed are not predetermined by a fixed signal path (e.g., a stereo coding tree), but can be dynamically modified to adapt to the input signal characteristics. The input to the actual stereo box could be (1) raw channels such as channels CH1-CH3, (2) outputs of preceding stereo boxes such as processed signals P1-P4, or (3) a composite channel of raw channels and outputs of preceding stereo boxes.
[0127] The processing within stereo boxes 110 and 112 can be either prediction-based (like the complex prediction box in USAC) or KLT / PCA-based (the input channel is rotated in the encoder (e.g., via a 2x2 rotation matrix) to maximize energy compression, i.e., concentrate the signal energy into one channel, and in the decoder, the rotated signal is re-converted in the direction of the original input signal).
[0128] In a possible embodiment of encoder 100, (1) the encoder calculates the inter-channel correlation between each channel pair, selects one suitable signal pair from the input signal, and applies the stereo tool to the selected channel; (2) the encoder recalculates the inter-channel correlation between all channels (unprocessed channels and processed intermediate output channels), selects one suitable signal pair from the input signal, and applies the stereo tool to the selected channel; (3) the encoder repeats step (2) until all inter-channel correlations fall below a threshold or the maximum number of transformations have been applied.
[0129] As already mentioned, the signal pairs processed by the encoder 100, or more precisely, the iteration processing unit 102, can be dynamically modified to adapt to the input signal characteristics, rather than being predetermined by a fixed signal path (e.g., a stereo coding tree). Thus, the encoder 100 (or iteration processing unit 102) can be configured to construct a stereo tree depending on at least three channels CH1 to CH3 of the multi-channel (input) signal 101. In other words, the encoder 100 (or iteration processing unit 102) can be configured to construct a stereo tree based on inter-channel correlation (for example, by calculating the inter-channel correlation value between each pair of at least three channels CH1 to CH3 in the first iteration step to select the pair having the highest value or a value above a threshold, and further by calculating the inter-channel correlation value between each pair of at least three channels and the previously processed channel in the second iteration step to select the pair having the highest value or a value above a threshold). According to the one-step method, a correlation matrix may be calculated for each iteration, including the correlation of all channels in any previous iterations that may have been processed.
[0130] As described above, the iteration processing unit 102 can be configured to derive a multichannel parameter MCH_PAR1 for the selected pair in the first iteration step and a multichannel parameter MCH_PAR2 for the selected pair in the second iteration step. The multichannel parameter MCH_PAR1 may include a first channel pair identifier (or index) that identifies (or signals) the channel pair selected in the first iteration step, and the multichannel parameter MCH_PAR2 may include a second channel pair identifier (or index) that identifies (or signals) the channel pair selected in the second iteration step.
[0131] The following describes efficient indexing of input signals. For example, channel pairs can be efficiently transmitted using a unique index for each pair, depending on the total number of channels. For example, the indexing of six channel pairs may be as shown in the following table.
[0132] JPEG0007885285000001.jpg6876
[0133] For example, in the table above, index 5 can transmit a signal through a pair consisting of the first channel and the second channel. Similarly, index 6 can transmit a signal through a pair consisting of the first channel and the third channel.
[0134] The total number of possible channel pair indices for n channels can be calculated as follows: numPairs=numChannels*(numChannels-1) / 2 Therefore, the number of bits required to transmit a signal through one channel pair is: numBits=floor(log2(numPairs-1))+1
[0135] Furthermore, encoder 100 may use a channel mask. The configuration of a multi-channel tool may include a channel mask that indicates the active channel of the tool. Thus, LFE (Low-Frequency Effect / Enhancement Channel) can be removed from the channel pair index, enabling more efficient encoding. For example, in the 11.1 setup, this reduces the number of channel pair indices from 12 × 11 / 2 = 66 to 11 × 10 / 2 = 55, allowing signal transmission with 6 bits instead of 7 bits. This mechanism can also be used to exclude channels intended for mono-objects (e.g., multiple language tracks). In decoding the channel mask, a channel map can be generated to allow remapping of channel pair indices to decoder channels.
[0136] Furthermore, the iteration processing unit 102 can be configured to derive a plurality of selected pair displays for the first frame, and the output interface 106 can be configured to include a hold indicator in the multi-channel signal 107 that indicates that for a second frame following the first frame, the second frame has the same plurality of selected pair displays as the first frame.
[0137] A hold indicator or hold tree flag can be used to signal that no new tree will be transmitted, but the last stereo tree should be used. This can be used to avoid multiple transmissions of the same stereo tree configuration if the channel correlation characteristics remain static for a longer period of time.
[0138] Figure 8 shows schematic block diagrams of stereo boxes 110 and 112. Stereo boxes 110 and 112 have inputs for a first input signal I1 and a second input signal I2, and outputs for a first output signal O1 and a second output signal O2. As shown in Figure 8, the dependence of output signals O1 and O2 on input signals I1 and I2 can be described by s-parameters S1 to S4.
[0139] The iteration processing unit 102 may use (or include) stereo boxes 110 and 112 to perform multi-channel processing operations on the input channel and / or the processed channel in order to derive (further) processed channels. For example, the iteration processing unit 102 may be configured to use general prediction-based or KLT (Karhunen-Loeve transform)-based rotation stereo boxes 110 and 112.
[0140] A general-purpose encoder (or encoder-side stereo box) can be configured to encode input signals I1 and I2 in order to obtain output signals O1 and O2 based on the following equation. TIFF0007885285000002.tif1039
[0141] A general-purpose decoder (or decoder-side stereo box) can be configured to decode input signals I1 and I2 in order to obtain output signals O1 and O2 based on the following equation. TIFF0007885285000003.tif1044
[0142] A prediction-based encoder (or encoder-side stereo box) can be configured to encode input signals I1 and I2 in order to obtain output signals O1 and O2 based on the following equations. TIFF0007885285000004.tif1068 Here, p is the prediction coefficient.
[0143] A prediction-based decoder (or decoder-side stereo box) can be configured to decode input signals I1 and I2 to obtain output signals O1 and O2 based on the following equations. TIFF0007885285000005.tif1047
[0144] A KLT-based rotary encoder (or encoder-side stereo box) can be configured to encode input signals I1 and I2 in order to obtain output signals O1 and O2 based on the following equations. TIFF0007885285000006.tif1054
[0145] A KLT-based rotary decoder (or decoder-side stereo box) can be configured to decode input signals I1 and I2 to obtain output signals O1 and O2 based on the following equation (reverse rotation): TIFF0007885285000007.tif1054
[0146] The following section explains how to calculate the rotation angle α for rotation based on KLT. The rotation angle α of the KLT base rotation can be defined as follows: TIFF0007885285000008.tif1142C xy is an entry in the unnormalized correlation matrix, where C 11 and C 22 This is the channel energy.
[0147] This can be done using the atan2 function to enable differentiation between the negative correlation in the numerator and the negative energy difference in the denominator. α=0.5*atan2(2*correlation[ch1][ch2], (correlation[ch1][ch1]-correlation[ch2][ch2]))
[0148] Furthermore, the iterative processing unit 102 can be configured to calculate inter-channel correlation using frames from each channel containing multiple bandwidths, thereby obtaining a single inter-channel correlation value for multiple bandwidths. The iterative processing unit 102 can also be configured to perform multi-channel processing for each of the multiple bandwidths, thereby obtaining multi-channel parameters from each of the multiple bandwidths.
[0149] This allows the iterative processing unit 102 to be configured to calculate stereo parameters in multichannel processing, and to be configured to perform only stereo processing in the bandwidth, with the stereo parameters being higher than the zero quantization threshold defined by the stereo quantizer (e.g., a KLT-based rotation encoder). The stereo parameters could be, for example, MS on / off, rotation angle, or prediction coefficient.
[0150] For example, the iterative processing unit 102 can be configured to calculate the rotation angle in multichannel processing, or iterative processing unit 102 can be configured to perform only rotation processing in the bandwidth, and the rotation angle is higher than the zero quantization threshold defined by the rotation angle quantizer (e.g., a KLT-based rotation encoder).
[0151] Therefore, the encoder 100 (or output interface 106) can be configured to transmit the conversion / rotation information as one parameter for any complete spectrum (full bandbox) or as multiple frequency-dependent parameters for any part of the spectrum.
[0152] The encoder 100 can be configured to generate a bitstream 107 based on the following table.
[0153] [Table 1]
[0154] [Table 2]
[0155] [Table 3]
[0156] [Table 4]
[0157] [Table 5]
[0158] [Table 6]
[0159] [Table 7]
[0160] Figure 9 shows a schematic block diagram of the repetition processing unit 102 according to one embodiment. In the embodiment shown in Figure 9, the multichannel signal 101 is a 5.1 channel signal having six channels: left channel L, right channel R, left surround channel Ls, right surround channel Rs, center channel C, and low-frequency sound effect channel LFE.
[0161] As shown in Figure 9, the LFE channel is not processed by the iteration processing unit 102. This may be because the inter-channel correlation values between the LFE channel and each of the other five channels L, R, Ls, Rs, and C are small, or because the channel mask assumed below does not process the LFE channel.
[0162] In the first iteration step, the iteration processing unit 102 calculates the inter-channel correlation values between each pair of five channels L, R, Ls, Rs, and C in order to select pairs that have the maximum value or a value above a threshold in the first iteration step. Assuming that the left channel L and the right channel R have the maximum value in Figure 9, the iteration processing unit 102 processes the left channel L and the right channel R using a stereo box (or stereo tool) 110 that performs multi-channel operation to derive the first and second processed channels P1 and P2.
[0163] In the second iteration step, the iteration processing unit 102 calculates the inter-channel correlation values between the five channels L, R, Ls, Rs, C and each pair of processed channels P1 and P2 in order to select pairs that have the maximum value or a value above a threshold in the second iteration step. Assuming that the left surround channel Ls and the right surround channel Rs have the maximum value in Figure 9, the iteration processing unit 102 processes the left surround channel Ls and the right surround channel Rs using the stereo box (or stereo tool) 112 to derive the third and fourth processed channels P3 and P4.
[0164] In the third iteration step, the iteration processing unit 102 calculates the inter-channel correlation values between the five channels L, R, Ls, Rs, C and each pair of processed channels P1 to P4 in order to select pairs that have the maximum value or a value above a threshold in the third iteration step. In Figure 9, assuming that the first processed channel P1 and the third processed channel P3 have the maximum value, the iteration processing unit 102 processes the first processed channel P1 and the third processed channel P3 using the stereo box (or stereo tool) 114 to derive the fifth and sixth processed channels P5 and P6.
[0165] In the fourth iteration step, the iteration processing unit 102 calculates the inter-channel correlation values between the five channels L, R, Ls, Rs, C and each pair of processed channels P1 to P6 in order to select pairs that have the maximum value or a value above a threshold in the fourth iteration step. In Figure 9, assuming that the fifth processed channel P5 and the central channel C have the maximum value, the iteration processing unit 102 processes the fifth processed channel P5 and the central channel C using the stereo box (or stereo tool) 115 to derive the seventh and eighth processed channels P7 and P8.
[0166] Stereo boxes 110-116 may be MS stereo boxes, i.e., mid / side stereo sound boxes configured to provide a mid channel and a side channel. The mid channel may be the sum of the input channels of the stereo box, and the side channel may be the difference between the input channels of the stereo box. Furthermore, stereo boxes 110 and 116 may be rotating boxes or stereo prediction boxes.
[0167] In Figure 9, the first processed channel P1, the third processed channel P3, and the fifth processed channel P5 may be mid-channels, and the second processed channel P2, the fourth processed channel P4, and the sixth processed channel P6 may be side channels.
[0168] Furthermore, as shown in Figure 9, the iterative processing unit 102 can be configured, if applicable in the second iteration step, to calculate, select, and process using only the input channels L, R, Ls, Rs, C and the mid-channels P1, P3, and P5 of the processed channel in subsequent iteration steps. In other words, the iterative processing unit 102 can be configured, if applicable in the second iteration step, to not use the side channels P1, P3, and P5 of the processed channel when calculating, selecting, and processing in subsequent iteration steps.
[0169] Figure 11 shows a flowchart of a method 300 for encoding a multichannel signal having at least three channels. The method 300 includes a first iteration step of selecting a pair having the highest value or a pair having a value above a threshold, processing the selected pair using a multichannel processing operation to derive a multichannel parameter MCH_PAR1 for the selected pair, and a first processed channel, in which step 302 calculates an interchannel correlation value between each pair of at least three channels in order to derive a first processed channel; a second iteration step of using at least one of the processed channels to perform calculation, selection and processing to derive a multichannel parameter MCH_PAR2 and a second processed channel, in which step 304 calculates, selects and processes in a second iteration step to obtain an encoded channel, in which step 306 encodes the channels resulting from the iterations performed by the iteration processing unit to obtain an encoded channel, and a third step of generating an encoded multichannel signal having the encoded channel and the first and multichannel parameters MCH_PAR2.
[0170] The following section explains multi-channel decoding. Figure 10 shows a schematic block diagram of a decoder 200 that decodes an encoded multichannel signal 107 having encoded channels E1 to E3 and at least two multichannel parameters MCH_PAR1 and MCH_PAR2.
[0171] The device 200 includes a channel decoder 202 and a multichannel processing unit 204. The channel decoder 202 is configured to decode the encoded channels E1 to E3 to obtain the decoded channels D1 to D3.
[0172] For example, the channel decoder 202 may comprise at least three mono decoders (or monoboxes or monotools) 206_1 to 206_3, each of which can be configured to decode one of at least three encoded channels E1 to E3 to obtain its respective decoded channel E1 to E3. The mono decoders 206_1 to 206_3 may, for example, be conversion-based audio decoders.
[0173] The multichannel processing unit 204 is configured to perform multichannel processing using a second pair of decoded channels identified by the multichannel parameter MCH_PAR2 and using the multichannel parameter MCH_PAR2 to obtain processed channels, and to perform further multichannel processing using a first pair of channels identified by the multichannel parameter MCH_PAR1 and using the multichannel parameter MCH_PAR1, such that the first pair of channels includes at least one processed channel.
[0174] As shown in Figure 10 as an example, the multichannel parameter MCH_PAR2 can indicate (or signal) that the second decoded channel pair consists of the first decoded channel D1 and the second decoded channel D2. Therefore, the multichannel processing unit 204 uses the second decoded channel pair, consisting of the first decoded channel D1 and the second decoded channel D2 (identified by the multichannel parameter MCH_PAR2), and the multichannel parameter MCH_PAR2 to perform multichannel processing and obtain processed channels P1* and P2*. The multichannel parameter MCH_PAR1 can indicate that the first decoded channel pair consists of the first processed channel P1* and the third decoded channel D3. Therefore, the multi-channel processing unit 204 uses the first decoded channel pair, consisting of the first processed channel P1* and the third decoded channel D3 (identified by the multi-channel parameter MCH_PAR1), and uses the multi-channel parameter MCH_PAR1 to perform further multi-channel processing to obtain processed channels P3* and P4*.
[0175] Furthermore, the multi-channel processing unit 204 can provide a third processed channel P3* as the first channel CH1, a fourth processed channel P4* as the third channel CH3, and a second processed channel P2* as the second channel CH2.
[0176] Assuming that the decoder 200 shown in Figure 10 receives a multi-channel signal 107 encoded from the encoder 100 shown in Figure 7, the first decoded channel D1 of the decoder 200 may be equivalent to the third processed channel P3 of the encoder 100, the second decoded channel D2 of the decoder 200 may be equivalent to the fourth processed channel P4 of the encoder 100, and the third decoded channel D3 of the decoder 200 may be equivalent to the second processed channel P2 of the encoder 100. Furthermore, the first processed channel P1* of the decoder 200 may be equivalent to the first processed channel P1 of the encoder 100.
[0177] Furthermore, the encoded multichannel signal 107 may also be a serial signal, and the multichannel parameter MCH_PAR2 is received by the decoder 200 before the multichannel parameter MCH_PAR1. In that case, the multichannel processing unit 204 can be configured to process the decoded channels in the order in which the multichannel parameters MCH_PAR1 and MCH_PAR2 are received by the decoder. In the example shown in Figure 10, the decoder receives the multichannel parameter MCH_PAR2 before the multichannel parameter MCH_PAR1, thereby performing multichannel processing using the second decoded channel pair identified by the multichannel parameter MCH_PAR2 (consisting of the first and second decoded channels D1 and D2) before performing multichannel processing using the first decoded channel pair identified by the multichannel parameter MCH_PAR1 (consisting of the first processed channel P1* and the third decoded channel D3).
[0178] In Figure 10, the multi-channel processing unit 204 performs two multi-channel processing operations as an example. For illustrative purposes, the multi-channel processing operations performed by the multi-channel processing unit 204 are shown in Figure 10 by processing boxes 208 and 210. Processing boxes 208 and 210 can be implemented in hardware or software. Processing boxes 208 and 210 may be stereo boxes such as a general-purpose decoder (or decoder-side stereo box), a prediction-based decoder (or decoder-side stereo box), or a KLT-based rotation decoder (or decoder-side stereo box), as described above with reference to encoder 100.
[0179] For example, encoder 100 can use a KLT-based rotary encoder (or a stereo box on the encoder side). In that case, encoder 100 can derive the multichannel parameters MCH_PAR1 and MCH_PAR2 such that the multichannel parameters MCH_PAR1 and MCH_PAR2 include the rotation angle. The rotation angle can be differentially encoded. Therefore, the multichannel processing unit 204 of decoder 200 can be equipped with a differential decoder for differentially decoding the differentially encoded rotation angle.
[0180] The device 200 may further include an input interface 212 configured to receive and process an encoded multichannel signal 107, provide encoded channels E1 to E3 to a channel decoder 202, and provide multichannel parameters MCH_PAR1 and MCH_PAR2 to a multichannel processing unit 204.
[0181] As already mentioned, a hold indicator (or hold tree flag) may be used to signal that no new tree will be transmitted, but the last stereo tree should be used. This can be used to avoid multiple transmissions of the same stereo tree configuration if the channel correlation characteristics remain static for a longer period of time.
[0182] Therefore, if the encoded multichannel signal 107 includes multichannel parameters MCH_PAR1 and MCH_PAR2 for the first frame and includes a hold indicator for the second frame following the first frame, the multichannel processing unit 204 can be configured to perform multichannel processing or further multichannel processing in the second frame on the same second channel pair or the same first channel pair used in the first frame.
[0183] Multichannel processing and further multichannel processing may include stereo processing using stereo parameters, where for each scale factor band or group of scale factor bands of the decoded channels D1-D3, a first stereo parameter is included in the multichannel parameter MCH_PAR1 and a second stereo parameter is included in the multichannel parameter MCH_PAR2. Thus, the first and second stereo parameters may be of the same type, such as rotation angle or prediction coefficient. Of course, the first and second stereo parameters may be of different types. For example, the first stereo parameter may be a rotation angle and the second stereo parameter may be a prediction coefficient, and vice versa.
[0184] Furthermore, the multi-channel parameters MCH_PAR1 and MCH_PAR2 may include a multi-channel processing mask that indicates which scale factor bandwidths are processed by multi-channel processing and which are not. This allows the multi-channel processing unit 204 to be configured not to perform multi-channel processing in the scale factor bandwidths indicated by the multi-channel processing mask.
[0185] The multichannel parameters MCH_PAR1 and MCH_PAR2 may each include a channel pair identifier (or index), and the multichannel processing unit 204 may be configured to decode the channel pair identifier (or index) using a predetermined decoding rule or a decoding rule indicated in the encoded multichannel signal.
[0186] For example, channel pairs can efficiently transmit signals using a unique index for each pair, depending on the total number of channels, as described above with reference to encoder 100.
[0187] Furthermore, the decoding rule can be a Huffman decoding rule that can be configured so that the multi-channel processing unit 204 performs Huffman decoding for channel pair identification.
[0188] The encoded multichannel signal 107 may further include a multichannel processing enable indicator that indicates only the subgroup of decoded channels for which multichannel processing is permitted, and at least one decoded channel for which multichannel processing is not permitted. This allows the multichannel processing unit 204 to be configured not to perform any multichannel processing on at least one decoded channel for which multichannel processing is not permitted, as indicated by the multichannel processing enable indicator.
[0189] For example, if the multichannel signal is a 5.1 channel signal, the multichannel processing enable indicator may indicate that multichannel processing is permitted only for five channels, namely right R, left L, right surround Rs, left surround LS, and center C, and that multichannel processing is not permitted for the LFE channel.
[0190] The following C code can be used for the decoding process (decoding of channel pair indices). This requires, for all channel pairs, the number of channels using active KLT processing (n channels) and the number of channel pairs in the current frame (numPairs).
[0191] maxNumPairIdx = nChannels*(nChannels-1) / 2 - 1; numBits = floor(log2(maxNumPairIdx)+1; pairCounter = 0; for (chan1=1; chan1 < nChannels; chan1++) { for (chan0=0; chan0 < chan1; chan0++) { if (pairCounter == pairIdx) { channelPair[0] = chan0; channelPair[1] = chan1; return; } else pairCounter++; } } }
[0192] The following C code can be used to decode the prediction coefficients for non-band angle.
[0193] for(pair=0; pair <numPairs; pair++) { mctBandsPerWindow = numMaskBands[pair] / windowsPerFrame; if(delta_code_time[pair] > 0) { lastVal = alpha_prev_fullband[pair]; else { lastVal = DEFAULT_ALPHA; } newAlpha = lastVal + dpcm_alpha[pair][0]; if(newAlpha >= 64) { newAlpha -= 64; } for (band=0; band < numMaskBands; band++){ / * set all angles to fullband angle * / pairAlpha[pair][band] = newAlpha; / * set previous angles according to mctMask * / if(mctMask[pair][band] > 0) { alpha_prev_frame[pair][band%mctBandsPerWindow] = newAlpha; } else { alpha_prev_frame[pair][band%mctBandsPerWindow] = DEFAULT_ALPHA; } } alpha_prev_fullband[pair] = newAlpha; for(band=bandsPerWindow; band <MAX_NUM_MC_BANDS; band++) { alpha_prev_frame[pair][band] = DEFAULT_ALPHA; } }
[0194] The following C code can be used to decode the prediction coefficients for non-band KLT angles.
[0195] for(pair=0; pair<numPairs; pair++) { mctBandsPerWindow = numMaskBands[pair] / windowsPerFrame; for(band=0; band<numMaskBands[pair]; band++) { if(delta_code_time[pair] > 0) { lastVal = alpha_prev_frame[pair][band%mctBandsPerWindow]; } else { if ((band % mctBandsPerWindow) == 0) { lastVal = DEFAULT_ALPHA; } } if (msMask[pair][band] > 0 ) { newAlpha = lastVal + dpcm_alpha[pair][band]; if(newAlpha >= 64) { newAlpha -= 64; } pairAlpha[pair][band] = newAlpha; alpha_prev_frame[pair][band%mctBandsPerWindow] = newAlpha; lastVal = newAlpha; } else { alpha_prev_frame[pair][band%mctBandsPerWindow] = DEFAULT_ALPHA; / * -45° * / } / * reset fullband angle * / alpha_prev_fullband[pair] = DEFAULT_ALPHA; } for(band=bandsPerWindow; band <MAX_NUM_MC_BANDS; band++) { alpha_prev_frame[pair][band] = DEFAULT_ALPHA; } }
[0196] To avoid floating-point differences in trigonometric functions across different platforms, use the following lookup table to directly convert angle indices to sin / cos.
[0197] tabIndexToSinAlpha
[64] = { -1.000000f,-0.998795f,-0.995185f,-0.989177f,-0.980785f,-0.970031f,-0.956940f,-0.941544f, -0.923880f,-0.903989f,-0.881921f,-0.857729f,-0.831470f,-0.803208f,-0.773010f,-0.740951f, -0.707107f,-0.671559f,-0.634393f,-0.595699f,-0.555570f,-0.514103f,-0.471397f,-0.427555f, -0.382683f,-0.336890f,-0.290285f,-0.242980f,-0.195090f,-0.146730f,-0.098017f,-0.049068f, 0.000000f, 0.049068f, 0.098017f, 0.146730f, 0.195090f, 0.242980f, 0.290285f, 0.336890f, 0.382683f, 0.427555f, 0.471397f, 0.514103f, 0.555570f, 0.595699f, 0.634393f, 0.671559f, 0.707107f, 0.740951f, 0.773010f, 0.803208f, 0.831470f, 0.857729f, 0.881921f, 0.903989f, 0.923880f, 0.941544f, 0.956940f, 0.970031f, 0.980785f, 0.989177f, 0.995185f, 0.998795f }; tabIndexToCosAlpha
[64] = { 0.000000f, 0.049068f, 0.098017f, 0.146730f, 0.195090f, 0.242980f, 0.290285f, 0.336890f, 0.382683f, 0.427555f, 0.471397f, 0.514103f, 0.555570f, 0.595699f, 0.634393f, 0.671559f, 0.707107f, 0.740951f, 0.773010f, 0.803208f, 0.831470f, 0.857729f, 0.881921f, 0.903989f, 0.923880f, 0.941544f, 0.956940f, 0.970031f, 0.980785f, 0.989177f, 0.995185f, 0.998795f, 1.000000f, 0.998795f, 0.995185f, 0.989177f, 0.980785f, 0.970031f, 0.956940f, 0.941544f, 0.923880f, 0.903989f, 0.881921f, 0.857729f, 0.831470f, 0.803208f, 0.773010f, 0.740951f, 0.707107f, 0.671559f, 0.634393f, 0.595699f, 0.555570f, 0.514103f, 0.471397f, 0.427555f, 0.382683f, 0.336890f, 0.290285f, 0.242980f, 0.195090f, 0.146730f, 0.098017f, 0.049068f };
[0198] For decoding multichannel coding, the following C code can be used in a KLT rotation-based method.
[0199] decode_mct_rotation() { for (pair=0; pair < self->numPairs; pair++) { mctBandOffset = 0; / * Inverse MCT credit * / for (win = 0, group = 0; group <num_window_groups; group++) { for (groupwin = 0; groupwin < window_group_length[group]; groupwin++, win++) { *dmx = spectral_data[ch1][win]; *res = spectral_data[ch2][win]; apply_mct_rotation_wrapper(self,dmx,res,&alphaSfb[mctBandOffset], &mctMask[mctBandOffset],mctBandsPerWindow, alpha, totalSfb,pair,nSamples); } mctBandOffset += mctBandsPerWindow; } } }
[0200] For bandwidth processing, the following C code can be used. apply_mct_rotation_wrapper(self, *dmx, *res, *alphaSfb, *mctMask, mctBandsPerWindow, alpha, totalSfb, pair, nSamples) { sfb = 0; if (self->MCCSignalingType == 0) { } else if (self->MCCSignalingType == 1) { / * Apply fullband box * / if (!self->bHasBandwiseAngles[pair] && !self->bHasMctMask[pair]) { apply_mct_rotation(dmx, res, alphaSfb[0], nSamples); } else { / * apply bandwise processing * / for (i = 0; i< mctBandsPerWindow; i++) { if (mctMask[i] == 1) { startLine = swb_offset [sfb]; stopLine = (sfb+2 <totalSfb)? swb_offset [sfb+2] :swb_offset [sfb+1]; nSamples = stopLine-startLine; apply_mct_rotation(&dmx[startLine], &res[startLine], alphaSfb[i], nSamples); } sfb += 2; / * break condition * / if (sfb >= totalSfb) { break; } } } } else if (self->MCCSignalingType == 2) { } else if (self->MCCSignalingType == 3) { apply_mct_rotation(dmx, res, alpha, nSamples); } }
[0201] To apply the KLT rotation, the following C code can be used. apply_mct_rotation(*dmx, *res, alpha, nSamples) {<00,00905>for (n = 0; n < nSamples; n++) { L = dmx[n] * tabIndexToCosAlpha [alphaIdx] - res[n] * tabIndexToSinAlpha [alphaIdx]; R = dmx[n] * tabIndexToSinAlpha [alphaIdx] + res[n] * tabIndexToCosAlpha [alphaIdx]; dmx[n] = L; res[n] = R; } }
[0202] Figure 12 shows a flowchart of a method 400 for decoding an encoded multichannel signal having encoded channels and at least two multichannel parameters MCH_PAR1 and MCH_PAR2. The method 400 comprises the steps of: decoding the encoded channels to obtain decoded channels; and performing multichannel processing using a second pair of decoded channels identified by the multichannel parameter MCH_PAR2 and using the multichannel parameter MCH_PAR2 to obtain processed channels; and performing further multichannel processing using a first pair of channels identified by the multichannel parameter MCH_PAR1 and using the multichannel parameter MCH_PAR1, wherein the first pair of channels includes at least one processed channel.
[0203] The following describes stereo filling in multichannel coding according to an embodiment.
[0204] As already outlined, an undesirable effect of spectral quantization is that it can produce spectral holes. For example, all spectral values within a particular frequency band may be set to zero on the encoder side as a result of quantization. For instance, the exact values of such spectral lines before quantization may be relatively low, and quantization can result in a situation where the spectral values of all spectral lines within a particular frequency band are set to zero. On the decoder side, this can produce undesirable spectral holes during decoding.
[0205] While the Multi-Channel Coding Tool (MCT) in MPEG-H allows adaptation to changing inter-channel dependencies, it does not enable stereo filling because it uses single-channel elements in its normal operating configuration.
[0206] As can be seen in Figure 14, multi-channel coding tools combine three or more hierarchically coded channels. However, during coding, the way in which the multi-channel coding tool (MCT) combines different channels changes from frame to frame depending on the current signal characteristics of the channels.
[0207] For example, in scenario (a) of Figure 14, the multi-channel coding tool (MCT) may combine the first channel Ch1 and the second channel CH2 to obtain a first composite channel (processed channel) P1 and a second composite channel P2 in order to generate a first coded audio signal frame. Next, the multi-channel coding tool (MCT) can combine the first composite channel P1 and the third channel CH3 to obtain a third composite channel P3 and a fourth composite channel P4. Then, the multi-channel coding tool (MCT) can encode the second composite channel P2, the third composite channel P3, and the fourth composite channel P4 to generate the first frame.
[0208] Next, for example, in scenario (b) of Figure 14, in order to generate a second encoded audio signal frame (temporarily) following the first encoded audio signal frame, the multi-channel coding tool (MCT) may combine the first channel CH1' and the third channel CH3' to obtain a first composite channel P1' and a second composite channel P2'. Next, the multi-channel coding tool (MCT) can combine the first composite channel P1' and the second channel CH2' to obtain a third composite channel P3' and a fourth composite channel P4'. Then, the multi-channel coding tool (MCT) can encode the second composite channel P2', the third composite channel P3', and the fourth composite channel P4' to generate a second frame.
[0209] As can be seen from Figure 14, the method by which the second, third, and fourth composite channels of the first frame were generated in the scenario of Figure 14(a) differs significantly from the method by which the second, third, and fourth composite channels of the second frame were generated in the scenario of Figure 14(b), with different combinations of channels being used to generate the respective composite channels P2, P3, and P4, as well as P2', P3', and P4'.
[0210] In particular, embodiments of the present invention are based on the following findings. As shown in Figures 7 and 14, the combined channels P3, P4 and P2 (or P2', P3' and P4' in scenario (b) of Figure 14) are supplied to the channel encoder 104. In particular, the channel encoder 104 can perform quantization such that the spectral values of channels P2, P3 and P4 are set to zero for quantization. Spectrally neighboring spectral samples may be encoded as spectral bands, each spectral band can contain a number of spectral samples.
[0211] The number of spectral samples in a given frequency band may differ for different frequency bands. For example, a frequency band in a lower frequency range may contain fewer spectral samples (e.g., 4 spectral samples) than a frequency band in a higher frequency range, which may contain, for example, 16 frequency samples. For example, the critical band of the Burke scale can define the frequency band used.
[0212] A particularly undesirable situation can arise when all spectral samples in a frequency band are set to zero after quantization. When such a situation can occur, according to the present invention, stereo filling is recommended. Furthermore, the present invention does not merely generate at least (pseudo) random noise based on the findings.
[0213] According to embodiments of the present invention, instead of or in addition to adding (pseudo) random noise, for example in scenario (b) of Figure 14, if all spectral values in the frequency band of channel P4' are set to zero, the composite channel that would be generated in the same or similar manner as channel P3' would be a very suitable basis for generating noise to fill the zero-quantized frequency band.
[0214] However, according to embodiments of the present invention, it is preferable not to use the spectral values of the current frame's P3' composite channel at the current time as the basis for filling the frequency band of the P4' composite channel, as this frequency band contains only zero spectral values, and both the composite channel P3' and the composite channel P4' are generated based on channels P1' and P2', and therefore using the current P3' composite channel would simply be panning.
[0215] For example, if P3' is the mid-channel of P1' and P2' (e.g., P3' = 0.5 * (P1' + P2')) and P4' is the side-channel of P1' and P2' (e.g., P4' = 0.5 * (P1' - P2')), then introducing the attenuated spectral value of P3' into the frequency band of P4', for example, will simply result in panning.
[0216] Alternatively, it is preferable to use the channel from the previous time point to generate spectral values for filling the spectral holes in the current P4' synthesis channel. According to the findings of the present invention, the combination of the channel from the previous frame corresponding to the P3' synthesis channel in the current frame provides a desirable basis for generating a spectral sample for filling the P4' spectral holes.
[0217] However, the composite channel P3 generated in the scenario of Figure 10(a) for the previous frame does not correspond to the composite channel P3' of the current frame because the composite channel P3 of the previous frame was generated in a different way than the composite channel P3' of the current frame.
[0218] According to the findings of the embodiments of the present invention, the approximation of the P3' synthesis channel should be generated based on the reconstructed channels of the previous frame on the decoder side.
[0219] Figure 10(a) shows an encoder scenario in which channels CH1, CH2, and CH3 are encoded for the previous frame by generating E1, E2, and E3. The decoder receives channels E1, E2, and E3 and reconstructs the encoded channels CH1, CH2, and CH3. Although some encoding losses may occur, the generated channels CH1*, CH2*, and CH3* that approximate CH1, CH2, and CH3 are very similar to the original channels CH1, CH2, and CH3, so CH1*≈CH1, CH2*≈CH2, and CH3*≈CH3. According to an embodiment, the decoder maintains the generated channels CH1*, CH2*, and CH3* for the previous frame in a buffer for use in noise filling in the current frame.
[0220] Figure 1a shows an apparatus 201 for decoding according to an embodiment, which will be described in more detail here.
[0221] The apparatus 201 in Figure 1a is adapted to decode a previous encoded multi-channel signal of a previous frame to obtain three or more previous audio output channels, and is configured to decode a current encoded multi-channel signal 107 of the current frame to obtain three or more current audio output channels.
[0222] The apparatus includes an interface 212, a channel decoder 202, a multi-channel processing unit 204 for generating three or more current audio output channels CH1, CH2, CH3, and a noise filling module 220.
[0223] The interface 212 is adapted to receive the current encoded multi-channel signal 107 and receive side information including a first multi-channel parameter MCH_PAR2.
[0224] The channel decoder 202 is adapted to decode the current encoded multichannel signal of the current frame and obtain a set of three or more decoded channels D1, D2, D3 of the current frame.
[0225] The multi-channel processing unit 204 is adapted to select a first selected pair of two decoded channels D1, D2 from a set of three or more decoded channels D1, D2, D3, depending on the first multi-channel parameter MCH_PAR2.
[0226] As an example, this is shown in Figure 1a by two channels D1 and D2 supplied to a (optional) processing box 208.
[0227] Furthermore, the multi-channel processing unit 204 is adapted to generate a first group of two or more processed channels P1*, P2* based on the first selected pair of two decoded channels D1, D2, and to obtain an updated set of three or more decoded channels D3, P1*, P2*.
[0228] In the example, two channels D1 and D2 are fed into (optional) box 208, and two processed channels P1* and P2* are generated from the two selected channels D1 and D2. The updated set of three or more decoded channels includes the remaining, unmodified channel D3, and further includes the P1* and P2* generated from D1 and D2.
[0229] Before the multi-channel processing unit 204 generates a first pair of two or more processed channels P1*, P2* based on a first selected pair of two decoded channels D1, D2, the noise filling module 220 identifies one or more frequency bands in at least one of the two channels of the first selected pair of two decoded channels D1, D2 where all spectral lines are quantized to zero, and generates a mixing channel using two or more of the three or more pre-audio output channels, rather than all of them, and fills the spectral lines of one or more frequency bands where all spectral lines are quantized to zero with noise generated using the spectral lines of the mixing channel, and the noise filling module 220 is adapted to select two or more pre-audio output channels to be used to generate the mixing channel from the three or more pre-audio output channels according to the side information.
[0230] Therefore, the noise-filling module 220 analyzes whether there are any frequency bands that have only zero spectral values, and further fills any empty frequency bands found with generated noise. For example, a frequency band may have, for example, 4, 8, or 16 spectral lines, and if all spectral lines in the frequency band are quantized to zero, the noise-filling module 220 fills it with generated noise.
[0231] A particular concept in an embodiment that may be used by the noise filling module 220, which specifies how to generate and fill in noise, is called stereo filling.
[0232] In the embodiment shown in Figure 1a, the noise filling module 220 interacts with the multi-channel processing unit 204. For example, in one embodiment, if the noise filling module wants to process two channels, for example, by a processing box, these channels are supplied to the noise filling module 220, which checks whether the frequency band is quantized to zero and, if detected, fills such a frequency band.
[0233] In another embodiment shown in Figure 1b, the noise-filling module 220 interacts with the channel decoder 202. For example, when the channel decoder decodes an encoded multichannel signal to obtain three or more decoded channels D1, D2, D3, the noise-filling module checks, for example, whether a frequency band has already been quantized to zero, and if detected, fills such a frequency band. In such an embodiment, the multichannel processing unit 204 may ensure that all spectral holes are already closed before filling the noise.
[0234] In further embodiments (not shown), the noise-filling module 220 can interact with both the channel decoder and the multi-channel processing unit. For example, when the channel decoder 202 generates decoded channels D1, D2, and D3, the noise-filling module 220 may already check whether the frequency bands are quantized to zero immediately after the channel decoder 202 generates them, but can only generate noise and fill their respective frequency bands when the multi-channel processing unit 204 actually processes these channels.
[0235] For example, random noise, computationally inexpensive operations, can be inserted into any of the frequency bands quantized to zero, but the noise-filling module may fill with noise generated from previously generated audio output channels only if they are actually processed by the multi-channel processing unit 204. However, in such embodiments, before inserting random noise, it is necessary to detect whether or not spectral holes exist, and this information should be maintained in memory, because after inserting random noise, each frequency band will have a non-zero spectral value due to the insertion of random noise.
[0236] In this embodiment, in addition to the noise generated based on the previous audio output signal, random noise is inserted into a frequency band quantized to zero.
[0237] In some embodiments, interface 212 may be adapted to receive, for example, the current encoded multichannel signal 107 and side information including a first multichannel parameter MCH_PAR2 and a second multichannel parameter MCH_PAR1.
[0238] The multichannel processing unit 204 may be adapted, for example, to select a second selected pair of two decoded channels P1*, D3 from an updated set of three or more decoded channels D3, P1*, P2*, depending on a second multichannel parameter MCH_PAR1, such that at least one channel P1* of the second selected pair of two decoded channels (P1*, D3) is one channel of the first pair of two or more processed channels P1*, P2*.
[0239] The multi-channel processing unit 204 may be adapted to generate a second group of two or more processed channels P3*, P4* based on the second selected pair of two decoded channels P1, D3, and to further update the updated set of three or more decoded channels.
[0240] An example of such an embodiment is shown in Figures 1a and 1b, in which an (optional) processing box 210 receives channel D3 and processed channel P1*, processes them to obtain processed channels P3* and P4*, and a further updated set of three decoded channels, including P2* which has not been modified by the processing box 210, and the generated P3* and P4*.
[0241] Processing boxes 208 and 210 are marked as optional in Figures 1a and 1b. This is to indicate that while processing boxes 208 and 210 may be used to implement the multi-channel processing unit 204, there are various possibilities for how the multi-channel processing unit 204 can be implemented. For example, instead of using different processing boxes 208 and 210 for different processing of two (or more) channels, the same processing boxes can be reused, or the multi-channel processing unit 204 may perform processing for two channels without using processing boxes 208 and 210 (as a subunit of the multi-channel processing unit 204).
[0242] In a further embodiment, the multi-channel processing unit 204 may be adapted to generate two or more first groups of processed channels P1*, P2* by generating a first group of exactly two processed channels P1*, P2* based on the first selected pair of two decoded channels D1*, D2*. The multi-channel processing unit 204 may be adapted to replace the first selected pair of two decoded channels D1*, D2* in a set of three or more decoded channels D1, D2, D3 with a first group of exactly two processed channels P1*, P2* to obtain an updated set of three or more decoded channels D3, P1*, P2*. The multi-channel processing unit 204 may be adapted to generate two or more second groups of processed channels P3*, P4* by generating a second group of exactly two processed channels P3*, P4* based on the second selected pair of two decoded channels P1*, D3*. Furthermore, the multi-channel processing unit 204 may be adapted to further update the updated set of three or more decoded channels by, for example, replacing the second selected pair of two decoded channels P1*, D3 in an updated set of three or more decoded channels D3, P1*, P2* with a second group of exactly two processed channels P3*, P4*.
[0243] In such embodiments, exactly two processed channels are generated from two selected channels (for example, two input channels of processing box 208 or 210), and these exactly two processed channels replace the selected channels in a set of three or more decoded channels. For example, processing box 208 of the multichannel processing unit 204 replaces selected channels D1 and D2 with P1* and P2*.
[0244] However, in other embodiments, upmixing may be performed within the device 201 for decoding, and three or more processed channels may be generated from two selected channels, or not all of the selected channels may be removed from the updated set of decoded channels.
[0245] A further challenge is the method for generating the mixing channels used to generate the noise produced by the noise-filling module 220.
[0246] According to some embodiments, the noise filling module 220 may be adapted to generate a mixing channel using exactly two of the three or more pre-audio output channels as two or more pre-audio output channels, for example, or the noise filling module 220 may be adapted to select exactly two pre-audio output channels from the three or more pre-audio output channels depending on side information.
[0247] Using only two of three or more pre-output channels helps reduce the complexity of the calculations required to determine the mixing channels.
[0248] However, in other embodiments, three or more pre-audio output channels are used to generate the mixing channel, but the number of pre-audio output channels considered is less than the total number of three or more pre-audio output channels.
[0249] In an embodiment where only two of the pre-output channels are considered, the mixing channel may be calculated, for example, as follows:
[0250] In one embodiment, the noise filling module 220 is, TIFF0007885285000016.tif1993 or formula Based on TIFF0007885285000017.tif2094, it has been adapted to generate a mixing channel using exactly two pre-audio output channels. Here D ch This is a mixing channel, TIFF0007885285000018.tif75 is the first of two exact pre-audio output channels, TIFF0007885285000019.tif65 is the second of two exact pre-audio output channels, and unlike the first of two exact pre-audio output channels, d is a real positive scalar.
[0251] In a typical situation, the midchannel TIFF0007885285000020.tif1989 may be the appropriate mixing channel. This method calculates the mixing channel as the mid channel of the two pre-audio output channels being considered.
[0252] However, in some scenarios, When applying TIFF0007885285000021.tif1991, for example, In the case of TIFF0007885285000022.tif618, a mixing channel close to zero may occur. Next, for example, It may be preferable to use TIFF0007885285000023.tif1989 as a mixing signal. Therefore, a side channel (for phase-shifted input channels) is used.
[0253] In an alternative method, the noise-filling module 220 is given by equation TIFF0007885285000024.tif20162 or formula Based on TIFF0007885285000025.tif14116, it has been adapted to generate a mixing channel using exactly two pre-audio output channels. Here TIFF0007885285000026.tif66 is a mixing channel, TIFF0007885285000027.tif65 is the first of two exact pre-audio output channels, TIFF0007885285000028.tif65 is the second of two exact pre-audio output channels, and unlike the first of two exact pre-audio output channels, α is the rotation angle.
[0254] This method calculates the mixing channel by rotating the two pre-audio output channels being considered.
[0255] The rotation angle α may be in the range, for example, -90° < α < 90°. In one embodiment, the rotation angle may be, for example, within the range of 30° < α < 60°.
[0256] Again, in typical situations, channel TIFF0007885285000029.tif19155 may be the appropriate mixing channel. This method calculates the mixing channel as the mid channel of the two pre-audio output channels being considered.
[0257] However, in some scenarios, When applying TIFF0007885285000030.tif19157, for example, In the case of TIFF0007885285000031.tif16120, a mixing channel close to zero may occur. Next, for example, It may be preferable to use TIFF0007885285000032.tif20169 as a mixing signal.
[0258] According to a particular embodiment, the side information may be, for example, the current side information assigned to the current frame, and the interface 212 may be adapted to receive, for example, previous side information assigned to the previous frame, which includes a previous angle, and the interface 212 may be adapted to receive, for example, current side information, which includes the current angle, and the noise filling module 220 may be adapted to use, for example, the current angle of the current side information as the rotation angle α, and not to use the previous angle of the previous side information as the rotation angle α.
[0259] Therefore, in such embodiments, even when the mixing channel is calculated based on the previous audio output channel, the current angle transmitted in the side information is used as the rotation angle, rather than the previously received rotation angle, but the mixing channel is calculated based on the previous audio output channel generated based on the previous frame.
[0260] Another aspect of some embodiments of the present invention relates to a scale factor. The frequency band may be, for example, the scale factor band.
[0261] According to some embodiments, before the multi-channel processing unit 204 generates a first pair of two or more processed channels P1*, P2* based on a first selected pair (D1, D2) of two decoded channels, the noise filling module (220) may, for example, be adapted to identify one or more scale factor bands for at least one of the two channels of the first selected pair D1, D2 of the two decoded channels, which are one or more frequency bands in which all spectral lines are quantized to zero, and may be adapted to generate a mixing channel using two or more of the three or more pre-audio output channels, rather than all of them, and may be adapted to fill the spectral lines of one or more frequency bands in which all spectral lines are quantized to zero with noise generated using the spectral lines of the mixing channel, depending on the scale factor of each of the one or more scale factor bands in which all spectral lines are quantized to zero.
[0262] In such embodiments, the scale factor may be assigned to each of the scale factor bands, for example, and the scale factor is taken into consideration when generating noise using the mixing channel.
[0263] In certain embodiments, the receiving interface 212 is configured, for example, to receive each of the scale factors of the one or more scale factor bands, where each scale factor of the one or more scale factor bands represents the energy of the spectral lines of the scale factor band before quantization. The noise-filling module 220 may be adapted, for example, to generate noise for each of the one or more scale factor bands, where all spectral lines are quantized to zero, and as a result, after adding the noise to one of the frequency bands, the energy of the spectral lines corresponds to the energy represented by the scale factor for the scale factor band.
[0264] For example, a mixing channel may show spectral values for four spectral lines in a scale factor band into which noise is inserted, and these spectral values may be, for example, 0.2, 0.3, 0.5, and 0.1.
[0265] The energy of the mixing channel's scale factor bandwidth may be calculated, for example, as follows: TIFF0007885285000033.tif11136
[0266] However, the scale factor relative to the bandwidth of the channel filled with noise may be as small as, for example, 0.0039.
[0267] The damping coefficient can be calculated, for example, as follows:
[0268]
number
[0269] Therefore, in the above example,
[0270]
number
[0271] In one embodiment, each spectral value in the scale factor band of the mixing channel used as noise is multiplied by the attenuation factor.
[0272] Therefore, each of the four spectral values in the scale factor band in the example above is multiplied by the attenuation factor to obtain the attenuated spectral value. 0.2 * 0.01 = 0.002 0.3 * 0.01 = 0.003 0.5 * 0.01 = 0.005 0.1 * 0.01 = 0.001
[0273] These attenuated spectral values may, for example, be inserted into the scale factor band of the channel that is filled with noise.
[0274] The above examples can be equally applied to logarithmic values by replacing the operations with their corresponding logarithmic operations, such as replacing multiplication with addition.
[0275] Furthermore, in addition to the description of the specific embodiments described above, other embodiments of the noise-filling module 220 apply one, some, or all of the concepts described with reference to Figures 2 to 6.
[0276] Another embodiment of the present invention relates to a problem based on the selection of an information channel from a pre-audio output channel to be used to generate a mixing channel to obtain inserted noise.
[0277] According to one embodiment, the device with the noise filling module 220 may be adapted to select exactly two pre-audio output channels from three or more pre-audio output channels, for example, depending on a first multi-channel parameter MCH_PAR2.
[0278] Therefore, in such embodiments, the first multi-channel parameter that adjusts which channel to select for processing also adjusts which pre-audio output channel to use to generate the mixing channel for generating the noise to be inserted.
[0279] In one embodiment, the first multichannel parameter MCH_PAR2 may indicate, for example, two decoded channels D1, D2 from a set of three or more decoded channels, and the multichannel processing unit 204 is adapted to select a first selected pair of two decoded channels D1, D2, D3 from a set of three or more decoded channels D1, D2, D3 by selecting the two decoded channels D1, D2 indicated by the first multichannel parameter MCH_PAR2. Furthermore, the second multichannel parameter MCH_PAR1 may indicate, for example, two decoded channels P1*, D3 from an updated set of three or more decoded channels. The multichannel processing unit 204 may be adapted to select a second selected pair of two decoded channels P1*, D3 from an updated set of three or more decoded channels D3, P1*, P2* by selecting the two decoded channels P1*, D3 indicated by the second multichannel parameter MCH_PAR1.
[0280] Therefore, in such embodiments, the channels selected for the first processing, for example, the processing box 208 in Figure 1a or Figure 1b, do not depend solely on the first multi-channel parameter MCH_PAR2. Furthermore, these two selected channels are explicitly specified in the first multi-channel parameter MCH_PAR2.
[0281] Similarly, in such embodiments, the channels selected for the second processing, for example, the processing box 210 in Figure 1a or Figure 1b, do not depend solely on the second multi-channel parameter MCH_PAR1. Furthermore, these two selected channels are explicitly specified in the second multi-channel parameter MCH_PAR1.
[0282] Embodiments of the present invention introduce a sophisticated indexing scheme for multichannel parameters, as described with reference to Figure 15.
[0283] Figure 15(a) shows the encoding of five channels, namely the left channel, right channel, center channel, left surround channel, and right surround channel, on the encoder side. Figure 15(b) shows the decoding of the encoded channels E0, E1, E2, E3, and E4 in order to reconstruct the left channel, right channel, center channel, left surround channel, and right surround channel.
[0284] Let's assume that an index is assigned to each of the five channels: left, right, center, left surround, and right surround. Index channel name 0 Left 1 Right 2 center 3 Left Surround 4 Right Surround
[0285] In Figure 15(a), on the encoder side, the first operation performed within the processing box 192 may be, for example, mixing channel 0 (left) and channel 3 (left surround), resulting in two processed channels. It can be assumed that one of the processed channels is the mid channel and the other is the side channel. However, other concepts that form the two processed channels, such as determining the two processed channels by performing a rotational operation, may also be applied.
[0286] Thus, the two generated and processed channels acquire the same index as the channel used for processing. That is, the first channel of the processed channel has index 0, and the second channel of the processed channel has index 3. The multi-channel parameters determined for this processing may be, for example, (0;3).
[0287] The second operation performed on the encoder side may be, for example, mixing channel 1 (right) and channel 4 (right surround) in the processing box 194 to obtain two further processed channels. Again, the two further generated and processed channels acquire the same index as the channel used for processing. That is, the first of the further processed channels has index 1, and the second of the processed channels has index 4. The multi-channel parameters determined for this processing may be, for example, (1;4).
[0288] The third operation performed on the encoder side may, for example, involve mixing processed channel 0 and processed channel 1 in the processing box 196 to obtain two other processed channels. Again, these two generated and processed channels acquire the same index as the channel used for processing. That is, the first of the further processed channels has index 0, and the second of the processed channels has index 1. The multi-channel parameters determined for this processing may, for example, be (0;1).
[0289] The encoded channels E0, E1, E2, E3, and E4 are distinguished by their indices, namely E0 has index 0, E1 has index 1, and E2 has index 2.
[0290] The encoder performs three calculations, resulting in three multi-channel parameters. (0;3),(1;4),(0;1)
[0291] Since the decoding device is supposed to perform the encoder operations in reverse order, the order of the multichannel parameters may be reversed, for example, when they are sent to the device for decoding. (0;1),(1;4),(0;3)
[0292] In the decoding device, (0;1) can be called the first multichannel parameter, (1,4) the second multichannel parameter, and (0,3) the third multichannel parameter.
[0293] On the decoder side shown in Figure 15(b), upon receiving the first multi-channel parameter (0;1), the decoder determines that this is the first processing operation on the decoder side and processes channel 0 (E0) and channel 1 (E1). This is done in box 296 in Figure 15(b). Both generated and processed channels inherit the indices from channels E0 and E1 used to generate them, and therefore the generated and processed channels also have indices 0 and 1.
[0294] When the decoder receives the second multi-channel parameters (1;4), it determines that this is the second processing operation on the decoder side and processes the processed channels 1 and 4 (E4). This is done in box 294 of Figure 15(b). Both generated and processed channels inherit the indices from channels 1 and 4 used to generate them, and therefore the generated and processed channels also have indices 1 and 4.
[0295] When the decoder receives the third multi-channel parameter (0;3), it determines that this is the third processing operation on the decoder side and processes the processed channels 0 and 3 (E3). This is done in box 292 of Figure 15(b). Both generated and processed channels inherit the indices from channels 0 and 3 used to generate them, and therefore the generated and processed channels also have indices 0 and 3.
[0296] As a result of processing by the decoding device, the left (index 0), right (index 1), center (index 2), left surround (index 3), and right surround (index 4) channels are reconstructed.
[0297] On the decoder side, for quantization purposes, it is assumed that all values of channel E1 (index 1) within a specific scale factor band are quantized to zero. A noise-filled channel 1 (channel E1) is desirable if the decoder wants to perform the processing in box 296.
[0298] As already outlined, the embodiment uses two pre-audio output signals to fill the spectral hole in channel 1 with noise.
[0299] In certain embodiments, if the channel on which the operation is performed has a scale factor bandwidth that is quantized to zero, two pre-audio output channels are used to generate noise having the same index numbers as the two channels that must be processed. In this example, if a spectral hole is detected in channel 1 before processing in processing box 296, the pre-audio output channels having index 0 (previous left channel) and index 1 (previous right channel) are used to generate noise to fill the spectral hole in channel 1 on the decoder side.
[0300] Since the index is consistently inherited by the processed channel resulting from the processing, if the previous output channel becomes the current audio output channel, it can be inferred that the previous output channel plays a role in generating the channel involved in the actual processing on the decoder side. Therefore, a good estimation of the zero-quantized scale factor bandwidth can be achieved.
[0301] According to the embodiment, the device may be adapted to assign an identification unit from a set of identification units to each of three or more pre-audio output channels, so that each of the three or more pre-audio output channels is assigned to exactly one identification unit from the set of identification units, and each identification unit from the set of identification units is assigned to exactly one of the three or more pre-audio output channels. Furthermore, the device may be adapted to assign an identification unit from the set of identification units to each of three or more decoded channels, so that each of the three or more decoded channels is assigned to exactly one identification unit from the set of identification units, and each identification unit from the set of identification units is assigned to exactly one of the three or more decoded channels.
[0302] Furthermore, the first multichannel parameter MCH_PAR2 can, for example, indicate a first pair of two identification units from a set of three or more identification units. The multichannel processing unit 204 may be adapted to select a first selected pair of two decoded channels D1, D2 from a set of three or more decoded channels D1, D2, D3, for example, by selecting two decoded channels D1, D2 to be assigned to the two identification units of the first pair of two identification units.
[0303] The device may be adapted, for example, to assign the first identification unit of a first pair of two identification units to a first processed channel of a first group of exactly two processed channels P1*, P2*. Furthermore, the device may be adapted, for example, to assign the second identification unit of a first pair of two identification units to a second processed channel of a first group of exactly two processed channels P1*, P2*.
[0304] The set of identification units may be, for example, a set of indices, for example, a set of non-negative integers (for example, a set including identification units 0, 1, 2, 3 and 4).
[0305] In certain embodiments, the second multichannel parameter MCH_PAR1 may, for example, indicate a second pair of two identification units of a set of three or more identification units. The multichannel processing unit 204 may be adapted to select a second selected pair of two decoded channels P1*, D3 from an updated set of three or more decoded channels D3, P1*, P2* by, for example, selecting two decoded channels (D3, P1*) to be assigned to the two identification units of the second pair of two identification units. Furthermore, the device may be adapted to assign, for example, the first identification unit of the two identification units of the second pair of two identification units to the first processed channel of a second group of exactly two processed channels P3*, P4*. Furthermore, the device may be adapted to assign, for example, the second identification unit of the two identification units of the second pair of two identification units to the second processed channel of a second group of exactly two processed channels P3*, P4*.
[0306] In certain embodiments, the first multichannel parameter MCH_PAR2 may, for example, indicate the first pair of two identification units of a set of three or more identification units. The noise-filling module 220 may be adapted to select exactly two pre-audio output channels from three or more pre-audio output channels, for example, by selecting two pre-audio output channels to be assigned to the two identification units of the first pair of two identification units.
[0307] As already outlined, Figure 7 shows an apparatus 100 for encoding a multichannel signal 101 having at least three channels (CH1 to CH3) according to one embodiment.
[0308] The device includes an iterative processing unit 102 adapted to calculate inter-channel correlation values between each pair of at least three channels (CH~CH3) in the first iteration step, in order to select a pair having the highest value or a pair having a value above a threshold, and to process the selected pair using multi-channel processing operations 110, 112 to derive an initial multi-channel parameter MCH_PAR1 for the selected pair, and to derive first processed channels P1, P2.
[0309] The iterative processing unit 102 is adapted to use at least one of the processed channels P1 to perform calculation, selection, and processing in a second iteration step to derive further multichannel parameters MCH_PAR2 and second processed channels P3, P4.
[0310] Furthermore, the device includes a channel encoder adapted to encode channels (P2-P4) resulting from iterative processing performed by the iteration processing unit 104 in order to obtain encoded channels (E1-E3).
[0311] Furthermore, the device includes an output interface 106 adapted to generate an encoded channel signal 107 having encoded channels (E1-E3), initial multichannel parameters, and further multichannel parameters MCH_PAR1 and MCH_PAR2.
[0312] Furthermore, the device includes an output interface 106 adapted to generate an encoded multichannel signal 107 that includes information indicating whether or not the decoder should fill in spectral lines in one or more frequency bands where all spectral lines are quantized to zero, using noise generated based on previously decoded audio output channels that have been previously decoded by the decoder.
[0313] Therefore, the encoding device can signal whether or not the decoding device should fill spectral lines in one or more frequency bands where all spectral lines are quantized to zero, using noise generated based on previously decoded audio output channels that were previously decoded by the decoding device.
[0314] According to one embodiment, each of the initial multichannel parameter and the further multichannel parameters MCH_PAR1 and MCH_PAR2 indicates exactly two channels, and each of the exactly two channels is either one of the encoded channels (E1-E3), one of the first or second processed channels P1, P2, P3, P4, or one of at least three channels (CH1-CH3).
[0315] The output interface 106 is adapted, for example, to generate an encoded multichannel signal 107, and includes information indicating whether the decoder should fill in spectral lines in one or more frequency bands where all spectral lines are quantized to zero, for each of the initial and multichannel parameters MCH_PAR1 and MCH_PAR2, for at least one channel of exactly two channels indicated by one of the initial and further multichannel parameters MCH_PAR1 and MCH_PAR2, using spectral data generated based on previously decoded audio output channels that the decoder has previously decoded, to indicate whether the decoder should fill in spectral lines in one or more frequency bands where all spectral lines of the at least one channel are quantized to zero.
[0316] Furthermore, the following describes a specific embodiment in which such information is transmitted using a hasStereoFilling[pair] value that indicates whether or not stereo filling should be applied to the MCT channel pair currently being processed.
[0317] Figure 13 shows a system according to an embodiment. This system comprises an encoding device 100 as described above and a decoding device 201 according to one of the embodiments described above.
[0318] The decoding device 201 is configured to receive the encoded multi-channel signal 107 generated by the encoding device 100 from the encoding device 100.
[0319] Furthermore, an encoded multi-channel signal 107 is provided. The encoded multichannel signal is - Encoded channels (E1~E3), - Multi-channel parameters MCH_PAR1, MCH_PAR2, -The spectral lines in one or more frequency bands where all spectral lines are quantized to zero, using spectral data generated based on previously decoded audio output channels that were previously decoded by the decoding device, to indicate whether or not the decoding device should fill them. Includes.
[0320] According to one embodiment, the encoded multichannel signal may include two or more multichannel parameters, for example, as multichannel parameters MCH_PAR1, MCH_PAR2, etc.
[0321] Each of two or more multichannel parameters MCH_PAR1, MCH_PAR2 can, for example, represent exactly two channels, each of which may be one of the encoded channels (E1-E3), one of several processed channels P1, P2, P3, P4, or one of at least three original (e.g., unprocessed) channels (CH-CH3).
[0322] Information indicating whether the decoding device should fill in spectral lines in one or more frequency bands where all spectral lines are quantized to zero may include, for example, information for each of two or more multichannel parameters MCH_PAR1, MCH_PAR2, for at least one channel of exactly two channels indicated by one of the two or more multichannel parameters, using spectral data generated based on previously decoded audio output channels that the decoding device has previously decoded, indicating whether the decoding device should fill in spectral lines in one or more frequency bands where all spectral lines of at least one channel are quantized to zero.
[0323] As already outlined, the following describes a specific embodiment in which such information is transmitted using a hasStereoFilling[pair] value indicating whether or not stereo filling should be applied to the MCT channel pair currently being processed.
[0324] The following sections will describe the general concepts and specific embodiments in more detail. The embodiment achieves a combination of stereo filling and MCT with the flexibility of using an arbitrary stereo tree for a parametric low bitrate encoding mode.
[0325] Inter-channel signal dependencies are exploited by hierarchically applying known combined stereo coding tools. For lower bitrates, embodiments extend MCT to use a combination of discrete stereo coding boxes and stereo fill boxes. Thus, semiparametric coding can be applied, for example, to channels with similar content, i.e., the channel pair with the highest correlation, while different channels can be coded independently or via nonparametric representations. Accordingly, the MCT bitstream syntax is extended to allow signaling when stereo fill is permitted and active.
[0326] The embodiment enables the generation of a previous downmix for any stereo filling pair.
[0327] Stereo filling relies on the use of a downmix of the previous frame to improve the filling of spectral holes through quantization in the frequency domain. However, in combination with MCT, the set of coupled-encoded stereo pairs can now change over time. As a result, two coupled-encoded channels may not have been coupled-encoded in the previous frame, i.e., when the tree configuration changed.
[0328] To estimate the previous downmix, the previously decoded output channels are saved and processed in inverse stereo operation. For a given stereo box, this is done using the parameters of the current frame and the decoded output channels of the previous frame corresponding to the channel indices of the processed stereo box.
[0329] If the previous output channel signal is unavailable due to an independent frame (a frame that can be decoded without considering the previous frame data) or a change in conversion length, the previous channel buffer for the corresponding channel is set to zero. Therefore, a non-zero previous downmix can be calculated as long as at least one of the previous channel signals is available.
[0330] If the MCT is configured to use a prediction-based stereo box, the pre-downmix is calculated using the inverse MS operation specified for the stereo-filling pair, preferably using one of the following two expressions based on the prediction direction flag (pred_dir in MPEG-H syntax): TIFF0007885285000036.tif535TIFF0007885285000037.tif535, Here, TIFF0007885285000038.tif43 is an arbitrary real scalar and a positive scalar.
[0331] If the MCT is configured to use a rotation-based stereo box, the pre-downmix is calculated using rotation with a negative rotation angle.
[0332] Therefore, for a rotation given as follows, The reverse rotation of TIFF0007885285000039.tif1054 is calculated as follows: TIFF0007885285000040.tif1254 and TIFF0007885285000041.tif54 are the front output channels. TIFF0007885285000042.tif55 and This is the desired pre-downmix of TIFF0007885285000043.tif55.
[0333] The embodiment realizes the application of stereo filling in MCT. Methods for applying stereo filling to a single stereo box are described in [1] and [5].
[0334] For a single stereo box, stereo filling is applied to the second channel of a given MCT channel pair.
[0335] In particular, the differences in stereo filling when combined with MCT are as follows: The MCT tree configuration is extended by one signaling bit per frame to allow signaling whether or not stereo filling is permitted in the current frame.
[0336] In a preferred embodiment, if stereo filling is permitted in the current frame, an additional bit is sent to each stereo box to activate stereo filling in the stereo box. This is a preferred embodiment because the encoder can control which box should have the stereo filling applied in the decoder.
[0337] In the second embodiment, if stereo filling is permitted for the current frame, stereo filling is permitted in all stereo boxes, and no additional bits are transmitted per individual stereo box. In this case, the selective application of stereo filling in individual MCT boxes is controlled by the decoder.
[0338] Further concepts and detailed embodiments are described below. The embodiment improves the quality of low-bitrate multi-channel operating point.
[0339] In frequency-domain (FD) encoded channel pair elements (CPEs), for perceptually improved filling of spectral holes caused by very coarse quantization in the encoder, the MPEG-H 3D audio standard enables the use of the stereo filling tool described in section 5.5.5.4.9 of [1]. This tool has been shown to be particularly useful for 2-channel stereo encoded at medium and low bitrates.
[0340] The Multi-Channel Coding Tool (MCT), described in Section 7 of [2], is introduced, which allows for flexible signal-adaptive definition of frame-by-frame coupled-coded channel pairs to take advantage of time-varying inter-channel dependencies in multi-channel setups. The benefits of MCT are particularly significant when used for efficient dynamic coupled coding of multi-channel setups where each channel resides in individual single-channel elements (SCEs), unlike the conventional CPE+SCE(+LFE) configuration which must be established a priori, as it allows coupled channel coding to be carried over and / or reconfigured from one frame to the next.
[0341] Encoding multichannel surround sound without using CPE has the disadvantage of not being able to utilize the combined stereo tools available only with CPE—predictive M / S encoding and stereo filling—which is particularly detrimental at medium and low bitrates. While MCT can serve as a substitute for M / S tools, there are currently no alternatives available for stereo filling tools.
[0342] The embodiment extends the MCT bitstream syntax with respect to each signal-transmitting bit, generalizing the application of stereo filling to any channel pair regardless of the channel element type, thereby enabling the use of stereo filling tools even within channel pairs of MCTs.
[0343] Some embodiments can achieve stereo filling signal transmission in MCT as follows, for example.
[0344] In CPE, the use of the stereo filling tool is signaled within the FD noise filling information of the second channel, as described in section 5.5.5.4.9.4 of [1]. When using MCT, all channels are potentially "second channels" (because there are possible channel pairs between elements). Therefore, it is proposed to explicitly signal stereo filling with an additional bit for each MCT-encoded channel pair. The presence of the aforementioned per-channel-pair addition is signaled using the two currently reserved entries of the MCTSignalingType element of MultichannelCodingFrame()[2] so that this additional bit is unnecessary if stereo filling is not used for any channel pair in a particular MCT "tree" instance.
[0345] A detailed explanation follows. Some embodiments can implement pre-downmix calculations, for example, as follows:
[0346] Stereo filling in CPE fills a specific "empty" scale factor band of a second channel by adding the respective MDCT coefficients of the downmix of the previous frame, scaled according to the transmit scale factor of the corresponding band (which is unused because the aforementioned band is fully quantized to zero). A weighted addition process controlled using the scale factor band of the target channel can be used similarly in the context of MCT. However, the source spectrum for stereo filling, i.e., the downmix of the previous frame, must be calculated in a different way than in CPE, particularly because the MCT "tree" configuration can change over time.
[0347] In MCT, the pre-downmix may be derived from the decoded output channels of the last frame (stored after MCT decoding) using the MCT parameters of the current frame for a given coupled channel pair. For pairs applying predictive M / S-based coupled coding, the pre-downmix will be the same as in the case of CPE stereo filling, either by the sum or difference of the appropriate channel spectra, depending on the directional indicator of the current frame. For stereo pairs using Karhunen-Loeve rotation-based coupled coding, the pre-downmix represents the inverse rotation calculated at the rotation angle of the current frame. Again, a detailed explanation is provided below.
[0348] In complexity assessments, stereo filling in the MCT, a medium and low bitrate tool, is not considered to increase the worst complexity when measured at both low / medium and high bitrates. Furthermore, using stereo filling typically coincides with more spectral coefficients being quantized to zero, thereby reducing the complexity of the context-based arithmetic decoder algorithm. Assuming a maximum of N / 3 stereo-filled channels are used in an N-channel surround configuration and an additional 0.2 WMOPS is used per stereo filling operation, with the coder's sampling rate at 48 kHz and the IGF tool operating only above 12 kHz, the peak complexity increases by only 0.4 WMOPS for 5.1 and 0.8 WMOPS for 11.1 channels. This is less than 2% of the overall decoder complexity.
[0349] The embodiment implements the MultichannelCodingFrame() element as follows.
[0350] [Table 8]
[0351] According to some embodiments, stereo filling in MCT may be carried out as follows:
[0352] Similar to the stereo filling of channel pair elements in IGF as described in section 5.5.5.4.9 of [1], stereo filling in multichannel coding tools (MCTs) fills the “empty” scale factor band (quantized to exactly zero) above the noise filling start frequency using a downmix of the output spectrum of the previous frame.
[0353] When stereo filling is active in an MCT-coupled channel pair (hasStereoFilling[pair]≠0 in Table AMD4.4), all "empty" scale factor bands in the noise filling region of the second channel of the pair (i.e., starting above noiseFillingStartOffset) are filled up to a specific target energy using a downmix of the corresponding output spectrum (after MCT application) of the previous frame. This is done after FD noise filling (see section 7.2 of ISO / IEC 23003-3:2012) and before scale factor and MCT-coupled stereo application. All output spectra after MCT processing are saved for potential stereo filling in the next frame.
[0354] An operational constraint may be, for example, that cascading execution of the stereo filling algorithm (hasStereoFilling[pair]≠0) in the free bandwidth of the second channel is not supported for any subsequent MCT stereo pair using hasStereoFilling[pair]≠0 if the second channel is the same. For channel pair elements, active IGF stereo filling of the second (residual) channel according to section 5.5.5.4.9 of [1] takes precedence over any subsequent application of MCT stereo filling on the same channel in the same frame, and is therefore invalid.
[0355] Terms and definitions can be defined, for example, as follows: hasStereoFilling[pair] Indicates the use of stereo filling for currently processed MCT channel pairs. ch1, ch2 Channel indices of the currently processed MCT channel pair spectral_data[][] Spectral coefficients of channels in the currently processed MCT channel pair spectral_data_prev[][] Output spectrum after MCT processing is completed in the previous frame downmix_prev[][] Estimated downmix of the output channel of the previous frame using the index given by the currently processed MCT channel pair. num_swb Total number of scale factor bandwidths; see ISO / IEC 23003-3, section 6.2.9.4. Refer to ccfl coreCoderFrameLength, conversion length, ISO / IEC 23003-3, section 6.1. noiseFillingStartOffset is the noise filling start line defined according to ccfl in Table 109 of ISO / IEC 23003-3. igf_WhiteningLevel Spectral whitening in IGF, see ISO / IEC 23008-3, section 5.5.5.4.7. seed[] Noise-filling seed used by randomSign(), see ISO / IEC 23003-3, section 7.2.
[0356] In some specific embodiments, the decryption process may be described as follows, for example:
[0357] MCT stereo filling is performed using the following four sequential operations. Step 1: Prepare the spectrum of the second channel for the stereo filling algorithm. If the stereo filling indicator hasStereoFilling[pair] for a given MCT channel pair is 0, stereo filling is not used and the following steps are not performed. Otherwise, if a scale factor application has been previously applied to the second channel spectrum of the pair, spectral_data[ch2], then the scale factor application is not performed.
[0358] Step 2: Generation of pre-downmix spectra for a given MCT channel pair The pre-downmix is estimated from the output signal spectral_data_prev[][] of the previous frame stored after the application of MCT processing. If the pre-output channel signal is unavailable, for example, in an independent frame (indepFlag>0), a conversion length change, or core_mode==1, the pre-channel buffer for the corresponding channel is set to zero.
[0359] For predicted stereo pairs, i.e., MCTSignalingType==0, the pre-downmix is calculated from the pre-output channel as downmix_prev[][] defined in step 2 of section 5.5.5.4.9.4 of [1], and spectrum[window][] is represented by spectral_data[][window].
[0360] For rotated stereo pairs, i.e., when MCTSignalingType==1, the pre-downmix is calculated from the pre-output channel by inverting the rotation operation defined in section 5.5.X.3.7.1 of [2].
[0361] apply_mct_rotation_inverse(*R, *L, *dmx, aIdx, nSamples) { for(n=0;n <nSamples;n++){ dmx=L[n]*tabIndexToCosAlpha[aIdx]+R[n]*tabIndexToSinAlpha[aIdx]; } } The L=spectral_data_prev[ch1][], R=spectral_data_prev[ch2][], and dmx=downmix_prev[] values from the previous frame are used, along with the aIdx and n samples from the current frame and the MCT pair.
[0362] Step 3: Execute the stereo filling algorithm in the available bandwidth of the second channel. Stereo filling is applied to the second channel of the MCT pair, as in step 3 of section 5.5.5.4.9.4 of [1], and spectrum[window] is It is represented by spectral_data[ch2][window], and max_sfb_ste is given by num_swb.
[0363] Step 4: Apply the scale factor and adaptive synchronization of the noise-filling seed. After step 3 of section 5.5.5.4.9.4 of [1], the scale factor is applied to the resulting spectrum as in ISO / IEC 23003-3 7.3, and the scale factor for empty bandwidths is handled as a normal scale factor. If the scale factor is not defined, its value may be equal to zero, for example, because it is above max_sfb. If IGF is used and igf_WhiteningLevel is equal to 2 on any of the tiles of the second channel, and neither channel uses eight short transformations, the spectral energies of both channels of the MCT pair are calculated in the range from index noiseFillingStartOffset to index ccfl / 2-1 before decode_mct() is performed. If the calculated energy of the first channel is more than eight times the energy of the second channel, the seed of the second channel [ch2] is set to be equal to the seed of the first channel [ch1].
[0364] Some embodiments are described in the context of apparatus, but these embodiments also represent a description of the corresponding method, and it is clear that the block or apparatus corresponds to a method step or a feature of a method step. Similarly, embodiments described in the context of a method step also represent a description of an item or feature of the corresponding block or corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device, such as a microprocessing unit, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.
[0365] Depending on the specific implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. Embodiments may be implemented using digital storage media such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, which have electronically readable control signals stored therein and cooperate with (or are capable of cooperating with) a computer system that is programmable for each method to be performed. Thus, the digital storage media may be computer-readable.
[0366] Some embodiments of the present invention include a data carrier having an electronically readable control signal, such that one of the methods described herein is performed in cooperation with a programmable computer system.
[0367] Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to execute one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a machine-readable carrier.
[0368] Other embodiments include a computer program for performing one of the methods described herein, which is stored on a machine-readable carrier.
[0369] In other words, embodiments of the method of the present invention are computer programs having program code for performing one of the methods of the present invention when the computer program is executed on a computer.
[0370] Accordingly, further embodiments of the methods of the present invention include a computer program for performing one of the methods described herein, and a data carrier (or digital storage medium or computer-readable medium) recorded therein. The data carrier, digital storage medium or recording medium is typically tangible and / or non-temporary.
[0371] Accordingly, a further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals may be configured to be transmitted, for example, over a data communication connection, such as the Internet.
[0372] Further embodiments include processing means configured to perform or applied to perform one of the methods described herein, such as a computer or a programmable logic device.
[0373] Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.
[0374] Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0375] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessing unit to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0376] The apparatus described herein can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.
[0377] The methods described herein may be performed using hardware devices, or using a computer, or using a combination of hardware devices and a computer.
[0378] The embodiments described above are merely illustrative of the principles of the present invention. Modifications and variations of the configurations and details described herein will be obvious to those skilled in the art. Accordingly, it is intended that the invention is limited only by the immediate claims and not by any specific details shown in the description and explanation of the embodiments herein.
Claims
1. A device (201) that decodes the previous encoded multichannel signal of the previous frame to obtain three or more previous audio output channels, and decodes the current encoded multichannel signal (107) of the current frame to obtain three or more current audio output channels, The apparatus (201) includes an interface (212), a channel decoder (202), a multi-channel processing unit (204) for generating the three or more current audio output channels, and a noise filling module (220), The interface (212) is adapted to receive the current encoded multichannel signal (107) and to receive side information including a first multichannel parameter (MCH_PAR2), The channel decoder (202) is adapted to decode the current encoded multichannel signal of the current frame to obtain a set of three or more decoded channels (D1, D2, D3) of the current frame. The multi-channel processing unit (204) is adapted to select a first selected pair (D1, D2) of two decoded channels from the set of three or more decoded channels (D1, D2, D3) according to the first multi-channel parameter (MCH_PAR2), The multi-channel processing unit (204) is adapted to generate a first group of two or more processed channels (P1*, P2*) based on the first selected pair of two decoded channels (D1, D2), and to obtain an updated set of three or more decoded channels (D3, P1*, P2*), Before the multi-channel processing unit (204) generates the first group of two or more processed channels (P1*, P2*) based on the first selected pair of two decoded channels (D1, D2), the noise filling module (220) identifies one or more frequency bands in at least one of the two channels of the first selected pair of two decoded channels (D1, D2) where all spectral lines are quantized to zero, and is adapted to generate a mixing channel using two or more of the three or more pre-audio output channels, rather than all of them, and to fill the spectral lines in the one or more frequency bands where all spectral lines are quantized to zero with noise generated using the spectral lines of the mixing channel, and the noise filling module (220) is adapted to select two or more pre-audio output channels to be used to generate the mixing channel from the three or more pre-audio output channels according to the side information. The noise filling module (220) is configured to use the output signal of the previous frame stored after the application of multi-channel coding tool processing, If the pre-output channel signal is unavailable, the noise-filling module (220) is configured to set the pre-channel buffer for the corresponding channel to zero. Device.
2. The noise filling module (220) is adapted to generate the mixing channel using exactly two of the three or more pre-audio output channels as the two or more pre-audio output channels among the three or more pre-audio output channels, The noise filling module (220) is adapted to select exactly two pre-audio output channels from the three or more pre-audio output channels according to the side information. The apparatus (201) according to claim 1.
3. The apparatus (201) according to claim 2, wherein the noise filling module (220) is adapted to select exactly two pre-audio output channels from the three or more pre-audio output channels in accordance with the first multi-channel parameter (MCH_PAR2).
4. The interface (212) is adapted to receive the current encoded multichannel signal (107) and the side information, including the first multichannel parameter (MCH_PAR2) and the second multichannel parameter (MCH_PAR1). The multichannel processing unit (204) is adapted to select a second selected pair of two decoded channels (P1*, D3) from the updated set of three or more decoded channels (D3, P1*, P2*) in accordance with the second multichannel parameter (MCH_PAR1), wherein at least one channel (P1*) of the second selected pair of two decoded channels (P1*, D3) is one channel of the first group of two or more processed channels (P1*, P2*), The multi-channel processing unit (204) is adapted to generate a second group of two or more processed channels (P3*, P4*) based on the second selected pair of two decoded channels (P1, D3), and to further update the updated set of three or more decoded channels. The apparatus (201) according to claim 2.
5. The multi-channel processing unit (204) is adapted to generate two or more first groups of processed channels (P1*, P2*) by generating a first group of exactly two processed channels (P1*, P2*) based on the first selected pair of two decoded channels (D1, D2), The multi-channel processing unit (204) is adapted to replace the first selected pair of two decoded channels (D1, D2) in the set of three or more decoded channels (D1, D2, D3) with the first group of exactly two processed channels (P1*, P2*), thereby obtaining the updated set of three or more decoded channels (D3, P1*, P2*), The multi-channel processing unit (204) is adapted to generate two or more second groups of processed channels (P3*, P4*) by generating the second group of exactly two processed channels (P3*, P4*) based on the second selected pair of two decoded channels (P1*, D3), The multi-channel processing unit (204) is adapted to replace the second selected pair of two decoded channels (P1*, D3) in the updated set of three or more decoded channels (D3, P1*, P2*) with the second group of exactly two processed channels (P3*, P4*), and to further update the updated set of three or more decoded channels. The apparatus (201) according to claim 4.
6. The first multi-channel parameter (MCH_PAR2) indicates two decoded channels (D1, D2) from the set of three or more decoded channels. The multichannel processing unit (204) is adapted to select the first selected pair of two decoded channels (D1, D2) from the set of three or more decoded channels (D1, D2, D3) by selecting the two decoded channels (D1, D2) indicated by the first multichannel parameter (MCH_PAR2), The second multi-channel parameter (MCH_PAR1) indicates two decoded channels (P1*, D3) from the updated set of three or more decoded channels. The multichannel processing unit (204) is adapted to select the second selected pair of the two decoded channels (P1*, D3) from the updated set of three or more decoded channels (D3, P1*, P2*) by selecting the two decoded channels (P1*, D3) indicated by the second multichannel parameter (MCH_PAR1). The apparatus (201) according to claim 5.
7. The device (201) is adapted to assign an identification unit from the set of identification units to each of the three or more pre-audio output channels, so that each of the three or more pre-audio output channels is assigned to exactly one identification unit from the set of identification units, and each identification unit from the set of identification units is assigned to exactly one of the three or more pre-audio output channels. The device (201) is adapted to assign an identification unit from the set of identification units to each channel of the set of three or more decoded channels (D1, D2, D3), so that each channel of the set of three or more decoded channels is assigned to exactly one identification unit from the set of identification units, and each identification unit of the set of identification units is assigned to exactly one channel of the set of three or more decoded channels (D1, D2, D3), The first multi-channel parameter (MCH_PAR2) indicates a first pair of two identification units in the set of three or more identification units, The multi-channel processing unit (204) is adapted to select the first selected pair of the two decoded channels (D1, D2) from the set of three or more decoded channels (D1, D2, D3) by selecting two decoded channels (D1, D2) to be assigned to the two identification units of the first pair of the two identification units. The apparatus (201) is adapted to assign the first identification unit of the first pair of two identification units to the first processed channel of the first group of two processed channels (P1*, P2*), The apparatus (201) is adapted to precisely assign the second of the two identification units of the first pair of two identification units to the second processed channel of the first group of two processed channels (P1*, P2*). The apparatus (201) according to claim 6.
8. The second multichannel parameter (MCH_PAR1) indicates a second pair of two identification units in the set of three or more identification units, The multi-channel processing unit (204) is adapted to select the second selected pair of the two decoded channels (P1*, D3) from the updated set of three or more decoded channels (D3, P1*, P2*) by selecting the two decoded channels (D3, P1*, P2*) to be assigned to the two identification units of the second pair of the two identification units. The apparatus (201) is adapted to assign the first of the two identification units of the second pair of two identification units to the first processed channel of the second group of two processed channels (P3*, P4*), The apparatus (201) is adapted to precisely assign the second identification unit of the second pair of identification units of the two identification units to the second processed channel of the second group of the two processed channels (P3*, P4*). The apparatus (201) according to claim 7.
9. The first multi-channel parameter (MCH_PAR2) indicates the first pair of two identification units in the set of three or more identification units, The apparatus (201) according to claim 7, wherein the noise filling module (220) is adapted to select the two pre-audio output channels from the three or more pre-audio output channels by selecting the two pre-audio output channels assigned to the two pre-audio output channels of the first pair of the two identification units.
10. Before the multi-channel processing unit (204) generates the first group of two or more processed channels (P1*, P2*) based on the first selected pair (D1, D2) of the two decoded channels, the noise filling module (220) identifies one or more scale factor bands for at least one of the two channels of the first selected pair (D1, D2) of the two decoded channels, which are one or more frequency bands where all spectral lines are quantized to zero, and generates the mixing channel using two or more pre-audio output channels, rather than all of the three or more pre-audio output channels, and adapts to filling the spectral lines of the one or more frequency bands where all spectral lines are quantized to zero with the noise generated using the spectral lines of the mixing channel, depending on the scale factor of each of the one or more scale factor bands where all spectral lines are quantized to zero. The apparatus (201) according to claim 1.
11. The interface (212) is configured to receive each of the scale factors of the one or more scale factor bandwidths, Each of the scale factor in the one or more scale factor bands represents the energy of the spectral line in the scale factor band before quantization. The noise-filling module (220) is adapted to generate the noise for each of the one or more scale factor bands in which all spectral lines are quantized to zero, and as a result the energy of the spectral lines corresponds to the energy indicated by the scale factor of the scale factor band after the noise has been added to one of the frequency bands. The apparatus (201) according to claim 10.
12. A method for decoding the previous encoded multichannel signal of the previous frame to obtain three or more previous audio output channels, and decoding the current encoded multichannel signal (107) of the current frame to obtain three or more current audio output channels, wherein the method is The current encoded multichannel signal (107) is received, and side information including the first multichannel parameter (MCH_PAR2) is received, Decode the current encoded multichannel signal of the current frame to obtain a set of three or more decoded channels (D1, D2, D3) of the current frame, In accordance with the first multi-channel parameter (MCH_PAR2), a first selected pair (D1, D2) of two decoded channels is selected from the set of three or more decoded channels (D1, D2, D3), Based on the first selected pair of two decoded channels (D1, D2), a first group of two or more processed channels (P1*, P2*) is generated, and an updated set of three or more decoded channels (D3, P1*, P2*) is obtained. Includes, Before the first selected pair of two decoded channels (D1, D2) is generated, For at least one of the two channels of the first selected pair of two decoded channels (D1, D2), identify one or more frequency bands in which all spectral lines are quantized to zero; generate a mixing channel using two or more of the three or more pre-audio output channels, rather than all of them; fill the spectral lines of the one or more frequency bands in which all spectral lines are quantized to zero with noise generated using the spectral lines of the mixing channel; and select the two or more pre-audio output channels from the three or more pre-audio output channels to be used to generate the mixing channel, depending on the side information. The method includes using the output signal of the previous frame stored after the application of multi-channel coding tool processing, The method includes setting the pre-channel buffer of the corresponding channel to zero when the pre-output channel signal is unavailable. method.
13. A computer program for carrying out the method described in claim 12, when executed on a computer or signal processing unit.
Citation Information
Patent Citations
Noise filling in multichannel audio coding
WO2015011061A1
Methods and devices for joint multichannel coding
WO2015036351A1