Apparatus and method for MDCT M / S stereo with comprehensive ILD having improved mid / side determination

The MDCT M/S stereo with integrated ILD processing addresses inefficiencies in existing encoders by using normalization and frequency domain noise shaping to optimize encoding decisions, enhancing efficiency and reducing computational costs for panned signals.

JP7704802B2Active Publication Date: 2025-07-08FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023078313
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-11-21
Filing Date
2023-05-11
Publication Date
2025-07-08
Estimated Expiration
2037-01-20

AI Technical Summary

Technical Problem

Existing MDCT-based audio encoders struggle with efficient stereo processing, particularly for panned signals, due to high computational costs and inefficient bit allocation in M/S encoding, which is not effective when inter-level differences (ILD) are present.

Method used

An apparatus and method for MDCT M/S stereo with integrated ILD processing that employs normalization, frequency domain noise shaping, and a single ILD parameter to determine optimal M/S or L/R encoding based on perceptually whitened signals, reducing bit requirements and simplifying the encoding process.

Benefits of technology

This approach efficiently handles panned signals with minimal side information, simplifies the encoding structure, and achieves effective bit reduction by optimizing bit allocation and encoding decisions based on perceptual whitening, thus improving encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704802000042
    Figure 0007704802000042
  • Figure 0007704802000043
    Figure 0007704802000043
  • Figure 0007704802000044
    Figure 0007704802000044
Patent Text Reader

Abstract

To provide an apparatus, system, method and program for handling panned signals with minimal side information.SOLUTION: The apparatus is provided for encoding an audio input signal of first and second channels including two or more channels to obtain an encoded audio signal, including a normalizer for determining a normalization value for the audio input signal depending on the audio input signal of the first channel and the audio input signal of the second channel, and an encoding unit for encoding a normalized audio signal to obtain the normalized audio signal. The normalizer modulates at least one of the audio signals of the first and second channels depending on the normalization value to determine the first and second channels for the normalized audio signal.SELECTED DRAWING: Figure 1a
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to audio signal encoding and decoding, and more particularly, to an apparatus and method for MDCT M / S stereo with an integrated ILD having improved mid / side determination.

Background Art

[0002] Band-wise M / S (M / S = mid / side) processing in an MDCT-based encoder (MDCT = modified discrete cosine transform) is a known and effective method for stereo processing. However, it is still not sufficient for panned signals, and additional processing such as complex prediction or encoding of the angle between the mid channel and the side channel is required.

[0003] [1], [2], [3] and [4] describe M / S processing in windowed, transformed and unnormalized (not whitened) signals.

[0004] [7] describes prediction between the mid channel and the side channel. [7] discloses an encoder for encoding an audio signal based on the combination of two audio channels. The audio encoder obtains a combined signal that is the mid signal, and further obtains a prediction residual signal that is the predicted side signal derived from the mid signal. The first combined signal and the prediction residual signal are encoded and recorded in a data stream together with prediction information. Further, [7] discloses a decoder that generates a decoded first audio channel and a second audio channel using the prediction residual signal, the first combined signal, and the prediction information.

[0005] In [5], an application of M / S stereo that is coupled after being separately normalized for each band is described. In particular, [5] relates to the Opus codec. Opus encodes the mid signal and the side signal as the normalized signals m = M / ||M|| and s = S / ||S||. To reproduce M and S from m and s, the angle θ s = arctan(||S|| / ||M||) is encoded. Depending on N, the size of the band, and a, the total number of bits available for m and s, the optimal allocation for m is a mid = (a - (N - 1)log2tanθ s ).

[0006] In known approaches (e.g., [2] and [4]), complex rate / distortion loops are combined by a decision that the band channels should be transformed (e.g., using M / S as followed from M to S by prediction residual calculation in [7]) to reduce the inter-channel correlation. This complex structure has high computer processing costs. Separating the perceptual model from the rate loop (as in [6a], [6b], and

[13] ) considerably simplifies the system.

[0007] Also, the encoding of the prediction coefficients or angles for individual bands requires a large number of bits (as in [5] and [7] for example).

[0008] In [1], [3], and [5], only a single decision is made across the entire spectrum to determine whether the entire spectrum is M / S encoded or L / R encoded.

[0009] If an ILD (inter-level difference) exists, i.e., if the channels are panned, M / S encoding is not efficient.

[0010] As outlined above, in an MDCT-based encoder, M / S processing with respect to bands is known to be an effective method for stereo processing. The M / S processing coding gain varies from 0% for uncorrelated channels to 50% for monaural or a phase difference of π / 2 between channels. For stereo non-masking and inverse non-masking (see [1]), it is important to have a robust M / S decision.

[0011] In [2] (for each band where the masking threshold between left and right changes by less than 2 dB), M / S coding is selected as the coding method.

[0012] In [1], the M / S decision is based on the estimated bit consumption for M / S coding and L / R coding (L / R = left / right) of the channels. The bitrate requirements for M / S coding and L / R coding are estimated from the spectrum and the masking threshold using the perceptual entropy (PE). The masking threshold is calculated for the left and right channels. The masking thresholds for the mid and side channels are assumed to be the minimum of the left and right thresholds.

[0013] Furthermore, [1] describes how the coding thresholds for the individual channels to be coded are derived. In particular, the coding thresholds for the left and right channels are calculated by individual perceptual models for these channels. In [1], the coding thresholds for the M channel and the S channel are equally selected and derived as the minimum of the left coding threshold and the right coding threshold.

[0014] Furthermore, [1] explains how to decide between L / R coding and M / S coding so that good coding performance is achieved. In particular, the perceptual entropy is estimated for L / R coding and M / S coding using the thresholds.

[0015] [3] and [4], in [1] and [2], the M / S process is performed on the window-displayed, transformed, and unnormalized (not whitened) signal, and the M / S decision is based on the masking threshold and the perceptual entropy estimate.

[0016] [5], the energies of the left and right channels are explicitly encoded, and the encoded angles preserve the energies of different signals. Even if L / R encoding is more efficient, it is assumed in [5] that M / S encoding is safe. According to [5], L / R encoding is only chosen when the correlation between channels is not strong enough.

[0017] Furthermore, the encoding of the prediction coefficients or angles of individual bands requires a large number of bits (see, for example, [5] and [7]). SUMMARY OF THE INVENTION PROBLEM TO BE SOLVED BY THE INVENTION

[0018] Therefore, if an improved concept for audio encoding and audio decoding is provided, it would be highly appreciated.

[0019] Therefore, an object of the present invention is to provide an improved concept for audio signal encoding, audio signal processing, and audio signal decoding. The object of the present invention is solved by an audio decoder according to claim 1, and an apparatus according to claim 23, and a method according to claim 37, and a method according to claim 38, and a computer program according to claim 39. MEANS FOR SOLVING THE PROBLEM

[0020] According to an embodiment, an apparatus for encoding a first channel and a second channel of an audio input signal including two or more channels to obtain an encoded audio signal is provided.

[0021] The apparatus for symbolization includes a normalizer configured to determine a normalization value for an audio input signal depending on a first channel of the audio input signal and depending on a second channel of the audio input signal. The normalizer is configured to determine a first channel and a second channel of a normalized audio signal by modulating at least one of the first channel and the second channel of the audio input signal depending on the normalization value.

[0022] Furthermore, the apparatus for symbolization includes an encoding unit configured to generate a processed audio signal having a first channel and a second channel such that one or more spectral bands of the first channel of the processed audio signal are the same as one or more spectral bands of the first channel of the normalized audio signal, and one or more spectral bands of the second channel of the processed audio signal are the same as one or more spectral bands of the second channel of the normalized audio signal, and at least one spectral band of the first channel of the processed audio signal is a spectral band of a mid signal depending on a spectral band of the first channel of the normalized audio signal and depending on a spectral band of the second channel of the normalized audio signal, and at least one spectral band of the second channel of the processed audio signal is a spectral band of a side signal depending on a spectral band of the first channel of the normalized audio signal and depending on a spectral band of the second channel of the normalized audio signal. The encoding unit is configured to encode the processed audio signal to obtain an encoded audio signal.

[0023] Furthermore, an apparatus for decoding an encoded audio signal including a first channel and a second channel is provided to obtain a first channel and a second channel of a decoded audio signal including two or more channels.

[0024] The apparatus for decoding includes a decoding unit configured to determine, for each of the individual spectral bands of a plurality of spectral bands, whether the spectral band of the first channel of the encoded audio signal and the spectral band of the second channel of the encoded audio signal are encoded using dual-mono encoding or mid-side encoding.

[0025] When dual-mono encoding is used, the decoding unit is configured to use the spectral band of the first channel of the encoded audio signal as the spectral band of the first channel of the intermediate audio signal and to use the spectral band of the second channel of the encoded audio signal as the spectral band of the second channel of the intermediate audio signal.

[0026] Furthermore, when mid-side encoding is used, the decoding unit is configured to generate the spectral band of the first channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal, and to generate the spectral band of the second channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal.

[0027] Furthermore, the apparatus for decoding including a denormalizer is configured to modulate at least one of the first channel and the second channel of the intermediate audio signal depending on a denormalization value to obtain the first channel and the second channel of the decoded audio signal.

[0028] Furthermore, a method for encoding the first channel and the second channel of an audio input signal including two or more channels to obtain an encoded audio signal is provided. The method includes the following. -Determining a normalization value for an audio input signal that depends on a first channel of the audio input signal and also depends on a second channel of the audio input signal. -Determining first and second channels of a normalized audio signal by modulating at least one of the first and second channels of the audio input signal depending on the normalization value. -Generating a processed audio signal having first and second channels such that one or more spectral bands of the first channel of the processed audio signal are the same as one or more spectral bands of the first channel of the normalized audio signal, and one or more spectral bands of the second channel of the processed audio signal are the same as one or more spectral bands of the second channel of the normalized audio signal, and at least one spectral band of the first channel of the processed audio signal is a spectral band of a mid signal that depends on a spectral band of the first channel of the normalized audio signal and also depends on a spectral band of the second channel of the normalized audio signal, and at least one spectral band of the second channel of the processed audio signal is a spectral band of a side signal that depends on a spectral band of the first channel of the normalized audio signal and also depends on a spectral band of the second channel of the normalized audio signal, and encoding the processed audio signal to obtain an encoded audio signal.

[0029] Furthermore, a method for decoding an encoded audio signal including first and second channels to obtain first and second channels of a decoded audio signal including two or more channels is provided. The method includes the following. - Determining, for each individual spectral band of a plurality of spectral bands, whether the spectral band of the first channel of the encoded audio signal and the spectral band of the second channel of the encoded audio signal are encoded using dual-mono encoding or mid-side encoding. - When dual-mono encoding is used, using the spectral band of the first channel of the encoded audio signal as the spectral band of the first channel of the intermediate audio signal and using the spectral band of the second channel of the encoded audio signal as the spectral band of the second channel of the intermediate audio signal. - When mid-side encoding is used, generating the spectral band of the first channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal, and generating the spectral band of the second channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal. And - Modulating at least one of the first and second channels of the intermediate audio signal depending on the unnormalized values to obtain the first and second channels of the decoded audio signal.

[0030] Furthermore, a computer program is provided. Each of the computer programs is configured to execute one of the methods described above when executed in a computer or a signal processor.

Advantages of the Invention

[0031] According to the embodiment, a new concept is provided that can handle panned signals using minimal side information.

[0032] According to some embodiments, FDNS (FDNS = Frequency Domain Noise Shaping) with a rate loop is used as described in [8], combined by spectral envelope distortion and used as described in [6a] and [6b]. In some embodiments, a single ILD parameter of the FDNS-whitened spectrum is followed and used by a determination regarding the band as to whether M / S coding or L / R coding is used for encoding. In some embodiments, the M / S determination is based on an estimated bit reduction. In some embodiments, the bit rate distribution between the M / S processing channels regarding the band is, for example, energy dependent.

[0033] Some embodiments are followed by M / S processing regarding the band having an efficient M / S determination mechanism and a rate loop that controls the only comprehensive gain, providing the combination of a single comprehensive ILD applied to the whitened spectrum.

[0034] Some embodiments particularly employ FDNS with a rate loop based on [6a] or [6b], combined with spectral envelope distortion based on, for example, [8]. These embodiments provide an efficient and highly effective method for separating quantization noise and perceptual shaping of the rate loop. When there are the advantages of M / S processing as described above, using a single ILD parameter of the FDNS-whitened spectrum allows for a simple and effective way of determination. Whitening the spectrum and removing the ILD allows for efficient M / S processing. Encoding a single comprehensive ILD for the described system is sufficient, and thus bit reduction is achieved in contrast to known approaches.

[0035] According to an embodiment, the M / S processing is made based on a perceptually whitened signal. The embodiment determines an encoding threshold and optimally determines a decision as to whether L / R coding or M / S coding is employed when processing a perceptually whitened and ILD-corrected signal.

[0036] Furthermore, according to the embodiment, a new bitrate estimation is provided.

[0037] Compared with [1] - [5], in the embodiment, the perceptual model is separated from the rate loops in [6a], [6b] and

[13] .

[0038] Even if the M / S decision is based on the estimated bitrate as proposed in [1], compared with [1], the difference in bitrate requirements between M / S encoding and L / R encoding does not depend on the masking threshold determined by the perceptual model. Instead, the bitrate requirement is determined by the lossless entropy coder being used. That is, instead of deriving the bitrate requirement from the entropy of the perceptual original signal, the bitrate requirement is derived from the entropy of the perceptually whitened signal.

[0039] Compared with [1] - [5], in the embodiment, the M / S decision is made based on the perceptually whitened signal, and a good estimate of the required bitrate is obtained. For this purpose, the arithmetic coder bit consumption estimation is applied as described in [6a] or [6b]. The masking threshold does not need to be explicitly considered.

[0040] In [1], the masking thresholds for the mid - channel and side - channel are assumed to be the minimum of the left and right masking thresholds. Spectral noise shaping is performed in the mid - channel and side - channel, for example, based on these masking thresholds.

[0041] According to the embodiment, spectral noise shaping can be performed, for example, on the left and right channels, and the perceptual envelope is accurately applied where it is estimated in such an embodiment.

[0042] Furthermore, the embodiments are based on the discovery that M / S coding is not efficient when an ILD is present, i.e., when the channel is panned. To avoid this, the embodiments use a single ILD parameter for a perceptually whitened spectrum.

[0043] According to some embodiments, a new concept for M / S determination for processing perceptually whitened signals is provided.

[0044] According to some embodiments, the coder uses a new concept that is not part of a classical audio coder as described, for example, in [1].

[0045] According to some embodiments, perceptually whitened signals are used for another coding, for example, in a similar way as they are used in a speech coder.

[0046] Such an approach has several advantages. For example, the coder structure is simplified. Compact representations of noise shaping characteristics and masking thresholds are achieved, for example, as LPC coefficients. Furthermore, the transform and speech coder structures are integrated, and thus, combined audio / speech coding is possible.

[0047] Some embodiments employ a comprehensive ILD parameter to efficiently code panned sources.

[0048] In an embodiment, the coder employs frequency domain noise shaping (FDNS) to perceptually whiten a signal with a rate loop, as described, for example, in [6a] or [6b] combined with the spectral envelope distortion described in [8]. In such an embodiment, the coder further uses a single ILD parameter of the FDNS-whitened spectrum, followed, for example, by an M / S vs L / R determination for the band. The M / S determination for the band is based, for example, on the estimated bitrate of the individual bands when encoded in L / R mode and M / S mode. The mode with at least the necessary bits is selected. The bitrate allocation between the M / S processed channels for the band is energy based.

[0049] Some embodiments apply the M / S determination for the band to the perceptually whitened and ILD corrected spectrum, using the estimated number of bits per band for the entropy coder.

[0050] In some embodiments, for example, FDNS with a rate loop is employed as described in [6a] or [6b] combined with the spectral envelope distortion described in [8]. This provides an efficient and very effective way to separate quantization noise and the perceptual shaping of the rate loop. When there are the advantages of the M / S processing as described, using a single ILD parameter of the FDNS-whitened spectrum allows a simple and effective way of determination. Whitening the spectrum and removing the ILD allows efficient M / S processing.

[0051] Encoding a single comprehensive ILD for the described system is sufficient, and thus bit reduction is achieved as compared to known approaches.

[0052] The embodiments modify the concepts provided in [1] when processing perceptually whitened and ILD-corrected signals. In particular, the embodiments employ equal overall gains for L, R, M, and S that form the encoding thresholds together with the FDNS. The overall gain is derived from SNR estimation or some other concept.

[0053] The M / S decision for the proposed band accurately estimates the number of bits required for encoding each band with an arithmetic coder. This is possible because the M / S decision is performed on the whitened spectrum and is directly followed by quantization, eliminating the need for an experimental search for thresholds.

[0054] In the following, embodiments of the present invention will be described in more detail with reference to the drawings.

Brief Description of the Drawings

[0055]

Figure 1a

Figure 1b

Figure 1c

Figure 1d

Figure 1e

Figure 1f

Figure 2a

Figure 2b

Figure 2c

Figure 2d

Figure 2e

Figure 2f

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

[0056] FIG. 1a illustrates an apparatus for encoding a first channel and a second channel of an audio input signal including two or more channels to obtain an encoded audio signal according to an embodiment.

[0057] The apparatus includes a normalizer 110 configured to determine a normalization value for the audio input signal depending on a first channel of the audio input signal and depending on a second channel of the audio input signal. The normalizer 110 is configured to determine a first channel and a second channel of the normalized audio signal by modulating at least one of the first channel and the second channel of the audio input signal depending on the normalization value.

[0058] For example, in an embodiment, the normalizer 110 is configured to determine a normalization value for the audio input signal depending on a plurality of spectral bands of the first channel and the second channel of the audio input signal. The normalizer 110 is configured to determine a first channel and a second channel of the normalized audio signal by modulating a plurality of spectral bands of at least one of the first channel and the second channel of the audio input signal depending on the normalization value.

[0059] Alternatively, for example, the normalizer 110 is configured to determine a normalization value for the audio input signal depending on the first channel of the audio input signal represented in the time domain and depending on the second channel of the audio input signal represented in the time domain. Further, the normalizer 110 is configured to determine the first and second channels of the normalized audio signal by modulating at least one of the first and second channels of the audio input signal represented in the time domain depending on the normalization value. The apparatus further includes a conversion unit (not shown in FIG. 1a) configured to convert the normalized audio signal from the time domain to the spectral domain such that the normalized audio signal is represented in the spectral domain. The conversion unit is configured to supply the normalized audio signal represented in the spectral domain to the encoding unit 120. For example, the audio input signal is a time-domain residual signal resulting from LPC filtering (LPC = linear predictive coding) of two channels of a time-domain audio signal.

[0060] Furthermore, the apparatus is configured to generate a processed audio signal having a first channel and a second channel such that one or more spectral bands of the first channel of the processed audio signal are the same as one or more spectral bands of the first channel of the normalized audio signal, and one or more spectral bands of the second channel of the processed audio signal are the same as one or more spectral bands of the second channel of the normalized audio signal, and at least one spectral band of the first channel of the processed audio signal depends on the spectral bands of the first channel of the normalized audio signal and the spectral bands of the second channel of the normalized audio signal and is a spectral band of a mid signal, and at least one spectral band of the second channel of the processed audio signal depends on the spectral bands of the first channel of the normalized audio signal and the spectral bands of the second channel of the normalized audio signal and is a spectral band of a side signal. The encoding unit 120 is configured to encode the processed audio signal to obtain an encoded audio signal.

[0061] In an embodiment, the encoding unit 120 is configured to select, for example, from a full-mid-side encoding mode, a full-dual-mono encoding mode, and a band-wise encoding mode, depending on a plurality of spectral bands of the first channel of the normalized audio signal and a plurality of spectral bands of the second channel of the normalized audio signal.

[0062] In such an embodiment, when, for example, the full mid-side encoding mode is selected, the encoding unit 120 is configured to generate a mid signal as the first channel of the mid-side signal from the first and second channels of the normalized audio signal, and to generate a side signal as the second channel of the mid-side signal from the first and second channels of the normalized audio signal, and to encode the mid-side signal to obtain an encoded audio signal.

[0063] According to such an embodiment, when, for example, the full dual-mono encoding mode is selected, the encoding unit 120 is configured to encode the normalized audio signal to obtain an encoded audio signal.

[0064] Furthermore, in such an embodiment, when, for example, an encoding mode related to a band is selected, one or more spectral bands of the first channel of the processed audio signal are the same as one or more spectral bands of the first channel of the normalized audio signal, and one or more spectral bands of the second channel of the processed audio signal are the same as one or more spectral bands of the second channel of the normalized audio signal, and at least one spectral band of the first channel of the processed audio signal depends on the spectral band of the first channel of the normalized audio signal and the spectral band of the second channel of the normalized audio signal to be the spectral band of the mid signal, and at least one spectral band of the second channel of the processed audio signal depends on the spectral band of the first channel of the normalized audio signal and the spectral band of the second channel of the normalized audio signal to be the spectral band of the side signal, and the encoding unit 120 is configured to generate a processed audio signal. The encoding unit 120 is configured to encode the processed audio signal to obtain an encoded audio signal.

[0065] According to an embodiment, the audio input signal is, for example, an audio stereo signal that strictly includes two channels. For example, the first channel of the audio input signal is the left channel of the audio stereo signal, and the second channel of the audio input signal is the right channel of the audio stereo signal.

[0066] In an embodiment, when a coding mode regarding a band is selected, the encoding unit 120 is configured to determine whether mid-side encoding or dual-mono encoding is adopted for each of a plurality of spectral bands of the processed audio signal.

[0067] When mid-side encoding is adopted for the spectral band, the encoding unit 120 is configured to generate, as the spectral band of the mid signal, the spectral band of the first channel of the processed audio signal based on the spectral band of the first channel of the normalized audio signal and based on the spectral band of the second channel of the normalized audio signal. The encoding unit 120 is configured to generate, as the spectral band of the side signal, the spectral band of the second channel of the processed audio signal based on the spectral band of the first channel of the normalized audio signal and based on the spectral band of the second channel of the normalized audio signal.

[0068] When dual - mono coding is adopted for the spectral band, the coding unit 120 is configured to use, for example, the spectral band of the first channel of the normalized audio signal as the spectral band of the first channel of the processed audio signal, and to use the spectral band of the second channel of the normalized audio signal as the spectral band of the second channel of the processed audio signal. Alternatively, the coding unit 120 is configured to use, for example, the spectral band of the second channel of the normalized audio signal as the spectral band of the first channel of the processed audio signal, and to use the spectral band of the first channel of the normalized audio signal as the spectral band of the second channel of the processed audio signal.

[0069] According to an embodiment, the coding unit 120 determines, for example, a first estimate for estimating the number of first bits required for coding by determining a first estimate for estimating the number of first bits required for coding when a full mid - side coding mode is adopted, and a second estimate for estimating the number of second bits required for coding when a full dual - mono coding mode is adopted, and a third estimate for estimating the number of third bits required for coding when a coding mode related to the band is adopted, and selects one of the full mid - side coding mode, the full dual - mono coding mode, and the coding mode related to the band by selecting the coding mode having the smallest number of bits among the first estimate, the second estimate, and the third estimate.

[0070] TIFF0007704802000001.tif79170

[0071] In an embodiment, for example, objective quality means for selecting from among a full mid-side encoding mode, a full dual-mono encoding mode, and an encoding mode related to bandwidth is adopted.

[0072] According to an embodiment, when encoding in the full mid-side encoding mode, for example, the encoding unit 120 determines a first estimate for estimating the number of first bits to be reduced, and when encoding in the full dual-mono encoding mode, determines a second estimate for estimating the number of second bits to be reduced, and when encoding in the encoding mode related to bandwidth, determines a third estimate for estimating the number of third bits to be reduced, and selects the encoding mode having the largest number of bits to be reduced among the first estimate, the second estimate, and the third estimate from among the full mid-side encoding mode, the full dual-mono encoding mode, and the encoding mode related to bandwidth, so as to be configured to select from among the full mid-side encoding mode, the full dual-mono encoding mode, and the encoding mode related to bandwidth.

[0073] In another embodiment, when the full mid-side encoding mode is adopted, for example, the encoding unit 120 estimates a first signal-to-noise ratio that occurs, and when encoding in the full dual-mono encoding mode, estimates a second signal-to-noise ratio that occurs, and when encoding in the encoding mode related to bandwidth, estimates a third signal-to-noise ratio that occurs, and selects the encoding mode having the largest signal-to-noise ratio among the first signal-to-noise ratio, the second signal-to-noise ratio, and the third signal-to-noise ratio from among the full mid-side encoding mode, the full dual-mono encoding mode, and the encoding mode related to bandwidth, so as to be configured to select from among the full mid-side encoding mode, the full dual-mono encoding mode, and the encoding mode related to bandwidth.

[0074] In an embodiment, the normalizer 110 is configured to determine a normalization value for the audio input signal depending on, for example, the energy of the first channel of the audio input signal and depending on the energy of the second channel of the audio input signal.

[0075] According to an embodiment, the audio input signal is represented, for example, in the spectral domain. The normalizer 110 is configured to determine a normalization value for the audio input signal depending on, for example, a plurality of spectral bands of the first channel of the audio input signal and depending on a plurality of spectral bands of the second channel of the audio input signal. Further, the normalizer 110 is configured to determine a normalized audio signal by modulating, for example, a plurality of spectral bands of at least one of the first channel and the second channel of the audio input signal depending on the normalization value.

[0076] TIFF0007704802000002.tif99168

[0077] According to the embodiment illustrated by FIG. 1b, the apparatus for encoding further includes, for example, a conversion unit 102 and a preprocessing unit 105. The conversion unit 102 is configured to convert a time-domain audio signal from the time domain to the frequency domain, for example, to obtain a converted audio signal. The preprocessing unit 105 is configured to generate the first channel and the second channel of the audio input signal by applying, for example, an encoder-side frequency-domain noise shaping operation to the converted audio signal.

[0078] In a particular embodiment, the preprocessing unit 105 is configured to generate the first channel and the second channel of the audio input signal by applying, for example, an encoder-side time-domain noise shaping operation to the converted audio signal before applying an encoder-side frequency-domain noise shaping operation to the converted audio signal.

[0079] FIG. 1c illustrates an apparatus for encoding according to another embodiment further including a conversion unit 115. The normalizer 110 is configured to determine a normalization value for the audio input signal depending on, for example, the first channel of the audio input signal represented in the time domain and depending on the second channel of the audio input signal represented in the time domain. Further, the normalizer 110 is configured to determine the first and second channels of the normalized audio signal by modulating at least one of the first and second channels of the audio input signal represented in the time domain depending on, for example, the normalization value. The conversion unit 115 is configured to convert the normalized audio signal from the time domain to the spectral domain such that, for example, the normalized audio signal is represented in the spectral domain. Further, the conversion unit 115 is configured to supply the normalized audio signal represented in the spectral domain to the encoding unit 120, for example.

[0080] FIG. 1d illustrates an apparatus for encoding according to another embodiment. The apparatus further includes a preprocessing unit 106 configured to receive a time domain audio signal including a first channel and a second channel. The preprocessing unit 106 is configured to apply a filter to, for example, the first channel of the time domain audio signal to create a first perceptually whitened spectrum in order to obtain the first channel of the audio input signal represented in the time domain. Further, the preprocessing unit 106 is configured to apply a filter to, for example, the second channel of the time domain audio signal to create a second perceptually whitened spectrum in order to obtain the second channel of the audio input signal represented in the time domain.

[0081] In the embodiment described with reference to FIG. 1e, the conversion unit 115 is configured to convert the normalized audio signal from the time domain to the spectral domain, for example, to obtain the converted audio signal. In the embodiment of FIG. 1e, the apparatus further includes a spectral domain pre-processor 118 configured to perform encoder-side temporal noise shaping on the converted audio signal to obtain the normalized audio signal represented in the spectral domain.

[0082] According to an embodiment, the encoding unit 120 is configured to obtain the encoded audio signal, for example, by applying encoder-side stereo intelligent gap filling to the normalized audio signal or the processed audio signal.

[0083] In another embodiment described with reference to FIG. 1f, a system for encoding four channels of an audio input signal including four or more channels is provided to obtain the encoded audio signal. The system includes a first apparatus 170 according to one of the embodiments described above for encoding the first channel and the second channel of the four or more channels of the audio input signal to obtain the first channel and the second channel of the encoded audio signal. Further, the system includes a second apparatus 180 according to one of the embodiments described above for encoding the third channel and the fourth channel of the four or more channels of the audio input signal to obtain the third channel and the fourth channel of the encoded audio signal.

[0084] FIG. 2a illustrates an apparatus for decoding an encoded audio signal including a first channel and a second channel to obtain a decoded audio signal according to an embodiment.

[0085] The apparatus for decoding includes a decoding unit 210 configured to determine, for each of the individual spectral bands of a plurality of spectral bands, whether the spectral band of the first channel of the encoded audio signal and the spectral band of the second channel of the encoded audio signal are encoded using dual-mono encoding or mid-side encoding.

[0086] When dual-mono encoding is used, the decoding unit 210 is configured to use the spectral band of the first channel of the encoded audio signal as the spectral band of the first channel of the intermediate audio signal, and to use the spectral band of the second channel of the encoded audio signal as the spectral band of the second channel of the intermediate audio signal.

[0087] Furthermore, when mid-side encoding is used, the decoding unit 210 is configured to generate the spectral band of the first channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal, and to generate the spectral band of the second channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal.

[0088] Furthermore, the apparatus for decoding includes a denormalizer 220 configured to modulate at least one of the first channel and the second channel of the intermediate audio signal depending on a denormalized value to obtain the first channel and the second channel of the decoded audio signal.

[0089] In an embodiment, the decoding unit 210 is configured to determine, for example, whether the encoded audio signal is encoded in a full mid-side encoding mode, a full dual-mono encoding mode, or an encoding mode related to a band.

[0090] Furthermore, in such an embodiment, the decoding unit 210 is configured to generate, for example, a first channel of the intermediate audio signal from the first and second channels of the encoded audio signal, and generate a second channel of the intermediate audio signal from the first and second channels of the encoded audio signal when it is determined that the encoded audio signal is encoded in the full mid-side encoding mode.

[0091] According to such an embodiment, the decoding unit 210 is configured to use, for example, the first channel of the encoded audio signal as the first channel of the intermediate audio signal and use the second channel of the encoded audio signal as the second channel of the intermediate audio signal when it is determined that the encoded audio signal is encoded in the full dual-mono encoding mode.

[0092] Furthermore, in such an embodiment, the decoding unit 210 is configured to, for example, when it is determined that the encoded audio signal is encoded in an encoding mode related to a band, - for each of the individual spectral bands of a plurality of spectral bands, determine whether the spectral band of the first channel of the encoded audio signal and the spectral band of the second channel of the encoded audio signal are encoded using a dual-mono encoding or a mid-side encoding mode. -When dual-mono encoding is used, the spectral band of the first channel of the intermediate audio signal is the spectral band of the first channel of the encoded audio signal, and the spectral band of the second channel of the intermediate audio signal is the spectral band of the second channel of the encoded audio signal, and it is configured to use them, -When mid-side encoding is used, the spectral band of the first channel of the intermediate audio signal is generated based on the spectral band of the first channel of the encoded audio signal and the spectral band of the second channel of the encoded audio signal, and the spectral band of the second channel of the intermediate audio signal is generated based on the spectral band of the first channel of the encoded audio signal and the spectral band of the second channel of the encoded audio signal, and it is configured to do so.

[0093] For example, in the full mid-side encoding mode, the following equations are applied to obtain the first channel L of the intermediate audio signal and the second channel R of the intermediate audio signal from M which is the first channel of the encoded audio signal and S which is the second channel of the encoded audio signal. L=(M + S) / sqrt(2) R=(M - S) / sqrt(2)

[0094] According to an embodiment, the decoded audio signal is, for example, an audio stereo signal that strictly includes two channels. For example, the first channel of the decoded audio signal is the left channel of the audio stereo signal, and the second channel of the decoded audio signal is the right channel of the audio stereo signal.

[0095] According to an embodiment, the denormalizer 220 is configured to modulate at least one of a plurality of spectral bands of the first channel and the second channel of the intermediate audio signal depending on a denormalization value, for example, to obtain the first channel and the second channel of the decoded audio signal.

[0096] In another embodiment shown in FIG. 2b, the denormalizer 220 is configured to modulate at least one of a plurality of spectral bands of the first channel and the second channel of the intermediate audio signal depending on a denormalization value, for example, to obtain the denormalized audio signal. In such an embodiment, the apparatus further includes, for example, a post-processing unit 230 and a conversion unit 235. The post-processing unit 230 is configured to perform at least one of decoder-side time-domain noise shaping and decoder-side frequency-domain noise shaping on the denormalized audio signal, for example, to obtain a post-processed audio signal. The conversion unit (235) is configured to convert the post-processed audio signal from the spectral domain to the time domain, for example, to obtain the first channel and the second channel of the decoded audio signal.

[0097] According to the embodiment illustrated by FIG. 2c, the apparatus further includes a conversion unit 215 configured to convert the intermediate audio signal from the spectral domain to the time domain. The denormalizer 220 is configured to modulate at least one of the first channel and the second channel of the intermediate audio signal represented in the time domain depending on a denormalization value, for example, to obtain the first channel and the second channel of the decoded audio signal.

[0098] In a similar embodiment described by FIG. 2d, the conversion unit 215 is configured to convert, for example, an intermediate audio signal from the spectral domain to the time domain. The denormalizer 220 is configured to modulate, for example, at least one of the first channel and the second channel of the intermediate audio signal represented in the time domain depending on the denormalization value to obtain a denormalized audio signal. The apparatus further includes a post-processing unit 235 configured to process the denormalized audio signal, which is a perceptually whitened audio signal, for example, to obtain the first channel and the second channel of the decoded audio signal.

[0099] According to another embodiment described by FIG. 2e, the apparatus further includes a spectral domain post-processor 212 configured to perform decoder-side temporal noise shaping on the intermediate audio signal. In such an embodiment, the conversion unit 215 is configured to convert the intermediate audio signal from the spectral domain to the time domain after decoder-side temporal noise shaping is performed on the intermediate audio signal.

[0100] In another embodiment, the decoding unit 210 is configured to apply, for example, decoder-side stereo intelligent gap filling to the encoded audio signal.

[0101] Furthermore, as described in FIG. 2f, a system for decoding an encoded audio signal including four or more channels to obtain four channels of the decoded audio signal is provided. The system includes a first device 270 for decoding the first and second channels of the four or more channels of the encoded audio signal to obtain the first and second channels of the decoded audio signal, according to one of the embodiments described above. Further, the system includes a second device 280 for decoding the third and fourth channels of the four or more channels of the encoded audio signal to obtain the third and fourth channels of the decoded audio signal, according to one of the embodiments described above.

[0102] FIG. 3 illustrates a system for generating an encoded audio signal from an audio input signal and generating a decoded audio signal from the encoded audio signal, according to an embodiment.

[0103] The system includes an encoding device 310 according to one of the embodiments described above. The encoding device 310 is configured to generate an encoded audio signal from the audio input signal.

[0104] Furthermore, the system includes a decoding device 320 as described above. The decoding device 320 is configured to generate a decoded audio signal from the encoded audio signal.

[0105] Similarly, a system is provided for generating an encoded audio signal from an audio input signal and for generating a decoded audio signal from the encoded audio signal. The system includes the system described in the embodiment of FIG. 1f (where the system described in the embodiment of FIG. 1f is configured to generate an encoded audio signal from an audio input signal) and the system described in the embodiment of FIG. 2f (where the system described in the embodiment of FIG. 2f is configured to generate a decoded audio signal from the encoded audio signal).

[0106] Preferred embodiments are described below.

[0107] FIG. 4 illustrates an apparatus for encoding according to another embodiment. In particular, a preprocessing unit 105 and a conversion unit 102 according to a specific embodiment are described. The conversion unit 102 is specifically configured to perform the conversion of the audio input signal from the time domain to the spectral domain. The conversion unit is configured to perform encoder-side time noise shaping and encoder-side frequency domain noise shaping on the audio input signal.

[0108] Furthermore, FIG. 5 illustrates a stereo processing module in an apparatus for encoding according to an embodiment. FIG. 5 illustrates a normalizer 110 and an encoding unit 120.

[0109] Furthermore, FIG. 6 illustrates an apparatus for decoding according to another embodiment. In particular, FIG. 6 illustrates a postprocessing unit 230 according to a specific embodiment. The postprocessing unit 230 is specifically configured to obtain the processed audio signal from the denormalizer 220. The postprocessing unit 230 is configured to perform at least one of decoder-side time noise shaping and decoder-side frequency domain noise shaping on the processed audio signal.

[0110] The time-domain transient detector (TD TD) and windowing (windowing) and MDCT and MDST and OLA are performed as described, for example, in [6a] or [6b]. MDCT and MDST form a modulated complex lapped transform (MCLT). Performing MDCT and MDST separately is equivalent to performing MCLT. "From MCLT to MDCT" exactly means taking the MDCT part of MCLT and discarding MDST (see

[12] ).

[0111] Selecting different window lengths in the left and right channels, for example, enforces dual-mono coding within that frame.

[0112] Time noise shaping (TNS) is performed as described, for example, in [6a] or [6b].

[0113] Frequency-domain noise shaping (FDNS) and the calculation of FDNS parameters are similar to the procedure described, for example, in [8]. One difference is that, for example, the FDNS parameters for frames where TNS is inactive are calculated from the MCLT spectrum. In frames where TNS is active, MDST is estimated from MDCT, for example.

[0114] FDNS is also replaced with a perceptual spectrum that whitens in the time domain (as described, for example, in

[13] ).

[0115] Stereo processing includes comprehensive ILD processing and M / S processing related to the band and bitrate distribution between channels.

[0116] TIFF0007704802000003.tif138170

[0117] TIFF0007704802000004.tif49170

[0118] When a perceptual spectrum whitened in the time domain is used (e.g., as described in

[13] ), a single comprehensive ILD is calculated and applied in the time domain before the conversion from the time domain to the frequency domain (i.e., before the MDCT). Alternatively, instead, the whitened perceptual spectrum is followed by a conversion from the time domain to the frequency domain followed by a single comprehensive ILD in the frequency domain. Alternatively, instead, a single comprehensive ILD is calculated in the time domain before the conversion from the time domain to the frequency domain and applied in the frequency domain after the conversion from the time domain to the frequency domain.

[0119] TIFF0007704802000005.tif33169

[0120] Comprehensive gain G est is estimated in a signal including the concatenated left and right channels. Thus, it is different from [6b] and [6a]. For example, the first estimation of the gain described in Section 5.3.3.2.8.1.1 "Comprehensive Gain Estimator" of [6b] or [6a] is used from scalar quantization, assuming a SNR gain of 6 dB per bit per sample.

[0121] The estimated gain is multiplied by a constant to obtain an under - estimation or over - estimation in the final gain G est The signals in the left channel, right channel, mid - channel or side - channel are then quantized with a quantization step size of 1 / G est which is G est used.

[0122] The quantized signal is then encoded using an arithmetic coder, a Huffman coder or another entropy coder in order to obtain the required number of bits. For example, the context based on the arithmetic coder described in Sections 5.3.3.2.8.1.3 to 5.3.3.2.8.1.7 of [6b] or [6a] is used. Since the rate loop (e.g., Section 5.3.3.2.8.1.2 of [6b] or [6a]) is executed after stereo encoding, the estimation of the required bits is sufficient.

[0123] As an example, for each quantized channel, the number of bits required for the context based on arithmetic coding is estimated as described in Sections 5.3.3.2.8.1.3 to 5.3.3.2.8.1.7 of [6b] or [6a].

[0124] According to an embodiment, the bit estimation for each individual quantized channel (left channel, right channel, mid channel or side channel) is determined based on the code of the following example. int context_based_arihmetic_coder_estimate( int spectrum[], int start_line, int end_line, int lastnz, / / lastnz=last non-zero spectrum line int&ctx, / / ctx=context int&probability, / / 14 bit fixed point probability const unsigned int cum_freq[N_CONTEXTS][] / / cum_freq=cumulative frequency tables,14 bit fixed point ) int nBits=0; ​for(int k = start_line; k < min(lastnz, end_line); k += 2) int a1 = abs(spectrum[k]); int b1 = abs(spectrum[k + 1]); / *Signs Bits* / nBits += min(a1, 1); nBits += min(b1, 1); while(max(a1, b1) >= 4) probability *= cum_freq[ctx][VAL_ESC]; int nlz = Number_of_leading_zeros(probability); nBits += 2 + nlz; probability >>= 14 - nlz; a1 >>= 1; b1 >>= 1; ctx = update_context(ctx, VAL_ESC); int symbol = a1 + 4 * b1; probability *= (cum_freq[ctx][symbol] - cum_freq[ctx][symbol + 1]); int nlz = Number_of_leading_zeros(probability); nBits += nlz; hContextMem->proba >>= 14 - nlz; ctx = update_context(ctx, a1 + b1); return nBits; ​​​​​ Here, spectrum is set to indicate the quantized spectrum to be coded. start_line is set to 0. end_line is set to the length of the spectrum. lastnz is set to the index of the last non-zero element of the spectrum. ctx is set to 0. The probability is set to 1 in 14-bit fixed-point notation (16384 = 1 << 14).

[0125] As outlined, the code of the above example is used to obtain bit estimates for at least one of, for example, the left channel, the right channel, the mid channel or the side channel.

[0126] Some embodiments use an arithmetic coder as described in [6b] and [6a]. Further details can be found, for example, in chapter 5.3.3.2.8 "Arithmetic coder" of [6b].

[0127] The number of bits estimated for "full dual - mono" (b LR ) is equal to the sum of the bits required for the right and left channels.

[0128] The number of bits estimated for "full M / S" (b MS ) is equal to the sum of the bits required for the mid and side channels.

[0129] TIFF0007704802000006.tif42170

[0130] TIFF0007704802000007.tif42170

[0131] TIFF0007704802000008.tif53170

[0132] TIFF0007704802000009.tif53170

[0133] The "M / S related to band" mode requires additional nBands bits for signaling in each individual band, regardless of whether L / R or M / S encoding is used. The selection between "M / S related to band" and "Full dual-mono" and "Full M / S" is encoded as a stereo mode, for example, in the bitstream. And for signaling, "Full dual-mono" and "Full M / S" do not require additional bits compared to "M / S related to band".

[0134] TIFF0007704802000010.tif66170

[0135] TIFF0007704802000011.tif42169

[0136] TIFF0007704802000012.tif42168

[0137] In some embodiments, for example, first the gain G is estimated and the quantization step size is estimated. For this, it is expected that there are sufficient bits to encode the L / R channels.

[0138] TIFF0007704802000013.tif23170

[0139] As already outlined, according to certain embodiments, for each individual quantized channel, the number of bits required for arithmetic coding is estimated, for example, as described in section 5.3.3.2.8.1.7 "Bit consumption estimation" of [6b], or in a similar section of [6a].

[0140] TIFF0007704802000014.tif37168

[0141] 4 contexts (ctx L , ctx R , ctx M , ctx M ) and 4 probabilities (p L , pR , p M , p M ) is initialized and then repeatedly updated.

[0142] At the beginning of the estimation (for i = 0), for each context (ctx L , ctx R , ctx M , ctx M ) is set to 0, and each probability (p L , p R , p M , p M ) is set to 1 in 14-bit fixed-point notation (16384 = 1 << 14).

[0143] TIFF0007704802000015.tif51169

[0144] TIFF0007704802000016.tif51169

[0145] TIFF0007704802000017.tif17169

[0146] TIFF0007704802000018.tif15169

[0147] In an alternative embodiment, the bit estimation regarding the band can be obtained as follows.

[0148] When the M / S process is executed, the spectrum is divided into bands, and for each individual band, it is determined. For all bands where M / S is used, MDCT L,k and MDCT R,k are replaced with MDCT M,k = 0.5(MDCT L,k + MDCT R,k ) and MDCT S,k = 0.5(MDCT L,k - MDCT R,k ).

[0149] TIFF0007704802000019.tif68169

[0150] TIFF0007704802000020.tif66169

[0151] FIG. 7 illustrates calculating the bit rate for M / S determination regarding the band according to the embodiment.

[0152] In particular, in FIG. 7, a process for calculating b BW is described. To reduce complexity, the arithmetic coder context for encoding the spectrum up to band i - 1 is reduced and reused in band i.

[0153] TIFF0007704802000021.tif21170

[0154] FIG. 8 illustrates the determination of the stereo mode according to the embodiment.

[0155] When "full dual - mono" is selected, the complete spectrum consists of MDCT L,k and MDCT R,k When "full M / S" is selected, the complete spectrum consists of MDCT M,k and MDCT S,k When "M / S regarding the band" is selected, some bands of the spectrum consist of MDCT L,k and MDCT R,k and the other bands consist of MDCT M,k and MDCT S,k respectively.

[0156] The stereo mode is encoded in the bitstream. Also in the "M / S regarding the band" mode, the M / S determination regarding the band is encoded in the bitstream.

[0157] The spectral coefficients in the two channels after stereo processing are shown as MDCT LM,k and MDCT RS,k Depending on the stereo mode and the M / S determination regarding the band, MDCTLM,k is equal to the MDCT in the M / S band M,k or the MDCT in the L / R band L,k and the MDCT RS,k is equal to the MDCT in the M / S band S,k or the MDCT in the L / R band R,k and is equal to. The MDCT LM,k The spectrum consisting of is, for example, referred to as the combined and encoded channel 0 (combined channel 0), or as the first channel. The MDCT RS,k The spectrum consisting of is, for example, referred to as the combined and encoded channel 1 (combined channel 1), or as the second channel.

[0158] TIFF0007704802000022.tif73170

[0159] TIFF0007704802000023.tif73168

[0160] TIFF0007704802000024.tif41168

[0161] TIFF0007704802000025.tif43170

[0162] Quantization including a rate loop and noise filling and entropy coding are as described in 5.3.3.2 "General coding procedure" of 5.3.3 "MDCT based on TCX" in [6b] or [6a]. The rate loop is the estimated G estIt can be optimized using. The power spectrum P (magnitude of MCLT) is used for the color / tone means in quantization and intelligent gap filling (IGF) as described in [6a] or [6b]. Since the whitened and band-related M / S processed MDCT spectrum is used for the power spectrum, the same FDNS and M / S processing should be performed on the MDST spectrum. The same scaling based on the more comprehensive ILD of the larger channels should be performed for MDST as is done for MDCT. For frames where TNS is active, the MDST spectrum used for power spectrum calculation is the whitened and M / S processed MDCT spectrum: P k =MDCT k 2 +(MDCT k+1 -MDCT k-1 ) 2 is estimated from.

[0163] The decoding process begins with the decoding and inverse quantization of the spectrum of the combined and encoded channels, followed by noise filling, as described in 6.2.2 "MDCT based on TCX" in [6b] or [6a]. The number of bits assigned to each individual channel is determined based on the window length, stereo mode, and bit rate split ratio encoded in the bitstream. The number of bits assigned to each individual channel must be known before the bitstream can be fully decoded.

[0164] Within the intelligent gap filling (IGF) block, lines that are quantized to zero in a specific range of the spectrum (referred to as target tiles) are filled by content processed from different ranges of the spectrum, referred to as source styles. For stereo processing with respect to the band, the stereo representation (i.e., either L / R or M / S) is different for the source style and the target tile. To guarantee good quality, if the representation of the source style is different from that of the target tile, the source style is processed to convert it to the representation of the target tile before gap filling in the decoder. This procedure has already been described in [9]. The IGF itself, as compared to [6a] and [6b], is applied to the whitened spectral region instead of the original spectral region. Compared to known stereo coders (e.g., [9]), the IGF is applied in the whitened and ILD-corrected spectral region.

[0165] TIFF0007704802000026.tif26170

[0166] ratio ILD If > 1, the right channel is scaled by ratio ILD Otherwise, the left channel is scaled by 1 / ratio ILD

[0167] For individual cases where division by zero occurs, a small epsilon is added to the denominator.

[0168] For example, for an intermediate bitrate of 48 kbps, MDCT-based coding causes very poor quantization of the spectrum in order to match the bit consumption target. It raises the need for parametric coding that is applied on a frame-by-frame basis in combination with discrete coding in the same spectral region and faithfully increases.

[0169] ​In the following, some aspects of those embodiments that employ stereo filling are described. It should be noted that it is not necessary for stereo filling to be employed for the above embodiments. Therefore, only some of the embodiments described above employ stereo filling. Other embodiments of the embodiments described above do not employ stereo filling at all.

[0170] Stereo frequency filling in MPEG-H frequency domain stereo is described, for example, in

[11] . In

[11] , the target energy for each individual band is achieved by utilizing the band energy sent from the encoder in the form of a magnification (e.g., in AAC). When frequency domain noise shaping (FDNS) is applied and the spectral envelope is encoded using LSF (line spectral frequency) ([6a], [6b], and [8] are referred to), it is not possible to change the scaling for only some frequency bands (spectral bands) as required by the stereo filling algorithm described in

[11] .

[0171] First, some preliminary information is provided.

[0172] When mid / side coding is employed, it is possible to encode the side signal in different ways.

[0173] According to the first group of embodiments, the side signal S is encoded in the same way as the mid signal M. Quantization is performed, but no other steps are executed to reduce the required bitrate. Generally, such an approach aims to allow a very precise reconstruction of the side signal S on the decoder side, but on the other hand, requires a large amount of bits for encoding.

[0174] According to the second group of embodiments, the residual side signal S res is generated from the original side signal S based on the M signal. In the embodiment, the residual side signal is calculated, for example, according to the following formula. S res = S - g·M

[0175] Another embodiment adopts, for example, a different definition for the residual side signal.

[0176] Residual signal S res is quantized and transmitted to the decoder together with the parameter g. By quantizing the residual signal S instead of the original side signal S res generally, more spectral values are quantized to zero. This generally reduces the amount of bits required for encoding and transmission compared to the quantized original side signal S.

[0177] In some of these embodiments of the second group of embodiments, a single parameter g is determined for the complete spectrum and transmitted to the decoder. In another embodiment of the second group of embodiments, each of a plurality of frequency bands / spectral bands of the frequency spectrum includes, for example, two or more spectral values. The parameter g is determined for each of the frequency bands / spectral bands and transmitted to the decoder.

[0178] FIG. 12 illustrates encoder-side stereo processing according to the first group or the second group of embodiments that do not employ stereo padding.

[0179] FIG. 13 illustrates decoder-side stereo processing according to the first group or the second group of embodiments that do not employ stereo padding.

[0180] According to the third group of embodiments, stereo padding is employed. In some of these embodiments, on the decoder side, the side signal S for a particular time point t is generated from the mid signal at the immediately previous time point t - 1.

[0181] Generating the side signal S for a specific time point t from the mid signal at the time point t-1 immediately before the decoder side is performed according to the following equation. S(t)=h b ·M(t-1)

[0182] On the encoder side, the parameter h b is determined for each of the individual frequency bands of the plurality of frequency bands of the spectrum. After determining the parameter h b , the encoder transmits the parameter h b to the decoder. In some embodiments, the side signal S itself or the spectral values of its residuals are not transmitted to the decoder. Such an approach aims to reduce the number of bits required.

[0183] In some other embodiments of the third group of embodiments, for those frequency bands where the side signal is larger than the mid signal, at least the spectral values of the side signal of those frequency bands are explicitly encoded and transmitted to the decoder.

[0184] According to the fourth group of embodiments, some of the frequency bands of the side signal S are encoded by explicitly encoding the original side signal S (see the first group of embodiments) or the residual side signal S res . On the other hand, for another frequency band, stereo filling is employed. Such an approach combines the first group or the second group of embodiments with the third group of embodiments that employ stereo filling. For example, the lower frequency bands are encoded by quantizing the original side signal S or the residual side signal S res . On the other hand, for another higher frequency band, stereo filling is employed.

[0185] FIG. 9 illustrates the stereo processing on the encoder side according to the third group or the fourth group of embodiments that employ stereo filling.

[0186] FIG. 10 illustrates the decoder-side stereo processing according to the third or fourth group of embodiments employing stereo padding.

[0187] Those of the embodiments described above that employ stereo padding employ stereo padding as described, for example, in MPEG-H. See MPEG-H frequency-domain stereo (e.g., see

[11] ).

[0188] Some of the embodiments that employ stereo padding apply the stereo padding algorithm described in

[11] in a system where, for example, the spectral envelope is encoded as LSF combined with noise padding. Encoding the spectral envelope is performed, for example, as described in the examples in [6a], [6b], and [8]. Noise padding is performed as described, for example, in [6a] and [6b].

[0189] In some specific embodiments, the stereo padding process including stereo padding parameter calculation is performed in the M / S band in the frequency domain from a lower frequency such as 0.08F s (F s = sampling frequency) to a higher frequency (e.g., IGF crossover frequency).

[0190] For example, for a frequency portion lower than a lower frequency (e.g., 0.08F s ), the original side signal S or a residual side signal derived from the original side signal S is quantized and transmitted to the decoder. For a frequency portion higher than a higher frequency (e.g., IGF crossover frequency), intelligent gap filling (IGF) is performed.

[0191] More specifically, in some of the embodiments, the side channel (second channel) is filled from the whitened MDCT spectrum downmix of the previous frame using "copy-over" for those frequency bands within the stereo fill range that are fully quantized to zero (e.g., from 0.08 times the sampling frequency to the IGF crossover frequency) (IGF = intelligent gap fill). "Copy-over" is applied, for example, for free to noise fill and is scaled accordingly depending on the correction factor transmitted from the encoder. In another embodiment, the low frequencies are 0.08F s may represent a different value.

[0192] 0.08F s Instead, in some embodiments, the low frequencies are from 0 to 0.50F s within the range of values. In a particular embodiment, the low frequencies are from 0.01F s to 0.50F s within the range of values. For example, the low frequencies are 0.12F s , 0.20F s or 0.25F s .

[0193] In another embodiment, in addition to or instead of intelligent gap fill, noise fill is performed for frequencies greater than the upper frequency.

[0194] In another embodiment, in the absence of an upper frequency, stereo fill is performed for individual frequency portions greater than the lower frequency.

[0195] In yet another embodiment, in the absence of a lower frequency, stereo fill is performed for the frequency portion from the lowest frequency band to the upper frequency.

[0196] In yet another embodiment, in the absence of a lower frequency and an upper frequency, stereo fill is performed for the entire frequency spectrum.

[0197] In the following, specific embodiments adopting stereo padding are described.

[0198] In particular, stereo padding with a correction factor according to a specific embodiment is described. Stereo padding with a correction factor is adopted, for example, in the embodiments of the stereo padding processing blocks of FIGS. 9 (encoder side) and 10 (decoder side).

[0199] In the following, -Dmx R represents, for example, the mid signal of the whitened MDCT spectrum. -S R represents, for example, the side signal of the whitened MDCT spectrum. -Dmx I represents, for example, the mid signal of the whitened MDST spectrum. -S I represents, for example, the side signal of the whitened MDST spectrum. -prevDmx R represents, for example, the mid signal of the whitened MDCT spectrum delayed by one frame. -prevDmx I represents, for example, the mid signal of the whitened MDST spectrum delayed by one frame.

[0200] Stereo padding encoding is applied when the stereo decision is M / S (full M / S) for all bands, or when it is M / S (M / S with respect to the band) for all stereo padding bands.

[0201] When it is determined to apply full dual - mono processing, the stereo padding is bypassed. Further, when L / R encoding is selected for some of the spectral bands (frequency bands), the stereo padding is also bypassed for these spectral bands.

[0202] Now, certain embodiments employing stereo filling are considered. Thus, the processing within the block is executed, for example, as follows.

[0203] TIFF0007704802000027.tif121170

[0204] TIFF0007704802000028.tif68170

[0205] - From these calculated energies (ERes fb , EprevDmx fb ), a stereo filling correction factor is calculated and sent to the decoder as side information. correction_factor fb =ERes fb / (EprevDmx fb +ε)

[0206] In an embodiment, ε = 0. In another embodiment, for example, 0.1 > ε > 0 to avoid division by 0.

[0207] - The magnification factor for the band is calculated, for example, depending on the calculated stereo filling correction factor for each spectral band to which stereo filling is applied. On the decoder side, since there is no inverse complex prediction operation for reconstructing the side signal from the residual (a R =a I =0), scaling for the band of the output mid signal and the output side (residual) signal by the magnification factor is introduced to compensate for the energy loss.

[0208] TIFF0007704802000029.tif53170

[0209] TIFF0007704802000030.tif73169

[0210] Therefore, more bits are spent encoding the downmix of the residue and the lower frequency bins, enhancing the overall quality.

[0211] In an alternative embodiment, all bits of the residue (side) are set to, for example, 0. Such alternative embodiments are based, for example, on the assumption that the downmix is, in most cases, larger than the residue.

[0212] Figure 11 illustrates the stereo filling of the side signal according to some specific embodiments on the decoder side.

[0213] Stereo filling is applied to the side channel after decoding, inverse quantization, and noise filling. For frequency bands within the stereo filling range that are quantized to zero, if the band energy after noise filling does not reach the target energy, "copy-over" from the whitened MDCT spectrum downmix of the last frame is applied, for example, as seen in Figure 11. The target energy for each frequency band is calculated from the stereo correction factor transmitted as a parameter from the encoder, for example, according to the following equation. ET fb =correction_factor fb ·EprevDmx fb

[0214] TIFF0007704802000031.tif48169

[0215] TIFF0007704802000032.tif48168

[0216] On the encoder side, alternative embodiments do not consider the MDST spectrum (or the MDCT spectrum). In those embodiments, for example, the encoder-side procedure is applied as follows.

[0217] For a frequency band (fb), it is the lower frequency (e.g., 0.08Fs (F s (=sampling frequency)) starts and enters the frequency region that rises to the upper frequency (for example, IGF crossover frequency). -side signal S R The residual Res of is calculated, for example, according to the following formula. Res = S R -a R Dmx R Here, a R is a prediction coefficient (for example, a real number).

[0218] TIFF0007704802000033.tif49168

[0219] -These calculated energies (ERes fb , EprevDmx fb ) are used to calculate the stereo fill correction factor, which is then sent to the decoder as side information. correction_factor fb = ERes fb / (EprevDmx fb + ε)

[0220] In an embodiment, ε = 0. In another embodiment, for example, to avoid division by zero, 0.1 > ε > 0.

[0221] -The magnification for the band is calculated, for example, depending on the calculated stereo fill correction factor for each spectral band where stereo fill is employed.

[0222] TIFF0007704802000034.tif56170

[0223] TIFF0007704802000035.tif68168

[0224] Thus, more bits are spent encoding the downmix of the residue and the lower frequency bins, improving the overall quality.

[0225] In an alternative embodiment, all bits of the residue (side) are set to, for example, 0. Such an alternative embodiment is based, for example, on the assumption that the downmix is, in most cases, larger than the residue.

[0226] According to some of the embodiments, means are provided for applying stereo filling in a system having, for example, FDNS. Therein, the spectral envelope is encoded using LSF (or similar coding that scales in a single band and cannot be changed independently).

[0227] According to some of the embodiments, means are provided for applying stereo filling in a system, for example, without complex / real prediction.

[0228] Some of the embodiments employ parametric stereo filling to control the stereo filling of the whitened left and right MDCT spectra (e.g., by the downmix of the previous frame) in such a way that explicit parameters (stereo filling correction factors) are sent from the encoder to the decoder.

[0229] More generally, in some of the embodiments, the encoding unit 120 of FIGS. 1a to 1e is configured to generate a processed audio signal such that, for example, at least one spectral band of the first channel of the processed audio signal is the spectral band of the mid signal, and at least one spectral band of the second channel of the processed audio signal is the spectral band of the side signal. To obtain the encoded audio signal, the encoding unit 120 is configured to encode the spectral band of the side signal, for example, by determining a correction factor for the spectral band of the side signal. The encoding unit 120 is configured to determine the correction factor for the spectral band of the side signal, for example, depending on the residual and depending on the spectral band of a previous mid signal corresponding to the spectral band of the mid signal. The previous mid signal precedes the mid signal in time. Further, the encoding unit 120 is configured to determine the residual, for example, depending on the spectral band of the side signal and depending on the spectral band of the mid signal.

[0230] According to some of the embodiments, the encoding unit 120 is configured to determine the correction factor for the spectral band of the side signal, for example, according to the following equation. correction_factor fb =ERes fb / (EprevDmx fb +ε) Here, correction_factor fb represents the correction factor for the spectral band of the side signal. ERes fb represents the residual energy that depends on the energy of the spectral band of the residual corresponding to the spectral band of the mid signal. EprevDmx fbindicates the prior energy that depends on the energy in the spectral band of the prior mid-signal. ε = 0, or 0.1 > ε > 0.

[0231] TIFF0007704802000036.tif58168

[0232] According to some of the embodiments, the residual is defined according to the following formula. Res R =S R -a R Dmx R -a I Dmx I Here, Res R is the residual. S R is the side signal. a R is the real part of the complex (prediction) coefficient, and a I is the imaginary part of the complex (prediction) coefficient. Dmx R is the mid-signal. Dmx I is another mid-signal that depends on the first channel of the normalized audio signal and also depends on the second channel of the normalized audio signal. Another side signal S that depends on the first channel of the normalized audio signal and also depends on the second channel of the normalized audio signal I Another residual of is defined according to the following formula. Res I =S I -a R Dmx R -a I Dmx I

[0233] TIFF0007704802000037.tif54169

[0234] In some of the embodiments, the decoding unit 210 in FIGS. 2a to 2e is configured to determine, for example, for each of the individual spectral bands of the plurality of spectral bands, whether the spectral band of the first channel of the encoded audio signal and the spectral band of the second channel of the encoded audio signal are encoded using dual-mono encoding or mid-side encoding. Further, the decoding unit 210 is configured to obtain, for example, the spectral band of the second channel of the encoded audio signal by reconstructing the spectral band of the second channel. When mid-side encoding is used, the spectral band of the first channel of the encoded audio signal is the spectral band of the mid signal, and the spectral band of the second channel of the encoded audio signal is the spectral band of the side signal. Further, when mid-side encoding is used, the decoding unit 210 is configured to reconstruct the spectral band of the side signal depending on, for example, a correction factor for the spectral band of the side signal and depending on the spectral band of a preceding mid signal corresponding to the spectral band of the mid signal. The preceding mid signal precedes the mid signal in time.

[0235] TIFF0007704802000038.tif99170

[0236] In some of the embodiments, the residual is derived, for example, from a complex stereo prediction algorithm on the encoder side. On the other hand, stereo prediction (real or complex) does not exist on the decoder side.

[0237] According to some of the embodiments, energy correction scaling of the spectrum on the encoder side is used, for example, to compensate for the fact that inverse prediction processing does not exist on the decoder side.

[0238] Although several aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of a method in which a block or device corresponds to a method step or a function of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or function of a corresponding apparatus. Some or all of the method steps are performed by (or with) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps are performed by such a device.

[0239] Depending on certain implementation requirements, embodiments of the invention are realized in hardware, software, at least in part in hardware or at least in part in software. The realization is carried out using a digital storage medium having electronically readable control signals stored thereon, for example, a floppy disk, a DVD, a Blu-ray disk, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory. They cooperate with, or can cooperate with, a programmable computer system such that the respective methods are executed. Accordingly, the digital storage medium is computer-readable.

[0240] Some embodiments according to the invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system such that one of the methods described herein is executed.

[0241] In general, embodiments of the invention are executed as a computer program product having program code. The program code serves to execute one of the methods when the computer program product runs on a computer. The program code is stored, for example, on a machine-readable carrier.

[0242] Another embodiment includes a computer program for performing one of the methods described herein. The computer program is stored on a machine-readable carrier.

[0243] That is, an embodiment of the method of the present invention is a computer program having program code for performing one of the methods described herein when the computer program runs on a computer.

[0244] Accordingly, another embodiment of the method of the present invention includes a data carrier (or digital storage medium or computer-readable medium) that contains a computer program for performing one of the methods described herein recorded thereon.

[0245] Accordingly, another embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals is configured to be transmitted, for example, via a data communication connection (e.g., via the Internet).

[0246] Another embodiment includes a processing means, for example, a computer or programmable logic device configured or adapted to perform one of the methods described herein.

[0247] Another embodiment includes a computer on which a computer program for performing one of the methods described herein is installed.

[0248] Another embodiment according to the invention includes an apparatus or system configured to transmit a computer program for performing at least one of the methods described herein to a receiver. The receiver is, for example, a computer or a mobile device or a memory device or a similar device. The apparatus or system includes, for example, a file server for transmitting the computer program to the receiver.

[0249] In some embodiments, a programmable logic device (e.g., a field programmable gate array, FPGA) is used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array cooperates with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0250] The apparatuses described herein are implemented using a hardware device, or using a computer, or by using a combination of a hardware device and a computer.

[0251] The methods described herein are executed using a hardware device, or using a computer, or by using a combination of a hardware device and a computer.

[0252] The above-described embodiments merely illustrate the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Accordingly, the invention is intended to be limited only by the scope of the appended patent claims and not by the specific details shown and described in these embodiments.

[0253] References [1] J. Herre, E. Eberlein and K. Brandenburg, “Combined Stereo Coding”, in 93rd AES Convention, San Francisco, 1992. [2] J. D. Johnston and A. J. Ferreira, “Sum-difference stereo transform coding”, in Proc. ICASSP, 1992. [3] ISO / IEC 11172-3, Information technology - Coding of moving pictures and associated audio for digital storage media at up to about 1,5 Mbit / s - Part 3 : Audio, 1993. [4] ISO / IEC 13818-7, Information technology - Generic coding of moving pictures and associated audio information - Part 7: Advanced Audio Coding (AAC), 2003. [5] J.-M. Valin, G. Maxwell, T. B. Terriberry and K. Vos, “High-Quality, Low-Delay Music Coding in the Opus Codec”, in Proc. AES 135th Convention, New York, 2013. [6a] 3GPP TS 26.445, Codec for Enhanced Voice Services (EVS); Detailed algorithmic description, V 12.5.0, Dezember 2015. [6b] 3GPP TS 26.445, Codec for Enhanced Voice Services (EVS); Detailed algorithmic description, V 13.3.0, September 2016. [7] H. Purnhagen, P. Carlsson, L. Villemoes, J. Robilliard, M. Neusinger, C. Helmrich, J. Hilpert, N. Rettelbach, S. Disch and B. Edler, “Audio encoder, audio decoder and related methods for processing multi-channel audio signal s using complex prediction”. US Patent 8,655,670 B2, 18 February 2014. [8] G. Markovic, F. Guillaume, N. Rettelbach, C. Helmrich and B. Schubert, “ Linear prediction based coding scheme using spectral domain noise shaping”. European Patent 2676266 B1, 14 February 2011. [9] S. Disch, F. Nagel, R. Geiger, B. N. Thoshkahna, K. Schmidt, S. Bayer, C. Neukam, B. Edler and C. Helmrich, “Audio Encoder, Audio Decoder and Relat ed Methods Using Two-Channel Processing Within an Intelligent Gap Filling Fr amework”. International Patent PCT / EP2014 / 065106, 15 07 2014.

[10] C. Helmrich, P. Carlsson, S. Disch, B. Edler, J. Hilpert, M. Neusinger, H. Purnhagen, N. Rettelbach, J. Robilliard and L. Villemoes, “Efficient Transform Coding Of Two-channel Audio Signals By Means Of Complex-valued Stereo Prediction”, in Acoustics, Speech and Signal Processing (ICASSP), 2011 IEEE International Conference on, Prague, 2011.

[11] C. R. Helmrich, A. Niedermeier, S. Bayer and B. Edler, “Low-complexity semi-parametric joint-stereo audio transform coding”, in Signal Processing Conference (EUSIPCO), 2015 23rd European, 2015.

[12] H. Malvar, "A Modulated Complex Lapped Transform and its Applications to Audio Processing", in Acoustics, Speech, and Signal Processing (ICASSP), 1999. Proceedings., 1999 IEEE International Conference on, Phoenix, AZ, 1999.

[13] B. Edler and G. Schuller, “Audio coding using a psychoacoustic pre- and post-filter” Acoustics, Speech, and Signal Processing, 2000. ICASSP ’00.

Claims

1. An apparatus for encoding a first channel and a second channel of an audio input signal including two or more channels to obtain an encoded audio signal, wherein the first audio signal depends on the audio input signal, the apparatus comprising: an encoding unit (120) configured to generate a processed audio signal having the first channel and the second channel, wherein one or more spectral bands of the first channel of the processed audio signal are one or more spectral bands of the first channel of the first audio signal, and at least one spectral band of the first channel of the processed audio signal is a spectral band of a mid signal that depends on the spectral band of the first channel of the first audio signal and on the spectral band of the second channel of the first audio signal, and the encoding unit (120) is configured to encode the processed audio signal to obtain the encoded audio signal, the apparatus being characterized thereby.

2. The encoding unit is configured to generate the processed audio signal such that one or more spectral bands of the second channel of the processed audio signal are one or more spectral bands of the second channel of the first audio signal, and at least one spectral band of the second channel of the processed audio signal is a spectral band of a side signal that depends on the spectral band of the first channel of the first audio signal and on the spectral band of the second channel of the first audio signal, the apparatus according to claim 1, characterized thereby.

3. The encoding unit (120) is configured to select from a full mid-side encoding mode, a full dual-mono encoding mode, and a per-band encoding mode depending on a plurality of spectral bands of the first channel of the first audio signal and on a plurality of spectral bands of the second channel of the first audio signal, When the full mid-side encoding mode is selected, the encoding unit (120) generates a mid signal from the first channel and the second channel of the first audio signal as the first channel of the mid-side signal, and generates a side signal from the first channel and the second channel of the first audio signal as the second channel of the mid-side signal, and is configured to encode the mid-side signal to obtain the encoded audio signal. When the full dual-mono encoding mode is selected, the encoding unit (120) is configured to encode the first audio signal to obtain the encoded audio signal. When the per-band encoding mode is selected, the encoding unit (120) is configured to encode the first audio signal to obtain the encoded audio signal. When the per-band encoding mode is selected, the encoding unit (120) is configured to generate the processed audio signal such that one or more spectral bands of the first channel of the processed audio signal are one or more spectral bands of the first channel of the first audio signal, and one or more spectral bands of the second channel of the processed audio signal are one or more spectral bands of the second channel of the first audio signal, and at least one spectral band of the first channel of the processed audio signal is a spectral band of a mid signal that depends on the spectral band of the first channel of the first audio signal and the spectral band of the second channel of the first audio signal, and at least one spectral band of the second channel of the processed audio signal is a spectral band of a side signal that depends on the spectral band of the first channel of the first audio signal and the spectral band of the second channel of the first audio signal. The encoding unit (120) is configured to encode the processed audio signal to obtain the encoded audio signal. The apparatus according to claim 1 or claim 2, characterized in that.

4. When the per-band encoding mode is selected, the encoding unit (120) is configured to determine whether to employ mid-side encoding or dual-mono encoding for each spectral band of a plurality of spectral bands of the processed audio signal. When mid-side encoding is employed for the spectral band, the encoding unit (120) is configured to generate the spectral band of the first channel of the processed audio signal as the spectral band of a mid signal based on the spectral band of the first channel of the first audio signal and the spectral band of the second channel of the first audio signal. Also, the encoding unit (120) is configured to generate the spectral band of the second channel of the processed audio signal as the spectral band of a side signal based on the spectral band of the first channel of the first audio signal and the spectral band of the second channel of the first audio signal. When dual-mono encoding is employed for the spectral band the encoding unit (120) is configured to use the spectral band of the first channel of the first audio signal as the spectral band of the first channel of the processed audio signal and use the spectral band of the second channel of the first audio signal as the spectral band of the second channel of the processed audio signal, or the encoding unit (120) is configured to use the spectral band of the second channel of the first audio signal as the spectral band of the first channel of the processed audio signal and use the spectral band of the first channel of the first audio signal as the spectral band of the second channel of the processed audio signal. The apparatus according to claim 3, characterized by the above. **Claim 5** The encoding unit (120) determines a first estimation for estimating the number of first bits required for encoding when the full mid-side encoding mode is adopted, determines a second estimation for estimating the number of second bits required for encoding when the full dual-mono encoding mode is adopted, determines a third estimation for estimating the number of third bits required for encoding when the per-band encoding mode is adopted, and selects, from among the full mid-side encoding mode, the full dual-mono encoding mode, and the per-band encoding mode, the encoding mode with the smallest number of bits among the first estimation, the second estimation, and the third estimation, and is configured to select from among the full mid-side encoding mode, the full dual-mono encoding mode, and the per-band encoding mode. The apparatus according to claim 3 or claim 4, characterized in that.

6. The encoding unit (120) determines a first estimation for estimating the number of first bits to be saved when encoding in the full mid-side encoding mode, determines a second estimation for estimating the number of second bits to be saved when encoding in the full dual-mono encoding mode, determines a third estimation for estimating the number of third bits to be saved when encoding in the per-band encoding mode, and selects, from among the full mid-side encoding mode, the full dual-mono encoding mode, and the per-band encoding mode, the encoding mode with the largest number of bits to be saved among the first estimation, the second estimation, and the third estimation, and is configured to select from among the full mid-side encoding mode, the full dual-mono encoding mode, and the per-band encoding mode, or The encoding unit (120) estimates a first signal-to-noise ratio that occurs when the full mid-side encoding mode is adopted, estimates a second signal-to-noise ratio that occurs when the full dual-mono encoding mode is adopted, estimates a third signal-to-noise ratio that occurs when the per-band encoding mode is adopted, and selects, from among the full mid-side encoding mode, the full dual-mono encoding mode, and the per-band encoding mode, the encoding mode having the largest signal-to-noise ratio among the first signal-to-noise ratio, the second signal-to-noise ratio, and the third signal-to-noise ratio, and is configured to select from among the full mid-side encoding mode, the full dual-mono encoding mode, and the per-band encoding mode. The apparatus according to claim 3 or claim 4, characterized in that.

7. The encoding unit (120) is configured to generate the processed audio signal such that at least one spectral band of the first channel of the processed audio signal is the spectral band of the mid signal and at least one spectral band of the second channel of the processed audio signal is the spectral band of the side signal. To obtain the encoded audio signal, the encoding unit (120) is configured to encode the spectral band of the side signal by determining a correction factor for the spectral band of the side signal. The encoding unit (120) is configured to determine the correction factor for the spectral band of the side signal depending on the residual and depending on the spectral band of a preceding mid signal corresponding to the spectral band of the mid signal, the preceding mid signal preceding the mid signal in time. The encoding unit (120) is configured to determine the residual depending on the spectral band of the side signal and depending on the spectral band of the mid signal. The apparatus according to claim 3 or claim 4, characterized in that.

8. The encoding unit (120) is configured to determine the correction factor for the spectral band of the side signal according to the formula correction_factor fb = ERes fb / (EprevDmx fb + ε) ​ Here, correction_factor fb represents the correction factor for the spectral band of the side signal, ERes fb indicates the residual energy that depends on the energy of the spectral band of the residual corresponding to the spectral band of the mid-signal, EprevDmx fb indicates the previous energy that depends on the energy of the spectral band of the previous mid signal, ε = 0 or 0.1 > ε > 0, The apparatus according to claim 7, characterized in that.

9. The residual is defined by the formula Res R = S R - a R Dmx R - a I Dmx I and is defined according to Here, Res R is the residual, S R is the side signal, a R is a coefficient, Dmx R is the mid signal, The encoding unit (120) is configured to determine the residual energy according to the formula The apparatus according to claim 7 or claim 8, characterized in that.

10. The residual is defined by the formula and is defined according to Res R = S R - a R Dmx R - a I Dmx I The encoding unit (120) is configured to determine the residual energy according to the formula Here, Res R is the residual, S R is the side signal, a R is the real part of the complex coefficient, a I is the imaginary part of the complex coefficient, Dmx R is the mid signal, Dmx I is another mid signal that depends on the first channel of the first audio signal and also depends on the second channel of the first audio signal Another side signal S that depends on the first channel of the first audio signal and depends on the second channel of the first audio signal I Another residual of is given by the formula Res R = S R Res = S R - a R Dmx R - a I Dmx I The encoding unit (120) is configured to determine the residual energy according to the formula The encoding unit (120) depends on the energy of the spectrum band of the residual corresponding to the spectrum band of the mid-signal and determines the prior energy that depends on the energy of the spectrum band of the other residual corresponding to the spectrum band of the mid-signal. The apparatus according to claim 8, characterized in that.

11. The apparatus includes a normalizer (110) configured to determine a normalization value for the audio input signal depending on the first channel of the audio input signal and depending on the second channel of the audio input signal. The normalizer (110) is configured to determine the first channel and the second channel of the first audio signal, which is the normalized audio signal, by modulating at least one of the first channel and the second channel of the audio input signal depending on the normalization value. The apparatus according to any one of claims 1 to 10, characterized in that.

12. The normalizer (110) is configured to determine the normalization value for the audio input signal depending on the energy of the first channel of the audio input signal and depending on the energy of the second channel of the audio input signal. The apparatus according to claim 11, characterized in that.

13. The audio input signal is represented in the spectral domain, The normalizer (110) is configured to determine the normalization value for the audio input signal depending on a plurality of spectral bands of the first channel of the audio input signal and depending on a plurality of spectral bands of the second channel of the audio input signal. ​ ​ The normalizer (110) is configured to determine the first audio signal by modulating a plurality of spectral bands of at least one of the first channel and the second channel of the audio input signal depending on the normalization value. The apparatus according to claim 11, characterized in that.

14. The normalizer (110) is configured to determine the normalization value based on the formula and is configured to determine the normalization value by quantizing the ILD. Here, MDCT L,k is the k-th coefficient of the MDCT spectrum of the first channel of the audio input signal, and MDCT R,k is the k-th coefficient of the MDCT spectrum of the second channel of the audio input signal, The apparatus according to claim 13, characterized in that.

15. The apparatus for encoding further includes a conversion unit (102) and a preprocessing unit (105), The conversion unit (102) is configured to convert a time-domain audio signal from the time domain to the frequency domain to obtain a converted audio signal. The preprocessing unit (105) is configured to generate the first channel and the second channel of the audio input signal by applying an encoder-side frequency-domain noise shaping operation to the converted audio signal. The apparatus according to claim 13 or claim 14, characterized in that.

16. The preprocessing unit (105) is configured to generate the first channel and the second channel of the audio input signal by applying an encoder-side time-domain noise shaping operation to the converted audio signal before applying the encoder-side frequency-domain noise shaping operation to the converted audio signal. The apparatus according to claim 15, characterized in that.

17. The normalizer (110) is configured to determine a normalization value for the audio input signal depending on the first channel of the audio input signal represented in the time domain and depending on the second channel of the audio input signal represented in the time domain. The normalizer (110) is configured to determine the first channel and the second channel of the first audio signal by modulating at least one of the first channel and the second channel of the audio input signal represented in the time domain depending on the normalization value. ​ The apparatus further includes a conversion unit (115) configured to convert the first audio signal from the time domain to the spectral domain such that the first audio signal is represented in the spectral domain. The conversion unit (115) is configured to supply the first audio signal represented in the spectral domain to the encoding unit (120). The apparatus according to any one of claims 1 to 10, characterized in that.

18. The apparatus further includes a preprocessing unit (106) configured to receive a time-domain audio signal including a first channel and a second channel. The preprocessing unit (106) is configured to apply a filter that creates a first perceptually whitened spectrum to the first channel of the time-domain audio signal to obtain the first channel of the audio input signal represented in the time domain. The preprocessing unit (106) is configured to apply a filter that creates a second perceptually whitened spectrum to the second channel of the time-domain audio signal to obtain the second channel of the audio input signal represented in the time domain. The apparatus according to claim 17, characterized in that.

19. The conversion unit (115) is configured to convert the first audio signal from the time domain to the spectral domain to obtain a converted audio signal. The apparatus further includes a spectral domain preprocessor (118) configured to perform encoder-side time noise shaping on the converted audio signal to obtain the first audio signal represented in the spectral domain. The apparatus according to claim 17 or claim 18, characterized in that.

20. The encoding unit (120) is configured to obtain the encoded audio signal by applying encoder-side stereo intelligent gap filling to the first audio signal or the processed audio signal. The apparatus according to any one of claims 1 to 19, characterized in that.

21. The audio input signal is an audio stereo signal that strictly includes two channels. The apparatus according to any one of claims 1 to 20, characterized in that.

22. A system for encoding four channels of an audio input signal including four or more channels to obtain an encoded audio signal, the system comprising: A first device (170) according to any one of claims 1 to 20 for encoding a first channel and a second channel of the four or more channels of the audio input signal to obtain a first channel and a second channel of the encoded audio signal; A second device (180) according to any one of claims 1 to 20 for encoding a third channel and a fourth channel of the four or more channels of the audio input signal to obtain a third channel and a fourth channel of the encoded audio signal, and comprising: A system, characterized by:

23. A device for decoding an encoded audio signal including a first channel and a second channel to obtain a first channel and a second channel of a decoded audio signal including two or more channels, The device comprises a decoding unit (210), In a first case, the decoding unit (210) is configured to use the spectral band of the first channel of the encoded audio signal as the spectral band of the first channel of the intermediate audio signal, and use the spectral band of the second channel of the encoded audio signal as the spectral band of the second channel of the intermediate audio signal. In a second case, the decoding unit (210) is configured to generate the spectral band of the first channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal. The device is configured to obtain the decoded audio signal from the intermediate audio signal by denormalizing at least one of the first channel and the second channel of the intermediate audio signal. A device, characterized by:

24. For each spectral band of a plurality of spectral bands, the decoding unit (210) determines whether the spectral band of the first channel of the encoded audio signal and the spectral band of the second channel of the encoded audio signal are encoded using dual-mono encoding or mid-side encoding, when the mid-side encoding is used, the decoding unit (210) is configured to generate a spectral band of the second channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal, The apparatus according to claim 23, characterized in that

25. the decoding unit (210) is configured to determine whether the encoded audio signal is encoded in a full mid-side encoding mode, or in a full dual-mono encoding mode, or in a per-band encoding mode, when it is determined that the encoded audio signal is encoded in the full mid-side encoding mode, the decoding unit (210) generates the first channel of the intermediate audio signal from the first channel and the second channel of the encoded audio signal, and generates the second channel of the intermediate audio signal from the first channel and the second channel of the encoded audio signal, when it is determined that the encoded audio signal is encoded in the full dual-mono encoding mode, the decoding unit (210) uses the first channel of the encoded audio signal as the first channel of the intermediate audio signal and uses the second channel of the encoded audio signal as the second channel of the intermediate audio signal, when it is determined that the encoded audio signal is encoded in the per-band encoding mode, For each of the plurality of spectral bands, configured to determine whether the spectral band of the first channel of the encoded audio signal and the spectral band of the second channel of the encoded audio signal are encoded using dual-mono encoding or encoded using a mid-side encoding mode, When the dual-mono encoding is used, configured to use the spectral band of the first channel of the encoded audio signal as the spectral band of the first channel of the intermediate audio signal, and use the spectral band of the second channel of the encoded audio signal as the spectral band of the second channel of the intermediate audio signal, When the mid-side encoding is used, generate the spectral band of the first channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal, and generate the spectral band of the second channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal configured to be such that, The apparatus according to claim 23 or claim 24, characterized in that.

26. The decoding unit (210) is configured to determine whether the spectral band of the first channel of the encoded audio signal and the spectral band of the second channel of the encoded audio signal are encoded using the dual-mono encoding or encoded using the mid-side encoding for each of the plurality of spectral bands, The decoding unit (210) is configured to obtain the spectral band of the second channel of the encoded audio signal by reconstructing the spectral band of the second channel, If mid-side coding is used, the spectral band of the first channel of the encoded audio signal is the spectral band of the mid signal, and the spectral band of the second channel of the encoded audio signal is the spectral band of the side signal. If mid-side coding is used, the decoding unit (210) is configured to reconstruct the spectral band of the side signal depending on a correction factor for the spectral band of the side signal and depending on the spectral band of a preceding mid signal corresponding to the spectral band of the mid signal, where the preceding mid signal precedes the mid signal in time. The apparatus according to claim 25, characterized in that. **Claim 27** The apparatus includes a denormalizer (220) configured to modulate at least one of the first channel and the second channel of the intermediate audio signal depending on a denormalization value to obtain the first channel and the second channel of the decoded audio signal. The apparatus according to any one of claims 23 to 26, characterized in that. **Claim 28** The denormalizer (220) is configured to modulate the plurality of spectral bands of at least one of the first channel and the second channel of the intermediate audio signal depending on the denormalization value to obtain the first channel and the second channel of the decoded audio signal. The denormalizer (220) is configured to modulate the plurality of spectral bands of at least one of the first channel and the second channel of the intermediate audio signal depending on the denormalization value to obtain a denormalized audio signal. The apparatus further includes a post-processing unit (230) and a conversion unit (235). The post-processing unit (230) is configured to perform at least one of decoder-side time noise shaping and decoder-side frequency domain noise shaping on the denormalized audio signal to obtain a post-processed audio signal. The conversion unit (235) is configured to convert the post-processed audio signal from the spectral domain to the time domain to obtain the first channel and the second channel of the decoded audio signal. The apparatus according to claim 27, characterized in that.

29. The apparatus further includes a conversion unit (215) configured to convert the intermediate audio signal from the spectral domain to the time domain. The denormalizer (220) is configured to modulate at least one of the first channel and the second channel of the intermediate audio signal represented in the time domain depending on the denormalization value to obtain the first channel and the second channel of the decoded audio signal. The apparatus according to claim 27, characterized in that.

30. The apparatus further includes a conversion unit (215) configured to convert the intermediate audio signal from the spectral domain to the time domain. The denormalizer (220) is configured to modulate at least one of the first channel and the second channel of the intermediate audio signal represented in the time domain depending on the denormalization value to obtain a denormalized audio signal. The apparatus further includes a post-processing unit (235) configured to process the denormalized audio signal, which is a perceptually whitened audio signal, to obtain the first channel and the second channel of the decoded audio signal. The apparatus according to claim 27, characterized in that.

31. The apparatus further includes a spectral domain post-processor (212) configured to perform decoder-side time noise shaping on the intermediate audio signal. The conversion unit (215) is configured to convert the intermediate audio signal from the spectral domain to the time domain after performing decoder-side time noise shaping on the intermediate audio signal. The apparatus according to claim 29 or claim 30, characterized in that.

32. The decoding unit (210) is configured to apply decoder-side stereo intelligent gap filling to the encoded audio signal. The apparatus according to any one of claims 23 to 31, characterized in that.

33. The decoded audio signal is an audio stereo signal that strictly includes two channels. The apparatus according to any one of claims 23 to 32, characterized in that.

34. A system for decoding an encoded audio signal including four or more channels to obtain four channels of a decoded audio signal including four or more channels, the system comprising: A first device (270) according to any one of claims 23 to 32 for decoding a first channel and a second channel of the four or more channels of the encoded audio signal to obtain a first channel and a second channel of the decoded audio signal; A second device (280) according to any one of claims 23 to 32 for decoding a third channel and a fourth channel of the four or more channels of the encoded audio signal to obtain a third channel and a fourth channel of the decoded audio signal, comprising: A system, characterized in that.

35. A system for generating an encoded audio signal from an audio input signal and for generating a decoded audio signal from the encoded audio signal, the system comprising: A device (310) according to any one of claims 1 to 21, wherein the device (310) according to any one of claims 1 to 21 is configured to generate the encoded audio signal from the audio input signal; A device (320) according to any one of claims 23 to 33, wherein the device (320) according to any one of claims 23 to 33 is configured to generate the decoded audio signal from the encoded audio signal; Including. A system, characterized in that.

36. A system for generating an encoded audio signal from an audio input signal and for generating a decoded audio signal from the encoded audio signal, the system comprising: A system according to claim 22, wherein the system according to claim 22 is configured to generate the encoded audio signal from the audio input signal. The system according to claim 34, wherein the system according to claim 34 is configured to generate the decoded audio signal from the encoded audio signal, comprising a system, characterized by

37. A method for encoding a first channel and a second channel of an audio input signal including two or more channels to obtain an encoded audio signal, wherein the first audio signal depends on the audio input signal, and the method comprises: generating a processed audio signal having a first channel and a second channel, wherein one or more spectral bands of the first channel of the processed audio signal are one or more spectral bands of the first channel of the first audio signal, and at least one spectral band of the first channel of the processed audio signal is a spectral band of a mid-signal that depends on the spectral band of the first channel of the first audio signal; encoding the processed audio signal to obtain the encoded audio signal; comprising a method, characterized by

38. A method for decoding a first channel and a second channel of an encoded audio signal including a first channel and a second channel to obtain a decoded audio signal including two or more channels, in a first case, the spectral band of the first channel of the encoded audio signal is used as the spectral band of the first channel of the intermediate audio signal, and the spectral band of the second channel of the encoded audio signal is used as the spectral band of the second channel of the intermediate audio signal; in a second case, the method includes generating a spectral band of the first channel of the intermediate audio signal based on the spectral band of the first channel of the encoded audio signal and based on the spectral band of the second channel of the encoded audio signal; the method includes obtaining the decoded audio signal from the intermediate audio signal by denormalizing at least one of the first channel and the second channel of the intermediate audio signal; a method, characterized by Claim 39 A computer program for executing the method according to claim 37 when executed on a computer or a signal processor. Claim 40 A computer program for executing the method according to claim 38 when executed on a computer or a signal processor.

Citation Information

Patent Citations

  • Encoding method and decoding method of signal and encoder and decoder using the same

    JP1996095599A

  • An audio encoder, encoding method, decoder, decoding method, and encoded audio signal for encoding an audio signal having an impulsive portion and a steady portion.

    JP2010530079A

  • Advanced stereo coding based on adaptively selectable left / right or mid / side stereo coding and parametric stereo coding combinations.

    JP2012521012A

  • MDCT-based complex predictive stereo coding

    JP2013524281A

  • Linear prediction-based coding scheme using spectral noise shaping

    JP2014510306A