Title - APPARATUS AND METHOD FOR DECODING AN ENCODED AUDIO SIGNAL, AND APPARATUS AND METHOD FOR ENCODING AN AUDIO SIGNAL
Patent Information
- Application Number
- ARP20220100163
- Authority / Receiving Office
- AR · AR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-11-10
- Filing Date
- 2022-01-27
- Publication Date
- 2026-08-26
- Estimated Expiration
- 2038-11-09
AI Technical Summary
Existing audio coding techniques face challenges in achieving low bit rates without significant quality loss, particularly due to high bit requirements for scale factors and complexity in spectral noise modeling, and are inflexible in adjusting perceptual filters.
An apparatus and method for encoding and decoding audio signals using subsampling and interpolation of scale parameters, which involves converting audio signals to a spectral representation, calculating a first set of scale parameters, subsampling to a second set, and interpolating back to a higher number for decoding, utilizing logarithmic domains for efficient vector quantization and spectral noise shaping.
This approach achieves low bit rates with high-quality spectral processing, minimizing perceptual noise and reducing complexity, while allowing flexible adjustments in perceptual filtering.
Abstract
Description
APPARATUS AND METHOD FOR ENCODING AND DECODING AN AUDIO SIGNAL USING SUBSAMPLING OR SCALE PARAMETER INTERPOLATION Descriptive report The present invention relates to audio processing and, in particular, to audio processing that operates in a spectral domain using scaling parameters for spectral bands. Known technique 1: Advanced Audio Coding (AAC) In one of the most widely used state-of-the-art perceptual audio codecs, Advanced Audio Coding (AAC) [1-2], spectral noise modeling is done with the help of so-called scaling factors. In this approach, the MDCT (Modified Discrete Cosine Transform) spectrum is divided into a number of non-uniform scaling factor bands. For example, at 48 kHz, the MDCT has 1,024 coefficients and is divided into 49 scaling factor bands. In each band, a scaling factor is used to scale the MDCT coefficients for that band. A scalar quantizer with a constant step size is then used to quantize the scaled MDCT coefficients. On the decoder side, inverse scaling is performed in each band, modeling the quantization noise introduced by the scalar quantizer. 238337 1651425 of 41 The 49 scaling factors are encoded in the bit stream as side information. This usually requires a considerably large number of bits to encode the scaling factors, due to the relatively large number of scaling factors and the high precision required. This can become a problem with low bit rates and / or low latency. Known technique 2: TCX (Transform Coded Excitation) based on MDCT In MDCT-based TCX, a transform-based audio codec used in the USAC (Unified Speech and Audio Coding), MPEG-D (Motion Picture Experts Group [3], and 3GPP (3rd Generation Partnership Project) standards) and EVS (Enhanced Voice Services [4] standards, spectral noise modeling is performed with the help of an LPC-based perceptual filter (Linear Predictive Coding), the same perceptual filter used in recent ACELP-based speech codecs (e.g.,AMR-WB (Adaptive Multi-Rate Wideband). In this approach, a set of 16 LPCs is first estimated in a 238337 1651425 of 41 pre-emphasized input signal. The LPCs are then weighted and quantized. The frequency response of the weighted and quantized LPCs is then computed in 64 evenly spaced bands. The MDCT coefficients are then scaled in each band using the computed frequency response. The scaled MDCT coefficients are then quantized using a scalar quantizer with a step size controlled by an overall gain. In the decoder, inverse scaling is performed on each of the 64 bands, modeling the quantization noise introduced by the scalar quantizer. This approach has a distinct advantage over the AAC approach: it requires encoding only 16 (LPC) + 1 (global gain) parameters as side information (as opposed to the 49 parameters in AAC). Furthermore, the 16 LPC parameters can be efficiently encoded with a small number of bits by employing an LSF (Line Spectral Frequencies) representation and a vector quantizer. Consequently, the approach described in technique 2 requires fewer bits of side information than the approach described in technique 1, which can be a significant difference at low bit rates and / or with low latency. However, this approach also has some drawbacks. The first drawback is that the frequency scale of the noise modeling is limited to being linear (i.e., using uniformly spaced bands) because the LPCs are estimated in the time domain. This is a disadvantage since the human ear is more sensitive to low frequencies than to high frequencies. The second drawback is the high complexity required by 238337 1651425 of 41 This approach. LPC estimation (autocorrelation, Levinson-Durbin), LPC quantization (LPC<->LSF conversion, vector quantization), and LPC frequency response calculation are all very costly operations. The third drawback is that this approach is not very flexible because the LPC-based perceptual filter cannot be easily modified, and this prevents some specific adjustments that would be required for critical audio items. Known technique 3: Enhanced MDCT-based TCX Some recent work has overcome the first drawback and partially the second drawback of the known technique. This work was published in documents US 9595262 B2 and EP2676266 B1. In this new approach, autocorrelation (for estimating LPCs) is no longer performed in the time domain but is instead calculated in the MDCT domain using an inverse transform of the MDCT coefficient energies. This allows the use of a non-uniform frequency scale by simply grouping the MDCT coefficients into 64 non-uniform bands and computing the energy of each band. This reduces the complexity required to compute the autocorrelation. However, most of the second and third drawbacks remain, even with the new approach. It is an object of the present invention to provide an improved concept for processing an audio signal. 238337 1651425 of 41 This object is achieved by means of an apparatus for encoding an audio signal of claim 1, a method for encoding an audio signal of claim 24, an apparatus for decoding an encoded audio signal of claim 25, a method for decoding an encoded audio signal of claim 40 or a computer program of claim 41. An apparatus for encoding an audio signal comprises a converter for converting the audio signal into a spectral representation. A scale parameter calculator is also provided for calculating a first set of scale parameters from the spectral representation. Furthermore, in order to keep the bit rate as low as possible, the first set of scale parameters is subsampled to obtain a second set of scale parameters, wherein a second number of scale parameters in the second set of scale parameters is less than a first number of scale parameters in the first set of scale parameters.Additionally, a scale parameter encoder is provided to generate a coded representation of the second set of scale parameters, along with a spectral processor to process the spectral representation using a third set of scale parameters. This third set of scale parameters has a third set of scale parameters that is greater than the second set. Specifically, the spectral processor is configured to use either the first set of scale parameters or to derive the third set of scale parameters from the second set of scale parameters or from the coded representation of the second set of scale parameters using an interpolation operation to obtain a representation. 238337 1651425 of 41 encoded spectral representation. An output interface is also provided to generate an encoded output signal comprising information about the encoded representation of the spectral representation and also comprising information about the encoded representation of the second set of scale parameters. The present invention is based on the finding that a low bit rate can be obtained without substantial loss of quality by scaling, on the encoder side, with a larger number of scaling factors and by subsampling the scaling parameters on the encoder side into a second set of scaling parameters or scaling factors, wherein the scaling parameters in the second set that is then encoded and transmitted or stored via an output interface are smaller than the first set of scaling parameters. This results in, on the one hand, fine scaling and, on the other hand, a low bit rate on the encoder side. On the decoder side, the small number of transmitted scale factors is decoded by a scale factor decoder to obtain a first set of scale factors where the number of scale factors or scale parameters in the first set is greater than the number of scale factors or scale parameters in the second set, and then, once again, fine scaling is performed using the higher number of scale parameters on the decoder side within a spectral processor to obtain a finely scaled spectral representation. This results in a low bit rate, on the one hand, and yet a 238337 1651425 of 41 high-quality spectral processing of the audio signal spectrum, on the other hand. Spectral noise modeling, as used in preferred implementations, is implemented using only a very low bit rate. Thus, this spectral noise modeling can be an essential tool even in a low-bit-rate transform-based audio codec. Spectral noise modeling shapes quantization noise in the frequency domain in such a way that the quantization noise is minimally perceived by the human ear, thereby maximizing the perceptual quality of the decoded output signal. Preferred embodiments refer to spectral parameters calculated from amplitude-related measurements, such as the energies of a spectral representation. Specifically, band energies, or more generally, amplitude-related energy measurements, are calculated as the basis for scaling parameters. The bandwidths used to calculate these amplitude-related energy measurements increase from low to high bands to approximate the characteristics of human hearing as closely as possible. Preferably, the division of the spectral representation into bands is performed according to the widely known Bark scale. In further embodiments, the scale parameters are calculated in the linear domain and are calculated particularly for the first set of scale parameters with the high number of scale parameters, and this high number of scale parameters is converted into a domain of the type 238337 1651425 of 41 logarithmic. A logarithmic domain is generally a domain in which small values are expanded and large values are compressed. The subsampling or decimation operation of the scale parameters is then performed in the logarithmic domain, which can be a base-10 logarithmic domain or a base-2 logarithmic domain, the latter being preferred for implementation purposes. The second set of scale factors is then calculated in the logarithmic domain. Preferably, a vector quantization of the second set of scale factors is performed, where the scale factors are in the logarithmic domain. Therefore, the result of the vector quantization indicates the scale parameters of the logarithmic domain.The second set of scaling factors or scaling parameters has, for example, a number of scaling factors that is half the number of scaling factors in the first set, or even a third, or even more preferably, a fourth set. The small number of quantized scaling parameters in the second set of scaling parameters are then passed into the bit stream and subsequently transmitted from the encoder side to the decoder side, or stored as an encoded audio signal along with a quantized spectrum that has also been processed using these parameters, where this processing further involves quantization using an overall gain.Preferably, however, the encoder derives from these second scale factors of the logarithmic domain once again a set of scale factors of the linear domain, which is the third set of scale factors, and the number of scale factors in the third set of scale factors is greater than the number in the second and is preferably even equal to the first number of scale factors in the first set of the first scale factors. Therefore, on the encoder's side, these. 238337 1651425 of 41 interpolated scaling factors are used to process the spectral representation, where the processed spectral representation is finally quantized, and entropically encoded in any way, such as by Huffman coding, arithmetic coding, or vector quantization-based coding, etc. In the decoder that receives an encoded signal with a small number of spectral parameters along with the encoded representation of the spectral signal, the small number of scale parameters is interpolated to a large number of scale parameters. This results in a first set of scale parameters where the number of scale parameters in the second set of scale factors or scale parameters is less than the number of scale parameters in the first set—that is, the set calculated by the scale factor / parameter decoder. Then, a spectral processor located within the device decoding the encoded audio signal processes the decoded spectral representation using this first set of scale parameters to obtain a scaled spectral representation.Then, a converter to convert the scaled spectral representation operates to finally obtain a decoded audio signal that is preferably in the time domain. Further embodiments result in additional advantages as set out below. In preferred embodiments, spectral noise modeling is performed with the help of 16 scaling parameters similar to the scaling factors used in known technique 1. These parameters are obtained in the encoder by first computing the spectrum energy 238337 1651425 of 41 MDCT is performed on 64 non-uniform bands (similar to the 64 non-uniform bands of known technique 3). After applying some processing to the 64 energies (smoothing, pre-emphasis, background noise, logarithmic conversion), the 64 processed energies are then subsampled by a factor of 4 to obtain 16 parameters, which are finally normalized and scaled. These 16 parameters are then quantized using vector quantization (similar to that employed in known technique 2 / 3). The quantized parameters are then interpolated to obtain 64 interpolated scale parameters. These 64 scale parameters are then used to directly model the MDCT spectrum in the 64 non-uniform bands. Similar to known techniques 2 and 3, the scaled MDCT coefficients are then quantized using a scalar quantizer with a step size controlled by an overall gain.In the decoder, inverse scaling is performed on each of the 64 bands, modeling the quantization noise introduced by the scalar quantizer. As in known technique 2 / 3, the preferred embodiment uses only 16+1 parameters as side information, and these parameters can be efficiently encoded with a low number of bits using vector quantization. Consequently, the preferred embodiment has the same advantages as known technique 2 / 3: it requires fewer bits of side information than the approach in known technique 1, which can be a significant difference with low bit rates and / or low latency. As in known technique 3, the preferred embodiment uses nonlinear frequency scaling and therefore does not have the first drawback of known technique 2. 238337 1651425 of 41 Contrary to the known 2 / 3 technique, the preferred implementation does not employ any of the high-complexity LPC-related functions. Comparatively, the required processing functions (smoothing, pre-emphasis, background noise, logarithmic conversion, normalization, scaling, interpolation) entail much lower complexity. Only vector quantization still has relatively high complexity. However, some low-complexity vector quantization techniques can be used with minimal performance loss (multi-split / multi-stage approaches). The preferred implementation, therefore, does not suffer from the second complexity drawback of the known 2 / 3 technique. Unlike the known 2 / 3 technique, the preferred implementation does not rely on a perceptual filter based on LPC. It employs 16 scaling parameters that can be computed very freely. The preferred implementation is more flexible than the known 2 / 3 technique and therefore does not have the third drawback of the known 2 / 3 technique. In conclusion, the preferred embodiment has all the advantages of the known technique 2 / 3 and none of its disadvantages. Preferred embodiments of the present invention are described below in greater detail with reference to the accompanying figures, in which: Fig. 1 is a block diagram of a device for encoding an audio signal; 238337 1651425 of 41 Fig. 2 is a schematic representation of a preferred implementation of the scale factor calculator of Fig. 1; Fig. 3 is a schematic representation of a preferred implementation of the subsampler of Fig. 1; Fig. 4 is a schematic representation of the scale factor encoder of Fig. 4; Figure 5 is a schematic illustration of the spectral processor of the Fig. 1; Fig. 6 illustrates a general representation of an encoder on one side, and a decoder on the other, implementing Spectral Noise Shaping (SNS); Fig. 7 illustrates a more detailed representation of the encoder side on one hand, and the decoder side on the other, where Temporal Noise Shaping (TNS) is implemented together with Spectral Noise Shaping (SNS); Fig. 8 illustrates a block diagram of a device for decoding an encoded audio signal; Figure 9 provides a schematic illustration showing the details of the 238337 1651425 of 41 scale factor decoder, spectral processor and spectral decoder of Fig. 8; Fig. 10 illustrates a subdivision of the spectrum into 64 bands; Fig. 11 illustrates a schematic representation of the subsampling operation on the one hand, and the interpolation operation on the other hand; Fig. 12a illustrates an audio signal in the time domain with overlapping frames; Fig. 12b illustrates an implementation of the converter in Fig. 1; and Fig. 12c illustrates a schematic representation of the converter in Fig. 8. Figure 1 illustrates an apparatus for encoding an audio signal 160. The audio signal 160 is preferably available in the time domain, although other representations of the audio signal, such as a prediction domain or any other domain, would also be useful. The apparatus comprises a converter 100, a scale factor calculator 110, a spectral processor 120, a subsampler 130, a scale factor encoder 140, and an output interface 150. The converter 100 is configured to convert the audio signal 160 into a spectral representation. The scale factor calculator 110 is configured to calculate a first set of scale parameters or scale factors of the spectral representation. 238337 1651425 of 41 Throughout this descriptive document, the terms "scale factor" or "scale parameter" are used to refer to the same parameter or value; that is, a parameter or value that, after some processing, is used to weight a class of spectral values. This weighting, when performed in the linear domain, is actually a multiplication operation by a scaling factor. However, when the weighting is performed in a logarithmic domain, the weighting operation by a scale factor is performed by an actual addition or subtraction operation. Therefore, for the purposes of this application, scaling not only means multiplication or division but also, depending on the domain, addition or subtraction, or generally means any operation by which the spectral value, for example, is weighted or modified using the scale factor or scale parameter. The subsampler 130 is configured to subsample the first set of scale parameters to obtain a second set of scale parameters, where a second number of the scale parameters in the second set of scale parameters is less than a first number of scale parameters in the first set of scale parameters. This is also indicated in the box in Fig. 1, showing that the second number is less than the first number. As illustrated in Fig. 1, the scale factor encoder is configured to generate a coded representation of the second set of scale factors, and this coded representation is sent to the output interface 150.Due to the fact that the second set of scale factors has a smaller number of scale factors than the first set of scale factors, the bit rate for transmitting or storing the encoded representation of the second set of scale factors is lower compared to a situation in which the subsampling of the factors had not been performed. 238337 1651425 of 41 scale performed on subsampler 130. Furthermore, the spectral processor 120 is configured to process the output of the spectral representation by the converter 100 in Fig. 1 using a third set of scaling parameters. The third set of scaling parameters, or scaling factors, has a third number of scaling factors that is greater than the second set of scaling factors. The spectral processor 120 is configured to use, for the purposes of spectral processing, the first set of scaling factors as already available from block 110 via line 171. Alternatively, the spectral processor 120 is configured to use the second set of scaling factors as output by the subsampler 130 for the calculation of the third set of scaling factors, as illustrated by line 172.In a further implementation, the spectral processor 120 uses the encoded representation output from the scale factor / parameter encoder 140 for the purpose of calculating the third set of scale factors, as illustrated by line 173 in Fig. 1. Preferably, the spectral processor 120 does not use the first set of scale factors, but instead uses the second set of scale factors calculated by the subsampler, or even more preferably, it uses the encoded representation, or generally, the second set of quantized scale factors, and then performs an interpolation operation to interpolate the second set of quantized spectral parameters to obtain the third set of scale parameters, which has a higher number of scale parameters due to the interpolation operation. Therefore, the coded representation of the second set of scale factors coming out of block 140 comprises a codebook index for a preferably used scale parameter codebook or a set of 238337 1651425 of 41 corresponding codebook indices. In other embodiments, the coded representation comprises the quantized scale parameters of the quantized scale factors obtained when the codebook index or set of codebook indices, or generally, the coded representation is entered into the decoder side of a vector decoder or any other decoder. Preferably, the spectral processor 120 uses the same set of scaling factors that is also available on the decoder side, i.e., it uses the second set of quantized scaling parameters together with an interpolation operation to finally obtain the third set of scaling factors. In a preferred embodiment, the third number of scale factors in the third set of scale factors is equal to the first number of scale factors. However, a smaller number of scale factors is also useful. For example, 64 scale factors could be derived in block 110, and these 64 scale factors could then be downsampled to 16 scale factors for transmission. Interpolation could then be performed not necessarily to 64 scale factors, but to 32 scale factors in spectral processor 120. Alternatively, interpolation could be performed to an even higher number, such as more than 64 scale factors as needed, provided that the number of scale factors transmitted in the encoded output signal 170 is less than the number of scale factors calculated in block 110 or calculated and used in block 120 of Fig. 1. Preferably, the scale factor calculator 110 is configured to 238337 1651425 of 41 perform various operations illustrated in Fig. 2. These operations relate to the calculation 111 of a band-related amplitude measurement. A preferred band-related amplitude measurement is band energy, but other amplitude-related measurements can also be used, for example, the sum of the magnitudes of the band amplitudes or the sum of the squares of the amplitudes corresponding to the energy. However, in addition to the power of 2 used to calculate band energy, other powers could also be used, such as a power of 3 that would reflect the subjective intensity of the signal, and even powers other than integers, such as powers of 1.5 or 2.5, can also be used to calculate band-related amplitude measurements.Even powers less than 1.0 can be used as long as it is ensured that the values processed by such powers are positively valued. An additional operation performed by the scaling factor calculator can be interband smoothing 112. This interband smoothing is preferably used to smooth out any instabilities that may appear in the vector of amplitude-related measurements obtained in step 111. If this smoothing were not performed, these instabilities would be amplified when subsequently converted to a logarithmic domain as illustrated in 115, especially at spectral values where the energy is close to 0. However, in other embodiments, interband smoothing is not performed. An additional preferred operation performed by the scale factor calculator 110 is the pre-emphasis operation 113. This pre-emphasis operation has a similar purpose to that of a pre-emphasis operation used in a filter. 238337 1651425 of 41 perceptual based on the LPCs of the MDCT-based TCX processing as previously discussed with respect to the known technique. This procedure increases the amplitude of the modeled spectrum at low frequencies, resulting in a reduction of quantization noise at low frequencies. However, depending on the implementation, the pre-emphasis operation - as well as the other specific operations - does not necessarily have to be carried out. An additional optional processing operation is background noise summation processing 114.This procedure improves the quality of signals containing very high spectral dynamics, such as the glockenspiel, by limiting the amplification of the amplitude of the modeled spectrum in the valleys. This has the indirect effect of reducing quantization noise at the peaks, at the cost of increased quantization noise in the valleys. In any case, the quantization noise is not perceptible due to the masking properties of the human ear, such as the absolute threshold of hearing, pre-masking, post-masking, or the general masking threshold. This generally indicates that a low-intensity tone relatively close in frequency to a high-intensity tone is not perceptible at all; that is, it is completely masked or only barely perceived by the human hearing mechanism, so this spectral contribution can be roughly quantized. However, the background noise summation operation 114 does not necessarily have to be performed. 238337 1651425 of 41 Block 115 also indicates a logarithmic domain conversion. Preferably, a transformation of an output from one of blocks 111, 112, 113, or 114 in Figure 2 is performed in a logarithmic domain. A logarithmic domain is one in which values close to zero are expanded and large values are compressed. Preferably, the logarithmic domain has a base of 2, but other logarithmic domains can also be used. However, a logarithmic domain with a base of 2 is better suited for implementation in a fixed-point signal processor. The output of the scale factor calculator 110 is a first set of scale factors. As illustrated in Fig. 2, each of blocks 112 to 115 can be bypassed; that is, the output of block 111, for example, could already be the first set of scale factors. However, all processing operations, and particularly the conversion to the logarithmic domain, are preferred. Therefore, the scale factor calculator could still be implemented by performing only steps 111 and 115 without the procedures in steps 112 to 114, for example. Therefore, the scale factor calculator is configured to perform one, two, or more of the procedures illustrated in Fig. 2 as indicated by the input / output lines connecting various blocks. Figure 3 illustrates a preferred implementation of the subsampler 130 of Figure 1. Preferably, low-pass filtering or, more generally, high-pass filtering is performed. 238337 1651425 of 41 with a certain window w(k) in step 131, and then a subsampling / decimation operation is performed on the filtered result. Because the low-pass filtering 131 and, in preferred embodiments, the subsampling / decimation operation 132 are arithmetic operations, filtering 131 and subsampling 132 can be performed within a single operation, as will be noted later. Preferably, the subsampling / decimation operation is performed such that an overlap is made between the individual groups of scale parameters of the first set of scale parameters. Preferably, a superposition of a scale factor is made in the filtering operation between two calculated decimated parameters. Therefore, step 131 performs low-pass filtering on the vector of scale parameters before decimation. This low-pass filter has an effect similar to that of the dispersion function used in psychoacoustic models.This filter reduces quantization noise at the peaks, at the cost of increased quantization noise around the peaks where it is already perceptually masked at least to a greater degree than the quantization noise at the peaks. Furthermore, the subsampler also performs a mean removal step 133 and an additional scaling step 134. However, the low-pass filtering operation 131, the mean removal step 133, and the scaling step 134 are only optional steps. Therefore, the subsampler illustrated in Fig. 3 or Fig. 1 can be implemented to perform only step 132 or to perform two of the steps illustrated in Fig. 3, such as step 132 and one of steps 131, 133, or 134. Alternatively, the subsampler can perform all four steps or only three of the four steps illustrated in Fig. 3, provided that the subsampling / decimation operation 132 is carried out. 238337 1651425 of 41 corporal. As noted in Fig. 3, the audio operations in Fig. 3 performed by the subsampler are carried out in the logarithmic domain in order to obtain better results. Figure 4 illustrates a preferred implementation of the scale factor encoder 140. The scale factor encoder 140 receives the second set of scale factors, preferably in the logarithmic domain, and performs vector quantization as illustrated in block 141 to ultimately produce one or more indices per frame. These one or more indices per frame can be sent to the output interface and written to the bitstream, i.e., introduced into the encoded output audio signal 170 by means of any available output interface procedure. Preferably, the vector quantizer 141 additionally produces the second set of scale factors from the quantized logarithmic domain. Therefore, this information can be output directly from block 141, as indicated by arrow 144.However, alternatively, the decoder codebook 142 is also available separately in the encoder. This decoder codebook receives one or more frame indices and derives from these one or more frame indices the second set of quantized logarithmic domain scaling factors, preferably as indicated in line 145. In typical implementations, the decoder codebook 142 will be integrated within the vector quantizer 141. Preferably, the vector quantizer 141 is a multi-stage or level vector quantizer, or a combination of a multi-stage / level vector quantizer, as used, for example, in any of the procedures of the known art. 238337 1651425 of 41 indicated. Therefore, it is ensured that the second set of scaling factors is the same second set of quantized scaling factors that is also available on the decoder side, i.e., in the decoder that only receives the encoded audio signal that has the single or multiple per-frame indices produced by block 141 via line 146. Figure 5 illustrates a preferred implementation of the spectral processor. The spectral processor 120, included within the encoder in Figure 1, comprises an interpolator 121 that receives the second set of quantized scale parameters and produces the third set of scale parameters, where the third number is greater than the second number and preferably equal to the first number. The spectral processor also comprises a linear domain converter 120. Spectral modeling is then performed in block 123, using the linear scale parameters on the one hand and the spectral representation obtained from the converter 100 on the other. Preferably, a subsequent time-noise modeling operation is performed, i.e., a prediction on the frequency, to obtain spectral residual values at the output of block 124, while the TNS side information is sent to the output interface as indicated by arrow 129. Finally, the spectral processor 125 has a scalar quantizer / encoder configured to receive a single global gain for the entire spectral representation, i.e., for a complete frame. Preferably, the global gain is derived based on certain bit rate considerations. Therefore, the global gain is set such that the representation 238337 The 1651425 of 41 encoded spectral representation generated by block 125 satisfies certain requirements, such as a bit rate requirement, a quality requirement, or both. The overall gain can be calculated iteratively or from a pre-fed measurement, as appropriate. Generally, the overall gain is used in conjunction with a quantizer, and a high overall gain typically results in coarser quantization, while a lower overall gain results in finer quantization. Therefore, in other words, a higher overall gain results in a larger quantization step size, while a lower overall gain results in a smaller quantization step size when using a fixed quantizer.However, other quantizers can be used in conjunction with the global gain functionality, such as a quantizer with some kind of compression functionality for high values—that is, some kind of nonlinear compression functionality—so that, for example, higher values are more compressed than lower values. The noted dependence between the global gain and the quantization approximation holds true when the global gain is multiplied by the pre-quantization values in the linear domain, corresponding to a sum in the logarithmic domain. However, if the global gain is applied by division in the linear domain, or by subtraction in the logarithmic domain, the dependence is reversed. The same is true when the global gain represents an inverse value. The following are preferred implementations of the individual procedures described with respect to Fig. 1 to Fig. 5. Detailed step-by-step description of preferred realizations: 238337 1651425 of 41 ENCODER: • Step 1: Energy per band (111) The band energies EB(n) are calculated as follows: / 7id(b+1)-1 EB(b) = > ,,, .--- for b = 0 ...¾ - 1 Ind(b + 1) — Ind(b) k=M(b) where X(k) are the MDCT coefficients, NB= 64 is the number of bands and lnd(n) are the band indices. The bands are not uniform and follow the perceptually relevant Bark scale (lower at low frequencies, higher at high frequencies). • Step 2: Smoothing (112) The band energy EB(b) is smoothed using '0.75 E^O) + 0.25 Eb(1) Es(b) = 0.25 Εβ(62) + 0.75 Εβ(63) ;0.25 EB(b - 1) + 0.5 EB(b) + 0.25 EB(b + 1) ,yes b = 0 ,yes b = 63 , yes no Note: This step is primarily used to smooth out any instabilities that may appear in the EB(b) vector. If not smoothed, these instabilities are amplified when converted to the logarithmic domain (see step 5), especially in the valleys where the energy is close to 0. • Step 3: Pre-emphasis (113) The smoothed band energy Es(b') is then pre-emphasized using EP(b) = Es(b) 10 10-63 for b = 0. .63con9tnt controls the pre-emphasis tilt and depends on the frequency of 238337 1651425 of 41 sampling. It is, for example, 18 at 16kHz and 30 at 48kHz. The pre-emphasis used in this step has the same purpose as the pre-emphasis used in the LPC-based perceptual filter of the known technique 2, increasing the amplitude of the modeled spectrum at low frequencies, resulting in a reduction of quantization noise at low frequencies. • Step 4: Background noise (114) A background noise of -40dB is added to Ep(b) using EP(b) = max(Ep(b),backgroundnoise') for b = 0..63, the background noise being calculated by backgroundnoise = max 'Σ^ΕΡ(^ 10, 2“32 This step improves the quality of signals containing very high spectral dynamics such as, for example, the glockenspiel, by limiting the amplification of the amplitude of the modeled spectrum in the valleys, which has the indirect effect of reducing quantization noise at the peaks, at the cost of an increase in quantization noise in the valleys where it is not noticeable anyway. • Step 5: Logarithm (115) Then a transformation is performed in the logarithmic domain using „ ... log2(^(^)) ,n_ EL(b) =--------- for b — 0..63 ¿i • Step 6: Subsampling (131, 132) 238337 1651425 of 41 The vector EL(b) is then undersampled by a factor of 4 using , 5 w(0)El(0) + y , w(k) EL(4b + k — 1) k=l 4 , if b — 0 £4(¿) = iw(k) EL(4b + k — 1) + w(5)El(63) k=0 5 , if b = 15 w(k) EL(4b + k — 1) ^=0 , if not With _ ( 1 2 3 3 2 1 jw(k) = 1— — — — — —ív 7112 12 12 12 12 12J This step applies a low-pass filter (w(k)) to the EL(b) vector before decimation. This low-pass filter has an effect similar to that of the dispersion function used in psychoacoustic models: it reduces quantization noise at the peaks, at the cost of increased quantization noise around the peaks where it is already perceptually masked. • Step 7: Mean removal and scaling (133, 134) The final scaling factors are obtained after removing the mean and scaling by a factor of 0.85 « Λ nor ( λn_ scp{n) = 0.85 E4(n)--—---- for n = 0. .15 \ 16 / Since the codec has additional global gain, the average can 238337 1651425 of 41 remove without any loss of information. Removing the mean also allows for more efficient vector quantization. The scaling of 0.85 slightly compresses the amplitude of the noise modeling curve. This has a perceptual effect similar to that of the dispersion function mentioned in Step 6: reduced quantization noise at the peaks and increased quantization noise at the valleys. • Step 8: Quantization (141, 142) Scale factors are quantized using vector quantization, producing indices that are then packed into the bit stream and sent to the decoder, and quantized scale factors scfQ(n). • Step 9: Interpolation (121, 122) The quantized scale factors scfQ(n) are interpolated using scfQint(0) = scfQ(0) scfQint(l) = scfQ(0) scfQint(4n + 2) = scfQ(n) + - (scfQ(n + 1) — scfQ(n)) for n = 0..14 z scfQint(4n + 3) = scfQ(n) + - (scfQ(n + 1) — scfQ(n)) for n = 0..14 z scfQint(4n + 4) = scfQ(n) + - (scfQ(n + 1) — scfQ(n)) for η = 0..14 scfQint(4n + 5) = scfQ(n) + - (scfQ(n + 1) — scfQ(n)) for η = 0..14 scfQint(62) = sc / Q(15) +-(sc / Q(15) - sefQ (14)) scfQint(63) = scfQ(15) + -(sc / Q(15) - sefQ (14)) 238337 1651425 of 41 and are transformed back to the linear domain using gSNS(b) - 2sc™int(b)for b = 0..63 Interpolation is used to obtain a smoothed noise modeling curve and thus avoid any large amplitude jumps between adjacent bands. • Step 10: Spectral Modeling (123) The SNS scaling factors gSNsW are applied to the MDCT frequency lines for each band separately in order to generate the modeled spectrum Xs(k) Xs(k) =--( 2 for k= Ind(b).. lnd(b + 1) — 1, for b= 0..63 PsnsW Figure 8 illustrates a preferred implementation of an apparatus for decoding an encoded audio signal 250 comprising information about an encoded spectral representation and information about an encoded representation of a second set of scaling parameters. The decoder comprises an input interface 200, a spectral decoder 210, a scale factor / parameter decoder 220, a spectral processor 230, and a converter 240. The input interface 200 is configured to receive the encoded audio signal 250 and to extract the encoded spectral representation, which is sent to the spectral decoder 210, and to extract the encoded representation of the second set of scaling factors, which is sent to the scale factor decoder 220. Likewise, the spectral decoder 210 is configured to decode the encoded spectral representation to obtain a decoded spectral representation, which is sent 238337 1651425 of 41 to the spectral processor. The scale factor decoder 220 is configured to decode the second set of encoded scale parameters to obtain a first set of scale parameters sent to the spectral processor 230. The first set of scale factors has a greater number of scale factors or parameters than the second set. The spectral processor 230 is configured to process the decoded spectral representation using the first set of scale parameters to obtain a scaled spectral representation. The scaled spectral representation is then converted by the converter 240 to finally obtain the decoded audio signal 260. Preferably, the scale factor decoder 220 is configured to operate substantially in the same manner as described with respect to the spectral processor 120 of Fig. 1 in relation to the calculation of the third set of scale factors or scale parameters as described in connection with blocks 141 or 142 and, in particular, with respect to blocks 121, 122 of Fig. 5. In particular, the scale factor decoder is configured to perform substantially the same procedure for interpolation and transformation back into the linear domain as described above with respect to step 9. Therefore, as illustrated in Fig. 9, the scale factor decoder 220 is configured to apply a decoder codebook 221 to the single or multiple indices per frame that represent the encoded scale parameter representation.Next, an interpolation is performed in block 222 which is substantially the same interpolation mentioned with respect to block 121 in Fig. 5. Then, a linear domain converter 223 is used which is substantially the same converter of. 238337 1651425 of 41 linear domain 122 mentioned with respect to Fig. 5. However, in other implementations, blocks 221, 222, 223 may operate differently than described with respect to the corresponding blocks on the encoder side. Likewise, the spectral decoder 210 illustrated in Fig. 8 comprises a quantizer / decoder block that receives, as an input, the encoded spectrum and produces a dequantized spectrum that is preferably dequantized using the overall gain that is additionally transmitted from the encoder side to the decoder side within the encoded audio signal. The dequantizer / decoder 210 may, for example, comprise an arithmetic or Huffman decoder functionality that receives, as an input, some kind of code and produces quantization indices that represent spectral values.These quantization indices then enter a dequantizer along with the overall gain, and the output are dequantized spectral values that can then be subjected to TNS processing, such as an inverse frequency prediction, in a TNS decoder processing block 211, which is optional. Specifically, the TNS decoder processing block also receives the TNS side information generated in block 124 of Fig. 5, as indicated by line 129. The output of the TNS decoder processing step 211 enters the spectral modeling block 212, where the first set of scaling factors calculated by the scaling factor decoder is applied to the decoded spectral representation, which may or may not have been TNS processed, and the output is the scaled spectral representation that then enters the converter 240 of Fig. 8. 238337 1651425 of 41 The following are additional procedures for preferred embodiments of the decoder. DECODER: • Step 1: Quantization (221) The vector quantizer indices produced in the encoder step 8 are read from the bit stream and used to decode the quantized scaling factors scfQ(n). • Step 2: Interpolation (222, 223) Same as encoder step 9. • Step 3: Spectral Modeling (212) The SNS scaling factors gSNS(b) are applied to the MDCT frequency lines quantized for each band separately in order to generate the decoded spectrum X(k) as outlined by the following code. X(k) = Xs(k) · gSNS(b) for k= ¡ηά^).. Ind(b + 1) — 1, for b= 0. .63 Figures 6 and 7 illustrate a general encoder / decoder configuration, where Figure 6 represents an implementation without TNS processing, while Figure 7 illustrates an implementation that includes TNS processing. Similar functionalities illustrated in Figures 6 and 7 correspond to similar functionalities in the other figures when identical part numbers are indicated. In particular, as illustrated in Figure 6, the input signal 160 enters a transformation stage 110 and is then... 238337 1651425 of 41 performs spectral processing 120. Specifically, the spectral processing is carried out by an SNS encoder indicated by reference numbers 123, 110, 130, and 140, which indicate that the SNS encoder block implements the functionalities indicated by these reference numbers. Following the SNS encoder block, a quantization operation 125 is performed, and the encoded signal enters the bit stream as indicated in 180 in Fig. 6. The bit stream 180 then reaches the decoder side, and following inverse quantization and decoding illustrated by reference number 210, the SNS decoder operation illustrated by blocks 210, 220, and 230 in Fig. 8 is performed, so that, finally, following an inverse transform 240, the decoded output signal 260 is obtained. Fig. 7 illustrates a representation similar to that in Fig. 6, but indicates that, preferably, TNS processing is performed after SNS processing on the encoder side and, correspondingly, TNS processing 211 is performed before SNS processing 212 with respect to the processing sequence on the decoder side. Preferably, the additional TNS tool is used between Spectral Noise Modeling (SNS) and quantization / encoding (see the block diagram below). TNS (Time Noise Modeling) also models quantization noise but performs time-domain modeling (as opposed to the frequency-domain modeling of SNS). TNS is useful for signals containing sharp attacks and for speech signals. TNS is usually applied (for example in AAC) between the transform and SNS. 238337 1651425 of 41 However, TNS is preferably applied to the modeled spectrum. This avoids some defects that were produced by the TNS decoder when operating the codec at low bit rates. Figure 10 illustrates a preferred subdivision of the spectral coefficients or spectral lines obtained by block 100 on the band-encoder side. Specifically, it indicates that the lower bands have a smaller number of spectral lines than the higher bands. Specifically, the X-axis in Fig. 10 corresponds to the band index and illustrates the preferred embodiment of 64 bands, and the Y-axis corresponds to the spectral line index, illustrating 320 spectral coefficients in a frame. Fig. 10 particularly illustrates the situation in the case of super wideband (SWB) with a sampling frequency of 32 kHz. In the case of super broadband, the situation with respect to the individual bands is such that a frame results in 160 spectral lines and the sampling frequency is 16 kHz so that, for both cases, a frame has a length in time of 10 milliseconds. Fig. 11 illustrates further details of the preferred subsampling performed in the subsampler 130 of Fig. 1 or the corresponding oversampling or interpolation performed in the scale factor decoder 220 of Fig. 8 or as illustrated in block 222 of Fig. 9. 238337 1651425 of 41 Along the X-axis, the index for bands 0 to 63 is provided. Specifically, there are 64 bands ranging from 0 to 63. The 16 subsampling points corresponding to scfQ(i) are illustrated as vertical lines 1100. In particular, Fig. 11 illustrates how a certain grouping of scale parameters is carried out to finally obtain the subsampled point 1100. For example, the first block of four bands consists of (0, 1, 2, 3) and the midpoint of this first block is at 1.5 indicated by item 1100 at index 1.5 on the X-axis. Accordingly, the second block of four bands is (4, 5, 6, 7), and the midpoint of the second block is 5,5. The 1110 windows correspond to the w(k) windows indicated with respect to the subsampling in step 6 described above. It can be observed that these windows are centered on the subsampled points and there is an overlap of one block on each side, as noted previously. The interpolation step 222 in Fig. 9 recovers the 64 bands from the 16 subsampled points. This is shown in Fig. 11 when the position of any of the lines 1120 is computed as a function of the two subsampled points indicated in 1100 around a given line 1120. The example below illustrates this. The position of the second band is calculated as a function of the two vertical lines around it (1.5 and 5.5): 2=1.5+1 / 8x(5.5-1.5). 238337 1651425 of 41 Correspondingly, the position of the third band as a function of the two vertical lines 1100 around it (1.5 and 5.5): 3=1.5+3 / 8x(5.5-1.5). A specific procedure is performed for the first two bands and the last two bands. For these bands, interpolation cannot be performed because there would be no vertical lines or values corresponding to the 1100 vertical lines outside the range of 0 to 63. Therefore, in order to solve this problem, extrapolation is performed as described with respect to step 9: interpolation as outlined above for the two bands 0 and 1 on the one hand, and 62 and 63 on the other. The following discusses a preferred implementation of the converter 100 of Fig. 1 on the one hand and the converter 240 of Fig. 8 on the other hand. Specifically, Fig. 12a illustrates a schematic to indicate the lattice made on the encoder side within converter 100. Fig. 12b illustrates a preferred implementation of converter 100 of Fig. 1 on the encoder side, and Fig. 12c illustrates a preferred implementation of converter 240 on the decoder side. The encoder-side converter 100 is preferably implemented to perform framing with overlapping frames, such as a 50% overlap where frame 2 overlaps with frame 1, and frame 3 overlaps with both frame 2 and frame 4. However, other overlaps or non-overlapping processing are also possible, but a 50% overlap in conjunction with an MDCT algorithm is preferred. To this end, converter 100 comprises an analysis window 101 and a connected spectral converter. 238337 1651425 of 41 subsequently 102 to perform FFT (Fast Fourier Transform) processing, MDCT processing, or any other type of time-to-spectrum conversion to obtain a frame sequence corresponding to a sequence of spectral representations as input in Fig. 1 to the blocks following converter 100. Correspondingly, one or more spectral representations are fed into converter 240 of Fig. 8. In particular, the converter comprises a time converter 241 that implements an inverse FFT operation, an inverse MDCT operation, or a corresponding time-to-spectrum conversion operation. The output is fed into a synthesis window 242, and the output of the synthesis window 242 is fed into an overlap-sum processor 243 to perform an overlap-sum operation to finally obtain the decoded audio signal.In particular, the overlapsum processing in block 243, for example, performs a sample-by-sample sum between the corresponding samples of the second half of, for example, frame 3 and the first half of frame 4 so that the audio sampling values are obtained for the overlap between frame 3 and frame 4 as indicated in item 1200 in Fig. 12a. Similar overlapsum operations are performed in a sample-by-sample manner to obtain the remaining audio sampling values of the decoded audio output signal. An audio signal encoded with the invention can be stored on a digital storage medium or a non-transient storage medium or can be transmitted by a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet. Although some aspects have been described in the context of a device, it is 238337 1651425 of 41 Clearly, these aspects also represent a description of the corresponding method, where a block or device corresponds to a step of a method or a feature of a step of a method. Similarly, the aspects described in the context of a step of a method also represent a description of a corresponding block, item, or feature of a device. Depending on certain implementation requirements, the embodiments of the invention can be implemented in hardware or software. Implementation can be carried out using a digital storage medium, for example, a floppy disk, a DVD, a CD, a ROM (Read Only Memory), PROM (Programmable ROM), EPROM (Erasable PROM), EEPROM (Electronically Erasable EPROM), or a FLASH memory card, which has electronically readable control signals stored therein, which cooperate (or are capable of cooperating) with a programmable computer system in such a way that the respective method is carried out. Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, in such a way that one of the methods described herein is carried out. Generally, embodiments of the present invention can be implemented as a computer program product with program code, the 238337 1651425 of 41 Program code is operational to carry out one of the methods when the product of a computer program is executed on a computer. The program code can, for example, be stored on a machine-readable medium. Other embodiments include the computer program for carrying out one of the methods described herein, stored on a machine-readable carrier or on a non-transient storage medium. In other words, an embodiment of the method of the invention is, therefore, a computer program that has program code to carry out one of the methods described herein, when the program is executed on a computer. A further embodiment of the method of the invention is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for carrying out one of the methods described herein. A further embodiment of the method of the invention is, therefore, a data stream or a sequence of signals representing the computer program for carrying out one of the methods described herein. The data stream or the sequence of signals may, for example, be configured to be transferred via a data communication connection, e.g., via the Internet. A further embodiment comprises a processing means, for example, 238337 1651425 of 41 a computer, or a programmable logic device, configured or adapted to carry out one of the methods described herein. A further embodiment comprises a computer having installed therein the computer program to carry out one of the methods described herein. In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably implemented using any hardware device. The embodiments described above are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will become apparent to others skilled in the art. Therefore, the intention is to limit the invention only to the scope of the patent claims to be granted and not to the specific details presented by way of description and explanation of the embodiments herein. 238337 1651425 of 41 Literature [1] ISO / IEC 14496-3:2001; Information technology - Coding of audio-visual objects - Part 3: Audio. [2] 3GPP TS 26.403; General audio codec audio processing functions; Enhanced 5 aacPlus general audio codec; Encoder specification; Advanced Audio Coding (AAC) part. [3] ISO / IEC 23003-3; Information technology - MPEG audio technologies - Part 3: Unified speech and audio coding. [4] 3GPP TS 26.445; Codec for Enhanced Voice Services (EVS); Detailed 10 algorithmic description.
Claims
1. An apparatus for decoding an encoded audio signal comprising information about an encoded spectral representation and information about an encoded representation of a second set of scale parameters, comprising: an input interface (200) for receiving the encoded audio signal and extracting the encoded spectral representation and the encoded representation of the second set of scale parameters; a spectral decoder (210) for decoding the encoded spectral representation to obtain a decoded spectral representation; the apparatus being characterized in that it further comprises: a scale parameter decoder (220) for decoding the second set of encoded scale parameters to obtain a first set of scale parameters,where a number of scale parameters from the second set is less than a number of scale parameters from the first set; a spectral processor (230) for processing the decoded spectral representation using the first set of scale parameters to obtain a scaled spectral representation; and a converter (240) for converting the scaled spectral representation to obtain a decoded audio signal, wherein the scale parameter decoder (220) is configured to determine an interpolated scale parameter based on a quantized scale parameter and a difference between the quantized scale parameter and a subsequent quantized scale parameter in an ascending sequence of frequency-quantized scale parameters,or wherein the spectral processor (230) is configured to apply (211) a time-noise modeling (TNS) decoder operation to the decoded spectral representation to obtain a TNS decoded spectral representation, and to weight (212) the TNS decoded spectral representation using the first set of scale parameters, or wherein the scale parameter decoder (220) is configured to interpolate quantized scale parameters such that the interpolated quantized scale parameters have values within a range of ±20% of the values obtained using the following equations: (FORMULA) wherein scfQ(n) is the quantized scale parameter for an index n, and wherein scfQint(k) is the interpolated scale parameter for an index k, or wherein the scale parameter decoder (220) is configured to perform an interpolation (222) to obtain scale parameters within,with respect to frequency, of the first set of scale parameters and to perform an extrapolation operation to obtain scale parameters at the edges, with respect to frequency, of the first set of scale parameters. Thirteen claims follow.