Audio encoder and decoder using frequency domain and time domain processors with full band gap filling
The combined time-domain and frequency-domain encoding/decoding processor with intelligent gap-filling addresses bandwidth limitations, ensuring high-quality audio reproduction and efficient bit rate usage by encoding the entire audio spectrum and restoring spectral gaps.
Patent Information
- Application Number
- JP2025169316
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-07-28
- Filing Date
- 2025-10-07
- Publication Date
- 2026-01-21
AI Technical Summary
Existing audio coding methods degrade audio quality due to band-limitation and bandwidth extension techniques, particularly for signals with significant high-frequency content, leading to reduced accuracy and inflexibility.
A combined time-domain and frequency-domain encoding/decoding processor with a gap-filling function that operates over the entire audio signal bandwidth, allowing precise encoding and decoding of the entire audio spectrum, and intelligent gap-filling to restore spectral gaps using parametric data and frequency tiles.
This approach enables high-quality audio reproduction by encoding the entire audio domain with high resolution, minimizing perceptual disruption at low bit rates and ensuring seamless switching between coding schemes, thus improving audio quality and efficiency.
Smart Images

Figure 2026010016000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to audio signal encoding and decoding, and in particular to audio signal processing using parallel frequency domain and time domain encoder / decoder processors. [Background technology]
[0002] Perceptual coding of audio signals for the purpose of data reduction for efficient storage or transmission is a widely used task. In particular, when a minimum bit rate is to be achieved, the coding used results in a degradation of audio quality, which is mainly caused by the limitation of the audio signal bandwidth to be transmitted at the coding side. In this case, the audio signal is typically low-pass filtered so that no spectral waveform content remains above a certain predetermined cutoff frequency.
[0003] In modern codecs, there are known methods for decoder-side signal restoration via audio signal bandwidth extension (BWE), such as Spectral Bandwidth Replication (SBR), which operates in the frequency domain, or the so-called Time Domain Bandwidth Extension (TD-BWE), which is a post-processor within the speech coder operating in the time domain.
[0004] In addition, there are several joint time-domain / frequency-domain coding concepts known by terms such as AMR-WB+ or USAC.
[0005] The common feature of these joint time domain / frequency domain coding concepts is that the frequency domain coder relies on a bandwidth extension technique, which brings about a band limit on the input audio signal, and the part above the crossover or boundary frequency is coded with a lower resolution coding concept and then synthesized at the decoder side. Therefore, such concepts mainly rely on the technique of a pre-processor at the encoder side and the corresponding post-processing function at the decoder side.
[0006] Typically, a time domain coder is selected for useful signals to be coded in the time domain, such as speech signals, and a frequency domain coder is selected for non-speech signals, musical sounds, etc. However, for non-speech signals that have significant harmonics, especially in the high frequency band, prior art frequency domain coders result in reduced accuracy and therefore degraded audio quality, because such significant harmonics can only be parametrically coded separately or completely excluded in the coding / decoding process.
[0007] Furthermore, there is also the concept that the time domain coding / decoding branch may further rely on bandwidth extension, where the high frequency region is parametrically coded while the low frequency region is typically coded using ACELP or any other CELP related coder, e.g., a speech coder. Such bandwidth extension feature increases bit rate efficiency, but on the other hand introduces further inflexibility, since both coding branches, i.e., the frequency domain coding branch and the time domain coding branch, are band-limited due to the bandwidth extension or spectral band replication process operating above a predetermined crossover frequency that is substantially lower than the maximum frequency contained in the input audio signal.
[0008] Relevant items in the current state of the art include: -SBR as a post-processing unit for waveform decoding (Non-Patent Documents 1-3) -MPEG-D USAC Core Switching (Non-Patent Document 4) -MPEG-H 3D IGF (Patent Document 1)
[0009] The following publications and patent documents disclose methods that are believed to constitute prior art to the present application:
[0010] MPEG-D USAC describes a switchable core encoder. However, in USAC, the band-limited core is always constrained to carry a low-pass filtered signal. Therefore, certain musical signals with significant high-frequency content, such as full-band sweeps and triangle tones, cannot be faithfully reproduced. [Prior art documents] [Patent documents]
[0011] [Patent Document 1] [5] PCT / EP2014 / 065109 [Non-patent literature]
[0012] [Non-Patent Document 1] [1] M. Dietz, L. Liljeryd, K. Kjoerling and O. Kunz, “Spectral Band Replication, a novel approach in audio coding,” in 112th AES Convention, Munich, Germany, 2002. [Non-patent document 2] [2] S. Meltzer, R. Boehm and F. Henn, “SBR enhanced audio codecs for digital broadcasting such as “Digital Radio Mondiale”(DRM),” in 112th AES Convention, Munich, Germany, 2002. [Non-patent document 3] [3] T. Ziegler, A. Ehret, P. Ekstrand and M. Lutzky, “Enhancing mp3 with SBR: Features and Capabilities of the new mp3PRO Algorithm,” in 112th AES Convention, Munich, Germany, 2002. [Non-patent document 4] [4] MPEG-D USAC Standard Summary of the Invention [Problem to be solved by the invention]
[0013] The object of the present invention is to provide an improved concept for audio coding.
[0014] This object is achieved by an audio encoder according to claim 1, an audio decoder according to claim 11, an audio encoding method according to claim 20, an audio decoding method according to claim 21 or a computer program according to claims 22 and 23.
[0015] The present invention is based on the following finding: a time-domain encoding / decoding processor can be combined with a frequency-domain encoding / decoding processor having a gap-filling function, but this gap-filling function for filling spectral holes operates over the entire bandwidth of the audio signal, or at least above a predetermined gap-filling frequency. What is important is that the frequency-domain encoding / decoding processor is in a position to perform encoding / decoding of precise waveforms or spectral values up to the maximum frequency, and not just up to the crossover frequency. Furthermore, the ability of a frequency-domain coder to encode the entire bandwidth with high resolution makes it possible to integrate the gap-filling function within the frequency-domain coder.
[0016] Thus, according to the present invention, by using a full-spectrum encoder / decoder processor, the challenges associated with separating bandwidth extension on the one hand and core coding on the other hand can be addressed and overcome by performing bandwidth extension in the same spectral domain in which the core decoder operates. Therefore, a full-rate core decoder is provided that encodes and decodes the entire audio signal domain. This does not require an encoder-side downsampler and a decoder-side upsampler. Instead, the entire processing is performed in the full sampling rate or full bandwidth domain. To achieve high coding gain, the audio signal is analyzed to find a first set of first spectral portions to be coded with high resolution. In one embodiment, this first set of first spectral portions may include the tonal portions of the audio signal. Meanwhile, the non-tonal or noisy components of the audio signal, constituting the second set of second spectral portions, are parametrically coded with lower spectral resolution. The coded audio signal then simply requires the first set of first spectral portions coded in a waveform-preserving manner with high spectral resolution, and the second set of second spectral portions coded with lower resolution parametrically using additional frequency "tiles" originating from the first set. On the decoder side, the core decoder, a full-band decoder, reconstructs a first set of first spectral portions in a waveform-preserving manner, i.e., without knowledge of whether there is any additional frequency regeneration. However, the spectrum so generated has many spectral gaps. These gaps are subsequently filled using the inventive intelligent gap-filling (IGF) technique, which uses, on the one hand, frequency regeneration applying parametric data, and, on the other hand, the source spectral region, i.e., the first spectral portions reconstructed by the full-rate audio decoder.
[0017] In a further embodiment, the spectral portions restored by only noise filling and not by bandwidth replication or frequency tile filling constitute a third set of third spectral portions. Due to the fact that the coding concept operates in a single domain, with core encoding / decoding on the one hand and frequency regeneration on the other, the IGF is not limited to filling high frequency regions but can also fill low frequency regions, which is achieved either by noise filling without frequency regeneration or by frequency regeneration using one frequency tile for different frequency regions.
[0018] Furthermore, it should be emphasized here that the information on spectral energy, individual energy, enduring energy, tile energy or loss energy may include not only energy values but also (e.g. absolute) amplitude values, level values or any other values from which a final energy value can be derived. Thus, the information on energy may include, for example, the energy values themselves and / or level values and / or amplitude values and / or absolute amplitude values.
[0019] A further aspect is based on the finding that the correlation state is not only important for the source region but also for the target region. Furthermore, the present invention recognizes that different correlation states may occur in the source and target regions. For example, when considering a speech signal with high-frequency noise, the state may be such that the low-frequency band containing the speech signal with a small number of overtones is highly correlated in the left and right channels when the loudspeaker is centrally located. However, the high-frequency portion may be strongly decorrelated due to the fact that there may be a different high-frequency noise on the left side compared to the possibility of a different high-frequency noise or no high-frequency noise on the right side. Therefore, if a simple gap-filling operation is performed that ignores this state, the high-frequency portion may also be correlated, which may result in severe spatial separation artifacts in the restored signal. To address this issue, parametric data for the reconstruction band, or more generally for a second set of second spectral portions to be reconstructed using the first set of first spectral portions, is calculated to identify either a first or a second different two-channel representation for the second spectral portion, or in other words, for the reconstruction band. On the encoder side, a two-channel identification is calculated for the second spectral portion, i.e., for which reconstruction band energy information is also calculated. On the decoder side, a frequency regeneration unit then regenerates the second spectral portion, which regeneration depends on the first portion, i.e., the source region, of the first set of first spectral portions and the parametric data for the second portion, such as spectral envelope energy information or any other spectral envelope data, as well as on the two-channel identification for the second portion, i.e., this reconstruction band under consideration.
[0020] The two-channel identification is preferably transmitted as one flag for each restoration band, and this data is transmitted from the encoder to the decoder, which then decodes the core signal as indicated by the suitably calculated flag for the core band. Then, in one embodiment, the core signal is stored into both (e.g., left / right and center / side) stereo representations, and for IGF frequency tile filling, a source tile representation is selected for the intelligent gap filling or restoration band, i.e., target region, that matches the target tile representation as indicated by the two-channel identification flag.
[0021] It should be emphasized here that this processing is not only useful for stereo signals, i.e., left and right channels, but also works for multi-channel signals. In the case of multi-channel signals, multiple pairs of different channels can be processed, for example, the left and right channels as a first pair, the left surround channel and the right surround channel as a second pair, and the center channel and the LFE channel as a third pair. For more advanced output channel formats, such as 7.1 or 11.1, other pairings can also be determined.
[0022] A further aspect is based on the finding that the audio quality of the reconstructed signal can be improved through IGF because the entire spectrum is accessible to the core encoder, and as a result, perceptually significant tonal parts, for example, in the high spectral region, can also be coded by the core encoder without parametric substitution. Additionally, a gap-filling operation is performed using frequency tiles from a first set of first spectral parts, e.g., a set of tonal parts typically from the low-frequency region, but possibly also a set of tonal parts from the high-frequency region. However, for decoder-side spectral envelope adjustment, spectral parts from the first set of spectral parts located within the reconstruction band are not further post-processed, e.g., by spectral envelope adjustment. Only the remaining spectral values in the reconstruction band that do not originate from the core decoder are envelope-adjusted using the envelope information. Preferably, the envelope information is full-band envelope information indicating the energy of a first set of first spectral portions within a reconstruction band and a second set of second spectral portions within the same reconstruction band, the latter spectral values in the second set of second spectral portions being indicated as zero and therefore not coded by the core encoder but rather parametrically coded using the low-resolution energy information.
[0023] Absolute energy values, whether normalized to the bandwidth of the corresponding band or not, have been found to be useful and very efficient in decoder-side applications, which is especially important when gain factors have to be calculated based on the residual energy in the reconstruction band, the lost energy in the reconstruction band, and the frequency tile information in the reconstruction band.
[0024] Furthermore, it is desirable that the encoded bitstream not only covers energy information for the reconstruction bands, but also additionally covers scale factors for scale factor bands extending up to the highest frequency. This ensures that for each reconstruction band in which a predetermined tonal portion, i.e., a first spectral portion, is available, a first set of this first spectral portion can actually be decoded with the correct amplitude. Furthermore, in addition to the scale factor for each reconstruction band, the energy for this reconstruction band is generated in the encoder and transmitted to the decoder. Furthermore, it is desirable that the reconstruction bands coincide with the scale factor bands, or in the case of energy grouping, at least the boundaries of the reconstruction bands coincide with the boundaries of the scale factor bands.
[0025] A further aspect is based on the finding that certain degradations in audio quality can be repaired by applying a signal-adaptive frequency tile-filling scheme. For this purpose, an analysis is performed on the encoder side to find the best matching source region candidate for a given target region. Matching information identifying a source region for the target region, together with optional additional information, is generated and transmitted to the decoder as side information. The decoder then applies the frequency tile-filling operation using the matching information. For this purpose, the decoder reads the matching information from the transmitted data stream or data file, accesses the identified source region for a given restoration band, and, if indicated by the matching information, performs additional processing of this source region data to generate raw spectral data for the restoration band. The result of the frequency tile-filling operation, i.e., the raw spectral data for the restoration band, is then shaped using spectral envelope information to finally obtain a restoration band that also includes first spectral portions, such as tonal portions. However, these tonal portions are not generated by an adaptive tile-filling scheme; these first spectral portions are directly output by the audio decoder or core decoder.
[0026] The adaptive spectral tile selection scheme may operate at a low granularity. In this embodiment, a source region is subdivided into multiple, typically overlapping, source regions, and a target region or restoration band is given by a non-overlapping frequency target region. The similarity between each source region and each target region is then determined at the encoder side, and the best matching pair of source and target regions is identified by match information. At the decoder side, the source regions identified in the match information are used to generate raw spectral data for the restoration band.
[0027] To obtain a high degree of granularity, each source region is allowed to shift to obtain the lag at which the similarity is maximized. This lag can be as fine as one frequency bin, allowing for a better match between the source and target regions.
[0028] Furthermore, in addition to identifying the best match pair, this correlation lag can also be transmitted in the match information, and even the sign can be transmitted. If the sign is determined to be negative at the encoder side, the corresponding flag is also transmitted in the match information, and at the decoder side, the spectral values in the source domain are multiplied by "-1" or "rotated" by 180 degrees in complex representation.
[0029] A further embodiment of the present invention applies a tile whitening operation. Spectral whitening removes coarse spectral envelope information and emphasizes the spectral fine structure that is most important for assessing tile similarity. Therefore, before calculating the cross-correlation measure, the frequency tiles, on the one hand, and / or the source signal, on the other hand, are whitened. When only the tiles are whitened using a predefined process, a whitening flag is transmitted to the decoder, indicating that the same predefined whitening process should be applied to the frequency tiles within the IGF.
[0030] For tile selection, it is desirable to use the correlation lag to spectrally shift the regenerated spectrum by an integer number of transform bins. Depending on the underlying transform, the spectral shift may require additional correction. For odd lags, the tiles are additionally modulated through multiplication by an alternating time sequence of -1 / 1 to compensate for the frequency-reversed representation of every other band in the MDCT. Furthermore, the sign of the correlation result is applied when generating the frequency tiles.
[0031] Furthermore, it is desirable to use tile pruning and stabilization to ensure that artifacts caused by rapid changes in source regions relative to the same reconstruction or destination region are avoided. To this end, a similarity analysis is performed between differently identified source regions, and if a source tile is similar to another source tile above a certain threshold, this source tile may be removed from the set of potential source tiles because it is highly correlated with the other source tiles. Furthermore, as a type of tile selection stabilization, it is desirable to maintain the tile order from the previous frame if none of the source tiles in the current frame are correlated (above a given threshold) with the destination tile in the current frame.
[0032] A further aspect is based on the finding that quality improvement and bitrate reduction can be achieved by combining temporal noise shaping (TNS) or temporal tile shaping (TTS) techniques with high-frequency restoration, particularly for signals containing transients that frequently occur in audio signals. The encoder-side TNS / TTS processing, performed by frequency-wide prediction, restores the temporal envelope of the audio signal. Depending on the configuration, i.e., if the temporal noise shaping filter is determined in the frequency domain covering not only the source frequency domain but also the target frequency domain to be restored in the frequency regeneration decoder, the temporal envelope is not only applied to the core audio signal up to the gap-filling start frequency, but also to the spectral domain of the reconstructed second spectral portion. In this way, pre-echoes or post-echoes that would occur without temporal tile shaping are reduced or eliminated. This is achieved by applying inverse prediction across frequency not only in the core frequency domain up to the predetermined gap-filling start frequency, but also in the frequency domain above the core frequency domain. For this purpose, frequency regeneration or frequency tile generation is performed at the decoder side before applying frequency-wide prediction. However, depending on whether the energy information calculation is performed on the spectral residual values after filtering or on the (full) spectral values before envelope shaping, the prediction across frequencies can be applied before or after spectral envelope shaping.
[0033] TTS processing across one or more frequency tiles further achieves correlation between the source and reconstructed regions, correlation in two adjacent reconstructed regions, or continuity of correlation between frequency tiles.
[0034] In one embodiment, it is desirable to use complex TNS / TTS filtering, which prevents (temporal) aliasing artifacts in critically sampled real representations such as the MDCT. The complex TNS filter can be calculated at the encoder side by additionally applying not only the modified discrete cosine transform but also the modified discrete sine transform to obtain a complex modified transform. Nevertheless, only the modified discrete cosine transform values, i.e., the real part of the complex transform, are transmitted. However, at the decoder side, it is possible to estimate the imaginary part of the transform using the MDCT spectrum of the previous or subsequent frame, so that at the decoder side, the complex filter can be reapplied for inverse prediction across frequency, specifically across the boundary between the source domain and the reconstruction domain, and across the boundary between frequency-adjacent frequency tiles within the reconstruction domain.
[0035] Our audio coding system efficiently encodes arbitrary audio signals over a wide range of bit rates. It converges to transparency at high bit rates while minimizing perceptual disruption at low bit rates. Thus, at the encoder, most of the available bit rate is used to waveform-code only the perceptually most significant structures of the signal, and at the decoder, the resulting spectral gaps are filled with signal content that roughly approximates the original spectrum. A very limited bit budget is consumed for parameter-driven so-called spectral intelligent gap filling (IGF), controlled by dedicated side information transmitted from the encoder to the decoder.
[0036] In a further embodiment, the time domain encoding / decoding processor relies on a low sampling rate and a corresponding bandwidth extension feature.
[0037] In a further embodiment, a cross processor is provided to initialize the time-domain coder / decoder with initialization data derived from the currently processed frequency-domain coder / decoder signal. Thus, if the currently processed audio signal portion is processed by a frequency-domain coder, a parallel time-domain coder is initialized so that when a switchover from the frequency-domain coder to the time-domain coder occurs, the time-domain coder can start processing because all initialization data related to the previous signal is already present by the cross processor. This cross processor is preferably applied on the coder side and additionally on the decoder side, and preferably uses a frequency-to-time transform that additionally performs a highly efficient downsampling from a high output or input sampling rate to a low time-domain core coder sampling rate by simply selecting a predetermined low-band portion of the time-domain signal with a predetermined reduced transform size. In this way, the sampling rate conversion from a high sampling rate to a low sampling rate is performed very efficiently, and this signal obtained by the conversion with a reduced transform size can then be used to initialize the time domain coder / decoder so that it is ready to immediately perform time domain coding if time domain coding is signaled by the controller and the previous audio signal portion was coded in the frequency domain.
[0038] Thus, preferred embodiments of the present invention allow seamless switching between a perceptual audio coder with spectral gap filling and a time domain coder with or without bandwidth extension.
[0039] Thus, the present invention is not limited to removing high frequency content above a cutoff frequency from an audio signal in a frequency domain coder, but rather relies on signal-adaptive removal of spectral bandpass regions leaving spectral gaps in the coder, and subsequently restoring these spectral gaps in the decoder. Preferably, an integrated solution such as intelligent gap filling is used, which efficiently combines full-bandwidth audio coding and spectral gap filling, especially in the MDCT transform domain.
[0040] Thus, the present invention provides an improved concept for combining speech coding with subsequent time-domain bandwidth expansion and full-band waveform decoding including spectral gap filling into a switchable perceptual coder / decoder.
[0041] Thus, in contrast to existing methods, the new concept utilizes full-band audio signal waveform coding in a transform domain coder, while at the same time allowing a seamless switchover to a speech coder, preferably followed by time domain bandwidth extension.
[0042] A further embodiment of the present invention avoids the aforementioned problems caused by fixed bandwidth limitations. This concept allows a switchable combination of a frequency-domain full-bandwidth waveform coder / decoder with spectral gap filling and a low-sampling-rate speech coder / decoder and time-domain bandwidth extension. Such a coder / decoder can waveform-code the aforementioned problematic signals, providing the full audio bandwidth up to the Nyquist frequency of the audio input signal. However, seamless switching between both coding schemes is ensured by embodiments that include a cross-processor. For this seamless switching, the cross-processor represents a cross-connection between a full-bandwidth, full-rate (input sampling rate) frequency-domain coder and a low-rate ACELP coder with a low sampling rate, both in the coder and the decoder. This cross-processor properly initializes ACELP parameters and buffers, particularly in the adaptive codebook, LPC filter, or resampling stage, when switching from a frequency-domain coder such as TCX to a time-domain coder such as ACELP.
[0043] Embodiments of the present invention will now be described with reference to the accompanying drawings. [Brief explanation of the drawings]
[0044] [Figure 1a] 1 shows an apparatus for encoding an audio signal. [Figure 1b] 1 shows a decoder for decoding an encoded audio signal, which is compatible with the encoder of FIG. 1a; [Figure 2a] 1 shows a preferred configuration of the decoder. [Figure 2b] 1 shows a preferred configuration of the encoder. [Figure 3a] 1b shows a schematic representation of the spectrum produced by the spectral domain decoder of FIG. 1b; [Figure 3b] 10 is a table showing the relationship between scale factors for scale factor bands, energy for restoration bands, and noise filling information for noise filling bands. [Figure 4a]Show the functionality of the spectral domain coder applying the spectral portion selection to the first and second sets of spectral portions [Figure 4b] The functional configuration is shown in Figure 4a. [Figure 5a] The function of the MDCT encoder is shown. [Figure 5b] The function of the decoder with MDCT technology is shown. [Figure 5c] 1 shows the configuration of a frequency regeneration unit. [Figure 6] 1 shows the configuration of an audio encoder. [Figure 7a] 1 shows a cross processor in an audio encoder. [Figure 7b] 1 shows the configuration of an inverse or frequency-to-time transform that additionally provides sampling rate reduction within the cross-processor. [Figure 8] 7 illustrates a preferred embodiment of the controller of FIG. 6. [Figure 9] 1 illustrates a further embodiment of a time domain coder with bandwidth extension capabilities. [Figure 10] A preferred method of using the pre-treatment unit is shown. [Figure 11a] 1 shows a schematic configuration of an audio decoder. [Figure 11b] 1 illustrates a cross processor within a decoder that provides initialization data for a time domain decoder. [Figure 12] A preferred configuration of the time domain decoding processor of FIG. 11a is shown. [Figure 13] 10 shows a further configuration of time domain bandwidth extension. [Figure 14a-1] 1 shows a part of a preferred configuration of an audio encoder. [Figure 14a-2] 4 shows the remainder of the preferred structure of the audio encoder. [Figure 14b] 1 shows a preferred configuration of an audio decoder. [Figure 14c] 1 illustrates an inventive configuration of a time domain decoder with sample rate conversion and bandwidth extension.
[0045] 6 shows an audio encoder for encoding an audio signal, including a first encoding processor 600 for encoding a first audio signal portion in the frequency domain. The first encoding processor 600 includes a time-to-frequency transform unit 602 for transforming a first input audio signal portion into a frequency-domain representation having spectral lines up to the maximum frequency of the input signal. The first encoding processor 600 further includes an analyzer 604 for analyzing the frequency-domain representation up to the maximum frequency, which analyzer determines a first spectral region to be encoded with a first spectral resolution and a second spectral region to be encoded with a second spectral resolution lower than the first spectral resolution. In particular, the full-band analyzer 604 determines which frequency lines or spectral values in the time-to-frequency transform unit spectrum should be encoded for each spectral line, and which other spectral portions should be encoded in a parametric manner, which latter spectral portions are then reconstructed at the decoder side using a gap-filling process. The actual encoding operation is performed by a spectral coder 606, which parametrically encodes a first spectral region or portion with a first resolution and a second spectral region or portion with a second spectral resolution.
[0046] The audio encoder of Fig. 6 further comprises a second encoding processor 610 for encoding audio signal portions in the time domain. Furthermore, the audio encoder comprises a controller 620 configured to analyze the audio signal at the audio signal input 601 and determine which portions of the audio signal are first audio signal portions to be encoded in the frequency domain and which portions of the audio signal are second audio signal portions to be encoded in the time domain. Furthermore, an encoded signal former 630, which may be configured, for example, as a bitstream multiplexer, is provided and is configured to form an encoded audio signal comprising a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion. The important point is that the encoded signal only comprises either a frequency-domain representation or a time-domain representation from one and the same audio signal portion.
[0047] Therefore, the controller 620 ensures that only one time-domain or frequency-domain representation for a single audio portion is present in the coded signal. There are several ways in which this can be achieved by the controller 620. One way is that for one and the same audio signal portion, both representations reach the block 630, and the controller 620 controls the coded signal former 630 to introduce only one of the both representations into the coded signal. However, alternatively, the controller 620 can also control the inputs to the first coding processor and the inputs to the second coding processor in such a way that, based on an analysis of the corresponding signal portion, only one of the both blocks 600 and 610 is activated to actually perform the entire coding operation, while the other block is deactivated.
[0048] Such deactivation can be inactivity, or it can be a kind of "initialization" mode, as shown, for example, with respect to FIG. 7a. In that initialization mode, the other encoding processor is activated only to receive and process initialization data to initialize its internal memory, and does not perform any special encoding operations at all. Such activation can be performed by a predetermined switch at the input (not shown in FIG. 6), or preferably by control lines 621 and 622. Thus, in this embodiment, when the controller 620 determines that the current audio signal portion should be encoded by the first encoding processor, the second encoding processor 610 does not output anything; instead, the second encoding processor is provided with initialization data so that it can be instantly switched on and activated in the future. On the other hand, the first encoding processor is configured not to need any data from the past to update any internal memory, and therefore, when the current audio signal portion should be encoded by the second encoding processor 610, the controller 620 can control, via the control line 621, that the first encoding processor 610 is completely inactive. This means that the first encoding processor 600 does not need to be in an initialization or standby state, but can remain completely inactive, which is particularly advantageous for mobile devices where power consumption and therefore battery life are an issue.
[0049] In a further specific configuration of the second encoding processor operating in the time domain, the second encoding processor comprises a downsampler 900 or sampling rate converter for converting the audio signal portion into a representation having a lower sampling rate, which is lower than the sampling rate at the input to the first encoding processor. This is shown in FIG. 9. In particular, if the input audio signal comprises a low band and a high band, the low sampling rate representation at the output of block 900 preferably comprises only the low band of the input audio signal portion, which low band is then coded by a time-domain low-band coder 910. This coder 910 is configured to time-domain code the low sampling rate representation provided by block 900. Furthermore, a time-domain bandwidth extension coder 920 is provided for parametrically coding the high band. For this purpose, the time-domain bandwidth extension coder 920 receives at least the high band of the input audio signal, or the low band and the high band of the input audio signal.
[0050] In a further embodiment of the present invention, the audio encoder further comprises a pre-processing unit 1000 (not shown in Fig. 6 but as shown in Fig. 10) configured to pre-process the first and second audio signal portions. In one embodiment, this pre-processing unit comprises a prediction analysis unit for determining prediction coefficients. This prediction analysis unit may be configured as an LPC (Linear Predictive Coding) analysis unit for determining LPC coefficients. However, other analysis units may also be configured. Furthermore, the pre-processing unit, also depicted in Fig. 14b, comprises a prediction coefficient quantization unit 1010, and the apparatus depicted in Fig. 14a receives prediction coefficient data from a prediction analysis unit, designated 1002 in Fig. 14a.
[0051] Furthermore, the pre-processing unit additionally comprises an entropy coder for generating a coded version of the quantized prediction coefficients. What is important is that the coded signal forming unit 630 or a specific configuration, i.e. the bitstream multiplexer 613, ensures that a coded version of the quantized prediction coefficients is included in the coded audio signal 632. Preferably, the LPC coefficients are not directly quantized, but are converted, for example, into ISFs or any other representation more suitable for quantization. This conversion is preferably performed by the LPC coefficients determination block 1002 or in the block 1010 for quantizing the LPC coefficients.
[0052] Furthermore, the pre-processing unit may include a resampler 1004 that resamples the audio input signal at the input sampling rate to a lower sampling rate for the time-domain coder. If the time-domain coder is an ACELP coder with an ACELP sampling rate, downsampling is preferably performed to 12.8 kHz or 16 kHz. The input sampling rate may be any specific sampling rate, such as 32 kHz or a higher sampling rate. On the other hand, the sampling rate of the time-domain coder will be predetermined by certain constraints, and the resampler 1004 performs this resampling to output a lower sampling rate representation of the input signal. Thus, the resampler 1004 may perform a function similar to the downsampler 900 described in the context of FIG. 9, or may even be the same component as the downsampler 900.
[0053] Furthermore, it is advisable to apply pre-emphasis in the pre-emphasis block 1005 shown in Fig. 14a. Pre-emphasis processing is well known in the art of time domain coding and has been presented in the literature referring to AMR-WB+ processing, and is specifically adapted to compensate for the spectral tilt, which allows for a better calculation of the LPC parameters for a given LPC order.
[0054] Furthermore, the pre-processing unit may additionally include a TCX-LTP parameter extraction unit for controlling an LTP (long-term prediction) post-filter, shown in Fig. 14b as 1420. This block is shown in Fig. 14a as 1006. In addition, the pre-processing unit may additionally include other functions, shown as 1007, which may include a pitch search function, a voice activity detection (VAD) function, or any other function known in the art of time domain or speech coding.
[0055] As mentioned above, the result of block 1006 is input into the coded signal, i.e., as shown in the embodiment of Fig. 14a, into the bitstream multiplexer 630. Furthermore, if required, data from block 1007 can also be input into the bitstream multiplexer, or alternatively, can be used for time domain coding in a time domain coder.
[0056] To summarize, both paths share a pre-processing operation 1000 in which commonly used signal processing operations are performed. These operations include resampling to the ACELP sampling rate (12.8 or 16 kHz) for one parallel path, which is always performed. Furthermore, TCX LTP parameter extraction is performed as shown in block 1006, as well as pre-emphasis and LPC coefficient determination. As mentioned above, pre-emphasis compensates for spectral tilt, thereby making the calculation of LPC parameters for a given LPC order more efficient.
[0057] Please refer now to Fig. 8, which shows a preferred embodiment of the controller 620. The controller receives at its input the audio signal portion under consideration. Preferably, as shown in Fig. 14a, the controller receives any signal available in the pre-processing unit 1000, which may be either the original input signal at the input sampling rate, a resampled version at a lower time-domain coder sampling rate, or a signal obtained after pre-emphasis processing in block 1005.
[0058] Based on this audio signal portion, the controller 620 commands the frequency domain coder simulator 621 and the time domain coder simulator 622 to calculate an estimated signal-to-noise ratio for each coder. Then, the selector 623 selects the coder that provided the better signal-to-noise ratio, taking into account the given bit rate. The selector then identifies the corresponding coder via a control output. If it is determined that the audio signal portion under consideration should be coded using a frequency domain coder, the time domain coder is set to an initialization state, or in other embodiments, does not need to be instantly switched to a completely deactivated state. However, if it is determined that the audio signal portion under consideration should be coded by a time domain coder, the frequency domain coder is deactivated.
[0059] Next, a preferred embodiment of the controller shown in Figure 8 will be described. The decision of whether to choose the ACELP or TCX path is made in the switching decision unit by simulating the ACELP and TCX encoders and switching to the branch that performs better. To this end, the SNRs of the ACELP and TCX branches are estimated based on ACELP and TCX encoder / decoder simulations. The TCX encoder / decoder simulations are performed without using TNS / TTS analysis, an IGF encoder, a quantization loop / arithmetic encoder, or any TCX decoder. Instead, the TCX SNR is estimated using an estimate of the quantizer distortion in the shaped MDCT domain. The ACELP encoder / decoder simulations are performed using only adaptive and innovative codebook simulations. The ACELP SNR is estimated simply by calculating the distortion introduced by the LTP filter in the weighted signal domain (adaptive codebook) and scaling this distortion by a constant factor (innovative codebook). In this way, the complexity is significantly reduced compared to approaches where TCX and ACELP coding are performed in parallel: the branch with the higher SNR is selected for the subsequent full coding operation.
[0060] If the TCX branch is selected, the TCX decoder runs each frame and outputs a signal at the ACELP sampling rate. This signal is used to update the memories used for the ACELP coding paths (LPC residual, Memwe, memory de-emphasis), allowing instantaneous switching from TCX to ACELP. Memory updates are performed within each TCX path.
[0061] Alternatively, a full analysis-by-synthesis process can be performed, i.e., both coder simulators 621, 622 perform the actual coding operation and their results are compared by the selector 623. Alternatively, a full feedforward calculation can also be performed by performing a signal analysis. For example, if the signal classifier determines that the signal is a speech signal, a time domain coder is selected, and if the signal classifier determines that the signal is a tone signal, a frequency domain coder is selected. Other techniques for distinguishing between both coders based on a signal analysis of the audio signal portion under consideration are also applicable.
[0062] Preferably, the audio encoder may additionally include a cross processor 700 as shown in Fig. 7a. When the frequency domain encoder 600 is active, the cross processor 700 provides initialization data to the time domain encoder 610, enabling the time domain encoder to support seamless switching in future signal portions. In other words, if the controller determines that a current signal portion should be coded using a frequency domain encoder and that the immediately following audio signal portion should be coded by the time domain encoder 610, such instantaneous seamless switching would not be possible without the above-mentioned cross processor. However, the cross processor provides a signal derived from the frequency domain encoder 600 to the time domain encoder 610 for the purpose of initializing the memory within the time domain encoder. This is because the time domain encoder 610 has a dependency for the current frame from the input signal or coded signal of the temporally previous frame.
[0063] In this way, the time domain encoder 610 is configured to be initialized with initialization data so that it can encode in an efficient manner the portion of the audio signal that follows the previous portion of the audio signal that was encoded by the frequency domain encoder 600.
[0064] In particular, the cross processor includes a frequency-to-time transform unit that converts the frequency domain representation into a time domain representation that can be sent to a time domain coder directly or after some further processing. This transform unit is shown in Figure 14a as an IMDCT (inverse modified discrete cosine transform) block 702. However, this block has a different transform size than the time-to-frequency transform block 602, which is shown in Figure 14a as a modified discrete cosine transform block. As shown in block 602, the time-to-frequency transform unit 602 operates at the input sampling rate, while the inverse modified discrete cosine transform unit 702 operates at the lower ACELP sampling rate.
[0065] The ratio between the time-domain coder sampling rate or ACELP sampling rate and the frequency-domain coder sampling rate or input sampling rate can be calculated, and this ratio becomes the downsampling factor DS shown in FIG. 7b. Block 602 has a large transform size, and IMDCT block 702 has a small transform size. Thus, as shown in FIG. 7b, IMDCT block 702 includes a selector 726 that selects the lower spectral portion of the input to IMDCT block 702. That portion of the full-band spectrum is defined by the downsampling factor DS. For example, if the lower sampling rate is 16 kHz and the input sampling rate is 32 kHz, the downsampling factor is 0.5, and therefore selector 726 selects the lower half of the full-band spectrum. For example, if the spectrum has 1024 MDCT lines, the selector selects the lower 512 MDCT lines.
[0066] This low frequency part of the full regional spectrum is input to a small size transform and foldout block 720, as shown in Figure 7b. The transform size is also selected according to the downsampling factor and is 50% of the transform size in block 602. Next, synthesis windowing is performed using a window with a small number of coefficients. The number of coefficients of the synthesis window is equal to the downsampling factor multiplied by the number of coefficients of the analysis window used by block 602. Finally, an overlap-add operation is performed with a small number of operations per block, which is also the number of operations per block in the full rate configuration MDCT multiplied by the downsampling factor.
[0067] In this way, a very efficient downsampling operation can be applied since the downsampling is included in the IMDCT structure. It should be emphasized in this context that block 702 can be constituted by an IMDCT, but also by any other transform or filter bank structure that can be appropriately sized for practical transform kernels and other transform-related operations.
[0068] In a further embodiment shown in Figure 14a, the time-frequency transform unit includes additional functionality in addition to the analyzer unit: the analyzer unit 604 of Figure 6 may in the embodiment of Figure 14a include a temporal noise shaping / temporal tile shaping analysis block 604a, which operates as described in the context of block 222 of Figure 2b as the TNS / TTS analysis block 604a, and the IGF encoder 604b in Figure 14a operates as described with respect to its corresponding tonality mask 226 of Figure 2b.
[0069] Furthermore, the frequency-domain coder preferably includes a noise shaping block 606a, which is controlled by the quantized LPC coefficients generated by block 1010. The quantized LPC coefficients used for noise shaping 606a perform spectral shaping of high-resolution spectral values or directly (as opposed to parametrically) coded spectral lines, and the result of block 606a resembles the spectrum of the signal after an LPC filtering stage operating in the time domain, such as LPC analysis filtering block 704, described below. Furthermore, the result of noise shaping block 606a is then quantized and entropy coded, as shown in block 606b. The result of block 606b (together with other side information) corresponds to the coded first audio signal portion or the frequency-domain coded audio signal portion.
[0070] The cross processor 700 includes a spectral decoder that computes a decoded version of the first encoded signal portion. In the embodiment of FIG. 14a, the spectral decoder 701 includes an inverse noise shaping block 703, a gap filling decoder 704, a TNS / TTS synthesis block 705, and the aforementioned IMDCT block 702. These blocks reverse certain operations performed by blocks 602-606b. In particular, the noise shaping block 703 reverses the noise shaping performed by block 606a based on the quantized LPC coefficients 1010. The IGF decoder 704 operates as described for blocks 202 and 206 with respect to FIG. 2A, the TNS / TTS synthesis block 705 operates as described in the context of block 210 of FIG. 2A, and the spectral decoder additionally includes the IMDCT block 702. Furthermore, the cross processor 700 of Figure 14a may additionally or alternatively include a delay stage 707 which supplies a delayed version of the decoded version obtained by the spectral decoder 701 to the de-emphasis stage 617 of the second encoding processor for initializing the de-emphasis stage 617.
[0071] Further, the cross processor 700 may additionally or alternatively include a weighted prediction coefficient analysis filtering stage 708, which filters the decoded version and provides the filtered decoded version to the codebook determiner 613, shown as "MMSE" in FIG. 14a, of the second encoding processor for initializing this block. Alternatively or additionally, the cross processor may include an LPC analysis filtering stage, which filters the decoded version of the first encoded signal portion output by the spectral decoder 701 and provides it to the adaptive codebook stage 612 for initializing this block. Alternatively or additionally, the cross processor may include a pre-emphasis stage 709, which performs pre-emphasis processing on the decoded version output by the spectral decoder 701 before LPC filtering. The output of the pre-emphasis stage may also be provided to an additional delay stage 710 for initializing the LPC synthesis filtering block 616 in the time-domain encoder 610.
[0072] The time-domain coding processor 610 includes a pre-emphasis block 612 operating at a low ACELP sample rate, as shown in Figure 14a. As shown, this pre-emphasis block is implemented in a pre-processing stage 1000 and is designated by the reference numeral 1005. The pre-emphasis data is input to an LPC analysis filtering stage 611 operating in the time domain, and this filter is controlled by the quantized LPC coefficients 1010 obtained by the pre-processing stage 1000. As is known from AMR-WB+, USAC or other CELP coders, the residual signal generated by block 611 is fed to an adaptive codebook 612, which is further connected to an innovative codebook stage 614, and the codebook data from the adaptive codebook 612 and the innovative codebook are input to the aforementioned bitstream multiplexer.
[0073] Furthermore, the ACELP gain / encoding stage 612 is provided in series with the innovative codebook stage 614, and the result of this block is input to the codebook determination block 613, shown as MMSE in FIG. 14a. This block cooperates with the innovative codebook block 614. Furthermore, the time-domain coder additionally includes a decoder portion having an LPC synthesis filtering block 616, a de-emphasis block 617, and an adaptive bass post-filter stage 618 for calculating parameters for the adaptive bass post-filter, where this adaptive bass post-filter is applied on the decoder side. If there is no adaptive bass post-filtering on the decoder side, blocks 616, 617, 618 would be unnecessary for the time-domain coder 610.
[0074] As shown, a plurality of blocks of the time-domain coder are dependent on the preceding signal, and these blocks are the adaptive codebook block, the codebook determination unit 613, the LPC synthesis filtering block 616, and the de-emphasis block 617. These blocks are supplied with data from the cross-processor, derived from the data of the frequency-domain encoding processor, to initialize these blocks in preparation for an instantaneous switch from the frequency-domain coder to the time-domain coder. As can be further seen from FIG. 14a, no dependency on previous data is required for the frequency-domain coder. Thus, the cross-processor 700 does not provide any memory initialization data to the frequency-domain coder with respect to the time-domain coder. However, for other embodiments of the frequency-domain coder where there is a dependency from the past and memory initialization data is required, the cross-processor 700 is configured to operate in both directions.
[0075] Thus, a preferred embodiment of the audio coder includes the following components.
[0076] A preferred audio decoder is described below: The waveform decoder section consists of a full-band TCX decoder path and an IGF, both operating at the input sampling rate of the codec. In parallel with this there is an alternative ACELP decoder path at a lower sampling rate, which is further augmented downstream by a TD-BWE.
[0077] For ACELP initialization when switching from TCX to ACELP, there is a cross-path (consisting of a shared TCX decoder front-end that provides additional output at a lower sampling rate and some post-processing) that performs the ACELP initialization of the present invention. In LPC, sharing the same sampling rate and filter order between TCX and ACELP allows for easier and more efficient ACELP initialization.
[0078] To visualize the switching, two switches are shown in Fig. 14b: the second switch selects between the output of TCX / IGF or ACELP / TD-BWE downstream, while the first switch 1480 either pre-updates the buffer in the resampling QMF stage downstream of the ACELP path with the output of the cross path, or simply passes the ACELP output through.
[0079] Next, the configuration of an audio decoder according to an embodiment of the present invention will be described with reference to Figs. 11a to 14c.
[0080] The audio decoder for decoding the encoded audio signal 1101 includes a first decoding processor 1120 for decoding a first encoded audio signal portion in the frequency domain. The first decoding processor 1120 includes a spectral decoder 1122 for decoding a first spectral region with high spectral resolution and synthesizing the second spectral region using a parametric representation of the second spectral region and at least one decoded first spectral region to obtain a decoded spectral representation. This decoded spectral representation is a full-band decoded spectral representation, as described in connection with FIG. 6 and also in connection with FIG. 1a. Thus, the first decoding processor typically includes a full-band configuration with gap-filling processing in the frequency domain. The first decoding processor 1120 further includes a frequency-to-time transform unit 1124 for transforming the decoded spectral representation into the time domain to obtain a decoded first audio signal portion.
[0081] The audio decoder further includes a second decoding processor 1140 that decodes the second encoded audio signal portion in the time domain to obtain a decoded second signal portion. The audio decoder further includes a combiner 1160 that combines the decoded first signal portion and the decoded second signal portion to obtain a decoded audio signal. The decoded signal portions are combined sequentially, as also illustrated by the switch arrangement 1160 of Figure 14b, which represents one embodiment of the combiner 1160 of Figure 11a.
[0082] Preferably, the second decoding processor 1140 is a time-domain bandwidth extension processor, and includes a time-domain low-band decoder 1200 for decoding a low-band time-domain signal, as shown in Fig. 12. This configuration further includes an upsampler 1210 for upsampling the low-band time-domain signal. In addition, a time-domain bandwidth extension decoder 1220 is provided for synthesizing a high-band of the output audio signal. A mixer 1230 is further provided, which mixes the synthesized high-band of the time-domain output signal with the upsampled low-band time-domain signal to obtain a time-domain decoder output. Thus, the block 1140 in Fig. 11a may be configured by the functions of Fig. 12 in a preferred embodiment.
[0083] Figure 13 shows a preferred embodiment of the time-domain bandwidth extension decoder 1220 of Figure 12. Preferably, a time-domain upsampler 1221 is provided, which receives as input the LPC residual signal from a time-domain low-band decoder, which is included in block 1140, designated 1200 in Figure 12, and further illustrated in the context of Figure 14b. The time-domain upsampler 1221 generates an upsampled version of the LPC residual signal. This version is then input to a nonlinear distortion block 1222, which generates an output signal with higher frequency values based on its input signal. The nonlinear distortion may be copying, mirroring, frequency shifting, or a nonlinear device such as a diode or transistor operated in a nonlinear region. The output signal of block 1222 is input to an LPC synthesis filtering block 1223, which is controlled by the LPC data also used for the lowband decoder, or by specific envelope data generated for example by the time domain bandwidth extension block 920 on the encoder side of Fig. 14a. The output of the LPC synthesis block is then input to a band pass or high pass filter 1224 to finally obtain the high band, which is then input to a mixer 1230 shown in Fig. 12.
[0084] A preferred embodiment of the upsampler 1210 of Figure 12 will now be described with reference to Figure 14b. This upsampler preferably includes an analysis filterbank operating at a first time-domain low-band decoder sampling rate. One specific implementation of such an analysis filterbank is the QMF analysis filterbank 1471 shown in Figure 14b. Furthermore, this upsampler includes a synthesis filterbank 1473 operating at a second output sampling rate that is higher than the first time-domain low-band sampling rate. Thus, the QMF synthesis filterbank 1473, which is a preferred implementation of a general filterbank, operates at the output sampling rate. 7b, the QMF analysis filter bank 1471 has, for example, only 32 filter bank channels, and the QMF synthesis filter bank 1473 has, for example, 64 QMF channels, but the higher half of the filter bank channels, i.e., the upper 32 filter bank channels, are fed with zeros or noise, while the lower 32 filter bank channels are fed with the corresponding signal provided by the QMF analysis filter bank 1471. However, the bandpass filtering 1472 is preferably performed within the QMF filter bank domain, which ensures that the QMF synthesis output 1473 is an upsampled version of the ACELP decoder output while not introducing any artifacts above the maximum frequency of the ACELP decoder.
[0085] In addition to or alternatively to bandpass filtering 1472, further processing operations may be performed in the QMF domain. If no processing is performed, the QMF analysis and QMF synthesis constitute an efficient upsampler 1210.
[0086] The construction of the individual elements of Figure 14b will now be described in more detail.
[0087] The full-band frequency-domain decoder 1120 includes a first decoding block 1122a that decodes the high-resolution spectral coefficients and also performs noise filling in the low-band part, as known for example from the USAC technique. Furthermore, the full-band decoder includes an IGF processor 1122b for filling spectral holes using synthesized spectral values that were only parametrically coded at the encoder side, and therefore at low resolution. Then, in block 1122c, an inverse noise shaping is performed, the result of which is input to the TNS / TTS synthesis block 705, which provides as its final output an input to a frequency-to-time transformer 1124, preferably configured as an inverse modified discrete cosine transform operating at the output sampling rate, i.e., the higher sampling rate.
[0088] Furthermore, a harmonic or LTP post-filter is used, which is controlled by the data obtained by the TCX LTP parameter extraction block 1006 of Fig. 14a. The result is a decoded first audio signal portion at the output sampling rate, which, as can be seen from Fig. 14b, has a high sampling rate and therefore no additional frequency reinforcement is required at all, since the decoding processor is a frequency-domain full-band decoder, preferably operating using the intelligent gap-filling technique described in the context of Figs. 1a-5c.
[0089] Some components in Figure 14b are very similar to their counterparts in the cross processor 700 of Figure 14a, particularly the IGF decoder 704, which corresponds to the IGF processing 1122b; the inverse noise shaping operation controlled by the quantized LPC coefficients 1145, which corresponds to the inverse noise shaping 703 of Figure 14a; and the TNS / TTS synthesis block 705 of Figure 14b, which corresponds to the TNS / TTS synthesis block 705 of Figure 14a. Importantly, however, the IMDCT block 1124 of Figure 14b operates at a high sampling rate, while the IMDCT block 702 of Figure 14a operates at a low sampling rate. Accordingly, the block 1124 of Figure 14b includes the larger-sized transform and convolution block 710, the synthesis window of block 712, and the overlap-add stage 714, which have a larger number of operations, a larger number of window coefficients, and a larger transform size than the corresponding features 720, 722, and 724 operated in block 701. This point is also discussed below with respect to block 1171 of cross processor 1170 in Figure 14b.
[0090] The time domain decoding processor 1140 preferably includes an ACELP or time domain lowband decoder 1200, which includes an ACELP decoder stage 1149 that receives the decoded gain and innovative codebook information. An ACELP adaptive codebook stage 1141 is then provided, followed by an ACELP post-processing stage 1142 and a final synthesis filter, such as an LPC synthesis filter 1143, controlled by quantized LPC coefficients 1145 obtained from the bitstream demultiplexer 1100, which corresponds to the coded signal analysis unit 1100 of Figure 11a. The output of the LPC synthesis filter 1143 is input to a de-emphasis stage 1144, which cancels or reverses the processing introduced by the pre-emphasis stage 1005 of the pre-processing unit 1000 of Figure 14a. The result is a time domain output signal at a low sampling rate and low bandwidth; when a time domain output is required, switch 1480 is in the position shown and the output of de-emphasis stage 1144 is input to upsampler 1210 and then mixed with the high bandwidth from the time domain bandwidth extension decoder 1220.
[0091] According to an embodiment of the present invention, the audio decoder further comprises a cross processor 1170, shown in Figures 11b and 14b, which calculates initialization data for the second decoding processor from the decoded spectral representation of the first encoded audio signal portion, thereby initializing the second decoding processor for decoding the second encoded audio signal portion that temporally follows the first audio signal portion in the encoded audio signal, i.e. the time domain decoding processor 1140 is prepared to switch instantly from one audio signal portion to the next without loss in quality or efficiency.
[0092] Preferably, the cross processor 1170 includes an additional frequency-to-time transform unit 1171 operating at a lower sampling rate than the frequency-to-time transform unit of the first decoding processor to obtain an additional decoded first signal portion in the time domain. The additional decoded first signal portion can be used as an initialization signal or any initialization data can be derived therefrom. This IMDCT or low-sampling-rate frequency-to-time transform unit is preferably configured as shown in FIG. 7b, including item 726 (selection unit), item 720 (small-size transform and folding), synthesis windowing with a small number of window coefficients as indicated by 722, and an overlap-add stage with a small number of operations as indicated by 724. Thus, the IMDCT block 1124 in the frequency-domain fullband decoder is configured as indicated by blocks 710, 712, and 714, and the IMDCT block 1171 is configured as indicated by blocks 726, 720, 722, and 724 in FIG. 7b. Again, the downsampling factor is the ratio of the time domain coder sampling rate or lower sampling rate to the higher frequency domain coder sampling rate or output sampling rate, and this downsampling factor can be any number less than 1, greater than 0 and less than 1.
[0093] 14b, the cross processor 1170 may further include a delay stage 1172, alone or in addition to other components, for delaying the additional decoded first signal portion and providing the delayed decoded first signal portion to the de-emphasis stage 1144 of the second decoding processor for initialization. Additionally or alternatively, the cross processor may further include a pre-emphasis filter 1173 and a delay stage 1175 for filtering and delaying the additional decoded first signal portion, the delayed output of which is provided to the LPC synthesis filtering stage 1143 of the ACELP decoder for initialization.
[0094] Furthermore, the cross processor may alternatively or in addition to the other components mentioned above include an LPC analysis filter 1174 which generates a prediction residual signal from the additional decoded first signal portion or the pre-emphasized additional decoded first signal portion and supplies this data to the codebook synthesis unit and preferably to the adaptive codebook stage 1141 of the second decoding processor. Furthermore, the output of the frequency-to-time transform unit 1171 with low sampling rate is also input to the QMF analysis stage 1471 of the upsampler 1210 for initialization purposes, i.e. when the audio signal portion currently being decoded is provided by the frequency-domain fullband decoder 1120.
[0095] A preferred audio decoder is described below. The waveform decoder section consists of a full-band TCX decoder path and an IGF, both operating at the input sampling rate of the codec. In parallel with this there is an alternative ACELP decoder path at a lower sampling rate, which is further augmented downstream by a TD-BWE.
[0096] For ACELP initialization when switching from TCX to ACELP, there is a cross-path (consisting of a shared TCX decoder front-end that provides additional output at a lower sampling rate and some post-processing) that performs the ACELP initialization of the present invention. In LPC, sharing the same sampling rate and filter order between TCX and ACELP allows for easier and more efficient ACELP initialization.
[0097] To visualize the switching, two switches are shown in Fig. 14b. The second switch selects between the TCX / IGF or ACELP / TD-BWE outputs downstream, while the first switch either pre-updates the buffer in the resampling QMF stage downstream of the ACELP path with the output of the cross path, or simply passes the ACELP output through.
[0098] In summary, preferred aspects of the present invention, which can be used alone or in combination, relate to the combination of ACELP and TD-BWE coders with full-bandwidth TCX / IGF techniques, preferably also using cross signals.
[0099] A further specific feature is the cross signal path for ACELP initialization, allowing for seamless switching.
[0100] A further aspect is that the short IMDT is fed with the lower portion of the high rate long MDCT coefficients in order to efficiently perform the sample rate conversion in the cross path.
[0101] A further feature is the efficient realization of a full-band TCX / IGF and partially shared cross-path in the decoder.
[0102] A further feature is the cross signal path for QMF initialization, allowing seamless switching from TCX to ACELP.
[0103] An additional feature is a cross signal path to the QMF that allows the delay gap between the ACELP resampled output and the filter bank-TCX / IGF output to be compensated when switching from ACELP to TCX.
[0104] A further aspect is that the LPC is provided for both the TCX and ACELP encoders at the same sampling rate and filter order, even though the TCX / IGF encoder / decoder is full-band capable.
[0105] Next, FIG. 14c will be described as a preferred implementation of a time domain decoder, operating either as a stand-alone decoder or in combination with a full-bandwidth frequency domain decoder.
[0106] Generally, the time domain decoder includes an ACELP decoder followed by a resampler or upsampler and a time domain bandwidth extension function. In particular, the ACELP decoder includes an ACELP decoding stage 1149 for recovering the gain and the innovative codebook, an ACELP adaptive codebook stage 1141, an ACELP post-processing unit 1142, an LPC synthesis filter 1143 controlled by the quantized LPC coefficients from the bitstream demultiplexer or coded signal analysis unit, followed by a de-emphasis stage 1144. Preferably, the time domain residual signal at the ACELP sampling rate is input to a time domain bandwidth extension decoder 1220, which provides a high bandwidth at its output.
[0107] To upsample the output of the de-emphasis 1144, an upsampler including a QMF analysis block 1471 and a QMF synthesis block 1473 is provided. Preferably, a band-pass filter is applied within the filter bank domain defined by blocks 1471 and 1473. In particular, as mentioned above, the same functions as in the blocks described above with the same reference numerals may be used. Furthermore, the time-domain bandwidth extension decoder 1220 may be configured as shown in Figure 13 and generally involves upsampling the ACELP residual signal or the time-domain residual signal at the ACELP sampling rate to the output sampling rate of the final bandwidth extension signal.
[0108] Details regarding a full-band frequency domain encoder and decoder will now be described with reference to Figures 1a to 5c.
[0109] FIG. 1a shows an apparatus for encoding an audio signal 99. The audio signal 99 is input to a time-spectral transform unit 100, which converts the audio signal having a certain sampling rate into an output spectral representation 101. The spectrum 101 is input to a spectral analysis unit 102, which analyzes the spectral representation 101. The spectral analysis unit 102 is configured to determine a first set 103 of first spectral portions to be coded with a first spectral resolution and a second set 105 of second spectral portions to be coded with a different second spectral resolution, the second spectral resolution being smaller than the first spectral resolution. The second set 105 of second spectral portions is input to a parameter calculation unit or parametric coder 104 for calculating spectral envelope information having the second spectral resolution. A spectral-domain audio coder 106 is further provided for generating a first coded representation 107 of the first set of first spectral portions having the first spectral resolution. Furthermore, the parameter calculator / parametric coder 104 is configured to generate a second coded representation 109 of a second set of second spectral portions. The first coded representation 107 and the second coded representation 109 are input to a bitstream multiplexer or bitstream former 108, which finally outputs a coded audio signal for transmission or storage in a storage device.
[0110] Typically, a first spectral portion such as 306 in Figure 3a would be surrounded by two second spectral portions such as 307a and 307b, but this is not the case for HE-AAC, where the core encoder frequency range is band-limited.
[0111] Figure 1b shows a decoder compatible with the encoder of Figure 1a. The first coded representation 107 is input to a spectral domain audio decoder 112, which generates a first decoded representation of a first set of first spectral portions, the first decoded representation having a first spectral resolution. Furthermore, the second coded representation 109 is input to a parametric decoder 114, which generates a second decoded representation of a second set of second spectral portions, the second decoded representation having a second spectral resolution that is lower than the first spectral resolution.
[0112] The decoder includes a frequency regeneration unit 116 that uses the first spectral portion to regenerate a reconstructed second spectral portion having a first spectral resolution. The frequency regeneration unit 116 performs a tile-filling operation, i.e., uses tiles or portions of the first spectral portion, copies the first set of first spectral portions into a reconstruction region or band having the second spectral portion, and typically performs spectral envelope shaping or other operations using information indicated by the decoded second representation output by the parametric decoder 114, i.e., the second set of second spectral portions. The decoded first set of first spectral portions and the reconstructed second set of spectral portions, indicated by line 117 at the output of the frequency regeneration unit 116, are input to a spectrum-to-time conversion unit 118, which converts the first decoded representation and the reconstructed second spectral portion into a time representation 119, i.e., a time representation having a higher sampling rate.
[0113] Figure 2b shows an embodiment of the encoder of Figure 1a. An audio input signal 99 is input to an analysis filter bank 220, which corresponds to the time-to-frequency transform unit 100 of Figure 1a. Next, a temporal noise shaping operation is performed in a TNS block 222. Thus, the input to the spectral analysis unit 102 of Figure 1a, which corresponds to the tonality mask block 226 of Figure 2b, can be full spectral values if no temporal noise shaping / temporal tile shaping operation is applied, or spectral residual values if a TNS operation is applied as shown in block 222 of Figure 2b. For two-channel or multi-channel signals, a joint channel coding 228 can additionally be performed, and the spectral domain encoder 106 of Figure 1a can include the joint channel coding block 228. Furthermore, an entropy encoder 232 for performing lossless data compression is provided and is also part of the spectral domain encoder 106 of Figure 1a.
[0114] Spectral analyzer / tonal mask 226 separates the output of TNS block 222 into core band and tonal components corresponding to first set of first spectral portions 103 in Figure 1a and residual components corresponding to second set of second spectral portions 105 in Figure 1a. Block 224, designated IGF parameter extraction coding, corresponds to parametric coder 104 in Figure 1a, and bitstream multiplexer 230 corresponds to bitstream multiplexer 108 in Figure 1a.
[0115] Preferably, the analysis filter bank 222 is configured as an MDCT (Modified Discrete Cosine Transform filter bank), which is used to transform the signal 99 into the time-frequency domain using the Modified Discrete Cosine Transform, which acts as a frequency analysis tool.
[0116] The spectral analyzer 226 preferably applies a tonal mask. This tonal mask estimation stage is used to separate tonal components from noise-like components in the signal, allowing the core encoder 228 to encode all tonal components using a psychoacoustic module. The tonal mask estimation stage can be configured in a number of different ways, and is preferably configured to be similar in function to the sinusoidal track estimation stages used in sine wave for speech / audio coding and noise modeling or HILN model-based audio encoders. Preferably, an easy-to-build configuration is used, without the need to maintain birth-death trajectories, although any other tonal or noise detector could also be used.
[0117] The IGF module calculates the similarity between the source and target regions. The target region will be represented by the spectrum from the source region. The measurement of similarity between the source and target regions is performed by the method of cross-correlation. The target region is divided into nTar non-overlapping frequency tiles. For every tile in the target region, nSrc source tiles are created from a fixed starting frequency. These source tiles overlap by a factor between 0 and 1, where 0 means 0% overlap and 1 means 100% overlap. Each of these source tiles is correlated with the target tile at various lags to find the source tile that best matches the target tile. The best-matching tile number is stored in tileNum[idx_tar], the lag at which it best correlates with the target is stored in xcorr_lag[idx_tar][idx_src], and the sign of the correlation is stored in xcorr_sign[idx_tar][idx_src]. If the correlation is highly negative, the source tile must be multiplied by -1 before the tile filling process in the decoder. Since the tonal components are preserved using the tonal mask, the IGF module also manages not to overwrite the tonal components in the spectrum. The per-band energy parameters are used to store the energy of the target region, which allows for an accurate reconstruction of the spectrum.
[0118] This method has the advantage over the classical SBR method described in [1] that the harmonic grid of the multi-tone signal is maintained by the core coder, while only the gaps between the sinusoids are filled with the best-matched "shaped noise" from the source domain. Another advantage of this system compared to Accurate Spectral Replacement (ASR) [2-4] is that there is no signal synthesis stage in the decoder that creates the important parts of the signal. Instead, this task is carried out by the core coder, allowing the preservation of important components of the spectrum. Another advantage of the proposed system is the continuous scalability that its features offer. Using tileNum[idx_tar] and xcorr_lag=0 for all tiles, referred to as gross granularity matching, can be used for low bitrates, while using the variable xcorr_lag for all tiles allows for a better match between the target and source spectrum.
[0119] Additionally, we propose a tile-selective stabilization technique to remove frequency-domain artifacts such as trilling and music noise.
[0120] In the case of stereo channel pairs, additional joint stereo processing is applied. This processing is necessary because, for certain destination ranges, signals may be highly correlated panned sources. If the source regions selected for this particular range are not well correlated, the spatial image may be adversely affected by uncorrelated source regions, even if the energy matches the destination range. The encoder typically performs cross-correlation of spectral values to analyze the energy bands of each destination range and sets a joint flag for this energy band if it exceeds a certain threshold. In the decoder, if the joint stereo flag is not set, the left and right channel energy bands are processed separately. If the joint stereo flag is set, both the energy and patching are performed in the joint stereo domain. The joint stereo information for the IGF region is signaled similarly to the joint stereo information for core coding, and for prediction, it includes a flag indicating whether the prediction direction is from downmix to residual or vice versa.
[0121] The energy can be calculated from the energy transmitted in the L / R domain. TIFF2026010016000002.tif21161 where k is the frequency index in the transform domain.
[0122] Another solution is to directly calculate and transmit the energy in the joint stereo domain for the bands where the joint stereo is active, so that no additional energy transformation is required at the decoder side.
[0123] The source tiles are always created according to the Mid / Side matrix. TIFF2026010016000003.tif21161
[0124] The energy adjustment is as follows: TIFF2026010016000004.tif21161
[0125] The joint stereo to LR conversion is as follows:
[0126] If no additional prediction parameters are coded: TIFF2026010016000005.tif21164
[0127] If an additional prediction parameter is coded and its signaled direction is mid to side: TIFF2026010016000006.tif26161
[0128] If the signaled direction is side to mid: TIFF2026010016000007.tif26161
[0129] Such processing ensures that from the tiles used to recreate highly correlated target regions and panned target regions, the resulting left and right channels represent correlated and panned sound sources, preserving the stereo image for such regions, even if the source regions are uncorrelated.
[0130] In other words, a joint stereo flag is transmitted in the bitstream to indicate whether, for example, L / R or M / S should be used for general joint stereo coding. At the decoder, the core signal is first decoded as indicated for the core band by the joint stereo flag. Then, the core signal is stored in both L / R and M / S representations. For IGF tile filling, the source tile representation is selected to match the target tile representation as indicated for the IGF band by the joint stereo information.
[0131] Temporal noise shaping (TNS) is a standard technique and is part of AAC [11-13]. TNS can be seen as an extension of the basic scheme of perceptual coders, inserting an optional processing step between the filter bank and the quantization stage. The main role of the TNS module is to hide the quantization noise of transient signals generated in the temporal masking domain, thereby resulting in a more efficient coding scheme. First, TNS calculates a set of prediction coefficients using "forward prediction" in the transform domain, e.g., MDCT. These coefficients are then used to flatten the temporal envelope of the signal. Since quantization affects the TNS-filtered spectrum, the quantization noise is also temporally flat. By applying inverse TNS filtering at the decoder side, the quantization noise is shaped according to the temporal envelope of the TNS filter, and thus the quantization noise is masked by the transients.
[0132] The IGF is based on the MDCT representation. For efficient coding, long blocks of about 20 ms should preferably be used. If the signal within such a long block contains transients, audible pre- and post-echoes will occur within the IGF spectral band due to tile filling.
[0133] This pre-echo effect is reduced by using TNS in the context of IGF. In this case, TNS is used as a temporal tile shaping (TTS) tool so that the spectral reconstruction at the decoder side is performed on the TNS residual signal. The required TTS prediction coefficients are calculated and applied using the full spectrum at the encoder side as usual. The start and stop frequencies of the TNS / TTS are determined by the IGF start frequency f IGFstart Compared to the legacy TNS, the stopping frequency of the TTS is increased to the stopping frequency of the IGF tool, which is f IGFstartAt the decoder side, the TNS / TTS coefficients are again applied to the whole spectrum, i.e., the core spectrum, the regenerated spectrum and the tonal components from the tonality mask (see Fig. 2a). The application of TTS is again necessary to make the temporal envelope of the regenerated spectrum match that of the original signal. In this way, the pre-echoes shown are reduced. In addition, f IGFstart Quantization noise in the signal below σ is also shaped as usual using TNS.
[0134] In legacy decoders, spectral patching of an audio signal destroys spectral correlations at patch boundaries and consequently impairs the temporal envelope of the audio signal by introducing variance. Therefore, another advantage of performing IGF tile filling on the residual signal is that after application of the shaping filter, the tile boundaries are seamlessly correlated, resulting in a more faithful temporal reproduction of the signal.
[0135] In the inventive encoder, the spectrum that has undergone TNS / TTS filtering, tonal masking, and IGF parameter estimation is devoid of any signal above the IGF start frequency, except for the tonal components. This sparse spectrum is then coded by a core encoder using principles of arithmetic and predictive coding. These coded components, together with their signaling bits, form the audio bitstream.
[0136] Figure 2a shows the corresponding decoder architecture. The bitstream of Figure 2a, corresponding to the coded audio signal, is input to a demultiplexer / decoder, which may be connected to blocks 112 and 114 in Figure 1b. The bitstream demultiplexer separates the input audio signal into a first coded representation 107 of Figure 1b and a second coded representation 109 of Figure 1b. The first coded representation, which comprises a first set of first spectral portions, is input to a joint channel decoding block 204, which corresponds to the spectral domain decoder 112 of Figure 1b. The second coded representation is input to a parametric decoder 114, not shown in Figure 2a, and then to an IGF block 202, which corresponds to the frequency regeneration unit 116 of Figure 1b. The first set of first spectral portions required for frequency regeneration are input to the IGF block 202 via line 203. Further, following the joint channel decoding 204, a specific core decoding is applied in a tonal mask block 206, whose output corresponds to the output of the spectral domain decoder 112. Next, combining, i.e., frame construction, is performed by a combiner 208, whose output now has a full spectrum, but still in the TNS / TTS filtered domain. Next, in block 210, an inverse TNS / TTS operation is performed using the TNS / TTS filter information provided via line 109. That is, the TTS side information is preferably included in the first coded representation generated by the spectral domain coder 106, which may be, for example, a simple AAC or USAC core coder, or it may be included in the second coded representation. At the output of block 210, the complete spectrum up to the maximum frequency is provided, which is the full frequency defined by the sampling rate of the original input signal. Next, a spectral-to-temporal transformation is performed in a synthesis filter bank 212, finally obtaining the audio output signal.
[0137] FIG. 3a shows a schematic representation of a spectrum. The spectrum is divided into multiple scale factor bands (SCBs), seven in the example shown in FIG. 3a: SCB1-SCB7. The scale factor bands may be AAC scale factor bands defined in the AAC standard, with the upper frequencies having larger bandwidths, as shown diagrammatically in FIG. 3a. Rather than performing intelligent gap filling from the beginning of the spectrum, i.e., at low frequencies, intelligent gap filling desirably begins its IGF operation at the IGF start frequency, indicated by reference numeral 309. Thus, the core frequency band extends from the lowest frequency to the IGF start frequency. Above the IGF start frequency, spectral analysis is applied to separate high-resolution spectral components 304, 305, 306, and 307 (the first set of first spectral portions) from low-resolution components represented by the second set of second spectral portions. FIG. 3a shows the spectrum input to, for example, the spectral-domain encoder 106 or the joint channel encoder 228. That is, the core encoder operates in the full range, but encodes a significant number of zero spectral values, which are either quantized to or set to zero before or after quantization. In either case, the core encoder operates in the full range, i.e., as if the spectrum were as shown. Meanwhile, the core decoder does not necessarily need to know about intelligent gap filling or about encoding a second set of second spectral portions with lower spectral resolution.
[0138] Preferably, the high resolution is defined by a line-by-line encoding of spectral lines, such as MDCT lines, while the second or lower resolution is defined, for example, by calculating only a single spectral value per scale factor band, where each scale factor band covers multiple frequency lines. Thus, the second or lower resolution is significantly lower in terms of its spectral resolution than the first or higher resolution, which is defined by a line-by-line encoding applied by a core encoder, such as an AAC or USAC core encoder.
[0139] Figure 3b shows the state for the scale factor or energy calculation. Due to the fact that the encoder is a core encoder, and that there may be, but is not necessarily, a first set of spectral components in each band, the core encoder not only calculates scale factors for each band in the core region below the IGF start frequency 309, but also for bands above the IGF start frequency, at half the sampling frequency, i.e., f S / 2 A maximum frequency F that is less than or equal to IGFstop Thus, the encoded tonality portions 302, 304, 305, 306, 307 of Fig. 3a and, in this embodiment, the scale factor bands SCB1 to SCB7, together correspond to the high-resolution spectral data. The low-resolution spectral data are calculated starting from the IGF start frequency and correspond to the energy information values E1, E2, E3, E4 transmitted together with the scale factors SF4 to SF7.
[0140] Particularly when the core encoder is in a low bitrate state, an additional noise-filling operation may be applied within the core band, i.e., at frequencies lower than the IGF start frequency, i.e., the scale factor bands SCB1-SCB3. In the noise-filling, multiple adjacent spectral lines are quantized to zero. At the decoder side, these zero-quantized spectral values are recombined, and their magnitudes are adjusted using a noise-filling energy, such as NF2, shown as 308 in FIG. 3b. The noise-filling energy can be given in absolute terms or relative to the scale factor, as in USAC, and corresponds to the energy of the set of zero-quantized spectral values. These noise-filling spectral lines may also be considered a third set of third spectral portions. These spectral portions are regenerated by simple noise-filling synthesis without any IGF operation, relying on frequency regeneration using frequency tiles from other frequencies to reconstruct the frequency tiles using spectral values and energy information E1, E2, E3, and E4 from the source domain.
[0141] Preferably, the bands for which the energy information is calculated coincide with the scale factor bands. In other embodiments, a grouping of energy information values is applied, e.g., only a single energy information value is transmitted for scale factor bands 4 and 5. However, also in this embodiment, the boundaries of the grouped restoration bands coincide with the boundaries of the scale factor bands. If a different band separation is applied, some recalculation or synchronization calculation may be applied, which may be reasonable depending on the given configuration.
[0142] Preferably, the spectral domain encoder 106 of Fig. 1a is a psychoacoustically driven encoder as shown in Fig. 4. Typically, the audio signal to be coded (401 in Fig. 4a) after being converted to the spectral domain, as in the MPEG2 / 4 AAC standard or the MPEG1 / 2 Layer 3 standard, is sent to a scale factor calculation unit 400. The scale factor calculation unit is controlled by a psychoacoustic model and additionally receives the audio signal to be quantized, or, as in the MPEG1 / 2 Layer 3 or MPEG AAC standard, receives a complex spectral representation of the audio signal. The psychoacoustic model calculates, for each scale factor band, a scale factor representing the psychoacoustic threshold. In addition, the scale factor is then adjusted by the cooperation of known inner and outer iterative loops or by any other suitable encoding process so that a predetermined bit rate requirement is met. Next, the spectral values to be quantized on the one hand and the calculated scale factor on the other hand are both input to a quantization unit 404. In a simple audio coder operation, the spectral values to be quantized are weighted by a scale factor, and the weighted spectral values are then input to a fixed quantizer, which typically has a compression function for the upper amplitude region. At the output of the quantizer, quantization indexes are then present, and these quantization indexes are then input to an entropy coder, which typically has a unique and highly efficient coding for sets of zero quantization indexes relative to adjacent frequency values, or "runs" of zero values as they are called in the industry.
[0143] However, in the audio encoder of Figure 1a, the quantizer 404 typically receives information about the second spectral portion from the spectrum analyzer. In this way, the quantizer 404 ensures that, at its output, the second spectral portion identified by the spectrum analyzer 102 is either zero or has a representation that is recognized by the encoder or decoder as a zero representation, and that these zeros can be coded very efficiently, especially when there are "runs" of zero values in the spectrum.
[0144] 4b shows a configuration of the quantization processor. MDCT spectral values may be input to a zeroing block 410. Thus, the second spectral portion is already set to zero before weighting by the scale factor is performed in block 412. In an additional configuration, block 410 is not provided, and the zeroing operation is performed in block 418 following the weighting block 412. In yet another configuration, the zeroing operation may also be performed in a zeroing block 422 following quantization in quantization block 420. In this configuration, blocks 410 and 418 would not exist. Generally, at least one of blocks 410, 418, and 422 will be provided depending on the particular configuration.
[0145] A quantized spectrum is then obtained at the output of block 422, which corresponds to the one shown in Figure 3a. This quantized spectrum is then input to an entropy coder such as 232 in Figure 2b, which may be a Huffman coder or an arithmetic coder as defined for example in the USAC standard.
[0146] The zeroing blocks 410, 418, 422, which are provided alternatively or in parallel with one another, are controlled by a spectrum analyzer 424. This spectrum analyzer preferably comprises any configuration of a known tonality detector or any different type of detector operable to separate the spectrum into components to be coded with high resolution and components to be coded with low resolution. Other such algorithms implemented in the spectrum analyzer could be a voice activity detector, a noise detector, a speech detector, or any other detector that relies on spectral information or associated metadata to decide on resolution requirements for different spectral portions.
[0147] FIG. 5a shows a preferred configuration of the time-spectral transform unit 100 of FIG. 1a, implemented, for example, in AAC or USAC. The time-spectral transform unit 100 includes a windowing unit 502 controlled by a transient detector 504. When the transient detector 504 detects a transient, a switch from a long window to a short window is signaled to the windowing unit. The windowing unit 502 calculates windowed frames for overlapping blocks, each windowed frame typically having 2N values, such as 2048 values. Next, a transform is performed in a block transform unit 506, which typically also provides truncation. Thus, a combined truncation / transform is performed to obtain a spectral frame having N values, such as MDCT spectral values. Thus, for a long windowing operation, the frame at the input of block 506 contains 2N values, such as 2048 values, and the spectral frame then has 1024 values. However, if a switch to short blocks is then made and eight short blocks are performed, each short block will have 1 / 8 the windowed time domain values compared to the long window, and each spectral block will have 1 / 8 the spectral values compared to the long block. Thus, when truncation is combined with a 50% windowing overlap operation, the spectrum becomes a critically sampled version of the time domain audio signal 99.
[0148] Next, reference is made to FIG. 5b, which shows a specific configuration of the frequency regeneration unit 116 and the spectral-to-time conversion unit 118 in FIG. 1b, or a specific configuration of the combined operations of blocks 208 and 212 in FIG. 2a. Consider a specific reconstruction band, such as scale factor band 6 in FIG. 3a, in FIG. 5b. A first spectral portion within this reconstruction band, i.e., first spectral portion 306 in FIG. 3a, is input to frame constructor / adjuster block 510. Furthermore, a reconstructed second spectral portion for scale factor band 6 is also input to frame constructor / adjuster block 510. Furthermore, energy information, such as E3 in FIG. 3b, for scale factor band 6 is also input to block 510. The reconstructed second spectral portion within the reconstruction band has already been generated by frequency tile filling using the source region, so the reconstruction band corresponds to the target region. Now, frame energy adjustment is performed to finally obtain a fully reconstructed frame having N values, such as that obtained at the output of combiner 208 in FIG. 2a. Next, in block 512, an inverse block transform / interpolation is performed, e.g., to obtain 248 time-domain values for the 124 spectral values at the input of block 512. Next, in block 514, a synthesis windowing operation is performed, which is also controlled by the long / short window indication transmitted as side information in the encoded audio signal. Next, in block 516, an overlap / add operation with the previous time frame is performed. Preferably, the MDCT applies a 50% overlap, so that N time-domain values are ultimately output for each new time frame of 2N values. The 50% overlap is highly preferred due to the fact that the overlap / add operation in block 516 provides critical sampling and continuous crossover from one frame to the next.
[0149] As shown in Figure 3a by reference numeral 301, the noise filling operation can be applied not only below the IGF start frequency, but also above the IGF start frequency, such as in the restoration band under consideration that corresponds to scale factor band 6 in Figure 3a. The noise filling spectral values can also be input to the frame builder / adjuster 510, and adjustment of the noise filling spectral values can also be applied within this block, or the noise filling spectral values can already be adjusted using the noise filling energy before being input to the frame builder / adjuster 510.
[0150] Preferably, the IGF operation, i.e., the frequency tile filling operation using spectral values from other parts, can be applied to the entire spectrum. Thus, the spectral tile filling operation can be applied not only to the high frequency band above the IGF start frequency, but also to the low frequency band. Furthermore, noise filling without frequency tile filling can also be applied not only to the low frequency band below the IGF start frequency, but also to the high frequency band above the IGF start frequency. However, it has been found that high-quality and efficient audio coding can be achieved when the noise filling operation is limited to the frequency region below the IGF start frequency and the frequency tile filling operation is limited to the frequency region above the IGF start frequency, as shown in Figure 3a.
[0151] Preferably, the target tile (TT) (with a frequency greater than the IGF start frequency) is bounded by the scale factor band boundary of the full-rate coder. The source tile (ST) (with a frequency lower than the IGF start frequency, i.e., the source tile) is not bounded by a scale factor band. The size of the ST should correspond to the size of the associated TT. An example is given below: TT[0] has a length of 10 MDCT bins. This corresponds exactly to the length of two consecutive SCBs (e.g., 4 + 6). In that case, all possible STs to be correlated with TT[0] also have a length of 10 bins. The second target tile TT[1] adjacent to TT[0] has a length of 15 bins (with SCBs of length 7 + 8). In that case, the ST for it has a length of 15 bins, rather than the 10 bins for TT[0].
[0152] If no TT with the length of the target tile for ST is found (for example, if the length of the TT is longer than the valid source area), the correlation is not calculated and the source area is copied into this TT again and again, successively in frequency such that the lowest frequency line of the second copy is next to the highest frequency line of the first copy, until the TT is completely filled.
[0153] Next, referring to FIG. 5c, a further preferred embodiment of the frequency regeneration unit 116 of FIG. 1b or the IGF block 202 of FIG. 2a will be described. Block 522 is a frequency tile generator that receives not only a target band ID but also a source band ID. For example, suppose the encoder determined that scale factor band 3 of FIG. 3a is a very good fit for reconstructing scale factor band 7. In that case, the source band ID would be 3 and the target band ID would be 7. Based on this information, frequency tile generator 522 applies a copy-up, harmonic tile-filling operation, or any other tile-filling operation to generate a second raw portion of spectral components 523. This second raw portion of spectral components has a frequency resolution equal to the frequency resolution contained in the first set of first spectral components.
[0154] Next, the first spectral portion of the reconstruction band, such as 307 in Fig. 3a, is input to a frame constructor 524, and the raw second portion 523 is also input to the frame constructor 524. The reconstructed frame is then adjusted by an adjuster 526 using a gain factor for the reconstruction band calculated by a gain factor calculator 528. Importantly, however, the first spectral portion in the frame is not affected by the adjuster 526; only the raw second portion for the reconstructed frame is affected by the adjuster 526. To this end, the gain factor calculator 528 analyzes the source band or raw second portion 523 and further analyzes the first spectral portion in the reconstruction band to finally find the correct gain factor 527, so that the energy of the adjusted frame output by the adjuster 526 has energy E4 when scale factor band 7 is considered.
[0155] In this context, it is very important to evaluate the accuracy of our high-frequency reconstruction in comparison with HE-AAC. This is illustrated with respect to scale factor band 7 in FIG. 3a. Suppose a prior art encoder detects the spectral portion 307 to be coded at high resolution as a "missing harmonic." The energy of this spectral component is then transmitted to the decoder along with the spectral envelope information for the reconstruction band, such as scale factor band 7. The decoder will then reconstruct this missing harmonic. However, the spectral value at which the missing harmonic 307 is reconstructed by the prior art decoder will be located in the center of frequency band 7, as indicated by reconstruction frequency 390. Thus, our invention avoids the frequency error 391 that would be introduced by the prior art decoder.
[0156] In one embodiment, the spectral analyzer is also configured to calculate a similarity between the first and second spectral portions and, based on the calculated similarity, determine, for the second spectral portion in the restoration domain, a first spectral portion that matches the second spectral portion as closely as possible. With this variable source / destination domain configuration, the parametric encoder will then additionally introduce matching information into the second encoded representation, indicating the matching source domain for each destination domain. At the decoder side, this information will be used by the frequency tile generator 522 of FIG. 5c, which generates the raw second portion 523 based on the source band ID and the destination band ID.
[0157] Furthermore, as shown in Figure 3a, the spectral analysis unit is configured to analyze the spectral representation up to a maximum analysis frequency, which is slightly less than half the sampling frequency and preferably at least one-quarter the sampling frequency, or typically greater.
[0158] As mentioned above, the encoder operates without downsampling and the decoder operates without upsampling, i.e., the spectral domain audio codec is configured to generate a spectral representation having a Nyquist frequency defined by the sampling rate of the original input audio signal.
[0159] As shown in Figure 3a, the spectral analysis unit is configured to analyze the spectral representation starting from the gap filling start frequency and stopping at a maximum frequency represented by the maximum frequency included in the spectral representation, wherein the spectral portion extending from the minimum frequency to the gap filling start frequency belongs to a first set of spectral portions, and further spectral portions such as 304, 305, 306, 307 having frequencies higher than the gap filling frequency are also included in the first set of first spectral portions.
[0160] As mentioned above, the spectral-domain audio decoder 112 is configured such that the maximum frequency represented by a spectral value in the first decoded representation is equal to the maximum frequency contained in the time representation having a certain sampling rate, and such that the spectral value for the maximum frequency in the first set of first spectral portions is zero or different from zero. In any case, for this maximum frequency in the first set of spectral components there is a scalefactor for the scalefactor band, which scalefactor is generated and transmitted regardless of whether all spectral values in this scalefactor band are set to zero, as described above in the context of Figures 3a and 3b.
[0161] The present invention therefore has the following advantages: in contrast to other parametric techniques for increasing compression efficiency, such as noise substitution and noise filling, which are exclusively used to efficiently represent noise-like signal content, the present invention allows accurate frequency reproduction of tonal components. To date, no current technology has disclosed a method for efficiently parametrically representing arbitrary signal content by spectral gap filling without the restriction of a fixed a priori division into low frequency bands (LF) and high frequency bands (HF).
[0162] Embodiments of the system of the present invention improve upon state-of-the-art approaches, resulting in high compression efficiency, zero or little perceptual disturbance, and full audio bandwidth even at low bit rates.
[0163] The overall system includes the following components: Full-band core coding Intelligent gap filling (tile filling or noise filling) Sparse tonal parts in the core selected by the tonal mask Joint stereo pair coding for all bands, including tile filling TNS on tiles Spectral whitening in the IGF region
[0164] The first step toward a more efficient system is to eliminate the need to transform spectral data into a second transform domain different from the transform domain of the core coder. Because mainstream audio codecs, such as AAC, use the MDCT as the basic transform, it would be beneficial to perform BWE in the MDCT domain as well. A second requirement for a BWE system would be the need to preserve the tonal grid. This ensures that even HF tonal components are preserved, resulting in superior quality of the coded audio compared to existing systems. Taking both of these requirements for a BWE scheme into account, we propose a novel system called Intelligent Gap Filling (IGF). Figure 2b shows the block diagram of the encoder side of the proposed system, while Figure 2a shows the decoder side.
[0165] Next, a full-band frequency domain first encoding processor and a full-band frequency domain decoding processor incorporating gap-filling operations, which may be configured separately or jointly, are described and defined.
[0166] In particular, the spectral domain decoder 112, corresponding to block 1122a, is configured to output a sequence of decoded frames of spectral values, where the decoded frames are first decoded representations, comprising spectral values for the first set of spectral portions and zero indications for the second spectral portions. The decoding apparatus further includes a combiner 208. Spectral values are generated by a frequency regenerator for the second set of second spectral portions, both of which are included in block 1122b. In this way, by combining the second spectral portions with the first spectral portions, a reconstructed spectral frame is obtained, comprising spectral values for the first set of first spectral portions and the second set of spectral portions. Then, the spectral-to-time transformer 118, corresponding to the IMDCT block 1124 of Fig. 14b, transforms the reconstructed spectral frame into a time representation.
[0167] As mentioned above, the spectral-to-temporal transform unit 118 or 1124 is configured to perform an inverse modified discrete cosine transform 512, 514 and further includes an overlap-add stage 516 for overlapping and adding subsequent time-domain frames.
[0168] In particular, the spectral-domain audio decoder 1122a is configured to generate a first decoded representation, the first decoded representation being configured to have a Nyquist frequency that defines a sampling rate equal to the sampling rate of the temporal representation generated by the spectral-to-temporal converter 1124.
[0169] Furthermore, the decoder 1112 or 1122a is configured to generate the first decoded representation such that the first spectral portion 306 is positioned in frequency between the two second spectral portions 307a and 307b.
[0170] In a further embodiment, the maximum frequency represented by the spectral value for the maximum frequency in the first decoded representation is equal to the maximum frequency included in the temporal representation produced by the spectral-to-temporal converter, and the spectral value for the maximum frequency in that first representation is zero or different from zero.
[0171] Further, as shown in FIG. 3, the encoded first audio signal portion further includes an encoded representation of a third set of third spectral portions to be restored by noise filling, and the first decoding processor 1120 further includes a noise filling unit included in block 1122b, which extracts noise filling information 308 from the encoded representation of the third set of third spectral portions and applies a noise filling operation on the third set of third spectral portions without using the first spectral portions in a different frequency region.
[0172] Furthermore, the spectral-domain audio decoder 112 is configured to generate a first decoded representation, the first decoded representation having a first spectral portion with frequency values greater than a frequency equal to a frequency located in the center of the frequency range covered by the time representation output by the spectral-to-time transform unit 118 or 1124.
[0173] Furthermore, the spectral analysis or fullband analysis unit 604 is configured to analyze the representation produced by the time-to-frequency transform unit 602 to determine a first set of first spectral portions to be coded at a first, higher spectral resolution and a second set of different second spectral portions to be coded at a second spectral resolution lower than the first spectral resolution, whereby the first spectral portion 306 is determined to be between two second spectral portions in terms of frequency, as shown at 307a and 307b in Figure 3.
[0174] In particular, the spectral analysis unit is configured to analyze the spectral representation up to a maximum analysis frequency that is at least 1 / 4 of the sampling frequency of the audio signal.
[0175] In particular, the spectral domain audio coder is configured to process a sequence of frames of spectral values for quantization and entropy coding, where in a frame the spectral values of a second set of a second portion are set to zero, or where in a frame there are a first set of spectral values of a first spectral portion and a second set of spectral values of a second spectral portion, and during subsequent processing the spectral values in the second set of spectral portions are set to zero as exemplarily shown at 410, 418, and 422.
[0176] The spectral domain audio encoder is configured to generate a spectral representation having a Nyquist frequency defined by the sampling rate of the audio input signal, or a first portion of the audio signal processed by a first encoding processor operating in the frequency domain.
[0177] The spectral-domain audio encoder 606 is further configured to provide a first encoded representation, where for a frame of the sampled audio signal, the encoded representation comprises a first set of first spectral portions and a second set of second spectral portions, and the spectral values in the second set of spectral portions are encoded as zero or noise values.
[0178] The full-band analyzer 604 or 102 calculates the maximum frequency f starting from the gap-filling start frequency 309 and represented by the maximum frequency contained in the spectral representation. max and a spectral portion extending from the minimum frequency to a gap filling start frequency 309 that belongs to the first set of first spectral portions.
[0179] In particular, the analysis unit applies a tonal masking operation to at least a portion of the spectral representation such that tonal and non-tonal components are separated from one another, wherein a first set of first spectral portions comprises the tonal components and a second set of second spectral portions comprises the non-tonal components.
[0180] Although the invention has been described above in the context of block diagrams, with each block representing an actual or logical hardware element, the invention may also be implemented as a computer-implemented method, with each block representing a corresponding method step, which represents a function performed by a corresponding logical or physical hardware block.
[0181] Although some aspects have been described above in the context of an apparatus, these aspects also represent a description of a corresponding method, and it is clear that a block or apparatus corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0182] The transmitted or encoded signals of the present invention can be stored on a digital storage medium or can be transmitted over a transmission medium, such as a wireless transmission medium like the Internet or a wired transmission medium.
[0183] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. This implementation can be implemented using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, flash memory, or the like, having electronically readable control signals stored therein and cooperating (or capable of cooperating) with a programmable computer system to perform the methods of the present invention. Thus, the digital storage medium can be computer-readable.
[0184] Some embodiments according to the invention include a data carrier having electronically readable control signals cooperable with a computer system programmable to carry out one of the methods described above.
[0185] Generally, embodiments of the present invention may be configured as a computer program product having program code operable to perform one of the methods of the present invention when the computer program product runs on a computer, the program code may for example be stored on a machine readable carrier.
[0186] Other embodiments of the invention comprise the computer program stored on a machine readable carrier for performing one of the methods described above.
[0187] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the above described methods, when the computer program runs on a computer.
[0188] Another embodiment of the invention is a data carrier (or a non-transitory storage medium such as a digital storage medium or a computer readable medium) comprising a computer program recorded thereon for performing one of the methods described above. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory.
[0189] Another embodiment of the invention represents a computer program for carrying out one of the methods described above. The data stream or sequence of signals may be adapted to be transmitted over a data communication connection, such as the Internet.
[0190] Another embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described above.
[0191] Another embodiment comprises a computer having installed thereon the computer program for performing one of the methods described above.
[0192] Further embodiments according to the invention include a device or a system configured to transmit (e.g. electronically or optically) a computer program for performing one of the above-mentioned methods to a receiver, which may be e.g. a computer, a mobile device, a memory device, etc. The device or system may for example comprise a file server for transmitting the computer program to the receiver.
[0193] In some embodiments, a programmable logic device (such as a field programmable gate array) may be used to perform some or all of the functions of the methods described above. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described above. In general, such methods may be suitably performed by any hardware apparatus.
[0194] The above-described embodiments are merely illustrative of the principles of the present invention. Modifications and variations in the arrangements and details described herein will be apparent to those skilled in the art. Therefore, the present invention is not to be limited by the specific details presented herein for purposes of illustration and description of the embodiments, but should be limited only by the scope of the appended claims. [remarks] [Claim 1] 1. An audio encoder for encoding an audio signal, comprising: a first encoding processor (600) for encoding a first audio signal portion in the frequency domain, a time-to-frequency transformer (602) for transforming the first audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the first audio signal portion; an analysis unit (604) that analyzes the frequency domain representation up to the maximum frequency to determine a first spectral portion to be coded at a first spectral resolution and a second spectral portion to be coded at a second spectral resolution lower than the first spectral resolution, the analysis unit (604) determining one first spectral portion (306) from the first spectral portions, the one first spectral portion being located between two second spectral portions (307a, 307b) from the second spectral portions in terms of frequency; a spectral encoder (606) for encoding the first spectral portion at the first spectral resolution and encoding the second spectral portion at the second spectral resolution, the spectral encoder (606) including a parametric encoder for calculating spectral envelope information having the second spectral resolution from the second spectral portion; a first encoding processor (600) having a second encoding processor (610) for encoding a second different audio signal portion in the time domain; a controller (620) that analyzes the audio signal and determines which portions of the audio signal are the first audio signal portions to be coded in the frequency domain and which portions of the audio signal are the second audio signal portions to be coded in the time domain; an encoded signal forming unit (630) for forming an encoded audio signal having a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion; an audio encoder including: [Claim 2] 2. The audio encoder according to claim 1, the input signal includes a high band and a low band; The second encoding processor (610) a sampling rate conversion unit (900) for converting the second audio signal portion to a representation with a lower sampling rate, the lower sampling rate being lower than the sampling rate of the audio signal, the lower sampling rate representation not including the higher frequency band of the input signal; a time domain lowband encoder (910) for time domain encoding said low sampling rate representation; a time domain bandwidth extension encoder (920) for parametrically encoding the high band; 1. An audio encoder having: [Claim 3] 3. An audio encoder according to claim 1, a pre-processing unit (1000) configured to pre-process the first audio signal portion and the second audio signal portion; The preprocessing unit includes a prediction analysis unit (1002) that determines prediction coefficients; the second encoding processor includes a prediction coefficient quantizer (1010) that generates a quantized version of the prediction coefficients, and an entropy encoder that generates a coded version of the quantized prediction coefficients; The encoded signal forming unit (630) is configured to introduce the encoded version into the encoded audio signal. [Claim 4] 4. An audio encoder according to claim 1, wherein: The pre-processing unit (1000) includes a resampler (1004) that resamples the audio signal to the sampling rate of the second encoding processor; and the prediction analyzer is configured to determine prediction coefficients using the resampled audio signal; or The pre-processing section (1000) further comprises a long-term prediction analysis stage (1006) for determining one or more long-term prediction parameters for the first audio signal portion. [Claim 5] 5. An audio encoder according to claim 1, wherein: an audio encoder further comprising a cross processor (700) for calculating initialization data for the second encoding processor (610) from the encoded spectral representation of the first audio signal portion so that the second encoding process (610) is initialized for encoding the second audio signal portion that immediately follows the first audio signal portion in time within the audio signal. [Claim 6] 6. The audio encoder of claim 5, wherein the cross processor (700) includes any of the following components: a spectral decoder (701) for computing a decoded version of the first encoded signal portion; a delay stage (707) that provides a delayed version of the decoded version to a de-emphasis stage (617) of the second encoding processor for initialization; a weighted prediction coefficient analysis filtering block (708) that provides the filter output to the codebook determination unit (613) of the second encoding processor (610) for initialization; an analysis filtering stage (706) for filtering the decoded or pre-emphasized (709) version and providing the filter residual to an adaptive codebook determiner (612) of the second encoding processor for initialization; or A pre-emphasis filter (709) that filters the decoded version and provides a delayed or pre-emphasized version to the synthesis filtering stage (616) of the second encoding processor (610) for initialization. [Claim 7] 7. An audio encoder according to claim 1, wherein: the analyzer (604) is configured to perform a temporal tile shaping, a temporal noise shaping analysis, or an operation of setting spectral values in the second spectral portion to zero; the first encoding processor (600) is configured to perform shaping (606a) of the spectral values of the first spectral portion using prediction coefficients (1010) derived from the first audio signal portion, and to perform quantization and entropy coding operations (606b) of the shaped spectral values of the first spectral portion; The spectral values of the second spectral portion are set to zero. [Claim 8] 8. The audio encoder of claim 7, further comprising a cross processor, the cross processor (700) comprising: a noise shaping unit (703) for shaping the quantized spectral values of the first spectral portion using LPC coefficients (1010) derived from the first audio signal portion; a spectral decoder (704, 705) for decoding the spectrally shaped spectral portion of the first spectral portion with high spectral resolution and synthesizing a second spectral portion using the parametric representation of the second spectral portion and the at least one decoded first spectral portion to obtain a decoded spectral representation; a frequency-to-time transform unit (702) for transforming the spectral representation into a time domain to obtain a decoded first audio signal portion, wherein a sampling rate associated with the decoded first audio signal portion is different from a sampling rate of the audio signal, and a sampling rate associated with an output signal of the frequency-to-time transform unit (702) is different from a sampling rate of the audio signal input to the frequency-to-time transform unit (602); an audio encoder, [Claim 9] 9. An audio encoder according to any one of claims 1 to 8, an audio encoder, wherein the second encoding processor includes at least one of the following blocks: Predictive Analytics Filter (611); an adaptive codebook stage (612); Innovative codebook stage (614); an estimation unit (613) for estimating innovative codebook entries; ACELP / Gain Encoding Stage (615); a predictive synthesis filtering stage (616); De-emphasis stage (617); Bass post-filter analysis stage (618). [Claim 10] 10. An audio encoder according to any one of claims 1 to 9, the time domain coding processor has an associated second sampling rate; the frequency domain coding processor has an associated first sampling rate that is higher than the second sampling rate; the audio encoder further comprises a cross processor (700) for calculating initialization data for the second encoding processor from the encoded spectral representation of the first audio signal portion; the cross processor includes a frequency-to-time converter (702) for generating a time domain signal at the second sampling rate; The frequency-time conversion unit (702) a selection unit (726) that selects a low-frequency portion of the spectrum input to the frequency-to-time conversion unit in accordance with a ratio of the first sampling rate to the second sampling rate that is less than 1; a transform processor (720) having a transform length smaller than the transform length of the time-frequency transform unit (602); a synthesis windowing unit (712) that performs windowing using a window having fewer window coefficients than the window used by the time-frequency conversion unit (602); Audio encoder. [Claim 11] 1. An audio decoder for decoding an encoded audio signal, the audio decoder comprising: a first decoding processor (1120) for decoding the first encoded audio signal portion in the frequency domain, a spectral decoder (1122) configured to obtain a decoded spectral representation by decoding first spectral portions with high spectral resolution and synthesizing second spectral portions using parametric representations of the second spectral portions and at least one decoded first spectral portion, the spectral decoder (1122) being configured to generate the first decoded representation such that one first spectral portion (306) is located between two second spectral portions (307a, 307b) in terms of frequency; a frequency-to-time transform unit (1120) for transforming the decoded spectral representation into the time domain to obtain a decoded first audio signal portion; a first decoding processor (1120) including: a second decoding processor (1140) for decoding the second encoded audio signal portion in the time domain to obtain a decoded second audio signal portion; a combiner (1160) for combining the decoded first spectral portion and the decoded second spectral portion to obtain a decoded audio signal; [Claim 12] 12. The audio decoder of claim 11, wherein the second decoding processor: a time domain low-band decoder (1200) for decoding a low-band time domain signal; an upsampler (1210) for upsampling the low-band time domain signal; a time domain bandwidth extension decoder (1220) for synthesizing a high band of the time domain output signal; a mixer (1230) for mixing the synthesized high-band of the time-domain signal with the upsampled low-band time-domain signal; an audio decoder, [Claim 13] 13. An audio decoder according to claim 12, An audio decoder, wherein the upsampler (1210) includes an analysis filter bank (1471) operating at a first time-domain low-band decoder sampling rate and a synthesis filter bank (1473) operating at a second output sampling rate higher than the first time-domain low-band sampling rate. [Claim 14] 14. An audio decoder according to claim 12 or 13, The time-domain low-band decoder (1200) includes a residual signal, decoders (1149, 1141, 1142), and a synthesis filter (1143) that filters the residual signal using synthesis filter coefficients (1145); an audio decoder, wherein the time-domain bandwidth extension decoder (1220) is configured to upsample (1221) the residual signal, process (1222) the upsampled residual signal using a nonlinear operation to obtain a high-band residual signal, and spectrally shape (1223) the high-band residual signal to obtain a synthesized high-band; [Claim 15] 15. An audio decoder according to any one of claims 11 to 14, An audio decoder, wherein the first decoding processor (1120) includes an adaptive long-term prediction postfilter (1420) that post-filters the first decoded first signal portion, the filter (1420) being controlled by one or more long-term prediction parameters included in the encoded audio signal. [Claim 16] 16. An audio decoder according to any one of claims 11 to 15, An audio decoder further comprising a cross processor (1170) for calculating initialization data for the second decoding processor (1140) from the decoded spectral representation of the first encoded audio signal portion so that the second decoding processor (1140) is initialized for decoding the encoded second audio signal portion that temporally follows the first audio signal portion within the encoded audio signal. [Claim 17] 17. The audio decoder of claim 16, wherein the cross processor further comprises the following components: an additional frequency-to-time conversion unit (1170) operating at a lower sampling rate than the frequency-to-time conversion unit (1124) of the first decoding processor (1120) to obtain an additional decoded first signal portion in the time domain, wherein the signal output by the frequency-to-time conversion unit (1171) has a second sampling rate that is lower than a first sampling rate associated with the output of the frequency-to-time conversion unit (1124) of the second decoding processor, the additional frequency-to-time conversion unit (1171) including a selection unit (726) that selects a lower portion of the spectrum to be input to the additional frequency-to-time conversion unit (1171) in accordance with a ratio of the first sampling rate to the second sampling rate that is less than 1; a transform processor (720) having a transform length smaller than the transform length (710) of the frequency-to-time transform unit (1124); A synthesis windowing unit (722) that uses a window having a smaller number of coefficients than the window used by the frequency-to-time transform unit (1124). [Claim 18] 18. An audio decoder according to claim 16 or 17, wherein the cross processor (1170) comprises the following components: a delay stage (1172) for delaying the additional decoded first signal portion and providing a delayed version of the decoded first signal portion to a de-emphasis stage (1144) of the second decoding processor for initialization; a pre-emphasis filter (1173) and delay stage (1175) for filtering and delaying the additional decoded first signal portion for initialization and providing the delay stage output to the predictive synthesis filter (1143) of the second decoding processor; a predictive analysis filter (1174) for generating a predictive residual signal from the additional decoded first signal portion or the pre-emphasized (1173) additional decoded first signal portion, and supplying the predictive residual signal to a codebook synthesis unit (1141) of the second decoding processor (1200); or A switch (1480) that supplies the additional decoded first signal portion to an analysis stage (1471) of a resampler (1210) of the second decoding processor for initialization. [Claim 19] 19. An audio decoder according to any one of claims 11 to 18, An audio decoder, wherein the second decoding processor (1200) includes at least one of the following blocks: ACELP decoding gain and innovative codebooks; Adaptive codebook synthesis stage (1141); ACELP post-processing unit (1142); PredictionSynthesisFilter(1143); De-emphasis stage (1144). [Claim 20] 1. A method for encoding an audio signal, the method comprising the steps of: A step (600) of first encoding a first audio signal portion in the frequency domain, comprising: a substep (602) of transforming the first audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the first audio signal portion; a substep (604) of analyzing the frequency domain representation up to the maximum frequency and determining a first spectral portion to be coded at a first spectral resolution and a second spectral portion to be coded at a second spectral resolution lower than the first spectral resolution, wherein one first spectral portion (306) is determined from the first spectral portions, and the one first spectral portion is determined to be located between two second spectral portions (307a, 307b) from the second spectral portions in terms of frequency; a substep (606) of encoding the first spectral portion at the first spectral resolution and encoding the second spectral portion at the second spectral resolution, wherein encoding the second spectral portion comprises calculating spectral envelope information having the second spectral resolution from the second spectral portion; a first encoding step (600); second encoding (610) a second different audio signal portion in the time domain; analyzing the audio signal to determine which portions of the audio signal are the first audio signal portions to be coded in the frequency domain and which portions of the audio signal are the second audio signal portions to be coded in the time domain (620); Forming (630) an encoded audio signal having a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion. [Claim 21] 1. A method for decoding an encoded audio signal, comprising the steps of: A first decoding (1120) of the first encoded audio signal portion in the frequency domain, comprising: a sub-step (1122) of decoding first spectral portions with high spectral resolution and synthesizing second spectral portions using parametric representations of the second spectral portions and at least one decoded first spectral portion to obtain a decoded spectral representation, the sub-step (1122) comprising generating the first decoded representation such that one first spectral portion (306) is located between two second spectral portions (307a, 307b) in terms of frequency; a substep (1120) of transforming the decoded spectral representation into the time domain to obtain a decoded first audio signal portion; a first decoding step (1120) having second decoding the second encoded audio signal portion in the time domain to obtain a decoded second audio signal portion (1140); Combining (1160) the decoded first spectral portion and the decoded second spectral portion to obtain a decoded audio signal. [Claim 22] 22. A computer program for carrying out the method according to claim 20 or claim 21 when running on a computer or processor.
Claims
1. 1. An audio encoder for encoding an audio signal, comprising: A first encoding processor (600) for encoding a first audio signal portion in the frequency domain, comprising: a time-to-frequency transformer (602) for transforming the first audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the first audio signal portion; an analysis unit (604) that analyzes the frequency domain representation up to the maximum frequency to determine a first spectral portion to be coded at a first spectral resolution and a second spectral portion to be coded at a second spectral resolution lower than the first spectral resolution, wherein the analysis unit (604) determines one first spectral portion (306) from the first spectral portions, and determines that the one first spectral portion is located between two second spectral portions (307 a, 307 b) from the second spectral portions in terms of frequency; a spectral encoder (606) for encoding the first spectral portion at the first spectral resolution and encoding the second spectral portion at the second spectral resolution, the spectral encoder (606) including a parametric encoder for calculating spectral envelope information having the second spectral resolution from the second spectral portion; a first encoding processor (600) having: a second encoding processor (610) for encoding a second different audio signal portion in the time domain; a controller (620) that analyzes the audio signal and determines which portions of the audio signal are the first audio signal portions to be coded in the frequency domain and which portions of the audio signal are the second audio signal portions to be coded in the time domain; an encoded signal forming unit (630) for forming an encoded audio signal having a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion; an audio encoder including:
2. 2. The audio encoder of claim 1, the input signal includes a high band and a low band; The second encoding processor (610) a sampling rate conversion unit (900) for converting the second audio signal portion to a representation at a lower sampling rate, the lower sampling rate being lower than the sampling rate of the audio signal, the lower sampling rate representation not including the higher band of the input signal; a time-domain low-band coder (910) for time-domain coding said low-sampling-rate representation; a time domain bandwidth extension encoder (920) for parametrically encoding the high band; 1. An audio encoder having:
3. 3. An audio encoder according to claim 1, a pre-processing unit (1000) configured to pre-process the first audio signal portion and the second audio signal portion; The preprocessing unit includes a prediction analysis unit (1002) for determining prediction coefficients; the second encoding processor includes a prediction coefficient quantizer (1010) that generates a quantized version of the prediction coefficients, and an entropy encoder that generates a coded version of the quantized prediction coefficients; An audio encoder, wherein the encoded signal former (630) is configured to introduce the encoded version into the encoded audio signal.
4. 4. An audio encoder according to claim 1, wherein: The pre-processing unit (1000) includes a resampler (1004) for resampling the audio signal to the sampling rate of the second encoding processor; and the prediction analyzer is configured to determine prediction coefficients using the resampled audio signal; or The audio encoder, wherein the pre-processing section (1000) further comprises a long-term prediction analysis stage (1006) for determining one or more long-term prediction parameters for the first audio signal portion.
5. 5. An audio encoder according to claim 1, wherein:
1. An audio encoder comprising: a cross processor (700) for calculating initialization data for the second encoding processor (610) from the encoded spectral representation of the first audio signal portion, such that the second encoding process (610) is initialized for encoding the second audio signal portion that immediately follows the first audio signal portion in time within the audio signal.
6. 6. The audio encoder of claim 5, wherein the cross processor (700) includes any of the following components: a spectral decoder (701) for computing a decoded version of said first encoded signal portion; a delay stage (707) for providing a delayed version of the decoded version to a de-emphasis stage (617) of the second encoding processor for initialization; a weighted prediction coefficient analysis filtering block (708) that provides the filter output to the codebook determination unit (613) of the second encoding processor (610) for initialization; an analysis filtering stage (706) for filtering the decoded or pre-emphasized (709) version and providing the filter residual to the adaptive codebook determiner (612) of the second encoding processor for initialization; or A pre-emphasis filter (709) that filters the decoded version and provides a delayed or pre-emphasized version to the synthesis filtering stage (616) of the second encoding processor (610) for initialization.
7. 7. An audio encoder according to claim 1, wherein: the analyzer (604) is configured to perform a temporal tile shaping, a temporal noise shaping analysis, or an operation of setting spectral values in the second spectral portion to zero; the first encoding processor (600) is configured to perform shaping (606a) of the spectral values of the first spectral portion using prediction coefficients (1010) derived from the first audio signal portion, and to perform quantization and entropy coding operations (606b) of the shaped spectral values of the first spectral portion; The spectral values of the second spectral portion are set to zero.
8. 8. The audio encoder of claim 7, further comprising a cross processor, the cross processor (700) comprising: a noise shaping unit (703) for shaping quantized spectral values of the first spectral portion using LPC coefficients (1010) derived from the first audio signal portion; a spectral decoder (704, 705) for decoding the spectrally shaped spectral portion of the first spectral portion with high spectral resolution and synthesizing a second spectral portion using the parametric representation of the second spectral portion and the at least one decoded first spectral portion to obtain a decoded spectral representation; a frequency-to-time transform unit (702) for transforming the spectral representation into a time domain to obtain a decoded first audio signal portion, wherein a sampling rate associated with the decoded first audio signal portion is different from a sampling rate of the audio signal, and a sampling rate associated with an output signal of the frequency-to-time transform unit (702) is different from a sampling rate of the audio signal input to the frequency-to-time transform unit (602); an audio encoder,
9. 9. An audio encoder according to any one of claims 1 to 8, an audio encoder, wherein the second encoding processor includes at least one of the following blocks: Predictive analytics filter (611); Adaptive codebook stage (612); Innovative Codebook Stage (614); an estimator (613) for estimating innovative codebook entries; ACELP / Gain Encoding Stage (615); a predictive synthesis filtering stage (616); De-emphasis stage (617); Bass Post-Filter Analysis Stage (618).
10. 10. An audio encoder according to any one of claims 1 to 9, the time domain coding processor has an associated second sampling rate; the frequency domain coding processor has an associated first sampling rate that is higher than the second sampling rate; the audio encoder further comprises a cross processor (700) for calculating initialization data for the second encoding processor from the encoded spectral representation of the first audio signal portion, the cross processor includes a frequency-to-time converter (702) for generating a time domain signal at the second sampling rate; The frequency-time conversion unit (702) a selection unit (726) for selecting a low frequency portion of the spectrum input to the frequency-to-time conversion unit in accordance with a ratio of the first sampling rate to the second sampling rate that is less than 1; a transform processor (720) having a transform length smaller than that of the time-frequency transform unit (602); a synthesis windowing unit (712) that performs windowing using a window having fewer window coefficients than the window used by the time-frequency transform unit (602); Audio encoder.
11. 1. An audio decoder for decoding an encoded audio signal, the audio decoder comprising: a first decoding processor (1120) for decoding a first encoded audio signal portion in the frequency domain, the first decoding processor comprising: a spectral decoder (1122) configured to obtain a decoded spectral representation by decoding first spectral portions with high spectral resolution and synthesizing second spectral portions using parametric representations of the second spectral portions and at least one decoded first spectral portion, the spectral decoder (1122) being configured to generate the first decoded representation such that one first spectral portion (306) is located between two second spectral portions (307a, 307b) in terms of frequency; a frequency-to-time transform unit (1120) for transforming the decoded spectral representation into the time domain to obtain a decoded first audio signal portion; a first decoding processor (1120) including: a second decoding processor (1140) for decoding the second encoded audio signal portion in the time domain to obtain a decoded second audio signal portion; A combiner (1160) for combining the decoded first spectral portion and the decoded second spectral portion to obtain a decoded audio signal.
12. 12. The audio decoder of claim 11, wherein the second decoding processor: a time domain low-band decoder (1200) for decoding a low-band time domain signal; an upsampler (1210) for upsampling the low-band time domain signal; a time domain bandwidth extension decoder (1220) for synthesizing the high band of the time domain output signal; a mixer (1230) for mixing the synthesized high-band of the time-domain signal with the upsampled low-band time-domain signal; an audio decoder,
13. 13. An audio decoder according to claim 12, 1. An audio decoder, wherein the upsampler (1210) includes an analysis filter bank (1471) operating at a first time-domain low-band decoder sampling rate and a synthesis filter bank (1473) operating at a second output sampling rate higher than the first time-domain low-band sampling rate.
14. 14. An audio decoder according to claim 12 or 13, the time-domain low-band decoder (1200) includes a residual signal, decoders (1149, 1141, 1142), and a synthesis filter (1143) that filters the residual signal using synthesis filter coefficients (1145); 12. An audio decoder, wherein the time-domain bandwidth extension decoder (1220) is configured to upsample (1221) the residual signal, process (1222) the upsampled residual signal using a nonlinear operation to obtain a high-band residual signal, and spectrally shape (1223) the high-band residual signal to obtain a synthesized high-band.
15. 15. An audio decoder according to any one of claims 11 to 14, An audio decoder, wherein the first decoding processor (1120) includes an adaptive long-term prediction postfilter (1420) that post-filters the first decoded first signal portion, the filter (1420) being controlled by one or more long-term prediction parameters included in the encoded audio signal.
16. 16. An audio decoder according to any one of claims 11 to 15, An audio decoder further comprising a cross processor (1170) for calculating initialization data for the second decoding processor (1140) from the decoded spectral representation of the first encoded audio signal portion so that the second decoding processor (1140) is initialized for decoding the encoded second audio signal portion that temporally follows the first audio signal portion within the encoded audio signal.
17. 17. The audio decoder of claim 16, wherein the cross processor further comprises the following components: a frequency-to-time conversion unit (1170) operating at a lower sampling rate than the frequency-to-time conversion unit (1124) of the first decoding processor (1120) to obtain an additional decoded first signal portion in the time domain, wherein the signal output by the frequency-to-time conversion unit (1171) has a second sampling rate lower than a first sampling rate associated with the output of the frequency-to-time conversion unit (1124) of the second decoding processor, the additional frequency-to-time conversion unit (1171) including a selection unit (726) for selecting a lower part of the spectrum to be input to the additional frequency-to-time conversion unit (1171) in accordance with a ratio of the first sampling rate to the second sampling rate that is less than 1; a transform processor (720) having a transform length smaller than the transform length (710) of said frequency-to-time transform unit (1124); A synthesis windowing unit (722) that uses a window that has a smaller number of coefficients than the window used by the frequency-to-time transform unit (1124).
18. 18. An audio decoder according to claim 16 or 17, wherein the cross processor (1170) comprises the following components: a delay stage (1172) for delaying the additional decoded first signal portion and providing a delayed version of the decoded first signal portion to a de-emphasis stage (1144) of the second decoding processor for initialization; a pre-emphasis filter (1173) and delay stage (1175) for filtering and delaying the additional decoded first signal portion for initialization and providing a delay stage output to a predictive synthesis filter (1143) of the second decoding processor; a predictive analysis filter (1174) for generating a predictive residual signal from the additional decoded first signal portion or the pre-emphasized (1173) additional decoded first signal portion and feeding the predictive residual signal to a codebook synthesis unit (1141) of the second decoding processor (1200); or A switch (1480) for supplying the additional decoded first signal portion to the analysis stage (1471) of the resampler (1210) of the second decoding processor for initialization.
19. 19. An audio decoder according to any one of claims 11 to 18, An audio decoder, wherein the second decoding processor (1200) includes at least one of the following blocks: ACELP decoding gain and innovative codebook; Adaptive codebook synthesis stage (1141); ACELP post-processing unit (1142); Prediction synthesis filter (1143); De-emphasis stage (1144).
20. 1. A method for encoding an audio signal, comprising the steps of: A step (600) of first encoding a first audio signal portion in the frequency domain, comprising: transforming (602) the first audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the first audio signal portion; a substep (604) of analyzing the frequency domain representation up to the maximum frequency and determining a first spectral portion to be coded at a first spectral resolution and a second spectral portion to be coded at a second spectral resolution lower than the first spectral resolution, wherein one first spectral portion (306) is determined from the first spectral portions, and the one first spectral portion is determined to be located between two second spectral portions (307a, 307b) from the second spectral portions in terms of frequency; a sub-step (606) of encoding the first spectral portion at the first spectral resolution and encoding the second spectral portion at the second spectral resolution, wherein encoding the second spectral portion comprises calculating spectral envelope information having the second spectral resolution from the second spectral portion; a first encoding step (600) comprising: second encoding (610) a second different audio signal portion in the time domain; analyzing the audio signal to determine (620) which portions of the audio signal are the first audio signal portions to be coded in the frequency domain and which portions of the audio signal are the second audio signal portions to be coded in the time domain; Forming (630) an encoded audio signal having a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion.
21. 1. A method for decoding an encoded audio signal, comprising the steps of: A first decoding (1120) of the first encoded audio signal portion in the frequency domain, comprising: a sub-step (1122) of decoding first spectral portions with high spectral resolution and synthesizing second spectral portions using parametric representations of the second spectral portions and at least one decoded first spectral portion to obtain a decoded spectral representation, the sub-step (1122) comprising generating the first decoded representation such that one first spectral portion (306) is located between two second spectral portions (307 a, 307 b) in terms of frequency; a substep (1120) of transforming the decoded spectral representation into the time domain to obtain a decoded first audio signal portion; a first decoding step (1120) having second decoding (1140) the second encoded audio signal portion in the time domain to obtain a decoded second audio signal portion; Combining (1160) the decoded first spectral portion and the decoded second spectral portion to obtain a decoded audio signal.
22. A computer program for carrying out the method according to claim 20 or 21 when running on a computer or processor.
Citation Information
Patent Citations
Apparatus and method for encoding or decoding an audio signal with intelligent gap filling in the spectral domain
WO2015010948A1