Audio encoders and decoders with frequency domain processors and time domain processors

Through the intelligent gap filling technology that combines the full-band spectrum encoder with the time domain encoder, the problems of reduced audio quality and low coding efficiency in the high frequency band of the audio encoder in the existing technology are solved, and efficient coding and seamless switching of full-band audio signals are achieved.

CN113963704BActive Publication Date: 2025-10-10FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111184409.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2014-07-28
Filing Date
2015-07-24
Publication Date
2025-10-10
Estimated Expiration
2035-07-24

AI Technical Summary

Technical Problem

Existing audio coding technologies degrade audio quality when processing non-speech signals with prominent harmonics in the high-frequency band, and frequency domain encoders introduce inflexibility and band limitations when expanding bandwidth, resulting in low coding efficiency.

Method used

A full-band spectrum encoder is combined with a time domain encoder. The high frequency part is encoded in the frequency domain through intelligent gap filling technology. The full-band spectrum encoder is used for high-resolution encoding. The spectrum gaps are filled in a parametric manner on the decoder side to achieve efficient encoding and decoding of the entire audio signal range.

Benefits of technology

It achieves efficient encoding of full-band audio signals, improves audio quality, avoids band limitations and inflexibility, supports seamless switching to time domain bandwidth expansion, and improves coding efficiency and audio reconstruction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113963704B_ABST
    Figure CN113963704B_ABST
Patent Text Reader

Abstract

An audio encoder for encoding an audio signal, comprising a first encoding processor (600) for encoding a first audio signal portion in the frequency domain, the first encoding processor (600) comprising a time-to-frequency converter (602), an analyzer (604) for analyzing a frequency domain representation up to a maximum frequency to determine a first spectral portion to be encoded with a first spectral resolution and a second spectral region to be encoded with a second spectral resolution, the second spectral resolution being lower than the first spectral resolution, a spectral encoder (606) for encoding the first spectral portion with the first spectral resolution and for encoding the second spectral portion with the second spectral resolution, a second encoding processor (610) for encoding a second, different audio signal portion in the time domain, a controller (620), and an encoded signal former (630).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese invention patent application "Audio encoder and decoder using a frequency domain processor with full bandgap filling and a time domain processor" with application number 201580049740.7 and application date on March 15, 2017. Technical Field

[0002] The present invention relates to audio signal encoding and decoding, and in particular to audio signal processing using parallel frequency domain and time domain encoder / decoder processors. Background Art

[0003] Perceptual coding of audio signals is a widely used practice for the purpose of reducing data for efficient storage or transmission of audio signals. In particular, when the goal is to achieve the lowest possible bit rate, the coding employed results in a reduction in audio quality, which is often primarily due to encoder-side limitations on the bandwidth of the audio signal to be transmitted. Here, the audio signal is typically low-pass filtered so that no spectral waveform content remains above a predetermined cutoff frequency.

[0004] In contemporary codecs, there are well-known methods for decoder-side signal recovery by audio signal bandwidth extension (BWE), e.g. spectral band replication (SBR) operating in the frequency domain or so-called time-domain bandwidth extension (TD-BWE) a post-processor in speech encoders operating in the time domain.

[0005] Additionally, there are several combined time / frequency domain coding concepts, such as those known under the terms AMR-WB+ or USAC.

[0006] All of these combined time-domain / coding concepts have the following in common: the frequency-domain encoder relies on bandwidth extension techniques that introduce band limitations into the input audio signal, and the portion above the crossover or boundary frequencies is encoded using low-resolution coding concepts and synthesized at the decoder. Therefore, these concepts rely primarily on pre-processor techniques on the encoder side and corresponding post-processing functions on the decoder side.

[0007] Typically, a time domain coder is selected for useful signals encoded in the time domain (e.g., speech signals), and a frequency domain coder is selected for non-speech signals, music signals, etc. However, in particular for non-speech signals having prominent harmonics in the high frequency band, prior art frequency domain coders have reduced accuracy and, therefore, reduced audio quality due to the fact that such prominent harmonics can only be separately coded parametrically or completely eliminated in the encoding / decoding process.

[0008] Furthermore, there are concepts in which the time-domain coding / decoding branch additionally relies on bandwidth extension, which also parametrically encodes the higher frequency range, while the lower frequency range is typically encoded using ACELP or any other CELP-related coder (e.g., a speech coder). This bandwidth extension functionality increases bit rate efficiency, but on the other hand introduces further inflexibility due to the fact that both coding branches, i.e., the frequency-domain coding branch and the time-domain coding branch, are band-limited due to the spectral band replication process or bandwidth extension process operating above a certain crossover frequency that is substantially lower than the maximum frequency included in the input audio signal.

[0009] Related topics in the prior art include

[0010] - SBR as a post-processor for waveform decoding [1-3]

[0011] - MPEG-D USAC core switching[4]

[0012] - MPEG-H 3D IGF [5]

[0013] The following papers and patents describe methods that are believed to constitute prior art for this application:

[0014] [1] M. Dietz, L. Liljeryd, K. Kjörling, and O. Kunz, “Spectral Band Replication, a novel approach in audio coding,” in 112th AES Convention, Munich, Germany, 2002.

[0015] [2] S. Meltzer, R. Böhm, and F. Henn, “SBR enhanced audio codecs for digital broadcasting such as “Digital Radio Mondiale” (DRM),” in 112th AES Convention, Munich, Germany, 2002.

[0016] [3] T. Ziegler, A. Ehret, P. Ekstrand, and M. Lutzky, “Enhancing mp3 with SBR: Features and Capabilities of the new mp3PRO Algorithm,” in 112th AES Convention, Munich, Germany, 2002.

[0017] [4] MPEG-D USAC standard.

[0018] [5] PCT / EP2014 / 065109.

[0019] MPEG-D USAC describes a switchable core encoder. However, in USAC, the band-limited core is constrained to always transmit a low-pass filtered signal. Consequently, certain music signals containing prominent high-frequency content, such as full-band sweeps and triangular sounds, cannot be faithfully reproduced. Summary of the Invention

[0020] It is an object of the present invention to provide an improved concept for audio coding.

[0021] This object is achieved by an audio encoding device encoder, an audio decoder, an audio encoding method, an audio decoding method or a machine-readable storage medium according to an embodiment of the present application.

[0022] The present invention is based on the discovery that a time-domain encoding / decoding processor can be combined with a frequency-domain encoding / decoding processor having a gap-filling function, but where this gap-filling function operates over the entire frequency band of the audio signal, or at least above a certain gap-filling frequency, for filling spectral holes. Importantly, the frequency-domain encoding / decoding processor is particularly capable of performing encoding / decoding with precise waveform or spectral values ​​up to a maximum frequency, not just up to a crossover frequency. Furthermore, the full-band capability of the frequency-domain encoder for encoding at high resolution allows the gap-filling function to be integrated into the frequency-domain encoder.

[0023] Therefore, according to the present invention, by using a full-band spectral encoder / decoder processor, the problems associated with separate bandwidth extension, on the one hand, and core coding, on the other hand, can be solved and overcome by performing bandwidth extension in the same spectral domain in which the core decoder operates. Thus, a full-rate core decoder is provided that encodes and decodes the full audio signal range. This eliminates the need for a downsampler on the encoder side and an upsampler on the decoder side. Instead, the entire processing is performed in the full sampling rate or full bandwidth domain. To achieve high coding gain, the audio signal is analyzed to find a first set of first spectral portions that must be encoded at high resolution, wherein this first set of first spectral portions may, in one embodiment, include the tonal portions of the audio signal. On the other hand, non-tonal or noise components in the audio signal, which constitute a second set of second spectral portions, are parametrically encoded at low spectral resolution. The encoded audio signal then requires only the first set of first spectral portions encoded in a waveform-preserving manner with high spectral resolution, and, in addition, a second set of second spectral portions parametrically encoded at low resolution using frequency "tiles" derived from the first set. On the decoder side, the core decoder, acting as a full-band decoder, reconstructs the first set of first spectral portions in a waveform-preserving manner, i.e., without any knowledge of the presence of any additional frequency regeneration. However, the resulting spectrum has many spectral gaps. These gaps are then filled using the inventive Intelligent Gap Filling (IGF) technique, using frequency regeneration using the application parameter data on the one hand and the source spectral range (i.e., the first spectral portion reconstructed by the full-rate audio decoder) on the other.

[0024] In a further embodiment, the spectrum portions reconstructed by only noise filling, rather than bandwidth replication or frequency patch filling, constitute a third group of third spectrum portions. Due to the fact that the coding concept operates in a single domain for core encoding / decoding on the one hand and frequency regeneration on the other hand, the IGF is not only limited to filling the upper frequency range, but can also fill the lower frequency range by noise filling without frequency regeneration or by frequency regeneration using frequency patches of different frequency ranges.

[0025] Furthermore, it is emphasized that information about spectral energy, information about individual energies or individual energy information, information about surviving energy or surviving energy information, information about patch energy or patch energy information, or information about missing energy or missing energy information may include not only energy values, but also (e.g., absolute) amplitude values, level values, or any other values ​​from which the final energy value can be derived. Thus, the information about energy may, for example, include the energy value itself, and / or a value of a level and / or an amplitude and / or an absolute amplitude.

[0026] Another aspect is based on the discovery that correlation is important not only for the source range but also for the target range. Furthermore, the present invention recognizes that different correlations may occur in the source and target ranges. For example, when considering a speech signal with high-frequency noise, it may be the case that a low-frequency band of the speech signal, including a small number of overtones, is highly correlated in the left and right channels when the loudspeaker is positioned centrally. However, due to the presence of high-frequency noise on the left side that differs from another high-frequency noise, or the absence of high-frequency noise on the right side, the high-frequency components may be strongly uncorrelated. Therefore, when performing a straightforward gap-filling operation that ignores this situation, the high-frequency components will also be correlated, potentially resulting in severe spatial segregation artifacts in the reconstructed signal. To address this issue, parameter data is calculated for the reconstructed frequency band, or more generally, for a second set of second spectral components that must be reconstructed using the first set of first spectral components, to identify first or second different binaural representations for the second spectral components, or in other words, first or second different binaural representations for the reconstructed frequency band. Therefore, on the encoder side, binaural identification is calculated for the second spectral components, i.e., for the components for which energy information for the reconstructed frequency band is also calculated. A frequency regenerator on the decoder side then regenerates the second spectral portion based on the first part of the first set of first spectral portions (i.e. the source range and parameter data for the second part, such as spectral envelope energy information or any other spectral envelope data) and additionally based on the binaural identification for the second part (i.e. for the reconstructed frequency band under reconsideration).

[0027] The binaural identification is preferably sent as a flag for each reconstruction band, and this data is sent from the encoder to the decoder, which then decodes the core signal as indicated by the preferably calculated flag for the core band. In implementation, the core signal is then stored in a stereo representation (e.g., left / right and mid / side), and for IGF frequency patch filling, the source patch representation is selected to fit the target patch representation as indicated by the binaural identification flag for the smart gap filling or reconstruction band (i.e., for the target range).

[0028] It is important to emphasize that this process works not only for stereo signals, i.e., for the left and right channels, but also for multichannel signals. In the case of multichannel signals, several different pairs of channels can be processed in this way, for example, left and right channels as a first pair, left surround and right surround as a second pair, and center channel and LFE channel as a third pair. Other pairings can be determined for higher output channel formats such as 7.1, 11.1, etc.

[0029] Another aspect is based on the discovery that IGF can improve the audio quality of the reconstructed signal because the entire spectrum is accessible to the core encoder, allowing, for example, perceptually important tonal portions in the high spectral range to be encoded by the core encoder rather than by parametric substitution. Furthermore, a gap-filling operation is performed using frequency tiles from a first set of first spectral portions, e.g., a set of tonal portions typically from the lower frequency range, but also from the higher frequency range (if available). However, for decoder-side spectral envelope adjustment, spectral portions from the first set of spectral portions located in the reconstructed band are not further post-processed, e.g., by spectral envelope adjustment. Only the remaining spectral values ​​in the reconstructed band that do not originate from the core decoder are envelope adjusted using envelope information. Preferably, the envelope information is full-band envelope information that accounts for the energy of the first set of first spectral portions in the reconstructed band and the energy of the second set of second spectral portions in the same reconstructed band, wherein the spectral values ​​in the second set of second spectral portions are indicated as zero and therefore not encoded by the core encoder, but are instead parametrically encoded using low-resolution energy information.

[0030] It has been found that the absolute energy values, normalized or not normalized with respect to the bandwidth of the corresponding frequency band, are useful and very efficient in decoder-side applications. This is particularly applicable when the gain factors have to be calculated based on the residual energy in the reconstructed frequency band, the missing energy in the reconstructed frequency band, and the frequency patch information in the reconstructed frequency band.

[0031] Furthermore, it is preferred that the encoded bitstream not only includes the energy information for the reconstruction bands, but also the scale factors for the scale factor bands extending up to the maximum frequency. This ensures that for each reconstruction band available for a certain tonal portion (i.e., first spectral portion), the first set of first spectral portions can actually be decoded with the correct amplitude. Furthermore, in addition to the scale factors for each reconstruction band, the energy for that reconstruction band is generated in the encoder and sent to the decoder. Furthermore, it is preferred that the reconstruction bands coincide with the scale factor bands, or, in the case of energy grouping, that at least the boundaries of the reconstruction bands coincide with the boundaries of the scale factor bands.

[0032] Another aspect is based on the discovery that certain impairments in audio quality can be remedied by applying a signal-adaptive frequency patch filling scheme. To this end, an encoder-side analysis is performed to identify the best matching source region candidate for a target region. Matching information identifying a source region for the target region, along with optional additional information, is generated and sent to the decoder as auxiliary information. The decoder then uses this matching information to apply a frequency patch filling operation. To do this, the decoder reads the matching information from the transmitted data stream or data file, accesses the source region identified for a reconstructed frequency band, and, if indicated in the matching information, performs additional processing on this source region data to generate raw spectral data for the reconstructed frequency band. The result of the frequency patch filling operation (i.e., the raw spectral data of the reconstructed frequency band) is then shaped using spectral envelope information to ultimately obtain a reconstructed frequency band that also includes first spectral components, such as tonal components. However, these tonal components are not generated by the adaptive patch filling scheme; instead, these first spectral components are directly output by the audio decoder or core decoder.

[0033] The adaptive spectral patch selection scheme can operate at a low granularity. In this implementation, the source region is subdivided into generally overlapping source regions, and the target region or reconstruction band is given by non-overlapping frequency target regions. Then, at the encoder side, the similarity between each source region and each target region is determined, and the best matching pair of source and target regions is identified through matching information. At the decoder side, the source region identified in the matching information is used to generate the original spectral data of the reconstructed band.

[0034] For the purpose of obtaining higher granularity, each source region is allowed to be shifted so as to obtain a certain lag where the similarity is maximum. This lag can be as fine as a frequency bin and allows for an even better match between the source and target regions.

[0035] Furthermore, in addition to identifying only the best matching pair, the correlation lag can also be sent within the match information, and furthermore, the sign can even be sent. If the sign is determined to be negative on the encoder side, then a corresponding sign flag is also sent within the match information, and on the decoder side, the source region spectrum value is multiplied by "-1," or "rotated" by 180 degrees in the complex representation.

[0036] Another implementation of the present invention uses a patch whitening operation. Spectral whitening removes coarse spectral envelope information and emphasizes the fine spectral structure that is of greatest interest for evaluating patch similarity. Thus, the frequency patches on the one hand and / or the source signal on the other hand are whitened before calculating the cross-correlation measure. When only a predefined process is used to whiten the patch, a whitening flag is sent to indicate that the decoder should apply the same predefined whitening process to the frequency patches within the IGF.

[0037] Regarding patch selection, a lag of the correlation is preferably used to spectrally shift the regenerated spectrum by an integer number of transform bins. Depending on the underlying transform, the spectral shift may require additional correction. In the case of an odd lag, the patch is additionally modulated by multiplying by an alternating time sequence of -1 / 1 to compensate for the frequency-inverted representation of every other band within the MDCT. In addition, the sign of the correlation result is applied when generating the frequency patches.

[0038] Furthermore, patch pruning and stabilization are preferably used to ensure that artifacts created by rapidly changing source regions for the same reconstructed or target region are avoided. To this end, a similarity analysis is performed between the different identified source regions, and when a source patch is similar to other source patches with a similarity above a threshold, the source patch can be discarded from the set of potential source patches because it is highly correlated with the other source patches. Furthermore, as a form of patch selection stabilization, if none of the source patches in the current frame are correlated with the target patch in the current frame (better than a given threshold), the patch order from the previous frame is preferably maintained.

[0039] Another aspect is based on the discovery that improved quality and reduced bitrate, particularly for signals containing transients (as they often occur in audio signals), can be achieved by combining temporal noise shaping (TNS) or temporal patch shaping (TTS) techniques with high-frequency reconstruction. TNS / TTS processing at the encoder, implemented via frequency-dependent prediction, reconstructs the temporal envelope of the audio signal. According to the implementation, when the temporal noise shaping filter is determined within a frequency range that covers not only the source frequency range but also the target frequency range to be reconstructed in the frequency reconstruction decoder, the temporal envelope is applied not only to the core audio signal up to a gap-filling start frequency but also to the spectral range of the reconstructed secondary spectral portion. Consequently, pre-echoes and post-echoes that would occur without temporal patch shaping are reduced or eliminated. This is achieved by applying frequency-dependent inverse prediction not only within a core frequency range up to a certain gap-filling start frequency but also within a frequency range above the core frequency range. To this end, frequency reconstruction or frequency patch generation is performed at the decoder before applying the frequency-dependent prediction. However, the frequency-dependent prediction can be applied before or after the spectrum envelope shaping, depending on whether the energy information calculation has been performed on the spectrum residual values ​​after filtering or on the (entire) spectrum values ​​before envelope shaping.

[0040] The TTS processing with respect to one or more frequency patches further establishes the continuity of the correlation between the source range and the reconstruction range or in two adjacent reconstruction ranges or frequency patches.

[0041] In implementation, complex TNS / TTS filtering is preferably used. This avoids (temporal) aliasing artifacts of critically sampled real-number representations (such as the MDCT). In addition to obtaining a complex modified transform, the complex TNS filter can be calculated on the encoder side by applying not only a modified discrete cosine transform but also a modified discrete sine transform. However, only the modified discrete cosine transform values, i.e., the real part of the complex transform, are transmitted on the decoder side. However, it is possible to use the MDCT spectrum of the previous or subsequent frame to estimate the imaginary part of the transform, so that on the decoder side, the complex filter can again be applied for inverse prediction with respect to frequency, and in particular, with respect to the boundaries between the source range and the reconstruction range, as well as with respect to the boundaries between frequency-adjacent frequency blocks within the reconstruction range.

[0042] The audio coding system of the present invention efficiently encodes arbitrary audio signals over a wide range of bit rates. However, for high bit rates, the system converges to transparency, and for low bit rates, perceptual annoyances are minimized. Consequently, the majority of the available bit rate is used to waveform-code only the most perceptually relevant structure of the signal at the encoder, and the resulting spectral gaps are filled at the decoder with a signal content that roughly approximates the original spectrum. Parameter-driven, so-called spectrally intelligent gap filling (IGF) is controlled using dedicated side information sent from the encoder to the decoder, consuming the very limited bit budget.

[0043] In alternative embodiments, the time-domain encoding / decoding processor relies on a lower sampling rate and corresponding bandwidth extension functionality.

[0044] In another embodiment, a crossover processor is provided to initialize the time domain encoder / decoder using initialization data derived from the currently processed frequency domain encoder / decoder signal. This allows the parallel time domain encoder to be initialized while the currently processed audio signal portion is being processed by the frequency domain encoder, so that when a switch occurs from the frequency domain encoder to the time domain encoder, the time domain encoder can immediately begin processing because all initialization data associated with the earlier signal is already present due to the crossover processor. The crossover processor is preferably applied on the encoder side and also on the decoder side, and preferably uses a frequency-time transform that, among other things, performs very efficient downsampling from a higher output or input sampling rate to a lower time domain core encoder sampling rate by selecting only a certain low-frequency band portion of the domain signal and a certain reduced transform size. Thus, the sample rate conversion from a high sampling rate to a low sampling rate is performed very efficiently, and the signal obtained by the transform with the reduced transform size can then be used to initialize the time domain encoder / decoder, so that the time domain encoder / decoder is ready to perform time domain encoding immediately when this condition is signaled by the controller and the immediately preceding audio signal portion is encoded in the frequency domain.

[0045] Therefore, preferred embodiments of the present invention allow for seamless switching of a perceptual audio coder including spectral gap filling and a time domain coder with or without bandwidth extension.

[0046] Therefore, the present invention relies on a method that is not limited to removing high-frequency content above a cutoff frequency from the audio signal in a frequency domain encoder, but rather signal-adaptively removes spectral bandpass regions leaving spectral gaps in the encoder and subsequently reconstructs these spectral gaps in the decoder. Preferably, an integrated solution such as smart gap filling is used, which effectively combines full-bandwidth audio coding and spectral gap filling, particularly in the MDCT transform domain.

[0047] Therefore, the present invention provides an improved concept for combining speech coding and subsequent time-domain bandwidth extension with full-band waveform decoding including spectral gap filling into a switchable perceptual encoder / decoder.

[0048] Thus, in contrast to already existing approaches, the new concept exploits full-band audio signal waveform coding in a transform domain coder and at the same time allows for a seamless switch to a speech coder, preferably followed by a time domain bandwidth extension.

[0049] Other embodiments of the present invention avoid the interpretation problems that arise due to fixed frequency band limitations. This concept implements a switchable combination of a full-band waveform encoder in the frequency domain equipped with spectral gap filling and a lower sampling rate speech encoder and time domain bandwidth extension. This encoder is capable of waveform encoding the problematic signal described above, thereby providing the full audio bandwidth up to the Nyquist frequency of the audio input signal. Nevertheless, seamless and instantaneous switching between the two coding strategies is ensured, in particular, by embodiments having a crossover processor. For this seamless switching, the crossover processor represents a cross-connection at both the encoder and decoder between a full-band capable full-rate (input sampling rate) frequency domain encoder and a low-rate ACELP encoder with a lower sampling rate, so as to appropriately initialize ACELP parameters and buffers, particularly within the adaptive codebook, LPC filter, or resampling stage, when switching from a frequency domain encoder such as TCX to a time domain encoder such as ACELP. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The invention is then discussed with respect to the accompanying drawings, in which:

[0051] Figure 1a An apparatus for encoding an audio signal is shown;

[0052] Figure 1b Shown with Figure 1a a decoder matched to the encoder for decoding the encoded audio signal;

[0053] Figure 2a A preferred implementation of the encoder is shown;

[0054] Figure 2b A preferred implementation of the encoder is shown;

[0055] Figure 3a Shown by Figure 1b Schematic representation of the spectrum produced by the frequency domain decoder;

[0056] Figure 3b showing a table indicating the relationship between the scale factors for the scale factor bands and the energy for the reconstruction bands and the noise filling information for the noise filling bands;

[0057] Figure 4a The functionality of a spectral domain encoder for applying the selection of spectral portions to the first and second groups of spectral portions is shown;

[0058] Figure 4b Shown Figure 4a Implementation of the function;

[0059] Figure 5a shows the functionality of the MDCT encoder;

[0060] Figure 5b shows the functionality of a decoder with MDCT technology;

[0061] Figure 5c shows the implementation of a frequency regenerator;

[0062] Figure 6 shows the implementation of an audio encoder;

[0063] Figure 7a shows a cross-processor within an audio encoder;

[0064] Figure 7b An implementation is shown that additionally provides an inverse sampling rate reduction or frequency-to-time transform within the crossover processor;

[0065] Figure 8 Shown Figure 6 The preferred implementation of the controller;

[0066] Figure 9 Another embodiment of a time domain encoder with bandwidth extension function is shown;

[0067] Figure 10 shows the preferred use of a pre-processor;

[0068] Figure 11a shows a schematic implementation of an audio decoder;

[0069] Figure 11bA cross-processor within a decoder for providing initialization data for a time domain decoder is shown;

[0070] Figure 12 A preferred implementation of a time domain decoding processor is shown Figure 11a

[0071] Figure 13 A further implementation of a time domain bandwidth extension is shown

[0072] Fig. 14a shows a preferred implementation of an audio encoder;

[0073] Figure 14b A preferred implementation of an audio decoder is shown

[0074] Figure 14c An inventive implementation of a time domain decoder with sample rate conversion and bandwidth extension is shown. DETAILED DESCRIPTION

[0075] Figure 6 An audio encoder for encoding an audio signal is shown, comprising a first encoding processor 600 for encoding a first audio signal portion in the frequency domain. The first encoding processor 600 comprises a time-to-frequency converter 602 for converting the first input audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the input signal. Further, the first encoding processor 600 comprises an analyzer 604 for analyzing the frequency domain representation up to the maximum frequency to determine a first spectral region to be encoded with a first spectral representation and to determine a second spectral region to be encoded with a second spectral resolution, the second spectral resolution being lower than the first spectral resolution. In particular, the full-band analyzer 604 determines which spectral lines or spectral values in the time-to-frequency converter spectrum are to be encoded in a spectral line fashion and which other spectral portions are to be encoded in a parametric fashion, the spectral values of these latter ones then being reconstructed at the decoder side with a gap filling process. The actual encoding operation is performed by a spectral encoder 606 for encoding the first spectral region or spectral portion in the first resolution and for encoding the second spectral region or portion in a parametric fashion with the second spectral resolution.

[0076] Figure 6 ​The audio encoder further comprises a second encoding processor 610 for encoding the audio signal portion in the time domain. In addition, the audio encoder comprises a controller 620 configured for analyzing the audio signal at the audio signal input 601 and for determining which portion of the audio signal is the first audio signal portion that is encoded in the frequency domain and which portion of the audio signal is the second audio signal portion that is encoded in the time domain. Furthermore, an encoded signal former 630, which can for example be implemented as a bitstream multiplexer, is provided, configured for forming an encoded audio signal comprising a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion. It is important that the encoded signal only has either the frequency domain representation or the time domain representation from the same audio signal portion.

[0077] The controller 620 thus ensures that for a single audio signal portion only the time domain representation or the frequency domain representation is in the encoded signal. This can be implemented by the controller 620 in several ways. One way would be that for the same audio signal portion both representations reach the block 630 and the controller 620 controls the encoded signal former 630 to only introduce one of the two representations into the encoded signal. However, alternatively, the controller 620 can control the input into the first encoding processor and the input into the second encoding processor such that based on the analysis of the respective signal portion only one of the blocks 600 or 610 is activated to actually perform the full encoding operation and the other block is deactivated.

[0078] This deactivation can be a deactivation, alternatively, for example, with respect to Figure 7a is merely an "initialization" mode in which the other encoding processor is only active for receiving and processing initialization data in order to initialize the internal memory, but does not perform any specific encoding operation at all. This activation can be done by some switch at the input not shown in Figure 6 or, preferably, by the control lines 621 and 622. Thus, in this embodiment, when the controller 620 has determined that the current audio signal portion should be encoded by the first encoding processor, the second encoding processor 610 does not output anything while the second encoding processor is still provided with initialization data to be active for a future instant switch. On the other hand, the first encoding processor is configured not to need any data from the past to update any internal memory and thus, when the current audio signal portion is to be encoded by the second encoding processor 610, then the controller 620 can control the first ending encoding processor 600 to be completely inactive via the control line 621. This means that the first encoding processor 600 does not need to be in an initialization state or a waiting state, but can be in a complete deactivation state. This is particularly preferred for mobile devices where power consumption and thus battery life is an issue.

[0079] In a further specific implementation of the second encoding processor operating in the time domain, the second encoding processor comprises a downsampler 900 or sample rate converter for converting the audio signal portion into a representation having a lower sampling rate, wherein the lower sampling rate is lower than the sampling rate at the input to the first encoding processor. Figure 9 900. In particular, when the input audio signal comprises low frequency band and high frequency band, preferably, the low frequency band that only has the input audio signal part is represented by the lower sampling rate at the output of piece 900, and this low frequency band is encoded by time domain low frequency band encoder 910 then, and time domain low frequency band encoder 910 is configured to represent that the lower sampling rate provided by piece 900 is carried out time domain coding. In addition, time domain bandwidth extension encoder 920 is provided, is used for high frequency band is encoded in parameter mode. For this reason, time domain bandwidth extension encoder 920 receives the high frequency band of input audio signal or low frequency band and the high frequency band of input audio signal at least.

[0080] In another embodiment of the present invention, the audio encoder further comprises (although in Figure 6 Not shown, but in Figure 10 14a ), is configured to preprocess a first audio signal portion and a second audio signal portion. In one embodiment, the preprocessor includes a prediction analyzer for determining prediction coefficients. The prediction analyzer can be implemented as an LPC (Linear Predictive Coding) analyzer for determining LPC coefficients. However, other analyzers are also possible. Furthermore, the preprocessor (also shown in FIG. 14a ) includes a prediction coefficient quantizer 1010 , wherein the device shown in FIG. 14a receives prediction coefficient data from the prediction analyzer, also shown at 1002a, 1002b in FIG. 14a .

[0081] Furthermore, the preprocessor further includes an entropy encoder for generating encoded versions of the quantized prediction coefficients. It is important to note that the coded signal former 630 or, in a specific implementation, the bitstream multiplexer 613 ensures that the encoded versions of the quantized prediction coefficients are included in the encoded audio signal 632. Preferably, the LPC coefficients are not directly quantized, but rather converted to, for example, an ISF, or any other representation more suitable for quantization. This conversion is preferably performed by determining the LPC coefficient blocks 1002a, 1002b or within the block 1010 for quantizing the LPC coefficients.

[0082] Furthermore, the pre-processor may comprise a resampler 1004 or 1021 in FIG14a for resampling the audio input signal at the input sampling rate to a lower sampling rate for the time domain encoder. When the time domain encoder is an ACELP encoder having a certain ACELP sampling rate, then the downsampling is performed preferably to 12.8 kHz or 16 kHz. The input sampling rate may be any one of a certain number of sampling rates, such as 32 kHz or even higher sampling rates. On the other hand, the sampling rate of the time domain encoder will be predetermined by certain constraints, and the resampler 1004 performs this resampling and outputs a lower sampling rate representation of the input signal. Thus, the resampler may perform a similar function and may even be a Figure 9 The same element as the downsampler 900 shown in the context of .

[0083] Furthermore, pre-emphasis is preferably applied in the pre-emphasis blocks 1005a, 1005b in Figure 14a. Pre-emphasis processing is well known in the art of time domain coding and is described in the literature with reference to AMR-WB+ processing, and pre-emphasis is particularly configured to compensate for spectral tilt and thus allow for better calculation of LPC parameters for a given LPC order.

[0084] In addition, the preprocessor may further include a Figure 14b The TCX-LTP parameter extraction of the LTP postfilter is shown at 1420 in FIG. This block is shown at 1024 in FIG. 14a. Furthermore, the preprocessor may additionally include other functions shown at 1007, and these other functions may include a pitch search function, a voice activity detection (VAD) function, or any other function known in the art of time domain or speech coding.

[0085] As shown, the result of block 1024 is input into the encoded signal, i.e., in the embodiment of Figure 14a, into the bitstream multiplexer 630. In addition, the data from block 1007 may also be introduced into the bitstream multiplexer if desired, or may alternatively be used for the purpose of time-domain encoding in a time-domain encoder.

[0086] Thus, in summary, both paths share a pre-processing operation 1000, where common signal processing operations are performed. These include resampling to the ACELP sampling rate (12.8 or 16 kHz) for one parallel path, which is always performed. Furthermore, TCX LTP parameter extraction, shown at block 1024, is performed, along with pre-emphasis and determination of the LPC coefficients (1002a, 1002b). As outlined, pre-emphasis compensates for spectral tilt, thus making the calculation of the LPC parameters for a given LPC order more efficient.

[0087] Then, refer to Figure 8 , in order to illustrate a preferred implementation of the controller 620. The controller receives at input the portion of the audio signal under consideration. Preferably, as shown in FIG14 a, the controller receives any signal available in the pre-processor 1000, which can be the original input signal at the input sampling rate or a resampled version at a lower time domain encoder sampling rate, or the signal obtained after the pre-emphasis processing in block 1005.

[0088] Based on the audio signal portion, the controller 620 addresses a frequency domain encoder simulator 621 and a time domain encoder simulator 622 so as to calculate the estimated signal to noise ratio for each encoder possibility. Subsequently, the selector 623 naturally selects the encoder that has provided a better signal to noise ratio under the consideration of the predefined bit rate. The selector then identifies the corresponding encoder by controlling the output. When it is determined that the audio signal portion under consideration is to be encoded using the frequency domain encoder, the time domain encoder is set to an initialized state, or in other embodiments, does not require very instantaneous switching in a completely deactivated state. However, when it is determined that the audio signal portion under consideration is to be encoded by the time domain encoder, the frequency domain encoder is deactivated.

[0089] Then, it is shown Figure 8 The preferred implementation of the controller shown in [ 1 ] is shown. The decision on whether to select the ACELP or TCX path in the switching decision is made by simulating ACELP and TCX encoders and switching to the better performing branch. To this end, the SNRs of the ACELP and TCX branches are estimated based on ACELP and TCX encoder / decoder simulations. The TCX encoder / decoder simulations are performed without TNS / TTS analysis, IGF encoders, quantization loops / arithmetic encoders, or any TCX decoders. Instead, the TCX SNR is estimated using an estimate of the quantizer distortion in the shaped MDCT domain. The ACELP encoder / decoder simulations are performed using only adaptive and innovative codebook simulations. The ACELP SNR is simply estimated by calculating the distortion introduced by the LTP filter in the weighted signal domain (adaptive codebook) and scaling it by a constant factor (innovative codebook). This significantly reduces complexity compared to methods that perform TCX and ACELP coding in parallel. The branch with the higher SNR is selected for the subsequent full coding run.

[0090] When the TCX branch is selected, the TCX decoder is run in each frame, outputting a signal at the ACELP sampling rate. This is used to update the memory used for the ACELP coding path (LPC residual, Mem w0, memory de-emphasis) to enable instantaneous switching from TCX to ACELP. Memory updates are performed in each TCX path.

[0091] Alternatively, a complete analysis by synthesis process can be performed, that is, both encoder simulators 621 and 622 implement the actual encoding operation, and the results are compared by selector 623. Alternatively, again, the complete feedforward calculation can be completed by performing signal analysis. For example, when the signal is determined to be a speech signal by a signal classifier, a time domain encoder is selected, and when the signal is determined to be a music signal, a frequency domain encoder is selected. Other processes can also be applied to distinguish between the two encoders based on signal analysis of the audio signal portion under consideration.

[0092] Preferably, the audio encoder further comprises Figure 7a . When the frequency domain encoder 600 is active, the cross processor 700 provides initialization data to the time domain encoder 610 so that the time domain encoder is ready for seamless switching in future signal portions. In other words, when the current signal portion is determined to be encoded using the frequency domain encoder, and when the controller determines that the immediately following audio signal portion is to be encoded by the time domain encoder 610, then without the cross processor, such immediate seamless switching would not be possible. However, for the purpose of initializing the memory in the time domain encoder, the cross processor provides the time domain encoder 610 with a signal derived from the frequency domain encoder 600 because the time domain encoder 610 has a dependency on the signal encoded from the current frame or the immediately preceding frame in time from the input.

[0093] Therefore, the time domain encoder 610 is configured to be initialized by the initialization data in order to encode audio signal portions subsequent to earlier audio signal portions encoded by the frequency domain encoder 600 in an efficient manner.

[0094] In particular, the crossover processor includes a time converter for converting the frequency domain representation into a time domain representation, which can be forwarded to the time domain encoder directly or after some further processing. This converter is shown in Figure 14a as an IMDCT (Inverse Modified Discrete Cosine Transform) block. However, compared to the time-frequency converter block 602 shown in Figure 14a, this block 702 has a different transform size (Modified Discrete Cosine Transform block). As shown in block 602, the time-frequency converter 602 operates at the input sampling rate, and the inverse modified discrete cosine transform 702 operates at the lower ACELP sampling rate.

[0095] The ratio of the time domain encoder sampling rate or ACELP sampling rate to the frequency domain encoder sampling rate or input sampling rate can be calculated and is Figure 7b The downsampling factor DS is shown. Block 602 has a large transform size, and IMDCT block 702 has a small transform size. Figure 7bAs shown, the IMDCT block 702 therefore includes a selector 726 for selecting the lower spectrum portion of the input to the IMDCT block 702. The portion of the full-band spectrum is defined by the downsampling factor DS. For example, when the lower sampling rate is 16kHz and the input sampling rate is 32kHz, the downsampling factor is 0.5, so the selector 726 selects the lower half of the full-band spectrum. When the spectrum has, for example, 1024 MDCT lines, the selector selects the lower 512 MDCT lines.

[0096] This low frequency portion of the full-band spectrum is input to a small-scale transform and foldout block 720, as Figure 7b As shown. The transform size is also selected based on the downsampling factor and is 50% of the transform size in block 602. Synthesis windowing is then performed, where the window has a small number of coefficients. The number of coefficients of the synthesis window is equal to the downsampling factor multiplied by the number of coefficients of the analysis window used in block 602. Finally, an overlap-add operation is performed with an even smaller number of operations per block, and the number of operations per block is again the number of operations per block in the full-rate implementation of the MDCT multiplied by the downsampling factor.

[0097] Therefore, a very efficient downsampling operation can be applied since downsampling is included in the IMDCT implementation. In this context, it is emphasized that block 702 can be implemented by IMDCT, but can also be implemented by any other transform or filter bank implementation that can be appropriately sized in the actual transform kernel and other transform related operations.

[0098] In another embodiment shown in FIG. 14 a , the time-to-frequency converter comprises additional functionality in addition to the analyzer. Figure 6 The analyzer 604 may include a temporal noise shaping / temporal patch shaping analysis block 604a in the embodiment of FIG. 14a, which is as described in relation to the TNS / TTS analysis block 604a. Figure 2b 14a and 14b. Figure 2b Proceed as shown.

[0099] Furthermore, the frequency domain encoder preferably includes a noise shaping block 606a. Noise shaping block 606a is controlled by the quantized LPC coefficients generated as in block 1010. The quantized LPC coefficients used for noise shaping 606a perform spectral shaping of the high-resolution spectral values ​​or lines that are directly coded (rather than parametrically coded), and the result of block 606a resembles the spectrum of the signal after an LPC filtering stage that operates in the time domain (e.g., the LPC analysis filter block 704, described later). Furthermore, the result of noise shaping block 606a is then quantized and entropy coded, as indicated by block 606b. The result of block 606b corresponds to the encoded first audio signal portion, or the frequency-domain encoded audio signal portion (along with other side information).

[0100] The crossover processor 700 includes a spectral decoder for computing a decoded version of the first coded signal portion. In the embodiment of FIG. 14 a, the spectral decoder 701 includes the previously discussed inverse noise shaping block 703, gap filling decoder 704, TNS / TTS synthesis block 705, and IMDCT block 702. These blocks undo certain operations performed by blocks 602 through 606 b. Specifically, the noise shaping block 703 undoes the noise shaping performed by block 606 a based on the quantized LPC coefficients 1010. The IGF decoder 704 is described with respect to FIG. Figure 2a Blocks 202 and 206 operate as discussed above, and the TNS / TTS synthesis block 705 operates as in Figure 2a 14 a and the spectral decoder additionally comprises an IMDCT block 702. Furthermore, the crossover processor 700 in FIG14 a additionally or alternatively comprises a delay stage 707 for feeding a delayed version of the decoded version obtained by the spectral decoder 701 into the de-emphasis stage 617 of the second encoding processor for the purpose of initializing the de-emphasis stage 617.

[0101] Furthermore, the crossover processor 700 may additionally or alternatively include a weighted prediction coefficient analysis filter stage 708 for filtering the decoded version and for feeding the filtered decoded version to a codebook determiner 613, indicated as "MMSE" in FIG14a, of the second encoding processor for initializing the block. Additionally or alternatively, the crossover processor includes an LPC analysis filter stage for filtering the decoded version of the first coded signal portion output by the spectrum decoder 700 to an adaptive codebook stage 712 for initializing the block 612. Additionally or alternatively, the crossover processor further includes a pre-emphasis stage 709 for performing a pre-emphasis process on the decoded version output by the spectrum decoder 701 prior to the LPC filtering 706. The pre-emphasis stage output may also be fed to a further delay stage 710 for the purpose of initializing an LPC synthesis filter block 616 within the time domain encoder 610, for the purpose of initializing the LPC analysis filter block 611.

[0102] As shown in FIG14 a, the time domain encoder processor 610 includes a pre-emphasis operation at the lower ACELP sampling rate. As shown, this pre-emphasis is the pre-emphasis performed in the pre-processing stage 1000 and has the reference numeral 1005. The pre-emphasized data is input to the LPC analysis filter stage 611, which operates in the time domain and is controlled by the quantized LPC coefficients 1010 obtained by the pre-processing stage 1000. As is known from AMR-WB+ or USAC or other CELP coders, the residual signal generated by block 611 is provided to an adaptive codebook 612. In addition, the adaptive codebook 612 is connected to the innovation codebook stage 614, and the codebook data from the adaptive codebook 612 and the innovation codebook are input to the bitstream multiplexer, as shown.

[0103] Furthermore, an ACELP gain / encoding stage 615 is provided in series with the innovative codebook stage 614, and the result of this block is input to the codebook determiner 613, indicated as MMSE in FIG14 a. This block cooperates with the innovative codebook block 614. Furthermore, the time-domain encoder further comprises a decoder portion having an LPC synthesis filter block 616, a de-emphasis block 617, and an adaptive bass post-filtering stage 618 for calculating the parameters of the adaptive bass post-filtering, which is, however, applied on the decoder side. In the absence of any adaptive bass post-filtering on the decoder side, blocks 616, 617, and 618 would not be necessary for the time-domain encoder 610.

[0104] As shown, several blocks of the time domain decoder depend on the previous signal, and these blocks are the adaptive codebook block, the codebook determiner 613, the LPC synthesis filter block 616 and the de-emphasis block 617. These blocks are provided with data from the cross processor derived from the frequency domain encoding processor data in order to be ready for the instantaneous switch from the frequency domain encoder to the time domain encoder (e.g. Figure 14a-2 1450 in the figure). As can also be seen in FIG14a, for the frequency domain encoder, any dependency on earlier data is not necessary. Therefore, the cross processor 700 does not provide any memory initialization data from the time domain encoder to the frequency domain encoder. However, for other implementations of the frequency domain encoder where dependencies from the past exist and where memory initialization data is required, the cross processor 700 is configured to operate in both directions.

[0105] Therefore, a preferred embodiment of the audio encoder comprises the following parts:

[0106] The preferred audio decoder is described in the following: The waveform decoder part consists of a full-band TCX decoder path and an IGF, both operating at the codec's input sampling rate. In parallel, there is an alternative ACELP decoder path at a lower sampling rate, which is further enhanced downstream by a TD-BWE.

[0107] For ACELP initialization upon switching from TCX to ACELP, there is a cross-path performing the inventive ACELP initialization (consisting of a shared TCX decoder front-end, but otherwise providing an output at a lower sampling rate and some post-processing). Sharing the same sampling rate and filtering order in the LPC between TCX and ACELP allows for easier and more efficient ACELP initialization.

[0108] For visualization of the switching, two switches are drawn in 14b. When the second switch downstream selects between the TCX / IGF or ACELP / TD-BWE output, the first switch either pre- updates the buffers in the resampling QMF stage downstream of the ACELP path by the output of the cross-path or simply passes the ACELP output.

[0109] Subsequently, in Figures 11a-14c the context of the audio decoder implementation according to aspects of the present application is discussed.

[0110] The audio decoder for decoding the encoded audio signal 1101 comprises a first decoding processor 1120 for decoding the first encoded audio signal portion in the frequency domain. The first decoding processor 1120 comprises a spectral decoder 1122 for decoding the first spectral region at a high spectral resolution and for synthesizing the second spectral region using a parametric representation of the second spectral region and the decoded first spectral region to obtain a decoded spectral representation. The decoded spectral representation is as discussed in the context of Figure 6 and also as discussed in the context of the full-band decoded spectral representation of Figure 1a . Thus, in general, the first decoding processor comprises a full-band implementation with a gap filling procedure in the frequency domain. The first decoding processor 1120 further comprises a frequency time converter 1124 for converting the decoded spectral representation into the time domain to obtain a decoded first audio signal portion.

[0111] Further, the audio decoder comprises a second decoding processor 1140 for decoding the second encoded audio signal portion in the time domain to obtain a decoded second signal portion. Further, the audio decoder comprises a combiner 1160 for combining the decoded first signal portion and the decoded second signal portion to obtain a decoded audio signal. The decoded signal portions are combined in order, which is also discussed in Figure 14b . Figure 11aThe switch implementation 1160 of the embodiment of the combiner 1160 is shown.

[0112] Preferably, the second decoding processor 1140 is a time domain bandwidth extension processor and comprises, as shown in Figure 12 the time domain low band decoder 1200 for decoding the low band time domain signal. The implementation further comprises an up-sampler 1210 for up-sampling the low band time domain signal. In addition, a time domain bandwidth extension decoder 1220 is provided for synthesizing the high band of the output audio signal. Furthermore, a mixer 1230 is provided for mixing the synthesized high band of the time domain output signal and the up-sampled low band time domain signal to obtain the time domain encoder output. Thus, in the preferred embodiment, Figure 11a the block 1140 in Figure 12 may be implemented by the functionality of

[0113] Figure 13 A preferred embodiment of the time domain bandwidth extension decoder 1220 of Figure 12 is shown. Preferably, a time domain up-sampler 1221 is provided which receives as input the LPC residual signal from the time domain low band decoder shown at 1200 within the block 1140 and in the context of Figure 12 as shown in Figure 14b The time domain up-sampler 1221 produces an up-sampled version of the LPC residual signal. This version is then input into a non-linear distortion block 1222 which produces an output signal with higher frequency values based on its input signal. The non-linear distortion can be a copy, a mirror, a frequency shift or a non-linear device such as a diode or transistor operating in a non-linear region. The output signal of the block 1222 is input into an LPC synthesis filtering block 1223 which is also controlled by the LPC data for the low band decoder or, for example, by specific envelope data produced by the encoder side time domain bandwidth extension block 920 of Fig. 14a. The output of the LPC synthesis block is then input into a band pass or high pass filter 1224 to finally obtain the high band which is then input into the mixer 1230 as shown in Figure 12 .

[0114] Subsequently, Figure 12 a preferred implementation of the up-sampler 1210 is discussed in the context of Fig. 14a. The up-sampler preferably comprises an analysis filter bank operating at the first time domain low band decoder sampling rate. A specific implementation of such an analysis filter bank is Figure 14b. In addition, the upsampler includes a synthesis filter bank 1473 operating at a second output sampling rate higher than the first time domain low frequency band sampling rate. Therefore, the QMF synthesis filter bank 1473, which is a preferred implementation of the general filter bank, operates at the output sampling rate. Figure 7b 5, then the QMF analysis filterbank 1471 has, for example, only 32 filterbank channels and the QMF synthesis filterbank 1473 has, for example, 64 QMF channels, but the upper half of the filterbank channels, i.e. the upper 32 filterbank channels, are fed with zeros or noise, whereas the lower 32 filterbank channels are fed with the corresponding signals provided by the QMF analysis filterbank 1471. However, preferably, the bandpass filtering 1472 is performed in the QMF filterbank domain in order to ensure that the QMF synthesis output 1473 is an upsampled version of the ACELP decoder output, but without any artifacts above the maximum frequency of the ACELP decoder.

[0115] Further processing operations may be performed in the QMF domain in addition to or as an alternative to bandpass filtering 1472. If no processing is performed at all, the QMF analysis and QMF synthesis constitute an efficient upsampler 1210.

[0116] Afterwards, Figure 14b The structure of each element is discussed in more detail.

[0117] Full-band frequency domain decoder 1120 comprises a first decoding block 1122a for decoding high-resolution spectral coefficients and for performing, for example, noise filling in the low-frequency band portion known from the USAC technique in addition. In addition, the full-band decoder comprises an IGF processor 1122b for filling spectrum holes with synthetic spectral values ​​that have been encoded only in a parameter manner and therefore at the encoder side at low resolution. Then, in block 1122c, inverse noise shaping is performed, and the result is input to the TNS / TTS synthesis block 705, which is provided as the input of the final output to a frequency-time converter 1124, which is preferably implemented as an inverse modified discrete cosine transform operating at the output, i.e., a high sampling rate.

[0118] In addition, using Figure 14b The result is then the first audio signal portion decoded at the output sampling rate and as obtained from the TCX LTP parameter extraction block 1024. Figure 14b As can be seen, the data has a high sampling rate and therefore does not require any further frequency enhancement at all due to the fact that the decoding processor is a frequency domain full band decoder which is preferably used in Figures 1a-5cThe smart gap filling technique is discussed in the context of FIG.

[0119] Figure 14b Several elements in FIG14a are very similar to corresponding blocks in the crossover processor 700 of FIG14a, in particular with respect to the IGF decoder 704 corresponding to the IGF processing 1122b, and the inverse noise shaping operation controlled by the quantized LPC coefficients 1145 corresponding to the inverse noise shaping 703 of FIG14a, and Figure 14b The TNS / TTS synthesis block 705 in FIG. 14a corresponds to the block TNS / TTS synthesis 705 in FIG. 14a . However, it is important to note that Figure 14b The IMDCT block 1124 in FIG. 14 a operates at a high sampling rate, while the IMDCT block 702 in FIG. 14 a operates at a low sampling rate. Figure 14b The block 1124 in FIG. 1 includes a large sized transform and expansion block 710, a synthesis window in block 712, and an overlap-add stage 714 having a correspondingly large number of operations, a large number of window coefficients, and a large transform size compared to the corresponding features 720, 722, 724, which is operated in block 702 and will be described later. Figure 14b The crossbar processor 1170 is outlined in block 1171 .

[0120] The time domain decoding processor 1140 preferably comprises an ACELP or time domain low band decoder 1200 including an ACELP decoder stage 1149 for obtaining decoding gain and innovative codebook information. In addition, an ACELP adaptive codebook stage 1141 is provided, followed by an ACELP post-processing stage 1142 and a final synthesis filter (e.g. an LPC synthesis filter 1143), which again consists of a code corresponding to Figure 11a The quantized LPC coefficients 1145 obtained by the bitstream demultiplexer 1100 of the encoded signal parser 1100 are controlled by the quantized LPC coefficients 1145. The output of the LPC synthesis filter 1143 is input to the de-emphasis stage 1144 for removing or undoing the processing introduced by the pre-emphasis stage 1005 of the pre-processor 1000 of FIG. 14 a. The result is a time domain output signal at a low sampling rate and a low frequency band, and in the case where a frequency domain output is required, the switch 1480 is in the indicated position and the output of the de-emphasis stage 1144 is introduced into the upsampler 1210 and then mixed with the high frequency band from the time domain bandwidth extension decoder 1220.

[0121] According to an embodiment of the present invention, the audio decoder further comprises Figure 11b and Figure 14bThe crossover processor 1170 shown in is used to calculate initialization data for the second decoding processor based on the decoded spectral representation of the first encoded audio signal portion, so that the second decoding processor is initialized to decode the encoded second audio signal portion of the encoded audio signal that temporally follows the first audio signal portion, that is, so that the time domain coding processor 1140 is ready for instantaneous switching from one audio signal portion to the next audio signal portion without any loss in quality or efficiency.

[0122] Preferably, the crossover processor 1170 comprises an additional frequency-to-time converter 1171 operating at a lower sampling rate than the frequency-to-time converter of the first decoding processor in order to obtain a further decoded first signal portion in the time domain to be used as an initialization signal or for which any initialization data can be derived. Preferably, the IMDCT or the low sampling rate frequency-to-time converter is implemented as Figure 7b , item 726 (selector), item 720 (small size transform and expansion), a synthesis window with a smaller number of window coefficients as shown at 722, and an overlap-add stage with a smaller number of operations as shown at 724. Thus, the IMDCT block 1124 in the frequency domain full-band decoder is implemented as shown by blocks 710, 712, 714, and the IMDCT block 1171 is as shown. Figure 7b This is shown as being implemented by blocks 726, 720, 722, 724. Again, the downsampling factor is the ratio between the time domain encoder sampling rate or low sampling rate and the higher frequency domain sampling rate or output sampling rate, and is less than 1 and can be any number greater than 0 and less than 1.

[0123] like Figure 14b As shown, the crossover processor 1170 comprises, alone or in addition to other elements, a delay stage 1172 for delaying the further decoded first signal portion and for feeding the delayed decoded first signal portion into the de-emphasis stage 1144 of the second decoding processor for initialization. Furthermore, the crossover processor additionally or alternatively comprises a pre-emphasis filter 1173 and a delay stage 1175 for filtering and delaying the further decoded first signal portion and for providing the delayed output of the block 1175 into the LPC synthesis filter stage 1143 of the ACELP decoder for initialization purposes.

[0124] Furthermore, the crossover processor may alternatively or in addition to the other mentioned elements comprise an LPC analysis filter 1174 for generating a prediction residual signal from the further decoded first signal portion or the pre-emphasized further decoded first signal portion and for feeding the data into a codebook synthesizer of the second decoding processor and preferably into the adaptive codebook stage 1141. Furthermore, the output of the frequency-time converter 1171 with a low sampling rate is also input into the QMF analysis stage 1471 of the upsampler 1210 for initialization purposes, i.e. when the currently decoded audio signal portion is delivered by the frequency domain full band decoder 1120.

[0125] The preferred audio decoder is described below: The waveform decoder part consists of a full-band TCX decoder path and an IGF, both of which operate at the codec's input sampling rate. In parallel, there is an alternative ACELP decoder path at a lower sampling rate, which is further enhanced downstream by TD-BWE.

[0126] For ACELP initialization when switching from TCX to ACELP, there is a cross-path that performs the ACELP initialization of the present invention (consisting of sharing the TCX decoder front end, but additionally providing output at a lower sampling rate and some post-processing). Sharing the same sampling rate and filter order between TCX and ACELP in LPC allows for easier and more efficient ACELP initialization.

[0127] To visualize the switch, Figure 14b Two switches are drawn in Figure 1. When the second switch downstream selects between TCX / IGF or ACELP / TD-BWE outputs, the first switch either pre-updates the buffer in the resampling QMF stage downstream of the ACELP path with the output of the crossover path or simply passes the ACELP output.

[0128] In summary, preferred aspects of the invention, which may be used alone or in combination, relate to the combination of ACELP and TD-BWE coders with full-band capable TCX / IGF techniques, preferably associated with the use of crossover signals.

[0129] Another specific feature is the crossover signal paths used for ACELP initialization to enable seamless handover.

[0130] Another aspect is that the short IMDCT is fed with the lower portion of the high rate long MDCT coefficients to achieve efficient sample rate conversion in the cross path.

[0131] Another feature is the efficient implementation of the crossover path shared with the full-band TCX / IGF part in the decoder.

[0132] Another feature is the crossover signal path for QMF initialization to enable seamless switching from TCX to ACELP.

[0133] An additional feature is a cross signal path to QMF which allows compensation of the delay gap between the ACELP resampled output and the filter bank - TCX / IGF output when switching from ACELP to TCX.

[0134] Another aspect is that LPC is provided for both TCX and ACELP coders at the same sampling rate and filtering order, although the TCX / IGF coder / decoder is full-band capable.

[0135] Then, Figure 14c A preferred implementation of a time domain decoder is discussed that operates either as a standalone decoder or in combination with a full-band capable frequency domain decoder.

[0136] Typically, the time domain decoder comprises an ACELP decoder 1500, followed by a resampler or upsampler and a time domain bandwidth extension function. In particular, the ACELP decoder comprises an ACELP decoding stage 1149 for recovering the gain and innovating the codebook, an ACELP adaptive codebook stage 1141, an ACELP post-processor 1142, an LPC synthesis filter 1143 controlled by quantized LPC coefficients from a bitstream demultiplexer or a coded signal parser, and a subsequent de-emphasis stage 1144. Preferably, the time domain residual signal at the ACELP sampling rate is input to a time domain bandwidth extension decoder 1220, which provides the high frequency band at its output.

[0137] In order to upsample the de-emphasis 1144 output, an upsampler is provided comprising a QMF analysis block 1471 and a QMF synthesis block 1473. Within the filter bank domain defined by blocks 1471 and 1473, a bandpass filter is preferably applied. In particular, as already discussed above, the same functions may also be used, which have been discussed with respect to the same reference numerals. Furthermore, the time domain bandwidth extension decoder 1220 may be as described above. Figure 13 The implementation is shown and generally includes upsampling the ACELP residual signal or the time domain residual signal at the ACELP sampling rate, which is ultimately upsampled to the output sampling rate of the bandwidth extended signal.

[0138] Then, about Figures 1a-5c Further details on full-band capable frequency domain encoders and decoders are discussed.

[0139] Figure 1aAn apparatus for encoding an audio signal 99 is shown. The audio signal 99 is input to a time-to-spectral converter 100, which converts the audio signal having a sampling rate into a spectral representation 101 output by the time-to-spectral converter. The spectrum 101 is input to a spectrum analyzer 102 for analyzing the spectral representation 101. The spectrum analyzer 101 is configured to determine a first group of first spectral portions 103 to be encoded at a first spectral resolution and a second, different group of second spectral portions 105 to be encoded at a second spectral resolution. The second spectral resolution is smaller than the first spectral resolution. The second group of second spectral portions 105 are input to a parameter calculator or parameter encoder 104 for calculating spectral envelope information having the second spectral resolution. Furthermore, a spectral domain audio encoder 106 is provided for generating a first encoded representation 107 of the first group of first spectral portions having the first spectral resolution. Furthermore, the parameter calculator / parameter encoder 104 is configured to generate a second encoded representation 109 of the second group of second spectral portions. The first coded representation 107 and the second coded representation 109 are input into a bitstream multiplexer or bitstream former 108 and the block 108 ultimately outputs the coded audio signal for transmission or storage on a storage device.

[0140] Typically, the first spectrum portion (e.g. Figure 3a 306) will be surrounded by two second spectral portions such as 307a, 307b. This is not the case in HE AAC, where the core encoder frequency range is band-limited.

[0141] Figure 1b Shown with Figure 1a The first encoded representation 107 is input to a spectral domain audio decoder 112 for generating first decoded representations of a first set of first spectral portions, the decoded representations having a first spectral resolution. Furthermore, the second encoded representation 109 is input to a parametric decoder 114 for generating second decoded representations of a second set of second spectral portions having a second spectral resolution lower than the first spectral resolution.

[0142] The decoder also includes a frequency regenerator 116 for regenerating reconstructed second spectral portions having a first spectral resolution using the first spectral portions. Frequency regenerator 116 performs a patch filling operation, i.e., uses patches or portions of the first set of first spectral portions and copies the first set of first spectral portions into a reconstruction range or reconstruction band having the second spectral portions, and typically performs spectral envelope shaping or another operation indicated by the decoded second representation output by parametric decoder 114 (i.e., by using information about the second set of second spectral portions). The decoded first set of first spectral portions and the reconstructed second set of spectral portions are input to a spectrum-to-time converter 118, as indicated at the output of frequency regenerator 116 on line 117. Spectrum-to-time converter 118 is configured to convert the first decoded representation and the reconstructed second spectral portions into a time representation 119 having a high sampling rate.

[0143] Figure 2b Shown Figure 1a The audio input signal 99 is input to the corresponding Figure 1a The analysis filter bank 220 of the time-to-spectral converter 100 is then used. A temporal noise shaping operation is then performed in the TNS block 222. Thus, the time noise shaping operation corresponding to Figure 2b Block tone mask 226 Figure 1a The input to the spectrum analyzer 102 of can be the full spectrum value when no temporal noise shaping / temporal block shaping operation is applied, or can be the full spectrum value when a temporal noise shaping / temporal block shaping operation is applied. Figure 2b , the TNS operation shown in block 222 may be a spectrum residual value. For a two-channel signal or a multi-channel signal, a joint channel coding 228 may be additionally performed so that Figure 1a The spectral domain encoder 106 may include a joint channel encoding block 228. In addition, an entropy encoder 232 is provided for performing lossless data compression, which is also Figure 1a Part of the spectral domain encoder 106.

[0144] The spectrum analyzer / tone mask 226 separates the output of the TNS block 222 into a core frequency band and tonal components corresponding to the first set of first spectrum portions 103 and tonal components corresponding to the first set of first spectrum portions 103. Figure 1a The block 224 indicating the encoding for the IGF parameter extraction corresponds to Figure 1a The parameter encoder 104 and the bitstream multiplexer 230 correspond to Figure 1a bit stream multiplexer 108.

[0145] Preferably, the analysis filterbank 222 is implemented as an MDCT (Modified Discrete Cosine Transform filterbank) and the MDCT is used to transform the signal 99 into the time-frequency domain with a Modified Discrete Cosine Transform used as a frequency analysis tool.

[0146] The spectrum analyzer 226 preferably applies a pitch mask. This pitch mask estimation stage is used to separate the tonal components from the noise-like components in the signal. This allows the core encoder 228 to encode all tonal components using the psychoacoustic module. The pitch mask estimation stage can be implemented in many different ways and is preferably similar in its functionality to the sinusoidal track estimation stage used in sinusoidal and noise modeling for speech / audio coding [8,9] or in the HILN model-based audio encoder described in

[10] . Preferably, an implementation that is easy to implement without the need to maintain birth and death tracks is used, but any other pitch or noise detector may also be used.

[0147] The IGF module calculates the similarity between the source region and the target region. The target region will be represented by the spectrum from the source region. The similarity between the source region and the target region is measured using the cross-correlation method. The target region is divided into Non-overlapping frequency patches. For each patch in the target region, create Source patches. These source patches overlap by a factor between 0 and 1, where 0 means 0% overlap and 1 means 100% overlap. Each of these source patches is correlated with the target patch at various lags to find the source patch that best matches the target patch. The best matching patch number is stored in The lag at which it best correlates with the target is stored in and the sign of the correlation is stored in In the case of very negative correlation, the source patch needs to be multiplied by -1 before patch filling at the decoder. The IGF module also takes care not to overwrite the tonal components in the spectrum, as a tone mask is used to preserve them. The band energy parameter is used to store the energy of the target region, allowing us to accurately reconstruct the spectrum.

[0148] This approach has certain advantages over conventional SBR [1] in that the harmonic grid of the multi-tone signal is preserved by the core encoder, while only the gaps between the sinusoids are filled with the best-matched "shaped noise" from the source region. Another advantage of this system compared to ASR (Exact Spectral Replacement) [2-4] is that there is no signal synthesis stage, which creates the important parts of the signal at the decoder. Instead, this task is taken over by the core encoder, making it possible to preserve the important components of the spectrum. Another advantage of the proposed system is the continuous scalability provided by the features. Simply use and , is called granular matching and can be used for low bit rates while using a variable This allows us to better match the target and source spectra.

[0149] Furthermore, a patch selection stabilization technique is proposed to remove frequency domain artifacts such as judder and musical noise.

[0150] In case of stereo channel pairs, an additional joint stereo processing is applied. This is necessary because for a certain destination range the signals can be highly correlated panned sources. In case the source regions selected for this particular area are not well correlated, the spatial image may be impaired due to the uncorrelated source regions, although the energy matches the destination area. The encoder analyses each destination area band, typically performing a cross correlation of the spectral values ​​and setting a joint flag for the band if a certain threshold is exceeded. In the decoder, if the joint stereo flag is not set, the left and right channel bands are processed separately. In case the joint stereo flag is set, both energy and patching are performed in the joint stereo domain. Similar to the joint stereo information for core coding, the joint stereo information for the IGF area is signaled, including a flag indicating in case of prediction whether the direction of prediction is from downmix to residual or vice versa.

[0151] The energy can be calculated based on the transmitted energy in the L / R domain.

[0152]

[0153]

[0154] in is the frequency index in the transform domain.

[0155] Another solution is to compute and send the energy directly in the joint stereo domain for the bands where joint stereo is active, so no additional energy transform is needed on the decoder side.

[0156] Source tiles are always created from a mid / side matrix:

[0157]

[0158]

[0159] Energy Adjustment:

[0160]

[0161]

[0162] Joint Stereo -> LR Transform:

[0163] If no additional prediction parameters are encoded:

[0164]

[0165]

[0166] If additional prediction parameters are coded and if the signaled direction is from the center to a side:

[0167]

[0168] If the signaling direction is from one side to the middle:

[0169]

[0170] This process ensures that, based on the tiles used to regenerate highly correlated destination regions and panned destination regions, even if the source regions are uncorrelated, the resulting left and right channels still represent correlated and panned sound sources, thereby preserving the stereo image for such regions.

[0171] In other words, in the bitstream, a joint stereo flag is sent that indicates whether L / R or M / S should be used as an example of general joint stereo coding. In the decoder, first, the core signal is decoded as indicated by the joint stereo flag for the core band. Second, the core signal is stored in both L / R and M / S representations. For IGF patch filling, the source patch representation is selected to fit the target patch representation as indicated by the joint stereo information for the IGF band.

[0172] Temporal noise shaping (TNS) is a standard technique and is part of AAC [11-13]. TNS can be considered an extension of the basic scheme of the perceptual encoder, inserting an optional processing step between the filter bank and the quantization stage. The main task of the TNS module is to hide the quantization noise generated in temporally masked regions of the transient-like signal, and thus it leads to a more efficient coding scheme. First, TNS calculates a set of prediction coefficients using "forward prediction" in the transform domain (e.g., MDCT). These coefficients are then used to flatten the temporal envelope of the signal. Since quantization affects the spectrum after TNS filtering, the quantization noise is also temporally flat. By applying an inverse TNS filter on the decoder side, the quantization noise is shaped according to the TNS-filtered temporal envelope, and the quantization noise is thus masked by transients.

[0173] IGF is based on MDCT representation. For efficient coding, preferably, long blocks of about 20ms must be used. If the signal within such long blocks contains transients, audible pre-echoes and post-echoes may occur in the IGF spectral band due to block filling.

[0174] This pre-echo effect is reduced by using TNS in the context of IGF. Here, TNS is used as a temporal tile shaping (TTS) tool, since the spectrum regeneration in the decoder is performed on the TNS residual signal. The full spectrum is used as usual on the encoder side and the required TTS prediction coefficients are applied. The TNS / TTS start and stop frequencies are not affected by the IGF start frequency of the IGF tool. IGFstart Compared with conventional TNS, the TTS stop frequency increased to that of the IGF tool, which is higher than f IGFstart On the decoder side, the TNS / TTS coefficients are applied again to the full spectrum, i.e. the core spectrum plus the regenerated spectrum plus the tonal component from the pitch mask. The application of TTS is necessary to shape the temporal envelope of the regenerated spectrum to match the envelope of the original signal again. Thus, the pre-echo shown is reduced. Furthermore, it is still as usual with TNS below f IGFstart The quantization noise is shaped in the signal.

[0175] In traditional decoders, spectral patching on an audio signal destroys spectral correlation at patch boundaries and thereby impairs the temporal envelope of the audio signal by introducing dispersion. Therefore, another benefit of performing IGF patch filling on the residual signal is that, after applying the shaping filter, the patch boundaries are seamlessly correlated, resulting in a more faithful temporal reproduction of the signal.

[0176] In the encoder of the present invention, the spectrum that has undergone TNS / TTS filtering, tone masking, and IGF parameter estimation contains no signals above the IGF start frequency except for the tonal components. This sparse spectrum is now encoded by the core encoder using the principles of arithmetic coding and predictive coding. These encoded components, together with the signaling bits, form the audio bitstream.

[0177] Figure 2a The corresponding decoder implementation is shown in FIG. Figure 2a The bit stream in is input to the demultiplexer / decoder 200, which converts the Figure 1b Connected to blocks 112 and 114. The bitstream demultiplexer separates the input audio signal into Figure 1b The first coded representation 107 and Figure 1b The first coded representation having the first set of first spectral portions is input to the second coded representation 109 corresponding to Figure 1b The second coded representation is input to the joint channel decoding block 204 of the spectral domain decoder 112 of Figure 2a The parameter decoder 114 not shown in FIG is then input to the corresponding Figure 1bThe first set of first spectral portions required for frequency regeneration is input to IGF block 202 via line 203. Furthermore, after joint channel decoding 204, specific core decoding is applied in tone mask block 206, so that the output of tone mask 206 corresponds to the output of spectral domain decoder 112. Combining, i.e., frame building, is then performed by combiner 208, where the output of combiner 208 now has a full-range spectrum, but still in the TNS / TTS filtered domain. In block 210, an inverse TNS / TTS operation is then performed using the TNS / TTS filtering information provided via line 109. The TTS side information is preferably included in the first encoded representation produced by spectral domain encoder 106 (e.g., a straight AAC or USAC core encoder), or may also be included in the second encoded representation. At the output of block 210, the complete spectrum is provided up to the maximum frequency, which is the full range of frequencies defined by the sampling rate of the original input signal. Then, a spectrum / time conversion is performed in the synthesis filter bank 212 to finally obtain an audio output signal.

[0178] Figure 3a The spectrum is subdivided into scale factor bands SCB, where Figure 3a In the example shown in FIG, there are seven scale factor bands SCB1 to SCB7. The scale factor bands may be AAC scale factor bands defined in the AAC standard and have an increased bandwidth for upper frequencies, such as Figure 3a Schematically shown. Preferably, rather than performing smart gap filling from the very beginning of the spectrum, i.e., at low frequencies, the IGF operation is initiated at the IGF start frequency shown at 309. Thus, the core frequency band extends from the lowest frequency to the IGF start frequency. Above the IGF start frequency, spectral analysis is applied to separate the high-resolution spectral components 304, 305, 306, 307 (the first group of first spectral portions) from the low-resolution components represented by the second group of second spectral portions. Figure 3a Spectra are shown as exemplary inputs to the spectral domain encoder 106 or the joint channel encoder 228, i.e., the core encoder operates in the full range, but encodes a large number of zero spectral values, i.e., these zero spectral values ​​are quantized to zero or set to zero before or after quantization. In any case, the core encoder operates in the full range, i.e., as the spectrum would be shown, i.e., the core decoder does not necessarily have to be aware of any intelligent gap filling or encoding of the second set of second spectral portions with lower spectral resolution.

[0179] Preferably, the high resolution is defined by line-wise coding of spectral lines, such as MDCT lines, while the second or low resolution is defined by, for example, computing only a single spectral value per scale factor band, where a scale factor band covers several frequency lines. Thus, with respect to its spectral resolution, the second or low resolution is much lower than the first or high resolution defined by line-wise coding typically applied by a core encoder (e.g., an AAC or USAC core encoder).

[0180] Regarding the scaling factor or energy calculation, the situation is Figure 3b Due to the fact that the encoder is a core encoder and due to the fact that there may be, but not necessarily have to be, components of the first set of spectral portions in each frequency band, the core encoder operates not only in the core range below the IGF start frequency 309, but also above the IGF start frequency up to a maximum frequency f IGFstop Calculate the scaling factor for each frequency band where the maximum frequency is less than or equal to half the sampling frequency, i.e., f s / 2 .therefore, Figure 3a The coded tone parts 302, 304, 305, 306, 307 and in this embodiment correspond to high-resolution spectral data together with scale factors SCB1 to SCB7. The low-resolution spectral data is calculated starting from the IGF start frequency and corresponds to energy information values ​​E1, E2, E3, E4, which are transmitted together with scale factors SF4 to SF7.

[0181] In particular, when the core encoder is in low bitrate conditions, an additional noise filling operation in the core band (i.e. at frequencies lower than the IGF start frequency, i.e. in the scale factor bands SCB1 to SCB3) can be applied. In the noise filling, there are several adjacent spectral lines that have been quantized to zero. At the decoder side, these quantized to zero spectral values ​​are resynthesized and used, for example, Figure 3b The resynthesized spectral values ​​are adjusted in terms of their amplitude by the noise filling energy of NF2 shown at 308 in FIG. The noise filling energy, which can be given in absolute terms or in relative terms, in particular with respect to a scaling factor as in USAC, corresponds to the energy of the set of spectral values ​​quantized to zero. These noise filled spectral lines can also be considered as a third set of third spectral portions that are regenerated by direct noise filling synthesis without any IGF operation that relies on frequency regeneration using frequency patches from other frequencies, said IGF operation being used to reconstruct the spectral patches using the spectral values ​​from the source range and the energy information E1, E2, E3, E4.

[0182] Preferably, the frequency bands for which the energy information is calculated coincide with the scale factor bands. In other embodiments, grouping of the energy information values ​​is applied, such that, for example, only a single energy information value is sent for scale factor bands 4 and 5, but even in this embodiment, the boundaries of the grouped reconstruction bands coincide with the boundaries of the scale factor bands. If a different frequency band spacing is applied, some recalculation or synchronization of the calculations may be applied, and this may be meaningful depending on the specific implementation.

[0183] Preferably, Figure 1a The spectral domain encoder 106 is as follows Figure 4a Typically, as shown in, for example, the MPEG2 / 4AAC standard or the MPEG1 / 2, layer 3 standard, the audio signal to be encoded ( Figure 4a 401) are forwarded to a scale factor calculator 400. The scale factor calculator is controlled by a psychoacoustic model 402, which in turn receives the audio signal to be quantized or, as in the MPEG1 / 2 Layer 3 or MPEG AAC standards, a complex spectral representation of the audio signal. The psychoacoustic model 402 calculates a scale factor representing a psychoacoustic threshold for each scale factor band. Furthermore, the scale factor is adjusted, through a well-known collaboration of inner and outer iterative loops or any other suitable coding process, so that certain bitrate conditions are met. The spectral values ​​to be quantized and the calculated scale factors are then input to a quantizer processor 404. In direct audio coder operation, the spectral values ​​to be quantized are weighted by the scale factors and then input to a fixed quantizer, which typically has a compression function for the upper amplitude range. The quantizer processor then outputs quantization indices, which are then forwarded to an entropy encoder, which typically has a specific and very efficient encoding for a set of zero quantization indices (or, as it is also known in the art, a "stretch" of zero values) for adjacent frequency values.

[0184] However, in Figure 1a In an audio encoder of the type described herein, the quantizer processor typically receives information about the second spectral portion from the spectrum analyzer. Thus, the quantizer processor 404 ensures that, in the output of the quantizer processor 404, the second spectral portion, as identified by the spectrum analyzer 102, is zero or has a representation that is confirmed by the encoder or decoder to be a zero representation, which can be encoded very efficiently, especially when there are "stretches" of zero values ​​in the spectrum.

[0185] Figure 4b4. The implementation of the quantizer processor is shown. The MDCT spectrum value can be input into a set to zero block 410. Then, before the weighting by the scale factor in execution block 412, the second spectrum portion has been set to zero. In an additional implementation, block 410 is not provided, but the set to zero collaboration is performed in block 418 after the weighting block 412. In an even further implementation, the set to zero operation can also be performed in a set to zero block 422 after the quantization in the quantizer block 420. In this implementation, blocks 410 and 418 will not exist. Typically, at least one of blocks 410, 418, and 422 is provided depending on the specific implementation.

[0186] Then, at the output of block 422, the value corresponding to Figure 3a The quantized spectrum is then input into a Figure 2b In an entropy encoder such as 232 in , it can be, for example, a Huffman encoder or an arithmetic encoder defined in the USAC standard.

[0187] The zero blocks 410, 418, 422 provided alternately or in parallel with each other are controlled by a spectrum analyzer 424. The spectrum analyzer preferably comprises any implementation of a known pitch detector, or any other kind of detector operable to separate the spectrum into a component to be coded at a high resolution and a component to be coded at a low resolution. Other such algorithms implemented in the spectrum analyzer may be a voice activity detector, a noise detector, a speech detector, or any other detector, depending on the spectrum information or associated metadata regarding the resolution requirements for the different spectral portions.

[0188] Figure 5a shows as implemented, for example, in AAC or USAC Figure 1aThe preferred implementation of the time spectrum converter 100 of FIG. Time spectrum converter 100 comprises the window adder 502 that is controlled by the transient detector 1020 of transient detector 504 or Figure 14 a. When transient detector 504 detects transient, then the switching from long window to short window is notified to the window adder with signal. Window adder 502 is then for overlapping block calculation windowing frame, and wherein each windowing frame has two N values ​​usually, for example 2048 values. Then, the conversion in the execution block converter 506, and this block converter provides extraction in addition usually, makes to carry out combined extraction / conversion to obtain the spectrum frame with N value (for example MDCT spectrum value). Therefore, for the long window operation, the frame at the input of piece 506 comprises two N values, for example 2048 values, and spectrum frame then has 1024 values. Then, however, when eight short blocks are executed, a switch is performed on the short blocks, where each short block has 1 / 8 the windowed time domain values ​​compared to the long window, and each spectral block has 1 / 8 the spectral values ​​compared to the long block. Thus, when this decimation is combined with the 50% overlap operation of the windower, the spectrum is a critically sampled version of the time domain audio signal 99.

[0189] Then, refer to Figure 5b , which shows Figure 1b A specific implementation of the frequency regenerator 116 and the spectrum-time converter 118, or Figure 2a The specific implementation of the combined operation of blocks 208 and 212. Figure 5b In , a specific reconstruction band is considered, e.g. Figure 3a The scaling factor of the band is 6. The first spectral part in the reconstructed band, i.e. Figure 3a The first spectrum portion 306 of the scale factor band 6 is input to the frame builder / adjuster block 510. In addition, the reconstructed second spectrum portion for scale factor band 6 is also input to the frame builder / adjuster 510. In addition, energy information (such as Figure 3b E3) is also input to block 510. The reconstructed second spectral portion in the reconstruction band has been generated by frequency patching using the source range, and the reconstruction band then corresponds to the target range. Now, energy adjustment of the frame is performed so that then finally an energy of Figure 2aA fully reconstructed frame with N values ​​is obtained at the output of combiner 208. Then, in block 512, an inverse block transform / interpolation is performed to obtain 248 time-domain values ​​for, for example, the 124 spectral values ​​at the input of block 512. A synthesis windowing operation is then performed in block 514, again controlled by the long / short window indication sent as auxiliary information in the encoded audio signal. Then, in block 516, an overlap / add operation with the previous time frame is performed. Preferably, the MDCT applies a 50% overlap, so that for each new time frame of 2N values, N time-domain values ​​are ultimately output. A 50% overlap is highly preferred because it provides key sampling and continuous crossover from one frame to the next due to the overlap / add operation in block 516.

[0190] like Figure 3a As shown in 301, for example, Figure 3a In order to obtain the desired reconstruction frequency band consistent with the scale factor band 6, a noise filling operation can be additionally applied not only below the IGF start frequency but also above the IGF start frequency. The noise filling spectrum value can then also be input into the frame builder / adjuster 510, and the adjustment of the noise filling spectrum value can also be applied within the block, or the noise filling spectrum value can be adjusted using the noise filling energy before being input into the frame builder / adjuster 510.

[0191] Preferably, the IGF operation can be applied in the complete spectrum, i.e. the frequency patch filling operation using spectrum values ​​from other parts. Thus, the spectrum patch filling operation can be applied not only to the high frequency band above the IGF start frequency, but also to the low frequency band. Furthermore, noise filling without frequency patch filling can be applied not only below the IGF start frequency, but also above the IGF start frequency. However, it has been found that high quality and high efficiency audio coding can be obtained when the noise filling operation is restricted to a frequency range below the IGF start frequency and when the frequency patch filling operation is restricted to a frequency range above the IGF start frequency, e.g. Figure 3a shown.

[0192] Preferably, the target patch (TT) (with frequencies greater than the IGF start frequency) is constrained to the scale factor band boundaries of the full rate encoder. The source patch (ST) from which information is obtained (i.e., for frequencies below the IGF start frequency) is not constrained to the scale factor band boundaries. The size of the ST should correspond to the size of the associated TT. This is illustrated using the following example. TT[0] has a length of 10 MDCT bins. This corresponds exactly to the length of the two subsequent SCBs (e.g., 4+6). Then, all possible STs associated with TT[0] also have a length of 10 bins. The second target patch TT[1] adjacent to TT[0] has a length of 15 bins (SCBs have a length of 7+8). Then, the ST for it has a length of 15 bins instead of 10 bins for TT[0].

[0193] If it happens that no TT of the ST with the length of the target patch can be found (when, for example, the length of the TT is larger than the available source range), the correlation is not calculated and the source range is copied to this TT multiple times (the copying is done one after another so that the frequency line of the lowest frequency of the second copy follows (in terms of frequency) the frequency line of the highest frequency for the first copy) until the target patch TT is completely filled.

[0194] Then, refer to Figure 5c , which shows Figure 1b Frequency regenerator 116 or Figure 2a Another preferred embodiment of the IGF block 202. Block 522 is a frequency patch generator that receives not only the target band ID but also the source band ID. Exemplarily, the Figure 3a The scale factor band of 1 is very well suited for reconstructing scale factor band 7. Therefore, the source band ID will be 2, and the target band ID will be 7. Based on this information, the frequency patch generator 522 applies an upward copy or harmonic patch filling operation or any other patch filling operation to generate the original second portion of spectral components 523. The original second portion of spectral components has the same frequency resolution as the frequency resolution included in the first set of first spectral portions.

[0195] Then, the first spectral portion of the band is reconstructed (e.g. Figure 3a307) is input to the frame builder 524, and the original second part 523 is also input to the frame builder 524. The reconstructed frame is then adjusted by the adjuster 526 using the gain factor of the reconstructed frequency band calculated by the gain factor calculator 528. However, it is important that the first spectral portion in the frame is not affected by the adjuster 526, but only the original second part of the reconstructed frame is affected by the adjuster 526. To this end, the gain factor calculator 528 analyzes the source frequency band or the original second part 523 and additionally analyzes the first spectral portion in the reconstructed frequency band to ultimately find the correct gain factor 527 so that the energy of the frame output after adjustment by the adjuster 526 has an energy E4 when considering the scale factor band 7.

[0196] In this context, it is very important to evaluate the high frequency reconstruction accuracy of the present invention compared to HE-AAC. Figure 3a This is explained using the scale factor Band 7 in the prior art. Assume that a prior art encoder detects a spectral portion 307 to be encoded at high resolution as a "missing harmonic." The energy of this spectral component is then sent to the decoder, along with the spectral envelope information for the reconstruction band (e.g., scale factor Band 7). The decoder will then recreate the missing harmonics. However, the spectral value at which the prior art decoder reconstructs missing harmonics 307 will be in the middle of Band 7, at a frequency indicated by reconstruction frequency 390. Thus, the present invention avoids frequency error 391 that would otherwise be introduced by the prior art decoder.

[0197] In one implementation, the spectrum analyzer is further implemented to calculate a similarity between the first spectral portion and the second spectral portion and to determine, for the second spectral portion in the reconstruction range, a first spectral portion that matches the second spectral portion as closely as possible based on the calculated similarity. Then, in this variable source range / destination range implementation, the parametric encoder will additionally introduce matching information into the second encoded representation, which matching information indicates the matching source range for each destination range. On the decoder side, this information will then be used by Figure 5c The frequency patch generator 522 uses, Figure 5c The generation of the original second portion 523 based on the source band ID and the target band ID is shown.

[0198] In addition, if Figure 3a As shown, the spectrum analyzer is configured to analyze the spectral representation up to a maximum analysis frequency which is just a small amount below half the sampling frequency and preferably at least a quarter of the sampling frequency or typically higher.

[0199] As shown, the encoder operates without downsampling and the decoder operates without upsampling.In other words, the spectral domain audio encoder is configured to generate a spectral representation having a Nyquist frequency defined by the sampling rate of the original input audio signal.

[0200] In addition, if Figure 3a As shown, the spectrum analyzer is configured to analyze a spectrum representation starting with a gap filling start frequency and ending with a maximum frequency represented by a maximum frequency included in the spectrum representation, wherein a spectrum portion extending from the minimum frequency up to the gap filling start frequency belongs to a first group of spectrum portions, and wherein another spectrum portion (such as 304, 305, 306, 307) having a frequency value higher than the gap filling frequency is additionally included in the first group of first spectrum portions.

[0201] As outlined, the spectral domain audio decoder 112 is configured such that the maximum frequency represented by the spectral values ​​in the first decoded representation is equal to the maximum frequency comprised in the time representation having the sampling rate, wherein the spectral value for the maximum frequency is zero or different from zero in the first set of first spectral components. In any case, for this maximum frequency in the first set of spectral components, there is a scale factor for a scale factor band, which is generated and transmitted regardless of whether all spectral values ​​in the scale factor band are set to zero, as Figure 3a and 3b discussed in the context of .

[0202] The present invention is therefore advantageous over other parametric techniques for increasing compression efficiency, such as noise substitution and noise filling (techniques specifically designed for efficient representation of noise-like local signal content), by allowing accurate frequency reproduction of tonal components. To date, no prior art technique has addressed the efficient parametric representation of arbitrary signal content by spectral gap filling without the constraint of a fixed a priori split between the low-frequency (LF) and high-frequency (HF) bands.

[0203] Embodiments of the inventive system improve upon prior art methods, providing high compression efficiency with little or no perceptual annoyance even for low bit rates and full audio bandwidth.

[0204] Typical systems include:

[0205] •Full-band core encoding

[0206] • Intelligent gap filling (block filling or noise filling)

[0207] • Sparse tonal fractions in the core selected by the tone mask

[0208] • Full-band joint stereo pair coding, including patch filling

[0209] • TNS on tiles

[0210] • Spectral whitening within the IGF range

[0211] The first step toward a more efficient system is to remove the need to transform the spectral data into a second transform domain, distinct from the one used by the core encoder. Since most audio codecs (such as AAC, for example) use the MDCT as the base transform, performing BWE in the MDCT domain is also beneficial. A second requirement for a BWE system is the need to preserve the pitch grid, thereby preserving even high-frequency tonal components and improving the quality of the encoded audio compared to existing systems. To address both of these requirements for a BWE solution, a new system called Intelligent Gap Filling (IGF) has been proposed. Figure 2b shows a block diagram of the proposed system on the encoder side, and Figure 2a The system at the decoder side is shown.

[0212] Subsequently, further optional features of the full-band frequency domain first encoding processor and the full-band frequency domain decoding processor incorporating a gap filling operation are discussed and defined, which may be implemented separately or together.

[0213] In particular, the spectral domain decoder 112, corresponding to block 1122a, is configured to output a sequence of decoded frames of spectral values, the decoded frames being a first decoded representation, wherein the frame comprises spectral values ​​for the first set of spectral portions and a zero indication for the second spectral portion. The apparatus for decoding further comprises a combiner 208. The spectral values ​​are generated by a frequency regenerator for the second set of second spectral portions, wherein both the combiner and the frequency regenerator are included in block 1122b. Thus, by combining the second spectral portions with the first spectral portions, a reconstructed spectral frame comprising spectral values ​​of the first set of first spectral portions and the second set of spectral portions is obtained, and corresponds to Figure 14b The spectrum-to-time converter 118 of the IMDCT block 1124 in then converts the reconstructed spectrum frame into a time representation.

[0214] As outlined, the spectrum-to-time converter 118 or 1124 is configured to perform an inverse modified discrete cosine transform 512, 514 and further comprises an overlap-add stage 516 for overlapping and adding subsequent time-domain frames.

[0215] In particular, the spectral domain audio decoder 1122a is configured to generate the first decoded representation such that the first decoded representation has a Nyquist frequency defining a sampling rate equal to the sampling rate of the time representation generated by the spectrum-to-time converter 1124 .

[0216] Furthermore, the decoder 1112 or 1122a is configured to generate the first decoded representation such that the first spectrum portion 306 is placed with respect to a frequency between the two second spectrum portions 307a, 307b.

[0217] In another embodiment, the maximum frequency represented by its spectral value in the first decoded representation is equal to the maximum frequency comprised in the time representation produced by the spectrum-to-time converter, wherein its spectral value is zero or different from zero in the first representation.

[0218] In addition, as in Figure 3a As shown in , the encoded first audio signal portion also includes an encoded representation of a third group of third spectral portions to be reconstructed by noise filling, and the first decoding processor 1120 further includes a noise filler included in block 1122b, for extracting noise filling information 308 from the encoded representation of the third group of third spectral portions and for applying a noise filling operation in the third group of third spectral portions without using the first spectral portions in a different frequency range.

[0219] Furthermore, the spectral domain audio decoder 112 is configured to generate a first decoded representation having a first spectral portion having a frequency value greater than a frequency that is equal to a frequency in the middle of the frequency range covered by the time representation output by the spectrum-to-time converter 118 or 1124.

[0220] Furthermore, a spectrum analyzer or full-band analyzer 604 is configured to analyze the representation generated by the time-to-frequency converter 602 for determining a first set of first spectral portions to be coded with a first high spectral resolution and a second, different set of second spectral portions to be coded with a second spectral resolution lower than the first spectral resolution, and to determine, by means of the spectrum analyzer, the frequency domains of the first and second spectral portions. Figure 3a A first spectrum portion 306 is provided between two second spectrum portions at 307a and 307b in FIG.

[0221] In particular, the spectrum analyzer is configured for analyzing the spectral representation up to a maximum analysis frequency, which is at least one quarter of the sampling frequency of the audio signal.

[0222] In particular, the spectral domain audio encoder is configured to process a frame sequence of spectral values ​​for quantization and entropy encoding, wherein, in the frame, the spectral values ​​of the second group of second parts are set to zero, or wherein, in the frame, there are spectral values ​​of the first group of first spectral parts and the second group of second spectral parts, and wherein, during subsequent processing, the spectral values ​​in the second group of spectral parts are set to zero, as exemplarily shown at 410, 418, 422.

[0223] The spectral domain audio encoder is configured to generate a spectral representation having a Nyquist frequency defined by a sampling rate of the audio input signal or a first portion of the audio signal processed by a first encoding processor operating in the frequency domain.

[0224] The spectral domain audio encoder 606 is further configured to provide a first encoded representation such that for a frame of the sampled audio signal, the encoded representation comprises a first set of first spectral portions and a second set of second spectral portions, wherein spectral values ​​in the second set of spectral portions are encoded as zero or noise values.

[0225] The full-band analyzer 604 or 102 is configured to analyze a spectral representation starting with the gap filling start frequency 309 and ending with a maximum frequency fmax represented by the maximum frequency included in the spectral representation, and the spectral portion extending from the minimum frequency up to the gap filling start frequency 309 belongs to the first group of first spectral portions.

[0226] In particular, the analyzer is configured to apply a tone masking process to at least a portion of the spectral representation such that tonal components and non-tonal components are separated from each other, wherein a first set of first spectral portions comprises tonal components and wherein a second set of second spectral portions comprises non-tonal components.

[0227] The present invention can be further implemented by the following embodiments, which can be combined with any of the examples and embodiments described and claimed herein:

[0228] 1. An audio encoder for encoding an audio signal, comprising:

[0229] A first encoding processor (600) is configured to encode a first audio signal portion in a frequency domain, wherein the first encoding processor (600) comprises:

[0230] a time-to-frequency converter (602) for converting the first audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the first audio signal portion;

[0231] an analyzer (604) for analyzing the frequency domain representation up to the maximum frequency to determine a plurality of first spectral portions to be encoded with a first spectral resolution and a plurality of second spectral portions to be encoded with a second spectral resolution, the second spectral resolution being lower than the first spectral resolution, wherein the analyzer (604) is configured to determine a first spectral portion (306) of the plurality of first spectral portions that is arranged between two second spectral portions (307a, 307b) of the plurality of second spectral portions with respect to frequency;

[0232] a spectral encoder (606) for encoding the plurality of first spectral portions with the first spectral resolution and encoding the plurality of second spectral portions with the second spectral resolution, wherein the spectral encoder comprises a parameter encoder for calculating spectral envelope information with the second spectral resolution from the plurality of second spectral portions;

[0233] A second encoding processor (610) for encoding different second audio signal portions in the time domain;

[0234] a controller (620) configured to analyze the audio signal and to determine which portion of the audio signal is a first audio signal portion encoded in the frequency domain and which portion of the audio signal is a second audio signal portion encoded in the time domain; and

[0235] A coded signal former (630) is provided for forming a coded audio signal, the coded audio signal comprising a first coded signal portion for the first audio signal portion and a second coded signal portion for the second audio signal portion.

[0236] 2. The audio encoder of embodiment 1, wherein the input signal has a high frequency band and a low frequency band,

[0237] Wherein, the second encoding processor (610) comprises: a sampling rate converter (900) for converting the second audio signal portion into a lower sampling rate representation, the lower sampling rate being lower than the sampling rate of the audio signal, wherein the lower sampling rate representation does not include a high frequency band of the input signal;

[0238] A time domain low-band encoder (910) for time domain encoding the lower sampling rate representation; and

[0239] A time-domain bandwidth extension encoder (920) is used to encode the high frequency band in a parametric manner.

[0240] 3. The audio encoder according to embodiment 1, further comprising:

[0241] A preprocessor (1000) configured to preprocess the first audio signal portion and the second audio signal portion,

[0242] Wherein, the preprocessor includes:

[0243] A prediction analyzer (1002) for determining prediction coefficients; and

[0244] Wherein, the second encoding processor includes:

[0245] a prediction coefficient quantizer (1010) for generating a quantized version of the prediction coefficient; and

[0246] an entropy encoder for producing an encoded version of the quantized prediction coefficients,

[0247] The coded signal former (630) is configured to introduce the coded version into the coded audio signal.

[0248] 4. The audio encoder according to embodiment 1,

[0249] wherein the pre-processor (1000) comprises a resampler (1004) for resampling the audio signal to a sampling rate of the second encoding processor; and

[0250] wherein the prediction analyzer is configured to determine prediction coefficients using the resampled audio signal, or

[0251] The pre-processor (1000) further comprises a long-term prediction analysis stage (1006) for determining one or more long-term prediction parameters for the first audio signal portion.

[0252] 5. The audio encoder according to embodiment 1 further comprises a cross processor (700) for calculating initialization data for a second encoding processor (610) based on the encoded spectral representation of the first audio signal portion, so that the second encoding process (610) is initialized to encode a second audio signal portion of the audio signal that temporally follows the first audio signal portion.

[0253] 6. The audio encoder according to embodiment 5, wherein the crossover processor (700) comprises:

[0254] a spectral decoder (701) for calculating a decoded version of the first encoded signal portion;

[0255] a delay stage (707) for feeding a delayed version of the decoded version into a de-emphasis stage (617) of a second encoding processor for initialization;

[0256] a weighted prediction coefficient analysis filter block (708) for feeding the filter output into a codebook determiner (613) of a second encoding processor (610) for initialization;

[0257] an analysis filtering stage (706) for filtering the decoded version or the pre-emphasized (709) version and for feeding the filtered residual into an adaptive codebook determiner (612) of a second encoding processor for initialization; or

[0258] A pre-emphasis filter (709) for filtering the decoded version and for feeding the delayed or pre-emphasized version to a synthesis filter stage (616) of a second encoding processor (610) for initialization.

[0259] 7. The audio encoder according to embodiment 1,

[0260] wherein the analyzer (604) is configured to perform a temporal block shaping or temporal noise shaping analysis or an operation of setting spectral values ​​in the second spectral portion to zero,

[0261] wherein the first encoding processor (600) is configured to perform a shaping (606a) of spectral values ​​of the first spectral portion using prediction coefficients (1010) derived from the first audio signal portion, and wherein the first encoding processor (600) is further configured to perform a quantization and entropy coding operation (606b) of the shaped spectral values ​​of the first spectral portion, and

[0262] The spectrum value of the second spectrum part is set to zero.

[0263] 8. The audio encoder according to embodiment 7, further comprising a crossover processor (700), wherein the crossover processor (700) comprises:

[0264] a noise shaper (703) for shaping the quantized spectral values ​​of the first spectral portion using LPC coefficients (1010) derived from the first audio signal portion;

[0265] a spectral decoder (704, 705) for decoding the spectrally shaped spectral portion of the first spectral portion with a high spectral resolution and for synthesizing the second spectral portion using the parametric representation of the second spectral portion and at least the decoded first spectral portion to obtain a decoded spectral representation;

[0266] A frequency-to-time converter (702) is configured to convert the spectral representation into the time domain to obtain a decoded first audio signal portion, wherein a sampling rate associated with the decoded first audio signal portion is different from a sampling rate of the audio signal, and a sampling rate associated with an output signal of the frequency-to-time converter (702) is different from a sampling rate of the audio signal input to the frequency-to-time converter (602).

[0267] 9. The audio encoder of embodiment 1, wherein the second encoding processor comprises at least one block from the following group of blocks:

[0268] Measure analysis filter (611);

[0269] Adaptive codebook stage (612);

[0270] Innovation codebook level (614);

[0271] an estimator (613) for estimating innovation codebook entries;

[0272] ACELP / gain coding stage (615);

[0273] prediction synthesis filter stage (616);

[0274] de-emphasis stage (617); and

[0275] low bass post-filter analysis stage (618).

[0276] 10. The audio encoder of embodiment 1,

[0277] wherein the time domain coding processor has an associated second sampling rate,

[0278] wherein the frequency domain coding processor has a first sampling rate associated therewith which is higher than the second sampling rate, wherein the audio encoder further comprises a cross-over processor (700) for calculating initialization data for the second coding processor from an encoded spectral representation of the first audio signal portion,

[0279] wherein the cross-over processor comprises a frequency-time converter (702) for generating a time domain signal at the second sampling rate,

[0280] wherein the frequency-time converter (702) comprises:

[0281] a selector (726) for selecting a low portion of the spectrum input into the frequency-time converter depending on a ratio of the first sampling rate and the second sampling rate, the ratio of the first sampling rate and the second sampling rate being smaller than one,

[0282] a transform processor (720) having a transform length which is smaller than a transform length of the time-frequency converter (602); and

[0283] a synthesis windower (712) for windowing using a window having a smaller number of window coefficients compared to a window used by the time-frequency converter (602).

[0284] 11. An audio decoder for decoding an encoded audio signal, comprising:

[0285] a first decoding processor (1120) for decoding a first encoded audio signal portion in the frequency domain, the first decoding processor (1120) comprising:

[0286] a spectral decoder (1122) for decoding the plurality of first spectral portions with a high spectral resolution and for synthesizing the plurality of second spectral portions using the parametric representation of the plurality of second spectral portions and at least the decoded first spectral portions to obtain a decoded spectral representation, wherein the spectral decoder (1122) is configured to produce a first decoded representation such that a first spectral portion (306) is arranged between two second spectral portions (307a, 307b) with respect to frequency; and

[0287] a frequency-time converter (1120) for converting the decoded spectral representation into the time domain to obtain a decoded first audio signal portion;

[0288] a second decoding processor (1140) for decoding the second encoded audio signal portion in the time domain to obtain a decoded second audio signal portion; and

[0289] a combiner (1160) for combining the decoded first spectral portion and the decoded second spectral portion to obtain a decoded audio signal.

[0290] 12. The audio decoder according to claim 11, wherein the second decoding processor comprises:

[0291] a time-domain low-band decoder (1200) for decoding a low-band time-domain signal;

[0292] an up-sampler (1210) for up-sampling the low-band time-domain signal;

[0293] a time-domain bandwidth extension decoder (1220) for synthesizing a high-band of the time-domain output signal; and

[0294] a mixer (1230) for mixing the synthesized high-band of the time-domain signal and the up-sampled low-band time-domain signal.

[0295] 13. The audio encoder according to claim 12,

[0296] wherein the up-sampler (1210) comprises an analysis filter bank (1471) operating at a first time-domain low-band decoder sampling rate and a synthesis filter bank (1473) operating at a second output sampling rate higher than the first time-domain low-band sampling rate.

[0297] 14. The audio decoder according to claim 12,

[0298] wherein the time-domain low-band decoder (1200) comprises a residual signal, a decoder (1149, 1141, 1142) and a synthesis filter (1143) for filtering the residual signal using synthesis filter coefficients (1145),

[0299] The time-domain bandwidth extension decoder (1220) is configured to upsample the residual signal (1221), process (1222) the upsampled residual signal using a nonlinear operation to obtain a high-band residual signal, and perform spectrum shaping (1223) on the high-band residual signal to obtain a synthesized high-band.

[0300] 15. The audio decoder according to embodiment 11,

[0301] The first decoding processor (1120) comprises an adaptive long-term prediction postfilter (1420) for post-filtering the first decoded first signal portion, wherein the filter (1420) is controlled by one or more long-term prediction parameters included in the encoded audio signal.

[0302] 16. The audio decoder according to embodiment 11, further comprising:

[0303] A cross processor (1170) is configured to calculate initialization data for a second decoding processor (1140) from a decoded spectral representation of the first encoded audio signal portion, such that the second decoding processor (1140) is initialized to decode an encoded second audio signal portion of the encoded audio signal that temporally follows the first audio signal portion.

[0304] 17. The audio decoder according to embodiment 16, wherein the cross processor further comprises:

[0305] a frequency-to-time converter (1170) operating at a lower sampling rate than the frequency-to-time converter (1124) of the first decoding processor (1120) to obtain a first signal portion for further decoding in the time domain,

[0306] wherein the signal output by the frequency-to-time converter (1171) has a second sampling rate that is lower than the first sampling rate associated with the output of the frequency-to-time converter (1124) of the second decoding processor,

[0307] wherein the additional frequency-to-time converter (1171) comprises: a selector (726) for selecting a low portion of a frequency spectrum input to the additional frequency-to-time converter (1171) based on a ratio between a first sampling rate and a second sampling rate, the ratio between the first sampling rate and the second sampling rate being less than 1;

[0308] a transform processor (720) having a transform length smaller than a transform length (710) of a time-to-frequency converter (1124); and

[0309] The synthesis windower (722) uses a window with a smaller number of coefficients than the window used by the frequency-to-time converter (1124).

[0310] 18. The audio decoder of embodiment 16, wherein the crossover processor (1170) comprises:

[0311] a delay stage (1172) for delaying the further decoded first signal portion and for feeding a delayed version of the decoded first signal portion into a de-emphasis stage (1144) of a second decoding processor for initialization;

[0312] a pre-emphasis filter (1173) and a delay stage (1175) for filtering and delaying the first signal portion for further decoding, and for feeding the delay stage output into a predictive synthesis filter (1143) of a second decoding processor for initialization;

[0313] a prediction analysis filter (1174) for generating a prediction residual signal from the further decoded first spectral portion or the pre-emphasized (1173) further decoded first signal portion and for feeding the prediction residual signal to a codebook synthesizer (1141) of a second decoding processor (1200); or

[0314] A switch (1480) for feeding the further decoded first signal portion into an analysis stage (1471) of a resampler (1210) of a second decoding processor for initialization.

[0315] 19. The audio decoder according to embodiment 11,

[0316] wherein the second decoding processor (1200) comprises at least one block in a block group, the block group comprising:

[0317] ACELP for decoding gain and innovation codebooks;

[0318] Adaptive codebook synthesis stage (1141);

[0319] ACELP postprocessor (1142);

[0320] Predictive synthesis filter (1143); and

[0321] De-emphasis level (1144).

[0322] 20. A method for encoding an audio signal, comprising:

[0323] A first encoding (600) of a first audio signal portion is performed in the frequency domain, wherein the first encoding (600) comprises:

[0324] converting (602) a first audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the first audio signal portion;

[0325] analyzing (604) the frequency domain representation up to the maximum frequency to determine a plurality of first spectral portions to be encoded with a first spectral resolution and a plurality of second spectral portions to be encoded with a second spectral resolution, the second spectral resolution being lower than the first spectral resolution, wherein the analyzing (604) determines a first spectral portion (306) of the plurality of first spectral portions to be arranged relative to frequency between two second spectral portions (307a, 307b) of the plurality of second spectral portions;

[0326] encoding (606) the plurality of first spectral portions with the first spectral resolution and the plurality of second spectral portions with the second spectral resolution, wherein encoding a second spectral portion comprises computing spectral envelope information having the second spectral resolution from the plurality of second spectral portions;

[0327] second encoding (610) a different second audio signal portion in the time domain;

[0328] analyzing (620) the audio signal and determining which portion of the audio signal is the first audio signal portion encoded in the frequency domain and which portion of the audio signal is the second audio signal portion encoded in the time domain; and

[0329] forming (630) an encoded audio signal comprising a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion.

[0330] 21. A method of decoding an encoded audio signal, comprising:

[0331] first decoding (1120) a first encoded audio signal portion in the frequency domain, the first decoding (1120) comprising:

[0332] decoding (1122) a plurality of first spectral portions with a high spectral resolution and synthesizing the plurality of second spectral portions using a parametric representation of the plurality of second spectral portions and at least the decoded first spectral portions to obtain a decoded spectral representation, wherein decoding (1122) comprises producing the first decoded representation such that a first spectral portion (306) is arranged relative to frequency between two second spectral portions (307a, 307b); and

[0333] converting (1120) the decoded spectral representation into the time domain to obtain a decoded first audio signal portion;

[0334] decoding (1140) the second encoded audio signal portion in the time domain to obtain a decoded second audio signal portion; and

[0335] combining (1160) the decoded first spectral portion and the decoded second spectral portion to obtain a decoded audio signal.

[0336] 22. A machine readable storage medium, storing a computer program for performing the method according to embodiment 20 or embodiment 21, when executed on a computer or processor.

[0337] Although the application has been described in the context of block diagrams in which the blocks represent real or logical hardware components, the application can also be implemented as a computer-implemented method. In the latter case, the blocks represent corresponding method steps, wherein these steps represent functional entities performed by corresponding logical or physical hardware blocks.

[0338] While some aspects have been described in the context of an apparatus, it is clear that other aspects also represent a description, possibly with different wording, of corresponding method steps for performing the same functions. In some embodiments, one or more of the method steps can be performed by hardware (e.g. by one or more microprocessors or other electronic circuits) or by a combination of hardware and software. Corresponding apparatus can perform one or more of the method steps.

[0339] The transmitted or encoded signal of the application can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0340] Depending on certain implementation requirements, embodiments of the application can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, and an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium can be computer readable.

[0341] Some embodiments according to the application comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0342] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer.The program code may, for example, be stored on a machine-readable carrier.

[0343] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0344] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0345] A further embodiment of the inventive method is, therefore, a data carrier (or a non-transitory storage medium, such as a digital storage medium or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium is typically tangible and / or non-transitory.

[0346] Therefore, a further embodiment of the inventive method is a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can, for example, be configured to be transmitted via a data communication connection (eg via the Internet).

[0347] A further embodiment comprises a processing means, for example a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0348] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0349] Another embodiment according to the present invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.

[0350] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0351] The above examples are merely illustrative of the principles of the application. It will be readily apparent to those skilled in the art that modifications and variations of the arrangements and details described herein can be made without departing from the spirit and scope of the application. Accordingly, it is intended that the scope of the application be limited only by the scope of the claims attached hereto.

Claims

1. An audio encoder for encoding an audio signal, comprising: A first encoding processor (600) for encoding a first audio signal portion in a frequency domain, the first audio signal portion having a sampling frequency associated therewith, wherein the first encoding processor (600) comprises: a time-to-frequency converter (602) for converting the first audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the first audio signal portion, wherein the maximum frequency is less than or equal to half the sampling frequency and is at least one quarter of the sampling frequency or higher; A spectral encoder (606) for encoding the frequency domain representation; A second encoding processor (610) for encoding different second audio signal portions in the time domain; wherein the second encoding processor (610) has an associated second sampling frequency, wherein the first encoding processor (600) has a first sampling frequency associated therewith, the first sampling frequency being different from the second sampling frequency; A crossover processor (700) for calculating initialization data for a second encoding processor (610) based on the encoded spectral representation of the first audio signal portion, such that the second encoding process (610) is initialized to encode a second, different audio signal portion of the audio signal that temporally follows the first audio signal portion, wherein the crossover processor (700) comprises a frequency-to-time converter (702) for generating a time domain signal at a second sampling frequency, wherein the frequency-to-time converter (702) comprises: a selector (726) for selecting a portion of the frequency spectrum input to the frequency-to-time converter (702) based on a ratio between the first sampling frequency and the second sampling frequency; a transform processor (720) having a transform length different from a transform length of the time-to-frequency converter (602); and a synthesis windower (712) for windowing using a window having a different number of window coefficients than the window used by the time-to-frequency converter (602); a controller (620) configured to analyze the audio signal and to determine which portion of the audio signal is a first audio signal portion encoded in the frequency domain and which portion of the audio signal is a different second audio signal portion encoded in the time domain; and A coded signal former (630) is provided for forming a coded audio signal, the coded audio signal comprising a first coded signal portion for a first audio signal portion and a second coded signal portion for a second, different audio signal portion.

2. The audio encoder according to claim 1, wherein The audio signal has a high frequency band and a low frequency band. The second encoding processor (610) includes: a sampling rate converter (900) for converting the different second audio signal portions into a lower sampling frequency representation, the lower sampling frequency being lower than the sampling frequency of the audio signal, wherein the lower sampling frequency representation does not include a high frequency band of the audio signal; A time domain low band encoder (910) for time domain encoding of the lower sampling frequency representation; and A time-domain bandwidth extension encoder (920) is used to encode the high frequency band in a parametric manner.

3. The audio encoder according to claim 1 , further comprising: A preprocessor (1000) configured to preprocess a first audio signal portion and a different second audio signal portion, Wherein, the pre-processor (1000) includes a prediction analyzer (1002) for determining a prediction coefficient; The encoding signal former (630) is configured to introduce an encoded version of the prediction coefficients into the encoded audio signal.

4. The audio encoder according to claim 1, wherein the pre-processor (1000) comprises a resampler (1004) for resampling the audio signal to a second sampling frequency of the second encoding processor (610); and wherein the prediction analyzer is configured to determine prediction coefficients using the resampled audio signal, or The pre-processor (1000) further comprises a long-term prediction analysis stage (1024) for determining one or more long-term prediction parameters for the first audio signal portion.

5. The audio encoder according to claim 1, wherein The cross processor (700) comprises: a spectral decoder for computing a decoded version of the first encoded signal portion; a delay stage (707) for feeding a delayed version of the decoded version into a de-emphasis stage (617) of a second encoding processor (610) for initialization; a weighted prediction coefficient analysis filter block (708) for feeding the filter output into a codebook estimator (613) of a second encoding processor (610) for initialization; an analysis filtering stage (706) for filtering the decoded version or the pre-emphasized version and for feeding the filtered residue into an adaptive codebook determiner (612) of a second encoding processor (610) for initialization; or A pre-emphasis filter (709) for filtering the decoded version and for feeding the delayed or pre-emphasized version to a synthesis filter stage (616) of a second encoding processor (610) for initialization.

6. The audio encoder according to claim 1, wherein the first encoding processor (600) is configured to perform a shaping (606a) of spectral values ​​of the frequency domain representation using prediction coefficients derived from the first audio signal portion, and wherein the first encoding processor (600) is further configured to perform quantization and entropy coding operations (606b) of the shaped spectral values ​​of the frequency domain representation.

7. The audio encoder according to claim 1, wherein The cross processor (700) comprises: a noise shaper (703) for shaping the quantized spectral values ​​of the frequency domain representation using LPC coefficients derived from the first audio signal portion; and A spectral decoder is configured to decode the spectrally shaped spectral portion of the frequency domain representation at a high spectral resolution to obtain a decoded spectral representation.

8. The audio encoder according to claim 1, wherein the second encoding processor (610) comprises at least one block from the following group of blocks: Predictive Analysis Filter (611); Adaptive codebook determiner (612); Innovation codebook level (614); A codebook estimator (613) for estimating innovative codebook entries; ACELP / Gain Coding Stage (615); predictive synthesis filter stage (616); De-emphasis stage (617); and Bass Post Filter Analysis Stage (618).

9. An audio decoder for decoding an encoded audio signal, comprising: a first decoding processor (1120) for decoding the first encoded audio signal portion in the frequency domain, the first decoding processor (1120) comprising a frequency-time converter (1124) for converting the decoded spectral representation into the time domain to obtain a decoded first audio signal portion, wherein the decoded spectral representation extends up to a maximum frequency of the time representation of the decoded audio signal, the spectral value of the maximum frequency being zero or different from zero; a second decoding processor (1140) for decoding the second encoded audio signal portion in the time domain to obtain a decoded second audio signal portion; a cross processor (1170) for calculating initialization data for a second decoding processor (1140) from the decoded spectral representation of the first encoded audio signal portion, such that the second decoding processor (1140) is initialized to decode a second encoded audio signal portion of the encoded audio signal that temporally follows the first encoded audio signal portion; a combiner (1160) for combining the decoded first audio signal portion and the decoded second audio signal portion to obtain a decoded audio signal, The cross processor (1170) further includes: A further frequency-to-time converter (1171) is configured to obtain a further decoded first audio signal portion in the time domain, wherein the further decoded first audio signal portion output by the further frequency-to-time converter (1171) has a second sampling frequency different from a first sampling frequency, wherein the first sampling frequency is associated with an output of the frequency-to-time converter (1124) of the first decoding processor (1120), wherein the further frequency-to-time converter (1171) comprises: a selector (726) for selecting a portion of the frequency spectrum input to the further frequency-to-time converter (1171) based on a ratio between the first sampling frequency and the second sampling frequency; a transform processor (720) having a transform length different from a transform length (710) of a frequency-to-time converter (1124) of said first decoding processor (1120); and A synthesis windower (722) uses a window having a different number of coefficients than the window used by the frequency-to-time converter (1124) of the first decoding processor (1120).

10. The audio decoder according to claim 9, wherein The second decoding processor (1140) includes: A time domain low frequency band decoder (1200), configured to decode to obtain a low frequency band time domain signal; A resampler (1210) for resampling the low-band time domain signal; A time-domain bandwidth extension decoder (1220) for synthesizing a high-frequency band of a time-domain output signal; and A mixer (1230) is used to mix the high frequency band of the synthesized time domain output signal and the resampled low frequency band time domain signal.

11. The audio decoder according to claim 9, The first decoding processor (1120) comprises an adaptive long-term prediction postfilter (1420) for post-filtering the decoded first audio signal portion, wherein the postfilter (1420) is controlled by one or more long-term prediction parameters included in the encoded audio signal.

12. The audio decoder of claim 9, wherein the crossover processor (1170) comprises: a first delay stage (1172) for delaying the further decoded first audio signal portion and for feeding a delayed version of the further decoded first audio signal portion into a de-emphasis stage (1144) of a second decoding processor (1140) for initialization; a pre-emphasis filter (1173) and a second delay stage (1175) for filtering and delaying the further decoded portion of the first audio signal, and for feeding the output of the second delay stage (1175) into a predictive synthesis filter (1143) of a second decoding processor (1140) for initialization; a prediction analysis filter (1174) for generating a prediction residual signal from the further decoded first audio signal portion or the pre-emphasized further decoded first audio signal portion and for feeding the prediction residual signal to a codebook synthesizer (1141) of a second decoding processor (1140); or A switch (1480) for feeding the further decoded first audio signal portion into an analysis stage (1471) of a resampler (1210) of a second decoding processor (1140) for initialization.

13. The audio decoder according to claim 9, The second decoding processor (1140) includes at least one block from the following block groups: A stage for decoding the ACELP gain and the innovative codebook; A codebook synthesizer (1141) for synthesizing an adaptive codebook; ACELP postprocessor (1142); Predictive synthesis filter (1143); and De-emphasis level (1144).

14. A method for encoding an audio signal, comprising: Encoding a first audio signal portion in the frequency domain, the first audio signal portion having a sampling frequency associated therewith, the encoding comprising: converting the first audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the first audio signal portion, wherein the maximum frequency is less than or equal to half the sampling frequency and at least one quarter of the sampling frequency or higher; and Encoding the frequency domain representation; encoding the different second audio signal parts in the time domain; wherein encoding different second audio signal portions has an associated second sampling frequency; wherein encoding the first audio signal portion has a first sampling frequency associated therewith, the first sampling frequency being different from the second sampling frequency; calculating initialization data for the step of encoding a different second audio signal portion based on the encoded spectral representation of the first audio signal portion, such that the step of encoding the different second audio signal portion is initialized to encode a different second audio signal portion that temporally follows the first audio signal portion in the audio signal; wherein the calculating comprises generating a time domain signal at a second sampling frequency by a frequency to time converter (702), wherein the generating comprises: selecting a portion of the frequency spectrum input to a frequency-to-time converter (702) according to a ratio between a first sampling frequency and a second sampling frequency, processing using a transform processor (720) having a transform length different from a transform length of a time-to-frequency converter (602) for converting the first audio signal portion, and performing synthesis windowing (712) using a window having a different number of window coefficients than a window used by a time-to-frequency converter (602) for converting the first audio signal portion, analyzing the audio signal and determining which portion of the audio signal is a first audio signal portion encoded in the frequency domain and which portion of the audio signal is a different second audio signal portion encoded in the time domain; and An encoded audio signal is formed, the encoded audio signal comprising a first encoded signal portion for a first audio signal portion and a second encoded signal portion for a second, different audio signal portion.

15. A method for decoding an encoded audio signal, comprising: decoding the first encoded audio signal portion in the frequency domain by a first decoding processor (1120), the decoding comprising: converting the decoded spectral representation into the time domain by a frequency-time converter (1124) to obtain a decoded first audio signal portion, wherein the decoded spectral representation extends up to a maximum frequency of the time representation of the decoded audio signal, the spectral value of the maximum frequency being zero or different from zero; decoding the second encoded audio signal portion in the time domain to obtain a decoded second audio signal portion; calculating initialization data for the step of decoding the second encoded audio signal portion based on the decoded spectral representation of the first encoded audio signal portion, such that the step of decoding the second encoded audio signal portion is initialized to decode the second encoded audio signal portion of the encoded audio signal that temporally follows the first encoded audio signal portion; and combining the decoded first audio signal portion and the decoded second audio signal portion to obtain a decoded audio signal, The calculation also includes: using a further frequency-time converter (1171) configured for obtaining a further decoded first audio signal portion in the time domain, wherein the further decoded first audio signal portion in the time domain output by the further frequency-to-time converter (1171) has a second sampling frequency different from a first sampling frequency, wherein the first sampling frequency is associated with the output of the frequency-to-time converter (1124) of the first decoding processor (1120), The use of another frequency-to-time converter (1171) includes: selecting a portion of the frequency spectrum input to the further frequency-to-time converter (1171) in accordance with a ratio between a first sampling frequency and a second sampling frequency; using a transform processor (720) having a transform length different from the transform length (710) of a frequency-to-time converter (1124) of the first decoding processor (1120); and A synthesis windower (722) is used that uses a window having a different number of coefficients than the window used by a frequency-to-time converter (1124) of the first decoding processor (1120).

Citation Information

Patent Citations

  • Audio encoder and decoder using a frequency domain processor, a time domain processor, and a cross processor for continuous initialization

    CN106796800A