Audio encoders and decoders with frequency domain processors and time domain processors

By combining a full-band spectrum encoder with a time-domain encoder, the intelligent gap-filling technology solves the audio quality problem of existing audio encoders when the high-frequency bandwidth is extended, realizing efficient encoding and high-quality reconstruction of full-band audio signals, and supporting seamless switching to time-domain bandwidth extension.

CN113963705BActive Publication Date: 2026-04-21FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2015-07-24
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing audio coding technologies suffer from reduced audio quality when processing high-frequency bandwidth extensions, especially non-speech signals. In particular, they lack sufficient coding accuracy for prominent harmonics, and frequency domain encoders rely on bandwidth extensions, resulting in inflexibility.

Method used

By combining a full-band spectrum encoder with a time-domain encoder, high-resolution encoding is performed in the frequency domain using intelligent gap filling (IGF) technology to fill spectral holes. Combined with an adaptive frequency patching scheme and time noise shaping technology, high-quality reconstruction of full-band audio signals is ensured.

Benefits of technology

It achieves efficient encoding and decoding of full-band audio signals, avoids the bandwidth limitations of frequency domain encoders, improves audio quality, especially the fidelity of the tonal part in the high-frequency band, and supports seamless switching to time domain bandwidth extension.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113963705B_ABST
    Figure CN113963705B_ABST
Patent Text Reader

Abstract

An audio encoder for encoding an audio signal, comprising a first encoding processor (600) for encoding a first audio signal portion in the frequency domain, the first encoding processor (600) comprising a time-to-frequency converter (602), an analyzer (604) for analyzing a frequency domain representation up to a maximum frequency to determine a first spectral portion to be encoded with a first spectral resolution and a second spectral region to be encoded with a second spectral resolution, the second spectral resolution being lower than the first spectral resolution, a spectral encoder (606) for encoding the first spectral portion with the first spectral resolution and for encoding the second spectral portion with the second spectral resolution, a second encoding processor (610) for encoding a second, different audio signal portion in the time domain, a controller (620), and an encoded signal former (630).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese invention patent application No. 201580049740.7, filed on March 15, 2017, entitled "Audio encoder and decoder using a frequency domain processor and a time domain processor with full bandgap filling". Technical Field

[0002] This invention relates to audio signal encoding and decoding, and particularly to audio signal processing using parallel frequency-domain and time-domain encoder / decoder processors. Background Technology

[0003] Perceptual coding of audio signals is a widely used practice for the purpose of efficient storage or data reduction in audio signal transmission. In particular, when aiming for the lowest possible bit rate, the coding employed results in a degraded audio quality, typically due to limitations on the encoder side of the bandwidth of the audio signal to be transmitted. Here, the audio signal is usually low-pass filtered so that no spectral waveform content remains above a predetermined cutoff frequency.

[0004] In contemporary codecs, there are known methods for decoder-side signal recovery via audio signal bandwidth extension (BWE), such as spectral band copying (SBR) operating in the frequency domain or so-called time-domain bandwidth extension (TD-BWE), which is a post-processor in a speech encoder operating in the time domain.

[0005] In addition, there are several combined time-domain / frequency-domain coding concepts, such as those known under the terms AMR-WB+ or USAC.

[0006] All these combined time-domain / coding concepts share the following commonality: the frequency-domain encoder relies on bandwidth-expanding techniques that introduce frequency band limitations into the input audio signal, and the portion above the crossover or boundary frequencies is encoded using a low-resolution coding concept and synthesized on the decoder side. Therefore, these concepts primarily depend on preprocessor techniques on the encoder side and corresponding post-processing functions on the decoder side.

[0007] Typically, time-domain encoders are chosen for useful signals (e.g., speech signals) encoded in the time domain, while frequency-domain encoders are chosen for non-speech signals, music signals, etc. However, especially for non-speech signals with prominent harmonics in the high-frequency band, existing frequency-domain encoders suffer from reduced accuracy and, consequently, reduced audio quality. This is because such prominent harmonics can only be encoded parametrically and separately, or completely eliminated during the encoding / decoding process.

[0008] Furthermore, there exists a concept where the time-domain coding / decoding branch additionally relies on bandwidth expansion, which also parametrically encodes the higher frequency range, while the lower frequency range is typically encoded using ACELP or any other CELP-related encoder (e.g., a speech encoder). This bandwidth expansion functionality increases bit-rate efficiency, but on the other hand, it introduces further inflexibility due to the fact that both coding branches—the frequency-domain coding branch and the time-domain coding branch—are band-limited by a spectral band copying or bandwidth expansion process operating substantially below a certain crossover frequency included in the maximum frequency of the input audio signal.

[0009] Relevant topics in existing technology include

[0010] - SBR as a post-processor for waveform decoding [1-3]

[0011] - MPEG-D USAC core switching[4]

[0012] - MPEG-H 3D IGF [5]

[0013] The following papers and patents describe methods that are considered to constitute the prior art of this application:

[0014] [1] M. Dietz, L. Liljeryd, K. Kjörling and O. Kunz, “Spectral Band Replication, a novel approach in audio coding,” at the 112th AES Congress, Munich, Germany, 2002.

[0015] [2] S. Meltzer, R. Böhm and F. Henn, “SBR enhanced audio codecs for digital broadcasting such as “Digital Radio Mondiale” (DRM),” at the 112th AES Congress, Munich, Germany, 2002.

[0016] [3] T. Ziegler, A. Ehret, P. Ekstrand and M. Lutzky, “Enhancing mp3 with SBR: Features and Capabilities of the new mp3PRO Algorithm,” at the 112th AES Congress, Munich, Germany, 2002.

[0017] [4] MPEG-D USAC standard.

[0018] [5] PCT / EP2014 / 065109.

[0019] In MPEG-D USAC, a switchable core encoder is described. However, in USAC, the band-limited core is restricted to always sending a low-pass filtered signal. Therefore, certain music signals containing prominent high-frequency content, such as full-band scanning and triangular sounds, cannot be faithfully reproduced. Summary of the Invention

[0020] The purpose of this invention is to provide an improved concept for audio coding.

[0021] This objective is achieved by an audio encoding device, an audio decoder, an audio encoding method, an audio decoding method, or a machine-readable storage medium according to embodiments of the present invention.

[0022] This invention is based on the discovery that a time-domain encoder / decoder can be combined with a frequency-domain encoder / decoder that has gap-filling capabilities, but the gap-filling function operates across the entire frequency band of the audio signal or at least above a certain gap-filling frequency. Importantly, the frequency-domain encoder / decoder is particularly capable of encoding / decoding with precise waveform or spectral values ​​up to the maximum frequency, not just up to the crossover frequency. Furthermore, the full-band capability of the frequency-domain encoder for high-resolution encoding allows the gap-filling function to be integrated into the frequency-domain encoder.

[0023] Therefore, according to the present invention, by using a full-bandwidth spectrum encoder / decoder processor, problems related to the separation of bandwidth expansion on the one hand and core encoding on the other hand can be solved and overcome by performing bandwidth expansion in the same spectral domain in which the core decoder operates. Thus, a full-rate core decoder is provided that encodes and decodes the full audio signal range. This eliminates the need for a downsampler on the encoder side and an upsampler on the decoder side. Instead, the entire process is performed in the full sampling rate or full bandwidth domain. To obtain high coding gain, the audio signal is analyzed to find a first set of first spectral portions that must be encoded at high resolution, wherein, in one embodiment, the first set of first spectral portions may include: the tone portion of the audio signal. On the other hand, non-tone or noise components in the audio signal constituting a second set of second spectral portions are parametrically encoded at low spectral resolution. The encoded audio signal then requires only the first set of first spectral portions encoded in a waveform-preserving manner with high spectral resolution, and, furthermore, the second set of second spectral portions encoded parametrically at low resolution using frequency “tiles” derived from the first set. On the decoder side, the core decoder, which is the full-band decoder, reconstructs the first set of first spectral portions in a waveform-preserving manner, i.e., without any knowledge of the existence of any additional frequency regeneration. However, the spectrum thus generated has many spectral gaps. These gaps are then filled with the intelligent gap-filling (IGF) technique of the present invention by using frequency regeneration using the applied parameter data on the one hand and the source spectral range (i.e., the first spectral portion reconstructed by the full-rate audio decoder) on the other hand.

[0024] In another embodiment, the spectral portion reconstructed by noise filling alone, rather than bandwidth replication or frequency patch filling, constitutes a third group of third spectral portions. Due to the fact that the coding concept operates in a single domain for both core encoding / decoding and frequency regeneration, the IGF is not only limited to filling the higher frequency range but can also fill the lower frequency range by using noise filling without frequency regeneration or by using frequency regeneration with frequency patches in different frequency ranges.

[0025] Furthermore, it should be emphasized that information regarding spectral energy, information regarding individual energies, information regarding surviving energies, information regarding piece energies, or information regarding missing energies can include not only energy values ​​but also (e.g., absolute) amplitude values, level values, or any other values ​​from which the final energy value can be derived. Therefore, information regarding energy can, for example, include the energy value itself, and / or the level and / or the amplitude and / or the absolute amplitude value.

[0026] Another aspect is based on the finding that correlation is important not only for the source range but also for the target range. Furthermore, the invention acknowledges situations where different correlations may occur in the source and target ranges. For example, when considering a speech signal with high-frequency noise, it's possible that the low-frequency band of the speech signal, including a small number of overtones, is highly correlated in the left and right channels when the speaker is placed in the center. However, the high-frequency portion may be strongly uncorrelated due to the fact that there may be high-frequency noise on the left that differs from another high-frequency noise, or no high-frequency noise on the right. Therefore, when performing a direct gap-filling operation that ignores this situation, the high-frequency portion will also be correlated, and this can produce severe spatial isolation artifacts in the reconstructed signal. To address this problem, parameter data for the reconstructed frequency band, or generally, parameter data for the second set of second spectral portions that must be reconstructed using the first set of first spectral portions, is calculated to identify a first or second distinct binocular representation for the second spectral portion, or in other words, a first or second distinct binocular representation for the reconstructed frequency band. Therefore, on the encoder side, binocular recognition is calculated for the second spectral portion, i.e., binocular recognition is calculated for the portion where energy information of the reconstructed frequency band is also calculated. The frequency regenerator on the decoder side then regenerates the second spectrum portion based on the first portion of the first set of first spectrum portions (i.e., source range and parameter data for the second portion, such as spectral envelope energy information or any other spectral envelope data) and additionally based on the dual-channel identification for the second portion (i.e., for the reconstructed frequency band under reconsideration).

[0027] The dual-channel identification is preferably sent as a flag for each reconstructed frequency band, and this data is sent from the encoder to the decoder, which then decodes the core signal as indicated by the flag preferably calculated for the core frequency band. In the implementation, the core signal is then stored in a stereo representation (e.g., left / right and center / side), and for IGF frequency patch filling, the source patch representation is selected to fit the target patch representation as indicated by the dual-channel identification flag for intelligent gap filling or reconstructed frequency bands (i.e., for the target range).

[0028] It's important to emphasize that this process works not only for stereo signals (i.e., the left and right channels) but also for multi-channel signals. In the case of multi-channel signals, several different channel pairs can be processed in this way, for example, the left and right channels as the first pair, the left surround and right surround channels as the second pair, and the center channel and LFE channel as the third pair. Other pairings can be determined for higher output channel formats such as 7.1 and 11.1.

[0029] Another aspect is based on the finding that IGF can improve the audio quality of the reconstructed signal because the entire spectrum is accessible to the core encoder, allowing perceptually important tonal portions, such as those in the high-frequency range, to still be encoded by the core encoder instead of by parametric substitution. Additionally, gap-filling is performed using frequency blocks from a first set of first spectral portions, which are, for example, a set of tonal portions typically from the lower frequency range, but also from the higher frequency range (if available). However, for spectral envelope adjustment on the decoder side, the spectral portions from the first set of spectral portions located in the reconstructed band are not further post-processed by, for example, spectral envelope adjustment. Only the remaining spectral values ​​in the reconstructed band that do not originate from the core decoder are enveloped using the envelope information. Preferably, the envelope information is a full-band envelope information taking into account the energy of the first set of first spectral portions in the reconstructed band and the energy of the second set of second spectral portions in the same reconstructed band, wherein the latter spectral values ​​in the second set of second spectral portions are indicated as zero and therefore not encoded by the core encoder, but rather parametrically encoded using low-resolution energy information.

[0030] It has been found that normalized or unnormalized absolute energy values ​​relative to the bandwidth of the corresponding frequency band are useful and highly efficient in decoder-side applications. This is particularly useful when gain factors must be calculated based on residual energy in the reconstructed frequency band, missing energy in the reconstructed frequency band, and frequency patch information in the reconstructed frequency band.

[0031] Furthermore, preferably, the encoded bitstream covers not only the energy information of the reconstructed frequency bands but also the scaling factor of the scaling factor bands extending up to the maximum frequency. This ensures that for each reconstructed frequency band available for a certain tone portion (i.e., the first spectral portion), the first set of first spectral portions can actually be decoded with the correct amplitude. In addition to the scaling factor for each reconstructed frequency band, energy for that reconstructed frequency band is generated in the encoder and sent to the decoder. Furthermore, it is preferable that the reconstructed frequency bands coincide with the scaling factor bands, or, in the case of energy grouping, at least the boundaries of the reconstructed frequency bands coincide with the boundaries of the scaling factor bands.

[0032] On the other hand, it is based on the finding that certain impairments in audio quality can be remedied by applying an adaptive frequency patching scheme. To this end, encoder-side analysis is performed to identify candidate source regions that best match a given target region. Matching information for identifying a source region for the target region, along with optional additional information, is generated and sent as auxiliary information to the decoder. The decoder then uses the matching information to apply the frequency patching operation. For this purpose, the decoder reads the matching information from the transmitted data stream or data file, accesses the source region identified for a given reconstructed frequency band, and, if indicated in the matching information, performs additional processing on the source region data to generate raw spectral data for the reconstructed frequency band. This result of the frequency patching operation (i.e., the raw spectral data for the reconstructed frequency band) is then shaped using spectral envelope information to ultimately obtain a reconstructed frequency band that also includes first spectral components such as tone components. However, these tone components are not generated by the adaptive patching scheme; instead, these first spectral components are directly output by the audio decoder or core decoder.

[0033] The adaptive spectrum patch selection scheme can operate with low granularity. In this implementation, the source region is subdivided into typically overlapping source regions, and the target region, or reconstructed band, is given by non-overlapping frequency target regions. Then, on the encoder side, the similarity between each source region and each target region is determined, and the best matching pair of source and target regions is identified through matching information. On the decoder side, the source regions identified in the matching information are used to generate the raw spectrum data for the reconstructed band.

[0034] To achieve a higher granularity, each source region is allowed to shift to obtain a hysteresis that maximizes similarity. This hysteresis can be as fine as a frequency bin, allowing for even better matching between source and target regions.

[0035] In addition to identifying only the best matching pair, the associated hysteresis can be sent within the matching information, and even the symbol can be sent. When the symbol is determined to be negative on the encoder side, the corresponding symbol flag is then sent within the matching information, and on the decoder side, the source region spectral value is multiplied by "-1", or "rotated" 180 degrees in the complex representation.

[0036] Another implementation of the invention applies a patch whitening operation. Spectral whitening removes coarse spectral envelope information and emphasizes the fine spectral structure most relevant to evaluating patch similarity. Therefore, frequency patches and / or source signals are whitened before calculating cross-correlation measurements. When the patch is whitened using only a predefined procedure, a whitening flag is sent, instructing the decoder to apply the same predefined whitening procedure to frequency patches within the IGF.

[0037] Regarding tile selection, a correlation-dependent hysteresis is preferably used to shift the regenerated spectrum across the spectrum using an integer number of transform bins. Depending on the fundamental transform, the spectrum shift may require additional corrections. In the case of odd hysteresis, the tiles are additionally modulated by multiplying by an alternating time series of -1 / 1 to compensate for the frequency inversion representation every other band within the MDCT. Furthermore, the sign of the correlation result is applied when generating frequency tiles.

[0038] Furthermore, tile pruning and stability are preferably employed to ensure that artifacts created by rapidly changing source regions used for the same reconstructed or target region are avoided. To this end, similarity analysis is performed between different identified source regions, and a source tile can be discarded from the group of potential source tiles if it is similar to other source tiles with a similarity above a threshold, because it is highly correlated with other source tiles. Additionally, as a form of tile selection stability, if none of the source tiles in the current frame are correlated with the target tile in the current frame (between a given threshold), the tile order from previous frames is preferably maintained.

[0039] Another aspect is based on the finding that combining time-noise shaping (TNS) or time-blocking (TTS) techniques with high-frequency reconstruction yields improved quality and reduced bit rate, particularly for signals including transient components (as these frequently occur in audio signals). TNS / TTS processing on the encoder side, implemented through frequency-relative prediction, reconstructs the temporal envelope of the audio signal. According to the implementation, when the time-noise shaping filter is determined to cover not only the source frequency range but also the target frequency range to be reconstructed in the frequency reproduction decoder, the time envelope is applied not only to the core audio signal up to the gap-filling start frequency but also to the spectral range of the reconstructed second spectral portion. Therefore, front echoes or back echoes that would occur without time-blocking are reduced or eliminated. This is achieved by applying inverse frequency prediction not only to the core frequency range up to a certain gap-filling start frequency but also to the frequency range above the core frequency range. For this purpose, frequency reproduction or frequency block generation is performed on the decoder side before applying the frequency-relative prediction. However, the prediction relative to frequency can be applied before or after spectral envelope shaping, depending on whether the energy information calculation has been performed on the filtered spectral residuals or on the (full) spectral values ​​before envelope shaping.

[0040] The TTS processing relative to one or more frequency blocks further establishes continuity of correlation between the source range and the reconstructed range, or between two adjacent reconstructed ranges or frequency blocks.

[0041] In implementation, complex TNS / TTS filtering is preferred. This avoids (temporal) aliasing artifacts in the real-number representations of critical samples (such as MDCT). Besides obtaining the complex-modified transform, the complex TNS filter can be computed on the encoder side by applying not only the modified discrete cosine transform but also the modified discrete sine transform. However, only the modified discrete cosine transform value, i.e., the real part of the complex transform, is transmitted. However, on the decoder side, it is possible to estimate the imaginary part of the transform using the MDCT spectrum of previous or subsequent frames, allowing the complex filter to be applied again on the decoder side for inverse prediction relative to frequency, and specifically, for prediction relative to the boundary between the source and reconstructed ranges, and also relative to the boundary between adjacent frequency patches within the reconstructed range.

[0042] The audio coding system of this invention efficiently encodes arbitrary audio signals over a wide range of bit rates. However, for high bit rates, the system converges to transparency, and for low bit rates, perceptual annoyance is minimized. Therefore, a major share of the available bit rate is used for waveform encoding only the perceptually most relevant structures of the signal in the encoder, and the resulting spectral gaps are filled in the decoder with a signal content that roughly approximates the original spectrum. This parameter-driven, so-called intelligent spectral gap filling (IGF) is controlled by consuming a very limited bit budget through dedicated auxiliary information sent from the encoder to the decoder.

[0043] In another embodiment, the time-domain encoder / decoder processor relies on a lower sampling rate and corresponding bandwidth expansion capabilities.

[0044] In another embodiment, a cross processor is provided to initialize the time-domain encoder / decoder using initialization data derived from the currently processed frequency-domain encoder / decoder signal. This allows the parallel time-domain encoder to be initialized while the currently processed audio signal portion is being processed by the frequency-domain encoder, so that when a switch from the frequency-domain encoder to the time-domain encoder occurs, the time-domain encoder can immediately begin processing because all initialization data associated with the earlier signal is already available due to the cross processor. This cross processor is preferably applied to the encoder side and additionally to the decoder side, and preferably uses a frequency-time transform, which further performs highly efficient downsampling from a higher output or input sample rate to a lower time-domain core encoder sample rate by selecting only a certain low-frequency band portion of the domain signal and a reduced transform size. Thus, the sample rate conversion from high to low sample rate is performed very efficiently, and the signal obtained through the transform with the reduced transform size can then be used to initialize the time-domain encoder / decoder, so that the time-domain encoder / decoder is ready to immediately perform time-domain encoding when this situation is signaled by the controller and the immediately preceding audio signal portion is encoded in the frequency domain.

[0045] Therefore, preferred embodiments of the present invention allow for seamless switching between a perceptual audio encoder including spectral gap filling and a temporal encoder with or without bandwidth extension.

[0046] Therefore, this invention relies on methods that are not limited to removing high-frequency content above the cutoff frequency from an audio signal in a frequency-domain encoder, but rather on adaptively removing spectral bandpass regions that leave spectral gaps in the encoder and subsequently reconstructing these spectral gaps in the decoder. Preferably, an integrated solution such as intelligent gap filling is used, which effectively combines full-bandwidth audio coding and spectral gap filling, particularly in the MDCT transform domain.

[0047] Therefore, the present invention provides an improved concept for combining speech coding and subsequent time-domain bandwidth extension with full-band waveform decoding including spectral gap filling into a switchable perceptual encoder / decoder.

[0048] Therefore, compared to existing methods, the new concept utilizes full-band audio signal waveform encoding in the transform domain encoder and simultaneously allows for seamless switching to the speech encoder, preferably followed by time-domain bandwidth extension.

[0049] Other embodiments of the invention avoid the interpretation problems caused by fixed frequency band limitations. This concept enables a switchable combination of a full-bandwidth waveform encoder and a lower sampling rate speech encoder with temporal bandwidth extension in the frequency domain, equipped with spectral gap filling. Such an encoder is capable of waveform encoding the aforementioned problematic signals, thereby providing the full audio bandwidth up to the Nyquist frequency of the audio input signal. Nevertheless, seamless instantaneous switching between the two encoding strategies is particularly guaranteed by embodiments with cross-processors. For this seamless switching, a cross-processor represents a cross-connection at both the encoder and decoder between a full-bandwidth, full-rate (input sampling rate) frequency domain encoder and a low-rate ACELP encoder with a lower sampling rate, to appropriately initialize ACELP parameters and buffers, particularly within adaptive codebooks, LPC filters, or resampling stages, when switching from a frequency domain encoder such as a TCX to a temporal encoder such as ACELP. Attached Figure Description

[0050] The invention will then be discussed with reference to the accompanying drawings, wherein:

[0051] Figure 1a A device for encoding audio signals is shown;

[0052] Figure 1b It shows the relationship with Figure 1a The encoder is matched with a decoder used to decode the encoded audio signal;

[0053] Figure 2a A preferred implementation of the decoder is shown;

[0054] Figure 2b A preferred implementation of the encoder is shown;

[0055] Figure 3a It shows the result of Figure 1b A schematic representation of the spectrum generated by the frequency domain decoder;

[0056] Figure 3b A table showing the relationship between the scaling factor for the scaling factor band and the energy for the reconstruction band and the noise filling information for the noise filling band is presented.

[0057] Figure 4a The functionality for applying the selection of a spectral portion to a spectral domain encoder in the first and second sets of spectral portions is shown.

[0058] Figure 4b It shows Figure 4a The implementation of the function;

[0059] Figure 5a The function of the MDCT encoder is shown;

[0060] Figure 5b The functionality of a decoder with MDCT technology is shown;

[0061] Figure 5c The implementation of the frequency regenerator is shown;

[0062] Figure 6 The implementation of the audio encoder is shown;

[0063] Figure 7a The cross processor within the audio encoder is shown;

[0064] Figure 7b An additional implementation of the inverse OR frequency-time transformation with reduced sampling rate is shown within the cross-processor;

[0065] Figure 8 It shows Figure 6 The preferred implementation of the controller;

[0066] Figure 9 Another embodiment of a time-domain encoder with bandwidth extension capability is shown;

[0067] Figure 10 The preferred use of the preprocessor is shown;

[0068] Figure 11a A schematic implementation of the audio decoder is shown;

[0069] Figure 11bThe cross processor within the decoder is shown, which provides initialization data for the time-domain decoder.

[0070] Figure 12 It shows Figure 11a A preferred implementation of the time-domain decoding processor;

[0071] Figure 13 Another implementation of time-domain bandwidth extension is shown;

[0072] Figure 14a illustrates a preferred implementation of the audio encoder;

[0073] Figure 14b A preferred implementation of the audio decoder is shown;

[0074] Figure 14c An innovative implementation of a time-domain decoder with sampling rate conversion and bandwidth extension is shown. Detailed Implementation

[0075] Figure 6 An audio encoder for encoding an audio signal is shown, including a first encoding processor 600 for encoding a portion of a first audio signal in the frequency domain. The first encoding processor 600 includes a time-to-frequency converter 602 for converting a portion of the first input audio signal into a frequency domain representation having spectral lines up to the maximum frequency of the input signal. Furthermore, the first encoding processor 600 includes an analyzer 604 for analyzing the frequency domain representation up to the maximum frequency to determine a first spectral region to be encoded using a first spectral representation, and to determine a second spectral region to be encoded using a second spectral resolution lower than the first spectral resolution. Specifically, the analyzer 604 (which may be a full-band analyzer) determines which frequency lines or spectral values ​​in the time-to-frequency converter spectrum should be encoded in a spectral line manner, and which other spectral portions should be encoded parametrically, and then these latter spectral values ​​are reconstructed on the decoder side using a gap-filling process. The actual encoding operation is performed by a spectrum encoder 606, which is used to encode the first spectral region or portion of the spectrum at the first resolution, and to encode the second spectral region or portion parametrically at the second spectral resolution.

[0076] Figure 6The audio encoder also includes a second encoding processor 610 for encoding portions of the audio signal in the time domain. Additionally, the audio encoder includes a controller 620 configured to analyze the audio signal at the audio signal input 601 and to determine which portion of the audio signal is a first audio signal portion encoded in the frequency domain and which portion is a second audio signal portion encoded in the time domain. Furthermore, an encoded signal shaper 630, which can be implemented, for example, as a bitstream multiplexer, is provided and configured to form an encoded audio signal including a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion. Importantly, the encoded signal has only a frequency domain representation or a time domain representation from the same audio signal portion.

[0077] Therefore, controller 620 ensures that for a single audio signal segment, only a time-domain representation or a frequency-domain representation exists in the encoded signal. This can be achieved by controller 620 in several ways. One way would be that for the same audio signal segment, two representations arrive at block 630, and controller 620 controls the encoded signal formler 630 to incorporate only one of the two representations into the encoded signal. Alternatively, however, controller 620 can control the inputs to the first encoding processor and the inputs to the second encoding processor such that, based on the analysis of the corresponding signal segment, only one of blocks 600 or 610 is activated to actually perform the full encoding operation, and the other blocks are deactivated.

[0078] Deactivation can be performed, or alternatively, such as relative to... Figure 7a The example shown is merely an "initialization" mode, where another encoding processor is only active for receiving and processing initialization data to initialize the internal memory, but does not perform any specific encoding operations. This activation can be achieved through... Figure 6 This is accomplished by a switch at an input not shown, or preferably via control lines 621 and 622. Therefore, in this embodiment, when the controller 620 has determined that the current audio signal portion should be encoded by the first encoding processor, while the second encoding processor is still provided with initialization data to be active for future transient switching, the second encoding processor 610 outputs nothing. On the other hand, the first encoding processor is configured not to require any data from the past to update any internal memory, and therefore, when the current audio signal portion is to be encoded by the second encoding processor 610, the controller 620 can control the first end encoding processor 600 to be completely inactive via control line 621. This means that the first encoding processor 600 does not need to be in an initialization or waiting state, but can be in a completely deactivated state. This is particularly preferred for mobile devices where power consumption and therefore battery life are issues.

[0079] In a further specific implementation of the second encoding processor operating in the time domain, the second encoding processor includes a downsampler 900 or a sample rate converter for converting portions of the audio signal into a representation with a lower sample rate, wherein the lower sample rate is lower than the sample rate at the input to the first encoding processor. This is in Figure 9 As shown in the diagram. Specifically, when the input audio signal includes both low-frequency and high-frequency bands, it is preferable that the lower sampling rate representation at the output of block 900 only has the low-frequency band portion of the input audio signal, which is then encoded by a time-domain low-frequency encoder 910, configured to perform time-domain encoding on the lower sampling rate representation provided by block 900. Furthermore, a time-domain bandwidth extension encoder 920 is provided for parametrically encoding the high-frequency band. For this purpose, the time-domain bandwidth extension encoder 920 receives at least the high-frequency band of the input audio signal or both the low-frequency and high-frequency bands of the input audio signal.

[0080] In another embodiment of the invention, the audio encoder further includes (although in Figure 6 Not shown in the image, but... Figure 10 A preprocessor 1000 (shown in FIG. 14a) is configured to preprocess a first audio signal portion and a second audio signal portion. In one embodiment, the preprocessor includes a predictive analyzer for determining predictive coefficients. This predictive analyzer can be implemented as an LPC (Linear Predictive Coding) analyzer for determining LPC coefficients. However, other analyzers may also be implemented. Furthermore, the preprocessor (also shown in FIG. 14a) includes a predictive coefficient quantizer 1010, wherein the device shown in FIG. 14a receives predictive coefficient data from the predictive analyzer, also shown at 1002a and 1002b in FIG. 14a.

[0081] Furthermore, the preprocessor additionally includes an entropy encoder for generating an encoded version of the quantized prediction coefficients. It is important to note that the encoded signal shaper 630, or a specific implementation, namely the bitstream multiplexer 613, ensures that the encoded version of the quantized prediction coefficients is included in the encoded audio signal 632. Preferably, the LPC coefficients are not directly quantized, but are converted to, for example, ISF, or any other representation more suitable for quantization. This conversion is preferably performed by determining LPC coefficient blocks 1002a, 1002b or within block 1010 for quantizing the LPC coefficients.

[0082] Furthermore, the preprocessor may include a resampler 1004 or 1021 in Figure 14a for resampling the audio input signal at the input sampling rate to a lower sampling rate for the time-domain encoder. When the time-domain encoder is an ACELP encoder with a certain ACELP sampling rate, downsampling is preferably performed to 12.8 kHz or 16 kHz. The input sampling rate can be any of a certain number of sampling rates (e.g., 32 kHz or even higher). On the other hand, the sampling rate of the time-domain encoder will be predetermined by certain constraints, and the resampler 1004 performs this resampling and outputs a lower sampling rate representation of the input signal. Therefore, the resampler can perform a similar function, and could even be as follows: Figure 9 The same element as the downsampler 900 shown in the context.

[0083] Furthermore, pre-emphasis is preferably applied in pre-emphasis blocks 1005a and 1005b in Figure 14a. Pre-emphasis processing is well known in the field of time-domain coding and is described in literature referring to AMR-WB+ processing. Pre-emphasis is specifically configured to compensate for spectral tilt and thus allows for better computation of LPC parameters in a given LPC order.

[0084] In addition, the preprocessor may include additional components for control. Figure 14b The TCX-LTP parameter extraction 1024 of the LTP post-filter is shown at 1420 in Figure 14a. This block is shown at 1024 in Figure 14a. Furthermore, the preprocessor may additionally include other functions shown at 1007, and these other functions may include tone search, voice activity detection (VAD), or any other functions known in the temporal or speech coding domain.

[0085] As shown, the result of block 1024 is input into the encoded signal, that is, in the embodiment of FIG14a, it is input into the bitstream multiplexer 630. Furthermore, if desired, data from block 1007 can also be introduced into the bitstream multiplexer, or alternatively, it can be used for time-domain encoding purposes in the time-domain encoder.

[0086] Therefore, in summary, both paths share preprocessing operation 1000, in which common signal processing operations are performed. These include resampling to the ACELP sampling rate (12.8 or 16 kHz) for one of the parallel paths, and this resampling is always performed. Additionally, TCX LTP parameter extraction, shown at block 1024 in Figure 14a, is performed, along with LPC coefficient pre-emphasis 1005a and determination 1002a. As summarized, pre-emphasis 1005a compensates for spectral tilt, thus making the calculation of LPC parameters in a given LPC order more efficient.

[0087] Subsequently, reference Figure 8 To illustrate a preferred implementation of controller 620. The controller receives the considered portion of the audio signal at the input. Preferably, as shown in FIG14a, the controller receives any signal available in preprocessor 1000, which may be the original input signal at the input sampling rate or a resampled version at a lower time-domain encoder sampling rate, or a signal obtained after pre-emphasis processing in block 1005.

[0088] Based on the audio signal portion, controller 620 addresses frequency domain encoder simulator 621 and time domain encoder simulator 622 to calculate an estimated signal-to-noise ratio for each encoder possibility. Subsequently, selector 623 naturally selects the encoder that already provides a better signal-to-noise ratio, taking into account a predefined bit rate. The selector then identifies the appropriate encoder via a control output. When it is determined that the considered audio signal portion will be encoded using the frequency domain encoder, the time domain encoder is set to an initialized state, or in other embodiments, a fully deactivated state where very instantaneous switching is not required. However, when it is determined that the considered audio signal portion will be encoded by the time domain encoder, the frequency domain encoder is deactivated.

[0089] Subsequently, it was shown Figure 8 The preferred implementation of the controller shown is as follows. The decision of whether to select the ACELP or TCX path is executed during the switching decision by simulating the ACELP and TCX encoders and switching to the better execution branch. To this end, the SNR of the ACELP and TCX branches is estimated based on the ACELP and TCX encoder / decoder simulations. The TCX encoder / decoder simulation is performed without TNS / TTS analysis, IGF encoder, quantization loop / arithmetic encoder, or without any TCX decoder. Alternatively, the TCX SNR is estimated using an estimate of the quantizer distortion in the shaped MDCT domain. The ACELP encoder / decoder simulation is performed using only the simulations of the adaptive codebook and the innovation codebook. The ACELP SNR is simply estimated by calculating the distortion introduced by the LTP filter in the weighted signal domain (adaptive codebook) and scaling this distortion by a constant factor (innovation codebook). Therefore, the complexity is significantly reduced compared to methods that perform TCX and ACELP encoding in parallel. The branch with the higher SNR is selected for subsequent full encoding runs.

[0090] When the TCX branch is selected, the TCX decoder runs in each frame, outputting the signal at the ACELT sampling rate. This is used to update the memory for the ACELP encoding paths (LPC remnants, Mem w0, memory de-emphasis) to achieve an instantaneous switch from TCX to ACELP. Memory updates are performed in each TCX path.

[0091] Alternatively, a full analysis can be performed through synthesis processing, where both encoder simulators 621 and 622 perform the actual encoding operations, and the results are compared by selector 623. Alternatively, the complete feedforward calculation can also be performed by performing signal analysis. For example, a time-domain encoder is selected when the signal is determined to be a speech signal by a signal classifier, and a frequency-domain encoder is selected when the signal is determined to be a music signal. Other processes can also be applied to distinguish between the two encoders based on signal analysis of the considered audio signal portion.

[0092] Preferably, the audio encoder further includes Figure 7a The cross processor 700 is shown. When the frequency domain encoder 600 is active, the cross processor 700 provides initialization data to the time domain encoder 610, preparing the time domain encoder for seamless switching in future signal sections. In other words, without the cross processor, such immediate seamless switching would be impossible when the frequency domain encoder determines that the current signal section is to be encoded, and when the controller determines that the immediately following audio signal section is to be encoded by the time domain encoder 610. However, for the purpose of initializing the memory in the time domain encoder, the cross processor provides the time domain encoder 610 with signals derived from the frequency domain encoder 600, because the time domain encoder 610 has a dependency on the signals encoded from the current frame of the input or the frame immediately preceding it in time.

[0093] Therefore, the time-domain encoder 610 is configured to be initialized with initialization data in order to encode the audio signal portion following the earlier audio signal portion encoded by the frequency-domain encoder 600 in an efficient manner.

[0094] Specifically, the cross-processor includes a time converter for converting the frequency domain representation to a time domain representation, which can be forwarded to the time domain encoder directly or after some further processing. This converter is shown in Figure 14a as an IMDCT (Inverse Modified Discrete Cosine Transform) block. However, this block 702 has a different transform size (modified discrete cosine transform block) compared to the time-frequency converter block 602 shown in Figure 14a. As shown in block 602, the time-frequency converter 602 operates at the input sampling rate, while the inverse modified discrete cosine transform 702 operates at a lower ACELP sampling rate.

[0095] It can calculate the ratio of the time-domain encoder sampling rate or ACELP sampling rate to the frequency-domain encoder sampling rate or input sampling rate, and it is Figure 7b The downsampling factor DS is shown. Block 602 has a large transform size, while IMDCT block 702 has a small transform size. For example... Figure 7bAs shown, the IMDCT block 702 therefore includes a selector 726 for selecting the lower spectral portion of the input to the IMDCT block 702. The portion of the full-band spectrum is defined by a downsampling factor DS. For example, when the lower sampling rate is 16 kHz and the input sampling rate is 32 kHz, the downsampling factor is 0.5, therefore, the selector 726 selects the lower half of the full-band spectrum. When the spectrum has, for example, 1024 MDCT lines, the selector selects the lower 512 MDCT lines.

[0096] The low-frequency portion of the full-band spectrum is input into the small-size transform and unfold block 720, such as... Figure 7b As shown. The transform size is also selected based on the downsampling factor and is 50% of the transform size in block 602. Then, synthesis windowing is performed, where the window has a small number of coefficients. The number of coefficients in the synthesis window is equal to the downsampling factor multiplied by the number of coefficients in the analysis window used in block 602. Finally, an overlap-add operation is performed with an even smaller number of operations per block, and the number of operations per block is again the number of operations per block in full-rate MDCT multiplied by the downsampling factor.

[0097] Therefore, highly efficient downsampling operations can be applied because downsampling is included in the IMDCT implementation. In this context, it should be emphasized that block 702 can be implemented by IMDCT, but it can also be implemented by any other transform or filter bank that can be appropriately sized in the actual transform kernel and other transform-related operations.

[0098] In another embodiment shown in Figure 14a, the time-to-frequency converter includes additional functions in addition to the analyzer. Figure 6 The analyzer 604 may include the time noise shaping / time patch shaping analysis block 604a in the embodiment of FIG14a, as in the analysis block 604a for TNS / TTS. Figure 2b Operate as discussed in the context of block 222, and regarding the tone mask 226 corresponding to the IGF encoder 604b in Figure 14a. Figure 2b Operate as shown.

[0099] Furthermore, the frequency domain encoder preferably includes a noise shaping block 606a. The noise shaping block 606a is controlled by quantized LPC coefficients generated as in block 1010. The quantized LPC coefficients used for noise shaping 606a perform spectral shaping of high-resolution spectral values ​​or spectral lines through direct encoding (rather than parametric encoding), and the result of block 606a resembles the spectrum of the signal after an LPC filtering stage, which operates in the time domain (e.g., LPC analysis filter block 704, described later). Furthermore, the result of noise shaping block 606a is then quantized and entropy encoded as shown in block 606b. The result of block 606b corresponds to the encoded first audio signal portion or the frequency-domain encoded audio signal portion (along with other auxiliary information).

[0100] The cross-processor 700 includes a decoded version of the spectrum decoder for computing the first coded signal portion. In the embodiment of FIG14a, the spectrum decoder 701 includes the previously discussed inverse noise shaping block 703, gap-filling decoder 704, TNS / TTS synthesis block 705, and IMDCT block 702. These blocks undo specific operations performed by blocks 602 to 606b. Specifically, the noise shaping block 703 undoes the noise shaping performed by block 606a based on the quantized LPC coefficients 1010. The IGF decoder 704 is as described above. Figure 2a Operate blocks 202 and 206 as discussed, and TNS / TTS composite block 705 as in Figure 2a The operation is as discussed in the context of block 210, and the spectrum decoder additionally includes IMDCT block 702. Furthermore, the cross processor 700 in FIG14a additionally or alternatively includes a delay stage 707 for feeding a delayed version of the decoded version obtained by the spectrum decoder 701 into the deemphasis stage 617 of the second encoding processor for the purpose of initializing the deemphasis stage 617.

[0101] In addition, the cross-processor 700 may additionally or alternatively include a weighted prediction coefficient analysis filter stage 708 for filtering the decoded version and feeding the filtered decoded version to the codebook determiner 613 (marked "MMSE" in FIG. 14a) of the second encoder processor for initializing the block. Additionally or alternatively, the cross-processor includes an LPC analysis filter stage 706 for filtering the decoded version of the first coded signal portion output by the spectrum decoder 701 and feeding the result to the adaptive codebook stage 612 for initializing block 612. Additionally or alternatively, the cross-processor also includes a pre-emphasis stage 709 for performing pre-emphasis processing on the decoded version output by the spectrum decoder 701 prior to the LPC filter 706. The output of the pre-emphasis stage may also be fed to an additional delay stage 710 for initializing the LPC synthesis filter block 616 within the time-domain encoder 610, for initializing the LPC analysis filter block 611.

[0102] As shown in Figure 14a, the time-domain encoder processor 610 includes a pre-emphasis operation at a lower ACELP sampling rate. This pre-emphasis is performed in preprocessing stage 1000, as shown, and is indicated by reference numeral 1005. The pre-emphasis data is input to an LPC analysis filter stage 611 that operates in the time domain, and this filter is controlled by quantization LPC coefficients 1010 obtained through preprocessing stage 1000. The residual signal generated by block 611, as known from AMR-WB+, USAC, or other CELP encoders, is provided to an adaptive codebook 612. Furthermore, the adaptive codebook 612 is connected to an innovation codebook stage 614, and codebook data from the adaptive codebook 612 and the innovation codebook are input to a bitstream multiplexer, as shown.

[0103] In addition, an ACELP gain / coding stage 615 is provided in series with the innovation codebook stage 614, and the result of this block is input into the codebook determiner 613 indicated as MMSE in Figure 14a. This block cooperates with the innovation codebook block 614. Furthermore, the time-domain encoder also includes a decoder section with an LPC synthesis filter block 616, a deemphasis block 617, and an adaptive bass post-filter stage 618 for calculating the parameters of the adaptive bass post-filter; however, the adaptive bass post-filter is applied on the decoder side. Without any adaptive bass post-filter on the decoder side, blocks 616, 617, and 618 are not necessary for the time-domain encoder 610.

[0104] As shown, several blocks of the time-domain decoder depend on the preceding signal, and these blocks are an adaptive codebook block, a codebook determiner 613, an LPC synthesis filter block 616, and a de-emphasis block 617. These blocks are provided with data derived from the cross-processor from the frequency-domain encoder data in order to prepare for the instantaneous switch from the frequency-domain encoder to the time-domain encoder (e.g., ...). Figure 14a-2 These blocks are initialized for the purpose of (as shown in Figure 1450). It can also be seen from Figure 14a that for the frequency domain encoder, any dependency on earlier data is not necessary. Therefore, the cross processor 700 does not provide any memory initialization data from the time domain encoder to the frequency domain encoder. However, for other implementations of the frequency domain encoder where there are dependencies from the past and where memory initialization data is required, the cross processor 700 is configured to operate in both directions.

[0105] Therefore, a preferred embodiment of the audio encoder includes the following parts:

[0106] The preferred audio decoder is described below: The waveform decoder section consists of a full-band TCX decoder path and an IGF, both of which operate at the input sampling rate of the codec. In parallel, an alternative ACELP decoder path exists at a lower sampling rate, which is further enhanced downstream by TD-BWE.

[0107] For ACELP initialization when switching from TCX to ACELP, there exists a cross-path that performs the ACELP initialization of this invention (consisting of a shared TCX decoder front end, but additionally providing output at a lower sampling rate and some post-processing). Sharing the same sampling rate and filtering order between TCX and ACELP in LPC allows for easier and more efficient ACELP initialization.

[0108] To visualize the switching, two switches are plotted in 14b. When the second switch downstream selects between the TCX / IGF or ACELP / TD-BWE output, the first switch either pre-updates the buffer in the resampled QMF stage downstream of the ACELP path via the cross-path output, or simply passes the ACELP output.

[0109] Subsequently, Figures 11a-14c The audio decoder implementation according to aspects of the invention is discussed in the context of this invention.

[0110] An audio decoder for decoding the encoded audio signal 1101 includes a first decoding processor 1120 for decoding a portion of the first encoded audio signal in the frequency domain. The first decoding processor 1120 includes a spectrum decoder 1122 for decoding a first spectral region at high spectral resolution and for synthesizing a second spectral region using a parameter representation of a second spectral region and at least the decoded first spectral region to obtain a decoded spectral representation. The decoded spectral representation is as follows: Figure 6 Discussed in the context of and also as Figure 1a The spectral representation of full-band decoding is discussed in the context of [the previous discussion]. Therefore, in general, the first decoding processor includes a full-band implementation with a gap-filling process in the frequency domain. The first decoding processor 1120 also includes a frequency-to-time converter 1124 for converting the decoded spectral representation to the time domain to obtain the decoded first audio signal portion.

[0111] Furthermore, the audio decoder includes a second decoding processor 1140 for decoding the second encoded audio signal portion in the time domain to obtain the decoded second signal portion. Additionally, the audio decoder includes a combiner 1160 for combining the decoded first signal portion and the decoded second signal portion to obtain the decoded audio signal. The decoded signal portions are combined sequentially, which also... Figure 14b Zhongyou means Figure 11a An embodiment of the combiner 1160 is shown with a switch implementation 1160.

[0112] Preferably, the second decoding processor 1140 is a time-domain bandwidth extension processor, and as follows: Figure 12 The implementation includes a time-domain low-frequency band decoder 1200 for decoding low-frequency band time-domain signals. It also includes an upsampler 1210 for upsampling the low-frequency band time-domain signals. Additionally, a time-domain bandwidth extension decoder 1220 is provided for synthesizing the high-frequency band of the output audio signal. Furthermore, a mixer 1230 is provided for mixing the high-frequency band of the synthesized time-domain output signal and the upsampled low-frequency band time-domain signal to obtain a time-domain encoder output. Therefore, in a preferred embodiment, Figure 11a Block 1140 in the middle can be passed Figure 12 This is achieved through the functionality.

[0113] Figure 13 It shows Figure 12 A preferred embodiment of the time-domain bandwidth-extended decoder 1220 is provided. Preferably, a time-domain upsampler 1221 is provided, which is included within block 1140 and... Figure 12 It is shown at 1200 and in Figure 14bThe time-domain low-band decoder, further illustrated in the context, receives the LPC vestigial signal as input. A time-domain upsampler 1221 generates an upsampled version of the LPC vestigial signal. This version is then fed into a nonlinear distortion block 1222, which generates an output signal with a higher frequency value based on its input signal. The nonlinear distortion can be a copy, mirror, frequency shift, or nonlinear device, such as a diode or transistor operating in a nonlinear region. The output signal of block 1222 is fed into an LPC synthesis filter block 1223, which is also controlled by LPC data for the low-band decoder, or, for example, by specific envelope data generated by the time-domain bandwidth extension block 920 on the encoder side of FIG. 14a. The output of the LPC synthesis block is then fed into a bandpass or high-pass filter 1224 to finally obtain the high-frequency band, and then fed into a mixer 1230, such as... Figure 12 As shown.

[0114] Subsequently, Figure 12 A preferred implementation of the upsampler 1210 is discussed in the context of FIG. 14a. The upsampler preferably comprises an analysis filter bank operating at the sampling rate of the first time-domain low-frequency band decoder. A specific implementation of this analysis filter bank is... Figure 14b The QMF analysis filter bank 1471 is shown. Furthermore, the upsampler includes a synthesis filter bank 1473 that operates at a second output sampling rate higher than the first time-domain low-frequency band sampling rate. Therefore, the synthesis filter bank 1473, preferably implemented as a QMF filter bank, operates at the output sampling rate. When as... Figure 7b When the downsampling factor T discussed in the context is 0.5, the QMF analysis filter bank 1471 has, for example, only 32 filter bank channels, and the synthesis filter bank 1473 has, for example, 64 QMF channels. However, the upper half of the filter bank channels, i.e., the upper 32 filter bank channels, are fed with zero or noise, while the lower 32 filter bank channels are fed with the corresponding signal provided by the QMF analysis filter bank 1471. Preferably, however, bandpass filtering 1472 is performed within the QMF filter bank domain to ensure that the QMF synthesis output is an upsampled version of the ACELP decoder output, but without any artifacts above the maximum frequency of the ACELP decoder.

[0115] As an addition to or replacement of the bandpass filter 1472, further processing operations can be performed within the QMF domain. If no processing is performed at all, QMF analysis and QMF synthesis constitute a highly efficient upsampler 1210.

[0116] Subsequently, Figure 14b The structure of each component will be discussed in more detail.

[0117] The full-band frequency domain decoder 1120 includes a first decoding block 1122a for decoding high-resolution spectral coefficients and for additionally performing noise filling in the low-frequency band portion, for example, as known from USAC technology. Furthermore, the full-band decoder includes an IGF processor 1122b for filling spectral holes using synthesized spectral values ​​that have been encoded parametrically and therefore at low resolution on the encoder side. Then, in block 1122c, inverse noise shaping is performed, and the result is input to a TNS / TTS synthesis block 705, which provides the final output as input to a frequency-to-time converter 1124, preferably implemented as an inverse modified discrete cosine transform operating at the output, i.e., a high sampling rate.

[0118] Furthermore, a harmonic or LTP post-filter (such as) controlled by data obtained from the TCX LTP parameter extraction block 1024 in Figure 14a is used. Figure 14b (As shown in 1420). The result is then the first audio signal portion decoded at the output sampling rate, and as from Figure 14b As can be seen, the data has a high sampling rate, therefore, no further frequency enhancement is needed at all, due to the fact that the decoding processor is a frequency-domain full-band decoder, which is preferably used in... Figures 1a-5c The intelligent gap-filling technique is discussed in the context of [the discussion].

[0119] Figure 14b Several elements in the diagram are very similar to the corresponding blocks in the cross-processor 700 of Figure 14a, particularly the IGF decoder 704 corresponding to the IGF processing 1122b, and the inverse noise shaping operation controlled by the quantization LPC coefficients 1145 corresponding to the inverse noise shaping 703 of Figure 14a. Figure 14b The TNS / TTS synthesis block 705 in the figure corresponds to the TNS / TTS synthesis block 705 in Figure 14a. However, it is important that... Figure 14b IMDCT block 1124 in Figure 14a operates at a high sampling rate, while IMDCT block 702 in Figure 14a operates at a low sampling rate. Therefore, Figure 14b Block 1124 includes a large fixed-size transform and expand block 710, a composition window in block 712, and an overlapping additive stage 714, which operates in block 702 and will be discussed later, compared to the corresponding features 720, 722, and 724, and has a correspondingly larger number of operations, a larger number of window coefficients, and a larger transform size. Figure 14b The cross processor 1170 is outlined in block 1171.

[0120] The time-domain decoding processor 1140 preferably includes an ACELP or time-domain low-frequency band decoder 1200, which includes an ACELP decoder stage 1149 for obtaining decoding gain and innovation codebook information. Additionally, an ACELP adaptive codebook stage 1141 is provided, along with a subsequent ACELP post-processing stage 1142 and a final synthesis filter (e.g., an LPC synthesis filter 1143), which again consists of a codebook corresponding to... Figure 11a The quantization LPC coefficients 1145 obtained by the bitstream multiplexer 1100 of the encoded signal parser 1100 are controlled. The output of the LPC synthesis filter 1143 is input to the deemphasis stage 1144 to cancel or undo the processing introduced by the preemphasis stage 1005 of the preprocessor 1000 of FIG. 14a. The result is a time-domain output signal at a low sampling rate and low frequency band, and when a frequency-domain output is required, switch 1480 is in the indicated position, and the output of the deemphasis stage 1144 is introduced into the upsampler 1210 and then mixed with the high-frequency band from the time-domain bandwidth-extended decoder 1220.

[0121] According to an embodiment of the invention, the audio decoder further includes Figure 11b and Figure 14b The cross processor 1170 shown is configured to calculate initialization data for the second decoding processor based on the spectral representation of the decoded first coded audio signal portion, such that the second decoding processor is initialized to decode the second audio signal portion of the encoded audio signal that follows the encoded first audio signal portion in time, i.e., to prepare the time-domain encoding processor 1140 for instantaneous switching from one audio signal portion to the next audio signal portion without any loss in quality or efficiency.

[0122] Preferably, the cross processor 1170 includes an additional frequency-time converter 1171 operating at a lower sampling rate than the frequency-time converter of the first decoding processor, in order to obtain a first signal portion for further decoding in the time domain, which can be used as an initialization signal or from which any initialization data can be derived. Preferably, this IMDCT or low sampling rate frequency-time converter is implemented as... Figure 7b The items 726 (selector), 720 (small-size transform and expand), the synthesis window with a small number of window coefficients as shown in 722, and the overlapping addition stage with a small number of operations as shown in 724. Therefore, the IMDCT block 1124 in the frequency domain full-band decoder is implemented as shown by blocks 710, 712, and 714, and the IMDCT block 1171 is as shown in... Figure 7bThe implementation shown is provided by blocks 726, 720, 722, and 724. Again, the downsampling factor is the ratio between the time-domain encoder sampling rate or low sampling rate and the higher frequency-domain sampling rate or output sampling rate, and this downsampling factor is less than 1 and can be any number greater than 0 and less than 1.

[0123] like Figure 14b As shown, the cross-processor 1170, either alone or among other components, includes a delay stage 1172 for delaying a first signal portion to be further decoded and for feeding the delayed decoded first signal portion to the deemphasis stage 1144 of the second decoding processor for initialization. Additionally, the cross-processor may also include a pre-emphasis filter 1173 and a delay stage 1175 for filtering and delaying the first signal portion to be further decoded, and for providing the delayed output of block 1175 to the LPC synthesis filter stage 1143 of the ACELP decoder for initialization purposes.

[0124] Furthermore, the cross-processor may alternatively include, or in addition to the other mentioned elements, an LPC analysis filter 1174 for generating a predictive residual signal based on a first signal portion that is further decoded or a pre-emphasized first signal portion that is further decoded, and for feeding data into the codebook synthesizer of the second decoding processor, and preferably into the adaptive codebook stage 1141. Additionally, the output of the frequency-to-time converter 1171 with a low sampling rate is also input to the QMF analysis filter bank 1471 of the upsampler 1210 for initialization purposes, i.e., when the currently decoded audio signal portion is delivered by the frequency domain full-band decoder 1120.

[0125] The preferred audio decoder is described below: The waveform decoder section consists of a full-band TCX decoder path and an IGF, both of which operate at the input sampling rate of the codec. In parallel, an alternative ACELP decoder path exists at a lower sampling rate, which is further enhanced downstream by TD-BWE.

[0126] For ACELP initialization when switching from TCX to ACELP, there exists a cross-path that performs the ACELP initialization of this invention (consisting of a shared TCX decoder front end, but additionally providing output at a lower sampling rate and some post-processing). Sharing the same sampling rate and filtering order between TCX and ACELP in LPC allows for easier and more efficient ACELP initialization.

[0127] For visual switching, in Figure 14bTwo switches are shown in the diagram. When the second switch downstream selects between the TCX / IGF or ACELP / TD-BWE output, the first switch either pre-updates the buffer in the resampled QMF stage downstream of the ACELP path via the cross path output, or simply passes the ACELP output.

[0128] In summary, preferred aspects of the invention, which can be used alone or in combination, involve the combination of ACELP and TD-BWE encoders with full-band TCX / IGF technology, preferably associated with the use of cross signals.

[0129] Another specific feature is the cross signal path used for ACELP initialization to enable seamless switching.

[0130] On the other hand, the lower portion of the short IMDCT is fed with high-rate long MDCT coefficients to achieve efficient sampling rate conversion in the cross path.

[0131] Another feature is the efficient implementation of the cross-path shared with the full-band TCX / IGF section in the decoder.

[0132] Another feature is the cross signal path used for QMF initialization to enable seamless switching from TCX to ACELP.

[0133] An additional feature is the cross signal path to the QMF, which allows compensation for the delay gap between the ACELP resampled output and the filter bank-TCX / IGF output when switching from ACELP to TCX.

[0134] On the other hand, LPC is provided for both the TCX and ACELP encoders with the same sampling rate and filtering order, even though the TCX / IGF encoder / decoder is capable of full-bandwidth operation.

[0135] Subsequently, Figure 14c It is discussed as a preferred implementation of either operating as a standalone decoder or operating in combination with a time-domain decoder capable of operating in the full-band frequency domain.

[0136] Typically, the time-domain decoder includes an ACELP decoder 1500, followed by a resampler or upsampler and a time-domain bandwidth extension function. Specifically, the ACELP decoder includes an ACELP decoding stage 1149 for gain recovery and codebook innovation, an ACELP adaptive codebook stage 1141, an ACELP post-processor 1142, an LPC synthesis filter 1143 or an encoded signal resolver controlled by quantized LPC coefficients from a bitstream multiplexer, and a subsequently connected de-emphasis stage 1144. Preferably, the time-domain residual signal at the ACELP sampling rate is input to a time-domain bandwidth extension decoder 1220, which provides a high-frequency band at the output.

[0137] To upsample the deemphasized 1144 output, an upsampler is provided comprising a QMF analysis filter bank 1471 and a synthesis filter bank 1473 implemented as a QMF filter bank. A bandpass filter is preferably applied within the filter bank domain defined by the AMR analysis filter bank 1471 and the synthesis filter bank 1473. In particular, the same functionality, as discussed previously and with reference to the same reference numerals, can also be used. Furthermore, the time-domain bandwidth-extended decoder 1220 can be as follows... Figure 13 The implementation is shown. It typically involves upsampling the ACELP residual signal or time-domain residual signal at the ACELP sampling rate, which ultimately becomes the output sampling rate of the bandwidth-extended signal.

[0138] Subsequently, regarding Figures 1a-5c Further details are discussed regarding frequency domain encoders and decoders capable of operating across the entire frequency band.

[0139] Figure 1a An apparatus for encoding an audio signal 99 is shown. The audio signal 99 is input to a time-to-spectrum converter 100, which converts the audio signal with a sampling rate into a spectral representation 101 output by the time-to-spectrum converter. The spectrum 101 is input to a spectrum analyzer 102 for analyzing the spectral representation 101. The spectrum analyzer 101 is configured to determine a first set of first spectral portions 103 to be encoded at a first spectral resolution and a different second set of second spectral portions 105 to be encoded at a second spectral resolution. The second spectral resolution is smaller than the first spectral resolution. The second set of second spectral portions 105 is input to a parameter calculator or parameter encoder 104 for calculating spectral envelope information with the second spectral resolution. Furthermore, a spectral domain audio encoder 106 is provided for generating a first encoded representation 107 of the first set of first spectral portions with the first spectral resolution. Furthermore, the parameter calculator / parameter encoder 104 is configured to generate a second encoded representation 109 of the second set of second spectral portions. The first encoding representation 107 and the second encoding representation 109 are input into a bitstream multiplexer or bitstream shaper 108, and block 108 ultimately outputs the encoded audio signal for transmission or storage on a storage device.

[0140] Typically, the first spectral portion (e.g.) Figure 3a 306 will be surrounded by two second spectral sections (such as 307a and 307b). This is not the case in HE AAC, where the core encoder frequency range is band-limited.

[0141] Figure 1b It shows the relationship with Figure 1aThe encoder is matched with a decoder. A first encoded representation 107 is input to a spectral domain audio decoder 112 to generate a first decoded representation of a first set of first spectral portions having a first spectral resolution. Furthermore, a second encoded representation 109 is input to a parameter decoder 114 to generate a second decoded representation of a second set of second spectral portions having a second spectral resolution lower than the first spectral resolution.

[0142] The decoder also includes a frequency regenerator 116 for regenerating a reconstructed second spectral portion having a first spectral resolution using a first spectral portion. The frequency regenerator 116 performs a tile-filling operation, i.e., using tiles or portions of the first set of the first spectral portions and copying that first set of the first spectral portions into a reconstructed range or reconstructed band having the second spectral portion, and typically performs spectral envelope shaping or another operation indicated by the decoded second representation output by the parametric decoder 114 (i.e., using information about the second set of the second spectral portions). The decoded first set of the first spectral portions and the reconstructed second set of the spectral portions are input to a spectrum-to-time converter 118, as indicated at the output of the frequency regenerator 116 on line 117, which is configured to convert the first decoded representation and the reconstructed second spectral portion into a time representation 119 having a certain high sampling rate.

[0143] Figure 2b It shows Figure 1a Encoder implementation. Audio input signal 99 is input to the corresponding... Figure 1a The time-to-spectrum converter 100 is analyzed in the filter bank 220. Then, a time noise shaping operation is performed in the TNS block 222. Therefore, to the corresponding Figure 2b 226 block tone mask Figure 1a The input to the spectrum analyzer 102 can be the full spectrum value when no time noise shaping / time patch shaping operation is applied, or when an operation such as... Figure 2b The TNS operation shown in block 222 can be a spectral residual value. For dual-channel or multi-channel signals, joint channel coding 228 can be performed separately to make... Figure 1a The spectral domain encoder 106 may include a joint channel coding block 228. Furthermore, an entropy encoder 232 is provided for performing lossless data compression, which is also... Figure 1a It is part of the spectrum domain encoder 106.

[0144] The spectrum analyzer / tone mask 226 separates the output of the TNS block 222 into the core frequency band and the tone components corresponding to the first group of the first spectrum portion 103 and the corresponding... Figure 1a The residual components of the second spectral portion 105 of the second group. Block 224, indicated as the IGF parameter extraction encoding, corresponds to... Figure 1a The parameter encoder 104, and the bitstream multiplexer 230 corresponding to Figure 1a 108 bitstream multiplexer.

[0145] Preferably, the analysis filter bank 222 is implemented as an MDCT (Modified Discrete Cosine Transform Filter Bank), and the MDCT is used to transform the signal 99 into the time-frequency domain using a modified discrete cosine transform as a frequency analysis tool.

[0146] The spectrum analyzer 226 preferably applies a tone mask. This tone mask estimation stage is used to separate tone components from noise-like components in the signal. This allows the core encoder 228 to encode all tone components using a psychoacoustic module. The tone mask estimation stage can be implemented in many different ways and is preferably functionally similar to the sinusoidal track estimation stage used in sine and noise modeling for speech / audio coding [8,9] or in the HILN-based audio encoder described in

[10] . Preferably, an implementation that is easy to implement and does not require maintaining life and death tracks is used, but any other tone or noise detector may also be used.

[0147] The IGF module calculates the similarity between the source and target regions. The target region is represented by the spectrum from the source region. The similarity between the source and target regions is measured using a cross-correlation method. The target region is divided into... Non-overlapping frequency tiles. For each tile in the target region, create tiles starting from a fixed frequency. Source tiles. These source tiles have an overlap factor between 0 and 1, where 0 means 0% overlap and 1 means 100% overlap. Each of these source tiles is associated with a target tile at various lags to find the source tile that best matches the target tile. The best-matching tile number is stored in [the database / source]. In this context, the lag that best relates it to the objective is stored. In the middle, and the symbols of correlation are stored In cases of very negative correlation, the source piece needs to be multiplied by -1 before piece padding at the decoder. The IGF module also considers not overwriting the tone components in the spectrum, as tone masks are used to preserve tone components. The bandgap energy parameter is used to store the energy of the target region, allowing us to accurately reconstruct the spectrum.

[0148] This approach has some advantages over traditional SBR [1]: the harmonic grid of the multi-tone signal is preserved by the core encoder, while only the gaps between the sine waves are filled with best-matched “shaping noise” from the source region. Another advantage of this system compared to ASR (Accurate Spectral Replacement) [2-4] is the absence of a signal synthesis stage, which creates important parts of the signal at the decoder. Instead, this task is taken over by the core encoder, making it possible to preserve important components of the spectrum. Another advantage of the proposed system is the continuous scalability provided by the features. Only one piece needs to be used. and This is known as granular matching and can be used at low bit rates, while using variables for each piece. This allows us to better match the target and source spectrum.

[0149] In addition, a piece selection stabilization technique for removing frequency domain artifacts such as jitter and musical noise is proposed.

[0150] In the case of stereo channel pairs, additional joint stereo processing is applied. This is necessary because, for a given destination range, the signal can be a highly correlated panned source. If the source regions selected for that particular region are not well correlated, the spatial picture may be compromised due to the uncorrelated source regions, even though the energy matches the destination region. The encoder analyzes the band structure for each destination region, typically performing cross-correlation of the spectral values, and sets a joint flag for that band if a certain threshold is exceeded. In the decoder, if the joint stereo flag is not set, the left and right channel band structures are processed separately. If the joint stereo flag is set, energy and patching are performed in the joint stereo domain. Similar to the joint stereo information used for core encoding, joint stereo information for the IGF region is signaled, including, in the case of prediction, indicating whether the prediction direction is from undermix to residual, or vice versa.

[0151] The energy can be calculated based on the transmitted energy in the L / R domain.

[0152]

[0153]

[0154] in It is the frequency index in the transform domain.

[0155] Another solution is to calculate and transmit energy directly in the joint stereo domain, which is the active frequency band for joint stereo, so no additional energy conversion is needed on the decoder side.

[0156] The source tile is always created based on the center / side matrix:

[0157]

[0158]

[0159] Energy Adjustment:

[0160]

[0161]

[0162] Combined stereo -> LR transformation:

[0163] If the additional prediction parameters are not encoded:

[0164]

[0165]

[0166] If the additional prediction parameters are encoded and if the signaling direction is from the middle to one side:

[0167]

[0168] If the signal is sent from one side to the middle:

[0169]

[0170] This process ensures that, based on the tiles used to regenerate highly correlated target regions and translations of target regions, the resulting left and right channels still represent correlated and translated sound sources, even if the source regions are uncorrelated, thus preserving the stereo image for such regions.

[0171] In other words, in the bitstream, a joint stereo flag indicating whether L / R or M / S should be used as an example of general joint stereo coding is transmitted. In the decoder, firstly, the core signal is decoded as indicated by the joint stereo flag for the core band. Secondly, the core signal is stored in both L / R and M / S representations. For IGF patch filling, the source patch representation is selected to suit the target patch representation as indicated by the joint stereo information for the IGF band.

[0172] Temporal noise shaping (TNS) is a standard technique and is part of AAC [11-13]. TNS can be considered an extension of the basic scheme of a perceptual encoder, inserting an optional processing step between the filter bank and the quantization stage. The main task of the TNS module is to hide the quantization noise generated in the time-masked region of transient similar signals, and thus it leads to a more efficient coding scheme. First, TNS uses "forward prediction" (e.g., MDCT) in the transform domain to compute a set of prediction coefficients. These coefficients are then used to flatten the temporal envelope of the signal. Since quantization affects the spectrum after TNS filtering, the quantization noise is also temporarily flattened. By applying inverse TNS filtering on the decoder side, the quantization noise is shaped according to the time envelope of the TNS filter, and thus the quantization noise is transiently masked.

[0173] IGF is based on MDCT representation. For efficient encoding, it is preferable to use long blocks of approximately 20 ms. If the signal within such a long block contains transients, audible front and back echoes occur in the IGF spectral bands due to block padding.

[0174] This pre-echo effect is reduced by using TNS in the IGF context. Here, TNS is used as a Time Patch Shaping (TTS) tool because spectral regeneration is performed in the decoder on the TNS residual signal. The required TTS prediction coefficients are calculated and applied using the full spectrum on the encoder side as usual. The TNS / TTS start and stop frequencies are independent of the IGF start frequency of the IGF tool. Impact. Compared to traditional TNS, the TTS stopping frequency increases to the stopping frequency of the IGF tool, which is higher. On the decoder side, the TNS / TTS coefficients are again applied to the full spectrum, i.e., the core spectrum plus the regenerated spectrum plus the tonal component from the tone mask. The application of TTS is necessary to form the temporal envelope of the regenerated spectrum to match the envelope of the original signal again. Therefore, the pre-echo is reduced. Furthermore, it still maintains its TNS at a level lower than... The quantization noise is shaped in the signal.

[0175] In traditional decoders, spectral patching on the audio signal disrupts the spectral correlation at the patch boundaries, thereby impairing the temporal envelope of the audio signal by introducing dispersion. Therefore, another benefit of performing IGF patch filling on the residual signal is that, after applying shaping filtering, the patch boundaries are seamlessly correlated, resulting in more faithful temporal reproduction of the signal.

[0176] In the encoder of this invention, the spectrum that has undergone TNS / TTS filtering, tone masking, and IGF parameter estimation contains no signals above the IGF start frequency except for the tone components. This sparse spectrum is now encoded by the core encoder using the principles of arithmetic coding and predictive coding. These encoded components, along with signaling bits, form the audio bitstream.

[0177] Figure 2a The corresponding decoder implementation is shown. This corresponds to the encoded audio signal. Figure 2a The bitstream is input to the multiplexer / decoder 200, which will... Figure 1b Connect to blocks 112 and 114. The bitstream multiplexer separates the input audio signal into... Figure 1b The first code represents 107 and Figure 1b The second coded representation is 109. The first coded representation, having a first set of first spectral portions, is input to the corresponding... Figure 1b The second encoded representation is input into the joint channel decoding block 204 of the spectrum domain decoder 112. Figure 2a The parameter decoder 114 (not shown) is then input into the corresponding... Figure 1b The frequency regenerator 116 is located in IGF block 202. The first set of first spectral portions required for frequency regeneration is input to IGF block 202 via line 203. Furthermore, after joint channel decoding 204, specific core decoding is applied in tone mask block 206 such that the output of tone mask 206 corresponds to the output of spectrum domain decoder 112. Combining, i.e., frame construction, is then performed by combiner 208, where the output of combiner 208 now has the full-range spectrum but remains in the TNS / TTS filtered domain. Then, in block 210, inverse TNS / TTS operation is performed using TNS / TTS filtered information provided via line 109; i.e., TTS auxiliary information is preferably included in a first coded representation generated by spectrum domain encoder 106 (e.g., spectrum domain encoder 106 may be a direct AAC or USAC core encoder), or may also be included in a second coded representation. At the output of block 210, a complete spectrum up to the maximum frequency is provided, which is the full-range frequency defined by the sampling rate of the original input signal. Then, a spectrum / time conversion is performed in the synthesis filter bank 212 to finally obtain the audio output signal.

[0178] Figure 3a A schematic representation of the spectrum is shown. The spectrum is subdivided by the scaling factor band SCB, where in Figure 3a The example shown contains seven scaling factor bands, SCB1 to SCB7. Scaling factor bands can be AAC scaling factor bands as defined in the AAC standard, and have increased bandwidth for the upper frequencies, such as... Figure 3aAs illustrated schematically. Preferably, instead of performing intelligent gap filling at low frequencies from the very beginning of the spectrum, IGF operation begins at the IGF start frequency shown at 309. Therefore, the core band extends from the lowest frequency to the IGF start frequency. Above the IGF start frequency, spectral analysis is applied to separate high-resolution spectral components 304, 305, 306, and 307 (first group, first spectral portion) from the low-resolution components represented by the second group, second spectral portion. Figure 3a The diagram illustrates the spectrum input to either the spectral domain encoder 106 or the combined channel encoder 228. Specifically, the core encoder operates across the full range but encodes a large number of zero-spectral values, i.e., these zero-spectral values ​​are quantized to zero or set to zero before or after quantization. Regardless, the core encoder operates across the full range, i.e., as the spectrum will be shown, meaning the core decoder does not necessarily need to know the second set of second frequencies with lower spectral resolution.

[0179] Any intelligent gap filling or encoding in the spectral portion.

[0180] Preferably, the high resolution is defined by linear encoding of spectral lines, such as MDCT lines, while the second resolution or low resolution is defined, for example, by calculating only a single spectral value for each scaling factor band, where the scaling factor band covers several frequency lines. Therefore, the second low resolution is much lower in spectral resolution than the first or high resolution defined by linear encoding typically applied by the core encoder (e.g., an AAC or USAC core encoder).

[0181] Regarding scaling factors or energy calculations, the situation is as follows: Figure 3b As shown in the diagram. Due to the fact that the encoder is the core encoder and because components of the first set of spectral parts in each band may, but are not necessarily, exist, the core encoder extends not only in the core range below the IGF starting frequency of 309, but also above the IGF starting frequency up to the maximum frequency. Calculate the scaling factor for each frequency band, where the maximum frequency is less than or equal to half the sampling frequency, i.e., f. s / 2 .therefore, Figure 3a The encoded tone portions 302, 304, 305, 306, and 307, along with scaling factors SCB1 to SCB7 in this embodiment, correspond to high-resolution spectral data. Low-resolution spectral data is calculated from the IGF starting frequency and corresponds to energy information values ​​E1, E2, E3, and E4, which are transmitted along with scaling factors SF4 to SF7.

[0182] Specifically, when the core encoder is operating at low bit rates, additional noise padding can be applied to the core band (i.e., a frequency lower than the IGF start frequency, specifically within the scaling factor bands SCB1 to SCB3). In this noise padding, several adjacent spectral lines have been quantized to zero. On the decoder side, these quantized-to-zero spectral values ​​are resynthesized and processed using methods such as... Figure 3b The noise-fill energy of NF2, shown at point 308, is used to adjust the resynthesized spectral values ​​in terms of their amplitude. The noise-fill energy, which can be given in absolute terms or in relative terms, particularly with respect to a scaling factor as in USAC, corresponds to the energy of this set of spectral values ​​quantized to zero. These noise-filled spectra can also be considered as a third set of third spectral portions, which are regenerated by direct noise-fill synthesis without relying on any IGF operation using frequency regeneration from frequency patches from other frequencies, which is used to reconstruct spectral patches using spectral values ​​and energy information E1, E2, E3, E4 from the source range.

[0183] Preferably, the frequency band for which the energy information is calculated coincides with the scaling factor frequency band. In other embodiments, energy information values ​​are grouped such that, for example, for scaling factor bands 4 and 5, only a single energy information value is transmitted; however, even in this embodiment, the boundary of the reconstructed frequency band of the group coincides with the boundary of the scaling factor frequency band. If different frequency band spacings are applied, some recalculation or synchronization calculation can be applied, and this may be meaningful depending on the specific implementation.

[0184] Preferably, Figure 1a The spectrum domain encoder 106 is as follows Figure 4a The image shows a psychoacoustic-driven encoder. Typically, as shown in, for example, the MPEG2 / 4 AAC standard or the MPEG1 / 2, Layer 3 standard, the audio signal to be encoded after being converted to a spectral range ( Figure 4aThe signal (401) is forwarded to a scaling factor calculator 400. The scaling factor calculator is controlled by a psychoacoustic model 402, which in turn receives the audio signal to be quantized, or a complex spectral representation of the audio signal as in MPEG1 / 2 Layer 3 or MPEG AAC standards. The psychoacoustic model 402 calculates a scaling factor representing a psychoacoustic threshold for each scaling factor band. Furthermore, the scaling factor is then adjusted through the cooperation of known internal and external iterative loops or through any other suitable encoding process to satisfy certain bit rate conditions. The spectral values ​​to be quantized and the calculated scaling factors are then input to a quantizer processor 404. In direct audio encoder operation, the spectral values ​​to be quantized are weighted by the scaling factor, and the weighted spectral values ​​are then input to a fixed quantizer, typically with compression capabilities up to the upper amplitude range. A quantization index is then present at the output of the quantizer processor and forwarded to an entropy encoder, which typically has specific and highly efficient encoding for a set of zero quantization indices (or, as referred to in the art, "extensions" of zero values) of adjacent frequency values.

[0185] However, in Figure 1a In an audio encoder, the quantizer processor typically receives information about a second spectral portion from a spectrum analyzer. Therefore, the quantizer processor 404 ensures that the second spectral portion in the output of the quantizer processor 404, such as that identified by the spectrum analyzer 102 as zero or having a representation confirmed as zero by the encoder or decoder, can be encoded very efficiently, especially when there is an "extension" of zero values ​​in the spectrum.

[0186] Figure 4b An implementation of the quantizer processor is shown. MDCT spectral values ​​can be input into the zero-setting block 410. Then, before weighting by the scaling factor in block 412, the second spectral portion has been set to zero. In an additional implementation, block 410 is not provided; instead, the zero-setting operation is performed in block 418 after weighting block 412. In even a further implementation, the zero-setting operation can also be performed in zero-setting block 422 after quantization in quantizer block 420. In this implementation, blocks 410 and 418 will not exist. Typically, at least one of blocks 410, 418, and 422 is provided depending on the specific implementation.

[0187] Then, at the output of block 422, the corresponding value is obtained. Figure 3a The quantized spectrum of the content shown is then input to a device such as... Figure 2b In an entropy encoder such as 232, it can be, for example, a Huffman encoder or an arithmetic encoder as defined in the USAC standard.

[0188] The zero-set blocks 410, 418, and 422, provided alternately or in parallel, are controlled by the spectrum analyzer 424. The spectrum analyzer preferably includes any implementation of a known tone detector, or any different kind of detector operable to separate the spectrum into components to be encoded at high resolution and components to be encoded at low resolution. Other such algorithms implemented in the spectrum analyzer may be speech activity detectors, noise detectors, speech detectors, or any other detector, depending on the spectral information or associated metadata regarding the resolution requirements of different spectral portions.

[0189] Figure 5a This is illustrated, for example, in AAC or USAC. Figure 1a A preferred implementation of the time-to-spectrum converter 100 is described. The time-to-spectrum converter 100 includes a windower 502 controlled by a transient detector 504 or the transient detector 1020 of FIG. 14a. When the transient detector 504 detects a transient, a switch from a long window to a short window is signaled to the windower. The windower 502 then calculates windowed frames for the overlapping blocks, where each windowed frame typically has two N values, for example, 2048 values. A transformation is then performed within a block transformer 506, which typically additionally provides decimation, such that a combined decimation / transform is performed to obtain a spectral frame with N values ​​(e.g., MDCT spectral values). Thus, for long window operation, the frame at the input of block 506 includes two N values, for example, 2048 values, while the spectral frame has 1024 values. Then, however, when executing eight short blocks, a switching is performed on the short blocks, where each short block has 1 / 8 of the windowed time-domain value compared to the long window, and each spectral block has 1 / 8 of the spectral value compared to the long block. Therefore, when this decimation is combined with the 50% overlap operation of the windower, the spectrum is a critically sampled version of 99% of the time-domain audio signal.

[0190] Subsequently, reference Figure 5b It shows Figure 1b The specific implementation of the frequency regenerator 116 and the spectrum-to-time converter 118, or Figure 2a The specific implementation of the combined operation of blocks 208 and 212. Figure 5b In the middle, considering a specific reconstruction frequency band, for example Figure 3a The scaling factor is band 6. The first spectral portion of this reconstructed band, namely... Figure 3a The first spectral portion 306 is input into the frame builder / adjuster block 510. Furthermore, the second spectral portion reconstructed for the scaling factor band 6 is also input into the frame builder / adjuster 510. Additionally, energy information (such as that used for the scaling factor band 6) is also input. Figure 3bThe E3) is also input into block 510. The reconstructed second spectral portion in the reconstructed band has been generated using the source range through frequency patching, and the reconstructed band then corresponds to the target range. Now, frame energy adjustment is performed so that the final result is obtained, for example, in... Figure 2a The complete reconstructed frame with N values ​​is obtained at the output of combiner 208. Then, in block 512, an inverse block transform / interpolation is performed to obtain 248 temporal values ​​for, for example, 124 spectral values ​​at the input of block 512. Then, in block 514, a synthesis windowing operation is performed, which is again controlled by a long window / short window indication sent as auxiliary information in the encoded audio signal. Then, in block 516, an overlap / addition operation with the previous time frame is performed. Preferably, MDCT applies 50% overlap, such that for each new time frame with 2N values, N temporal values ​​are ultimately output. The 50% overlap is particularly preferred because it provides key sampling and continuous crossover from one frame to the next due to the overlap / addition operation in block 516.

[0191] like Figure 3a As shown at position 301, for example, for... Figure 3a The expected reconstructed frequency band, consistent with the scaling factor band 6, can be further filled with noise not only below the IGF start frequency but also above the IGF start frequency. The noise fill spectrum value can then be input into the frame builder / adjuster 510, and adjustments to the noise fill spectrum value can be applied within this block, or the noise fill spectrum value can be adjusted using noise fill energy before being input into the frame builder / adjuster 510.

[0192] Preferably, the IGF operation can be applied across the entire spectrum, i.e., a frequency patching operation using spectral values ​​from other parts. Therefore, the frequency patching operation can be applied not only to the high-frequency band above the IGF start frequency but also to the low-frequency band. Furthermore, noise filling without frequency patching can be applied not only below the IGF start frequency but also above it. However, it has been found that high-quality and high-efficiency audio coding can be obtained when the noise filling operation is limited to a frequency range below the IGF start frequency and when the frequency patching operation is limited to a frequency range above the IGF start frequency, such as... Figure 3a As shown.

[0193] Preferably, the target piece (TT) (with a frequency greater than the IGF start frequency) is bound to the scaling factor band boundary of the full-rate encoder. The source piece (ST) from which information is obtained (i.e., for frequencies below the IGF start frequency) is not bound to the scaling factor band boundary. The size of the ST should correspond to the size of the associated TT. This is illustrated using the following example. TT[0] has a length of 10 MDCT cells. This corresponds exactly to the length of two subsequent SCBs (e.g., 4+6). Then, all possible STs associated with TT[0] also have a length of 10 cells. The second target piece TT[1] adjacent to TT[0] has a length of 15 cells (the SCBs have a length of 7+8). Then, the ST for it has a length of 15 cells instead of the 10 cells for TT[0].

[0194] If a TT with the length of the target piece cannot be found (when, for example, the length of the TT is greater than the available source range), the correlation is not calculated, and the source range is copied to the TT multiple times (one copy at a time, such that the lowest frequency line of the second copy follows the highest frequency line used for the first copy in terms of frequency) until the target piece TT is completely filled.

[0195] Subsequently, reference Figure 5c It shows Figure 1b Frequency regenerator 116 or Figure 2a Another preferred embodiment of IGF block 202. Block 522 is a frequency patch generator that receives not only the target frequency band ID but also the source frequency band ID. Exemplarily, the source frequency band ID has already been determined on the encoder side. Figure 3a The scaling factor band is well-suited for reconstructing the scaling factor band 7. Therefore, the source band ID will be 2, and the target band ID will be 7. Based on this information, the frequency patch generator 522 applies an up-copy or harmonic patch filling operation, or any other patch filling operation, to generate the original second portion of the spectral component 523. The original second portion of the spectral component has the same frequency resolution as that included in the first set of first spectral portions.

[0196] Then, the first spectral portion of the reconstructed frequency band (e.g.) Figure 3aThe 307) is input into the frame builder 524, and the original second part 523 is also input into the frame builder 524. Then, the adjuster 526 adjusts the reconstructed frame using the gain factor of the reconstructed frequency band calculated by the gain factor calculator 528. However, it is important that the first spectral portion of the frame is not affected by the adjuster 526, but only the original second part of the reconstructed frame is affected by the adjuster 526. For this purpose, the gain factor calculator 528 analyzes the source frequency band or the original second part 523, and additionally analyzes the first spectral portion in the reconstructed frequency band, to finally find the correct gain factor 527, such that the energy output of the frame adjusted by the adjuster 526 has energy E4 when the scale factor frequency band 7 is assumed.

[0197] In this context, evaluating the high-frequency reconstruction accuracy of the present invention compared to HE-AAC is of great importance. This is about Figure 3a This can be explained using the scaling factor band 7. Assume a prior art encoder detects a spectral portion 307 to be encoded as a "lost harmonic" at high resolution. The energy of this spectral component is then sent to the decoder along with spectral envelope information (e.g., scaling factor band 7) used to reconstruct the band. The decoder then recreates the lost harmonic. However, the spectral value of the lost harmonic 307 reconstructed by the prior art decoder will be in the middle of band 7 at the frequency indicated by the reconstructed frequency 390. Therefore, the present invention avoids the frequency error 391 introduced by the prior art decoder.

[0198] In one implementation, the spectrum analyzer is further implemented to calculate the similarity between a first spectral portion and a second spectral portion, and based on the calculated similarity, to determine a first spectral portion in the reconstructed range that matches the second spectral portion as closely as possible to the second spectral portion. Then, in this variable source range / destination range implementation, the parametric encoder further incorporates matching information into a second coded representation, which indicates the matching source range for each destination range. On the decoder side, this information is then... Figure 5c The frequency tile generator 522 is used. Figure 5c The generation of the original second part 523 based on the source band ID and the target band ID is shown.

[0199] In addition, such as Figure 3a As shown, the spectrum analyzer is configured to analyze the spectrum representation up to a maximum analysis frequency, which is only a small amount below half the sampling frequency, and preferably at least a quarter or usually higher than the sampling frequency.

[0200] As shown, the encoder operates without downsampling, and the decoder operates without upsampling. In other words, the spectral domain audio encoder is configured to produce a spectral representation with Nyquist frequencies defined by the sampling rate of the initial input audio signal.

[0201] In addition, such as Figure 3a As shown, the spectrum analyzer is configured to analyze a spectrum representation that begins with the gap fill start frequency and ends with the maximum frequency represented by the maximum frequency included in the spectrum representation, wherein the spectrum portion extending from the minimum frequency to the gap fill start frequency belongs to a first group of spectrum portions, and another spectrum portion (such as 304, 305, 306, 307) having a frequency value higher than the gap fill frequency is additionally included in the first group of first spectrum portions.

[0202] As outlined, the spectral domain audio decoder 112 is configured such that the maximum frequency represented by the spectral values ​​in the first decoded representation is equal to the maximum frequency included in the time representation having a sampling rate, wherein the spectral value for the maximum frequency is zero or different from zero in the first set of first spectral components. In any case, for this maximum frequency in the first set of spectral components, there exists a scaling factor for the scaling factor band, which is generated and transmitted regardless of whether all spectral values ​​in that scaling factor band are set to zero, as... Figure 3a and 3b This is discussed in the context of [the previous sentence].

[0203] Therefore, this invention is advantageous for other parametric techniques that increase compression efficiency, such as noise substitution and noise filling (these techniques are specifically designed for the efficient representation of noise in local signal content), and allows for accurate frequency reproduction of tone components. To date, no prior art has addressed the efficient parametric representation of arbitrary signal content through spectral gap filling without the constraint of fixed a priori segmentation in the low-frequency (LF) and high-frequency (HF) bands.

[0204] The embodiments of the present invention improve upon the methods of the prior art, thereby providing high compression efficiency with little or no perceptible disturbance and full audio bandwidth even at low bit rates.

[0205] A typical system includes:

[0206] • Full-band core coding

[0207] • Intelligent gap filling (patch filling or noise filling)

[0208] • Sparse tone components in the core selected via tone masking

[0209] • Full-band joint stereo pair encoding, including patch filling

[0210] • TNS on the puzzle pieces

[0211] • Spectral whitening within the IGF range

[0212] The first step toward a more efficient system is to eliminate the need to transform spectral data into a second transform domain different from one of the core encoders. Since most audio codecs (such as AAC, for example) use MDCT as the basic transform, performing BWE in the MDCT domain is also useful. A second requirement for a BWE system would be the need to preserve the pitch grid, thereby preserving even the HF pitch components, and thus resulting in higher quality encoded audio than existing systems. To address these two requirements of the BWE scheme, a new system called Intelligent Gap Fill (IGF) has been proposed. Figure 2b A block diagram of the proposed system on the encoder side is shown, and Figure 2a The system on the decoder side is shown.

[0213] Subsequently, additional optional features of the full-band frequency domain first encoding processor and the full-band frequency domain decoding processor incorporating gap-filling operations, which can be implemented separately or together, are discussed and defined.

[0214] Specifically, the spectrum domain decoder 112 corresponding to block 1122a is configured to output a sequence of decoded frame values ​​for the spectrum, the decoded frame being a first decoded representation, wherein the frame includes spectrum values ​​for a first set of spectrum portions and zero indications for a second set of spectrum portions. The decoding apparatus also includes a combiner 208. The spectrum values ​​are generated by a frequency regenerator for a second set of second spectrum portions, wherein both the combiner and the frequency regenerator are included within block 1122b. Therefore, by combining the second spectrum portion and the first spectrum portion, a reconstructed spectrum frame including the spectrum values ​​of the first set of first spectrum portions and the second set of spectrum portions is obtained, and corresponding to... Figure 14b The spectrum-to-time converter 118 of the IMDCT block 1124 then converts the reconstructed spectrum frame into a time representation.

[0215] As outlined, the spectrum-to-time converter 118 or 1124 is configured to perform an inversely modified discrete cosine transform 512, 514, and also includes an overlap-addition stage 516 for overlapping and adding subsequent time-domain frames.

[0216] Specifically, the spectrum domain audio decoder 1122a is configured to generate a first decoded representation such that the first decoded representation has a Nyquist frequency defined as equal to the sampling rate of the time representation generated by the spectrum-time converter 1124.

[0217] Furthermore, decoder 1112 or 1122a is configured to generate a first decoded representation such that the first spectral portion 306 is positioned with respect to the frequencies between the two second spectral portions 307a, 307b.

[0218] In another embodiment, the maximum frequency represented by the spectral value of the maximum frequency in the first decoded representation is equal to the maximum frequency included in the time representation generated by the spectrum-time converter, wherein the spectral value of the maximum frequency in the first representation is zero or different from zero.

[0219] In addition, such as in Figure 3a As shown, the encoded first audio signal portion also includes an encoded representation of a third set of third spectral portions to be reconstructed by noise filling, and the first decoding processor 1120 further includes a noise filler included in block 1122b for extracting noise filling information 308 from the encoded representation of the third set of third spectral portions and for applying noise filling operations in the third set of third spectral portions without using the first spectral portions in different frequency ranges.

[0220] Furthermore, the spectrum domain audio decoder 112 is configured to generate a first decoded representation having a first spectral portion, the frequency value of which is greater than a frequency equal to the midpoint of the frequency range covered by the time representation output by the spectrum-time converter 118 or 1124.

[0221] Furthermore, a spectrum analyzer or analyzer 604 (which may be a full-band analyzer) is configured to analyze the representation generated by the time-to-frequency converter 602 to determine a first set of first spectral portions to be encoded with a first high spectral resolution and a different second set of second spectral portions to be encoded with a second spectral resolution lower than the first spectral resolution, and, through the spectrum analyzer, to determine the frequency at... Figure 3a The first spectral portion 306 is located between the two second spectral portions at 307a and 307b.

[0222] Specifically, the spectrum analyzer is configured to analyze the spectrum representation up to a maximum analysis frequency, which is at least one-quarter of the sampling frequency of the audio signal.

[0223] Specifically, the spectral domain audio encoder is configured to process a sequence of frames of spectral values ​​for quantization and entropy coding, wherein, in a frame, the spectral values ​​of a second group of second portions are set to zero, or wherein, in a frame, there are spectral values ​​of a first group of first spectral portions and a second group of second spectral portions, and wherein, during subsequent processing, the spectral values ​​in the second group of spectral portions are set to zero, as exemplarily shown at 410, 418, 422.

[0224] The spectrum domain audio encoder is configured to produce a spectrum representation having a Nyquist frequency defined by the sampling rate of a first portion of an audio signal processed by either the audio input signal or a first encoding processor operating in the frequency domain.

[0225] The spectrum domain audio encoder 606 is also configured to provide a first coded representation such that, for a frame of the sampled audio signal, the coded representation includes a first set of first spectral portions and a second set of second spectral portions, wherein spectral values ​​in the second set of spectral portions are encoded as zero or noise values.

[0226] The analyzer 604 or 102 is configured to analyze a spectrum representation that begins at the gap-filling start frequency 309 and ends at the maximum frequency fmax, which is represented by the maximum frequency included in the spectrum representation, and the spectrum portion extending from the minimum frequency up to the gap-filling start frequency 309 belongs to the first group of first spectrum portions.

[0227] Specifically, the analyzer is configured to apply tone masking processing to at least a portion of the spectral representation, such that the tone components and non-tone components are separated from each other, wherein a first set of first spectral portions includes tone components, and wherein a second set of second spectral portions includes non-tone components.

[0228] The invention can be further implemented through the following embodiments, which can be combined with any examples and embodiments described and claimed herein:

[0229] 1. An audio encoder for encoding audio signals, comprising:

[0230] A first encoding processor (600) is configured to encode a portion of a first audio signal in the frequency domain, wherein the first encoding processor (600) includes:

[0231] A time-to-frequency converter (602) is used to convert a portion of a first audio signal into a frequency domain representation having spectral lines extending up to the maximum frequency of the first audio signal portion;

[0232] An analyzer (604) is configured to analyze the frequency domain representation up to the maximum frequency to determine a plurality of first spectral portions to be encoded with a first spectral resolution and a plurality of second spectral portions to be encoded with a second spectral resolution lower than the first spectral resolution, wherein the analyzer (604) is configured to determine a first spectral portion (306) of the plurality of first spectral portions, the first spectral portion being set relative to a frequency between two second spectral portions (307a, 307b) of the plurality of second spectral portions;

[0233] A spectrum encoder (606) is configured to encode the plurality of first spectrum portions with a first spectrum resolution and to encode the plurality of second spectrum portions with a second spectrum resolution, wherein the spectrum encoder includes a parameter encoder configured to calculate spectral envelope information having a second spectrum resolution based on the plurality of second spectrum portions;

[0234] The second encoding processor (610) is used to encode different portions of the second audio signal in the time domain;

[0235] The controller (620) is configured to analyze the audio signal and to determine which part of the audio signal is a first audio signal portion encoded in the frequency domain and which part of the audio signal is a second audio signal portion encoded in the time domain; and

[0236] The encoded signal former (630) is used to form an encoded audio signal, the encoded audio signal including a first encoded signal portion for a first audio signal portion and a second encoded signal portion for a second audio signal portion.

[0237] 2. The audio encoder according to Embodiment 1, wherein the input signal has a high-frequency band and a low-frequency band.

[0238] The second encoding processor (610) includes a sampling rate converter (900) for converting a portion of the second audio signal into a lower sampling rate representation, the lower sampling rate being lower than the sampling rate of the audio signal, wherein the lower sampling rate representation does not include the high-frequency band of the input signal;

[0239] A time-domain low-frequency encoder (910) is used for time-domain encoding of lower sampling rate representations; and

[0240] A time-domain bandwidth extension encoder (920) is used to encode high-frequency bands parametrically.

[0241] 3. The audio encoder according to Embodiment 1 further includes:

[0242] The preprocessor (1000) is configured to preprocess the first audio signal portion and the second audio signal portion.

[0243] The preprocessor includes:

[0244] Predictive analyzer (1002) is used to determine the prediction coefficients; and

[0245] The second encoding processor includes:

[0246] A predictor coefficient quantizer (1010) is used to generate a quantized version of the predictor coefficients; and

[0247] An entropy encoder is used to produce an encoded version of the quantized prediction coefficients.

[0248] The encoded signal shaper (630) is configured to incorporate the encoded version into the encoded audio signal.

[0249] 4. The audio encoder according to Embodiment 1,

[0250] The preprocessor (1000) includes a resampler (1004) for resampling the audio signal to the sampling rate of the second encoding processor; and

[0251] The predictive analyzer is configured to use the resampled audio signal to determine the prediction coefficients, or

[0252] The preprocessor (1000) also includes a long-term prediction analysis level for determining one or more long-term prediction parameters for the first audio signal portion.

[0253] 5. The audio encoder according to Embodiment 1 further includes a cross processor (700) for calculating initialization data of a second encoding processor (610) based on the encoded spectral representation of the first audio signal portion, such that the second encoding processor (610) is initialized to encode a second audio signal portion of the audio signal that follows the first audio signal portion in time.

[0254] 6. The audio encoder according to Embodiment 5, wherein the cross processor (700) comprises:

[0255] A spectrum decoder (701) is used to compute the decoded version of the first coded signal portion;

[0256] The delay stage (707) is used to feed the delayed version of the decoded version to the de-emphasis stage (617) of the second encoding processor for initialization;

[0257] The weighted prediction coefficient analysis filter block (708) is used to feed the filter output to the codebook determiner (613) of the second encoding processor (610) for initialization;

[0258] The analysis filter stage (706) is used to filter the decoded version or the pre-emphasized version (709), and to feed the filter residue into the adaptive codebook determiner (612) of the second encoding processor for initialization; or

[0259] A pre-emphasis filter (709) is used to filter the decoded version and to feed the delayed or pre-emphasis version to the synthesis filter stage (616) of the second encoding processor (610) for initialization.

[0260] 7. The audio encoder according to Embodiment 1,

[0261] The analyzer (604) is configured to perform time-blocking or time-noise-shaping analysis or to set the spectral values ​​in the second spectral portion to zero.

[0262] The first encoding processor (600) is configured to perform spectral shaping (606a) of the spectral values ​​of the first spectral portion using prediction coefficients (1010) derived from the first audio signal portion, and the first encoding processor (600) is further configured to perform quantization and entropy encoding operations (606b) on the quantized spectral values ​​of the first spectral portion.

[0263] In this case, the spectral value of the second spectral portion is set to zero.

[0264] 8. The audio encoder according to Embodiment 7 further includes a cross processor (700), wherein the cross processor (700) includes:

[0265] A noise shaper (703) is used to shape the quantized spectral values ​​of the first spectral portion using LPC coefficients (1010) derived from the first audio signal portion;

[0266] The spectrum decoder (704, 705) is used to decode the spectrum portion of the spectrum-shaped first spectrum portion with high spectral resolution, and to synthesize the second spectrum portion using the parameter representation of the second spectrum portion and at least the decoded first spectrum portion to obtain the decoded spectrum representation.

[0267] A frequency-to-time converter (702) is used to convert a spectral representation to the time domain to obtain a first audio signal portion that is decoded, wherein the sampling rate associated with the first audio signal portion that is decoded is different from the sampling rate of the audio signal, and the sampling rate associated with the output signal of the frequency-to-time converter (702) is different from the sampling rate of the audio signal input to the frequency-to-time converter (602).

[0268] 9. The audio encoder according to Embodiment 1, wherein the second encoding processor includes at least one block from the following block group:

[0269] Measurement and analysis filter (611);

[0270] Adaptive codebook level (612);

[0271] Innovation Codebook Level (614);

[0272] Estimator (613) is used to estimate innovative codebook entries;

[0273] ACELP / Gain Coding Level (615);

[0274] Predictive synthetic filter stage (616);

[0275] Remove the heavy level (617); and

[0276] Bass post-filter analysis stage (618).

[0277] 10. The audio encoder according to Embodiment 1,

[0278] The time-domain coding processor has an associated second sampling rate.

[0279] The frequency domain coding processor has a first sampling rate associated with it that is higher than the second sampling rate, and the audio encoder further includes a cross processor (700) for calculating initialization data for the second coding processor from the encoded spectral representation of the first audio signal portion.

[0280] The cross processor includes a frequency-to-time converter (702) for generating a time-domain signal at a second sampling rate.

[0281] The frequency-time converter (702) includes:

[0282] Selector (726) is used to select the lower portion of the spectrum input to the frequency-time converter based on the ratio of a first sampling rate and a second sampling rate, wherein the ratio of the first sampling rate and the second sampling rate is less than 1.

[0283] The converter processor (720) has a smaller conversion length than the time-to-frequency converter (602); and

[0284] A synthetic windower (712) is used for windowing with a smaller number of window coefficients compared to the window used by the time-to-frequency converter (602).

[0285] 11. An audio decoder for decoding encoded audio signals, comprising:

[0286] A first decoding processor (1120) is used to decode a portion of a first coded audio signal in the frequency domain. The first decoding processor (1120) includes:

[0287] A spectrum decoder (1122) is configured to decode a plurality of first spectral portions with high spectral resolution and to synthesize the plurality of second spectral portions using parameter representations of a plurality of second spectral portions and at least the decoded first spectral portions to obtain a decoded spectral representation, wherein the spectrum decoder (1122) is configured to generate a first decoded representation such that a first spectral portion (306) is positioned between two second spectral portions (307a, 307b) relative to a frequency; and

[0288] A frequency-to-time converter (1120) is used to convert the decoded spectral representation to the time domain to obtain a decoded first audio signal portion;

[0289] A second decoding processor (1140) is used to decode a portion of the second coded audio signal in the time domain to obtain the decoded second audio signal portion; and

[0290] A combiner (1160) is used to combine the first spectral portion of the decoded signal and the second spectral portion of the decoded signal to obtain the decoded audio signal.

[0291] 12. The audio decoder according to Embodiment 11, wherein the second decoding processor comprises:

[0292] A time-domain low-frequency band decoder (1200) is used to decode low-frequency band time-domain signals;

[0293] Upsampler (1210) is used to upsample low-frequency band time-domain signals;

[0294] A time-domain bandwidth-extended decoder (1220) is used to synthesize high-frequency bandgap time-domain output signals; and

[0295] A mixer (1230) is used to mix the high-frequency band of the synthesized time-domain signal and the upsampled low-frequency band time-domain signal.

[0296] 13. The audio encoder according to Embodiment 12,

[0297] The upsampler (1210) includes an analysis filter bank (1471) operating at a first time-domain low-frequency band decoder sampling rate and a synthesis filter bank (1473) operating at a second output sampling rate higher than the first time-domain low-frequency band sampling rate.

[0298] 14. The audio decoder according to Embodiment 12,

[0299] The time-domain low-frequency band decoder (1200) includes a residual signal, decoders (1149, 1141, 1142), and a synthesis filter (1143). The synthesis filter (1143) is used to filter the residual signal using synthesis filter coefficients (1145).

[0300] The time-domain bandwidth extension decoder (1220) is configured to upsample the residual signal (1221), process (1222) the upsampled residual signal using nonlinear operations to obtain a high-frequency band residual signal, and perform spectral shaping (1223) on the high-frequency band residual signal to obtain a synthesized high-frequency band.

[0301] 15. The audio decoder according to Embodiment 11,

[0302] The first decoding processor (1120) includes an adaptive long-term prediction post-filter (1420) for post-filtering a first signal portion of the first decoding, wherein the filter (1420) is controlled by one or more long-term prediction parameters included in the encoded audio signal.

[0303] 16. The audio decoder according to Embodiment 11 further includes:

[0304] A cross processor (1170) is configured to calculate initialization data for a second decoding processor (1140) from the decoded spectral representation of the first encoded audio signal portion, such that the second decoding processor (1140) is initialized to decode the encoded second audio signal portion of the encoded audio signal that follows the first audio signal portion in time.

[0305] 17. The audio decoder according to Embodiment 16, wherein the cross processor further includes:

[0306] The frequency-to-time converter (1170) operates at a lower sampling rate compared to the frequency-to-time converter (1124) of the first decoding processor (1120) to obtain a first portion of the signal for further decoding in the time domain.

[0307] The signal output by the frequency-to-time converter (1171) has a second sampling rate that is lower than the first sampling rate associated with the output of the frequency-to-time converter (1124) of the second decoding processor.

[0308] The additional frequency-to-time converter (1171) includes a selector (726) for selecting the lower portion of the spectrum input to the additional frequency-to-time converter (1171) according to a ratio of a first sampling rate and a second sampling rate, wherein the ratio of the first sampling rate and the second sampling rate is less than 1.

[0309] The conversion processor (720) has a smaller conversion length than the conversion length (710) of the time-to-frequency converter (1124); and

[0310] The synthesis windower (722) uses a window with a smaller number of coefficients compared to the window used by the frequency-time converter (1124).

[0311] 18. The audio decoder according to embodiment 16, wherein the cross processor (1170) includes:

[0312] The delay stage (1172) is used to delay the first signal portion for further decoding and to feed the delayed version of the first signal portion to the de-emphasis stage (1144) of the second decoding processor for initialization;

[0313] The pre-emphasis filter (1173) and delay stage (1175) are used to filter and delay the first signal portion for further decoding, and to feed the output of the delay stage into the predictive synthesis filter (1143) of the second decoding processor for initialization;

[0314] A predictive analysis filter (1174) is used to generate a predictive residual signal from a further decoded first spectral portion or a further decoded first signal portion of a pre-emphasized (1173) signal, and to feed the predictive residual signal into the codebook synthesizer (1141) of the second decoding processor (1200); or

[0315] A switch (1480) is used to feed the first signal portion of the further decoded signal to the analysis stage (1471) of the resampler (1210) of the second decoding processor for initialization.

[0316] 19. The audio decoder according to Embodiment 11,

[0317] The second decoding processor (1200) includes at least one block in a block group, the block group comprising:

[0318] ACELP is used for decoding gain and innovation codebooks;

[0319] Adaptive codebook synthesis level (1141);

[0320] ACELP post-processor (1142);

[0321] Predictive synthesis filter (1143); and

[0322] Remove the severity level (1144).

[0323] 20. A method for encoding an audio signal, comprising:

[0324] In the frequency domain, a first encoding (600) is performed on a portion of the first audio signal, wherein the first encoding (600) includes:

[0325] The first audio signal portion is converted (602) into a frequency domain representation with spectral lines up to the maximum frequency of the first audio signal portion;

[0326] Analysis (604) extends up to the frequency domain representation of the maximum frequency to determine a plurality of first spectral portions to be encoded with a first spectral resolution and a plurality of second spectral portions to be encoded with a second spectral resolution lower than the first spectral resolution, wherein the analysis (604) determines a first spectral portion (306) of the plurality of first spectral portions, the first spectral portion being set relative to a frequency between two second spectral portions (307a, 307b) of the plurality of second spectral portions;

[0327] Encoding the plurality of first spectral portions using the first spectral resolution (606), and encoding the plurality of second spectral portions using the second spectral resolution, wherein encoding the second spectral portions includes calculating spectral envelope information having the second spectral resolution based on the plurality of second spectral portions;

[0328] In the time domain, different portions of the second audio signal are encoded in a second way (610).

[0329] Analyze the (620) audio signal and determine which part of the audio signal is the first audio signal portion encoded in the frequency domain, and which part of the audio signal is the second audio signal portion encoded in the time domain; and

[0330] An encoded audio signal is formed (630), the encoded audio signal including a first encoded signal portion for a first audio signal portion and a second encoded signal portion for a second audio signal portion.

[0331] 21. A method for decoding an encoded audio signal, comprising:

[0332] In the frequency domain, a first decoding (1120) is performed on a portion of the first coded audio signal, the first decoding (1120) including:

[0333] Decoding (1122) a plurality of first spectral portions using high spectral resolution, and synthesizing the plurality of second spectral portions using parameter representations of the plurality of second spectral portions and at least the decoded first spectral portions to obtain a decoded spectral representation, wherein decoding (1122) includes generating a first decoded representation such that a first spectral portion (306) is positioned between two second spectral portions (307a, 307b) relative to a frequency; and

[0334] The decoded spectral representation is transformed (1120) into the time domain to obtain the decoded first audio signal portion;

[0335] In the time domain, the second coded audio signal portion is subjected to a second decoding (1140) to obtain the decoded second audio signal portion; and

[0336] Combine (1160) the first spectral portion of the decoded signal and the second spectral portion of the decoded signal to obtain the decoded audio signal.

[0337] 22. A machine-readable storage medium storing a computer program, which, when run on a computer or processor, is used to perform the method according to embodiment 20 or embodiment 21.

[0338] Although the invention has been described in the context of block diagrams (where the blocks represent real or logical hardware components), the invention can also be implemented as a computer-implemented method. In the latter case, the blocks represent corresponding method steps, where these steps represent functionality performed by the corresponding logical or physical hardware blocks.

[0339] Although some aspects have been described in the context of the apparatus, it will be clear that these aspects also represent a description of the corresponding method, wherein a block or device corresponds to a method step or a feature of a method step. Similarly, the scheme described in the context of method steps also represents a description of the features of the corresponding block or item or the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices (such as microprocessors, programmable computers, or electronic circuits). In some embodiments, one or more of the most important method steps may be performed by such devices.

[0340] The transmitted or encoded signals of the present invention can be stored on a digital storage medium or transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0341] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or software. The implementations can be performed using a digital storage medium (e.g., floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) on which electronically readable control signals are stored, the control signals cooperating with (or capable of cooperating with) a programmable computer system to cause the various methods to be performed. Therefore, the digital storage medium can be computer-readable.

[0342] Some embodiments of the invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0343] Typically, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer. The program code may, for example, be stored on a machine-readable medium.

[0344] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.

[0345] In other words, embodiments of the method of the present invention are therefore computer programs having program code for performing one of the methods described herein when the computer program is run on a computer.

[0346] Therefore, another embodiment of the method of the present invention is a data carrier (or a non-transitory storage medium such as a digital storage medium or a computer-readable medium) containing a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.

[0347] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).

[0348] Another embodiment includes a processing means, such as a computer or programmable logic device configured or adapted to perform one of the methods described herein.

[0349] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.

[0350] Another embodiment of the invention includes an apparatus or system configured to transmit a computer program to a receiver (e.g., electronically or optically) for performing one of the methods described herein. The receiver may be, for example, a computer, mobile device, storage device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.

[0351] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0352] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the invention is intended to be limited only by the scope of the appended claims and not by the specific details given by way of the description and explanation of the embodiments herein.

Claims

1. An audio encoder for encoding an audio signal, comprising: a first encoding processor for encoding a first audio signal portion of the audio signal in a frequency domain, wherein the first encoding processor comprises: a time-to-frequency converter for converting the first audio signal portion into a frequency domain representation comprising spectral lines up to a maximum frequency of the first audio signal portion; and a spectral encoder for encoding the frequency domain representation to obtain an encoded spectral representation of the first audio signal portion; a second encoding processor for encoding a second audio signal portion of the audio signal in a time domain, wherein the second audio signal portion is different from the first audio signal portion; a cross processor for computing initialization data for the second encoding processor from the encoded spectral representation of the first audio signal portion, such that the second encoding processor is initialized to encode a second audio signal portion of the audio signal in the time domain that temporally follows the first audio signal portion; a controller configured for analyzing the audio signal and for determining which portion of the audio signal is the first audio signal portion that is encoded in the frequency domain and which portion of the audio signal is the second audio signal portion that is encoded in the time domain; an encoded signal former for forming an encoded audio signal comprising a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion; and a pre-processor configured for pre-processing the first audio signal portion and the second audio signal portion, wherein the pre-processor comprises a resampler for resampling the audio signal to a sampling rate of the second encoding processor to obtain a resampled audio signal and a prediction analyzer configured to determine prediction coefficients using the resampled audio signal, or wherein the pre-processor comprises a long-term prediction analysis stage for determining one or more long-term prediction parameters for the first audio signal portion.

2. The audio encoder of claim 1, wherein the audio signal comprises a high frequency band and a low frequency band, and wherein the second encoding processor comprises: a sampling rate converter for converting the second audio signal portion into a representation having a lower sampling rate, the lower sampling rate being lower than a sampling rate of the audio signal, wherein the representation having the lower sampling rate does not comprise the high frequency band of the audio signal; a time domain low band encoder for time domain encoding the representation having the lower sampling rate; and a time domain bandwidth extension encoder for parametrically encoding the high frequency band.

3. The audio encoder of claim 1, wherein the pre-processor comprises a prediction analyzer for determining prediction coefficients; and wherein the encoded signal former is configured for introducing an encoded version of the prediction coefficients into the encoded audio signal.

4. The audio encoder of claim 1, wherein, the cross processor comprises: a spectral decoder for computing a decoded version of the first encoded signal portion; and a delay stage for delaying the decoded version of the first encoded signal portion to obtain a delayed version and feeding the delayed version into a de-emphasis stage of the second encoding processor for initialization.

5. The audio encoder of claim 1, wherein, the cross processor comprises: a spectral decoder for computing a decoded version of the first encoded signal portion; and a weighted prediction coefficient analysis filter block for filtering the decoded version of the first encoded signal portion to obtain a filter output and feeding the filter output to the innovation codebook determinator of the second encoding processor for initialization.

6. The audio encoder of claim 1, wherein, The cross-processor comprises: a spectral decoder for computing a decoded version of the first encoded signal portion; and an analysis filter stage for filtering the decoded version of the first encoded signal portion or a pre-emphasis version derived from the decoded version of the first encoded signal portion by the pre-emphasis stage to obtain a filter residual signal and feeding the filter residual signal to the adaptive codebook determinator of the second encoding processor for initialization.

7. The audio encoder of claim 1, wherein, The cross-processor comprises: a spectral decoder for computing a decoded version of the first encoded signal portion; and a pre-emphasis filter for filtering the decoded version of the first encoded signal portion to obtain a pre-emphasized version and feeding the pre-emphasized version or a delayed pre-emphasized version to the synthesis filter stage of the second encoding processor for initialization.

8. The audio encoder of claim 1, wherein the first encoding processor is configured to perform a shaping of spectral values of the frequency domain representation using prediction coefficients derived from the first audio signal portion to obtain shaped spectral values, and wherein the first encoding processor is further configured to perform a quantization and an entropy coding operation of the shaped spectral values of the frequency domain representation.

9. The audio encoder of claim 1, wherein, The cross-processor comprises: a noise shaper for shaping quantized spectral values of the frequency domain representation using LPC coefficients derived from the first audio signal portion; a spectral decoder for decoding spectral shaped spectral portions of the frequency domain representation at a high spectral resolution to obtain a decoded spectral representation; and a frequency-time converter for converting the decoded spectral representation into the time domain to obtain a decoded first audio signal portion, wherein a sampling rate associated with the decoded first audio signal portion is different from a sampling rate of the audio signal, and a sampling rate associated with an output signal of the frequency-time converter is different from a sampling rate associated with the audio signal input into the time-frequency converter.

10. The audio encoder of claim 1, wherein the second encoding processor comprises at least one element of the following element group: a prediction analysis filter; an adaptive codebook stage; an innovation codebook stage; an estimator for estimating innovation codebook entries; an ACELP / gain coding stage; a prediction synthesis filter stage; a de-emphasis stage; and a bass post-filter analysis stage.

11. The audio encoder of claim 1, wherein the second encoding processor comprises an associated second sampling rate, wherein the first encoding processor has a first sampling rate associated therewith, the first sampling rate being different from the second sampling rate, wherein the cross-processor comprises a frequency-time converter for producing a time domain signal at the second sampling rate, and wherein the frequency-time converter comprises: a selector for selecting a portion of a spectrum input into the frequency-time converter depending on a ratio of the first sampling rate and the second sampling rate, a transform processor comprising a transform length different from a transform length of the time-frequency converter; and a synthesis windower for windowing using a window comprising a different number of window coefficients compared to a window used by the time-frequency converter.

12. An audio decoder for decoding an encoded audio signal, comprising: a first decoding processor for decoding a first encoded audio signal portion of the encoded audio signal in a frequency domain, wherein the first decoding processor is configured to reconstruct the first set of first spectral portions in a waveform-preserving manner to generate a spectrum with gaps, wherein the gaps in the spectrum are filled with an intelligent gap filling, IGF, technique including a frequency regeneration using the application parameter data and using the reconstructed first spectral portions of the first set of first spectral portions to obtain a decoded spectral representation, and wherein the first decoding processor comprises a frequency-time converter for converting the decoded spectral representation into a time domain to obtain a decoded first audio signal portion; a second decoding processor for decoding a second encoded audio signal portion of the encoded audio signal in a time domain to obtain a decoded second audio signal portion; a cross processor for computing initialization data for the second decoding processor from the decoded spectral representation of the first encoded audio signal portion such that the second decoding processor is initialized to decode the second encoded audio signal portion of the encoded audio signal following the first encoded audio signal portion in time in the time domain; and a combiner for combining the decoded first audio signal portion and the decoded second audio signal portion to obtain a decoded audio signal.

13. The audio decoder of claim 12, wherein, The second decoding processor comprises: a time-domain low-band decoder for decoding to obtain a low-band time-domain signal; a resampler for resampling the low-band time-domain signal; a time-domain bandwidth extension decoder for synthesizing a high-band of the time-domain output signal; and a mixer for mixing the synthesized high-band of the time-domain output signal and the resampled low-band time-domain signal.

14. The audio decoder of claim 12, wherein the first decoding processor comprises an adaptive long-term prediction postfilter for postfiltering the decoded first audio signal portion, wherein the postfilter is controlled by one or more long-term prediction parameters included in the encoded audio signal.

15. The audio decoder of claim 12, wherein, The cross processor further comprises: a frequency-time converter operating at a first effective sampling rate different from a second effective sampling rate associated with the frequency-time converter of the first decoding processor to obtain the further decoded first audio signal portion in the time domain, wherein a signal output by the frequency-time converter comprises a second sampling rate different from a first sampling rate associated with an output of the frequency-time converter of the second decoding processor, wherein the additional frequency-time converter comprises a selector for selecting a portion of the spectrum input into the additional frequency-time converter depending on a ratio of the first sampling rate and the second sampling rate; a transform processor comprising a transform length different from a transform length of the frequency-time converter; and a synthesis windower for windowing using a window comprising a different number of window coefficients compared to a window used by the time-frequency converter. A synthesis windower uses windows comprising a different number of coefficients than the windows used by the frequency-time converter.

16. The audio decoder of claim 12, wherein the cross processor comprises: a delay stage for delaying the further decoded first audio signal portion and feeding a delayed version of the further decoded first audio signal portion into a de-emphasis stage of the second decoding processor for initialization.

17. The audio decoder of claim 12, wherein, The cross processor comprises a pre-emphasis filter and a delay stage for filtering and delaying the further decoded first audio signal portion and feeding a delay stage output into a prediction synthesis filter of the second decoding processor for initialization.

18. The audio decoder of claim 12, wherein, The cross processor comprises a prediction analysis filter for generating a prediction residual signal from the further decoded first audio signal portion or from the pre-emphasized further decoded first audio signal portion and feeding the prediction residual signal into a codebook synthesizer of the second decoding processor.

19. The audio decoder of claim 12, wherein, The cross processor comprises a switch for feeding the further decoded first audio signal portion into an analysis stage of a resampler of the second decoding processor for initialization.

20. The audio decoder of claim 12, wherein the second decoding processor comprises at least one element of a group of elements, the group of elements comprising: a stage for decoding an ACELP gain and an innovation codebook; an adaptive codebook synthesis stage; an ACELP post-processor; a prediction synthesis filter; and a de-emphasis stage.

21. A method of encoding an audio signal, comprising: encoding a first audio signal portion of the audio signal in a frequency domain, the encoding comprising: converting the first audio signal portion into a frequency domain representation comprising spectral lines up to a maximum frequency of the first audio signal portion; and encoding the frequency domain representation to obtain an encoded spectral representation of the first audio signal portion; encoding a second audio signal portion of the audio signal in a time domain, wherein the second audio signal portion is different from the first audio signal portion; computing initialization data for steps of encoding the second audio signal portion from the encoded spectral representation of the first audio signal portion, such that the steps of encoding the second audio signal portion are initialized to encode the second audio signal portion of the audio signal in the time domain that temporally follows the first audio signal portion; analyzing the audio signal and determining which portion of the audio signal is the first audio signal portion encoded in the frequency domain and which portion of the audio signal is the second audio signal portion encoded in the time domain; forming an encoded audio signal comprising a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion, and pre-processing the first audio signal portion and the second audio signal portion, wherein the pre-processing comprises resampling the audio signal to a sampling rate of the second encoding processor to obtain a resampled audio signal; and using the resampled audio signal to determine prediction coefficients, or wherein the pre-processing comprises determining one or more long-term prediction parameters for the first audio signal portion.

22. A method of decoding an encoded audio signal, comprising: decoding a first encoded audio signal portion of the encoded audio signal in a frequency domain, the decoding comprising: reconstructing the first set of first spectral portions in a waveform-preserving manner to generate a spectrum with gaps, and filling the gaps in the spectrum with an intelligent gap filling, IGF, technique including using frequency regeneration of applied parameter data and using the reconstructed first spectral portions of the first set of first spectral portions to obtain a decoded spectral representation, and converting the decoded spectral representation into a time domain to obtain a decoded first audio signal portion; decoding a second encoded audio signal portion of the encoded audio signal in a time domain to obtain a decoded second audio signal portion; computing initialization data for the step of decoding the second encoded audio signal portion from the decoded spectral representation of the first encoded audio signal portion, such that the step of decoding the second encoded audio signal portion is initialized to decode the second encoded audio signal portion in the time domain that temporally follows the first encoded audio signal portion in the encoded audio signal; and combining the decoded first audio signal portion and the decoded second audio signal portion to obtain a decoded audio signal.

Citation Information

Patent Citations

  • Integrated voice / audio encoding / decoding device and method whereby the overlap region of a window is adjusted based on the transition interval

    US20120209600A1

  • Audio signal encoder, audio signal decoder, method for encoding or decoding an audio signal using an aliasing-cancellation

    US20120271644A1