Audio encoder or audio decoder using a raw coding operation

The raw coding of windowed and time-domain aliased frames in audio codecs addresses the unpredictability of bitrate in lapped transform-based codecs, ensuring efficient encoding and decoding with a fixed maximum frame size and reduced computational complexity.

WO2026153655A1PCT designated stage Publication Date: 2026-07-23FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2025-05-30
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing lapped transform-based lossy and lossless audio codecs face challenges in predicting the required bitrate accurately, often resulting in excessive data rates due to time-frequency transforms and entropy encoding, particularly when encoding raw PCM signals.

Method used

Implementing a raw coding approach for windowed and time-domain aliased frames, which includes analysis windowing and time-domain aliasing processing, to ensure a maximum frame size and reduce computational complexity by encoding windowed and optionally folded audio signals without transforms.

Benefits of technology

This method achieves a predictable maximum bitrate per frame, reduces computational complexity, and enhances coding efficiency by minimizing the number of samples to be encoded, while maintaining lossless or lossy audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025065079_23072026_PF_FP_ABST
    Figure EP2025065079_23072026_PF_FP_ABST
Patent Text Reader

Abstract

An audio encoder for encoding an audio signal comprising a sequence of blocks of audio samples comprises: an analysis windower (100) for processing the sequence of blocks of audio samples using an analysis window, wherein the analysis window has a length being larger than a length of a block of audio samples of the blocks of audio samples, wherein the analysis windower (100) is configured to advance, in the sequence of blocks of audio samples, from a windowing operation to a following windowing operation by an amount of samples being smaller than a length of the analysis window so that each audio sample is windowed by at least two subsequent windowing operations; and a raw encoder (200) for raw encoding the samples of a processed block of a sequence of processed blocks to obtain raw-encoded data for the processed block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Audio Encoder or Audio Decoder Using a Raw Coding Operation

[0002] Specification

[0003] The present invention relates to audio processing and to coding or decoding using a raw coding operation or a raw decoding operation in a windowing domain or a time domain aliasing domain.

[0004] The goal of a lossless or lossy audio codec is to reduce the data rate, but there are always signals which will lead to a higher data rate than strictly necessary when encoded as a raw PCM (pulse code modulation) signal.

[0005] Typically, there do exist lapped transformed based lossless or lossy audio codecs such as an AAC codec, a USAC codec, an LC3 codec or other codecs.

[0006] Typically, for such a codec, it is not immediately obvious how such a lapped transform based lossy or lossless audio codec can have a maximum predictable bitrate. Particularly, due to the time-frequency transform and the subsequent entropy encoding, the number of bits required for a certain block cannot be easily predicted.

[0007] Furthermore, the situation can occur that the number of bits required for the time-frequency transform and the subsequent entropy encoding using certain audio tools or directly entropy encoding the result of the time-spectrum conversion can even result in a situation where the required bitrate can become excessive.

[0008] Therefore, it is an object of the present invention to provide an improved audio coding / de-coding concept.

[0009] This object is achieved by an audio encoder for encoding an audio signal of claim 1, an audio decoder for decoding an encoded audio data of claim 19, a method of encoding of claim 34, a method of decoding of claim 35, a computer program of claim 36 or an encoded audio signal of claim 37.

[0010] The present invention is based on the finding that a raw coding of windowed or windowed and time domain aliased frames is performed in order to avoid a situation, where an audiocodec results in a higher data rate than strictly necessary when applying the raw (e.g. PCM) coding. This leads to defined maximum frame size irrespective of the characteristics of the audio signal to be coded. In accordance with the invention, the analysis windowing operation is performed before performing the raw encoding. In a preferred embodiment, the analysis windowing operation not only comprises using the analysis window for weighting the time domain samples of a block of audio samples, but also comprises the application of the time domain aliasing (combination or folding or fold-in) processing to obtain the sequence of processed blocks which are then encoded by the raw encoder.

[0011] While the PCM coding of the windowed samples already provides an advantage over PCM coding the non-windowed samples of a block due to the fact that windowing typically results in smaller values, the application of the time domain aliasing processing enhances the coding efficiency significantly, since the application of the time domain aliasing processing reduces the amount of time domain samples to be encoded by a certain amount such as by a factor of two in the context of MDCT processing or, generally, in the context of lapped transform based lossy or lossless audio codecs.

[0012] The inventive procedure of raw encoding windowed or windowed and time domain aliasing processed samples is superior over the raw encoding of spectral lines, too. The specific raw encoding of spectral lines would raise the question how many bits are necessary to store each spectral line. When one assumes that the input signal is sampled at 16 bits per sample, it is not predictable how large each spectral line can get, in view of the energy compacting properties of the underlying transform. Although it might be possible to figure this out theoretically and to design an algorithm to optimally code the spectral lines with low overhead, the present invention is superior over this procedure, since it is not only simpler, but also less computationally complex.

[0013] In accordance with preferred embodiments, lapped transforms can be broken up into three steps, the windowing step, the time domain aliasing step and the transform itself. Rather than raw coding the spectral lines coming out of the transform or rather than raw encoding the unprocessed time domain samples of a block, the present invention relies on the raw coding of the windowed and preferably time domain aliased signal. This will by design result in a maximum overhead of only one bit per sample in case of time domain aliasing compared to the input signal. Therefore, by skipping the transform and performing the raw encoding without any transform, a less computationally complex procedure is achieved bothwithin the encoder and in the decoder compared to the application of the transform, i.e., compared to a raw coding of spectral lines.

[0014] A preferred embodiment of the present invention refers to the raw coding of a windowed signal or a windowed and time domain aliased signal.

[0015] Another embodiment refers to a method for ensuring a maximum frame size in a lapped transformed based lossless audio codec. If the entropy coded spectral lines would cause the frame to surpass a certain size threshold, the present invention is used as a fallback mode in which the windowed and preferably time domain aliased signal is raw coded with a maximum bitrate of the input signal bits per sample plus one bit per sample.

[0016] In an embodiment, the audio encoder encodes an audio signal comprising a sequence of blocks of audio samples. The audio encoder comprises an analysis windower (100) where the analysis window has a length being larger than a length of a block of audio samples of the blocks of audio samples. Furthermore, the analysis windower (100) is configured to advance, in the sequence of blocks of audio samples, from a windowing operation to a following windowing operation by an amount of samples being smaller than a length of the analysis window so that at least two audio samples are windowed by at least two subsequent windowing operations. The procedure of the windower results in a processed block of a sequence of processed blocks, which are input into a raw encoder for raw encoding the samples of such a processed block of the sequence of processed blocks to obtain raw-encoded data for the processed block.

[0017] Preferably, the windowing operation not only comprises using information on the analysis window, but also the application of a time domain aliasing processing. This time domain aliasing processing comprises a combination of certain portions of windowed samples taken from a sampling operation by a single window. A specific combination of such portions of samples comprises the folding in operation in the context of the MDCT processing (MDCT = modified discrete cosine transform).

[0018] On the decoder-side, a raw decoder is applied for decoding the raw-encoded data to obtain a processed block, and this processed block is fed to a processor for processing the processed block using a synthesis window and for overlapping and adding current block data and preceding block data to obtain a block of decoded audio samples. The processor may comprise a synthesis windower using a synthesis window to obtain a synthesized block orportions of the synthesized block which are then processed by an overlap-adder for overlapping and adding the portions of the synthesized block and portions of a preceding synthesized block to obtain a block of decoded audio samples. When the windowing and over-lap-adding operations are performed jointly, the processor does not have a separate windower and overlap-adder. Such a procedure may occur in an integer implementation using lifting steps, where the processor uses the current block data and the preceding block data within the lifting steps, where the weighting and adding operations performed in the lifting steps together implement the windowing and overlap adding.

[0019] Specifically, the processor not only comprises the application of the synthesis window but, preferably, the corresponding folding-out, unfolding or extension operation in the context of lapped transforms and the overlap add operation. By this extension operation, the foldingin or combination operation on the encoder-side is reversed so that the output of the extension operation has more samples than the input into this operation.

[0020] The raw encoding of the time domain windowed and optionally folded audio signal can be used as a normal coding tool or can be used as a fallback mode in a lossy or lossless audio coder having an entropy coder. In the normal coding tool, several blocks of the sequence of processed blocks are windowed and preferably folded blocks are then either raw encoded such as by a PCM coder or any other coder that encodes the individual time domain samples or groups of individual time domain samples, or are encoded by a well-known lossy or lossless coder having the entropy coder.

[0021] On the decoder-side, the encoded audio data is input into a corresponding raw decoder and the raw decoded analysis-windowed and preferably folded data are input into a processor performing the operation of the synthesis windower. The synthesis windower preferably performs the unfolding operation and the application of the synthesis window in order to obtain a synthesized block or portions of the synthesized block that can be used by an overlap-adder comprised by the processor to perform an overlap-adding operation with portions of a preceding synthesized block in order to obtain a block of decoded audio samples. However, the windowing and overlap-adding operations can also be performed jointly. Then, the processor does not have a separate windower and overlap-adder but performs an operation equivalent to both synthesis windowing and overlap adding.

[0022] In another embodiment, the raw encoding of time domain windowed and preferably folded data and the corresponding raw decoding of windowed and preferably folded data using araw decoder and the subsequent unfolding in a processor or synthesis windower are performed as a fallback mode. In such a situation, the analysis windower on the one hand and the processor or synthesis windower on the other hand are used for both coding modes, i.e. , a regular coding mode that uses an entropy encoder / entropy decoder and a time-frequency converter in the encoder and a frequency-time converter in the decoder. However, depending on the coding mode decision, either the branch with the raw encoder or the branch with the time-frequency converter and the entropy encoder or, with respect to the decoder, the branch with the entropy decoder and the frequency-time converter is used. Hence, it is made sure that a useful time domain aliasing cancellation in the decoder due to the functionality of the overlap-adder is obtained even between a frame that is encoded in the fallback mode and a frame that is encoded in the regular mode.

[0023] The decision which coding mode is to be taken, i.e., either the regular mode or the fallback mode is done either using a feedback loop that calculates the bit consumption required by the entropy encoder branch and compares this bit consumption with the maximum threshold given by the raw encoder performance for the coding of the processed block. In case of a PCM coder, the coding of the processed block requires, at the maximum, 1 bit more per sample due to the folding operation that, at the maximum, results in a doubling of the greatest number.

[0024] Exemplarily, for example, for a 16 bits PCM representation, each sample of the processed block requires 17 bits at the maximum in case of the folding operation. When, however, the embodiment is considered, in which the windowed data are PCM encoded, then only 16 bits are required due to the fact that windowing alone reduces the values or leaves the values as they are. However the price is that due to the not-performed folding operation, the number of samples in this embodiment is higher than in the other embodiment, since an advantageous side effect of the folding is the strong reduction of the number of samples to be encoded. This is for the price of one bit more per sample, but this “penalty” is much less than the savings by the reduction of the number of samples to be raw encoded in case of the folding operation.

[0025] In other embodiments, an open loop forward coding mode decision can be performed by means of signal characteristics detection of the audio signal in order to decide, whether, for a certain frame of the audio signal, the fallback mode is to be performed or the regular mode. In this mode an actual performing of the full coding with the entropy coder is not necessary.In a preferred embodiment, the regular coding mode is a lossless audio coded that preferably relies on an integer transform such as an integer MDCT as described in detail in “Audio Coding Based on Integer Transforms”, Geiger, R., PhD thesis, November 2, 2007.

[0026] In this application, the operations for analysis windowing, folding and performing the timefrequency transform (by DCT-IV transform) are performed using lifting decompositions in order to avoid any quantization or rounding losses. Correspondingly, the decoder-side frequency-time transform operations, the subsequent unfolding operations and the processing operations comprising synthesis windowing and final overlapping and adding are also performed using lifting decompositions so that any quantization or rounding losses do not occur in the decoder as well. Specifically in this implementation, current block data and preceding block data are used in at least three lifting steps, wherein the lifting steps together implement windowing and overlap adding.

[0027] The inventive fallback mode is preferably performed using an integer processing such as the integer processing of the well-known integer MDCT, but without the DCT-IV time-frequency transform.

[0028] In such a situation, a fully lossless audio encoder relying on the modified discrete cosine transform is obtained that has a fixed maximum bitrate per frame, i.e., the one given by the raw encoder for encoding the windowed and folded audio data per block in case the bitrate provided by the other branch with an entropy encoder would be higher.

[0029] In a further embodiment, the inventive fallback mode can also be used together with a lossy encoder such as an encoder comprising a quantizer. This quantizer can be implemented as a controllable quantizer in order to control the quantization action in case of a bitrate requirement provided from the outside such as by a given maximum output bitrate. Furthermore, the quantizer can be controlled by a quantizer controller that receives control data from a psychoacoustic model in order to perform a psychoacoustically motivated quantization in case the output bitrate requirements are not too tough. In such a situation, the fallback mode can also rely on a quantizer that quantizes the windowed and folded audio data as the case may be.

[0030] Preferred embodiments of the present invention are subsequently discussed with respect to the accompanying drawings, in which:Fig. 1 illustrates a preferred embodiment of an audio encoder;

[0031] Fig. 2 illustrates a preferred embodiment of an audio encoder in which the raw encoder operates in a fallback mode;

[0032] Fig. 3 illustrates a bitstream, in which raw encoded blocks and entropy encoded blocks occur side by side;

[0033] Fig. 4 illustrates a preferred implementation of a standalone inventive decoder;

[0034] Fig. 5 illustrates a preferred implementation of the usage of the raw decoder branch as a fallback coding mode together with an entropy decoder / frequency-time converter branch;

[0035] Fig. 6 illustrates an audio decoder with a regular branch and the fallback branch together with a quantizer and audio coding tools;

[0036] Fig. 7 illustrates a decoder with a dequantizer that can be controlled by a side information or by external settings;

[0037] Fig. 8 illustrates an implementation of a two-branch encoder relying on the integer MDCT processing;

[0038] Fig. 9 illustrates a two-branch decoder relying on an inverse integer MDCT processing;

[0039] Fig. 10a illustrates a first notation with respect to windows overlapping by 50 percent and time domain samples extending from earlier samples to later samples;

[0040] Fig. 10b illustrates an explanation of certain variables and the definition off the analysis window portions a, b, c, d;

[0041] Fig. 11a illustrates an illustration of the combination / fold-in procedures and a generation of the DCT-IV inputs;Figs. 11 b, 11 c illustrate a further implementation of the combination / fold-in operations and the corresponding extension / fold-out operations for the integer and DCT;

[0042] Fig. 11 d illustrates a further procedure for the folding and unfolding operations with a notation different from Fig. 11a;

[0043] Fig. 12 illustrates an integer MDCT based lossless audio codec;

[0044] Fig. 13 illustrates a procedure for performing the raw coding of spectral lines;

[0045] Fig. 14 illustrates a decomposition of the audio codec of Fig. 12 into windowing / time- domain aliasing procedures and transform procedures for the encoder side and the decoder side; and

[0046] Fig. 15 illustrates the “short cut” of the inventive fallback mode within the integer MDCT based lossless audio codec of Fig. 12.

[0047] Fig. 1 illustrates an audio encoder for encoding an audio signal 16 comprising a sequence of blocks of audio samples. The audio signal is input into an analysis windower 100 for processing the sequence of blocks of audio samples using an analysis window. The analysis window has a length being larger than a length of a block of audio samples of the blocks of audio samples as is, for example, illustrated in Fig. 10a or 11d, where a block has N audio samples and the window has a length of 2N audio samples. The number of samples of the window does not necessarily have to be two times the number of samples of a block, but can also be larger or shorter but at least longer than the length of the block of audio samples. The analysis windower is configured to advance, in the sequence of blocks of audio samples, from a windowing operation to a following windowing operation by an amount of samples, i.e., the advance amount which is smaller than the length of the analysis window. This results in the situation that at least two audio samples are windowed by at least two subsequent windowing operations.

[0048] In the embodiment in Fig. 10a or 11d, a 50 percent overlap is illustrated, but the overlap can also be larger, so that at least two audio samples are windowed by more than two subsequent windowing operations.The processed block of the sequence of processed blocks is input into a raw encoder 200 for encoding the samples of the processed block 20 of the sequence of processed blocks to obtain raw-encoded data 30 for the processed block. This raw-encoded data for the processed block represent the encoded audio data of the encoder in Fig. 1.

[0049] Preferably, the analysis windower 100 not only performs the analysis windowing using the analysis window illustrated, for example, in Fig. 10a or 11d, but additionally performs the combination or folding or fold-in operation in order to reduce the amount of samples from a block at the input into analysis windower 100 compared to the number samples in a block at the output of the analysis windower, i.e., in a processed block 20. In the context of lapped transforms, such as the MDCT transform, the preferred operation is the folding operation, where two consecutive portions of windowed samples are combined / added together. The time domain aliasing introduced by this folding operation is, however, not critical, since it is removed by corresponding processing operations such as an unfolding operation and an overlap-add operation in a corresponding decoder as known in the art particularly for the prominent MDCT operation that relies on this functionality. Therefore, the analysis windower processes each block of the sequence of blocks by a specific windowing operation that comprises using the information on the analysis window on the one hand and the application of a time domain aliasing processing to obtain the sequence of processed blocks so that each processed block has a number of processed samples which is smaller than the number of window samples of the analysis window.

[0050] In a specific embodiment which is related to MDCT or integer MDCT processing, the number of processed samples in a processed block is equal to the number of audio samples within the sequence of blocks. Hence, the number of samples in a block of the sequence of blocks 16 and the number of samples in a block of the sequence of processed blocks 20 in Fig. 1 is the same.

[0051] In a further embodiment, the raw encoder is configured to apply a PCM (Pulse Code Modulation) coding to the samples of the processed block i.e., to the output of the analysis windower 100. However, other coding operations can be performed, and it is preferred that coding tools are used for the raw encoder which provide a maximum bitrate fixed for all classes of signals. Another example of such a raw encoder is a grouped PCM coder which codes tuples or groups of individual samples of a processed block, for example.Fig. 2 illustrates an audio encoder which has two branches. One branch includes the raw encoder 200 and the other branch includes the time-frequency converter 300 and the entropy encoder 400 which is indicated as a lossless entropy encoder such as an arithmetic entropy encoder or a Huffman encoder or any other entropy encoder known to those skilled in the art.

[0052] Both coding branches preferably use the same analysis windower 100 that not only performs analysis windowing, but also the fold-in or folding operation. The folded data for a block, i.e., the processed block is either forwarded to the raw encoder 200 or to a timefrequency converter 300 by a switch 430 controlled by a controller 435. The controller additionally outputs control side information in an embodiment and this control side information is used by an output interface 440 to be included into the sequence of processed blocks 30 generated by the output interface 440.

[0053] The controller 435 can rely on a closed loop. In this case, the controller performs the coding operation and determines the output bitrate consumption of the entropy encoder 400 and in case this bit consumption is higher than the threshold bit consumption provided by the raw encoder 200, then the processed block is forwarded to the raw encoder 200.

[0054] When, however, it is determined that the bitrate provided by the entropy encoder 400 is lower than the bitrate incurred by the raw encoder 200, then the processed block generated by the analysis windower 100 is forwarded to the time-frequency converter 300. In an embodiment, which relies on an MDCT, the time-frequency converter is configured to perform a DCT- IV operation. However, in other embodiments, other algorithms different from the DCT-IV operation can be performed, such as DST operations or DCT-I, DCT-II, or DCT-III operations.

[0055] The time-frequency converter 300 generates spectral samples that are input into a general spectral coder 400 that can rely on a lossy coding portion 410 and a lossless coding portion 400, i.e., the entropy encoder. In an embodiment, the lossy block 410 is missing and only the lossless block 400 is there. However, in other embodiments, the lossy block 410 is used and typically comprises a quantizer that can be controlled by a psychoacoustic model or by other control data such as a maximum bitrate control.

[0056] Fig. 3 illustrates an exemplary sequence of processed blocks comprising four processed blocks 10, 11, 12, and 13, where each block has a side information SI. The payload data inthe raw encoded block are the PCM-coded time domain windowed and preferably folded audio samples. In the entropy encoded blocks such as blocks 11 and 13, the payload data consists of entropy-encoded spectral data derived from the spectral samples output by the time-frequency converter 300. Fig. 3 illustrates a situation, where a current block 12 is a raw encoded block, where the preceding block 11 is an entropy encoded block and the next block is again an entropy encoded block 13, while the first block in the illustration in Fig. 3 is again a raw-encoded block.

[0057] Thus, Fig. 3 illustrates switching from a raw-encoded block to an entropy-encoded block, back to the raw-encoded block and again to the entropy-encoded block. However, this switching is easily possible, since the overlap adder illustrated at 700 in Fig. 4 always receives synthesis-windowed data irrespective of whether the decoding to obtain this data was a raw decoding or an entropy decoding. The overlap-add operation to finally obtain a block of decoded samples is always the same irrespective of the selected coding branch.

[0058] Fig. 4 illustrates an audio decoder for decoding encoded audio data 30 comprising raw-encoded data. The audio decoder comprises a raw decoder 500 for decoding the raw-encoded data to obtain a processed block. The processed block is input into a processor 650 comprising a synthesis windower 600 that outputs a synthesis windowed and optionally folded-out or unfolded block of samples using a synthesis window in order to generate a synthesized block. The synthesized block at the output of the synthesis windower 600 is combined with a preceding synthesized block 601 by an overlap-adder 700 also comprised by the processor 650 which finally outputs the block 18 of decoded samples.

[0059] It is to be mentioned that the synthesis windower 600 generates a full synthesized block in one embodiment, and the overlap-adder combines two subsequent blocks in the overlap range. In other embodiments, it is not necessary to calculate a full synthesized block. Instead, only a portion of a synthesized block or a portion of a preceding synthesized block in an overlap range is calculated and these portions are added / combined without explicitly storing corresponding full blocks.

[0060] In further embodiments, the functionalities of the windower 600 and the overlap adder 700 are jointly performed by the processor 650 without explicitly calculating synthesized or windowed portions. Such an implementation can be the implementation illustrated in Fig. 11c, which can be turned into a lifting implementation, as described in the prior publication “Audio Coding Based on Integer Transforms” by R. Geiger, wherein the weights applied in the liftingsteps, the input into the lifting steps and the sequence of lifting steps together result in the decoded audio samples as are obtained by a separate synthesis windowing and a subsequent overlap adding. In otherwords, the weights used into the lifting steps jointly implement the synthesis windowing and overlap add operation.

[0061] Fig. 5 illustrates a situation, in which the decoder of Fig. 4 is integrated. Particularly, an upper decoding branch comprises the raw decoder 500, the synthesis windower 600 being connected to the windower 600 and the overlap adder 700. The upper coding branch comprises an entropy decoder 800, a frequency-time converter 810 and the subsequent synthesis windower 600 that feeds the overlap-adder 700. Furthermore, the raw decoder 500 on the one hand and the entropy decoder 800 on the other hand are fed by an input interface 820 which additionally forwards the side information from Fig. 3 to the encoded data parser 830 so that the block 830 controls the input interface 820 or either the raw decoder or the entropy decoder to receive and process a corresponding block of data in the encoded audio signal 30.

[0062] Although the synthesis windower 600 is illustrated in Fig. 5 as occurring two times, i.e. , in each branch, it is to be emphasized that in an embodiment, only a single synthesis windower is there and the single synthesis windower forwards its output to a single input of the overlap adder 700 in order to output the following block of decoded samples such as an output of block 11 or of block 13 of Fig. 3, when the entropy decoder 800 and the frequency-time converter 810 were active.

[0063] In an embodiment, the synthesis windower 600 only performs synthesis windowing rather than unfolding or the extension operation when the raw decoded block output by the raw decoder does not contain any time domain aliasing. However, in the preferred embodiment, the synthesis windowing operation performed by the synthesis windower 600 comprises using information on the synthesis window and the application of the unfolding or extension operation to obtain the synthesized block which has a number of samples being greater than the number of samples of the processed block output by the raw decoder 500.

[0064] In the other embodiment, where the functionalities of the windower 600 and the overlap adder 700 are jointly performed by the processor 650 without explicitly calculating synthesized or windowed portions, the overlap adder functionality indicated by block 700 occurs in the processor indicated by dotted lines in Fig. 5 as well as in Figs. 4, 7, 9. Such an implementation can be the implementation illustrated in Fig. 11c, which can be turned intoa lifting implementation, as described in the prior publication “Audio Coding Based on Integer Transforms” by R. Geiger, or any other implementation jointly realizing windowing and overlap adding.

[0065] Fig. 6 illustrates a preferred implementation of an encoder similar to the illustration of Fig.

[0066] 2. However, the functionality of the analysis windower 100 is illustrated as an analysis windower for performing analysis windowing with at least portions of a block illustrated at 110 and a subsequent folder 124 performing the folding or fold-in or, generally, combination information in order to reduce the number of time domain audio samples. Furthermore, the time-frequency converter 300 is illustrated as a DCT-IV processor. Additionally, an optional quantizer 410a generally illustrated at 410 in Fig. 2 is illustrated which is controlled by a quantizer controller 410c. Furthermore, certain audio coding tools 410b are illustrated such as tools occurring in the AAC codec, the MP3 codec, the EVS codec, the LISAC codec or the LC3 codec or any other codec. Such audio coding tools can comprise temporal noise shaping, time domain prediction, LPC transform coding, noise filling, bandwidth extension or any other corresponding audio coding tools either individually or in combination. The output of the audio coding tools block 410b is then input into the entropy encoder 400 being an arithmetic encoder or a Huffman encoder or any other lossless encoder.

[0067] The other coding branch is illustrated as comprising several options. One option is that the output of the folder 120 is directly forwarded to the raw encoder 200 and the raw encoded block is input into the output interface 440. However, another option in another embodiment is to quantize the output of block 120 using the quantizer 410a. This means that in the corresponding embodiment the time-frequency converter 300 is bypassed so that the folder output directly enters the quantizer 410a, and the quantized data is input into the raw encoder 200. Such a quantization could, for example, be a deletion of a least significant bit from each windowed and folded sample of the processed block, for example. This quantizer mode 410a can also be controlled by the quantizer controller 410c depending on certain quantization targets such as a maximum output bitrate also for the fallback mode. Since the maximum output bitrate for the fallback mode is always predetermined, the quantization of the output of the folder 120 results in a certain produced maximum bitrate. However, when the quantizer 410a is also applied to the folder 120 output then the fallback mode is not lossless anymore.

[0068] Fig. 7 illustrates a corresponding decoder that consists of the branch with the raw decoder 500 and the branch with the entropy and decoder 800. In contrast to the illustration in Fig.5, Fig. 7 illustrates only a single processor 650 indicated in dotted lines implementing the functionality of the synthesis windower 600 and the overlap adder 700 for both branches which feeds the overlap adder 700.

[0069] Furthermore, optional decoding tools 410b and an optional dequantizer 410a are illustrated that can be controlled using quantizer controller data provided as side information from the input interface or implicitly provided or set by an operator of the decoder.

[0070] In the preferred embodiment, the fold-out or unfolding operation is performed by the synthesis window before the synthesized block is obtained which is then input into the overlap adder 700 together with the preceding synthesized block 601 illustrated in Fig. 4.

[0071] Fig. 8 illustrates a preferred implementation of the present invention in the context of an integer MDCT operation where any lossy encoder tools do not occur. Instead, the analysis window application and folding 100a is performed using integer operations only which are preferably decomposed into lifting steps.

[0072] The result of block 100a is input into the switch 430 which either forwards the corresponding block to a DCT-IV 300a which is implemented by means of integer operations in lifting steps and the result is forwarded to the entropy encoder 400 which is lossless and, therefore, does not incur any distortions anymore. In case of the fallback mode, the switch 430 forwards the output of block 100a to the raw encoder 200.

[0073] A corresponding decoding implementation is illustrated in Fig. 9 showing the inventive two-branch decoder in the context of an inverse integer MDCT. Entropy decoded data output by block 800 are input into a frequency-time converter 810a which is implemented by integer operations only. Preferably, block 810a implements the same operation as block 300a.

[0074] Similarly, the raw decoder block 500 forwards its output into the processor 650 implementing the functionality of the synthesis windower 600a performing the fold-out and synthesis windowing operation again with integer operations in lifting steps so that any losses do not occur. In an embodiment, the output of block 600a is input into the overlap adder 700 which also does not incur any losses, since the input into block 700 only consists of integer data. In the integer implementation using lifting decompositions, the functionalities of the blocks 600a and 700 are jointly performed.Although it has been stated in Fig. 8 and Fig. 9 that certain blocks perform integer operations, it is to be emphasized that these blocks typically perform operations that result in noninteger data. Then, this non-integer data is rounded using a rounding function. However, due to the fact that the lifting steps are performed one after the other, the rounding error can always be recovered without any loss so that, although rounding functions are performed, the rounding errors are always avoided in the final result. Thus, the integer approximation of a lifting step can always be inverted without introducing any error. Applying such an approximation to each of, for example, three lifting steps, one can get an integer approximation of, for example, a given Givens rotation, and this rounded rotation can be reverted without introducing any error by applying the inverse rounded lifting steps in reverse order using the same rounding function. If the rounding function r is odd symmetric, the inverse rounded rotation is identical to the rounded rotation with the negative angle.

[0075] Fig. 10a illustrates one implementation of a notation matching with the encoder-side illustration in Fig. 11a for the purpose of illustrating the folding operation performed by the analysis windower 100 in a preferred implementation. In Fig. 10a, two overlapping windows are illustrated. One window extends from -N to +N and the second window extends between 0 and 2N. Furthermore, several parameters of Fig. 10a and variables are explained in Fig.

[0076] 10b. Particularly, reference is made to the reverse notation for n’ in the last line of Fig. 10b. Additionally, the Fig. 10a to Fig. 11a illustration refers to the application of the MDCT and a 50% overlap and the window is a symmetric window. However, similar calculations can also be performed for symmetric windows and for windows having different overlaps.

[0077] Fig. 11a illustrates which portions of the window are applied to which DCT operation. Fig.

[0078] 11a illustrates two DCT-IV operations 811 and 812. Particularly, the DCT-IV operation 812 illustrates the DCT-IV operation for the window extending from 0 to 2N in Fig. 10a. The folding operation is illustrated by the functionality of the adder 103a that combines the first portion b and the second portion a of the first window half extending from 0 to N so that, at the output of the adder 103a, there are only N / 2 values.

[0079] The adder 104a combines the samples of the third window portion d and the last window portion c in order to once again obtain N / 2 samples at the output of the combiner or adder 104a. Hence, it becomes clear how the window extending from 0 to 2N in Fig. 2a is applied and the result is folded-in to obtain the DCT-IV input. The DCT-IV operation 811 receives the data of the first window extending between -N and +N but, again, the first two portions a, b are combined and the third and fourth portions are combined as well.Another illustration for the combination or fold-in on the one hand and the extension or fold-out on the other hand is illustrated in Figs. 11b and 11c as shown in the prior publication “Audio Coding Based on Integer Transforms” by R. Geiger. In this embodiment, the functionalities of the windower 600 and the overlap adder 700 are jointly performed by the processor 650 without explicitly calculating synthesized or windowed portions. Such an implementation can be the implementation illustrated in Fig. 11c, which can be turned into a lifting implementation, where the weights applied in the lifting steps implementing the arrows to the right of the DCT-IV blocks, the input into the lifting steps and the sequence of lifting steps together result in the decoded audio samples as are obtained by a separate synthesis windowing and a subsequent overlap adding. In other words, the weights used into the lifting steps jointly implement the synthesis windowing and overlap add operation.

[0080] Fig. 11d illustrates a further implementation with a different notation in order to explain the combination or folding or fold-in on the encoder-side and the fold-out or unfolding or extension operation on the decoder-side.

[0081] Fig.11 d illustrates a window 70, which has an increasing portion to the left and a decreasing portion to the right, where one can divide this window into four portions: a, b, c, and d. Window 70 has, as can be seen from the figure only aliasing portions in the 50% overlap / add situation illustrated. Specifically, the first portion having samples from zero to N corresponds to the second portions of a preceding window 69, and the second half extending between sample N and sample 2N of window 70 is overlapped with the first portion of window 71, which is in the illustrated embodiment window i+1 , while window 70 is window i.

[0082] The MDCT operation can be seen as the cascading of the folding operation and a subsequent transform operation and, specifically, a subsequent DCT operation, where the DCT of type-IV (DCT-IV) is applied. Specifically, the folding operation is obtained by calculating the first portion N / 2 of the folding block as -CR-d, and calculating the second portion of N / 2 samples of the folding output as a-bR, where R is the reverse operator. Thus, the folding operation results in N output values while 2N input values are received.

[0083] A corresponding unfolding operation on the decoder-side is illustrated, in equation form, in Fig. 11D as well. Generally, an MDCT operation on (a,b,c,d) results in exactly the same output values as the DCT-IV of (-CR-d, a-bp) as indicated in Fig. 4A. Correspondingly, andusing the unfolding operation, an IMDCT operation results in the output of the unfolding operation applied to the output of a DCT-IV inverse transform.

[0084] Therefore, time aliasing is introduced by performing a folding operation on the encoderside. Then, the result of the folding operation is transformed into the frequency domain using a DCT-IV block transform requiring N input values. On the decoder-side, N input values are transformed back into the time domain using a DCT-IV-1operation, and the output of this inverse transform operation is thus changed into an unfolding operation to obtain 2N output values which, however, are aliased output values.

[0085] In order to remove the aliasing which has been introduced by the folding operation and which is still there subsequent to the unfolding operation, the overlap / add operation by the overlap-adder 700 of Fig. 4 is required.

[0086] Therefore, when the result of the unfolding operation is added with the previous IMDCT result in the overlapping half, the reversed terms are canceled and one obtains simply, for example, b and d, thus recovering the original data.

[0087] In order to obtain a TDAC for the windowed MDCT, a requirement exists, which is known as “Princen-Bradley” condition, which means that the window coefficients raised to2for the corresponding samples which are combined in the time domain aliasing canceller as to result in unity (1 ) for each sample. Fig. 11 D illustrates the window sequence as, for example, applied in the AAC-MDCT for long windows or short windows,

[0088] Thus, in the context of Fig. 11d and in the general MDCT operation, the MDCT is considered with 2 N inputs and N outputs. Particularly, the input is divided into four blocks a, b, c, d each of size N / 2. If one shifts these to the right by N / 2 (from the +N / 2 term in the MDCT definition) then (b, c, d) extend past the end of the N DCT-IV input. Hence, they are folded back according to the boundary conditions of the MDCT formulas, i.e., to the alternating even / odd boundary conditions, i.e., even at the left boundary, odd at the right boundary, and so on.

[0089] Similarly, the IMDCT formula also incurs the DCT-IV (which is its own inverse) where the output is extended or unfolded via the boundary conditions to a length 2 N and shifted back to the left by N / 2. The inverse DCT-IV would simply give back the inputs (-CR-d, a-bp) fromabove. When this is extended via the boundary conditions and shifted, one obtains the unfolding result illustrated in Fig. 11 d.

[0090] Particularly, the result of the folding operation in Fig. 11d is either input into the raw encoder 200 or the time-frequency converter 300 of Fig. 2, and the result of the unfolding or extension operation in Fig. 11d is input, in the decoder, in the overlap adder 700. Due to this overlap-add operation, a continuous cross-fade from one block to the other due to the applied synthesis window is obtained so that any blocking artifacts or so do not occur and have not to be addressed by additional procedures. Instead, irrespective of whether the data input into the synthesis windower 600 stem from the raw decoder 500 or the frequencytime converter 810, the procedure performed by the synthesis windower, i.e., the unfolding and application of the synthesis window, is always the same.

[0091] Fig. 12 illustrates a general representation of an integer MDCT based lossless audio codec relying on an MDCT 50 and a subsequent entropy coding 400 of the MDCT data and a subsequent entropy decoding 800 feeding an inverse MDCT block 52 in order to obtain the decoded data 18. The block 50 in the encoder comprises the blocks windower 100 and time-frequency converter 300 for the encoder in other figures. The block 52 represents the functionality of the frequency-time converter 810 and the functionality of the processor 650 in the decoder in other figures. The processor 650 comprises the synthesis windowing 600 and overlap adding 700 functionalities. The following problems can occur in such a situation. Particularly, the question is what is the maximum frame size, which buffers have to be allocated in the decoder and which upper limits of the bitrate can be defined.

[0092] Furthermore, high entropy signals exist and for such signals the bitstream is larger than the raw input signal. Hence, one option would be to perform a raw coding of the MDCT output data, i.e., to simply skip entropy coding. If the entropy coded frame size is larger than the fallback frame size then the fallback can be used. Although this procedure solves both problems, i.e., the maximum frame size is the one using the fallback mode and for high entropy signals which cause large frame sizes, these are also kept to the frame size with the fallback mode, this procedure is not very useful for lapped transform based lossless audio codecs. For non-lapped transform based lossless audio codecs, one could use a raw coding of the PCM signal. This, however, is also not useful since lapped transform based audio codecs such as the ones that rely on MDCT or other related procedures have to rely on the timedomain aliasing cancellation (TDAC). Regarding the option of raw coding spectral lines, itis not always clear how large the spectral lines can get due to the energy compaction characteristic of a typical time-frequency transform. Therefore, for a worst case assumption one must start from the point that all the energy in one spectral line is concentrated which would result in a large overhead for the fallback mode. Therefore, the procedure illustrated in Fig.

[0093] 13 is not advantageous.

[0094] Fig. 14 illustrates a decomposition of the audio codec of Fig. 12 into a windowing procedure 100 and a transform procedure 300 together forming the functionality of block 50 for the encoder side and a transform procedure 810 and a windowing / time-domain aliasing procedure 650 forming the functionality of block 52 for the decoder side.

[0095] However, the procedure illustrated in Fig. 15 applying a raw encoder 200 on the encoderside and a corresponding raw-decoder 500 on the decoder-side to the windowed and timedomain aliased data as a “shortcut” illustrated by the big arrow in Fig. 15 avoids the energy compaction procedure by means of the time-frequency transform and nevertheless guarantees a minimum bit consumption for a certain frame.

[0096] It is to be mentioned here that all alternatives or aspects as discussed before and all aspects as defined by independent claims in the following claims can be used individually, i.e. , without any other alternative or object than the contemplated alternative, object or independent claim. However, in other embodiments, two or more of the alternatives or the aspects or the independent claims can be combined with each other and, in other embodiments, all aspects, or alternatives and all independent claims can be combined to each other.

[0097] An inventively encoded signal can be stored on a digital storage medium ora non-transitory storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0098] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.

[0099] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a CD, a ROM, a PROM, anEPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed.

[0100] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed. Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier. Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier or a non-transitory storage medium. In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet. A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein. In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.

[0101] The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.

Claims

Claims1. Audio encoder for encoding an audio signal comprising a sequence of blocks of audio samples, comprising:an analysis windower (100) for processing the sequence of blocks of audio samples using an analysis window to obtain a processed block of a sequence of processed blocks, wherein the analysis window has a length being larger than a length of a block of audio samples of the blocks of audio samples,wherein the analysis windower (100) is configured to advance, in the sequence of blocks of audio samples, from a windowing operation to a following windowing operation by an amount of samples being smaller than a length of the analysis window so that each audio sample is windowed by at least two subsequent windowing operations; anda raw encoder (200) for raw encoding the samples of the processed block of the sequence of processed blocks to obtain raw-encoded data for the processed block.

2. Audio encoder of claim 1, wherein the analysis windower (100) is configured to process each block of the sequence of blocks by the windowing operation, the windowing operation comprising using information on the analysis window and applying a combination processing to windowed time domain samples to obtain the sequence of processed blocks, wherein the processed block comprises a time domain aliasing and has a number of processed samples being smaller than the number of window samples of the analysis window.

3. Audio encoder of claim 2, wherein the number of processed samples in a processed block is equal to the number of audio samples of a block in the sequence of blocks of audio samples, or wherein the combination processing is a folding or fold-in processing.

4. Audio encoder of one of the preceding claims, wherein the raw encoder (200) is configured to apply a PCM (pulse code modulation) coding to the samples of the processed block.

5. Audio encoder of one of the preceding claims, comprising:a time-frequency converter (300) for converting a further processed block of the sequence of processed blocks into a spectral representation of the processed block; andan entropy encoder (400) for entropy-encoding the spectral representation of the further processed block to obtain entropy-encoded data for the spectral representation of the further processed block.

6. Audio encoder of claim 5, configured for processing the block of the sequence of blocks with the analysis windower (100) and the raw encoder (200) to obtain raw- encoded data for the processed block and for processing the further block of the sequence of blocks with the analysis windower (100), the time-frequency converter (300), and the entropy encoder (400) to obtain entropy-encoded data for the further processed block, wherein the block follows the further block in the sequence of blocks, orconfigured for processing the further block of the sequence of blocks with the analysis windower (100), the time-frequency converter (300) and the entropy encoder (400) and for processing the block of the sequence of blocks with the analysis windower (100) and the raw encoder (200), wherein the further block follows the block in the sequence of blocks.

7. Audio encoder of claim 6, comprising a controller for feeding an output of the analysis windower (100) to either an input of the raw encoder (200) or an input of the timefrequency converter (300).

8. Audio encoder of one of claims 5 to 7, wherein the time-frequency converter (300) is configured to perform a DCT-IV operation.

9. Audio encoder of one of claims 6 to 8, comprising an encoded data former (440) for concatenating the raw-encoded data for the processed block and the entropy-encoded data for the spectral representation of the further processed block to obtain encoded audio data.

10. Audio encoder of claim 9, wherein the encoded data former (440) is configured to generate the encoded audio data so that the block includes information on whether the block is the raw-encoded block or the entropy-encoded block.

11. Audio encoder of one of claims 9 or 10, wherein the encoded data former (440) is configured to generate a bitstream as the encoded audio data comprising a first sequence of bits for the raw-encoded data and a second sequence of bits for the entropy-encoded data, wherein the first sequence of bits follows the second sequence of bits or vice-versa.

12. Audio encoder of one of claims 6 to 11, wherein the audio encoder is configured to determine (435) whether the block of the sequence of blocks is to be processed either using the raw encoder (200) or using the entropy encoder (400), based on an estimation of a number of information units necessary for encoding the block using the entropy encoder (400) so that the encoder requiring a lower number of bits is determined.

13. Audio encoder of one of claims 1 to 12, wherein the analysis windower (100) or the analysis windower (100) and the time-frequency converter (300) are configured to apply integer operations.

14. Audio encoder of claim 13, wherein the integer operations represent an integer MDCT (modified discrete cosine transform) operation.

15. Audio encoder of one of the preceding claims, further comprising a controllable quantizer (410a) for quantizing processed blocks of the sequence of processed blocks.

16. Audio encoder of claim 15, wherein the controllable quantizer (410a) is controllable so that a quantization is performed in such a way that an allowed amount of information units for a corresponding block or a plurality of corresponding blocks is maintained.

17. Audio encoder of claim 15 or 16, wherein the controllable quantizer (410a) is only operated for the further block of the sequence of blocks which is to be encoded by the entropy encoder (400).

18. Audio encoder of one of claims 15 to 17, wherein the audio encoder is configured to apply an integer MDCT coding with the entropy encoder (400) or an integer windowing using the analysis windower (100) or, when an allowed amount of information units is fulfilled, a quantization using a quantizer (410, 410a) and an encoding using the entropy encoder (400).

19. Audio decoder for decoding encoded audio data comprising raw-encoded data, the audio decoder comprising:a raw decoder (500) for decoding the raw-encoded data to obtain a processed block; anda processor (650) for processing the processed block using a synthesis window and for overlapping and adding current block data and preceding block data to obtain a block of decoded audio samples.

20. Audio decoder of claim 19, wherein the encoded audio data comprise entropy-encoded data, the audio decoder comprising:an entropy decoder (800) for entropy-decoding the entropy-encoded data to obtain a spectral representation for a preceding block; anda frequency-time converter (810) for converting the spectral representation into a time representation to obtain a time-domain representation for the preceding block,wherein the processor (650) is configured to process the time-domain representation for the preceding block to obtain at least a portion of the preceding block data.

21. Audio decoder of claim 19, wherein the encoded audio data comprise raw-encoded data for a preceding block, wherein the raw decoder (500) is configured to decode the raw-encoded data for the preceding block to obtain a preceding processed block, andwherein the processor (650) is configured to process the preceding processed block to obtain at least a portion of the preceding block data.

22. Audio decoder of claim 20 or 21 , wherein the audio decoder comprises entropy-encoded data for a following block, the audio decoder comprising:an entropy decoder (800) for entropy-decoding the entropy-encoded data for the following block to obtain a spectral representation for the following block;a frequency-time converter (810) for converting the spectral representation for the following block into a time representation to obtain a time-domain representation for the following block,wherein the processor (650) is configured to process the time-domain representation to obtain at least a portion of following block data, and to overlap and add the portion of the following block data and the portion of the current block data to obtain at least a portion of a following block of decoded audio samples.

23. Audio decoder of one of claims 19 to 22, wherein the processor (650) comprises a synthesis windower (600) being configured to process the processed block by a synthesis windowing operation, the synthesis windowing operation comprising using information on the synthesis window and applying an extension operation to obtain the current block data, the current block data having a number of samples being greater than the number of samples of the processed block.

24. Audio decoder of one of claims 19 to 23, wherein the synthesis window has a number of synthesis window samples, and wherein the number of synthesis window samples is greater than the number of samples of the block of decoded audio samples, or wherein the extension operation is an unfolding operation or a folding out operation.

25. Audio decoder of one claims 19 to 24, wherein the processor (650) is configured to overlap at least the portion of the preceding block data and at least the portion of the current block data by an overlap distance being greater than ten samples or being equal to or greater than a number of samples of the block of decoded audio samples.

26. Audio decoder of one of claims 19 to 25, wherein the raw decoder (500) is configured to apply a PCM decoding to obtain the samples of the processed block.

27. Audio decoder of one of claims 20 to 26, wherein the frequency-time converter (810) is configured to apply a DCT-IV operation (discrete cosine transform).

28. Audio decoder of one of claims 20 to 27, comprising an encoded data parser (830) for determining, from the encoded audio data, whether a block of encoded audio data is to be decoded using the raw decoder (500) or the entropy decoder (800), and wherein the audio decoder is configured to forward the block of encoded audio data to the entropy decoder (800) or the raw decoder (500) depending on a determination result.

29. Audio decoder of claim 28, wherein the encoded data parser (830) is configured to determine a state of an information unit associated to the block in accordance with an encoded data syntax.

30. Audio decoder of one of claims 28 or 29, wherein the encoded data is a bitstream having a sequence of blocks, wherein the encoded data parser (830) is configured to parse the sequence of blocks using a bitstream syntax.

31. Audio decoder of one of claims 19 to 30, wherein the processor (650) or the processor (650) and the frequency-time converter (810) are configured to perform integer operations.

32. Audio decoder of claim 31 , wherein the frequency-time converter (810) and the processor (650) are configured to apply an integer inverse MDCT operation using lifting steps using the current block data and the preceding block data or using the following block data and the current block data.

33. Audio decoder of one of claims 19 to 31, wherein the processor (650) comprises a synthesis windower (600) for processing the processed block using the synthesis window to obtain, as the current block data, synthesized current block data, and an overlap-adder (700) for overlapping and adding the synthesized current block data and, as the preceding block data, synthesized preceding block data to obtain a block of decoded audio samples.

34. Method of encoding an audio signal comprising a sequence of blocks of audio samples, the method comprising:processing the sequence of blocks of audio samples using an analysis window to obtain a processed block of a sequence of processed blocks, wherein the analysis window has a length being larger than a length of a block of audio samples of the blocks of audio samples,wherein the processing comprises advancing, in the sequence of blocks of audio samples, from a windowing operation to a following windowing operation by an amount of samples being smaller than a length of the analysis window so that each audio sample is windowed by at least two subsequent windowing operations; andraw encoding the samples of the processed block of the sequence of processed blocks to obtain raw-encoded data for the processed block.

35. Method of decoding encoded audio data comprising raw-encoded data, the method comprising:raw decoding the raw-encoded data to obtain a processed block; andprocessing the processed block using a synthesis window, and overlapping and adding current block data and preceding block data to obtain a block of decoded audio samples.

36. Computer program for performing, when running on a computer or a processor, the method of claim 34 or the method of claim 35.

37. Encoded audio signal comprising:raw-encoded data (10, 12) for a block of a sequence of blocks and first side information indicating that the raw-encoded data are raw-encoded; andentropy-encoded data (11, 13) for a spectral representation of a further block of the sequence of blocks and second side information indicating that the entropy-encoded data are entropy-encoded.