Frequency domain audio coding supporting transform length switching

The described frequency-domain audio codec supports additional transform lengths by interleaving coefficients, addressing quality issues in transient signals at low bit rates and ensuring compatibility with existing codecs.

JP7799005B2Active Publication Date: 2026-01-14FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024190397
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2013-10-18
Filing Date
2024-10-30
Publication Date
2026-01-14
Estimated Expiration
2034-07-15

AI Technical Summary

Technical Problem

Existing frequency domain audio codecs struggle to provide satisfactory quality for audio signals with significant transients, such as rain or applause, at low bit rates due to temporal blurring or increased data overhead, and introducing new codecs is challenging due to market adoption of well-known codecs.

Method used

A frequency-domain audio codec concept that supports additional transform lengths by interleaving frequency-domain coefficients, allowing existing codecs to switch between transform lengths without losing backward compatibility, using entropy coding and inverse transforms to maintain quality.

Benefits of technology

Enables existing codecs to achieve better quality with minimal coding efficiency penalty while maintaining compatibility, ensuring reasonable playback even for legacy decoders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799005000028
    Figure 0007799005000028
  • Figure 0007799005000029
    Figure 0007799005000029
  • Figure 0007799005000030
    Figure 0007799005000030
Patent Text Reader

Abstract

To provide a frequency domain audio codec extended in a downward compatible manner.SOLUTION: A frequency domain audio decoder 10 includes: a frequency domain (FD) coefficient extractor 12; a scaling coefficient extractor 14; an inverter 16; and a coupler 18. In input, the FD coefficient extractor and the scaling coefficient extractor access an incoming data stream 20. Outputs of the FD coefficient extractor and the scaling coefficient extractor are connected to inputs of the inverter, and output of the inverter is connected to input of the coupler. The coupler outputs a reconstructed audio signal 22. The FD coefficient extractor extracts FD coefficients 24 of a frame 26 of the audio signal from the data stream. The FD coefficient describes a spectrum of the audio signal in respective frames at various spectral-temporal resolutions and collectively indicates a spectrogram 28 of the audio signal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to frequency domain audio coding supporting transform length switching. [Background technology]

[0002] Modern frequency domain speech / audio coding systems, such as the Opus / Celt codecs of the IETF [1], MPEG-4(HE-)AAC [2], or especially MPEG-D xHE-AAC(USAC) [3], provide the means to encode an audio frame using either one long transform, i.e., a long block, or eight consecutive short transforms, i.e., short blocks, depending on the temporal stability of the signal.

[0003] For certain audio signals, such as rain or the applause of a large audience, neither long-block nor short-block coding leads to satisfactory quality at low bit rates. This can be explained by the significant density of transients in such recordings: coding with only long blocks can cause temporal blurring of frequent, audible coding errors, also known as pre-echoes, while coding with only short blocks is generally inefficient due to the increased data overhead that results in spectral holes.

[0004] It would therefore be desirable to have at hand a frequency-domain audio coding concept that is also suitable for the kind of audio signals just outlined. Of course, it is feasible to build new frequency-domain audio codecs that support, among other things, switching between a set of transform lengths that encompass specific desired transform lengths that are suitable for specific kinds of audio signals. However, introducing new frequency-domain audio codecs that are adopted by the market is not an easy task. Well-known codecs are already available and frequently used. It would therefore be desirable to have a concept that allows existing frequency-domain audio codecs to be extended to additionally support new desired transform lengths, but still maintain backward compatibility with existing coders and decoders. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] [1] Internet Engineering Task Force (IETF), RFC 6716, “Definition of the Opus Audio Codec,” Proposed Standard, Sep. 2012. Available online at http: / / tools.ietf.org / html / rfc6716. [Non-patent document 2] [2] International Organization for Standardization, ISO / IEC 14496-3:2009, “Information Technology - Coding of audio-visual objects - Part 3: Audio,” Geneva, Switzerland, Aug. 2009. [Non-patent document 3] [3] M. Neuendorf et al., “MPEG Unified Speech and Audio Coding - The ISO / MPEG Standard for High-Efficiency Audio Coding of All Content Types,” in Proc. 132nd Convention of the AES, Budapest, Hungary, Apr. 2012. Also to appear in the Journal of the AES, 2013. [Non-patent document 4] [4] International Organization for Standardization, ISO / IEC 23003-3:2012, “Information Technology - MPEG audio - Part 3: Unified speech and audio coding,” Geneva, Jan. 2012. [Non-Patent Document 5] [5] JDJohnston and AJFerreira, “Sum-Difference Stereo Transform Coding”, in Proc. IEEE ICASSP-92, Vol. 2, March 1992. [Non-patent document 6] [6] N.Rettelbach, et al., European Patent EP2304719A1, “Audio Encoder, Audio Decoder, Methods for Encoding and Decoding an Audio Signal, Audio Stream and Computer Program”, April 2011. Summary of the Invention [Problem to be solved by the invention]

[0006] It is therefore an object of the present invention to provide a concept that allows existing frequency domain audio codecs to be extended in a backward compatible manner to support additional transform lengths, such that they switch between transform lengths including this new transform length. [Means for solving the problem]

[0007] This object is achieved by the subject matter of the independent claims attached hereto.

[0008] The present invention is based on the observation that it is possible to provide a frequency-domain audio codec capable of additionally supporting specific transform lengths in a backward-compatible manner when the frequency-domain coefficients of each frame are transmitted interleaved, regardless of the signaling that signals for each frame which transform length is actually applied, and further when the frequency-domain coefficient extraction and scale factor extraction operate independently of that signaling. This approach allows older frequency-domain audio coders / decoders that do not support this signaling to nevertheless operate with error-free playback and reasonable quality. At the same time, frequency-domain audio coders / decoders that support switching to / from the additionally supported transform lengths achieve even better quality despite being backward-compatible. As far as the coding efficiency penalty due to the frequency-domain coefficients being coded transparently to older decoders is concerned, this is of a relatively minor nature due to the interleaving.

[0009] Advantageous embodiments of the present application are the subject matter of the dependent claims. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a schematic block diagram of a frequency domain audio decoder according to one embodiment; [Figure 2] FIG. 2 is a schematic diagram illustrating the function of the inverter of FIG. 1; [Figure 3]3 is a schematic diagram illustrating a possible upstream displacement of the inverse TNS filtering process of FIG. 2 according to one embodiment. [Figure 4] FIG. 10 illustrates window selection possibilities when using transform splitting of long stop-start windows in USAC, according to one embodiment. [Figure 5] 1 is a block diagram of a frequency domain audio encoder according to one embodiment; DETAILED DESCRIPTION OF THE INVENTION

[0011] In particular, preferred embodiments of the present application are described below with reference to the drawings.

[0012] Figure 1 shows a frequency domain audio decoder supporting transform length switching according to one embodiment of the present application. The frequency domain audio decoder of Figure 1 is generally indicated using reference numeral 10 and comprises a frequency domain coefficient extractor 12, a scaling coefficient extractor 14, an inverse transformer 16 and a combiner 18. At their inputs, the frequency domain coefficient extractor 12 and the scaling coefficient extractor 14 have access to an incoming data stream 20. The outputs of the frequency domain coefficient extractor 12 and the scaling coefficient extractor 14 are connected to respective inputs of the inverse transformer 16. The output of the inverse transformer 16 is connected to an input of the combiner 18. The combiner 18 Decoder The reconstructed audio signal is output at output 22 of 10 .

[0013] The frequency-domain coefficient extractor 12 is configured to extract frequency-domain coefficients 24 of frames 26 of the audio signal from the data stream 20. The frequency-domain coefficients 24 may be MDCT coefficients or may belong to some other transform, such as another lapped transform. As will be further explained below, the frequency-domain coefficients 24 belonging to a particular frame 26 describe the spectrum of the audio signal in the respective frame 26 with various spectro-temporal resolutions. The frames 26 represent time portions into which the audio signal is successively partitioned in time. Collectively, all frequency-domain coefficients 24 of all frames represent a spectrogram 28 of the audio signal. The frames 26 may, for example, be of equal length. Due to the nature of the audio content of an audio signal changing over time, it may be disadvantageous to describe the spectrum of each frame 26 with a continuous spectro-temporal resolution, for example, by using a transform with a constant transform length. The transform length, for example, spans the time length of each frame 26, i.e., includes the sample values ​​in this frame 26 of the audio signal as well as the time-domain samples preceding and following the respective frame. For example, a lossy transmission of the spectrum of each frame in the form of frequency-domain coefficients 24 may result in pre-echo artifacts. Therefore, in a manner outlined further below, the frequency-domain coefficients 24 of each frame 26 describe the spectrum of the audio signal in this frame 26 with a spectro-temporal resolution that is switchable by switching between different transform lengths. However, as far as the frequency-domain coefficient extractor 12 is concerned, the latter situation is transparent to this. The frequency-domain coefficient extractor 12 operates independently of any signaling that signals the just-mentioned switching between different spectro-temporal resolutions of the frame 26.

[0014] The frequency-domain coefficient extractor 12 may use entropy coding to extract the frequency-domain coefficients 24 from the data stream 20. For example, the frequency-domain coefficient extractor may use context-based entropy decoding, such as variable context arithmetic decoding, to extract the frequency-domain coefficients 24 from the data stream 20 by assigning the same context to each of the frequency-domain coefficients 24, regardless of the signaling described above that signals the spectro-temporal resolution of the frame 26 to which the respective frequency-domain coefficient belongs. Alternatively, as a second example, the extractor 12 may use Huffman decoding to define a set of Huffman code words, regardless of the signaling described above that specifies the resolution of the frame 26.

[0015] There are several different possibilities for how the frequency-domain coefficients 24 describe the spectrogram 28. For example, the frequency-domain coefficients 24 may simply represent some prediction residual. For example, the frequency-domain coefficients may represent a prediction residual obtained, at least in part, by stereo prediction from another audio signal representing the corresponding audio channel or downmix from the multi-channel audio signal to which the signal spectrogram 28 belongs. Alternatively, or in addition to a prediction residual, the frequency-domain coefficients 24 may represent a sum (middle) or difference (outer) signal according to the M / S stereo paradigm [5]. Furthermore, the frequency-domain coefficients 24 may have undergone temporal noise shaping.

[0016] Moreover, the frequency domain coefficients 24 are quantized and spectrally modified to maintain the quantization error below a psychoacoustic detection (or masking) threshold, e.g., the quantization step size is controlled via respective scaling factors associated with the frequency-domain coefficients 24. The scale factor extractor 14 is responsible for extracting the scaling factors from the data stream 20.

[0017] We will briefly discuss in more detail below the switching between different spectro-temporal resolutions from frame to frame. As explained in more detail below, switching between different spectro-temporal resolutions implies either that all frequency-domain coefficients 24 in a particular frame 26 belong to one transform, or that the frequency-domain coefficients 24 of each frame 26 actually belong to different transforms. The different transforms could be, for example, two transforms whose transform lengths are half the transform length of the single transform just mentioned. While the embodiments described below in conjunction with the figures assume switching between one transform on the one hand and two transforms on the other hand, in practice, switching between one transform and three or more transforms is also feasible in principle, and the embodiments given below can be easily converted to such alternative embodiments.

[0018] FIG. 1 uses hatching to illustrate an exemplary case where the current frame is of a type that is represented by two short transforms. One of the two short transforms is derived using the second half of the current frame 26 of the audio signal, and the other is obtained by transforming the first half of the current frame 26 of the audio signal. Due to the shortened transform length, the spectral resolution with which the frequency-domain coefficients 24 describe the spectrum of frame 26 is reduced, i.e., halved when using two short transforms, while the temporal resolution is increased, i.e., doubled in this case. In FIG. 1, for example, the frequency-domain coefficients 24 shown with hatching belong to the preceding transform, and the frequency-domain coefficients 24 without hatching belong to the subsequent transform. Thus, spectrally co-located frequency-domain coefficients 24 describe the same spectral components of the audio signal in frame 26, but at slightly different times, i.e., in two consecutive transform windows of the transform split frame.

[0019] In data stream 20, frequency-domain coefficients 24 are transmitted in an interleaved manner, such that spectrally corresponding frequency-domain coefficients of two different transforms immediately follow each other. In other words, if the frequency-domain coefficients 24 as received from frequency-domain coefficient extractor 12 are sequentially ordered as if they were the frequency-domain coefficients of a long transform, they are then transmitted in an interleaved manner in this sequence, such that spectrally co-located frequency-domain coefficients 24 are immediately adjacent to each other and pairs of such spectrally co-located frequency-domain coefficients 24 are ordered according to spectral / frequency order. Interestingly, when so ordered, the sequence of interleaved frequency-domain coefficients 24 appears similar to a sequence of frequency-domain coefficients 24 obtained by a single long transform. Again, as far as the frequency-domain coefficient extractor 12 is concerned, switching between different transform lengths or spectro-temporal resolutions in units of frames 26 is transparent to it, and therefore, as a result of the context selection for context-adaptive entropy coding of the frequency-domain coefficients 24, the same context will be selected, regardless of whether the current frame is actually a long transform or a split transform type, without the extractor 12 being aware of it. For example, the frequency-domain coefficient extractor 12 can select the context to be utilized for a particular frequency-domain coefficient based on its spectro-temporally neighboring already coded / decoded frequency-domain coefficients, where this spectro-temporal neighboring is defined in the interleaved state shown in FIG. 1. This has the following consequence: Assume that the currently coded / decoded frequency-domain coefficient 24 was part of a previous transform, which is indicated using hatching in FIG. 1. The spectrally immediately neighboring frequency-domain coefficients are then actually frequency-domain coefficients 24 of the same previous transform (i.e., the one with hatching in FIG. 1).However, for context selection, the frequency-domain coefficient extractor 12 nevertheless uses frequency-domain coefficients 24 belonging to a subsequent transform, i.e., spectrally adjacent (according to the reduced spectral resolution of the shortened transform), assuming them to be immediate spectral neighbors of one longer transform of the current frequency-domain coefficient 24. Similarly, in selecting a context for frequency-domain coefficients 24 of a subsequent transform, the frequency-domain coefficient extractor 12 uses frequency-domain coefficients 24 belonging to a preceding transform and actually spectrally co-located with that coefficient as immediate spectral neighbors. In particular, the decoding order defined among the coefficients 24 of the current frame 26 can, for example, proceed from lowest frequency to highest frequency. A similar observation is valid when the frequency-domain coefficient extractor 12 is configured to entropy decode frequency-domain coefficients 24 of the current frame 26 in groups / tuples of immediately consecutive frequency-domain coefficients 24 when ordered but not deinterleaved. Instead of using tuples of spectrally adjacent frequency-domain coefficients 24 that belong only to the same short transform, the frequency-domain coefficient extractor 12 may select a context for a particular spectrally adjacent tuple of mixed frequency-domain coefficients 24 that belong to different short transforms based on such a particular spectrally adjacent tuple of mixed frequency-domain coefficients 24 that belong to different short transforms.

[0020] As shown above, due to the fact that in the interleaved state the resulting spectrum as obtained by two short transforms looks very similar to the spectrum obtained by one long transform, the entropy coding penalty resulting from the operation of the frequency domain coefficient extractor 12 being independent of transform length switching is low.

[0021] As mentioned above, we resume the description of decoder 10 with scaling coefficient extractor 14, which is responsible for extracting scaling coefficients for frequency-domain coefficients 24 from data stream 20. The spectral resolution at which scale coefficients are assigned to frequency-domain coefficients 24 is coarser than the relatively fine spectral resolution supported by a long transform. As indicated by curly brackets 30, frequency-domain coefficients 24 can be grouped into multiple scale coefficient bands. The partitioning within the scale coefficient bands may be selected based on psychoacoustic considerations, for example, to coincide with so-called Bark (or critical) bands. Because scaling coefficient extractor 14 does not rely on transform length switching, just as frequency-domain coefficient extractor 12 does, scaling coefficient extractor 14 assumes that each frame 26 is partitioned into multiple equal scale coefficient bands 30 and extracts a scale coefficient 32 for each such scale coefficient band 30, regardless of transform length switching signaling. At the encoder side, the assignment of frequency-domain coefficients 24 to these scale coefficient bands 30 is performed in the uninterleaved state shown in FIG. 1. As a result, for a frame 26 corresponding to a split transform, each scale coefficient 32 belongs to a group that includes both frequency-domain coefficients 24 of the preceding transform and frequency-domain coefficients 24 of the succeeding transform.

[0022] The inverse transformer 16 is configured to receive, for each frame 26, the corresponding frequency-domain coefficients 24 and the corresponding scale factors 32, and to perform an inverse transform on the frequency-domain coefficients 24 of the frame 26, scaled according to the scale factors 32, to obtain a time-domain portion of the audio signal. The inverse transformer 16 may use, for example, a lapped transform such as a modified discrete cosine transform (MDCT). The combiner 18 combines the time-domain portions, for example, by using a suitable overlap-add technique, to obtain the audio signal. The overlap-add technique, for example, provides time-domain anti-aliasing among overlapping portions of the time-domain portions output by the inverse transformer 16.

[0023] Of course, the inverse transformer 16 is responsive to the aforementioned transform length switches signaled in the data stream 20 for the frames 26. The operation of the inverse transformer 16 will now be described in more detail with reference to FIG.

[0024] Figure 2 shows a possible internal structure of the inverse transformer 16 in more detail. As shown in Figure 2, the inverse transformer 16 receives, for a current frame, frequency-domain coefficients 24 associated with that frame and corresponding scale factors 32 for dequantizing the frequency-domain coefficients 24. Additionally, the inverse transformer 16 is controlled by signaling 34 present in the data stream 20 for each frame. The inverse transformer 16 may be further controlled via other components of the data stream 20 that are optionally included within the data stream 20. The following description provides more details regarding these additional parameters.

[0025] As shown in Figure 2, the inverse transformer 16 of Figure 2 comprises an inverse quantizer 36, an activatable deinterleaver 38, and an inverse transform stage 40. To facilitate understanding of the following description, the incoming frequency-domain coefficients 24 as derived for the current frame from the frequency-domain coefficient extractor 12 are shown labeled 0 to N-1. Again, because the frequency-domain coefficient extractor 12 is agnostic to the signaling 34, i.e., operates independently of the signaling 34, the frequency-domain coefficient extractor 12 provides the frequency-domain coefficients 24 to the inverse transformer 16 in the same manner regardless of whether the current frame is of a split transform type or a single transform type, i.e., whether the number of frequency-domain coefficients 24 is N in this example, and the association of the indices 0 to N-1 to the N frequency-domain coefficients 24 also remains the same regardless of the signaling 34. If the current frame is of 1 or long transform type, the indices 0 to N-1 correspond to the ordering of the frequency domain coefficients 24 from lowest frequency to highest frequency, and if the current frame is of split transform type, the indices correspond to the ordering for the frequency domain coefficients, but the frequency domain coefficients are then spectrally arranged according to spectral order, but interleaved such that every second and every other frequency domain coefficient 24 belongs to the subsequent transform, while the other frequency domain coefficients 24 belong to the preceding transform.

[0026] The same applies to the scale coefficients 32. Because the scale coefficient extractor 14 operates independently of the signaling 34, the number, order and values ​​of the scale coefficients 32 coming from the scale coefficient extractor 14 are independent of the signaling 34, and the scale coefficients 32 in FIG. 2 are exemplarily numbered S0 to S1 with indices corresponding to the sequential order among the scale coefficient bands to which they are associated. M is shown as:

[0027] Similar to the frequency-domain coefficient extractor 12 and the scale factor extractor 14, the inverse quantizer 36 can operate independently of or independent of the signaling 34. The inverse quantizer 36 inverse quantizes or scales the incoming frequency-domain coefficients 24 using scale factors associated with the scale factor band to which each frequency-domain coefficient belongs. Again, the membership of the incoming frequency-domain coefficients 24 to individual scale factor bands, and thus their association with scale factors 32, is independent of the signaling 34; thus, the inverse transformer 16 scales the frequency-domain coefficients 24 by scale factors 32 at a spectral resolution that is independent of the signaling 34. For example, the inverse quantizer 36 may assign frequency-domain coefficients indices 0-3 for the first scale factor band, thus resulting in the first scale factor S0, indices 4-9 for the second scale factor band, thus resulting in the scale factor S1, and so on, independent of the signaling 34. The scale factor boundaries are intended to be exemplary only. The inverse quantizer 36 may, for example, perform multiplications using the associated scale factors to inversely quantize the frequency-domain coefficients 24, i.e., x0 as x0·s0, x1 as x1·s0, ... x3 as x3·s0, x4 as x4·s1, ... x9 as x9·s1, and so on. Alternatively, the inverse quantizer 36 may perform interpolation of the scale factors actually used to inversely quantize the frequency-domain coefficients 24 from the coarse spectral resolution defined by the scale factor bands. The interpolation may be independent of the signaling 34. However, the latter interpolation may alternatively depend on the signaling to take into account different spectro-temporal sampling positions of the frequency-domain coefficients 24 depending on whether the current frame is of a split transform type or a 1 / long transform type.

[0028] FIG. 2 illustrates that the order among the frequency-domain coefficients 24 remains the same up to the input of the activatable deinterleaver 38, and the same is true, at least in part, for the overall operation up to that point. FIG. 2 also illustrates that further operations may be performed by the inverse transformer 16 upstream of the activatable deinterleaver 38. For example, the inverse transformer 16 may be configured to perform noise filling on the frequency-domain coefficients 24. For example, in the sequence of frequency-domain coefficients 24, scale factor bands, i.e., groups of incoming frequency-domain coefficients ordered according to indexes 0 through N-1, may be identified, where all frequency-domain coefficients 24 in each scale factor band are quantized to zero. Such frequency-domain coefficients may be filled using artificial noise generation, e.g., using a pseudorandom number generator. The intensity / level of the noise filled within the zero-quantized scale factor bands may be adjusted using the scale factor of the respective scale factor band, since the spectral coefficients therein are all zero and therefore not required for scaling. Such noise filling is shown at 40 in FIG. 2 and is described in more detail in one embodiment in European Patent Application Publication No. EP2304719A1 [6].

[0029] 2 further illustrates that the inverse transformer 16 can be configured to support joint stereo coding and / or inter-channel stereo prediction. In the framework of inter-channel stereo prediction, the inverse transformer 16 can predict 42 the spectrum of the deinterleaved sequence represented by the order of indices 0 to N-1, for example, from another channel of the audio signal. That is, this can mean that the frequency-domain coefficients 24 describe the spectrogram of a channel of the stereo audio signal, and that the inverse transformer 16 is configured to process the frequency-domain coefficients 24 as a prediction residual of a prediction signal derived from another channel of the stereo audio signal. This inter-channel stereo prediction can be performed, for example, at a certain spectral granularity independent of the signaling 34. Complex prediction parameters 44 controlling the complex stereo prediction 42 can trigger the complex stereo prediction 42, for example, for a particular one of the aforementioned scale factor bands. For each scale factor band for which complex prediction is initiated by the complex prediction parameters 44, the scaled frequency domain coefficients 24 ordered from 0 to N-1 present in the respective scale factor band are summed with an inter-channel predicted signal obtained from the other channels of the stereo audio signal, and the complex coefficients contained in the complex prediction parameters 44 for this respective scale factor band can control the predicted signal.

[0030] Furthermore, within the framework of joint stereo coding, the inverse transformer 16 can be configured to perform MS decoding 46. That is, the decoder 10 of FIG. 1 can perform the operations described above twice, once for the first channel of the stereo audio signal and once for the second channel. Controlled via MS parameters in the data stream 20, the inverse transformer 16 can MS-decode these two channels or leave them as they are, i.e., the left and right channels of the stereo audio signal. The MS parameters 48 can switch between MS coding at the frame level or even at some finer level, such as per scale factor band or group thereof. For example, in the case of MS decoding being initiated, the inverse transformer 16 can form the sum or difference of corresponding frequency-domain coefficients 24 in coefficient order 0 to N-1 with corresponding frequency-domain coefficients of the other channel of the stereo audio signal.

[0031] 2 shows that the activatable deinterleaver 38 responds to signaling 34 for the current frame as follows: if the current frame is signaled by signaling 34 to be a divided transform frame, it deinterleaves the incoming frequency-domain coefficients to obtain two transforms, i.e., an earlier transform 50 and a later transform 52; if signaling 34 indicates that the current frame is a long transform frame, it leaves the frequency-domain coefficients interleaved to yield one transform 54. When deinterleaving, the deinterleaver 38 forms one of the transforms 50 and 52, i.e., one short transform from the frequency-domain coefficients with even indices and the other short transform from the frequency-domain coefficients at odd index positions. For example, the even-indexed frequency-domain coefficients form the earlier transform (starting at index 0), while the other frequency-domain coefficients form the later transform. These transforms 50 and 52 undergo an inverse transform of the shorter transform length, resulting in time-domain portions 56 and 58, respectively. 1 positions the time-domain portions 56 and 58 correctly in time, i.e., positions the time-domain portion 56 resulting from the preceding transform 50 before the time-domain portion 58 resulting from the subsequent transform 52, and performs an overlap-add process therebetween with the time-domain portions resulting from the preceding and subsequent frames of the audio signal. If not deinterleaved, the frequency-domain coefficients arriving at the interleaver 38 would directly form the long transform 54, and the inverse transform stage 40 performs an inverse transform on the frequency-domain coefficients to result in a time-domain portion 60 that spans the entire time interval of the current frame 26 and beyond. Combiner 18 combines time-domain portion 60 with the respective time-domain portions resulting from the preceding and subsequent frames of the audio signal.

[0032] The frequency-domain audio decoders described thus far enable transform length switching to enable compatibility with frequency-domain audio decoders that do not support signaling 34. In particular, such “legacy” decoders may erroneously assume that frames signaled by signaling 34 are of a long transform type when in fact they are of a split transform type. That is, they may erroneously leave the split-type frequency-domain coefficients interleaved and perform a long transform length inverse transform. However, the resulting quality of the affected frames of the reconstructed audio signal is still quite reasonable.

[0033] Conversely, the coding efficiency penalty remains quite reasonable. The coding efficiency penalty results from ignoring signaling 34, since frequency-domain coefficients and scale factors are coded without taking into account the meaning of various coefficients and without exploiting this variation to increase coding efficiency. However, the latter penalty is relatively small compared to the advantage of enabling backward compatibility. The latter statement also applies to the restriction on activation and deactivation of noise filler 40, complex stereo prediction 42, and MS decoding 46 only within contiguous spectral portions (scale factor bands) in the deinterleaved state defined by indices 0 to N-1 in FIG. 2. While the opportunity to enable frame-type-specific control of these coding tools (e.g., with two noise levels) may provide advantages in some cases, these advantages are overcompensated by the advantage of having backward compatibility.

[0034] 2 illustrates that the decoder of FIG. 1 can be further configured to support TNS (Temporal Noise Shaping) coding while still maintaining backward compatibility with decoders that do not support signaling 34. In particular, FIG. 2 illustrates the possibility that inverse TNS filtering, if any, may occur after any complex stereo prediction 42 and MS decoding 46. To maintain backward compatibility, inverse transformer 16 is configured to perform inverse TNS filtering 62 on the sequence of N coefficients, regardless of signaling 34, using each TNS coefficient 64. With this approach, data stream 20 encodes TNS coefficients 64 equally regardless of signaling 34; i.e., the number of TNS coefficients and the manner in which they are encoded are the same. However, inverse transformer 16 is configured to apply TNS coefficients 64 differently. If the current frame is a long transform frame, the inverse TNS filtering is performed on the long transform 54, i.e., the interleaved sequence of frequency-domain coefficients, and if the current frame is signaled as a divided transform frame by signaling 34, the inverse transformer 16 inverse TNS filters 62 the concatenation of the preceding transform 50 and the succeeding transform 52, i.e., the sequence of frequency-domain coefficients with indices 0, 2, ..., N-2, 1, 3, 5, ..., N-1. The inverse TNS filtering 62 can, for example, include the inverse transformer 16 applying a filter whose transfer function is set in accordance with the TNS coefficients 64 for the deinterleaved or interleaved sequence of coefficients passed through the processing sequence upstream of the deinterleaver 38.

[0035] Thus, a "legacy" decoder that erroneously processes a split frame type frame as a long transform frame will apply the TNS coefficients 64 that have been generated by the encoder by analyzing the concatenation of two real-time transforms, 50 and 52, to transform 54, and will therefore generate an incorrect time-domain portion 60 by applying an inverse transform to transform 54. However, if the use of such split transform frames is restricted to signals representing rain or applause, etc., then even if this degradation in quality occurs in such a decoder, it may be tolerable to a listener.

[0036] For completeness, Figure 3 shows that the inverse TNS filtering 62 of the inverse transformer 16 can be inserted elsewhere in the processing sequence shown in Figure 2. For example, the inverse TNS filtering 62 could be located upstream of the complex stereo prediction 42. To preserve the deinterleaved domain downstream and upstream of the inverse TNS filtering 62, Figure 3 shows that if the frequency-domain coefficients 24 were only previously deinterleaved 66, then to perform the inverse TNS filtering 68 within the deinterleaved concatenation where the frequency-domain coefficients 24 as processed so far are in the order of indices 0, 2, 4, ..., N-2, 1, 3, ..., N-3, N-1, the deinterleaving is reversed 70 to obtain the inverse TNS filtered versions of the frequency-domain coefficients again in their interleaved order 0, 1, 2, ..., N-1. The position of the inverse TNS filtering 62 within the processing step sequence shown in Figure 2 may be fixed or may be signaled via the data stream 20, for example, frame by frame or at some other granularity.

[0037] It should be noted that, for ease of explanation, the above embodiments focus only on the juxtaposition of long transform frames and split transform frames. However, the embodiments of the present application can be similarly extended by introducing frames of other transform types, such as frames consisting of eight short transforms. In this regard, it should be noted that the aforementioned independence only relates to frames that are distinguished from such other frames of any third transform type by additional signaling, whereby a "legacy" decoder would mistakenly process a split transform frame as a long transform frame by examining the additional signaling contained in all frames; only frames that are distinguished from other frames (all except split transform and long transform frames) include signaling 34. As far as such other frames (all except split transform and long transform frames) are concerned, it should be noted that the operating modes of the extractors 12 and 14, such as context selection, may depend on the additional signaling, i.e., such operating modes may differ from the operating modes applied to split transform and long transform frames.

[0038] Before describing a suitable encoder compatible with the above-described decoder embodiment, we will describe an implementation of the above-described embodiment that is suitable for adaptively updating an xHE-AAC-based audio encoder / decoder to enable it to support backward-compatible transform splitting.

[0039] That is, the following describes a possible method for implementing transform length partitioning in an audio codec based on MPEG-D xHE-AAC (USAC) in order to improve the coding quality of certain audio signals at low bit rates. The transform partitioning tool is signaled in a semi-backward compatible manner so that a legacy xHE-AAC decoder can parse and decode the bitstream according to the above embodiment without obvious audio errors or loss. As shown below, this semi-backward compatible signaling makes use of the unused possible values ​​of frame syntax elements that control the usage of noise filling in a conditional coding manner. Legacy xHE-AAC decoders do not support these possible values ​​of the respective noise filling syntax elements, but improved audio decoders do.

[0040] In particular, the embodiment described below, in accordance with the above-described embodiment, allows for intermediate transform lengths for coded signals similar to rain or applause, preferably split long blocks, i.e., two consecutive transforms, each half or a quarter of the spectral length of the long block, with the maximum temporal overlap between these transforms being less than the maximum temporal overlap between consecutive long blocks. To enable coded bitstreams with transform splitting, i.e., signaling 34, to be read and parsed by legacy xHE-AAC decoders, the splitting should be used semi-backward compatible, and the presence of such a transform splitting tool should not cause legacy decoders to stop decoding or even to not start decoding. The readability of such bitstreams via the xHE-AAC infrastructure can also promote market adoption. To achieve the just-mentioned semi-backward compatibility goal for using transform splitting with xHE-AAC or its possible derivatives, the transform splitting is signaled via xHE-AAC noise-filled signaling. According to the above-described embodiment, to construct the transform partitioning for the xHE-AAC encoder / decoder, a partitioned transform consisting of two separate half-length transforms can be used instead of a frequency domain (FD) stop-start window sequence. The temporally consecutive half-length transforms are interleaved into a single stop-start-like block per coefficient for decoders that do not support transform partitioning, i.e., legacy xHE-AAC decoders. Signaling via noise-filling signaling is performed as described below. In particular, 8-bit noise-filling side information can be used to signal the transform partitioning. This is feasible because the MPEG-D standard [4] states that all 8 bits are transmitted even if the noise level to be applied is zero. In that situation, some of the noise-filling bits can be reused for the transform partitioning, i.e., signaling 34.

[0041] Semi-backward compatibility for bitstream parsing and playback by legacy xHE-AAC decoders can be ensured as follows: Transform splitting is signaled via a noise level of zero, i.e., the first three noise-filling bits have a value of all zero, followed by five non-zero bits (conventionally representing a noise offset) that contain side information about the transform splitting and the noise level to be lost. Because legacy xHE-AAC decoders ignore the value of the 5-bit offset when the 3-bit noise level is zero, the presence of transform splitting signaling 34 only affects noise filling in legacy decoders. That is, because the first three bits are zero, noise filling is turned off and the remaining decoding operations work as intended. In particular, the split transform is processed like a conventional stop-start block using a full-length inverse transform (due to the coefficient interleaving described above), and no deinterleaving is performed. Thus, a conventional decoder still allows for "graceful" decoding of the improved data stream / bitstream 20, since it does not need to attenuate the output signal 22 or even abort decoding when a transform split type frame arrives. Naturally, such a conventional decoder will not be able to provide an exact reconstruction of the split transform frames, resulting in a deterioration in quality for the affected frames compared to decoding by a proper decoder, for example, according to Fig. 1. Nevertheless, assuming that transform splitting is used as intended, i.e., only for transient or noisy inputs at low bit rates, the quality provided by the xHE-AAC decoder should be better than if the affected frames were dropped due to attenuation or would otherwise result in obvious playback errors.

[0042] Specifically, the extension of the xHE-AAC encoder / decoder towards transform splitting can be as follows:

[0043] According to the above description, the new tool to be used in xHE-AAC can be called transform splitting (TS). Transform splitting is a new tool in the frequency domain (FD) encoder of xHE-AAC or, for example, in MPEG-H 3D-Audio, which is based on USAC [4]. Transform splitting can then be used for specific transient signal passages as an alternative to the usual long transform (which leads to temporal blurring, especially pre-echoes, at low bit rates) or eight short transforms (which lead to spectral holes and bubble artifacts at low bit rates). Transform splitting can then be signaled semi-backward compatible by interleaving the FD coefficients into a long transform that can be accurately parsed by a conventional MPEG-D USAC decoder.

[0044] The description of this tool is similar to that above. When transform splitting is active in a long transform, two half-length MDCTs are used instead of one full-length MDCT, and the coefficients of the two MDCTs, i.e., 50 and 52, are transmitted line-by-line interleaved. Interleaved transmission has already been used, for example, in the case of frequency-domain (stop-start) transforms, where the coefficients of the first MDCT in time are located at even indices and the coefficients of the second MDCT in time are located at odd indices (when indexing starts at zero). However, decoders that are not capable of processing stop-start transforms are unable to correctly parse the data stream. That is, since the different contexts used to entropy code frequency-domain coefficients are valid for such stop-start transforms, i.e., a modified syntax streamlined to half the transform, any decoder that is not capable of supporting stop-start windows had to ignore the respective stop-start window frames.

[0045] Referring briefly back to the embodiment described above, this may enable the decoder of FIG. 1 to support partitioning of a particular frame 26 into more than two transforms beyond the description presented thus far, or using signaling that extends signaling 34. However, with regard to the concurrent notation of transform partitioning of frame 26 other than the partitioned transform initiated using signaling 34, FD coefficient extractor 12 and scaling coefficient extractor 14 respond to this signaling in that their operating modes change in response to further signaling in addition to signaling 34. Furthermore, streamlined transmission of TNS coefficients, MS parameters, and complex prediction parameters tailored to signaled transform types other than the partitioned transform types per 56 and 59 requires that each decoder must be able to respond to, i.e., understand, the signaling selection between frames containing these "known transform types" or long transform types per 60 and other transform types, such as, for example, one partitioned frame into eight short transforms as in the case of AAC. In that case, this "known signaling" identifies frames in which signaling 34 signals a split transform type as long transform type frames, so that decoders that are not capable of understanding signaling 34 process these frames as long transform frames rather than other types of frames, such as the eight short transform type frames.

[0046] Returning again to the discussion of possible extensions to xHE-AAC, incorporating transform splitting tools into this coding framework may result in certain operational restrictions. For example, transform splitting may be allowed to be used only in frequency-domain long start or stop-start windows. That is, the underlying syntax element window_sequence may be required to be equal to 1. Additionally, due to semi-backward compatibility signaling, there may be a requirement that transform splitting can only be applied when the syntax element noiseFilling is 1 in the syntax container UsacCoreConfig(). When transform splitting is signaled to be active, all frequency-domain tools except TNS and inverse MDCT operate on interleaved (long) TS coefficient sets. This allows for the reuse of scale factor band offsets and long transform arithmetic coder tables, as well as window shapes and overlap lengths.

[0047] In the following, we present the terms and definitions used below to explain how the USAC standard described in [4] can be extended to provide backward compatible transform splitting functionality. Interested readers may be referred to sections within that standard.

[0048] The new data elements may be: split_transform: A binary flag indicating whether a split transform is used for the current frame and channel.

[0049] The new auxiliary elements may be: window_sequence: Frequency domain window sequence type for the current frame and channel (Section 6.2.9) noise_offset: Noise filling offset to modify the scale factor of the zero quantization band (Section 7.2) noise_level: The noise filling level (Section 7.2) that represents the amount of spectral noise added. half_transform_length: half of coreCoderFrameLength (ccfl, Transform Length, section 6.1.1) half_lowpass_line: Half the number of MDCT lines transmitted for the current channel

[0050] Decoding of frequency domain (stop-)start transforms using transform splitting (TS) in the USAC framework can be performed in purely sequential steps as follows:

[0051] First, split_transform and half_lowpass_line decoding can be performed.

[0052] split_transform does not actually represent a separate bitstream element, but is derived from the noise filling factors, noise_offset and noise_level, and, in the case of UsacChannelPairElement(), the common_window flag in StereoCoreToolInfo(). If noiseFilling == 0, then split_transform is 0; otherwise:

number

[0053] In other words, if noise_level == 0, noise_offset contains the split_transform flag, followed by 4 bits of noise filling data, which are then rearranged. This operation must be performed before the noise filling process in section 7.2, as it changes the values ​​of noise_level and noise_offset. Furthermore, if common_window == 1 in UsacChannelPairElement(), split_transform is determined only on the left (first) channel, the right channel's split_transform is set equal to (and replicated from) the left channel's split_transform, and the above pseudocode is not performed on the right channel.

[0054] half_lowpass_line is determined from the "long" scale factor band offset table swb_offset_long_window and max_sfb of the current channel, or max_sfb_ste if stereo and common_window == 1.

[0055] In elements with StereoCoreToolInfo() and common_window == 1 it is max_sfb_ste, otherwise lowpass_sfb =max_sfb. Based on the igFilling flags half_lowpass_line is derived as follows:

number

[0056] Then, as a second step, half-length spectral deinterleaving for temporal noise shaping is performed.

[0057] After spectral dequantization, noise filling, and application of scale factors, but before application of Temporal Noise Shaping (TNS), the TS coefficients in spec[] are deinterleaved using a helper buffer[].

number

[0058] In-place deinterleaving effectively places two half-length TS spectra on top of each other, and TNS tools operate normally on the resulting full-length pseudospectrum.

[0059] See above, such a procedure is described in connection with FIG.

[0060] Then, as a third step, temporal re-interleaving is used with two successive inverse MDCTs.

[0061] If common_window == 1 in the current frame or stereo decoding is performed after TNS decoding (tns_on_lr == 0 in Section 7.8), spec[] must be time-reinterleaved into a full-length spectrum.

number

[0062] The resulting pseudospectrum is used for stereo decoding (Section 7.7) and dmx_re_prev[] is updated (Sections 7.7.2 and A.1.4). If tns_on_lr == 0, the stereo decoded full-length spectrum is deinterleaved again by repeating the process of Section A.1.3.2. Finally, two inverse MDCTs are computed using ccfl and the window_shape of that channel for the current and last frames. See Section 7.9 and Figure 1.

[0063] Some modifications can be made to the complex predictive stereo decoding of xHE-AAC.

[0064] To incorporate TS within xHE-AAC, an implicit semi-backward compatible signaling method can be used as an alternative.

[0065] The above describes a technique for using one bit in the bitstream to signal the use of the transform splitting of the present invention, contained in split_transform, to the decoder of the present invention. In particular, such signaling (called explicit semi-backward compatible signaling) allows subsequent legacy bitstream data (here, noise-filling side information) to be used independently of the signal of the present invention. That is, in an embodiment of the present invention, the noise-filling data does not depend on the transform splitting data, and vice versa. For example, noise-filling data consisting of all zeros (noise_level = noise_offset = 0) can be sent, while split_transform can hold any possible value (a binary flag of either 0 or 1).

[0066] Thus, if strict independence between the legacy bitstream data and the inventive bitstream data is not required and the inventive signaling is a binary decision, explicitly transmitting a signaling bit can be avoided; this binary decision can be signaled by the presence or absence of what may be called implicit semi-backward compatible signaling. Again taking the above embodiment as an example, the use of transform splitting can be signaled simply by using the inventive signaling. That is, if noise_level is zero and noise_offset is not zero at the same time, split_transform is set equal to 1. If both noise_level and noise_offset are not zero, split_transform is set equal to 0. When noise_level and noise_offset are both zero, a dependency of the inventive implicit signaling on the legacy noise-filling signal occurs. In this case, it is unclear whether the legacy implicit signaling or the inventive implicit signaling is used. To avoid such ambiguity, the value of split_transform must be specified in advance. In this example, if the noise filling data consists of all zeros, it is appropriate to specify split_transform = 0, since this is what a legacy coder without transform splitting should signal when noise filling should not be used in a frame.

[0067] The remaining problem to be solved in the case of implicit semi-backward compatible signaling is how to simultaneously signal split_transform == 1 and no noise filling. As mentioned before, the noise filling data must not be all zeros, and if a zero noise magnitude is desired, noise_level ((noise_offset & 14) / 2 as above) must equal 0. This allows for noise_offset greater than 0 ((noise_offset & 1) as above). *16) remains as the only solution. Advantageously, if noise filling is not performed in a decoder based on USAC [4], the value of noise_offset is ignored, so it can be seen that this approach is feasible in embodiments of the present invention. Therefore, the signaling of split_transform in the pseudocode above can be modified as follows, using the TS signaling bit reserved for transmitting noise_offset to transmit two bits (four values) instead of one bit for noise_offset:

number

[0068] Therefore, applying this alternative, the USAC description can be extended using the following explanation.

[0069] The tool descriptions are broadly the same: When transform splitting (TS) is active in a long transform, two half-length MDCTs are utilized instead of one full-length MDCT. The coefficients of the two MDCTs are transmitted line-by-line interleaved as in a conventional frequency-domain (FD) transform, with the coefficients of the first MDCT in time located at even indices and the coefficients of the second MDCT in time located at odd indices.

[0070] Operational restrictions may require that TS can only be used in FD long-start or stop-start windows (window_sequence == 1), and that TS can only be applied when noiseFilling is 1 in UsacCoreConfig(). When TS is signaled, all FD tools except TNS and inverse MDCT operate on interleaved (long) TS coefficient sets. This allows reusing scale factor band offsets and long transform arithmetic coder tables, as well as window shapes and overlap lengths.

[0071] The terms and definitions used below include the following subelements: common_window: Indicates if channel 0 and channel 1 of the CPE use the same window parameters (see ISO / IEC 23003-3:2012 section 6.2.5.1.1). window_sequence: FD window sequence type for the current frame and channel (see ISO / IEC 23003-3:2012 section 6.2.9). tns_on_lr: Indicates the operating mode of TNS filtering (see ISO / IEC 23003-3:2012 section 7.8.2). noiseFilling: This flag signals the usage of noise filling of spectral holes in the FD core encoder (see ISO / IEC 23003-3:2012 section 6.1.1.1). noise_offset: Noise filling offset to modify the scale factor of the zero quantization band (see ISO / IEC 23003-3:2012 section 7.2). noise_level: The noise filling level (see ISO / IEC 23003-3:2012 section 7.2) that represents the amount of spectral noise added. split_transform: A binary flag indicating whether TS is used in the current frame and channel. half_transform_length: Half of coreCoderFrameLength (ccfl, transform length, see ISO / IEC 23003-3:2012 section 6.1.1). half_lowpass_line: Half the number of MDCT lines transmitted for the current channel.

[0072] The decoding process involving TS can be described as follows: In particular, the decoding of the FD(Stop-)Start transformation with TS is performed in three successive steps as follows:

[0073] First, decoding of split_transform and half_lowpass_line is performed. The auxiliary element split_transform does not represent an independent bitstream element, but is derived from the noise filling elements noise_offset and noise_level and, in the case of UsacChannelPairElement(), the common_window flag in StereoCoreToolInfo(). If noiseFilling == 0, then split_transform is 0; otherwise:

number

[0074] In other words, if noise_level == 0, noise_offset contains the split_transform flag, followed by 4 bits of noise filling data, which are then rearranged. This operation must be performed before the noise filling process of ISO / IEC 23003-3:2012 section 7.2, as it changes the values ​​of noise_level and noise_offset.

[0075] Furthermore, if common_window == 1 in UsacChannelPairElement(), then split_transform is determined on the left (first) channel only, the right channel's split_transform is set equal to (and replicated from) the left channel's split_transform, and the above pseudocode is not executed on the right channel.

[0076] The auxiliary element half_lowpass_line is determined from the "long" scale factor band offset table, swb_offset_long_window and max_sfb of the current channel, or max_sfb_ste if stereo and common_window == 1.

number

[0077] Based on the igFilling flags, the half_lowpass_line is derived as follows:

number

[0078] Afterwards, half-length spectral deinterleaving for temporal noise shaping is performed.

[0079] After spectral dequantization, noise filling, and application of scale factors, but before application of temporal noise shaping (TNS), the TS coefficients in spec[ ] are deinterleaved using the helper buffer[ ].

number

[0080] In-place deinterleaving effectively places the two half-length TS spectra on top of each other, and the TNS tool then operates normally on the resulting full-length pseudospectrum.

[0081] Finally, temporal re-interleaving and two successive inverse MDCTs can be used.

[0082] If common_window == 1 in the current frame or stereo decoding is performed after TNS decoding (tns_on_lr == 0 in Section 7.8), spec[] must be time-reinterleaved into a full-length spectrum.

number

[0083] The resulting pseudospectrum is used for stereo decoding (ISO / IEC 23003-3:2012 section 7.7), dmx_re_prev[] is updated (ISO / IEC 23003-3:2012 section 7.7.2), and if tns_on_lr == 0, the stereo decoded full-length spectrum is again deinterleaved by repeating the process in that section. Finally, two inverse MDCTs are computed using ccfl and the window_shape of that channel for the current and last frames.

[0084] The processing for the TS shall follow the description given in ISO / IEC 23003-3:2012 section "7.9 Filter banks and block switching". The following additions shall be taken into consideration:

[0085] The TS coefficients in spec[] are deinterleaved using a helper buffer[] with a window length N based on the window_sequence value.

number

[0086] In this case, the IMDCT for the half-length TS is defined as follows:

number

[0087] The subsequent windowing and block switching steps are defined in the next subsections.

[0088] A transformation split by STOP_START_SEQUENCE looks like the following description:

[0089] The STOP_START_SEQUENCE combined with transform splitting is shown in Figure 2. It includes two overlapping and adding half-length windows 56, 58 with a length of N_l / 2, which is 1024 (960, 768), where N_s is set to 256 (240, 192), respectively.

[0090] The window (0,1) for the two half-length IMDCTs is given as follows:

number

number

number

[0091] The overlap and addition between two half-long windows resulting in the windowed time-domain values ​​zi,n is described as follows: where N_l is set to 2048 (1920, 1536) and N_s is set to 256 (240, 192), respectively.

number

[0092] A transformation split by LONG_START_SEQUENCE looks like the following description:

[0093] The LONG_START_SEQUENCE combined with transform division is shown in Figure 4. It contains three windows defined as follows, where N_l / is set to 1024 (960, 768) and N_s is set to 256 (240, 192):

number

number

[0094] The left / right window halves are given by:

number

number

[0095] The third window is equal to the left half of LONG_START_WINDOW.

number

number

[0096] Intermediate windowed time domain values

number

number

[0097] By applying W2, the final windowed time domain value Z i,n is obtained.

number

[0098] Regardless of whether the semi-backward compatible signaling used is explicit or implicit (both described above), some modifications may be necessary to the complex predictive stereo decoding of xHE-AAC to achieve meaningful operation on the interleaved spectrum.

[0099] A modification to the complex predictive stereo decoding can be implemented as follows.

[0100] When TS is active in a channel pair, no changes are needed to the underlying M / S or complex prediction process, since the FD stereo tool operates on an interleaved pseudospectrum. However, the derivation of the previous frame's downmix dmx_re_prev[] and the calculation of the downmix MDST dmx_im[] in ISO / IEC 23003-3:2012 section 7.7.2 need to be adapted if TS is used in either channel of the last or current frame.

[0101] If the TS has changed actively in any channel from the last to the current frame, use_prev_frame must be 0. In other words, dmx_re_prev[] must not be used in that case due to the transformation length.

[0102] If a TS was or is active, dmx_re_prev[] and dmx_re[] specify interleaved pseudospectrums that must be deinterleaved into their corresponding two half-length TS spectra for accurate MDST calculation.

[0103] When TS is activated, two half-length MDST downmixes are calculated using the adapted filter coefficients (Table 1 and Table 2) and interleaved into the full-length spectrum dmx_im[] (just like dmx_re[]).

[0104] window_sequence: Downmix MDST estimates are calculated for each group window pair. use_prev_frame is evaluated only for the first half-window pair of two half-window pairs. For the remaining window pairs, the previous window pair is always used for MDST estimation, which implies use_prev_frame = 1.

[0105] Window shape: The MDST estimation parameters for the current window are the filter coefficients as described below and depend on the shapes of the left and right window halves. For the first window, this means that the filter parameters are a function of the window_shape flags of the current and previous frames. The remaining windows are only affected by the current window_shape.

[0106] [Table 1]

[0107] [Table 2]

[0108] Finally, for the sake of completeness, Fig. 5 shows a possible frequency domain audio coder supporting transform length switching that is compatible with the embodiments outlined above. That is, the coder of Fig. 5, generally indicated using reference numeral 100, is capable of encoding an audio signal 102 into a data stream 20, which encoding is performed in a manner similar to that of the decoder of Fig. 1 described above and corresponding variants, in which the audio signal 102 is transformed for several frames. mode This is done so that TS frames can be used while "legacy" decoders can still process them without parsing errors etc.

[0109] The encoder 100 of Figure 5 comprises a transformer 104, an inverse scaler 106, a frequency-domain coefficient inserter 108, and a scale factor inserter 110. The transformer 104 is configured to receive an audio signal 102 to be coded and to transform a time-domain portion of the audio signal to obtain frequency-domain coefficients of frames of the audio signal. In particular, as has become clear from the above description, the transformer 104 decides on a frame-by-frame basis which partitioning of these frames 26 into transforms, or transform windows, will be used. As explained above, the frames 26 may be of equal length, and the transforms may be lapped transforms using overlapping transforms of different lengths. Figure 5 illustrates, for example, a case in which frame 26a undergoes one long transform, frame 26b undergoes transform splitting, i.e., two transforms of half length, and a further frame 26c undergoes two transforms of long transform length. -n length, three or more, i.e., 2 n >2 even shorter transforms. As mentioned above, this strategy allows the encoder 100 to adapt the spectro-temporal resolution of the spectrogram represented by the lapped transforms implemented by the transformer 104 to the time-varying audio content or type of audio content of the audio signal 102.

[0110] That is, frequency-domain coefficients representing the spectrogram of the audio signal 102 are provided at the output of the transformer 104. The inverse scaler 106 is connected to the output of the transformer 104 and is configured to inversely scale and simultaneously quantize the frequency-domain coefficients according to a scale factor. In particular, the inverse scaler operates on the frequency coefficients as they are obtained by the transformer 104. That is, the inverse scaler 106 necessarily needs to know the transform length assignment or transform mode assignment for the frame 26. It should also be noted that the inverse scaler 106 needs to determine the scale factor. For this purpose, the inverse scaler 106 is, for example, part of a feedback loop that evaluates the psychoacoustic masking threshold determined for the audio signal 102 and keeps the quantization noise, introduced by quantization and progressively set according to the scale factor, below the psychoacoustic detection threshold as much as possible, with or without any bitrate limitations.

[0111] At the output of the inverse scaler 106, the scale factors and the inversely scaled and quantized frequency domain coefficients are provided, the scale factor inserter 110 is configured to insert the scale factors into the data stream 20, and the frequency domain coefficient inserter 108 is configured to insert the frequency domain coefficients of the frames of the audio signal, inversely scaled and quantized according to the scale factors, into the data stream 20. To be decoder-compatible, both inserters 108 and 110 operate independently of the transform mode associated with the frame 26, as far as the juxtaposition of frame 26a in the long transform mode and frame 26b in the transform split mode is concerned.

[0112] In other words, inserters 110 and 108 operate independently of the signaling 34 that converter 104 is configured to signal in or insert into data stream 20 for frames 26a and 26b, respectively.

[0113] In other words, in the above embodiment, the transform coefficients of the long transform and split transform frames are properly adjusted, i.e., simple The one that arranges by serial arrangement or interleaving is the converter 104, and the inserter is actually Signaling 34 However, in a more general sense, it suffices that the independence of the frequency-domain coefficient inserter from the signaling is limited to inserting into the data stream a sequence of frequency-domain coefficients for each long transform and split transform frame of the audio signal that has been inversely scaled according to the scale factor, in that, depending on the signaling, if the frame is a long transform frame, the sequence of frequency-domain coefficients is formed by sequentially arranging the frequency-domain coefficients of one transform in an uninterleaved manner, and if each frame is a split transform frame, the sequence of frequency-domain coefficients is formed by interleaving the frequency-domain coefficients of two or more transforms of each frame.

[0114] As far as the frequency-domain coefficient inserter 108 is concerned, the fact that it operates independently of the signaling 34 that distinguishes between frame 26 a on the one hand and frame 26 b on the other hand means that the inserter 108 inserts into the data stream 20 the frequency-domain coefficients of the frames of the audio signal that have been inversely scaled according to the scale factor, either consecutively if one transform is performed for each frame without interleaving, or using interleaving to insert the frequency-domain coefficients of each frame if two or more transforms, i.e., two transforms in the example of Fig. 5, are performed for each frame. However, as already indicated above, the transform splitting mode can also be implemented differently, such as splitting one transform into three or more transforms.

[0115] Finally, it should be noted that the encoder of FIG. 5 can also be adapted to implement all other additional encoding means outlined above in relation to FIG. 2, such as MS encoding, complex stereo prediction 42 and TNS, and its respective parameters 44, 48 and 64 are determined for this purpose.

[0116] Although some aspects are described in terms of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, with blocks or devices corresponding to method steps or features of method steps. Similarly, aspects described in terms of a method step also represent a description of a corresponding block, item, or feature of a corresponding apparatus. Some or all of the method steps can be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, any one or more of the most important methods can be performed by such an apparatus.

[0117] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. The implementation can be realized using a digital storage medium, such as a floppy disk, DVD, Blu-Ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, on which electronically readable signals are stored that cooperate (or can cooperate) with a programmable computer system to implement the respective method. The digital storage medium can therefore be computer-readable.

[0118] Some embodiments according to the present invention include a data carrier having stored thereon an electronically readable signal capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0119] Generally, embodiments of the present invention can be realized as a computer program product having program code operable to perform one of the above methods when the computer program product is run on a computer, which program code can, for example, be stored on a machine-readable carrier.

[0120] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0121] In other words, one embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0122] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein, the data carrier, digital storage medium or computer-readable medium being generally tangible and / or non-transitory.

[0123] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, which data stream or sequence of signals can for example be arranged to be transmitted via a data communication connection, for example the Internet.

[0124] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0125] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0126] Further embodiments according to the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.

[0127] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods of the present invention are preferably performed by any hardware apparatus.

[0128] The above-described embodiments are merely illustrative of the principles of the present invention. Naturally, modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore the intention that the present invention be limited only by the scope of the appended claims and not by the specific details presented by the description and illustration of the embodiments herein.

[0129] [Claim 1] 1. A frequency domain audio decoder supporting transform length switching, comprising: a frequency domain coefficient extractor (12) configured to extract frequency domain coefficients (24) of frames of an audio signal from the data stream; a scale factor extractor (14) configured to extract scale factors from the data stream; an inverse transformer (16) configured to inverse transform the frequency domain coefficients of the frames scaled according to the scale factor to obtain a time domain portion of the audio signal; a combiner (18) configured to combine the time-domain portions to obtain the audio signal; The inverse transformer is responsive to a signaling in the frame of the audio signal, whereby, in response to the signaling, forming a transform by sequentially arranging the frequency domain coefficients of each frame scaled according to the scale factor in a non-deinterleaved manner, and performing an inverse transform of a first transform length on the transform; or forming two or more transforms by deinterleaving the frequency domain coefficients of the respective frames scaled according to the scale factor, and performing an inverse transform of each of the two or more transforms with a second transform length that is shorter than the first transform length; A frequency domain audio decoder, wherein the frequency domain coefficient extractor and the scale factor extractor operate independently of the signaling. [Claim 2] 2. The frequency-domain audio decoder of claim 1, wherein the scale factor extractor (14) is configured to extract the scale factors from the data stream with a spectro-temporal resolution that is independent of the signaling. [Claim 3] 3. The frequency-domain audio decoder of claim 1, wherein the frequency-domain coefficient extractor (12) uses context-based or codebook-based entropy decoding to extract the frequency-domain coefficients from the data stream by, for each frequency-domain coefficient, assigning the same context or codebook to the respective frequency-domain coefficient regardless of the signaling. [Claim 4] 4. The frequency domain audio decoder according to claim 1, wherein the inverse transformer is configured to scale the frequency domain coefficients by the scale factor at a spectral resolution independent of the signaling. [Claim 5] 5. The frequency-domain audio decoder according to claim 1, wherein the inverse transformer is configured to perform noise filling on the frequency-domain coefficients, the frequency-domain coefficients being arranged consecutively in a manner that is not deinterleaved and with a spectral resolution that is independent of the signaling. [Claim 6] The inverter is applying inverse temporal noise-shaping filtering to the frequency-domain coefficients in said forming the one transform, wherein the frequency-domain coefficients are arranged consecutively so as not to be deinterleaved; 6. The frequency domain audio decoder of claim 1, further comprising: a decoder configured to apply inverse temporal noise shaping filtering to the frequency domain coefficients in the formation of the two or more transforms, wherein the frequency domain coefficients are sequentially arranged to be deinterleaved, and the two or more transforms are spectrally concatenated accordingly. [Claim 7] 7. The frequency-domain audio decoder of claim 1, wherein the inverse transformer is configured to support joint stereo coding with or without inter-channel stereo prediction and to use the frequency-domain coefficients as sum (middle) or difference (outer) spectra or prediction residuals of the inter-channel stereo prediction, and the frequency-domain coefficients are arranged such that they are not deinterleaved regardless of the signaling. [Claim 8] 8. The frequency domain audio decoder of claim 1, wherein the number of the two or more transforms is equal to two, and the first transform length is twice the second transform length. [Claim 9] 9. The frequency domain audio decoder according to any one of claims 1 to 8, wherein the inverse transform is an inverse modified discrete cosine transform (MDCT). [Claim 10] 1. A frequency domain audio coder supporting transform length switching, comprising: a transformer (104) configured to transform a time domain portion of an audio signal to obtain frequency domain coefficients of frames of said audio signal; an inverse scaler (106) configured to inverse scale the frequency domain coefficients according to a scale factor; a frequency domain coefficient inserter (108) configured to insert the frequency domain coefficients of the frames of the audio signal, inversely scaled according to a scale factor, into the data stream; a scale factor inserter (110) configured to insert a scale factor into the data stream; the converter is configured to switch for the frames of the audio signal between performing at least one transform of a first transform length for each frame and performing two or more transforms of a second transform length for each frame that is shorter than the first transform length; the converter is further configured to signal the switching by signaling within the frame of the data stream; the frequency domain coefficient inserter is configured to insert, for each frame, a sequence of the frequency domain coefficients of the respective frame of the audio signal inversely scaled according to a scale factor into the data stream, independent of the signaling; the sequence of frequency-domain coefficients is formed, in response to the signaling, by sequentially arranging the frequency-domain coefficients of the one transform of each frame in a non-interleaved manner if one transform is performed for each frame, and by interleaving the frequency-domain coefficients of the two or more transforms of each frame if two or more transforms are performed for each frame; The scale factor inserter operates independently of the signaling of the frequency domain audio coder. [Claim 11] 1. A method for frequency domain audio decoding supporting transform length switching, comprising: extracting frequency domain coefficients of frames of an audio signal from the data stream; extracting scale factors from the data stream; inverse transforming the frequency domain coefficients of the frames scaled according to a scale factor to obtain a time domain portion of the audio signal; combining the time domain portions to obtain the audio signal; The inverse transforming step is responsive to signaling within the frame of the audio signal, whereby, in response to the signaling, forming a transform by sequentially arranging the frequency domain coefficients of each frame in a non-deinterleaved manner, and performing an inverse transform of a first transform length on the transform; or forming two or more transforms by deinterleaving the frequency domain coefficients of the respective frames, and performing an inverse transform of each of the two or more transforms with a second transform length that is shorter than the first transform length; The method, wherein the extraction of the frequency domain coefficients and the extraction of the scale factors are independent of the signaling. [Claim 12] 1. A method for frequency domain audio coding supporting transform length switching, comprising: performing a transform on a time domain portion of an audio signal to obtain frequency domain coefficients for frames of said audio signal; inverse scaling the frequency domain coefficients according to a scale factor; inserting the frequency domain coefficients of the frames of the audio signal, inversely scaled according to a scale factor, into a data stream; inserting a scale factor into the data stream; the step of performing the transform includes switching, for the frames of the audio signal, between performing one transform of at least a first transform length for each frame and performing two or more transforms of second transform lengths for each frame that are shorter than the first transform length; The method further includes signaling the switch by signaling within the frame of the data stream; the insertion of the frequency domain coefficients is performed by inserting, for each frame, a sequence of the frequency domain coefficients of the respective frame of the audio signal, inversely scaled according to a scale factor, into the data stream, independent of the signaling; the sequence of frequency-domain coefficients is formed, in response to the signaling, by sequentially arranging the frequency-domain coefficients of the one transform of the respective frame in a non-interleaved manner if one transform is performed for the respective frame, and by interleaving the frequency-domain coefficients of the two or more transforms of the respective frame if two or more transforms are performed for the respective frame; A method in which the insertion of the scale factor is performed independently of the signaling. [Claim 13] 13. A computer program having a program code for performing the method according to claim 11 or 12, when the computer program runs on a computer.

Claims

1. 1. A frequency domain audio decoder supporting transform length switching, comprising: a frequency domain coefficient extractor (12) configured to extract frequency domain coefficients (24) of frames of an audio signal from the data stream; a scale factor extractor (14) configured to extract scale factors from the data stream; an inverse transformer (16) configured to inverse transform the frequency domain coefficients of the frames scaled according to the scale factor to obtain a time domain portion of the audio signal; a combiner (18) configured to combine the time-domain portions to obtain the audio signal, The inverse transformer is responsive to a signaling in the frame of the audio signal, whereby, in response to the signaling, forming a transform by sequentially arranging the frequency domain coefficients of each frame scaled according to the scale factor in a non-deinterleaved manner, and performing an inverse transform of a first transform length on the one transform; or forming two or more transforms by deinterleaving the frequency domain coefficients of the respective frames scaled according to the scale factor, and performing an inverse transform of each of the two or more transforms with a second transform length that is shorter than the first transform length; the frequency domain coefficient extractor and the scale factor extractor operate independently of the signaling; the inverse transformer performs inverse temporal noise shaping filtering (62) on the sequence of N coefficients, regardless of the signaling, by applying a filter to the sequence of N coefficients, the filter having a transfer function set according to the TNS coefficients (64); performing the inverse temporal noise-shaping filtering on the non-interleaved consecutively arranged frequency-domain coefficients as the sequence of N coefficients in the formation of the one transform; wherein said forming of said two or more transforms is configured to perform said inverse temporal noise shaping filtering on said frequency domain coefficients arranged consecutively such that said two or more transforms are spectrally concatenated as a sequence of said N coefficients; The frequency domain coefficients (24) are grouped into several scale factor bands independent of the signaling, and the scale factor extractor (14) is configured to extract a scale factor (32) for each of the scale factor bands (30).

2. 2. The frequency domain audio decoder of claim 1, wherein the inverse transformer is configured to perform noise filling on the frequency domain coefficients, the frequency domain coefficients being arranged consecutively so as not to be deinterleaved and with a spectral resolution independent of the signaling.

3. 2. The frequency-domain audio decoder of claim 1, wherein the inverse transformer is configured to support joint stereo coding with or without inter-channel stereo prediction and to use the frequency-domain coefficients as sum (middle) or difference (outer) spectra or prediction residuals of the inter-channel stereo prediction, and the frequency-domain coefficients are arranged such that they are not deinterleaved regardless of the signaling.

4. 2. The frequency domain audio decoder of claim 1, wherein the number of the two or more transforms is equal to two, and the first transform length is twice the second transform length.

5. 1. A method for frequency domain audio decoding supporting transform length switching, comprising: extracting frequency domain coefficients of frames of an audio signal from the data stream; extracting scale factors from the data stream; inverse transforming the frequency domain coefficients of the frames scaled according to a scale factor to obtain a time domain portion of the audio signal; combining the time domain portions to obtain the audio signal; The inverse transforming step is responsive to signaling within the frame of the audio signal, whereby, in response to the signaling, forming a transform by sequentially arranging the frequency domain coefficients of each frame in a non-deinterleaved manner, and performing an inverse transform of a first transform length on the transform; or forming two or more transforms by deinterleaving the frequency domain coefficients of the respective frames, and performing an inverse transform of each of the two or more transforms with a second transform length that is shorter than the first transform length; the extraction of the frequency domain coefficients and the extraction of the scale factors are independent of the signaling; the inverse transforming step performs inverse temporal noise shaping filtering (62) on the series of N coefficients, regardless of the signaling, by applying a filter to the series of N coefficients whose transfer function is set according to the TNS coefficients (64); performing the inverse temporal noise-shaping filtering on the frequency-domain coefficients arranged consecutively in a non-deinterleaved manner as the sequence of N coefficients in the formation of the one transform; and performing the inverse temporal noise-shaping filtering on the frequency-domain coefficients arranged consecutively as the series of N coefficients such that the two or more transforms are spectrally coupled in the forming of the two or more transforms; The method, wherein the frequency domain coefficients (24) are grouped into several scale factor bands independent of the signaling, and a scale factor (32) is extracted for each of the scale factor bands (30).

6. A computer program having a program code for performing the method according to claim 5, when the computer program runs on a computer.

Citation Information

Patent Citations

  • Voice encoding method, voice decoding method, encoder and decoder

    JP1998293600A

  • lpc harmonic vocoder with superframe structure

    JP2003510644A

  • Apparatus and method for encoding and decoding audio signals

    JP2009500682A

  • Apparatus and method for encoding and decoding audio signals

    JP2009500683A

  • Audio signal encoder, audio signal decoder, method for encoding or decoding audio signals using aliasing erasure

    JP2013508765A