A decoder for decoding a symbolized audio signal and an encoder for encoding an audio signal

Adaptive switching between MDCT and MDST transform kernels addresses inefficiencies in MDCT coding, enhancing encoding efficiency and quality for harmonic and stereo signals with phase shifts.

JP7708937B2Active Publication Date: 2025-07-15FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024103916
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-06-17
Filing Date
2024-06-27
Publication Date
2025-07-15
Estimated Expiration
2036-03-08

AI Technical Summary

Technical Problem

Conventional MDCT coding in perceptual audio coders faces inefficiencies in encoding harmonic signals and stereo signals with phase shifts, leading to sub-optimal energy compression and increased complexity.

Method used

Adaptive switching between different transform kernels, such as MDCT-II, MDST-II, MDCT-IV, and MDST-IV, based on signal characteristics to improve encoding efficiency and handle phase shifts, using control information to switch between kernels with varying symmetries.

Benefits of technology

Enhances coding quality for harmonic signals and stereo signals with phase shifts, achieving improved energy compression and reduced complexity by adapting transform kernels dynamically.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708937000013
    Figure 0007708937000013
  • Figure 0007708937000014
    Figure 0007708937000014
  • Figure 0007708937000015
    Figure 0007708937000015
Patent Text Reader

Abstract

To provide a decoder, an encoder, a decoding method, an encoding method and a program for processing audio signals.SOLUTION: A decoder 2 includes an adaptive spectrum-to-time converter 6 and an overlap addition processor 8. The adaptive spectral-to-time converter 6 converts a block of consecutive spectral values 4' into consecutive blocks of time values 10, for example, via a frequency-to-time conversion, and receives control information 12 and, depending on the control information 12, switches between a first group of conversion kernels including one or more conversion kernels having different symmetries on both sides of the kernel, and a second group of conversion kernels including one or more conversion kernels having the same symmetry on both sides of the kernel. The overlap addition processor 8 overlaps and adds the consecutive blocks of the time values 10 to obtain a decoded audio value 14. The decoded audio value 14 may be a decoded audio signal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a decoder for decoding an encoded audio signal and an encoder for encoding an audio signal. Embodiments show methods and apparatus for signal adaptive transform kernel switching in audio encoding. In other words, the present invention relates to audio encoding, and in particular to perceptual audio encoding by lapped transforms such as, for example, the modified discrete cosine transform (MDCT) [1].

Background Art

[0002] All modern perceptual audio coders, including MP3, Opus, (Celt), the HE-AAC family, the new MPEG-H 3D audio, and the 3GPP enhanced voice service (EVS) codec, either employ the MDCT for quantization and encoding in the spectral domain or generate more channel waveforms. The synthesis version of this overlapping transform using the length-M spectrum spec[] is given by the following equation (1), where M = N / 2 is the length of the time window. JPEG0007708937000001.jpg11152After windowing, the time output x i,n is combined with the previous time output x i-1,n by an overlap-and-add (OLA) process. C may be a constant parameter greater than 0 or less than or equal to 1, for example, 2 / N.

[0003] The MDCT of equation (1) above is suitable for high-quality audio coding of any channel at various bitrates, but the coding quality may be insufficient in some cases. For example, · A harmonic signal having a specific fundamental frequency sampled via the MDCT such that each harmonic is represented by a plurality of MDCT bins. This leads to sub-optimal energy compression, i.e., a low coding gain, in the spectral domain. ​​· Generate a stereo signal with a phase shift of approximately 90 degrees between the MDCT bins of the channel, which is not available in the conventional M / S stereo-based joint channel coding. More advanced stereo coding, including the coding of the inter-channel phase difference (IPD), uses, for example, parametric stereo of HE-AAC or MPEG surround, but such tools operate in another filter bank domain and the complexity is increasing. Several academic papers and dissertations describe operations such as MDCT and MDST. These operations include "Lapped Orthogonal Transform (LOT)", "Extended Lapped Transform (ELT)", "Modulated Lapped Transform (MLT)", etc. Only [4] describes several different lapped transforms simultaneously, but it does not overcome the aforementioned drawbacks of MDCT. Therefore, an improved approach is needed.

Prior Art Documents

[0004]

[0005]

Non-Patent Documents

[0006]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

[0007] An object of the present invention is to provide an improved concept for processing audio signals. This object is solved by the subject matter of the independent claims. [Means for Solving the Problems]

[0008] The present invention is based on the finding that signal-adaptive changes or substitutions of the transform kernel may be able to overcome the aforementioned types of problems of this MDCT coding. According to an embodiment, the present invention addresses the above two problems regarding conventional transform coding by generalizing the MDCT coding principle to include three other similar transforms. According to the synthesis formula of Equation (1) described above, this proposed generalization is defined as the following Equation (2). JPEG0007708937000002.jpg14151

[0009] The 1 / 2 constant is replaced by the k0 constant, and the cos(...) function is replaced by the cs(...) function Please note that k0 and cs(...) are both selected adaptively with respect to the signal and the context.

[0010] According to an embodiment, the proposed modification of the MDCT coding paradigm can adapt to the instantaneous input characteristics for each frame, for example, so as to handle the aforementioned problems or cases.

[0011] An embodiment shows a decoder for decoding an encoded audio signal. The decoder includes an adaptive spectrum-time converter that is performed, for example, via a frequency-to-time conversion, to convert successive blocks of spectral values into successive blocks of time values. The decoder further includes an overlap-add processor that overlaps and adds successive blocks of time values to obtain the decoded audio values. The adaptive spectrum-time converter is configured to receive control information and switch according to the control information between a first group of conversion kernels including one or more conversion kernels having different symmetries on both sides of the kernel and a second group of conversion kernels including one or more conversion kernels having the same symmetry on both sides of the conversion kernel. The first group of conversion kernels can include one or more conversion kernels having odd symmetry on the left side of the conversion kernel and even symmetry on the right side of the conversion kernel, or vice versa, such as an inverse MDCT-IV transform or an inverse MDST-IV transform kernel, and vice versa. The second group of conversion kernels can include conversion kernels having even symmetry on both sides of the conversion kernel, such as an inverse MDCT-II transform kernel or an inverse MDST-II transform kernel, or conversion kernels having odd symmetry on both sides of the conversion kernel. The conversion kernel types II and IV will be described in more detail below.

[0012] Therefore, when encoding a signal as compared to encoding the signal with a classical MDCT Therefore, for a harmonic signal having a pitch that is at least approximately equal to an integer multiple of the frequency resolution of the transform, which can be the bandwidth of one transform bin in the spectral domain, it is advantageous to use a second group of transform kernels of the transform kernel, such as MDCT-II or MDST-II. In other words, using one of MDCT-II or MDST-II is advantageous for encoding a harmonic signal that is close to an integer multiple of the frequency resolution of the transform when compared to MDCT-IV.

[0013] A further embodiment shows that the decoder is configured to decode a multi-channel signal, such as a stereo signal. For example, in the case of a stereo signal, mid / side (M / S) stereo processing is typically superior to classical left / right (L / R) stereo processing. However, this approach does not work or is at least inferior when both signals have a 90-degree or 270-degree phase shift. According to an embodiment, it is advantageous to use MDST-IV based encoding to encode one of the two channels and conventional MDCT-IV encoding to encode the second channel. This results in a 90-degree phase shift between the two channels incorporated by an encoding scheme that compensates for a 90-degree or 270-degree phase shift of the audio channels.

[0014] A further embodiment shows an encoder for encoding an audio signal. The encoder includes an adaptive time-frequency converter for converting overlapping blocks of time values into consecutive blocks of spectral values. The encoder further comprises a controller for controlling the time-frequency converter so as to switch between a first group of conversion kernels and a second group of conversion kernels of the conversion kernels. For this purpose, the adaptive spectral-interconverter (6) receives control information (12) between a first group of conversion kernels including one or more conversion kernels having different symmetries on both sides of the kernel and a second group of conversion kernels including one or more conversion kernels having the same symmetry on both sides of the conversion kernel, and switches according to the control information. The encoder can be configured to apply different conversion kernels for the analysis of the audio signal. Thus, the encoder can apply the conversion kernels in the manner already described with respect to the decoder, and according to the embodiment, the encoder applies an MDCT or MDST operation, and the decoder applies the related inverse operation, i.e., an IMDCT or IMDST transform. Different conversion kernels will be described in detail below.

[0015] According to a further embodiment, the encoder comprises an output interface for generating an encoded audio signal having control information indicating the symmetry of the conversion kernel used to generate the current frame for the current frame. The output interface can generate control information for a decoder that can decode an audio signal encoded with the correct conversion kernel. In other words, the decoder needs to apply the inverse conversion kernel of the conversion kernel used by the encoder to encode the audio signal in each frame and channel. This information may be stored in the control information and transmitted from the encoder to the decoder, for example, using the control data section of the frame of the encoded audio signal.

[0016] Embodiments of the present invention will continue to be discussed with reference to the accompanying drawings.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 11C

Figure 12A

Figure 12B

Figure 12C

Figure 13A

Figure 13B

Figure 14A

Figure 14B

Figure 15

Figure 16

DETAILED DESCRIPTION OF THE INVENTION

[0018] Hereinafter, embodiments of the present invention will be described in more detail. Elements shown in each figure having the same or similar functions are associated with the same reference numerals.

[0019] FIG. 1 shows a schematic block diagram of a decoder 2 for decoding an encoded audio signal 4. The decoder includes an adaptive spectrum-time converter 6 and an overlap adder 8. The adaptive spectrum-time converter converts successive blocks of spectrum values 4' into successive blocks of time values 10, for example via a frequency-time conversion. Further, the adaptive spectrum-time converter (6) receives control information (12) and switches between a first group of conversion kernels including one or more conversion kernels having different symmetries on both sides of the kernel and a second group of conversion kernels including one or more conversion kernels having the same symmetry on both sides of the conversion kernel, according to the control information. Further, the overlap add processor 8 overlaps and adds successive time value blocks 10 to obtain a decoded audio value 14. The decoded audio value 14 may be a decoded audio signal.

[0020] According to an embodiment, the control information 12 can include a current bit indicating the current symmetry of the current frame, and the adaptive spectrum-time converter 6 is configured such that when the current bit indicates the same symmetry as that used in the previous frame, the current bit does not switch from the first group to the second group. In other words, for example, the control information 12 indicates using a conversion kernel of the first group for the previous frame, and if the current frame and the previous frame include the same symmetry, for example, when the current bit of the current frame and the previous frame have the same state, the conversion kernel of the first group indicated is applied, which means that the adaptive spectrum-time converter does not switch from the first conversion kernel group to the second conversion kernel group. In another way, that is, to stay in the second group or not to switch from the second group to the first group, the current bit indicating the current symmetry of the current frame indicates a symmetry different from that used in the previous frame. In other words, if the current symmetry and the previous symmetry are equal, if the previous frame was encoded using a conversion kernel from the second group, the current frame is decoded using the inverse conversion kernel of the second group.

[0021] Furthermore, if the current bit indicating the current symmetry of the current frame indicates a symmetry different from that used in the previous frame, the adaptive spectrum-time converter 6 is configured to switch from the first group to the second group. More specifically, when the current bit indicating the current symmetry of the current frame indicates a symmetry different from that used in the previous frame, the adaptive spectrum-time converter 6 is configured to switch the first group to the second group. Further, when the current bit indicating the current symmetry of the current frame indicates the same symmetry as that used in the previous frame, the adaptive spectrum-time converter 6 can switch the second group to the first group. More specifically, if the current frame and the previous frame include the same symmetry and the previous frame is encoded using the conversion kernel of the second group of conversion kernels, the current frame may be decoded using the conversion kernel of the first group of conversion kernels. The control information 12 may be derived from the encoded audio signal 4 as will be apparent below, or may be received via a separate transmission channel or carrier signal. Further, the current bit indicating the current symmetry of the current frame may be the symmetry on the right side of the conversion kernel.

[0022] In the 1986 paper [2] by Princen and Bradley, two lap transforms using trigonometric functions of the cosine or sine function are described. The first one, called "DCT-based" in that article, can be obtained by setting (2) cs() = cos() and k o = 0 , and the other is called "DST-based", where cs() = sin() and k o = 1 It is given and defined by (2). Due to the respective similarities between DCT-II and DST-II, which are commonly used in image coding, in this document, these specific cases of the general formulation of (2) are declared as the "MDCT type II" transform and the "MDST type II" transform, respectively. Princen and Bradley continued the investigation in their 1987 paper [3], proposing the common case of cs() = cos() and k o = 0.5, which was introduced in (1) and is generally known as "MDCT". For clarity of explanation and for the sake of its relationship with DCT-IV, this transform is referred to as "MDCT type IV" in this specification. Based on DST-IV, the observer has already identified the remaining possible combination obtained by using (2) with cs() = cos() and k = 0.5, which is called the "MDST type IV". Embodiments describe when to switch signal-adaptively among these four transforms. o using (2) with As pointed out in [1-3], it is valuable to define some rules regarding how to achieve an essential switch between four different transform kernels so that the complete reconstruction property (identical reconstruction of the input signal after analysis and synthesis transforms without spectral quantization or introduction of other distortions) is maintained. For this purpose, it is useful to examine the symmetric extension property of the synthesis transform according to (2), which is shown with respect to Figure 6.

[0023] · MDCT-IV shows odd symmetry on its left side and even symmetry on its right side. The synthesized signal is inverted on its left side during the inverse convolution of the signal of this transform. · MDST-IV shows even symmetry on its left side and even symmetry on its right side. The synthesized signal is inverted on its right side during the inverse convolution of the signal of this transform. · MDCT-II shows even symmetry on its left side and odd symmetry on its right side. The synthesized signal is not inverted on either side during the inverse folding of the signal of this transform. · MDST-II shows even symmetry on its left side and odd symmetry on its right side. The synthesized signal is not inverted on either side during the inverse folding of the signal of this transform. · MDCT-II shows even symmetry on its left side and odd symmetry on its right side. The synthesized signal is not inverted on either side during the inverse folding of the signal of this transform. ·MDST-II shows odd symmetry on its left side and even symmetry on its right side. The synthesized signal is inverted on both sides during the inverse convolution of the signals of this transformation.

[0024] Furthermore, two embodiments for deriving the control information 12 in the decoder will be described. The control information may include, for example, the value of k0 and cs () to indicate one of the four above-mentioned transformations. Accordingly, the adaptive spectrum-time conversion unit can read the control information of the previous frame and the control information following the previous frame from the encoded audio signal of the current frame's control data section. Optionally, the adaptive spectrum-time conversion unit 6 may read the control information 12 from the control data section of the current frame, or may read the control information for the previous frame from the control data section of the previous frame or from the decoder settings applied to the previous frame. In other words, the control information may be derived directly from the control data section or may be derived from the decoder settings of the current frame or the previous frame in the header.

[0025] Hereinafter, according to a preferred embodiment, the control information exchanged between the encoder and the decoder will be described. This section describes how the side information (i.e., the control information) is signaled and derived in the encoded bitstream, and how to derive and apply the appropriate conversion kernel in a robust (e.g., against frame loss) manner.

[0026] According to a preferred embodiment, the present invention is MPEG-D USAC (Extended HE-AAC) Or it can be integrated into the MPEG-H 3D audio codec. The determined side information can be transmitted within the so-called fd channel stream element available for each frequency domain (FD) channel and frame. More specifically, a 1-bit currAliasingSymmetry flag is written (by the encoder) and read (by the decoder) immediately before or after the scale_factor_data() bitstream element. When a given frame is an independent frame, i.e., indepFlag == 1, another bit prevAliasingSymmetry is written and read. Thereby, the symmetries on both the left and right sides, and the resulting transformation kernel are used within the frame and channel, and can be identified (and properly decoded) within the decoder even if the previous frame is lost during bitstream transmission. If the frame is not an independent frame, prevAliasingSymmetry is not written and not read, but is set equal to the value held by currAliasingSymmetry in the previous frame. According to a further embodiment, different bits or flags can be used to indicate control information (i.e., side information). Next, the respective values of cs() and k0 are derived from the currAliasingSymmetry and prevAliasingSymmetry flags (currAliasingSymmetry is abbreviated as symm

[0027] and prevAliasingSymmetry is abbreviated as symm i and prevAliasingSymmetry is abbreviated as symm i-1 ). In other words, symm i is the control information of the current frame at index i, and symm is the control information of the previous frame at index i - 1. Table 1 is based on side information regarding the symmetry derived by transmission and / or other means and a decoder-side decision matrix specifying the value of cs(...) i-1 and the symmetry derived by transmission and / or other means and a decoder-side decision matrix specifying the value of cs(...) and the symmetry derived by transmission and / or other means and a decoder-side decision matrix specifying the value of cs(...) is shown. Therefore, the adaptive spectrum-time converter can apply a conversion kernel based on Table 1 below. JPEG0007708937000003.jpg52158

[0028] Finally, when cs() and k0 are determined in the decoder, the inverse transform for a given frame and channel can be performed with the appropriate kernel using Equation (2). Before and after this synthesis transform, the decoder can operate as usual with respect to windowing, as in the prior art.

[0029] Figure 2 shows a schematic block diagram illustrating the signal flow in a decoder according to an embodiment, where solid lines indicate signals, dashed lines indicate side information, i indicates a frame index, and xi indicates a frame time-signal output. The bitstream demultiplexer 16 receives consecutive blocks of spectral values 4’ and control information 12. According to one embodiment, consecutive blocks of spectral values 4’’ and control information 12 are multiplexed onto a common signal, and the bitstream demultiplexer is configured to derive consecutive blocks of spectral values and control information from the common signal. The consecutive blocks of spectral values may further be input to a spectral decoder 18. Further, the control information of the current frame 12 and the previous frame 12’ is input to a mapper 20, which applies the mapping shown in Table 1. According to an embodiment, the control information of the previous frame 12’ may be derived using the encoded audio signal, i.e., the previous block of spectral values, or the current preset of the decoder applied to the previous frame. The spectrally decoded consecutive blocks of spectral values 4’’ and the processed control information 12’ including the parameters cs and k0 are input to an inverse kernel adaptation lap transform, which is the adaptive spectral-time converter 6 of FIG. 1. The output may be consecutive blocks of time values 10 that can optionally be processed using a synthesis window 7 to overcome discontinuities at the boundaries of consecutive blocks of time values, for example, and is input to an overlap-add processor 8 to derive the decoded audio values 14 by performing an overlap-add algorithm. The mapper 20 and the adaptive spectral-time converter 6 can further be moved to another position in the decoding of the audio signal. Thus, the positions of these blocks are merely proposed. Further, the control information may be calculated using the corresponding encoder, and an embodiment thereof is described, for example, with respect to FIG. 3.

[0030] FIG. 3 shows a schematic block diagram of an encoder for encoding an audio signal according to an embodiment. The encoder includes an adaptive time-spectrum converter 26 and a controller 28. The adaptive time-spectrum converter 26 converts overlapping blocks of time values 30, including, for example, blocks 30' and 30'', into consecutive blocks of spectral values 4'. Further, the adaptive spectrum-time converter (6) receives control information (12) and switches between a first group of conversion kernels including one or more conversion kernels having different symmetries on both sides of the kernel and a second group of conversion kernels including one or more conversion kernels having the same symmetry on both sides of the conversion kernel. Further, the controller 2 8 is configured to control a time-spectrum converter to switch between a first group of conversion kernels of the conversion kernel and a second group of conversion kernels of the conversion kernel. Optionally, the encoder 22 includes, for the current frame, an output interface 32 for generating an encoded audio signal to generate the encoded audio signal, and control information 12 indicating the symmetry of the conversion kernel used to generate the current frame. The current frame may be the current block of consecutive blocks of spectral values. The output interface can include symmetry information between the current frame and a previous frame that is a frame independent of the current frame in the control data section of the current frame, or can be included in the control data section of the current frame. And when the current frame is a dependent frame, only the symmetry information of the current frame exists, and the symmetry information of the previous frame does not exist. The output interface can include symmetry information for the current frame and the previous frame in the control data section of the current frame, the current frame is an independent frame, or only the symmetry information of the current frame is included in the control data section of the current frame, and when the current frame is a dependent frame, the symmetry information of the previous frame is not included. An independent frame includes, for example, an independent frame header, whereby the current frame can be reliably read without knowledge of the previous frame. A dependent frame is, for example, an audio file having variable bitrate switching. Therefore, a dependent frame can be read only with knowledge of one or more previous frames. An independent frame includes, for example, an independent frame header, whereby the current frame can be reliably read without knowledge of the previous frame. A dependent frame is, for example, an audio file having variable bitrate switching. Therefore, a dependent frame can be read only with knowledge of one or more previous frames.

[0031] The controller can be configured to analyze the audio signal 24, for example, with respect to a fundamental frequency that is at least close to an integer multiple of the conversion frequency resolution. Thus, the control device can derive the control information 12 for supplying to the adaptive time-frequency transformer 26 and optionally the output interface 32 using the control information 12. The control information 12 can indicate an appropriate conversion kernel of the first group of conversion kernels or the second group of conversion kernels. The first group of conversion kernels may have one or more conversion kernels with odd symmetry on the left side of the kernel and even symmetry on the right side of the kernel, or vice versa, or the second group of conversion kernels can include one or more conversion kernels with even symmetry on both sides of the kernel or odd symmetry on both sides of the kernel. In other words, the first group of conversion kernels can include MDCT-IV conversion kernels or MDST-IV conversion kernels, and the second group of conversion kernels can include MDCT-II conversion kernels or MDST-II conversion kernels. To decode the encoded audio signal, the decoder can apply the respective inverse transformation to the conversion kernels of the encoder. Thus, the decoder can have the first group of conversion kernels include inverse MDCT-IV conversion kernels or inverse MDST-IV conversion kernels, or the second group of conversion kernels can include inverse MDCT-II conversion kernels or inverse MDST-II conversion kernels.

[0032] In other words, the control information 12 can include a current bit indicating the current symmetry for the current frame. Further, the adaptive spectrum-time transformer 6 may be configured not to switch from the first group of conversion kernels to the second group of conversion kernels when the current bit indicates the same symmetry as that used in the previous frame, and when the current bit indicates a symmetry different from that used in the previous frame, the adaptive spectrum-time transformer is configured to switch from the first group of conversion kernels to the second group of conversion kernels.

[0033] Furthermore, when the current bit exhibits a symmetry different from that used in the previous frame, the adaptive spectrum-time converter 6 can be configured not to switch from the second group of conversion kernels to the first group of conversion kernels, and when the current bit exhibits the same symmetry as that used in the previous frame, the adaptive spectrum-time converter is configured to switch from the second group of conversion kernels to the first group of conversion kernels. To illustrate the relationship between the time portion and the block on either the encoder side or the analysis side or the decoder side or the synthesis side, reference is made to FIGS. 4A and 4B.

[0034] To illustrate the relationship between the time portion and the block on either the encoder side or the analysis side or the decoder side or the synthesis side, reference is made to FIGS. 4A and 4B.

[0035] FIG. 4B shows a schematic diagram of the 0th time portion to the 3rd time portion, and each of these next time portions has a certain overlap range 170. Based on these time portions, a series of consecutive blocks representing the overlapping time portions are generated by a process described in more detail with respect to FIG. 5A showing the analysis side of the aliasing-introducing conversion operation.

[0036] In particular, the time domain signal shown in FIG. 4B when FIG. 4B is applied to the analysis side is windowed by the windowing unit 201 that applies an analysis window. Thus, to obtain the 0th time portion, for example, an analysis window is applied to 2048 samples, particularly samples 1 to 2048. Therefore, N is equal to 1024, the windowing has a length of 2N samples, and this example is 2048. Next, the windowing unit applies a further analysis operation to sample 1025 as the first sample within the block to obtain the first time portion, rather than sample 2049 as the first sample of the block. Thus, a first overlap range 170 with a 1024-sample length for 50% overlap is obtained. This procedure is additionally applied to the second and third time portions, but always overlaps to obtain a certain overlap range 170.

[0037] The overlap does not necessarily have to be a 50% overlap, but it should be emphasized that the overlap can be higher or lower and can also be a multi - overlap. That is, an overlap of two or more windows is obtained such that samples of the audio signal in the time domain do not contribute to two windows and as a result to a block of spectral values, but the samples contribute to two or more windows / blocks of spectral values. On the other hand, one skilled in the art will further understand that there are other windowing shapes applicable by the windowing section 201 of FIG. 5A with portions having a value of 0 and / or 1. For such portions having a single value, such portions typically overlap with the 0 portions of the preceding or subsequent windows, and thus, specific audio samples located in a certain portion of the window having a single value contribute only to a block of a single spectral value.

[0038] The windowed (windowed) time portion obtained by FIG. 4B is transmitted to the folder 202 to perform a convolution operation. This convolution operation can be performed, for example, at the output of the folder 202 such that only blocks of sampling values having N samples per block are present. Then, following the folding operation by the folder 202, a time - frequency converter is applied, and it is a DCT - IV converter that converts N samples per block on the input side into N spectral values on the output side of the time - frequency converter 203.

[0039] Accordingly, a series of blocks of spectral values obtained at the output of block 203 are shown in FIG. 4A. Specifically, it shows a first block 191 having a first modified value associated with 102 in FIGS. 1A and 1B and a second modified value 192 associated with the second modified value shown in FIGS. 1A and 1B. Of course, the sequence further has blocks 193 or 194 that precede the second block or precede the first block as shown. The first and second blocks 191, 192 are obtained, for example, by converting the windowed first time portion of FIG. 4B to obtain the first block, and the second b The lock is obtained by the time-frequency converter 203 of FIG. 5A converting the windowed second time portion of FIG. 4B. Thus, in a block of a series of spectral values, both blocks of spectrally adjacent spectral values represent an overlap range covering the first and second time portions.

[0040] Subsequently, FIG. 5B is described to show the processing on the synthesis side or decoder side of the result of the encoder or analysis side processing of FIG. 5A. The block of a series of spectral values output by the frequency converter 203 of FIG. 5A is input to the modifier 211. As outlined, each block of spectral values has N spectral values for the example shown in FIGS. 4A - 5B (note that this is different from the M used in equations (1) and (2)). Each block associates modification values such as 102, 104 shown in FIGS. 1A and 1B. Next, in a typical IMDCT operation or redundancy reduction synthesis transform, a frequency-time converter 212, a folder 213 for inverse convolution, a windowing section 214 for applying a synthesis window, and an overlap / add operation are shown by a block 215 that is executed to obtain a time domain signal within the overlap range. In this example, since there are 2N values per block, after each overlap-and-operation, if the modification values 102, 104 are not variable over time or frequency, N new non-aliased time domain samples are obtained. However, if these values vary with time and frequency, the output signal of block 215 is not alias-free, and this issue is addressed by the first and second aspects of the present invention as discussed in the context of FIGS. 1B and 1A and in the context of the other figures of this specification.

[0041] Subsequently, a further explanation of the procedure performed by the blocks of FIGS. 5A and 5B is given.

[0042] This figure is M Although exemplified by reference to the DCT, other aliasing-introducing transforms can be processed in a similar, analogous way. As an overlapping transform, the MDCT is somewhat unusual compared to other Fourier-related transforms in that it has an output that is half the input (rather than the same number). In particular, it is a linear function F:R 2N → R N (where R represents the set of real numbers). 2N real numbers x0,..., x2N-1 are transformed into N real numbers X0,..., XN-1 according to the following formula. JPEG0007708937000004.jpg17158

[0043] (The normalization factor before this transform, where the unity is an arbitrary convention and varies from process to process. Only the product of the normalizations of the following MDCT and IMDCT is constrained.)

[0044] The inverse MDCT is known as the IMDCT. At first glance, since the number of inputs and outputs is different, the MDCT may seem non-invertible. However, perfect reversibility is achieved by adding overlapping IMDCTs of adjacent overlapping blocks in time, canceling the errors, and retrieving the original data. This technique is known as time-domain aliasing cancellation (TDAC).

[0045] The IMDCT follows the following formula to transform N real numbers X0,..., XN-1 into 2N real numbers y0,..., y2N-1. JPEG0007708937000005.jpg23157

[0046] (As in the case of the orthogonal transform DCT-IV, the inverse function is also in the same form as the forward transform.)

[0047] In the case of a windowed MDCT (windowed MDCT) with a normal window (see below), the normalization factor before the IMDCT should be doubled (i.e., it becomes 2 / N).

[0048] In a typical signal compression application, the conversion characteristics are further improved by using window functions \(w_n\) (\(n = 0,\ldots,2N - 1\)) that are multiplied with \(x_n\) and \(y_n\) in the MDCT and IMDCT formulas, and to avoid discontinuities at the \(n = 0\) and \(2N\) boundaries, the functions are made to smoothly approach zero at these points. (That is, data is windowed before the MDCT and after the IMDCT.) In principle, \(x\) and \(y\) can have different window functions, and the window function can also be changed from one block to the next (especially when data blocks of different sizes are concatenated), but for simplicity, the general case of the same window function for blocks of equal size is considered.

[0049] JPEG0007708937000006.jpg88148

[0050] The window applied to the MDCT must satisfy the Princen - Bradley condition and is thus different from the windows used for other types of signal analysis. One reason for this difference is that the MDCT window is applied twice, once for both the MDCT (analysis) and the IMDCT (synthesis). As can be seen by examining the definitions, for \(N\) too, the MDCT is essentially equivalent to the DCT - IV where the input is shifted by \(N / 2\) and two \(N\) - block data are transformed at once. By more carefully examining this equivalence, important properties such as those of the TDAC can be easily derived.

[0051]

[0052] To define the exact relationship with the DCT - IV, it must be recognized that the DCT - IV corresponds to alternating the even / odd boundary conditions (i.e., symmetry conditions). It is odd at the left boundary (around \(n=-1 / 2\)), the right boundary line (around \(n = N=-1 / 2\)), and may continue in place of the periodic boundary like the DFT. This is according to the following equation. JPEG0007708937000007.jpg17137 and JPEG0007708937000008.jpg16137

[0053] Thus, if the input is an array x of length N, it can be imagined that this array is extended to (x, -xR, -x, xR, ...). Here, xR represents x in reverse order.

[0054] Consider an MDCT with 2N inputs and N outputs. Here, the input is divided into four blocks (a, b, c, d) of size N / 2. Shifting the MDCT definition by N / 2 positions to the right from the +N / 2 term, (b, c, d) extends beyond the end of the N DCT-IV inputs and they need to be "folded" according to the above boundary conditions.

[0055] Thus, the MDCT of the 2N inputs (a, b, c, d) is exactly equivalent to the DCT-IV of N inputs (-cR - d, a - bR).

[0056] This is illustrated for the window function 202 of FIG. 5A. a is part 204b, b is part 205a, c is part 205b, and d is part 206a.

[0057] (In this way, the algorithm for calculating the DCT-IV can be trivially applied to the MDCT.) Similarly, the above formula for the IMDCT is exactly 1 / 2 of the DCT-IV (its own inverse), the output is extended to length 2N (via the boundary conditions) and shifted back by N / 2 positions to the left. The inverse DCT-IV simply returns the input (-cR - d, a - bR) from above. When this is extended and shifted by the boundary conditions, IMDCT(MDCT(a,b,c,d)) = (a - bR, b - aR, c + dR, d + cR) / 2 results.

[0058] Thus, half of the IMDCT output is redundant such as b - aR = -(a - bR)R, and the same is true for the last two terms. Grouping the input into larger blocks A, B of size N with A = (a, b) and B = (c, d), this result can be presented in a simpler way IMDCT(MDCT(A, B)) = (A - AR, B + BR) / 2 can be written as.

[0059] One can come to understand the mechanism of TDAC. Assume that the MDCT of two temporally adjacent and 50% overlapping 2N blocks (B, C) is calculated. The IMDCT becomes (B - BR, C + CR) / 2 as above. When this is added to half of the previous IMDCT result, the reverse terms cancel out, simply obtaining B and recovering the original data.

[0060] The origin of the term "time-domain aliasing cancellation" is not currently clear. The use of input data extending beyond the boundary of the logical DCT-IV causes aliasing in the same way that frequencies above the Nyquist frequency are aliased to lower frequencies (with respect to extended symmetry), and it is not possible to distinguish the contribution of (a, b, c, d) to the MDCT from the contribution of bR, or equivalently, convert to the result of IMDCT(MDCT(a, b, c, d)) = (a - b R, b - aR, c + dR, d + cR) / 2. Combinations such as c - dR etc. have exactly the correct signs to cancel when the combination is added.

[0061] In the case of odd N (which is actually rarely used), since N / 2 is not an integer, the MDCT is not just a shift substitution of DCT-IV. In this case, an additional shift of half of the samples means that the MDCT / IMDCT becomes equivalent to DCT-III / II, and the analysis is the same as above.

[0062] As seen above, the MDCT of 2N inputs (a, b, c, d) is equivalent to the DCT-IV of N inputs (-cR - d, a - bR). Since the DCT-IV is designed for the case where the function at the right boundary is odd, the values near the right boundary become close to 0. If the input signal is smooth, in the input sequence (a, b, c, d), the components at the right ends of a and bR are continuous, so their difference is small. Let's look at the center of the interval. Rewriting the above equation as (-cR - d, a - bR) = (-d, a) - (b, c)R, the second (b, c)R is in the middle. However, in the first term (-d, a), there is a discontinuity where the right end of -d coincides with the left end of a. This is the reason for using a window function that reduces the components near the boundary of the input sequence (a, b, c, d) towards 0.

[0063] As described above, in the normal MDCT, the TDAC property is proven, and it is shown that the original data is recovered when adding half of the overlapping IMDCTs of temporally adjacent blocks. The derivation of this inverse property for the windowed MDCT (windowed MDCT) is only slightly more complicated.

[0064] JPEG0007708937000009.jpg34154

[0065] JPEG0007708937000010.jpg29156.

[0066] Therefore, instead of performing MDCT(A, B), there is currently an MDCT S (WA, W R B) where all multiplications are performed element-wise. When this is input to the IMDCT and multiplied again (element-wise) by the window function, the last half of N becomes as follows. W R ·(W R B + (W R B) R ) = W R ·(W R B + WB R ) = W R 2 B + WW R BR

[0067] (Since the normalization of the IMDCT is twice different in the windowed case, the multiplication does not become 1 / 2.)

[0068] Similarly, the windowed MDCT and IMDCT of (B,C) are as follows for the first half of N. W·(WB - W R B R ) = W 2 B - WW R B R

[0069] Adding these two halves together restores the original data. Reconstruction is possible even at the discontinuity of the window switch when the two overlapping window halves satisfy the Princen - Bradley condition. Aliasing cancellation can be performed in exactly the same way as above in this case. For multiple overlapping transforms, three or more branches are required using all relevant gain values. Similarly, the windowed MDCT and IMDCT of (B,C) are as follows for the first half of N.

[0070] So far, the MDCT, more specifically the symmetry or boundary conditions of MDCT - IV, have been described. The explanations are also valid for other transform kernels such as MDCT - II, MDST - II, and MDST - IV. However, it must be noted that different symmetries or boundary conditions of other transform kernels need to be considered.

[0071] Figure 6 schematically shows the implicit decimation - in - time characteristics and symmetries (i.e., boundary conditions) of the four described overlapping transforms. The transforms are derived from (2) via the first synthesis basis function for each of the four transforms. IMDCT - IV34a, IMDCT - II34b, IMDST - IV34c, and IMDST - II34d are shown as schematic diagrams of the amplitude samples over time. Figure 6 clearly shows the even and odd symmetries of the transform kernels at the axis of symmetry 35 (i.e., the folding point) between the transform kernels as described above.

[0072] The time-domain aliasing cancellation (TDAC) property indicates that when even and odd symmetric extensions are summed during OLA (overlap and add) processing, the aliasing is cancelled. In other words, for TDAC to occur, a transform with even left symmetry must follow a transform with odd right symmetry, and vice versa. Therefore, · After (inverse) MDCT-IV, follow with inverse MDCT-IV or inverse MDST-II. · After (inverse) MDST-IV, follow with inverse MDST-IV or inverse MDCT-II. · After (inverse) MDCT-II, follow with inverse MDCT-IV or inverse MDST-II. · After (inverse) MDST-II, follow with inverse MDST-IV or inverse MDCT-II.

[0073] (a) and (b) of FIG. 7 schematically show two embodiments of a use case in which signal-adaptive transform kernel switching is applied to the transform kernel from one frame to the next while allowing for complete reconstruction. In other words, two possible sequences of the above-described transformation sequence are illustrated in FIG. 7. Here, the solid line (such as line 38c) indicates the transform window, the dashed line 38a indicates the left aliasing symmetry of the transform window, and the dotted line 38b indicates the right aliasing symmetry of the transform window. Further, the symmetric peak indicates even symmetry, and the symmetric valley indicates odd symmetry. In FIG. 7(a), 36a of frame i and 36b of frame i+1 are MDCT-IV transform kernels, and in 36c of frame i+2, MST-II is used as the transition to the MDCT-II transform kernel used in 36d of frame i+3. 36e of frame i+4 re-uses MDST-II, and re-uses MDST-IV for MDCT-II of frame i+5, which is not shown in FIG. 7(a) for example. However, FIG. 7(a) clearly shows that the dashed line 38a and the dotted line 38b compensate for the subsequent transform kernel. In other words, when the left aliasing symmetry of the current frame and the right aliasing symmetry of the previous frame are added together, the sum of the dotted lines is equal to 0, so that complete time-domain aliasing cancellation (TDAC) is obtained. The left and right aliasing symmetries (or boundary conditions) are related to the convolution characteristics described in FIGS. 5A and 5B, for example, and the MDCT generates an output containing N samples from an input containing 2N samples as a result.

[0074] FIG. 7(b) is the same as FIG. 7(a), except that a different series of transform kernels is used for frames i to i+4. In frame i 36a, MDCT-IV is used, and in 36b of frame i+1, MDST-II is used as the transition to MDST-IV used in 36c of frame i+2. Frame i+3 uses the MDCT-II transform kernel as the transition from the MDST-IV transform kernel used in 36d of frame i+2 to the MDCT-IV transform kernel of 36e of frame i+4.

[0075] Table 1 shows the correlation determination matrix for the conversion sequence.

[0076] The embodiments further show how the adaptive transform kernel switching proposed in audio coders such as HE - AAC can be advantageously employed to minimize or avoid the two problems stated at the beginning. The following addresses harmonic signals that are quasi - optimally encoded by the conventional MDCT. The adaptive transition to MDCT - II or MDST - II may be performed by the encoder, for example, based on the fundamental frequency of the input signal. More specifically, when the pitch of the input signal is exactly or very close to an integer multiple of the frequency resolution of the transform (i.e., the bandwidth of one transform bin in the spectral domain), MDCT - II or MDST - II may be used for the affected frames and channels. However, a direct transition from MDCT - IV to the MDCT - II transform kernel is not possible or at least does not guarantee time - domain aliasing cancellation (TDAC). Therefore, MDCT - II must be used as the transitional transform between the two in such cases. Conversely, for the transition from MDST - II to the traditional MDCT - IV (i.e., switching to traditional MDCT coding), the intermediate MDCT - II is advantageous.

[0077] So far, the proposed adaptive transform kernel switching has been described for a single audio signal to enhance the coding of harmonic audio signals. Furthermore, it can be easily adapted to multi - channel signals such as, for example, stereo signals. Here, for example, when two or more channels of a multi - channel signal have a phase shift of approximately ±90 degrees with respect to each other, the adaptive transform kernel switching is also advantageous.

[0078] In the case of multi-channel audio processing, it may be appropriate to use MDCT-IV coding for one audio channel and MDST-IV coding for a second audio channel. In particular, this concept is advantageous when both audio channels contain a phase shift of approximately ±90 degrees prior to coding. Since MDCT-IV and MDST-IV give a 90-degree phase shift to the coded signal when compared to each other, a ±90-degree phase shift between two channels of an audio signal is compensated after coding, i.e., it is converted to a 0-degree or 180-degree phase shift by the 90-degree phase difference between the cosine-based function of MDCT-IV and the sine function of MDST-IV. Thus, for example, in M / S stereo coding, both channels of the audio signal may be coded with an intermediate signal, and in the case of the above-described conversion to a 0-degree phase shift, it is necessary to code only the minimum residual information in the side signal, and in the case of an inversion to a 180-degree phase shift, the reverse (minimum information of the intermediate signal) is obtained, thereby achieving maximum channel compression. This may achieve a bandwidth reduction of up to 50% while using a lossless coding scheme compared to classical MDCT-IV coding of both audio channels. Furthermore, it is also conceivable to use MDCT stereo coding in combination with complex stereo prediction. Both approaches calculate, code, and transmit a residual signal from two channels of an audio signal. Furthermore, complex prediction calculates prediction parameters for coding the audio signal, and the decoder decodes the audio signal using the transmitted parameters. However, for example, 2 As already described above, for encoding two audio channels, only information regarding the encoding method (MDCT-II, MDST-II, MDCT-IV or MDST-IV) should be transmitted so that the decoder can apply the relevant encoding method. Since complex stereo prediction parameters should be quantized using a relatively high resolution, the information regarding the encoding method used may be encoded, for example, in 4 bits. Theoretically, the first and second channels may each be encoded using one of four different encoding methods, leading to 16 different possible states.

[0079] Accordingly, FIG. 8 shows a schematic block diagram of a decoder 2 for decoding a multi-channel audio signal. Compared with the decoder of FIG. 1, the decoder further comprises a multi-channel processor 40 for receiving blocks of spectral values 4a''', 4b''' representing the first and second multi-channels, and in order to obtain processed blocks of spectral values 4a', 4b' of the first multi-channel and the second multi-channel, the received blocks are processed according to joint multi-channel processing techniques, the adaptive spectral-time processor uses the processed block 4b' for the second multi-channel using the control information 12a for the first multi-channel and the control information 12b for the second multi-channel to obtain the processed block 4b' for the second multi-channel using the control information 12a for the first multi-channel and the control information 12b for the second multi-channel, the first multi- It is configured to process the processed block 4a' of the channel. The multi-channel processor 40 may apply, for example, left and right stereo processing, sum-difference stereo processing, or the multi-channel processor may apply complex prediction using complex prediction control information related to blocks of spectral values representing the first and second multi-channels. Thus, the multi-channel processor can include or obtain a fixed preset from the control information indicating which processing was used, for example, to encode an audio signal. In addition to separate bits or words within the control information, the multi-channel processor can obtain this information from the current control information, for example, by the absence or presence of multi-channel processing parameters. In other words, the multi-channel processor 40 can apply an inverse operation to the multi-channel processing performed by the encoder to recover the separate channels of the multi-channel signal. Further multi-channel processing techniques are described with respect to FIGS. 10 to 14. Further, reference numerals are applied to multi-channel processing, and reference numerals extended by the letter "a" indicate the first multi-channel, and reference numerals are extended by the letter "b" to indicate the second multi-channel. Further, the multi-channel is not limited to two channels or stereo processing, but can be applied to three or more channels by expanding the illustrated processing for two channels.

[0080] According to an embodiment, the multi-channel processor of the decoder can process the received block according to the joint multi-channel processing technique. Further, the received block can include the encoded residual signal of the first multi-channel representation and the second multi-channel representation. Further, the multi-channel processor may be configured to calculate the first multi-channel signal and the second multi-channel signal using the residual signal and the further encoded signal. In other words, the residual signal may be the side signal of the audio signal encoded in M / S, or the residual between the prediction of channel and channel of the audio signal based on further channels of the audio signal during use, for example, a complex stereo prediction. Thus, the multi-channel processor can convert the M / S or complex predicted audio signal into an L / R audio signal for further processing such as applying an inverse transform kernel. Therefore, the multi-channel processor can use the residual signal and an intermediate signal of the M / S encoded audio signal or a further encoded audio signal which may be a channel of the audio signal (e.g., MDCT encoded).

[0081] FIG. 9 shows the encoder 22 of FIG. 3 extended for multi-channel processing. Control information 12 Although it is predicted to be included in the encoded audio signal 4, the control information 12 may be further transmitted, for example, using a separate control information channel. The controller 28 of the multi-channel encoder can analyze the overlapping blocks of the time values 30a, 30b of the audio signal having the first channel and the second channel in order to determine the conversion kernels of the frames of the first channel and the corresponding frames of the second channel. Thus, the controller can try each combination of conversion kernels to derive options for the conversion kernels that minimize, for example, the M / S encoded or complex prediction residual signal (or the side signal with respect to M / S encoding). The minimized residual signal generates, for example, the residual signal having the lowest energy compared to the remaining residual signals. This is advantageous, for example, when further quantization of the residual signal uses fewer bits to quantize small signals compared to quantizing larger signals. Further, the controller 28 can determine the first control information 12a of the first channel and the second control information 12b of the second channel input to the adaptive time-frequency transformer 26 to which one of the aforementioned conversion kernels is applied. Thus, the time-frequency transformer 26 may be configured to process the first channel and the second channel of the multi-channel signal. Further, the multi-channel encoder can further include a multi-channel processor 42 for processing consecutive blocks of the spectral values 4a', 4b' of the first channel and the second channel, for example, using joint multi-channel processing techniques such as the following. For example, processed blocks of the spectral values 40a''', 40b''' can be obtained using sum-difference stereo encoding or complex prediction. The encoder can further include an encoding processor 46 for processing the processed blocks of the spectral values to obtain the encoded channels 40a''', 40b'''.The symbolization processor can encode an audio signal using, for example, a lossy audio compression or a lossless audio compression method, and can apply, for example, scalar quantization of spectral lines, entropy encoding, Huffman encoding, channel encoding, block encoding or convolutional encoding, or forward error correction or automatic repeat request. Further, irreversible audio compression may refer to using quantization based on a psychoacoustic model.

[0082] According to a further embodiment, a block of first processed spectral values represents a first encoded representation of joint multi-channel processing technology, and a block of second processed spectral values represents a second encoded representation of joint multi-channel processing technology. Thus, the symbolization processor 46 is configured to process the first processed block using quantization and entropy encoding to form a first encoded representation, and to process the second processed block using quantization and entropy encoding to form a second encoded representation. The first encoded representation and the second encoded representation may be formed within a bitstream representing the encoded audio signal. In other words, the first processing block can include an M / S encoded audio signal or an intermediate signal of an MDCT encoded channel of the encoded audio signal using complex stereo prediction. Further, the second processing block can include parameters or residual signals for complex prediction, or side signals of the M / S encoded audio signal.

[0083] Figure 10 shows an audio encoder for encoding a multi-channel audio signal 200 having two or more channel signals. The first channel signal is indicated by reference numeral 201 and the second channel is indicated by reference numeral 202. Both signals are input to an encoder calculator 203 for calculating a first composite signal 204 and a prediction residual signal 205 using the first channel signal 201, the second channel signal 202, and prediction information 206, resulting in the prediction residual signal 205. At this time, when combined with the prediction signal obtained from the first composite signal 204 and the prediction information 206, a second composite signal is obtained. In this regard, the first composite signal and the second composite signal can be derived from the first channel signal 201 and the second channel signal 202 using a combination rule.

[0084] The prediction information is generated by an optimizer 207 for calculating the prediction information 206 such that the prediction residual signal satisfies an optimization target 208. The first composite signal 204 and the residual signal 205 are input to a signal encoder 209 to encode the first composite signal 204, obtaining an encoded first composite signal 210, and encoding the residual signal 20 to obtain an encoded residual signal 211. Both the encoded signals 210 and 211 are input to an output interface 212 to obtain an encoded multi-channel signal 213 by combining the encoded first composite signal 210, the encoded prediction residual signal 211, and the prediction information 206.

[0085] Depending on the implementation, the optimizer 207 receives either the first channel signal 201 or the second channel signal 202, or as indicated by lines 214 and 215, the first composite signal 214 and the second composite signal 215 are obtained from a combiner 2031 in FIG. 11A described below.

[0086] FIG. 10 shows an optimization target in which the encoding gain is maximized, i.e., the bit rate is reduced as much as possible. In this optimization target, the residual signal D is minimized with respect to α. In other words, the prediction information α is ||S - αM||2 is selected so that This results in the solution for α shown in FIG. 10. The signals S, M are given in blocks and are spectral domain signals, with the notation ||...|| denoting the 2-norm of the arguments, and <...> denoting the dot product as usual. When the first channel signal 201 and the second channel signal 202 are input to the optimizer 207, the optimizer needs to apply a combination rule, and an exemplary combination rule is shown in FIG. 11C. However, when the first composite signal 214 and the second composite signal 215 are input to the optimizer 207, the optimizer 207 does not need to implement a combination rule by itself.

[0087] Other optimization targets may be related to perceptual quality. The optimization goal may be to obtain maximum perceptual quality. The optimizer then requires additional information from the perceptual model. Other implementations of optimization targets relate to obtaining a minimum or constant bit rate. The optimizer 207 is then implemented to perform a quantization / entropy coding operation to determine the bit rate required for a particular α value. Thus, α can be set to meet requirements such as a minimum or constant bit rate. Other implementations of optimization targets may relate to a minimum use of encoder or decoder resources. In case of such implementation of optimization targets, information about resources required for a certain optimization is available in the optimizer 207. Furthermore, combinations of these optimization targets or other optimization targets can be applied to control the optimizer 207 that calculates the prediction information 206.

[0088] The encoder calculator 203 in FIG. 10 can be implemented in different ways. An exemplary first embodiment is shown in FIG. 11A, where explicit combination rules are executed in the combiner 2031. An alternative exemplary implementation using the matrix computer 2039 is shown in FIG. 11B. The combiner 2031 in FIG. 11A may be implemented to execute the combination rules illustrated in FIG. 11C, which are well-known intermediate-side coding rules, and a weighting coefficient of 0.5 is applied to all branches. However, depending on the implementation, other weighting coefficients or no weighting coefficients at all can be implemented. Furthermore, it is also possible to apply other combination rules such as other linear combination rules or non-linear combination rules, and as long as there are corresponding inverse combination rules that can be applied to the decoder combiner 1162 shown in FIG. 12A, the encoder applies a combination rule that is the inverse of the combination rule applied by the encoder. For joint stereo prediction, since the influence on the waveform is "balanced" by the prediction, i.e., the error is included in the transmitted residual signal, any reversible prediction rule can be used. This is because the prediction operation between the encoder calculator 203 by the optimizer 207 is a waveform preservation process.

[0089] The combiner 2031 outputs a first combined signal 204 and a second combined signal 2032. The first combined signal is input to the predictor 2033, and the second combined signal 2032 is input to the residual calculator 2034. The predictor 2033 calculates a predicted signal 2035, which is combined with the second combined signal 2032 to finally obtain a residual signal 205. Specifically, the combiner 2031 is configured to combine two channel signals 201 and 202 of a multi-channel audio signal in two different ways to obtain a first combined signal 204 and a second combined signal 2032, and the two different ways are shown in the exemplary embodiment of FIG. 11C. The predictor 2033 is configured to apply prediction information to the first combined signal 204 or a signal obtained from the first combined signal to obtain the predicted signal 2035. The signal obtained from the combined signal can be derived by any non-linear or linear operation, and can be realized using a linear filter such as an FIR filter that performs weighted addition of certain values. A conversion from a real number to an imaginary number / a conversion from an imaginary number to a real number is advantageous.

[0090] The residual calculator 2034 in FIG. 11A can perform a subtraction operation such that the predicted signal 2035 is subtracted from the second combined signal. However, other operations in the remaining computers are also possible. Correspondingly, the combined signal calculator 1161 in FIG. 12A can perform an addition operation in which the decoded residual signal 114 and the predicted signal 1163 are added to obtain a second combined signal 1165.

[0091] The decoder calculator 116 can be implemented in different ways. The first implementation is shown in FIG. 12A. This embodiment includes a predictor 1160, a composite signal calculator 1161, and a combiner 1162. The predictor receives the decoded first composite signal 112 and the prediction information 108 and outputs a prediction signal 1163. Specifically, the predictor 1160 is configured to apply the prediction information 108 to the decoded first composite signal 112 or a signal derived from the decoded first composite signal. The derivation rule for deriving the signal to which the prediction information 108 is applied may be a conversion from a real number to an imaginary number, equivalently, an imaginary-real conversion or a weighting operation, or to the same extent, depending on the implementation, a phase shift operation, or a combined weighting / phase shift operation. The prediction signal 1163 is input to the composite signal calculator 1161 together with the decoded residual signal to calculate the decoded second composite signal 1165. The signals 112 and 1165 are respectively input to a combiner 1162 that combines the decoded first and second composite signals to obtain a decoded multi-channel audio signal having the decoded first channel signal and the decoded second channel signal on output lines 1166 and 1167. Alternatively, the decoder calculator is implemented as a matrix calculator 1168 that receives as inputs the decoded first composite signal or signal M, the decoded residual signal or signal D, and the prediction information α108. The matrix operator 1168 applies a transformation matrix shown as 1169 to the signals M and D to obtain output signals L and R. Here, L is the decoded first channel signal and R is the decoded second channel signal. The notation in FIG. 12B is similar to the stereo notation using the left channel L and the right channel R. This notation is applied for ease of understanding, but it is obvious to those skilled in the art that the signals L and R can be any combination of two channel signals within a multi-channel signal having three or more channel signals. The matrix operation 1169 unifies the operations of the blocks 1160, 1161, and 1162 in FIG. 12A into a kind of "single-shot" matrix calculation, and the inputs to the circuit in FIG. 12A and the outputs from the circuit in FIG. 12A are respectively the same as the inputs to the matrix operator 1168 and the outputs from the matrix operator 1168.

[0092] FIG. 12C shows an example of an inverse coupling rule applied by the combiner 1162 of FIG. 12A. In particular, the coupling rule is similar to the decoder-side coupling rule in well-known mid-side coding where L = M + S and R = M - S. It should be understood that the signal S used by the inverse coupling rule of FIG. 12C is the signal calculated by the composite signal calculator, that is, the combination of the predicted signal on line 1163 and the decoded residual signal on line 114. In this specification, the signals on the lines may sometimes be named by the reference numbers of the lines and sometimes be indicated by the reference numbers themselves resulting from the lines. Therefore, the notation where a line having a certain signal indicates the signal itself. The lines can be the physical lines of a hard-wired implementation. However, in a computerized implementation, there are no physical lines, but the signals represented by the lines are transmitted from one calculation module to another.

[0093] FIG. 13A shows an implementation of an audio encoder. Compared to the audio encoder shown in FIG. 11A, the first channel signal 201 is the spectral representation of the first channel signal 55a in the time domain. Similarly, the second channel signal 202 is the spectral representation of the time domain channel signal 55b. The conversion from the time domain to the spectral representation is performed by a time / frequency converter 50 for the first channel signal and a time / frequency converter 51 for the second channel signal. The spectral converters 50, 51 are preferably implemented as real number converters, but this is not necessarily the case. The conversion algorithm can be a discrete cosine transform, an FFT transform where only the real part is used, an MDCT, or other transforms that provide real-valued spectral values. Alternatively, both conversions can be performed as imaginary transforms such as DST, MDST, or FFT where only the imaginary part is used and the real part is discarded. Other transforms that provide only imaginary values can be used as well. One purpose of using a pure real-valued transform or a pure imaginary transform is computational complexity, because for each spectral value, only a single value such as magnitude or real part has to be processed, or alternatively, phase or imaginary part has to be processed. In contrast to a fully complex transform such as FFT, two values, i.e., the real and imaginary parts of each spectral line, have to be processed, which is an increase in computational complexity by at least two factors. Another reason to use a real-valued transform here is that such a transform sequence is typically critically sampled even in the presence of an inverse transform overlap, and thus provides a suitable (and commonly used) region for signal quantization and entropy coding (the standard “perceptual audio coding” paradigm implemented in “MP3”, AAC, or similar audio coding systems).

[0094] FIG. 13A further shows a residual calculator 2034 as an adder that receives a side signal at the "plus" input and a prediction signal output by a predictor 2033 at the "minus" input. Further, FIG. 13A shows a situation where predictor control information is transmitted to a multiplexer 212 that outputs a multiplexed bitstream representing a multi-channel audio signal encoded from an optimizer. In particular, the prediction operation is performed such that a side signal is predicted from an intermediate signal, as shown by the equation on the right side of FIG. 13A.

[0095] The predictor control information 206 is a factor as shown on the right side of FIG. 11B. In an embodiment where the prediction control information includes only the real part, such as the real part of the complex value α or the magnitude of the complex value α, when this part corresponds to a non-zero factor, the waveform structures of the intermediate signal and the side signal are similar, but a significant coding gain can be obtained when the amplitudes are different.

[0096] However, when the prediction control information includes only a second part that can be the imaginary part of a complex factor or the phase information of a complex factor, and the imaginary part or the phase information is different from zero, the present invention achieves a significant coding gain for signals that are phase-shifted from each other by a value different from 0 degrees or 180 degrees, and has similar waveform characteristics and similar amplitude relationships except for the phase shift. are.

[0097] The predictor control information is a complex value. And for signals with different amplitudes and phase shifts, a significant coding gain can be obtained. In a situation where the time / frequency conversion provides a complex spectrum, operation 2034 is a complex operation in which the real part of the predictor control information is applied to the real part of the complex spectrum M, and the imaginary part of the complex prediction information is applied to the imaginary part of the complex spectrum. Next, in the adder 2034, the result of this prediction operation is the predicted real spectrum and the predicted imaginary spectrum. The predicted real spectrum is subtracted from the real spectrum (band unit) of the secondary signal S, and the predicted imaginary spectrum is subtracted from the imaginary part of the spectrum of S to obtain a complex residual spectrum D.

[0098] The time-domain signals L and R are real-valued signals, but the frequency-domain signals can be either real or complex-valued. When the frequency-domain signal is real-valued, the conversion is a real-valued conversion. When the frequency-domain signal is complex, the conversion is a complex conversion. This means that the input to the time-frequency conversion and the output of the frequency-time conversion are real-valued, and the frequency-domain signal becomes, for example, a complex-valued QMF domain signal.

[0099] Figure 13B shows an audio decoder corresponding to the audio encoder shown in Figure 13A.

[0100] The bitstream output by the bitstream multiplexer 212 in Figure 13A is input to the bitstream demultiplexer 102 in Figure 13B. The bitstream demultiplexer 102 separates the bitstream into a downmix signal M and a residual signal D. The downmix signal M is input to the inverse quantizer 110a. The residual signal D is input to the inverse quantizer 110b. Further, the bitstream demultiplexer 102 demultiplexes the predictor control information 108 from the bitstream and inputs it to the predictor 1160. The predictor 1160 outputs a predicted side signal α·M, and the combiner 1161 synthesizes the residual signal output by the inverse quantizer 110b with the predicted side signal to finally obtain the reconstructed side signal S. Next, the side signal is input to a combiner 1162 that performs, for example, sum-difference processing as shown in Figure 12C for mid / side encoding. Specifically, block 1162 performs (inverse) mid / side decoding to obtain the frequency-domain representation of the left channel and the frequency-domain representation of the right channel. Next, the frequency-domain representation is converted to a time-domain representation by the corresponding frequency / time converters 52 and 53.

[0101] Depending on the implementation of the system, when the frequency-domain representation is a real-valued representation, the frequency / time converters 52, 53 are real-valued frequency / time converters, and when the frequency-domain representation is a complex-valued representation, they are complex-valued frequency / time converters.

[0102] However, to increase efficiency, it is advantageous to perform a real-valued transform, as shown in another embodiment in Fig. 14A for the encoder and in Fig. 14B for the decoder. The real-valued transforms 50 and 51 are realized by MDCT, i.e. MDCT-IV, or according to the invention MDCT-II or MDST-II or MDST-IV. Also, the prediction information is calculated as a complex value having a real part and an imaginary part. Since both spectra M, S are real-valued spectra, therefore there is no imaginary part of the spectrum, a real-to-imaginary converter 2070 is provided, which calculates an estimated imaginary spectrum 600 from the real spectrum of the signal M. This real-to-imaginary converter 2070 is part of the optimizer 207, where the imaginary spectrum 600 estimated in block 2070 is input together with the real spectrum M to an α optimizer stage 2071, which calculates the prediction information 206 with a real-valued factor indicated at 2073 and an imaginary factor indicated at 2074. Here, according to this embodiment, the real-valued spectrum of the first composite signal M is the side spectrum of the real part To obtain the predicted signal to be subtracted from R It is multiplied by 2073. The number spectrum 600 has an imaginary part α I is multiplied with to obtain a further prediction signal 13A, this prediction signal is then subtracted from the real-valued side spectrum as shown in 2034b. The prediction residual signal D is then quantized in quantizer 209b and the real-valued spectrum of M is quantized / encoded in block 209a. Furthermore, it is advantageous to quantize and code the prediction information α in quantizer / entropy encoder 2072 to obtain a coded complex α value that is transmitted to the bitstream multiplexer 212 of Fig. 13A, for example to be finally input as prediction information in the bitstream.

[0103] Regarding the location of the quantization / coding (Q / C) module 2072 for α, it should be noted that multipliers 2073 and 2074 use exactly the (quantized) α that is also used in the decoder. Therefore, 2072 is directly transferred to the output of 2071. Alternatively, it can be assumed that the quantization of α is already taken into account in the optimization process of 2071.

[0104] Although a complex spectrum could be calculated at the encoder side, since all the information is available, it is advantageous to perform a real to complex transform in block 2070 of the encoder, so that a similar condition for the decoder shown in Fig. 14B is generated. The decoder receives the real-valued coded spectrum of the first synthesis signal and a real-valued spectral representation of the coded residual signal. Furthermore, the coded complex prediction information is obtained in 108, which is entropy decoded and dequantized in block 65 to produce the real part αα shown in 1160b. R and the imaginary part α shown in 1160c I The weighting elements 1160b and 11 The intermediate signal output by 1160c is added to the decoded and dequantized prediction residual signal. In particular, the spectral values input to the weighter 1160c, whose weighting coefficient is the imaginary part of the complex prediction coefficient, are derived from the real-valued spectrum M by the real-to-imaginary converter 1160a, which is implemented in the same way as in block 2070 of Fig. 20 for the encoder side. On the decoder side, the complex-valued representation of the intermediate signal or the side signal is not available. In contrast to the encoder side, the reason is that only the coded real-valued spectrum was transmitted from the encoder to the decoder for bitrate and complexity reasons.

[0105] The real-to-imaginary transformer 1160a or corresponding block 2070 of FIG. 14A can be implemented as disclosed in WO 2004 / 013839 pamphlet or WO 2008 / 014853 pamphlet or U.S. Patent No. 6,980,933. Alternatively, any other implementation known in the art can be applied.

[0106] The embodiments further show how the proposed adaptive transform kernel switching can be advantageously used in audio codecs such as HE-AAC to minimize or avoid the two problems stated in the "Problem Statement" section. In the following, a stereo signal with an inter-channel phase shift of about 90 degrees is addressed. Here, the switch to MDST-IV-based coding can be used in one of the two channels, while the legacy MDCT-IV coding can be used in the other channel. Alternatively, MDCT-II coding can be used in one channel and MDST-II coding can be used in the other channel. Assuming that the cosine and sine functions are deformed with a 90-degree phase shift from each other (cos(x)=sin(x+π / 2)), the corresponding phase shift between the input channel spectra can thus be converted to a 0-degree or 180-degree phase shift, which can be very efficiently coded via the conventional M / S-based joint stereo coding. Similar to the case of a harmonic signal optimally coded with a conventional MDCT, it may be advantageous in the channel affected by the intermediate transition transform.

[0107] In both cases, for harmonic and stereo signals with an inter-channel phase shift of approximately 90 degrees, the encoder selects one of the four kernels for each transformation (see also Figure 7). Each decoder applying the transform kernel switching of the present invention can appropriately reconstruct the signal using the same kernel. In order for such a decoder to know which transform kernel to use for one or more inverse transforms within a given frame, side information explaining the selection of the transform kernel, or left-right symmetry, should be transmitted by the corresponding encoder at least once per frame. The next section explains the integration (i.e., modification) into the MPEG-H 3D audio codec 。

[0108] Further embodiments relate to audio coding, and in particular to low-rate perceptual audio coding using lap transforms such as the modified discrete cosine transform (MDCT). The embodiments address two specific problems with conventional transform coding by generalizing the MDCT coding principle to include three other similar transforms. The embodiments further show signal adaptation and context-adaptive switching between these four transform kernels in each encoded channel or frame, or for each transform in each encoded channel or frame. Each side information may be transmitted in the encoded bitstream to signal the kernel selection to the corresponding decoder.

[0109] Figure 15 shows a schematic block diagram of a method 1500 for decoding an encoded audio signal. Method 1500 includes step 1505 of converting consecutive blocks of spectral values into overlapping consecutive blocks of time values, step 1510 of adding by overlapping consecutive blocks of time values to obtain decoded audio values, and step 1515 of switching between a first group of transformation kernels including one or more transformation kernels having different symmetries on both sides of the kernel and a second group of transformation kernels including one or more transformation kernels having the same symmetry on both sides of the kernel in response to receiving control information.

[0110] FIG. 16 shows a schematic block diagram of a method 1600 for encoding an audio signal. Method 1600 includes step 1605 of converting overlapping blocks of time values into consecutive blocks of spectral values, step 1610 of controlling the time-frequency transformation to switch between the transformation kernels of the first group of transformation kernels and the transformation kernels of the second group of transformation kernels, and step 1615 of switching between a first group of transformation kernels including one or more transformation kernels having different symmetries on both sides of the kernel and a second group of transformation kernels including one or more transformation kernels having the same symmetry on both sides of the kernel in response to receiving control information.

[0111] In this specification, it should be understood that signals on a line may sometimes be named by the reference number of the line and sometimes be indicated by the reference number itself due to the line. Thus, a notation where a line having a certain signal indicates the signal itself. The line can be a physical line in a hardwired implementation. However, in a computerized implementation, there is no physical line, but the signal represented by the line is transmitted from one computing module to another computing module.

[0112] The present invention has been described in the context of block diagrams where blocks represent actual or logical hardware components, but the present invention can also be implemented by a computer-implemented method. In the latter case, the blocks represent corresponding method steps, and these steps represent functions that are executed by corresponding logical or physical hardware blocks.

[0113] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent descriptions of corresponding methods where the blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent descriptions of corresponding blocks or corresponding items or features of a device. Some or all of the method steps may be executed (or used) by a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of some of the most important method steps can be executed by such a device.

[0114] The transmitted or encoded signals of the present invention can be stored on a digital storage medium or transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0115] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. The implementation can be carried out using a digital storage medium such as a floppy disk, a DVD, a Blu-ray, a CD, a ROM, a PROM, and an EPROM, an EEPROM, or a flash memory on which electronically readable control signals are stored, and on which they cooperate (or can cooperate) with a programmable computer system such that the respective methods are executed. Thus, the digital storage medium may be computer-readable.

[0116] Some embodiments according to the present invention comprise a data carrier having an electronically readable control signal that can cooperate with a programmable computer system, and one of the methods described herein is executed.

[0117] In general, embodiments of the present invention can be implemented as a computer program product having program code that operates to execute one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a machine-readable readable carrier.

[0118] Other embodiments include a computer program for executing one of the methods described herein and is stored on a machine-readable carrier.

[0119] In other words, an embodiment of the method of the present invention is a computer program having program code for executing one of the methods described herein when the computer program is executed on a computer.

[0120] Accordingly, a further embodiment of the method of the present invention includes a data carrier (or a non-transitory storage medium such as a digital storage medium or a computer-readable medium) that records a computer program for executing one of the methods described herein. The data carrier, digital storage medium or recording medium is typically tangible and / or non-transitory.

[0121] Accordingly, a further embodiment of the method of the present invention is a data stream or a sequence of signals representing a computer program for executing one of the methods described herein. The data stream or sequence of signals can be configured to be transmitted, for example, via a data communication connection, such as via the Internet.

[0122] Further embodiments include processing means, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.

[0123] Further embodiments include a computer having installed thereon a computer program for performing one of the methods described herein.

[0124] Further embodiments according to the present invention include an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, or the like. The apparatus or system may include, for example, a file server for transmitting the computer program to the receiver.

[0125] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.

[0126] The above-described embodiments are merely illustrative of the principles of the present invention. It will be understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Accordingly, it is intended to be limited only by the scope of the impending claims and not by the specific details shown in the description and explanation of the embodiments herein.

[0127] References [1] H. S. Malvar, Signal Processing with Lapped Transforms, Norwood: Artech House, 1992. [2] J. P. Princen and A. B. Bradley, "Analysis / Synthesis Filter Bank Design Based on Time Domain Aliasing Cancellation," IEEE Trans. Acoustics, Speech, and Signal Proc., 1986. [3] J. P. Princen, A. W. Johnson, and A. B. Bradley, "Subband / transform coding using filter bank design based on time domain aliasing cancellation," in IEEE ICASSP, vol. 12, 1987. [4] H. S. Malvar, "Lapped Transforms for Efficient Transform / Subband Coding," IEEE Trans. Acoustics, Speech, and Signal Proc., 1990. [5] http: / / en.wikipedia.org / wiki / Modified_discrete_cosine_transform

Claims

1. A decoder (2) for decoding a coded audio signal (4), wherein the decoder comprises: an adaptive spectrum-time converter (6) for converting blocks of consecutive spectral values (4', 4'') into blocks of consecutive time values (10); and an overlap-add processor (8) for obtaining decoded audio values (14) by overlapping and adding consecutive blocks of the time values (10), wherein the adaptive spectrum-time converter (6) receives control information (12) and is configured to switch between a first conversion kernel group having conversion kernels with different symmetries on both sides, the conversion kernels of the first conversion kernel group including one or more conversion kernels having odd symmetry on the left side and even symmetry on the right side, or vice versa, and a second conversion kernel group having conversion kernels with equal symmetries on both sides, the conversion kernels of the second conversion kernel group including one or more conversion kernels having even symmetry or odd symmetry on both sides; a decoder.

2. The first conversion kernel group has one or more conversion kernels with odd symmetry on the left side and even symmetry on the right side of the kernel, or vice versa, or The second conversion kernel group has one or more conversion kernels with even symmetry or odd symmetry on both sides of the kernel, The decoder (2) according to claim 1.

3. The first conversion kernel group includes an inverse MDCT-IV conversion kernel or an inverse MDST-IV conversion kernel, or The second conversion kernel group includes an inverse MDCT-II conversion kernel or an inverse MDST-II conversion kernel, The decoder (2) according to claim 1.

4. The conversion kernels of the first conversion kernel group and the second conversion kernel group are based on wherein at least one conversion kernel of the first conversion kernel group has a parameter wherein at least one conversion kernel of the second conversion kernel group has a parameter cs( ) = cos( ) and k 0 = 0.5, or cs( ) = sin( ) and k 0 is based on 0.5, or The decoder (2) according to claim 1. cs( ) = cos( ) and k 0 = 0, or cs( ) = sin( ) and k 0 is based on = 1, Here, x i,n is the time-domain output, C is a constant parameter, N is the time window length, spec is the spectral value having M values for the block, M is equal to N / 2, i is the time block index, k is the spectral index indicating the spectral value, n is the time index indicating the time value in block i, n 0 is a constant parameter that is an integer or zero

5. The control information (12) includes a current bit indicating the current symmetry for the current frame, ​ When the current bit of the adaptive spectrum-time converter (6) exhibits the same symmetry as that used in the previous frame, it is configured not to switch from the first conversion kernel group to the second conversion kernel group. When the current bit of the adaptive spectrum-time converter (6) exhibits a symmetry different from that used in the previous frame, it is configured to switch from the first conversion kernel group to the second conversion kernel group. Decoder (2) according to claim 1.

6. When the current bit indicating the current symmetry of the current frame of the adaptive spectrum-time converter (6) exhibits the same symmetry as that used in the previous frame, it is configured to switch from the second conversion kernel group to the first conversion kernel group. When the current bit indicating the current symmetry of the current frame of the adaptive spectrum-time converter (6) is different in contrast from that used in the previous frame, it is configured not to switch from the second conversion kernel group to the first conversion kernel group. Decoder (2) according to claim 1.

7. The adaptive spectrum-time converter (6) is configured to read the control information (12) of the previous frame from the encoded audio signal (4), and the control information (12) of the current frame from the encoded audio signal (4) within the control data section of the current frame following the previous frame, or The adaptive spectrum-time converter (6) is configured to read the control information (12) from the control data section of the current frame, and retrieve the control information (12) of the previous frame from the control data section of the previous frame or from the decoder settings applied to the previous frame. Decoder (2) according to claim 1.

8. The adaptive spectrum-time converter (6) is configured to apply a conversion kernel based on the following table. Here, symm i is the control information (12) of the current frame at index i, and symm i-1 is the control information (12) of the previous frame at index i -1 ​ Decoder (2) according to claim 1. **Claim 9**: A multi-channel processor (40) further comprising receiving a block of spectral values representing a first multi-channel and a second multi-channel, processing the received block according to a combined multi-channel processing technique to obtain a block of processed spectral values for the first multi-channel and the second multi-channel, wherein the adaptive spectrum-time converter (6) is configured to process the block of processed spectral values for the first multi-channel using control information (12) for the first multi-channel and to process the block of processed spectral values for the second multi-channel using control information (12) for the second multi-channel. The decoder (2) according to claim 1. **Claim 10** The decoder (2) according to claim 9, wherein the multi-channel processor (40) is configured to apply complex prediction using complex prediction control information associated with the block of spectral values representing the first and second multi-channels. **Claim 11** The decoder (2) according to claim 9, wherein the multi-channel processor (40) is configured to process the received block according to the combined multi-channel processing technique, the received block including an encoded residual signal of the representation of the first multi-channel and the representation of the second multi-channel, and the multi-channel processor (40) is configured to calculate a block of processed spectral values for the first multi-channel and a block of processed spectral values for the second multi-channel using the encoded residual signal and another encoded signal. **Claim 12**: The first conversion kernel group includes an inverse MDCT-IV conversion kernel or an inverse MDST-IV conversion kernel, or the second conversion kernel group includes an inverse MDCT-II conversion kernel or an inverse MDST-II conversion kernel. The MDCT-IV conversion kernel is odd-symmetric on the left side and even-symmetric on the right side, and the composite signal is inverted on the left side during the signal convolution of this conversion, or The MDST-IV conversion kernel is even-symmetric on the left side and odd-symmetric on the right side, and the composite signal is inverted on the right side during the signal convolution of this conversion, or The MDCT-II transform kernel is even-symmetric on the left side and even-symmetric on the right side, and the composite signal is not inverted on either side during the signal convolution of this transform, or The MDST-II transform kernel is odd-symmetric on the left side and odd-symmetric on the right side, and the composite signal is inverted on both sides during the signal convolution of this transform, The decoder (2) according to claim 1.

13. The multi-channel processor (40) is configured to perform combined stereo processing or combined processing of two or more channels as the combined multi-channel processing technology, and the multi-channel signal has two channels or two or more channels. The decoder (2) according to claim 9.

14. The adaptive spectrum-time converter (6) is configured to use the transform kernel of the second transform kernel group for an encoded signal representing a harmonic signal whose pitch is at least approximately equal to an integer multiple of the frequency resolution of the transform, or The adaptive spectrum-time converter (6) is configured to use an MDST-IV based transform kernel for one of the two channels represented by the encoded signal, and an MDCT-IV based transform kernel for the second of the two channels. The decoder (2) according to claim 1.

15. An encoder (22) for encoding an audio signal (24), The encoder, An adaptive time-spectrum converter (26) for converting a block of overlapping time values into a block of consecutive spectral values (4', 4''), and A controller (28) for controlling the adaptive time-spectrum converter (26) to switch between the transform kernels of the first transform kernel group and the transform kernels of the second transform kernel group Including, The adaptive time-spectrum converter (26) receives control information (12) and, in response to the control information (12), the transform kernels of the first transform kernel group including one or more transform kernels with different symmetries on both sides of the transform kernels of the first transform kernel group and the second transform kernel group including one or more transform kernels with equal symmetries on both sides of the transform kernels of the second transform kernel group. It is configured to switch between the transform kernels. Encoder (22).

16. The encoder (22) according to claim 15, further comprising an output interface (32) for generating an encoded audio signal (4) having control information (12) indicating the symmetry of the conversion kernel used to generate the current frame for the current frame.

17. The output interface (32) is configured to include symmetry information of the current frame and the previous frame in the control data section of the current frame when the current frame is an independent frame, or when the current frame is a dependent frame, the control data section of the current frame is configured to include only the symmetry information of the current frame and not to include the symmetry information of the previous frame. The encoder (22) according to claim 16.

18. The first conversion kernel group has one or more conversion kernels that are odd-symmetric on the left side and even-symmetric on the right side, or vice versa, or the second conversion kernel group has one or more conversion kernels that are even-symmetric on both sides or odd-symmetric on both sides. The encoder (22) according to claim 15.

19. The first conversion kernel group includes an MDCT-IV conversion kernel or an MDST-IV conversion kernel, or the second conversion kernel group includes an MDCT-II conversion kernel or an MDST-II conversion kernel. The encoder (22) according to claim 15.

20. The controller (28) is configured such that an MDCT-IV conversion kernel or an MDST-II conversion kernel follows an MDCT-IV conversion kernel, or an MDST-IV conversion kernel or an MDCT-II conversion kernel follows an MDST-IV conversion kernel, or an MDCT-IV conversion kernel or an MDST-II conversion kernel follows an MDCT-II conversion kernel, or an MDST-IV conversion kernel or an MDCT-II conversion kernel follows the MDST-II conversion kernel. The encoder (22) according to claim 15.

21. The encoder (22) according to claim 15, wherein the controller (28) is configured to analyze the block (30) of the overlapping time values having a first channel and a second channel to determine the conversion kernel for the frame of the first channel and the corresponding frame of the second channel.

22. The encoder (22) according to claim 15, wherein the adaptive time-frequency converter (26) is configured to process a first channel and a second channel of a multi-channel signal, and the encoder (22) further includes a multi-channel processor (40) configured to process the blocks of the consecutive spectral values of the first channel and the second channel using a combined multi-channel processing technique to obtain a block of processed spectral values, and an encoding processor (46) configured to process the block of processed spectral values to obtain an encoded channel.

23. The first block of processed spectral values represents a first encoded representation of the combined multi-channel processing technique, the second block of processed spectral values represents a second encoded representation of the combined multi-channel processing technique, the encoding processor (46) is configured to process the first processed block using quantization and entropy encoding to form a first encoded representation, the encoding processor (46) is configured to process the second processed block using quantization and entropy encoding to form a second encoded representation, and the encoding processor (46) is configured to form a bitstream of the encoded audio signal (4) using the first encoded representation and the second encoded representation. The encoder (22) according to claim 22.

24. The first group of conversion kernels includes an MDCT-IV conversion kernel or an MDST-IV conversion kernel, or the second group of conversion kernels includes an MDCT-II conversion kernel or an MDST-II conversion kernel. The controller is configured such that an MDST-II conversion kernel follows an MDCT-IV conversion kernel, or an MDCT-II conversion kernel follows an MDST-IV conversion kernel, or an MDCT-IV conversion kernel follows an MDCT-II conversion kernel, or an MDST-IV conversion kernel follows the MDST-II conversion kernel, the encoder (22) according to claim 15.

25. The MDCT-IV conversion kernel is odd-symmetric on the left side and even-symmetric on the right side, and the composite signal is inverted to the left side during the signal convolution of this conversion, or The MDST-IV conversion kernel is even-symmetric on the left side and odd-symmetric on the right side, and the composite signal is inverted to the right side during the signal convolution of this conversion, or The MDCT-II conversion kernel is even-symmetric on the left side and even-symmetric on the right side, and the composite signal is not inverted on either side during the signal convolution of this conversion, or The MDST-II conversion kernel is odd-symmetric on the left side and odd-symmetric on the right side, and the composite signal is inverted on both sides during the signal convolution of this conversion, The encoder (22) according to claim 24.

26. The multi-channel processor (40) is configured to perform combined stereo processing or combined processing of two or more channels as the combined multi-channel processing technique, and the multi-channel signal has two channels or two or more channels, the encoder (22) according to claim 22.

27. The adaptive time-frequency transformer (26) is configured to use the conversion kernel of the second conversion kernel group for an audio signal (24) representing a harmonic signal whose pitch is at least approximately equal to an integer multiple of the frequency resolution of the conversion, or The adaptive time-frequency transformer (26) is configured to use an MDST-IV based conversion kernel for one of the two channels represented by the audio signal (24), and an MDCT-IV based conversion kernel for the second of the two channels, The encoder (22) according to claim 22.

28. A method (1500) for decoding an encoded audio signal (4), A step of performing a spectrum-time conversion on a block of consecutive spectrum values to a block of consecutive time values (10), Obtaining the decoded audio value (14) by superimposing and adding blocks (10) of consecutive time values Receiving control information (12) and, in response to the control information (12) and in the step of performing the spectrum-time conversion A method including a step of switching between a conversion kernel of a first conversion kernel group, where the conversion kernel of the first conversion kernel group includes one or more conversion kernels with different symmetries on both sides, and a conversion kernel of a second conversion kernel group, where the conversion kernel of the second conversion kernel group includes one or more conversion kernels with equal symmetries on both sides

29. A method (1600) for encoding an audio signal (24), Performing a time-spectrum conversion on a block (30) of overlapping time values to a block of consecutive spectrum values Controlling the time-spectrum conversion step to switch between a conversion kernel of a first conversion kernel group and a conversion kernel of a second conversion kernel group Receiving control information (12) and, in response to the control information (12) and in the step of performing the time-spectrum conversion A method including a step of switching between a conversion kernel of a first conversion kernel group, where the conversion kernel of the first conversion kernel group includes one or more conversion kernels with different symmetries on both sides, and a conversion kernel of a second conversion kernel group, where the conversion kernel of the second conversion kernel group includes one or more conversion kernels with equal symmetries on both sides

30. A computer program for executing the method according to any one of claims 28 or 29 when operating on a computer or a processor

Citation Information

Patent Citations

  • High quality audio encoder / decoder

    JP1993506345A

  • Audio encoder, audio decoder, and multichannel audio signal processing method using complex number prediction.

    JP2013528822A