Decoder for decoding encoded audio signal and encoder for encoding audio signal
The adaptive modification of MDCT transform kernels using MDCT-II, MDST-II, MDCT-IV, and MDST-IV transforms addresses suboptimal energy compaction and phase shift issues in MDCT-based audio codecs, improving coding efficiency and quality.
Patent Information
- Application Number
- JP2025112733
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2015-06-17
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-28
AI Technical Summary
Conventional MDCT-based audio codecs suffer from suboptimal energy compaction, low coding gain, and inadequate handling of stereo signals with phase shifts, leading to insufficient coding quality and increased complexity.
Adaptive modification of the MDCT transform kernel using signal-adaptive permutation, incorporating MDCT-II, MDST-II, MDCT-IV, and MDST-IV transforms, with symmetrical properties to handle harmonic signals and phase shifts, and encoding/decoding mechanisms to switch between these transforms based on instantaneous input characteristics.
Improves coding efficiency by addressing suboptimal energy compaction and phase shift issues, enhancing coding quality and reducing complexity in MDCT-based audio codecs.
Smart Images

Figure 2025163020000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a decoder and an audio decoder for decoding an encoded audio signal. The present invention relates to an encoder for encoding an audio signal. The present invention provides a method and apparatus for signal adaptive transform kernel switching in a The present invention relates to audio coding, in particular to coding techniques such as modified discrete cosine transform (M This paper deals with perceptual audio coding using lapped transforms such as the DCT [1]. [Background technology]
[0002] MP3, Opus, (Celt), HE-AAC family, new MPEG-H 3D audio Modern audio and 3GPP Enhanced Voice Services (EVS) codecs included All popular perceptual audio codecs use the MDCT for spectral domain quantization and coding. Generates waveforms with more than one channel using length - M spectrum spe The composite version of this lapped transform using c[] is given by: It is given by equation (1). JPEG2025163020000002.jpg11152After windowing, time output x i,n The overlap and add (OLA) process Previous time output x i-1,n C is greater than 0 or less than or equal to 1. It may be a constant parameter, for example 2 / N.
[0003] The MDCT in equation (1) above is a high-quality audio codec for any channel at various bit rates. It is suitable for loading, but the coding quality may be insufficient. for example, Sampling through the MDCT so that each harmonic is represented by multiple MDCT bins It is a harmonic signal with a specific fundamental frequency that is matched in the spectral domain. This leads to suboptimal energy compaction, i.e. low coding gain. - Channel coding is not available in conventional M / S stereo-based joint channel coding. Generates a stereo signal with approximately 90 degrees phase shift between the MDCT bins of each channel. More advanced stereo coding, including coding of inter-channel phase difference (IPD), is available, e.g., in HE- I am using AAC parametric stereo or MPEG surround, but Such tools operate in a different filter bank domain and have increased complexity.
[0004] Several journal articles and papers have described operations such as MDCT and MDST. These operations include Lapped Orthogonal Transform (LOT), Extended Lapped Transform (ELT), Modulation [4] is the only method that can simultaneously perform several different overlapped transformations. However, it does not overcome the aforementioned drawbacks of MDCT.
[0005] Therefore, an improved approach is needed. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] HS Malvar, Signal Processing with Lapped Transforms, Norwood: Artech House, 1992. [Non-patent document 2] JP Princen and AB Bradley, "Analysis / Synthesis Filter Bank Design Based on Time Domain Aliasing Cancellation," IEEE Trans. Acoustics, Speech, and Signal Proc., 1986. [Non-patent document 3] JP Princen, AW Johnson, and AB Bradley, "Subband / transform coding using filter bank design based on time domain aliasing ancellation," in IEEE ICASSP, vol. 12, 1987. [Non-patent document 4] HSMalvar, "Lapped Transforms for Efficient Transform / Subband Coding," IEEE Trans.Acoustics,Speech,and Signal Proc., 1990. [Non-Patent Document 5] http: / / en.wikipedia.org / wiki / Modified_discrete_cosine_transform Summary of the Invention [Problem to be solved by the invention]
[0007] The object of the present invention is to provide an improved concept for processing audio signals. This object is solved by the subject matter of the independent claims. [Means for solving the problem]
[0008] The present invention relates to a method for signal-adaptive modification or permutation of the transform kernel, which is suitable for the aforementioned types of MDCT coding. According to an embodiment, the present invention is based on the finding that it is possible to overcome the above-mentioned problems. by generalizing the MDCT coding principle to include three other similar transforms: This addresses the two problems with conventional transform coding. Therefore, this proposed generalization is defined as the following equation (2). JPEG2025163020000003.jpg14151
[0009] The 1 / 2 constant is replaced by the k0 constant, and the cos(...) function is replaced by the cs(...) function. Note that k0 and cs(...) are both signal and context appropriate. It is selected adaptively.
[0010] According to an embodiment, the proposed modification of the MDCT coding paradigm is e.g. It can adapt to instantaneous input characteristics on a frame-by-frame basis so that problems or cases can be handled .
[0011] The embodiment shows a decoder for decoding an encoded audio signal. To convert a successive block of vector values into a successive block of time values, e.g. It includes an adaptive spectrum-to-time converter, which is performed via frequency-to-time conversion. The decoder overlaps successive blocks of time values to obtain the decoded audio value. The adaptive spectrum-to-spectrum conversion further comprises an overlap-add processor for performing an overlap-add process. The transform kernel includes one or more transform kernels with different symmetries on either side of the kernel. A first group of transform kernels and one or more transform kernels with the same symmetry on both sides of the transform kernel. receiving control information from a second group of transformation kernels including a transform kernel; The first group of transform kernels is configured to switch between them as needed. Odd symmetry on the left side of the transformation kernel, such as the IV transformation or inverse MDST-IV transformation kernel one or more transformation kernels with even symmetry on the right side of the transformation kernel, or vice versa. The second group of transformation kernels can include, for example, the inverse M A transform kernel, such as a DCT-II transform kernel or an inverse MDST-II transform kernel. A transformation kernel with even symmetry on both sides, or a transformation kernel with odd symmetry on both sides For transformation kernels type II and IV, will be explained in more detail.
[0012] Therefore, when compared to encoding a signal with the classical MDCT, In order to do this, the frequency of the transform can be taken as the bandwidth of one transform bin in the spectral domain. For harmonic signals with pitches at least approximately equal to an integer multiple of the frequency resolution, Use the second group of transform kernels, e.g., MDCT-II or MDST-II. In other words, it is advantageous to use either MDCT-II or MDST-II. Compared with MDCT-IV, the use of MDCT-IV allows for the resolution of harmonics that are close to an integer multiple of the frequency resolution of the transform. It is advantageous to encode wave signals.
[0013] A further embodiment is where the decoder decodes a multi-channel signal, e.g. a stereo signal. For example, for a stereo signal, the Mid / Side (M / S) stereo processing is superior to classic left / right (L / R) stereo processing. However, if both signals have a phase shift of 90 degrees or 270 degrees, , this approach does not work, or at least is inferior. According to the embodiment, MDST - Encode one of the two channels using IV-based coding and the second channel It is advantageous to use conventional MDCT-IV coding to code Composed by a coding scheme that compensates for a 90-degree or 270-degree phase shift in the audio channels. This results in a 90 degree phase shift between the two embedded channels.
[0014] A further embodiment presents an encoder for encoding an audio signal. The reader uses an adaptive filter to convert overlapping blocks of time values into successive blocks of spectral values. The encoder includes a time-to-spectral converter of the first group of transform kernels. time to switch between the kernel and the transformation kernel of the second group of transformation kernels. The adaptive spectrum converter further includes a controller for controlling the spectrum converter. The rule-to-rule transformer (6) is a transform kernel with different symmetries on both sides of the kernel. The first group of transformation kernels includes one with the same symmetry on both sides of the transformation kernel. Control information (12) is exchanged between the second group of transformation kernels including the above transformation kernels. The encoder receives and switches depending on the control information. Thus, the encoder can be configured to apply a transform kernel such that The transform kernels can be applied in the manner already described for the decoder, and in accordance with the embodiment For example, the encoder applies the MDCT or MDST operation, and the decoder performs the associated inverse operation, i.e. applying the IMDCT or IMDST transform. For different transform kernels, This is explained in detail below.
[0015] According to a further embodiment, the encoder determines for the current frame: encoded with control information indicating the symmetry of the transformation kernel used to generate An output interface is provided for generating an audio signal. is a decoder that can decode an audio signal coded with the correct transform kernel. In other words, the decoder can generate control information for The inverse transform kernel of the transform kernel used by the This information is then stored in the encoded audio signal. The control data section of the frame of the audio signal is used to store and encode the control information. It may be transmitted from the reader to the decoder.
[0016] Embodiments of the present invention will continue to be discussed with reference to the accompanying drawings. [Brief explanation of the drawings]
[0017] [Figure 1] 1 shows a schematic block diagram of a decoder for decoding an encoded audio signal; [Figure 2] FIG. 2 is a schematic block diagram illustrating signal flow in a decoder according to one embodiment. [Figure 3] 1 shows a schematic block diagram of an encoder for encoding an audio signal according to one embodiment; [Figure 4A] 2 shows a schematic representation of a block of a series of spectral values obtained by an exemplary MDCT encoder. [Figure 4B] 1 shows a schematic diagram of a time domain signal input to an exemplary MDCT encoder. [Figure 5A] 1 shows a schematic block diagram of an exemplary MDCT encoder according to one embodiment; [Figure 5B]1 shows a schematic block diagram of an exemplary MDCT decoder according to one embodiment; [Figure 6] Schematically illustrates the implicit deconvolution properties and symmetries of the four described lapped transforms. [Figure 7] 10A-B illustrate two embodiments of a use case where signal adaptive transform kernel switching is applied to the transform kernel from one frame to the next while still allowing perfect reconstruction. [Figure 8] 1 shows a schematic block diagram of a decoder for decoding a multi-channel audio signal according to one embodiment; [Figure 9] FIG. 4 is a schematic block diagram of the encoder of FIG. 3 extended to multi-channel processing according to one embodiment; [Figure 10] 1 shows a schematic audio encoder for encoding a multi-channel audio signal having two or more channel signals according to one embodiment; [Figure 11A] FIG. 2 shows a schematic block diagram of an encoder calculator according to one embodiment. [Figure 11B] FIG. 10 shows a schematic block diagram of another encoder calculator according to one embodiment. [Figure 11C] 1 shows a schematic diagram of an exemplary combining rule for a first and second channel in a combiner according to one embodiment; [Figure 12A] 1 shows a schematic block diagram of a decoder calculator according to one embodiment; [Figure 12B] FIG. 1 shows a schematic block diagram of a matrix calculator according to one embodiment. [Figure 12C] 11D illustrates a schematic diagram of an exemplary anti-combination rule for the combination rule of FIG. 11C according to one embodiment. [Figure 13A] 1 shows a schematic block diagram of an implementation of an audio encoder according to one embodiment; [Figure 13B] 13B shows a schematic block diagram of an audio decoder corresponding to the audio encoder shown in FIG. 13A according to one embodiment. [Figure 14A] 10 shows a schematic block diagram of a further implementation of an audio encoder according to an embodiment; [Figure 14B] 14B shows a schematic block diagram of an audio decoder corresponding to the audio encoder shown in FIG. 14A according to one embodiment. [Figure 15] 1 is a schematic block diagram of a method for decoding an encoded audio signal; [Figure 16] 1 shows a schematic block diagram of a method for encoding an audio signal; DETAILED DESCRIPTION OF THE INVENTION
[0018] The following describes the embodiments of the present invention in more detail. Elements shown in each figure are associated with the same reference numerals.
[0019] FIG. 1 shows a schematic block diagram of a decoder 2 for decoding an encoded audio signal 4. The decoder includes an adaptive spectro-temporal converter 6 and an overlap adder 8. The type spectrum-to-time converter converts successive blocks of spectral values 4' into, for example, frequency-time Transformed into a successive block 10 of time values via a time-to-spectrum transform. The rule-to-rule transformer (6) is a transform kernel with different symmetries on both sides of the kernel. The first group of transformation kernels includes one with the same symmetry on both sides of the transformation kernel. Control information (12) is exchanged between the second group of transformation kernels including the above transformation kernels. The overlap-add processor 8 receives the control information and switches in response to the control information. , overlapping and adding successive blocks of time values 10 to obtain a decoded audio value 1 4. The decoded audio values 14 may be a decoded audio signal.
[0020] According to an embodiment, the control information 12 includes a current bit indicating the current symmetry of the current frame. and the adaptive spectrum-to-time converter 6 determines whether the current bit is When the current bit is shifted from the first group to the second group, it shows the same symmetry as was used In other words, for example, the control information 12 is configured not to switch to the previous frame. indicates that the first group of transformation kernels is to be used for the current frame and If the previous frame contains the same symmetry, e.g., the current bit of the current frame and the previous frame The first group of transformation kernels is applied when the frames have the same state. This is because the adaptive spectrum-to-time converter converts the first transformation kernel group into the second transformation kernel group. This means that the channel group will not be switched to another channel group. To stay in the current frame or not switch from the second group to the first group, The current bit indicating the current symmetry of the frame is a different pair than that used in the previous frame. In other words, if the current symmetry and the previous symmetry are equal, then the previous frame If the current frame was coded using a transform kernel from the second group, The image is decoded using the inverse transform kernel of the group.
[0021] Additionally, the current bit indicating the current symmetry of the current frame is used in the previous frame. If the first group exhibits a different symmetry from that of the first group, the adaptive spectrum-to-time converter 6 More specifically, the current frame is switched to the second group. The current bit indicating the current symmetry of the frame is a different symmetry than that used in the previous frame. When the first group is changed to the second group, the adaptive spectrum-to-time converter 6 changes the first group to the second group. Furthermore, a current bit indicating the current symmetry of the current frame is exhibits the same symmetry as used in the previous frame. The switch 6 can switch the second group to the first group. More specifically, , the current frame and the previous frame contain the same symmetry, and the previous frame is the first If the current frame is coded using a group of transform kernels, the The control information 12 may be decoded using a transform kernel of the first group of kernels. , may be derived from the encoded audio signal 4, as will become apparent below, or or may be received via a separate transmission channel or carrier signal. The current bit indicating the current symmetry of the frame may be the symmetry of the right side of the transformation kernel. stomach.
[0022] In a 1986 paper by Princen and Bradley [2], the trigonometric functions of the cosine and sine functions Two wrap transforms using numbers are described, called "DCT-based" in the article. The first one is (2)cs()=cos() and k o = 0. , the other is called "DST-based" and uses cs()=sin() and k o If =1 (2) is given and defined as follows. Due to the respective similarities with DST-II, in this paper we consider the general formulation of (2) These particular cases are the "MDCT Type II" transformation and the "MDST Type II" transformation, respectively. Princen and Bradley continued their investigation in a 1987 paper [3]. So, cs() = cos() and ko We propose a common case of =0.5, which is introduced in (1) and For clarity, we will refer to the DCT-IV as Due to the relationship between D and D, this transform is referred to herein as "MDCT Type IV." Based on ST-IV, cs() = cos() and k o =0.5 and (2) We have already identified the remaining possible combinations, called "MDST Type IV," obtained by The embodiment describes when to switch between these four transformations signal-adaptively. do.
[0023] As pointed out in [1-3], perfect reconstruction properties (without spectral quantization or other distortions) The four are used so that the identical reconstruction of the input signal after analysis and synthesis transformations (no introduction) is maintained. Some thoughts on how the essential switching between different transformation kernels is achieved For this purpose, it is worth defining the symmetric rules for the composition transformations that follow (2). It is useful to examine the expansion properties, which are shown with respect to FIG. MDCT-IV exhibits odd symmetry on its left side and even symmetry on its right side. The resulting signal is inverted on its left side during the deconvolution of the signal of this transformation. MDST-IV exhibits even symmetry on its left side and even symmetry on its right side. The resulting signal is inverted on its right side during the deconvolution of the signal of this transformation. MDCT-II exhibits even symmetry on its left side and odd symmetry on its right side. The resulting signal is not inverted on either side during the signal defolding of this transformation. MDST-II exhibits odd symmetry on its left side and even symmetry on its right side. The resulting signal is inverted on both sides during the deconvolution of this transformed signal.
[0024] Furthermore, two embodiments for deriving the control information 12 in the decoder are described. The control information may, for example, be the value of k0 and cs to indicate one of the four transformations mentioned above. () and (). Thus, the adaptive spectrum-to-time converter may The control information of the previous frame and the control information following the previous frame are extracted from the audio signal. The control data section of the frame can be read from the encoded audio signal. Optionally, the adaptive spectrum-to-time converter 6 may The control information 12 may be read from the control data section of the previous frame. Controls for the previous frame are taken from the decoder settings applied to the previous frame. In other words, the control information may be read from the control data section. It may be derived directly from the current frame or the previous frame's decoded data in the header. It may be derived from the reader settings.
[0025] The following describes the control information exchanged between the encoder and decoder according to a preferred embodiment: This section describes how the side information (i.e., control information) is encoded. signaled and derived in the generated bitstream, and robust How to derive and apply appropriate transformation kernels in a way that is suitable for frame loss (e.g., for frame loss) This article explains:
[0026] According to a preferred embodiment, the present invention provides MPEG-D USAC (extended HE-AAC) Or it can be integrated into the MPEG-H 3D audio codec. The information is available for each frequency domain (FD) channel and frame, so-called fd c It can be sent within the channel stream element. More specifically, scale_factor_data( ) bitstream element immediately preceded or followed by a 1-bit currAliasingSymmetry flag. It is written (by the encoder) and read (by the decoder). If the frame is an independent frame, i.e., indepFlag == 1, then another bit prevAliasingSymm The etry is written and read, which ensures both left and right symmetry and The resulting transformation kernel is used within the frame and channel to Even if a previous frame is lost during stream transmission, it can still be identified (and properly decoded) within the decoder. If the frame is not an independent frame, prevAliasingSymmetry is not written and is read. It is not read but set equal to the value that currAliasingSymmetry had on the previous frame. According to a further embodiment, different bits or flags are used to indicate the control information ( That is, sub-information).
[0027] Then, the values of cs() and k0 are calculated based on the currAliasingSymmetry and prevAliasingSymmetry. Derived from the ngSymmetry flag (currAliasingSymmetry is symmetry) i and prevAliasin gSymmetry is symm i-1 In other words, symm i is at index i The control information for the current frame in symm i-1 is the previous field at index i-1. Table 1 shows the control information for the frame. Decoder-side decision matrix that specifies the values of and cs(...) based on side information Therefore, the adaptive spectrum-to-time converter selects the conversion card according to Table 1 below. A panel can be applied. JPEG2025163020000004.jpg52158
[0028] Finally, once cs() and k0 are determined at the decoder, for a given frame and The inverse transform for the channel can be performed with the appropriate kernel using equation (2). Before and after the synthesis transform, the decoder operates normally as in the prior art with respect to windowing. It is possible to do this.
[0029] FIG. 2 shows a schematic block diagram illustrating the signal flow in a decoder according to one embodiment; where the solid line indicates the signal, the dashed line indicates the side information, and i indicates the frame index. , xi indicate frame time-signal output. The bitstream demultiplexer 16 Receives successive blocks of spectral values 4' and control information 12. According to one embodiment, Consecutive blocks of vector values 4'' and control information 12 are multiplexed into a common signal The bitstream demultiplexer separates blocks and sub-blocks of consecutive spectral values from a common signal. The successive blocks of spectral values are further configured to derive spectral and control information. The current frame 12 and the previous frame 13 may be input to a vector decoder 18. The control information of the system 12' is input to the mapper 20, which applies the mapping shown in Table 1. According to the embodiment, the control information of the previous frame 12' is included in the coded audio signal, i.e. The current decoder response for the previous block of spectral values or the previous frame. The spectrally decoded spectral value 4'' may be derived using a preset. The processed control information 12', including the successive blocks and the parameters cs and k0, is , which is input to the inverse kernel adaptive lap transformer, which is the adaptive spectrum-to-time transformer 6 in Figure 1. The output is then scaled to overcome discontinuities, e.g., at the boundaries of successive blocks of time values. , with successive blocks of time values 10 that can optionally be processed using a synthesis window 7. The audio values may be decoded by performing an overlap-add algorithm. 14 is input to the overlap-add processor 8. The mapper 20 and The adaptive spectral-to-temporal converter 6 can further move to another point in the decoding of the audio signal. Therefore, the locations of these blocks are merely suggestions. , the control information may be calculated using a corresponding encoder, the embodiment of which is e.g. For example, see FIG.
[0030] FIG. 3 is a schematic block diagram of an encoder for encoding an audio signal according to one embodiment. The encoder includes an adaptive time-to-spectral converter 26 and a controller 28. The adaptive time-to-spectral converter 26 comprises, for example, blocks 30' and 30''. The overlapping blocks of time values 30 containing Furthermore, the adaptive spectral-to-temporal converter (6) has different symmetries on both sides of the kernel. A first group of transformation kernels containing one or more transformation kernels and two groups of transformation kernels on either side of the transformation kernel. and a second group of transformation kernels that includes one or more transformation kernels with the same symmetry. Then, the controller 2 receives the control information (12) and switches the control information accordingly. 8 controls the time-spectrum converter to generate the first group of transform kernels; and a transformation kernel of the second group of transformation kernels. Optionally, the encoder 22 generates an encoded audio signal for the current frame. an output interface 32 for generating an encoded audio signal; and control information 1 indicating the symmetry of the transformation kernel used to generate the current frame. 2. The current frame is the current block of consecutive blocks of spectral values. The output interface may include the current It can contain symmetry information between the current frame and the previous frame, which is an independent frame, and or the control data section of the current frame. If the frame is a dependent frame, only the symmetric information of the current frame is present, and the symmetric information of the previous frame does not exist. The output interface must include the current It can contain symmetry information for the frame and the previous frame, and the current frame is independent frame, or the control data section of the current frame is symmetric to the current frame If the current frame is a dependent frame, it contains only the symmetric information of the previous frame. An independent frame does not include, for example, an independent frame header, which allows You can reliably read the current frame without knowledge of the previous frame. A file is, for example, an audio file with variable bit rate switching. Thus, a dependent frame can be read with only knowledge of one or more previous frames. An independent frame can include, for example, an independent frame header, which allows You can reliably read the current frame without knowledge of the previous frame. The program is, for example, an audio file with variable bit rate switching. Thus, a dependent frame can only be read with knowledge of one or more previous frames. Cut.
[0031] The controller may, for example, adjust the fundamental frequency to at least a multiple of the frequency resolution of the transform. The control device may be configured to analyze the audio signal 24. uses the control information 12 to control the adaptive time-to-spectral converter 26 and optionally the output interface The control information 12 can be derived to be supplied to the interface 32. Indicates the appropriate transformation kernel for the first group of kernels or the second group of transformation kernels. The first group of transformation kernels has odd symmetry on the left side of the kernel, and one or more transformation kernels with even symmetry on the right side of the kernel, or vice versa Alternatively, the second group of transformation kernels may have even symmetry on both sides of the kernel. or contains one or more transformation kernels with odd symmetry on both sides of the kernel In other words, the first group of transform kernels is the MDCT-IV transform kernel. A second group of transformation kernels may include a MDST-IV transformation kernel. The group can contain an MDCT-II transform kernel or an MDST-II transform kernel. To decode the encoded audio signal, the decoder performs each inverse transform. The decoder can then apply the transform kernel to the The first group of kernels is the inverse MDCT-IV transform kernel or the inverse MDST-IV transform kernel. or a second group of transform kernels may include the inverse MDCT-II transform The MDST-II transformation kernel may include a MDST-II transformation kernel.
[0032] In other words, the control information 12 is a current frame information indicating the current symmetry for the current frame. Furthermore, the adaptive spectrum-to-time converter 6 determines whether the current bit is When the same symmetry as used in the previous frame is shown, the first group is replaced by the second group. It may be configured not to switch to the transformation kernel of the group, and the current bit is When the system exhibits a different symmetry than that used in the previous system, the adaptive spectral-to-temporal converter It is configured to switch from one group of transform kernels to a second group of transform kernels.
[0033] Furthermore, the adaptive spectro-temporal converter 6 determines whether the current bit is used in the previous frame. When the second group exhibits a different symmetry from the first group, the transformation kernel can be configured not to switch to the current bit, and the current bit is used in the previous frame. When the second group exhibits the same symmetry as the first group, the adaptive spectrum-time converter can convert the It is configured to switch to a transformation kernel for the loop.
[0034] Time portion and block on either the encoder side, analysis side, decoder side, or synthesis side To illustrate the relationship between the clock and the clock, reference is made to Figures 4A and 4B.
[0035] FIG. 4B shows a schematic diagram of the 0th time portion to the 3rd time portion, and these subsequent time portions Each time portion of the part has a certain overlap range 170. Based on these time portions, the overlap A series of consecutive blocks representing the time portion indicates the analysis side of the aliasing-introducing transformation operation. It is generated by a process described in more detail with respect to FIG. 5A.
[0036] In particular, the time domain signal shown in FIG. 4B when applied to the analysis side is Therefore, to obtain the 0th time portion, For example, an analysis window is applied to 2048 samples, specifically samples 1 to 2048. Therefore, N is equal to 1024, and the windowing has a length of 2N samples, and this example is 48. Then the windower starts with sample 2049 as the first sample of the block. Instead, use sample 10 as the first sample in the block to get the first time portion. Further analysis operations are applied to 25. Therefore, for a 50% overlap, A first overlap range 170 is obtained that is 0.24 samples long. This procedure It is applied to three time segments additively, but always overlapped to obtain a certain overlap range 170. Meet.
[0037] The overlap does not necessarily have to be 50% overlap, but It is emphasized that the tops may be higher or lower and may be multi-overlapping. That is, the samples of the time domain audio signal should be split into two windows and the resulting The overlap of two or more windows is obtained so that they do not contribute to a block of spectral values as However, samples contribute to more than one window / block of spectral values. If so, the windowing section 201 of FIG. 5A may have portions with values of 0 and / or 1. It is further understood that there are other windowing shapes that may be applied. For portions that have values, such portions are typically zero portions of the preceding or succeeding window. A particular audio sample located in a certain portion of the window overlaps with the A pull only contributes to a single block of spectral values.
[0038] The windowed time portion obtained by FIG. 4B is used to perform the convolution operation. The folder 202 is then transferred to the folder 202 for folding. At the output, only blocks of sampled values with N samples per block are Convolution can be performed as it exists. And convolution by folder 202 Following the operation, a time-to-frequency transformer is applied, which is The N samples are converted into N spectral values at the output of the time-to-frequency converter 203. This is a DCT-IV converter.
[0039] Thus, the series of blocks of spectral values obtained at the output of block 203 is shown in FIG. 1A and 1B, and specifically, a first modification value 102 is associated with the 1A and 1B. Block 191 is shown. Of course, the sequence precedes the second block. or by adding block 193 or 194 preceding the first block as shown. The first and second blocks 191 and 192 may be, for example, the windowed blocks of FIG. The first time portion is converted to obtain a first block, and the second block is converted to obtain a second block. The lock is converted by the time-to-frequency converter 203 of FIG. 5A into the windowed second time signal of FIG. 4B. Therefore, in a series of blocks of spectral values, In this case, both blocks of temporally adjacent spectral values are in the first time portion and in the second time portion. Represents the overlapping range covering the time portion.
[0040] 5B shows the synthesis or decoding side of the results of the encoder or analysis side processing of FIG. 5A. 5A is explained to show the processing on the side of the frequency converter 203. The block of spectral values of the series is input to modifier 211. As outlined, Each block of spectral values has N spectral values for the example shown in FIGS. 4A-5B. (Note that this is different from equations (1) and (2) where M is used.) The locks have associated change values such as 102 and 104 shown in Figures 1A and 1B. In a typical IMDCT operation or redundancy reduction synthesis transform, the frequency-to-time transformer 212 , a folder 213 for deconvolution, a windowing unit 214 for applying a synthesis window, and ,The block in which the overlap / add operation is performed to obtain the time-domain signal in the overlapping,range. In this example, there are 2N values per block, so each After the lap and operation, the change values 102 and 104 are used for time or frequency. If the number of samples is not variable, N new alias-free time-domain samples are obtained. However, if these values vary with time and frequency, the output of block 215 The force signal is not aliasing-free, and this issue is discussed in the context of Figures 1B and 1A. and as discussed in the context of other figures herein, according to the first and second aspects of the invention. This will be dealt with.
[0041] Subsequently, a further description of the procedures performed by the blocks of Figures 5A and 5B is provided. can be done.
[0042] This figure is M Although exemplified by reference to the DCT, other aliasing-introducing transforms have similar It can be processed in a similar way. As a lapped transform, the MDCT has (but not the same number of) It is somewhat unusual compared to other Fourier-related transforms in that it has half the output as the input. In particular, it Linear function F:R 2N → R N (R represents the set of real numbers). 2N real numbers x0, ...,x2N-1 is transformed into N real numbers X0,...,XN-1 according to the formula . JPEG2025163020000005.jpg17158
[0043] (The normalization factor before this transformation, here unity, is an arbitrary convention and varies from process to process.) (Only the product of the MDCT and IMDCT normalizations below is constrained).
[0044] The inverse MDCT is also known as the IMDCT. At first glance, it has a different number of inputs and outputs. Therefore, it may seem that MDCT is not invertible. However, complete inversion is sometimes Add the overlapped IMDCTs of the adjacent overlapping blocks. This is achieved by canceling the errors and recovering the original data. This is known as time-domain aliasing cancellation (TDAC).
[0045] IMDCT converts N real numbers X0,...,XN-1 into 2N real numbers y0,...,y2 Follow the formula below to convert to N-1. JPEG2025163020000006.jpg23157
[0046] (As with the orthogonal transform DCT-IV, the inverse function has the same format as the forward transform.)
[0047] Windowed MDCT with a normalized window (see below) In this case, the normalization factor before the IMDCT should be doubled (i.e., to 2 / N).
[0048] In typical signal compression applications, the transform characteristics are determined by the MDCT and IMDCT formulas We use a window function wn (n=0,...,2N-1) that is multiplied with xn and yn in This is further improved by using (i.e., before the MDCT and after the IM) The data is windowed after the DCT.) In principle, x and y can have different window functions. The window function can also be changed from one block to the next (especially for different sizes). For simplicity, we will assume that the data blocks are of equal size. We consider the general case of identical window functions for all frames.
[0049] JPEG2025163020000007.jpg88148
[0050] The window applied to MDCT must satisfy the Princen-Bradley condition. This is different from the window used for other types of signal analysis. One reason for this difference is that MDCT (analysis The advantage of this approach is that the MDCT window is applied twice, for both the MDCT and the IMDCT (synthesis).
[0051] By examining the definition, we can see that for N, the MDCT also has N / 2 inputs. This is essentially the same as the DCT-IV, where two N blocks of data are transformed at once. By carefully examining this equivalence, important characteristics such as TDAC can be It can be easily derived.
[0052] To define the exact relationship with the DCT-IV, the DCT-IV is It must be recognized that this corresponds to alternating the left and right boundaries (i.e., symmetry conditions). (approximately n=-1 / 2), odd on the right boundary (around n=N=-1 / 2), and Instead of periodic boundaries, it may be followed by: JPEG2025163020000008.jpg17137 and JPEG2025163020000009.jpg16137
[0053] So if the input is an array x of length N, we can write this array as (x,-xR,-x, xR,...), where xR denotes x in reverse order.
[0054] Consider an MDCT with 2N inputs and N outputs. Here, we divide the inputs into blocks of size N Divide into four blocks (a, b, c, d) of +N / 2. Shifting right by / 2 moves (b,c,d) beyond the end of the N DCT-IV inputs. stretches and then we need to "convolve" them according to the boundary conditions above.
[0055] Therefore, the MDCT with 2N inputs (a, b, c, d) is exactly the same as the DCT-IV with N inputs. is equivalent to (-cR-d, a-bR).
[0056] This is illustrated for window function 202 in Figure 5A, where a is portion 204b and b is part 205a, c is part 205b, and d is part 206a.
[0057] (Thus, the algorithm for computing the DCT-IV can be trivially applied to the MDCT. Similarly, the IMDCT formula above is exactly 1:1 of the DCT-IV (its own inverse). / 2, and the output is expanded (via boundary conditions) to length 2N and moved back N / 2 to the left The inverse DCT-IV simply returns the input (-cR-d, a-bR) from above. When expanded and shifted by the boundary conditions, IMDCT(MDCT(a,b,c,d))=(a-bR,b-aR,c+dR,d+ cR) / 2 This becomes:
[0058] Therefore, half of the IMDCT output is redundant, as b-aR=-(a-bR)R. and similarly for the last two terms. Let the inputs be A = (a,b) and B = (c,d). A simpler way to achieve this result is to group the results into larger blocks A, B of size N. IMDCT(MDCT(A,B))=(A-AR,B+BR) / 2 It can be written as:
[0059] You will be able to understand how TDAC works. 2N blocks that are adjacent in time and overlap by 50% Assume that we are calculating the MDCT of block (B, C). The IMDCT is calculated as above using the R,C+CR) / 2. This is added to the previous IMDCT result in the overlapping half. , the inverse terms cancel out and we simply take B to recover the original data.
[0060] The origin of the term "time domain aliasing cancellation" is now clear. The use of input data that extends beyond the boundaries of the DCT-IV allows for frequencies above the Nyquist frequency. The wavenumbers are aliased to lower frequencies in the same way (with respect to extended symmetry). To distinguish the contribution of (a, b, c, d) to MDCT from that of bR or equivalently, IMDCT(MDCT(a,b,c,d))=(ab R, b-aR, c+dR, d+cR) / 2 result. , have exactly the correct symbols to cancel when the combination is added.
[0061] For odd N (rarely used in practice), N / 2 is not an integer, so MDC T is not just a shift permutation of the DCT-IV. In this case, an additional shift of half a sample means that MDCT / IMDCT is equivalent to DCT-III / II, and the analysis The same as above.
[0062] The MDCT of 2N inputs (a, b, c, d) is the N-input (-cR-d, a-bR) We have seen above that the DCT-IV is equivalent to the DCT-IV of Since it is designed for odd numbers, the values near the right boundary are close to 0. If so, then in the input sequence (a, b, c, d), the rightmost elements of a and b are consecutive. Therefore, the difference is small. Let's look at the middle of the interval. If we rewrite it as (-d,a)-(b,c)R, the second (b,c)R is in the middle. However, in the first term (-d, a), there is a discontinuity where the right end of -d coincides with the left end of a. This is a window function that reduces the boundary components of the input sequence (a, b, c, d) towards zero. That's why we use numbers.
[0063] As mentioned above, the TDAC property is proven in conventional MDCT, and temporally adjacent Adding the IMDCT of the block to the overlapping halves recovers the original data. The derivation of this inverse characteristic for the windowed MDCT (windowed MDCT) is The output is only slightly more complicated.
[0064] JPEG2025163020000010.jpg34154
[0065] JPEG2025163020000011.jpg29156.
[0066] So instead of doing MDCT(A,B), all multiplications are done element-wise. MDCT S (WA,W R B) is currently present. This is input to the IMDCT, and the window function When multiplied again (element-wise) by , the last half of N becomes: W R ·(W R B+(W R B) R )=W R ·(W R B+WB R )=W R 2 B+WW R B R
[0067] (The normalization of the IMDCT is two times different in the windowed case, so the multiplications are halved. (Not possible).
[0068] Similarly, the MDCT and IMDCT of windowed (B,C) are given by It will look like this. W·(WB-W R B R )=W 2 B-WW R B R
[0069] Adding these two halves together restores the original data. - When half of the window to be burlapped satisfies the Princen-Bradley condition, the window switching context De-aliasing can be done in this case in exactly the same way as above. Multiple lapped transforms can be used to create three or more branches using all relevant gain values. will be necessary.
[0070] So far, we have not discussed the symmetry or boundary conditions of MDCT, more specifically MDCT-IV. Other variants, MDCT-II, MDST-II, and MDST-IV, have been described. The explanation is also valid for the transformation kernel. However, the different symmetries or It should be noted that boundary conditions must be taken into account.
[0071] Figure 6 shows the implicit deconvolution properties and symmetries (i.e., boundaries) of the four described lapped transforms. The transformations are shown in Fig. 1. The first composite basis function for each of the four transformations is It is derived from (2) via IMDCT-IV34a, IMDCT-II34b, I MDST-IV34c and IMDST-II34d are schematic diagrams of amplitude samples over time. FIG. 6 shows the symmetry axis 35 (i.e., the folding) between the transformation kernels as described above. The even and odd symmetries of the transformation kernel at the fold point are clearly shown.
[0072] The Time Domain Aliasing Cancellation (TDAC) property is When even and odd symmetric extensions are summed during the add-and-add process, the aliasing In other words, for TDAC to occur, the odd right A transformation with symmetry must be followed by a transformation with even left-hand symmetry, and The reverse is also true. therefore, (Reverse) MDCT-IV followed by reverse MDCT-IV or reverse MDST-II . (Reverse) MDST-IV followed by reverse MDST-IV or reverse MDCT-II . (Reverse) MDCT-II followed by reverse MDCT-IV or reverse MDST-II . (Reverse) MDST-II followed by reverse MDST-IV or reverse MDCT-II .
[0073] Figure 7(a) and Figure 7(b) show the signal-adaptive conversion card while enabling perfect reconstruction. Uses kernel switching applied to the transform kernel from one frame to the next. In other words, the two cases of the above conversion sequence are shown in A possible sequence is illustrated in Figure 7, where solid lines (such as line 38c) indicate conversion windows. The dashed line 38a indicates the aliasing symmetry on the left side of the transform window, and the dotted line 38b indicates the aliasing symmetry on the right side of the transform window. It exhibits aliasing symmetry. Furthermore, symmetric peaks exhibit even symmetry, and symmetric valleys exhibit odd symmetry. In Figure 7(a), 36a in frame i and 36b in frame i+1 are MD The CT-IV transformation kernel is 36c in frame i+2 and 3 in frame i+3. MST-II is used as a transition to the MDCT-II transform kernel used in 6d. Frame i+4, 36e, uses MDST-II again, as shown in Figure 7(a). MDST-IV is used again for MDCT-II in frame i+5, which has not been 7(a) shows that the dashed line 38a and the dotted line 38b are used to compensate for the subsequent transformation kernel. In other words, the left-side aliasing symmetry of the current frame and the previous frame are clearly shown. Summing up the right-side aliasing symmetry of the frame, the sum of the dotted line and the dotted line is equal to 0, so , perfect time domain aliasing cancellation (TDAC) is obtained. The symmetry (or boundary conditions) can be expressed as the convolution properties described in, for example, Figures 5A and 5B. , and the MDCT produces an output containing N samples from an input containing 2N samples. This is the result.
[0074] FIG. 7(b) is similar to FIG. 7(a), and shows the time series from frame i to frame i+4. It just uses a different set of transform kernels. In frame i36a, MDCT-I V is used, and 36b in frame i+1 is the MDS used in 36c in frame i+2. Use MDST-II as the transition to T-IV. Frame i+3 is frame i+2. The MDST-IV transformation kernel used in 36d is the MDC in 36e of frame i+4. The MDCT-II transform kernel is used as a transition to the T-IV transform kernel.
[0075] The associated decision matrix for the transformation sequence is shown in Table 1.
[0076] The embodiment is based on the adaptive transform proposed in audio codecs such as HE-AAC. How can kernel switching be advantageously employed to solve the two challenges mentioned at the beginning? The following shows how conventional MDCT can minimize or avoid to deal with suboptimally coded harmonic signals. The adaptive transition to is performed by the encoder based on, for example, the fundamental frequency of the input signal. More specifically, if the pitch of the input signal is an integer multiple of the frequency resolution of the transform (i.e., (i.e., the bandwidth of one transform bin in the spectral domain), MDCT-II or MDST-II for affected frames and channels However, the MDCT-IV to MDCT-II transformation kernel Direct transition is not possible, or at least time-domain aliasing cancellation (TDAC) is required. Therefore, in such cases, MDCT-II uses the following as a transition transformation between the two. Conversely, the transition from MDST-II to the traditional MDCT-IV (i.e., switching to traditional MDCT coding) requires the intermediate MDCT-II is advantageous.
[0077] To date, adaptive conversion cards have been proposed to enhance the coding of harmonic audio signals. Channel switching has been described for a single audio signal. It can be easily adapted to multi-channel signals such as Leo signals, where, for example, Two or more channels of a multi-channel signal have a phase shift of approximately ±90 degrees relative to each other. In this case, adaptive transform kernel switching is also advantageous.
[0078] For multi-channel audio processing, MDCT is performed for one audio channel. -IV encoding for the second audio channel and MDST-IV encoding for the third audio channel In particular, it may be appropriate to set both audio channels at approximately ±90 degrees before encoding. This concept is advantageous when including a phase shift of . , which gives the coded signals a 90 degree phase shift compared to each other, so that the two channels of the audio signal The ±90 degree phase shift between channels is compensated after encoding, i.e., the phase shift of MDCT-IV is A 90 degree phase difference between the cosine basis function and the MDST-IV sine function results in a 0 degree or a 180 degree phase shift. So, for example, in M / S stereo coding: In this case, both channels of the audio signal may be encoded with an intermediate signal, with a 0 degree phase shift. For the above transformation to the 18 Inversion to a 0 degree phase shift yields the opposite (minimal information in the intermediate signal), which This achieves maximum channel compression. Compared to conventional MDCT-IV coding, it does not use a lossless coding scheme. While this may reduce bandwidth by up to 50%, it is still possible to achieve complex stereo prediction. It is also possible to use MDCT stereo coding in combination with both approaches. calculates, encodes, and transmits a residual signal from two channels of an audio signal. Complex prediction is the process of calculating prediction parameters for encoding an audio signal and decoding it. The decoder uses the transmitted parameters to decode the audio signal. MDCT-IV and MDST-IV for coding two audio channels have already been developed. As mentioned above, the code used is Regarding the method of MDCT-II, MDST-II, MDCT-IV, or MDST-IV Only information that is relevant to the stereo prediction should be transmitted. Since the image should be quantized using the image resolution, information about the coding scheme used is For example, it may be coded on 4 bits. Theoretically, the first and second channels can be coded on 4 bits. Each may be coded using one of the different coding schemes, resulting in 16 This leads to different possible states of
[0079] Therefore, FIG. 8 shows a schematic diagram of a decoder 2 for decoding a multi-channel audio signal. 2 shows a schematic block diagram of the decoder of FIG. 1. Compared to the decoder of FIG. 1, the decoder Multi-channel signal for receiving blocks of spectral values 4a''', 4b''' representing channels The system further comprises a channel processor 40 for processing the first multi-channel and second multi-channel To obtain the processed blocks of spectral values 4a' and 4b', the received blocks are According to the joint multi-channel processing technique, the adaptive spectro-temporal processor The control information 12a for the first multi-channel and the control information 12b for the second multi-channel are The processed block 4b' for the second multi-channel to be used is used to process the first multi-channel. The multi-channel processor is configured to process the processed blocks 4a' of the channels. The processor 40 may apply, for example, left-right stereo processing, sum-and-difference stereo processing, or Alternatively, the multi-channel processor may generate spectra representing the first and second multi-channels. Complex prediction is applied using complex prediction control information associated with the block of values. A multi-channel processor determines which processing is used to encode an audio signal, for example. This may include fixed presets or information derived from control information indicating whether In addition to separate bits or words in the control information, multi-channel processors The server may express this information, for example, by the absence or presence of multi-channel processing parameters. In other words, the multi-channel processor 40 Applying the inverse operation to the multi-channel processing performed in the encoder to separate the multi-channel signals Further multi-channel processing techniques are described in Figs. 14. Furthermore, references are made to multi-channel processing and are denoted by the letters " The reference letter extended by "a" indicates the first multichannel, and the reference letter extended by the letter "b" indicates the Therefore, it is expanded to show the second multi-channel. It is not limited to stereo or mono processing, but extends the illustrated processing to two channels. By doing so, it can be applied to more than two channels.
[0080] According to an embodiment, the multi-channel processor of the decoder performs joint multi-channel processing. The received block may be processed according to the technique. represents the coded residual signal of the first multi-channel representation and the second multi-channel representation as Further, the multi-channel processor may include a residual signal and a further code. Calculating a first multi-channel signal and a second multi-channel signal using the multiplied signal In other words, the residual signal may be an M / S encoded audio signal. It may be a side signal of the signal or, when used, an additional channel of the audio signal. The residual between the channel of the audio signal based on the signal and the prediction of the channel, e.g., complex sequence Therefore, the multi-channel processor may be, for example, an inverse transform card. Predict the M / S or complex predicted audio signal from the L / R audio signals. The difference signal and the intermediate signal of the M / S coded audio signal or the audio signal (e.g. , MDCT coded) channels. It can be used.
[0081] Figure 9 shows the encoder 22 of Figure 3 extended to multi-channel processing. is expected to be included in the coded audio signal 4, but the control information 12 is For example, the signal may be further transmitted using a separate control information channel. The controller 28 of the data processor 200 transfers a frame of the first channel and a corresponding frame of the second channel. To determine the transformation kernel of the frame, we have a first channel and a second channel. It is possible to analyze overlapping blocks of time values 30a, 30b of the audio signal. Therefore, the controller tries each combination of transformation kernels, e.g., M The residual signal for M / S coding or complex prediction (or the side signal for M / S coding) is We can derive options for the minimizing transform kernel. The minimized residual signal is For example, it generates a residual signal that has the lowest energy compared to the remaining residual signals. This means that, for example, the additional quantization of the residual signal compared to quantizing a larger signal It is advantageous to use fewer bits to quantize small signals. The controller 28 is an adaptive time-spectral transformer that applies one of the aforementioned transform kernels. The first control information 12a of the first channel and the second control information 12b of the second channel input to the converter 26 are Therefore, the time-spectrum converter 26 can determine the control information 12b of the second time-spectrum converter 26. configured to process a first channel and a second channel of a multi-channel signal, Furthermore, the multi-channel encoder may include a first channel and a second channel. The successive blocks of spectral values 4a' and 4b' are, for example, a multi-channel processor 42 for processing using multi-channel processing techniques; For example, spectral density can be obtained using sum-and-difference stereo coding or complex prediction. The processed blocks of block values 40a''' and 40b''' can be obtained. The spectral values are processed to obtain the encoded channels 40a''', 40b'''. The image processing system may further include an encoding processor 46 for processing the processed blocks. The encoding processor may use, for example, lossy or lossless audio compression methods. Audio signals can be coded using, for example, scalar quantization of spectral lines. , entropy coding, Huffman coding, channel coding, block codes or convolutional coding Codes, or forward error correction or automatic repeat request may be applied. Lossy audio compression may refer to the use of quantization based on psychoacoustic models. stomach.
[0082] According to a further embodiment, the first processed block of spectral values is a joint a first encoded representation of a multi-channel processing technique; and a second processed spectrum The block of values represents a second coded representation of the joint multi-channel processing technique. Therefore, the encoding processor 46 uses quantization and entropy coding to generate the first processing the processed blocks to form a first coded representation, and performing quantization and entropy. and processing the second processed block using bit coding to form a second coded representation. The first coded representation and the second coded representation are coded In other words, the audio signal may be formed into a bitstream representing the encoded audio signal. The first processing block uses complex stereo prediction to decode the encoded audio signal In an M / S encoded audio signal or MDCT encoded channel Furthermore, the second processing block can include a parameter for complex prediction. may contain a data or residual signal, or a side signal of an M / S coded audio signal. can.
[0083] FIG. 10 illustrates a multi-channel audio signal 200 having two or more channel signals. 2 shows an audio encoder for encoding a first channel signal, The first channel is shown at 201 and the second channel is shown at 202. Both signals are the first channel signal. A first composite signal 204 is generated using the signal 201, the second channel signal 202, and the prediction information 206. and a prediction residual signal 205 is input to an encoder calculator 203 for calculating the prediction residual signal 205. At this time, the predicted signal obtained from the first synthesized signal 204 and the prediction information 206 is When combined with the measured signal, a second composite signal is obtained, where the first composite signal and the second composite signal is obtained by combining the first channel signal 201 and the second channel signal 202 using a combining rule. The signal can be derived from the channel signal 202.
[0084] The prediction information is used to calculate the prediction information 206 so that the prediction residual signal meets the optimization target 208. The first composite signal 204 and the second composite signal 205 are generated by an optimizer 207 for calculating the first composite signal 204 and the second composite signal 205 . The residual signal 205 is input to a signal encoder 209 for encoding the first composite signal 204. 2. The residual signal 20 is coded to obtain a coded first composite signal 210. The coded first synthesis signal 210 is converted into a coded prediction residual signal 211. The multi-channel signal 2 is encoded by combining the residual signal 211 and the prediction information 206. 13, both coded signals 210 and 211 are sent to the output interface 21 It is entered into 2.
[0085] Depending on the implementation, the optimizer 207 may be configured to optimize the first channel signal 201 and the second channel signal 202. 202 or as indicated by lines 214 and 215. As shown, the first composite signal 214 and the second composite signal 215 are the same as those shown in FIG. 11A, which will be described later. The signal is obtained from the combiner 2031.
[0086] Figure 10 shows the coding gain being maximized, i.e. the bit rate being reduced as much as possible. The optimization target is shown as follows: the residual signal D is minimized with respect to α. In other words, the prediction information α is expressed as ||S-αM|| 2is selected so that This gives the solution for α shown in Figure 10. The signals S and M are Given a block-wise, spectral domain signal, the 2-norm of the argument of the notation ||…|| The first channel signal 201 and and the second channel signal 202 are input to the optimizer 207, Combination rules must be applied, and exemplary combination rules are shown in Figure 11C. The first composite signal 214 and the second composite signal 215 are input to the optimizer 207. In this case, the optimizer 207 does not need to implement the combination rules itself.
[0087] Other optimization targets may relate to perceptual quality. The optimization goal may be to maximize perceptual quality. Second, the optimizer needs additional information from the perceptual model. Other implementations of optimization targets use minimum or constant bitrates. Next, the optimizer 207 determines the required behavior for a particular α value. The method is implemented to perform a quantization / entropy coding operation to determine the bit rate. Therefore, α can be set to satisfy requirements such as minimum or constant bit rate. Other implementations of the optimization target can be set to In the case of the implementation of such optimization targets, The information about the resources required for the optimization is stored in the optimizer 207. Furthermore, combinations of these or other optimization targets are available. The matching is applied to control an optimizer 207 that calculates the forecast information 206. This can be done.
[0088] The encoder calculator 203 of FIG. 10 can be implemented in different ways. An embodiment is shown in FIG. 11A, where explicit combining rules are implemented in combiner 2031. An alternative exemplary implementation in which a matrix calculator 2039 is used is shown in FIG. The combiner 2031 in FIG. 11A implements the combining rules illustrated in FIG. 11C. This is a well-known intermediate encoding rule, and all A weighting factor of 0.5 is applied to the branch. However, depending on the implementation, other weighting factors may be used. It is not possible to implement any coefficients or weighting factors at all. Furthermore, other linear combination rules or nonlinear Other combination rules can also be applied, such as linear combination rules, and the decoder shown in FIG. As long as there is a corresponding inverse combining rule that can be applied to combiner 1162, We apply the opposite combination rule to that applied by the da. Therefore, the effect on the waveform is "balanced" by the prediction, i.e., the error is transmitted as a residual. Since the signal contains the reversible prediction rule, any reversible prediction rule can be used. This is because the predictive calculation by the encoder calculator 203 in accordance with the method 7 is a waveform saving process.
[0089] Combiner 2031 outputs first combined signal 204 and second combined signal 2032 . The first synthesis signal is input to a predictor 2033 and the second synthesis signal 2032 is input to a residual calculator 2034. 2034. The predictor 2033 calculates a predicted signal 2035, which is the second synthesis The signal 2031 is then combined with the signal 2032 to finally obtain the residual signal 205. The two channel signals 201 and 202 of the multi-channel audio signal are divided into two different and configured to combine the first and second composite signals 204 and 2032 in a manner Two different methods are shown in the exemplary embodiment of FIG. 11C. 3 is a signal obtained by combining the prediction information with the first synthesized signal 204 or the first synthesized signal 205 to obtain a predicted signal 2035. The signal derived from the composite signal may be applied to any It can be derived by any nonlinear or linear operation and performs weighted addition of certain values. Real to imaginary conversion / can be realized using linear filters such as The imaginary to real conversion is advantageous.
[0090] The residual calculator 2034 of FIG. 11A calculates the residual signal 2035 by subtracting the predicted signal 2035 from the second synthesized signal. However, other operations on the remaining calculators are also possible. Correspondingly, the combined signal calculator 1161 of FIG. 12A calculates the second combined signal 11 An addition operation is performed in which the decoded residual signal 114 and the prediction signal 1163 are added together to obtain . The calculation can be performed.
[0091] The decoder calculator 116 can be implemented in different ways. A first implementation is shown in FIG. This embodiment includes a predictor 1160, a composite signal calculator 1161, and a combiner 1162. The predictor combines the decoded first synthesized signal 112 and the prediction information 108 to generate a predictor 1162. The predictor 1160 receives the decoded 1st Prediction information 10 is added to the first synthesized signal 112 or to a signal derived from the decoded first synthesized signal. 8. Derivation for deriving a signal to which prediction information 108 is applied. The rule may be a real to imaginary transformation, or equivalently an imaginary to real transformation or weight Weighting operation, or equivalently, implementation, phase shift operation, or combined weighting / phase shift The predicted signal 1163 is used to calculate the decoded second synthesized signal 1165. To do this, the signal 112 and the decoded residual signal are input to the synthesis signal calculator 1161. and 1165 combine the decoded first composite signal and the second composite signal to generate a decoded The decoded first channel signal and the decoded second channel signal are output on output lines 1166 and 1167. 67 are input to a combiner 1162 which produces a decoded multi-channel audio signal. Alternatively, the decoder calculator may input the decoded first composite signal or signal M, a matrix calculator that receives as input the coded residual signal or signal D and the prediction information α 108; The matrix calculator 1168 calculates a transformation matrix shown as 1169 for the signal M, D to obtain output signals L, R, where L is the decoded first channel signal and R is the decoded second channel signal. The notation in FIG. 12B is This notation is used for ease of understanding. However, signals L and R are multi-channel signals with three or more channel signals. It will be apparent to those skilled in the art that the two channel signals may be any combination of two signals in a signal. The matrix operation 1169 is the same as the operations of blocks 1160, 1161, and 1162 in FIG. 12A. into a kind of "single shot" matrix calculation, and the input to the circuit in Figure 12A and Figure 1 The output from the circuit of 2A is input to the matrix calculator 1168 and to the matrix calculator 1 The outputs are identical to those from 168.
[0092] FIG. 12C shows an example of an inverse combining rule applied by combiner 1162 of FIG. 12A. In the well-known mid-side coding, the combining rule is L=M+S and R=MS. The combination rule used by the reverse combination rule in Figure 12C is similar to the decoder-side combination rule in The signal S to be calculated is the signal calculated by the composite signal calculator, i.e., the predicted signal on line 1163. It should be understood that the signal is a combination of the measured signal and the decoded residual signal on line 114. In this specification, signals on a line are sometimes named by the reference number of the line. It should be understood that there are lines sometimes indicated by the reference number itself. Therefore, the notation is such that a line carrying a signal indicates the signal itself. The lines can be hardwired physical lines, but they can also be computerized. In this implementation, there are no physical wires, but the signals represented by the wires are transmitted to some computing module. The data is transmitted from the module to other computing modules.
[0093] Figure 13A shows an implementation of the audio encoder. Compared to the first channel signal 55a in the time domain, the first channel signal 201 is Similarly, the second channel signal 202 is a time domain channel signal 55 b. The conversion from the time domain to the spectral representation is a time / frequency converter 50 for the first channel signal and a time / frequency converter 51 for the second channel signal. The spectral converters 50 and 51 are preferably implemented as real number converters. Transform algorithms include discrete cosine transform, real-valued Only a portion of the FFT, MDCT, or other transform that provides real-valued spectral values is used. Alternatively, both transforms can be expressed as follows: can be implemented as an imaginary transform like DST, MDST, or FFT, which are discarded Other transformations that provide only imaginary values can be used as well. One reason for using a pure imaginary transform is computational complexity, since each spectrum For tor values, should only a single value, such as magnitude or real part, be processed? This is because the phase or imaginary part must be processed. In contrast to a fully complex transformation, two values are used, namely the real and imaginary parts of each spectral line. Several parts must be processed, which increases the computational complexity by at least a factor of two. Another reason for using real-valued transforms here is that such transform sequences are usually It is critically sampled even in the presence of interconversion overlap, and Therefore, signal quantization and entropy coding (such as "MP3", AAC, or similar audio) (The standard "perceptual audio coding" paradigm implemented in audio coding systems) Provide an appropriate (and commonly used) area for
[0094] Figure 13A shows a side signal received at the "plus" input and a predictor at the "minus" input. a residual calculator 2034 as an adder that receives the prediction signal output by 2033; Furthermore, FIG. 13A shows that the predictor control information is encoded from the optimizer. A multi-channel audio signal generator that outputs a multiplexed bitstream representing the 13A. In particular, the prediction operation is performed by the equation on the right side of FIG. As shown, the side signal is predicted from the intermediate signal.
[0095] The predictor control information 206 is a factor as shown on the right side of FIG. In embodiments involving only real parts, such as the real part of the complex value α or the magnitude of the complex value α, If the part corresponds to a non-zero factor, the waveform structure of the intermediate signal and the side signal is similar. However, significant coding gains can be obtained when the amplitudes differ.
[0096] However, the predicted control information is the imaginary part of the complex factor or the phase information of the complex factor. If the imaginary part or phase information is different from zero, the present invention The brightness is significant for signals that are phase shifted from each other by values different from 0 or 180 degrees. Achieve coding gain and, except for phase shift, have similar waveform characteristics and similar amplitude relationships. do.
[0097] The predicted control information is complex-valued and is used for signals with different amplitudes and phase shifts. significant coding gain can be obtained. In this situation, operation 2034 determines whether the real part of the predictor control information is a part of the complex spectrum M. The complex prediction information is applied to the real part of the complex spectrum, and the complex prediction information is applied to the imaginary part of the complex spectrum. Next, in adder 2034, the result of this prediction calculation is added to the predicted actual spectrum. is the predicted imaginary spectrum of the subsignal S, and the predicted real spectrum is the real spectrum of the subsignal S. (band-wise) and the predicted imaginary spectrum is the imaginary part of the spectrum of S are subtracted to obtain the complex residual spectrum D.
[0098] The time domain signals L and R are real-valued signals, whereas the frequency domain signals can be real or complex-valued. If the frequency domain signal is real-valued, the transform is a real-valued transform. If the wavenumber domain signal is complex, the transform is a complex transform, which is called a time-frequency transform. This means that the input to and output of the frequency-time transform are real-valued, and the frequency domain signal is e.g. For example, it may be a complex-valued QMF domain signal.
[0099] FIG. 13B shows an audio decoder corresponding to the audio encoder shown in FIG. 13A. show.
[0100] The bitstream output by the bitstream multiplexer 212 of FIG. The bit stream demultiplexer 102 in FIG. The multiplexer 102 separates the bitstream into a downmix signal M and a residual signal D. The downmix signal M is input to the inverse quantizer 110a. The residual signal D is inversely quantized. Furthermore, the bitstream demultiplexer 102 demultiplexes the bitstream The predictor control information 108 from the stream is demultiplexed and input to the predictor 1160. The dequantizer 1160 outputs the predicted side signal α·M, and the combiner 1161 receives the signal output from the inverse quantizer 110b. The input residual signal is combined with the predicted side signal to obtain the final reconstructed side signal S. The side signal is then converted as shown in FIG. 12C for mid / side encoding. , are input to a combiner 1162 that performs sum-difference processing, for example. 62 to obtain a frequency domain representation of the left channel and a frequency domain representation of the right channel, Then, the frequency domain representation is converted into the corresponding frequency / time domain The signals are converted to a time domain representation by the time-domain converters 52 and 53.
[0101] Depending on the system implementation, if the frequency domain representation is a real-valued representation, the frequency / time transformation The converters 52 and 53 are real-valued frequency / time converters, and when the frequency domain representation is a complex-valued representation, In this case, it is a complex-valued frequency-to-time converter.
[0102] However, for efficiency, performing real-valued transforms is a 14A for the decoder and FIG. 14B for the decoder. The real-valued transforms 50 and 51 are MDCT, i.e., MDCT-IV, or According to the study, this is achieved by MDCT-II, MDST-II, or MDST-IV. The predicted information is calculated as a complex value having a real part and an imaginary part. Since M and S are real-valued spectra, the imaginary part of the spectrum does not exist, and the real A real-to-imaginary converter 2070 is provided to estimate the imaginary spectrum 6 from the real spectrum of the signal M. 00. This real-to-imaginary converter 2070 is part of the optimizer 207. The imaginary spectrum 600 estimated in block 2070 is then calculated together with the real spectrum M using the α option. The inputs to the optimizer stage 2071 are real-valued factors and and calculates prediction information 206 having an imaginary factor denoted by 2074. According to an embodiment, the real-valued spectrum of the first composite signal M is a side spectrum of the real part: To obtain the predicted signal that is subtracted from R It is multiplied by 2073. The number spectrum 600 has an imaginary part α indicated by 2074. I is multiplied with to obtain a further predicted signal This predicted signal is then subtracted from the real-valued side spectrum as shown in 2034b. Next, the prediction residual signal D is quantized in the quantizer 209b to obtain a real-valued space of M. The vector is quantized / encoded in block 209a. To obtain the coded complex α value transmitted to the stream multiplexer 212, The predictive information α is quantized and encoded in the quantizer / entropy encoder 2072. is advantageously provided, for example, and is ultimately entered into the bitstream as prediction information.
[0103] Regarding the position of the quantization / coding (Q / C) module 2072 for α, the multiplier 20 73 and 2074 are used to accurately set the (quantized) α, which is also used in the decoder. Therefore, 22072 is directly transferred to the output of 2071. Alternatively, the quantization of α can be considered in the optimization process of 2071. It can be thought that this is the case.
[0104] The encoder can calculate complex spectra, but all information is available. Therefore, the encoder is designed to generate a similar condition for the decoder shown in FIG. 14B. It is advantageous to perform a real to complex conversion in block 2070 of the coder. The real-valued coded spectrum of the first synthesis signal and the real-valued coded residual signal are Furthermore, the coded complex prediction information is obtained in block 65. The entropy decoding and inverse quantization are performed in R and the imaginary part α shown in 1160c I The weighting factors 1160b and 1160c are obtained. The intermediate signal output by 60c is added to the decoded and dequantized prediction residual signal. Specifically, a weighting unit 1160c uses the imaginary part of the complex prediction coefficient as a weighting coefficient. The spectral values input to are converted into a real-valued spectrum M by a real / imaginary converter 1160a. , which is performed in the same way as block 2070 of FIG. 20 for the encoder side. At the decoder side, the complex-valued representation of the intermediate or side signals is not available. This is in contrast to the coder side, where only the coded real-valued spectrum is This is because it is transmitted from the encoder to the decoder for speed and complexity reasons.
[0105] The real to imaginary transformer 1160a or the corresponding block 2070 in FIG. 14A is an international Publication No. 2004 / 013839 or International Publication No. 2008 / 014853 No. 6,980,933 or as disclosed in U.S. Pat. No. 6,980,933. Alternatively, any other implementation known in the art can be applied. Cut.
[0106] The embodiment is based on the idea that the proposed adaptive transform kernel switching is suitable for audio encoding such as HE-AAC. How can it be advantageously used in audio codecs? It further shows how to minimize or avoid the two issues mentioned above. , dealing with stereo signals with an inter-channel phase shift of about 90 degrees. Here, MDS The switch to T-IV based coding is used in one of the two channels. However, legacy MDCT-IV coding may be used in the other channel. In this case, MDCT-II coding is used in some channels and MDST-II coding is used in others. Cosine and sine functions are at 90 degrees to each other. Assuming that the phase-shifted transformation (cos(x)=sin(x+π / 2)) is The corresponding phase shift between the channel spectra is thus obtained in a conventional M / S-based 0 degrees or 1 degree, which can be coded very efficiently via joint stereo coding. This can be converted into an 80 degree phase shift, which is suboptimally coded by the conventional MDCT. As with harmonic signals, it may be advantageous in channels where mid-transition conversion is affected. be.
[0107] In both cases, harmonic and stereo signals with approximately 90 degrees inter-channel phase shift are generated. For the 1-bit signal, the encoder selects one of four kernels for each transform (see Figure 7). Each decoder that applies the transform kernel switching of the present invention uses the same Using the kernel, the signal can be properly reconstructed. To know which transformation kernel to use in one or more inverse transformations in a frame of The side information that explains the choice of transformation kernel, or left-right symmetry, is It should be transmitted by the corresponding encoder at least once per second. Now, let's explain the integration (i.e., modification) into the MPEG-H 3D audio codec. .
[0108] Further embodiments relate to audio coding, in particular to the Modified Discrete Cosine Transform (MDC The present invention relates to low-rate perceptual audio coding using lapped transforms such as T. By generalizing the MDCT coding principle to include three other similar transforms, The present invention is directed to two specific problems related to transform coding. The transformation between these four kernels in a channel or frame, or for each coded channel, Signal-adaptive and context-adaptive switching for each transition in a channel or frame To signal the kernel selection to the corresponding decoder, Side information may be transmitted in the coded bitstream.
[0109] FIG. 15 shows a schematic block diagram of a method 1500 for decoding an encoded audio signal. The method 1500 divides successive blocks of spectral values into overlapping successive blocks of time values. A step 1505 of converting and decoding successive blocks of time values to obtain decoded audio values is performed. Step 1510 of overlapping and adding the clocks; receiving control information and responding to the control information; A transformation kernel containing one or more transformation kernels with different symmetries on either side of the kernel is used. The first group of kernels and one or more transformation kernels with the same symmetry on both sides of the kernel are and a step 1515 of switching between a second group of transformation kernels including the first group of transformation kernels.
[0110] Figure 16 shows a schematic block diagram of a method 1600 for encoding an audio signal. 600 converts overlapping blocks of time values into contiguous blocks of spectral values. Step 1605 of converting the transformation kernels of the first group and the transformation kernels of the second group Control the time-spectral transformation to switch between group transformation kernels Step 1610 receiving the control information and, in response to the control information, a first group of transformation kernels including one or more transformation kernels having different symmetries; of a transformation kernel containing one or more transformation kernels with the same symmetry on both sides of the transformation kernel and switching 1615 between the first group and the second group.
[0111] In this specification, signals on a line are sometimes named by the reference number of the line. It should be understood that the lines are sometimes indicated by the reference numbers themselves. Therefore, the notation is such that a line carrying a signal indicates the signal itself. It can be a hardwired implementation of physical lines. In the implementation, there are no physical lines, but the signals represented by the lines are represented by some computational model. The data is transmitted from the module to other computing modules.
[0112] The present invention is based on the concept of block diagrams, where the blocks represent actual or logical hardware components. Although described in the context, the present invention may also be practiced by a computer-implemented method. In the latter case, the blocks represent corresponding method steps, and these steps are The functions performed by the corresponding logical or physical hardware blocks Represents Noh.
[0113] Although some aspects are described in the context of an apparatus, these aspects may also be implemented in block or digital form. If a device corresponds to a method step or feature of a method step, the corresponding method Similarly, the method steps described in the context of An aspect also represents a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the steps may be implemented using, for example, a microprocessor, a programmable computer, may be implemented by (or used in) a hardware device such as a computer or electronic circuit In some embodiments, one of the most important method steps Or several may be performed by such a device.
[0114] The transmitted or encoded signals of the present invention can be stored on a digital storage medium or or transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet. It is possible.
[0115] Depending on specific implementation requirements, embodiments of the invention may be implemented in hardware or in software. The implementation may be implemented using a floppy disk containing electronically readable control signals. Disk, DVD, Blu-ray, CD, ROM, PROM, and EPROM, EEPR It can be implemented using digital storage media such as ROM or flash memory, On top of that, they are programmed to run the respective methods. Therefore, the digital storage medium is a It may be computer readable.
[0116] Some embodiments of the present invention may be implemented in conjunction with a programmable computer system. a data carrier having an electrically readable control signal that can be read by the One of the methods described in the document is carried out.
[0117] Generally, embodiments of the present invention relate to a computer program product running on a computer. a computer having program code that operates to perform one of the methods when The program code can be implemented as a program product. It can be stored in a removable carrier.
[0118] Another embodiment is a computer program for carrying out one of the methods described herein. The program is stored on a machine readable carrier.
[0119] In other words, the method embodiment of the present invention is a computer program that runs on a computer A computer-readable medium having program code that, when executed, performs one of the methods described herein. It is a computer program that
[0120] Therefore, a further embodiment of the method of the present invention is to provide a data carrier (or digital storage medium) and non-transitory storage media or computer-readable media, such as storage media, a data carrier having recorded thereon a computer program for performing one of the methods described above; Digital or recording media are typically tangible and / or non-transitory.
[0121] Thus, a further embodiment of the method of the present invention is to carry out one of the methods described herein. A data stream or series of signals representing a computer program for The data stream or sequence of signals is transmitted over a data communication connection, for example. The information can be transmitted, for example, via the Internet.
[0122] Further embodiments are configured to perform one of the methods described herein. a processing means, e.g. a computer or programmable logic device, include.
[0123] A further embodiment is a computer for performing one of the methods described herein. This includes the computer on which the program is installed.
[0124] A further embodiment according to the invention is a device for carrying out one of the methods described herein. includes a device or system configured to transmit a computer program to a receiver (e.g., electronically or optically). The receiver may be, for example, a computer, a mobile device, The device or system may be, for example, a computer, a memory device, etc. A file server may be provided for transmitting computer programs to the receiver.
[0125] In some embodiments, a programmable logic device (e.g., a field programmable logic device) A programmable gate array (PGGA) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array (FPGA) may be used. The device may cooperate with a microprocessor to perform one of the methods described herein. In general, these methods are preferably implemented by any hardware device. It will be carried out.
[0126] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations in the specifications and details will be apparent to those skilled in the art. The present invention is limited only by the scope of the appended claims and is not to be construed as limiting the scope of the present invention. Thus, the intention is not to be limited by the specific details provided.
[0127] References [1] H. S. Malvar, Signal Processing with Lapped Transforms, Norwood: Artech Hous e, 1992. [2] J. P. Princen and A. B. Bradley, "Analysis / Synthesis Filter Bank Design Base d on Time Domain Aliasing Cancellation," IEEE Trans. Acoustics, Speech, and Signal Proc., 1986. [3] J. P. Princen, A. W. Johnson, and A. B. Bradley, "Subband / transform coding u sing filter bank design based on time domain aliasing cancellation," in IEEE ICASSP, vol. 12 , 1987. [4] H. S. Malvar, "Lapped Transforms for Efficient Transform / Subband Coding," IE EE Trans. Acoustics, Speech, and Signal Proc., 1990. [5] http: / / en.wikipedia.org / wiki / Modified_discrete_cosine_transform
Claims
1. A decoder (2) for decoding an encoded audio signal (4), comprising: The decoder A consecutive block of spectral values (4', 4'') is compared with a consecutive block of time values (10). an adaptive spectral-to-temporal converter (6) that converts To obtain the decoded audio value (14), a successive block of time values (10) an overlap-add processor (8) for overlapping and adding The adaptive spectrum-to-time converter (6) receives control information (12), Transformations that include one or more transformation kernels with different symmetries on either side of the kernel depending on the information a first group of kernels and one or more transformations that have the same symmetry on both sides of the transformation kernel; a second group of transformation kernels including the first group of transformation kernels, decoder.
2. The first group of transformation kernels has odd symmetry on the left side of the kernel and odd symmetry on the right side or one or more transformation kernels that have even symmetry in The second group of transformation kernels has even or odd symmetry on both sides of the kernel.
2. The decoder (2) of claim 1, comprising one or more transformation kernels having:
3. The first group of transform kernels is an inverse MDCT-IV transform kernel or an inverse MDS The second group of transformation kernels may include T-IV transformation kernels or may include inverse MDC 1 or claim 1, including a T-II transformation kernel or an inverse MDST-II transformation kernel. A decoder (2) according to paragraph 2.
4. The transformation kernels of the first group and the second group are based on the following equations: It's here, The at least one transformation kernel of the first group comprises: cs() = cos() and k 0 = 0.5 or cs() = sin() and k 0 = 0.5 It is based on the parameters or The at least one transformation kernel of the second group comprises: cs() = cos() and k 0 =0 or cs() = sin() and k 0 =1 It is based on the parameters Here, x i,n is the time domain output, C is a constant parameter, and N is the time window length. spec is the spectral value with M values for the block, where M is equal to N / 2. where i is the time block index and k is the spectral index indicating the spectral value. where n is a time index indicating the time value in block i, and n 0 is 4. The method according to claim 1, wherein the parameter is a constant parameter that is a number or zero. Decoder (2).
5. The control information (12) includes a current bit indicating the current symmetry for the current frame. Including, The adaptive spectro-temporal converter (6) determines whether the current bit is used in the previous frame. When the first group exhibits the same symmetry as the second group, the It is configured so that The adaptive spectro-temporal converter determines whether the current bit is used in the previous frame. When a compound exhibits a different symmetry from that of the compound, it is switched from the first group to the second group. A decoder (2) according to any one of claims 1 to 4, configured to switch 。
6. The adaptive spectral-to-temporal converter (6) generates a current symmetry-indicating current frame. When the current bit exhibits the same symmetry as that used in the previous frame, the second group configured to switch the loop to the first group; The adaptive spectro-temporal converter (6) determines whether the current bit is equal to the previous bit in the previous frame. Indicates the current symmetry of the current frame, which has a different symmetry than was used When the second group is not switched to the first group, A decoder (2) according to any one of claims 1 to 5.
7. The adaptive spectro-temporal converter (6) receives control information (12) about the previous frame. ) from the coded audio signal (4) and the current frame following said previous frame. The control information for the current frame is encoded in the control data section of the current frame. configured to read from a received audio signal; or The adaptive spectral-to-temporal converter (6) selects the control data set of the current frame. The control information (12) is read from the control data section of the previous frame. for the previous frame from the decoder settings applied to that previous frame. The method according to any one of claims 1 to 6, wherein the method is configured to extract the control information (12) for the A decoder (2) according to any one of claims 1 to 4.
8. The adaptive spectral-to-temporal converter (6) applies a conversion kernel according to the following table: It is configured to: Here, sym i is the control information of the current frame at index i, The sym i-1 is the index i -1 control information of the previous frame in A decoder (2) according to any one of claims 1 to 7.
9. The processed spectral values for the first multi-channel and the second multi-channel are spectral values representing the first and second multi-channels to obtain a block and receiving a block of the received block according to a joint multi-channel processing technique. and a multi-channel processor (40) for processing said adaptive spectrum A time processor (6) uses the control information for the first multi-channel to The processed blocks for the first multi-channel and the processed blocks for the second multi-channel are then the processed block for the second multi-channel using control information for the second multi-channel; A decoder according to any one of claims 1 to 8, configured to process locks. Da (2).
10. The multi-channel processor may further include: complex prediction control information associated with said block of spectral values is used to apply complex prediction. A decoder (2) according to claim 9, configured as follows:
11. The multi-channel processor performs the preceding processing in accordance with the joint multi-channel processing technique. and configured to process the received block, the received block being a coded residual signal of the multi-channel representation and said second multi-channel representation; and the multi-channel processor processes the residual signal and the further coded signal as and calculating the first multi-channel signal and the second multi-channel signal using 11. A decoder according to claim 9 or 10, configured to:
12. An encoder (22) for encoding an audio signal (24), comprising: The encoder comprises: Overlapping blocks of time values (30) into consecutive blocks of spectral values (4', 4'') an adaptive time-to-spectral converter for converting A transformation kernel of a first group of transformation kernels and a transformation kernel of a second group of transformation kernels a controller for controlling the time-spectrum converter to switch between a time-spectrum conversion kernel and a time-spectrum conversion kernel; (28), The adaptive time-to-spectral converter receives control information (12) and In response, a transformation kernel including one or more transformation kernels with different symmetries on either side of the kernel is created. a first group of transform kernels and one or more transform kernels with the same symmetry on both sides of the transform kernel; a second group of transformation kernels including the first group of transformation kernels, Coda.
13. For a current frame, the transformation coefficients used to generate the current frame are A coded audio signal (4) is generated having control information (12) indicating the symmetry of the channel.
13. The encoder of claim 12, further comprising an output interface (32) for generating (22)。
14. The output interface (32) is configured to output a signal indicating that the current frame is an independent frame. If the control data section of the current frame contains the Contains symmetry information for the previous frame, or If the current frame is a dependent frame, the control data of the current frame A section contains only symmetry information for the current frame and not the pair of the previous frame.
14. The encoder according to claim 12 or 13, which is configured not to include ambiguity information. 22)。
15. The first group of transformation kernels has odd symmetry on the left and even symmetry on the right. One or more transformation kernels that are symmetric or inversely symmetric, or The second group of panels consists of one or more transformation columns with even or odd symmetry on both sides. An encoder (22) according to any one of claims 12 to 14, having a channel.
16. The first group of transform kernels is the MDCT-IV transform kernel or MDST- Alternatively, the second group of transform kernels may include MDCT-IV transform kernels. The method according to any one of claims 12 to 15, further comprising: a MDST-II transformation kernel or an MDST-II transformation kernel.
10. The encoder according to claim 1,
17. The controller (28) performs MDCT-IV or MDST- II, or MDST-IV is followed by MDST-IV or or MDCT-II followed by MDCT-IV or MDCT-II. MDST-II is followed by MDST-IV or MDST-II. or MDCT-II. An encoder (22) according to paragraph 1.
18. The controller (28) synchronizes the frames of the first channel with the corresponding frames of the second channel. and a frame of the first channel to determine the transformation kernel. and configured to analyze overlapping blocks of said time values (30) having a first channel and a second channel. The encoder (22) according to any one of claims 12 to 17.
19. The time-to-spectral converter (26) converts a first channel and a second channel of a multi-channel signal into a time-to-spectral signal. The encoder (22) is configured to process a second channel, Using joint multi-channel processing techniques to obtain blocks of vector values, processing said successive blocks of spectral values of one channel and said second channel; a multi-channel processor (40) for decoding the encoded channels; an encoding processor (46) for processing the blocks of processed spectral values; The encoder (22) according to any one of claims 12 to 18, further comprising:
20. The first block of processed spectral values is subjected to the joint multi-channel processing the second block of processed spectral values represents a first coded representation of the technique, represents a second coded representation of a joint multi-channel processing technique, 46) processes the first processed block using quantization and entropy coding. and a coding processor (4) configured to process the first coded representation to form a first coded representation. 6) processing the second processed block using quantization and entropy coding; to form a second coded representation, and the coding processor is configured to: using the first coded representation and the second coded representation, and forming a bitstream of the resulting audio signal. An encoder (22) according to any one of claims 12 to 19.
21. A method (1500) for decoding an encoded audio signal, comprising: converting successive blocks of spectral values into successive blocks of time values; Overlap and add successive blocks of time values to obtain the decoded audio value Steps and receiving control information and generating a kernel having different symmetries on both sides according to the control information; a first group of transformation kernels including one or more transformation kernels; and a second group of transformation kernels that includes one or more transformation kernels with the same symmetry. The method includes the step of switching
22. A method (1600) for encoding an audio signal, comprising: A sequence that converts overlapping blocks of time values into successive blocks of spectral values. Tep and A transformation kernel of a first group of transformation kernels and a transformation kernel of a second group of transformation kernels controlling the time-to-spectral transform to switch between kernels; receiving control information and generating a kernel having different symmetries on both sides according to the control information; a first group of transformation kernels including one or more transformation kernels; and a second group of transformation kernels that includes one or more transformation kernels with the same symmetry. a method including the step of switching
23. When running on a computer or processor, the method according to claim 21 or claim 22 A computer program for carrying out the method of claim 1.
24. Multi-channel processing refers to joint stereo processing or the joint processing of two or more channels. Multi-channel signals refer to two or more channels. An apparatus, method or computer program according to any one of claims 1 to 23, comprising: Rum.