Encoder Using Forward Aliasing Elimination
By introducing a second syntax portion to predict forward aliasing cancellation data, the codec achieves error robustness and handles frame loss effectively when switching between different coding modes.
Patent Information
- Application Number
- JP2024064918
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2010-08-10
- Filing Date
- 2024-04-12
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2031-07-07
AI Technical Summary
Existing codecs struggle with error robustness and frame loss when switching between time-domain aliasing cancellation transform coding mode and time-domain coding mode, particularly due to the lack of explicit aliasing cancellation for transitions involving ACELP modes.
Incorporating a second syntax portion in the codec that allows the decoder to predict the presence of forward aliasing cancellation data, enabling error-robust operation even in the presence of frame loss by ensuring proper decoding without relying on speculative data processing.
The proposed solution enhances the codec's error robustness and ability to handle frame loss, preventing decoding failures and maintaining coding efficiency in error-prone environments.
Smart Images

Figure 0007693891000003 
Figure 0007693891000004 
Figure 0007693891000005
Abstract
Description
Technical Field
[0001] The present invention relates to a codec that supports a time-domain aliasing cancellation transform coding mode, a time-domain coding mode, and forward aliasing cancellation for switching between both modes.
Background Art
[0002] For example, in order to encode a general audio signal that represents a mixture of various types of audio signals such as voice and music, it is preferable to mix different coding modes. Each individual coding mode can be adapted to a specific audio type, and thus, a multi-mode audio encoder can utilize changing the coding mode over time in response to changes in the type of audio content. In other words, a multi-mode audio encoder can determine to use, for example, a coding mode specialized for encoding voice in particular, to encode a portion of an audio signal having voice content, and to use another coding mode to encode various portions of the audio content representing non-voice content such as music. Time-domain coding modes such as the codebook excited linear prediction coding mode tend to be more suitable for encoding voice content, but, for example, as far as encoding music is concerned, the transform coding mode tends to have better performance than the time-domain coding mode.
[0003] Solutions already exist to address the problem of handling the coexistence of various audio types within a single audio signal. The newly emerging USAC, for example, proposes a switch between, among other things, a frequency-domain coding mode that mainly follows the AAC standard and two further linear prediction modes similar to the subframe mode of the AMR-WB+ standard, namely, the MDCT (Modified Discrete Cosine Transformation)-based variants of the TCX (TCX = transform coded excitation) mode and the ACELP (adaptive codebook excitation linear prediction) mode. More precisely, in the AMR-WB+ standard, TCX is based on the DFT transform, while in USAC, TCX has an MDCT transform base. A specific framing structure is used to switch between an FD coding region similar to AAC and a linear prediction region similar to AMR-WB+. The AMR-WB+ standard itself uses its own framing structure that forms a subframing structure in relation to the USAC standard. The AMR-WB+ standard allows for a specific sub-division configuration that sub-divides the AMR-WB+ frame into smaller TCX and / or ACELP frames. Similarly, the AAC standard uses a base framing structure but allows for the use of different window lengths to transform-code the frame content. For example, a long transform length associated with a long window may be used, or eight short-length windows may be used with associated short-length transforms.
[0004] MDCT causes aliasing. This applies, for example, at the boundaries of TXC and FD frames. In other words, aliasing occurs in the overlapping region of the window and is canceled with the help of adjacent frames, just like in any frequency-domain coder that uses MDCT. That is, for the transition between two FD frames, or between two TCX (MDCT) frames, or the transition from FD to TCX, or from TCX to FD, there is implicit aliasing cancellation by the overlap / add process during reconstruction on the decoding side. After overlap / add, there is no aliasing left. However, in the case of transitions related to ACELP, there is no specific aliasing cancellation. Therefore, a new tool called FAC (forward aliasing cancellation) must be introduced. FAC will cancel the aliasing arising from adjacent frames when the adjacent frames are different from ACELP.
[0005] In other words, whenever a transition occurs between a transform coding mode and a time-domain coding mode such as ACELP, the problem of aliasing cancellation arises. To perform the conversion from the time domain to the spectral domain as efficiently as possible, time-domain aliasing cancellation transform coding such as MDCT, i.e., a coding mode using an overlapped transform, is used. Here, the overlapping window part of the signal is transformed using a transform where the number of transform coefficients per part is less than the number of samples per part, and as a result, aliasing occurs for each individual part. This aliasing is canceled by time-domain aliasing cancellation, i.e., by adding the overlapping aliasing parts of adjacent re-transformed signal parts. MDCT is this type of time-domain aliasing cancellation transform. Unfortunately, TDAC (time-domain aliasing cancellation) cannot be used for transitions between the TC coding mode and the time-domain coding mode.
[0006] To solve this problem, forward aliasing cancellation (FAC) can be used, whereby whenever a change occurs in the encoding mode from transform encoding to temporal domain encoding, within the data stream, the encoder sends a signal of additional FAC data in the current frame. However, this requires comparing the encoding modes of consecutive frames in order for the decoder to check whether the currently decoded frame contains FAC data in its syntax. This also means that there can be frames for which the decoder need not know whether it needs to read or parse FAC data from the current frame. In other words, if one or more of those frames are lost during transmission, the decoder is unaware of whether a change in the encoding mode has occurred or whether the bitstream of the encoded data of the current frame contains FAC data with respect to the immediately following (received) frame. Thus, the decoder has to discard the current frame and wait for the next frame. Alternatively, the decoder can parse the current frame by performing two decoding attempts, one assuming the presence of FAC data and the other assuming the absence of FAC data, and then can determine whether one of both options fails. The decoding process is likely to crash the decoder under one of two conditions. That is, in reality the latter option is not a viable approach. The decoder must always know how to interpret the data and must not rely on its own speculation regarding how to process the data. Summary of the Invention Problems to be Solved by the Invention
[0007] Accordingly, it is an object of the present invention to provide a codec that is error robust or robust to frame loss by supporting switching between the temporal domain aliasing cancellation transform encoding mode and the temporal domain encoding mode.
Means for Solving the Problem
[0008] This object is achieved by the subject matter of any of the independent claims appended hereto.
[0009] The present invention is based on the discovery that when a parser of a decoder predicts that the current frame contains forward aliasing cancellation data and thus reads the forward aliasing cancellation data from the current frame, and a second operation of not predicting that the current frame contains forward aliasing cancellation data and thus not reading the forward aliasing cancellation data from the current frame, depending on which is selected, when an additional syntax portion is added to the frame, a codec that is more error-robust or robust to frame loss, which supports switching between a temporal aliasing cancellation transform coding mode and a temporal coding mode, is achievable. In other words, while the supply of the second syntax portion slightly loses coding efficiency, the second syntax portion only provides the possibility of using the codec in the case of a communication channel having frame loss. Without the second syntax portion, the decoder will not be able to decode any data stream portion after the loss and will crash when attempting to resume syntax analysis. Thus, in an error-prone environment, the coding efficiency is prevented from becoming zero by the introduction of the second syntax portion.
[0010] Further preferred embodiments of the present invention are the subject matter of the dependent claims. Further, preferred embodiments of the present invention are described in more detail below with reference to the drawings.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16A
Figure 16B
Figure 17A
Figure 17B
Figure 18
Figure 19A
Figure 19B
Figure 20A
Figure 20B
Figure 21A
Figure 21B
Figure 22
DETAILED DESCRIPTION OF THE INVENTION
[0012] Figure 1 shows a decoder 10 according to an embodiment of the present invention. The decoder 10 is for decoding a data stream including a series of frames 14a, 14b, and 14c in which time segments 16a to 16c of an information signal 18 are encoded, respectively. As shown in FIG. 1, the time segments 16a to 16c are directly adjacent to each other and are non-overlapping segments that are ordered over time. As shown in FIG. 1, the time segments 16a to 16c may be of equal size, but other embodiments are also possible. Each of the time segments 16a to 16c is encoded in each one of the frames 14a to 14c. In other words, each time segment 16a to 16c is uniquely associated with one of the frames 14a to 14c, which also have a defined order among them according to the order of the segments 16a to 16c encoded in the frames 14a to 14c. FIG. 1 suggests that each frame 14a to 14c is of equal length, for example, measured in encoded bits, but this is of course not obligatory. Rather, the lengths of the frames 14a to 14c can vary depending on the complexity of the time segments 16a to 16c with which each frame 14a to 14c is associated.
[0013] To facilitate the description of the embodiments summarized below, it is assumed that the information signal 18 is an audio signal. However, it should be noted that the information signal can also be other signals such as signals output by physical sensors or other such things, for example, signals from optical sensors. In particular, the signal 18 can be sampled at a specific sampling rate, and the time segments 16a to 16c can cover directly consecutive portions of this signal 18 that are equal in time and number of samples, respectively. The number of samples for each of the time segments 16a to 16c can be, for example, 1024 samples.
[0014] The decoder 10 includes a parser 20 and a reconstructor 22. The parser 20 is configured to analyze the data stream 12 and, when analyzing the data stream 12, read a first syntax portion 24 and a second syntax portion 26 from the current frame 14b, i.e., the frame to be currently decoded. In FIG. 1, it is assumed by way of example that frame 14a is the frame that was decoded immediately before, while frame 14b is the frame to be currently decoded. Each of the frames 14a to 14c has a first syntax portion and a second syntax portion incorporated therein with the meaning outlined below or by that meaning. In FIG. 1, the first syntax portion in the frames 14a to 14c is indicated by a box having a "1" therein, and the second syntax portion is indicated by a box named "2".
[0015] Naturally, each of the frames 14a to 14c also has additional information incorporated therein, which is for indicating the associated time segments 16a to 16c in a way that will be outlined in more detail below. This information is shown in FIG. 1 by hatched blocks, and the reference numeral 28 is used for the additional information of the current frame 14b. The parser 20 is also configured to read the information 28 from the current frame 14b when analyzing the data stream 12.
[0016] The reconstructor 22 is configured to reconstruct the current time segment 16b of the information signal 18 associated with the current frame 14b based on further information 28, using one of a selected time-domain aliasing cancellation transform decoding mode and a time-domain decoding mode. That selection depends on the first syntax element 24. Both decoding modes also differ from each other by the presence or absence of a transition back from the spectral domain to the time domain using an inverse transform. (In addition to its corresponding transform) that inverse transform causes aliasing with respect to individual time segments, but for transitions at the boundaries between consecutive frames encoded in the time-domain aliasing cancellation transform coding mode, it is compensated by time-domain aliasing cancellation. The time-domain decoding mode does not require any inverse transform. Rather, its decoding maintains being in the time domain. Thus, generally speaking, the time-domain aliasing cancellation transform decoding mode of the reconstructor 22 is involved in the inverse transform being performed by the reconstructor 22. This inverse transform maps a first number of transform coefficients, such as obtained from the information 28 of the current frame 14b (which is the TDAC transform decoding mode), to a retransformed signal segment having a sample length of a second number of samples larger than the first number that thereby causes aliasing. The time-domain decoding mode may relate to a linear prediction decoding mode where the excitation and linear prediction coefficients are reconstructed from the information 28 of the current frame which is in that case the time-domain coding mode.
[0017] As thus made clear from the above description, in the time domain aliasing cancellation transform decoding mode, the reconstructor 22 obtains a signal segment for reconstructing an information signal in each time segment 16b by re-conversion from the information 28. The re-converted signal segment is longer than the current time segment 16b actually is, includes the time segment 16b, and is involved in the reconstruction of the information signal 18 in the time portion that extends beyond and spreads out. FIG. 1 shows a conversion window 32 used when converting the original signal, or when performing both conversion and re-conversion. As shown in the figure, the window 32 can include its starting zero portion 321 and its ending zero portion 322, and the aliasing portions 323 and 324 at the rise and fall of the current time segment 16b. The non-aliasing portion 325 where there is one window 32 can be located between the aliasing portions 323 and 324. The zero portions 321 and 322 are optional. It is also possible to have only one of the zero portions 321 and 322. As shown in FIG. 1, the window function can monotonically increase / decrease within the range of the aliasing portion. Aliasing occurs within the range of the aliasing portions 323 and 324 where the window 32 continuously rises from zero to 1 or vice versa. Also, as long as the previous and subsequent time segments are encoded in the time domain aliasing cancellation transform coding mode, the aliasing is not critical. This possibility is shown in FIG. 1 with respect to the time segment 16c. The dotted line indicates each conversion window 32' for the time segment 16c, and its aliasing portion occurs simultaneously with the aliasing portion 324 of the current time segment 16b. Adding the re-converted segment signals of the time segments 16b and 16c by the reconstructor 22 cancels out the aliasing of both re-converted signal segments with respect to each other.
[0018] However, when the previous or subsequent frame 14a or 14c is encoded in the time-domain coding mode, the transition between different coding modes occurs at the rising or falling edge of the current time segment 16b, and, to account for each aliasing, the data stream 12 includes forward aliasing cancellation data in each frame immediately following the transition to enable the decoder 10 to compensate for the aliasing occurring at each such transition. For example, it may also happen that the current frame 14b is for the time-domain aliasing cancellation transform coding mode, but the decoder 10 is unaware whether the previous frame 14a was for the time-domain coding mode. For example, frame 14a may be lost during transmission and thus the decoder 10 has no access to it. However, depending on the coding mode of frame 14a, the current frame 14b includes forward aliasing cancellation data to compensate for the aliasing occurring at the aliasing portion 323 or otherwise. Similarly, when the current frame 14b is for the time-domain coding mode and the previous frame 14a was not received by the decoder 10, the current frame 14b has forward aliasing cancellation data incorporated therein or independent of the mode of the previous frame 14a. In particular, when the previous frame 14a is for another coding mode, namely the time-domain aliasing cancellation transform coding mode, the forward aliasing cancellation data is present in the current frame 14b to cancel the aliasing that would occur at the boundary between time segments 16a and 16b if it were not there. However, when the previous frame 14a is for the same coding mode, namely the time-domain coding mode, the parser 20 need not predict the presence of forward aliasing cancellation data in the current frame 14b.
[0019] Therefore, the parser 20 utilizes the second syntax portion 26 to check whether the forward aliasing cancellation data 34 exists in the current frame 14b. In parsing the data stream 12, the parser 20 predicts whether the current frame 14b contains the forward aliasing cancellation data 34, and thus can select one of a first operation of reading the forward aliasing cancellation data 34 from the current frame 14b and a second operation of not predicting that the current frame 14b contains the forward aliasing cancellation data 34 and thus not reading the forward aliasing cancellation data from the current frame 14b, and the selection depends on the second syntax portion 26. If it exists, the reconstructor 22 is configured to perform forward aliasing cancellation at the boundary between the current time segment 16b and the previous time segment 16a of the previous frame 14a using the forward aliasing cancellation data.
[0020] Thus, compared to the situation without the second syntax portion, the decoder of FIG. 1 does not need to discard the current frame 14b or fail and interrupt the parsing even if, for example, the coding mode of the previous frame 14a is not known to the decoder 10 due to frame loss. Rather, the decoder 10 can utilize the second syntax portion 26 to check whether the current frame 14b has the forward aliasing cancellation data 34. In other words, the second syntax portion provides a clear criterion regarding whether one of two alternatives, i.e., the FAC data for the boundary with the previous frame, exists, is applied, and ensures that any decoder can perform the same operation regardless of their embodiments even in the case of frame loss. Thus, the embodiment outlined above introduces a mechanism for solving the problem of frame loss.
[0021] Before describing more detailed embodiments below, encoders capable of generating the data stream 12 of FIG. 1 are each illustrated by FIG. 2. The encoder of FIG. 2 is generally denoted by reference numeral 40 and is for encoding an information signal into the data stream 12 such that the data stream 12 includes a sequence of frames in which time segments 16a to 16c of the information signal are respectively encoded therein. The encoder 40 includes a constructor 42 and an inserter 44. The constructor is configured to encode the current time segment 16b of the information signal into the information of the current frame 14b using one of a first selected time-domain aliasing cancellation transform coding mode and a time-domain coding mode. The inserter 44 is configured to insert information 28 into the current frame 14b together with a first syntax portion 24 and a second syntax portion 26, where the first syntax portion indicates a first selection, i.e., a selection of the coding mode. The constructor 42 is alternatively configured to determine forward aliasing cancellation data for forward aliasing cancellation at the boundary between the current time segment 16b and the previous time segment 16a of the previous frame 14a, such that when the current frame 14b and the previous frame 14a are encoded using different ones of the time-domain aliasing cancellation transform coding mode and the time-domain coding mode, the forward aliasing cancellation data 34 is inserted into the current frame 14b, and when the current frame 14b and the previous frame 14a are encoded using the same ones of the time-domain aliasing cancellation transform coding mode and the time-domain coding mode, no forward aliasing cancellation data is inserted into the current frame 14b. That is, the constructor 42 of the encoder 40, in the sense of some optimizations, whenever it determines that it is preferable to switch from one of the two coding modes to the other, the constructor 42 and the inserter 44 are configured to determine and insert the forward aliasing cancellation data 34 into the current frame 14b, while when the coding mode is maintained between frames 14a and 14b, the FAC data 34 is not inserted into the current frame 14b.In order for the decoder to be able to obtain from the current frame 14b without knowing about the content of the previous frame 14a regarding whether the FAC data 34 is in the current frame 14b, a specific syntax part 26 is set depending on whether the current frame 14b and the previous frame 14a are encoded using the same or different time-domain aliasing cancellation transform coding mode and time-domain coding mode. A specific example for understanding the second syntax part 26 is outlined below.
[0022] In the following, an embodiment is described, whereby the codec belonging to the decoder and encoder of the above-described embodiment supports a special type of frame structure, whereby the frames 14a to 14c themselves have two different versions of the time-domain aliasing cancellation transform coding mode according to subframing. In particular, according to these embodiments shown further below, the first syntax part 24 associates each frame read thereby with a first frame type called the FD (frequency domain) coding mode below, or a second frame type called the LPD coding mode below, and when each frame is of the second frame type, the subframes of the sub-division of each frame consisting of several subframes are associated with one each of the first subframe type and the second subframe type. As outlined in more detail further below, the first subframe type is related to the corresponding subframe that is TCX encoded, while on the other hand, the second subframe type can be related to each of these subframes encoded using ACELP, i.e., Adaptive Codebook Excitation Linear Prediction. Also, any other codebook excitation linear prediction coding mode can be used in the same way.
[0023] The reconstructor 22 of FIG. 1 is configured to handle these different coding mode possibilities. For this purpose, the reconstructor 22 can be constructed as shown in FIG. 3. According to the embodiment of FIG. 3, the reconstructor 22 includes two switches 50 and 52 and these decoding modules 54, 56, and 58, each of which is configured to decode a specific type of frame and sub-frame, as will be described in more detail below.
[0024] Switch 50 has an input into which the information 28 of the currently decoded frame 14b enters and a control input through which the switch 50 is controllable depending on the first syntax part 25 of the current frame. Switch 50 has two outputs, one of which is connected to the input of the decoding module 54 that plays a role in FD decoding (FD = frequency domain), and the other one is connected to the input of the sub-switch 52, which also has two outputs, one of which is connected to the input decoding module 56 that plays a role in transform-coded excitation linear prediction decoding, and the other one is connected to the input of the module 58 that plays a role in codebook excitation linear prediction decoding. All the coding modules 54-58 output signal segments that reconstruct the time segments associated with each frame and sub-frame obtained by each decoding mode for these signal segments. And the transition handler 60 receives signal segments at each of its inputs to perform the transition processing and aliasing cancellation described above and further described in more detail below on its output of the reconstructed information signal. The transition handler 60 uses the forward aliasing cancellation data 34 as shown in FIG. 3.
[0025] According to the embodiment of FIG. 3, the reconstructor 22 operates as follows. When the first syntax portion 24 associates the current frame with the first frame type, i.e., the FD coding mode, the switch 50 transfers the information 28 to the FD decoding module 54 using frequency domain decoding as the first version of the time domain aliasing cancellation transform decoding mode to reconstruct the time segment 16b associated with the current frame 15b. Otherwise, i.e., when the first syntax portion 24 associates the current frame 14b with the second frame type, i.e., the LPD coding mode, the switch 50 transfers the information 28 operating in the sub-frame structure of the current frame 14 to the sub-switch 52 instead. More precisely, according to the LPD mode, the frame is divided into one or more sub-frames. Regarding the subsequent figures, as outlined in more detail below, the sub-division corresponds to the sub-division of the corresponding time segment 16b into non-overlapping sub-parts of the current time segment 16b. The syntax portion 24 indicates for each one or more sub-parts whether it is associated with the first sub-frame type or the second sub-frame type. If each sub-frame is for the first sub-frame type, the sub-switch 52 transfers each information 28 belonging to that sub-frame to the TCX decoding module 56 to use transform coding excited linear prediction decoding as the second version of the time domain aliasing cancellation transform decoding mode and reconstruct each sub-part of the current time segment 16b. However, if each sub-frame is for the second sub-frame type, the sub-switch 52 transfers the information 28 to the module 58 to perform codebook excited linear prediction coding as the time domain decoding mode and reconstruct each sub-part of the current time signal 16b.
[0026] The reconstructed signal segments output by the modules 54 to 58 are assembled by the transition handler 60 in the correct (display) time order regarding the execution of each transition process and overlap-and-add and time domain aliasing cancellation process as described above and as described in more detail below.
[0027] In particular, the FD decoding module 54 can be constructed as shown in Fig. 4 and can operate as described below. According to Fig. 4, the FD decoding module 54 includes an inverse quantizer 70 and a retransformer 72 serially connected to each other. As mentioned above, if the current frame 14b is an FD frame, it is transferred to the module 54, and the inverse quantizer 70 performs a spectrally variable inverse quantization of the transform coefficient information 74 within the information 28 of the current frame 14b, using the scale factor information 76 also contained by the information 28. The scale factor was determined at the encoder side, for example, using psychoacoustic principles to keep the quantization noise below the human masking threshold.
[0028] The retransformer 72 then performs a retransform on the inverse quantized transform coefficient information to obtain a retransformed signal segment 78 extending in time throughout and beyond the time segment 16b associated with the current frame 14b. As will be outlined in more detail below, the retransformation performed by the retransformer 72 is an IMDCT (Inverse Modified Discrete Cosine Transform) including a DCT IV followed by an unfolding operation in which windowing is performed using a retransform window that may be equal to or deviate from the transform window used in generating the transform coefficient information 74 by performing the steps described above in reverse order, i.e. a windowing and then a folding operation followed by a DCT IV, followed by a quantization that may be manipulated according to psychoacoustic principles to keep the quantization noise below a masking threshold.
[0029] It is worth noting that the amount of the conversion coefficient information 28 is due to the TDAC characteristic of the inverse conversion of the inverse converter 72, which is less than the number of samples of the reconstructed signal segment 78. In the case of the IMDCT, the number of conversion coefficients within the range of the information 47 is rather equal to the number of samples of the time segment 16b. That is, the underlying conversion requires time-domain aliasing cancellation in order to eliminate the aliasing generated by the conversion at the boundaries of the current time segment 16b, i.e., the rising and falling edges, and can be called a critically sampling conversion.
[0030] As a minor point, it should be noted that, similar to the sub-frame structure of the LPD frame, the FD frame can also be the subject of a sub-framing structure. For example, the FD frame can be for the long-window mode, in which a single window is used to window the signal portion that extends beyond the rising and falling edges of the current time segment in order to encode each time segment, or it can be for the short-window mode, in which each signal portion that extends beyond the boundary of the current time segment of the FD frame is sub-divided into smaller sub-portions that are each affected by each windowing and conversion. In that case, the FD decoding module 54 outputs the reconstructed signal segment for the sub-portion of the current time segment 16b.
[0031] After describing a possible embodiment of the FD decoding module 54, possible embodiments of the TCX LP decoding module and the codebook excited LP decoding modules 56 and 58 are described with respect to FIG. 5, respectively. In other words, FIG. 5 handles the case where the current frame is an LPD frame. In that case, the current frame 14b is constructed into one or more sub-frames. In this case, the structuring into three sub-frames 90a, 90b, and 90c is shown. It is possible that the structuring is by default restricted to certain sub-structuring possibilities. Each of the sub-parts is associated with one of the sub-parts 92a, 92b, and 92c of the current time segment 16b. That is, one or more sub-parts 92a to 92c cover the entire time segment 16b without overlap and without gaps. According to the order of the sub-parts 92a to 92c within the time segment 16b, the order is defined within the sub-frames 92a to 92c. As shown in FIG. 5, the current frame 14b is not completely sub-divided into sub-frames 90a to 90c. Further in other words, the LPC information is also sub-structured into individual sub-frames, but as will be described in more detail below, some parts of the current frame 14b belong to all sub-frames as LPC information, generally the first and second syntax parts 24 and 26, the FAC data 34, and possibly further data.
[0032] To process the TCX sub-frame, the TCX LP decoding module 56 includes a spectral weight extraction derivator 94, a spectral weighter 96, and a re-converter 98. For the sake of explanation, it is assumed that the first sub-frame 90a is a TCX sub-frame, while the second sub-frame 90b is an ACELP sub-frame.
[0033] To process the TCX subframe 90a, the extractor 94 derives a spectral weighting filter from the LPC information 104 in the information 28 of the current frame 14b, and the spectral weightor 96 uses the spectral weighting filter received from the extractor 94 to spectrally weight the transform coefficient information within the extent of the location of the subframe 90a, as indicated by the arrow 106.
[0034] Next, the reconverter 98 reconverts the spectrally weighted transform coefficient information to obtain a reconverted signal segment 108 that spreads over and beyond the entire subpart 92a of the current time segment at time t. The reconversion performed by the reconverter 98 is performed in the same manner as that performed by the reconverter 72. Substantially, the reconverters 72 and 98 can commonly have hardware, software routines, or programmable hardware parts.
[0035] The LPC information 104 included in the information 28 of the current LPD frame 16b can indicate the LPC coefficients for one point in the time segment 16b or for several points in the time segment 16b, such as a set of LPC coefficients for each of the subparts 92a to 92c. The spectral weighting filter extractor 94 converts the LPC coefficients to spectral weight coefficients that spectrally weight the transform coefficients in the information 90a by a transfer function derived from the LPC coefficients such that it is substantially close to an LPC synthesis filter or a modified version thereof. Any inverse quantization performed beyond the spectral weighting by the weightor 96 may not vary spectrally. Thus, different from the FD decoding mode, the quantization noise in the TCX coding mode is spectrally shaped using LPC analysis.
[0036] However, due to the use of reconversion, the reconverted signal segment 108 is subject to aliasing. However, by using the same reconversion, the reconverted signal segments 78 and 108 of consecutive frames and subframes can have their aliasing canceled by the transition handler 60 simply by adding their overlapping portions together.
[0037] (A) When processing the CELP subframe 90b, the excitation signal extractor 100 derives an excitation signal from the excitation update information in each subframe 90b, and the LPC synthesis filter 102 performs LPC synthesis filtering of the excitation signal using the LPC information 104 to obtain an LP synthesized signal segment 110 for the subportion 92b of the current time segment 16b.
[0038] The extractors 94 and 100 can be configured to perform some interpolation to adapt the LPC information 104 in the current frame 16b to the varying positions of the current subframe corresponding to the current subportion in the current time segment 16b.
[0039] Referring generally to FIGS. 3 - 5, the various signal segments 108, 110, and 78 then enter a transition handler 60 that assembles all the signal segments in the correct time order. In particular, the transition handler 60 performs time-domain aliasing cancellation within the range of the window portion that temporally overlaps at the boundaries between directly consecutive time segments of FD frames and TCX subframes to reconstruct the information signal across these boundaries. Thus, forward aliasing cancellation data for each of the boundaries between consecutive FD frames, between an FD frame followed by a TCX frame and an FD frame followed by a TCX subframe, is not necessary.
[0040] However, the situation changes whenever an FD frame or a TCX subframe (both indicating a variant of the transform coding mode) precedes an ACELP subframe (indicating a form of the time domain coding mode). In that case, the transition handler 16 derives a forward aliasing cancellation composite signal from the forward aliasing cancellation data of the current frame and adds the first forward aliasing cancellation composite signal to the reconstructed signal segment 100 or 78 of the immediately preceding time segment to reconstruct the information signal across each boundary. Since the boundaries between the TCX subframe and the ACELP subframe within the current frame define subportions of the associated time segment, when that boundary is partitioned into the current time segment 16b, the transition handler can verify that there is forward aliasing cancellation data for each of these transitions from the first syntax portion 24 and the subframing structure defined therein. The syntax portion 26 is not necessary. The previous frame 14a may or may not be gone.
[0041] However, when that boundary coincides with the boundary between consecutive time segments 16a and 16b, the parser 20 must examine the second syntax portion 26 within the current frame to determine whether the current frame 14b has forward aliasing cancellation data 34, and the FAC data 34 is for canceling the aliasing occurring at the front end of the current time segment 16b. This is because the previous frame is an FD frame or the last subframe of the previous LPD frame is a TCX subframe. At least, the parser 20 needs to know the syntax portion 26 in case the content of the previous frame is gone.
[0042] A similar description applies to other directions, i.e., the transition from an ACELP subframe to an FD frame or a TCX frame. As long as each boundary between each segment and sub - parts of the segment is partitioned inside the current time segment, the parser 20 has no problem determining that there is forward aliasing cancellation data 34 for these transitions from the current frame 14b itself, i.e., from the first syntax part 24. The second syntax part is not necessary and is even irrelevant. However, if the boundary occurs or coincides with the boundary between the previous time segment 16a and the current time segment 16b, the parser 20 needs to examine the second syntax part 26 to determine whether the forward aliasing cancellation data 34 exists at the front end of the current time segment 16b for the transition, at least when it cannot access the previous frame.
[0043] In the case of a transition from ACELP to FD or TCX, the transition handler 60 derives a second forward aliasing cancellation synthesis signal from the forward aliasing cancellation data 34 and adds the second forward aliasing cancellation synthesis signal to the re - transformed signal segment in the current time segment to reconstruct the information signal across the boundary.
[0044] Generally, after describing the embodiments related to FIGS. 3 - 5 related to embodiments where frames and sub - frames of different coding modes exist, specific embodiments of these embodiments will be outlined in more detail below. The description of these embodiments simultaneously includes possible means for generating each data stream containing this kind of frame and sub - frame. Below, although the principles outlined therein are also convertible to other signals, this specific embodiment is described as a unified speech and audio codec (USAC).
[0045] There are several purposes for the window switching in USAC. It mixes the FD frame, i.e., the frame encoded by frequency coding, and then the LPD frame constructed by the ACELP (sub) frame and the TCX (sub) frame. The ACELP frame (time domain coding) applies a window function processing that is rectangular and without overlap to the input samples, while the TCX frame (frequency domain coding) applies a non-rectangular and overlapping window function processing to the input samples and then encodes the signal using, for example, time domain aliasing cancellation (TDAC) transform, i.e., MDCT. To harmonize the overall window, the TCX frame can use a window with a centered uniform shape, and to manage the transition at the boundary of the ACELP frame, explicit information for eliminating time domain aliasing and the window function processing effect of the harmonized TCX window are transmitted. This additional information can be regarded as forward aliasing cancellation (FAC). The FAC data is quantized in the following embodiments in the LPC weighted region so that the FAC and the quantization noise of the decoded MDCT have the same nature.
[0046] FIG. 6 shows the processing in the encoder of the frame 120 encoded by transform coding (TC) that precedes and follows the frames 122, 124 encoded by ACELP. In accordance with the above description, the concept of TC includes MDCT over long and short blocks using AAC, as well as MDCT-based TCX. That is, the frame 120 may be an FD frame or a TCX (sub) frame, for example, as the sub-frames 90a, 92a in FIG. 5. FIG. 6 shows the time domain markers and the frame boundaries. The frame or time segment boundaries are indicated by dotted lines, while the time domain markers are short vertical lines along the horizontal axis. In the following description, it must be stated that the terms "time segment" and "frame" are sometimes used synonymously for their unique relevance there.
[0047] Thus, the vertical dotted lines in FIG. 6 indicate the beginning and end of a sub - part of a sub - frame / time segment or a frame / time segment that can be Frame 120. LPC1 and LPC2 indicate the centers of the analysis windows corresponding to the LPC filter coefficients or the LPC filter used to perform aliasing cancellation hereinafter. These filter coefficients are derived by the decoder by the reconstructor 22 or the extractors 90 and 100, for example, by using interpolation with LPC information 104 (see FIG. 5). The LPC filter includes LPC1 corresponding to its calculation at the beginning of Frame 120 and LPC2 corresponding to its calculation at the end of Frame 120. Frame 122 is assumed to be encoded by ACELP. The same applies to Frame 124.
[0048] FIG. 6 is composed of four numbered lines on the right side of FIG. 6. Each line indicates a step in the processing in the encoder. It should be understood that each line is time - aligned with the above - mentioned line.
[0049] Line 1 in FIG. 6 shows the original audio signal segmented in Frames 122, 120, and 124 as described above. Therefore, to the left of the marker "LPC1", the original signal is encoded by ACELP. Between the markers "LPC1" and "LPC2", the original signal is encoded using TC. As described above, in TC, noise shaping is applied directly in the transform domain rather than in the time domain. To the right of the marker LPC2, the original signal is encoded again by ACELP, i.e., by the time - domain coding mode. This sequence of coding modes (ACELP, then TC, then ACELP) is selected to show the processing of the FAC since the FAC is involved in both transitions (from ACELP to TC and from TC to ACELP).
[0050] However, it should be noted that the transitions at LPC1 and LPC2 in FIG. 6 can occur within the range inside the current time segment or at a point where they can coincide at its leading edge. In the former case, the determination of the presence of the relevant FAC data can be performed by the parser 20 based only on the first syntax portion 24, but in the case of frame disappearance, in the latter case, the parser 20 may require the syntax portion 26 to do so.
[0051] Line 2 of FIG. 6 corresponds to the decoded (synthesized) signals in each of frames 122, 120, and 124. Thus, reference numeral 110 in FIG. 5 is used within frame 122 to correspond to the possibility that the last sub - portion of frame 122 is an ACELP - encoded sub - portion as in 92b of FIG. 5. On the other hand, the combination of reference numerals 108 / 78 is used to indicate the signal payload portion for frame 120, similar to FIGS. 5 and 4. Further still, to the left of marker LPC1, the synthesis of that frame 122 is assumed to be encoded by ACELP. Therefore, the synthesized signal 110 to the left of marker LPC1 is identified as an ACELP - synthesized signal. Generally speaking, since ACELP processes the waveform encoding as accurately as possible, there is a high similarity between the ACELP synthesis and the original signal in that frame 122. Then, as seen by the decoder, the segment between markers LPC1 and LPC2 on line 2 of FIG. 6 indicates the output of the inverse MDCT of that segment 120. Further still, segment 120 may be a sub - portion of a TCX - encoded sub - frame such as, for example, time segment 16b of the FD frame or 90b of FIG. 5. In that figure, this segment 108 / 78 is called the "TC frame output". In FIGS. 4 and 5, this segment was called the re - transformed signal segment. When frame / segment 120 is a TCX segment sub - portion, the TC frame output indicates the TLP - synthesized signal that has been windowed again. Here, TLP represents "transform coding using linear prediction" in order to indicate that in the case of TCX, the noise shaping of each segment is accomplished in the time domain by filtering the MDCT coefficients that use the spectral information from LPC filters LPC1 and LPC2 respectively. It was also described above in relation to FIG. 5 regarding spectral weight - ing device 96. Also note that the synthesized signal, i.e., the preliminarily reconstructed signal including aliasing, i.e., signal 108 / 78, between markers "LPC1" and "LPC2" on line 2 of FIG. 6 includes windowing effects and time - domain aliasing at its beginning and end.In the case of MDCT as TDAC conversion, time-domain aliasing can be represented as expansions 126a and 126b, respectively. In other words, it extends from the beginning to the end of segment 120, and the upper curve of line 2 in FIG. 6 indicated by reference numerals 108 / 78 shows the windowing effect by a conversion window function processing that is flat in the middle but not at the beginning and end in order to leave the converted signal as it is. The folding effect is shown by lower curves 126a and 126b at the beginning and end of segment 120 having a minus sign at the beginning of the segment and a plus sign at the end of the segment. This windowing processing and time-domain aliasing (or folding) effect are inherent to MDCT functioning as an explicit example for TDAC conversion. As described above, aliasing can be eliminated when two consecutive frames are encoded using MDCT. However, when the "MDCT-encoded" frame 120 does not precede and / or follow other MDCT frames, its windowing processing and time-domain aliasing are not eliminated and remain in the time-domain signal after the inverse MDCT. As described above, forward aliasing cancellation (FAC) can then be used to correct these effects. Finally, segment 124 after marker LPC2 in FIG. 6 is also assumed to be encoded using ACELP. In order to obtain the synthesized signal in that frame, the filter state of LPC filter 102 (see FIG. 5) at the beginning of frame 124, i.e., the memory of the long-term and short-term predictors, must automatically be appropriate. It means that the time-domain aliasing and windowing effects at the end of the previous frame 120 between markers LPC1 and LPC2 must be eliminated by the application of FAC in a specific way described below. In summary, line 2 in FIG. 6 includes the synthesis of a signal preliminarily reconstructed from consecutive frames 122, 120, and 124, including the windowing effect on the time-domain aliasing of the output of the inverse MDCT for the frame between markers LPC1 and LPC2.
[0052] To obtain line 3 of FIG. 6, the difference between line 1 of FIG. 6, i.e., the original audio signal 18, and line 2 of FIG. 6, i.e., the difference between the composite signals 110 and 108 / 78, are each calculated as described above. This results in a first difference signal 128.
[0053] Further processing on the encoder side for frame 120 is described below with respect to line 3 of FIG. 6. At the beginning of frame 120, first, two contributions taken from the ACELP synthesis 110 to the left of marker LPC1 on line 2 of FIG. 6 are added to each as follows.
[0054] The first contribution 130 is the windowed and time-reversed (folded) version of the last ACELP synthesis sample, i.e., the last sample of the signal segment 110 shown in FIG. 5. The window length and shape for this time-reversed signal are the same as the aliasing part of the transform window to the left of frame 120. This contribution 130 can be regarded as a better approximation of the time-domain aliasing present in the MDCT frame 120 of line 2 of FIG. 6.
[0055] The second contribution 132 is the windowed zero-input response (ZIR) of the LPC1 synthesis filter with respect to the first state that was the final state of this filter at the end of the ACELP synthesis 110, i.e., at the end of frame 122. The window length and shape of this second contribution may be the same as those for the first contribution 130.
[0056] Regarding the new line 3 in FIG. 6, i.e., after adding the two contributions 130 and 132, a new difference is taken by the encoder in order to obtain line 4 in FIG. 6. Note that the difference signal 134 stops at marker LPC2. A plot of the approximation of the predicted envelope of the time-domain error signal is shown in line 4 of FIG. 6. The error in the ACELP frame 122 is predicted to be approximately flat in the time-domain amplitude. Then, the error in the TC frame 120 is predicted to exhibit a general shape, i.e., the time-domain envelope, as shown in this segment 120 of line 4 in FIG. 6. This predicted shape of the error amplitude is shown here only for the sake of explanation.
[0057] Note that if only the composite signal on line 3 in FIG. 6 is used for the decoder to generate or reconstruct the decoded audio signal, quantization noise will generally be present as the predicted envelope of the error signal 136 in line 4 of FIG. 6. Thus, it is understood that corrections must be sent to the decoder to compensate for this error at the beginning and end of the TC frame 120. This error results from the window function processing and time-domain aliasing effects inherent in the MDCT / inverse MDCT pair. The window function processing and time-domain aliasing were reduced at the beginning of the TC frame 120 by adding the cylindrical contributions 132 and 130 from the previous ACELP frame 122 as described above, but like the actual TDAC operations of successive MDCT frames, cannot be completely eliminated. Right at the TC frame 120 of line 4 in FIG. 6 before marker LPC2, all the window function processing and time-domain aliasing remain from the MDCT / inverse MDCT pair and thus must be completely eliminated by forward aliasing cancellation.
[0058] Before proceeding with the description of the encoding process for obtaining forward aliasing cancellation data, refer to FIG. 7 to briefly describe MDCT as an example of TDAC conversion processing. Both conversion directions are represented and described with respect to FIG. 7. The transition from the time domain to the conversion domain is shown in the upper half of FIG. 7, while the reconversion is represented in the lower part of FIG. 7.
[0059] When transitioning from the time domain to the conversion domain, the TDAC conversion is related to the window function processing 150 applied to the interval 152 of the signal to be converted, where the resulting conversion coefficients spread beyond the time segment 154 that is actually transmitted in the data stream. The window applied in the window function processing 150 is the non-aliasing portion M that spreads therebetween k together with the aliasing portion L that intersects the front end of the time segment 154 k and the aliasing portion R at the rear end of the time segment 154 k as shown in FIG. 7. The MDCT 156 is applied to the window function processed signal. That is, the folding 158 is performed to bend the first quarter of the interval 152 that extends between the front end of the interval 152 and the front end of the time segment 154 along the left (front) boundary of the time segment 154. The same is done for the aliasing portion R k . Thereafter, the DCT IV 160 is performed on the window function processed and folded signal resulting in a result having the same number of samples as the time signal 154 to obtain the same number of conversion coefficients. The conversation is then performed at 162. Of course, quantization 162 can also be regarded as not being included by the TDAC conversion.
[0060] The reconversion performs an inversion. That is, after the inverse quantization 164, first, the IMDCT 166 is used to obtain the same number of time samples as the number of samples of the time segment 154 to be reconstructed, for the DCT -1It is executed in relation to IV 168. Thereafter, the expansion process 168 is executed on the inverse-transformed signal portion received from module 168, which thereby doubles the length of the aliasing portion and extends over the time interval or number of time samples of the IMDCT result. Then, the window function processing is executed at 170 using a reconversion window 172, which may be the same as that used by the window function processing 150 but may also be different. The remaining blocks in FIG. 7 represent the TDAC or overlap / add processing executed on the overlapping portions of the successive segments 154, i.e., the addition of the expanded aliasing portions as executed by the transition handler in FIG. 3. As shown in FIG. 7, the TDAC by blocks 172 and 174 results in aliasing cancellation.
[0061] Now, proceed further with the description of FIG. 6. Assume that the TC frame 120 in FIG. 6 uses frequency-domain noise shaping (FDNS) in order to efficiently compensate for the window function processing and the time-domain aliasing effect at the beginning and end of the TC frame 120 in line 4 of FIG. 6. Then, the forward aliasing correction (FAC) is applied after the processing described in FIG. 8. First, note that FIG. 8 shows this processing with respect to both the left portion of the TC frame 120 near the marker LPC1 and the right portion of the TC frame 120 near the marker LPC2. Recall that the TC frame 120 in FIG. 6 was assumed to be preceded by the ACELP frame 122 at the LPC1 marker boundary and followed by the ACELP frame 124 at the LPC2 marker boundary.
[0062] In order to perform window function processing and compensate for the time-domain aliasing effect near marker LPC1, the processing is described in FIG. 8. First, a weighting filter W(z) is calculated from the LPC1 filter. The weighting filter W(z) may be the modified analysis or whitening filter A(z) of LPC1. For example, W(z)=A(z / λ), where λ is a predetermined weighting factor. The error signal at the beginning of the TC frame is denoted by reference numeral 138, just as in the case of line 4 of FIG. 6. This error is called the FAC target in FIG. 8. The error signal 138 is filtered by the filter W(z) at 140 using the initial state of this filter, i.e., the initial state when it is the filter memory of the ACELP error 141 in the ACELP frame 122 of line 4 of FIG. 6. The output of the filter W(z) then forms the input to the transform 142 of FIG. 6. The transform is shown as being, for example, an MDCT. The transform coefficients output by the MDCT are then quantized and encoded in the processing module 143. These encoded coefficients may form at least a part of the aforementioned FAC data 34. These encoded coefficients may be transmitted to the encoding side. The output of processing Q, i.e., the quantized MDCT coefficients, is then the input to an inverse transform such as IMDCT144 to form a time-domain signal filtered by the inverse filter 1 / W(z) at 145 with zero memory (zero initial state). The filtering by 1 / W(z) extends beyond the length of the FAC target using zero input for samples beyond the FAC target. The output of the filter 1 / W(z) is the FAC synthesis signal 146, which is a correction signal that can be applied here at the beginning of the TC frame 120 to compensate for the window function processing and the time-domain aliasing effect occurring therein.
[0063] Next, the processing for window function processing and time-domain aliasing correction at the end of the TC frame 120 (before marker LPC2) will be described. To achieve this purpose, refer to FIG. 9.
[0064] The error signal after the end of the TC frame 120 on line 4 of FIG. 6 is supplied with reference numeral 147 and indicates the FAC target of FIG. 9. The FAC target 147 follows the same processing sequence as the FAC target 138 of FIG. 8, using only processing that differs in the initial state of the weighting filter W(z) 140. The initial state of the filter 140 for filtering the FAC target 147 is the error in the TC frame 120 on line 4 of FIG. 6, indicated by reference numeral 148 in FIG. 6. Next, the further processing steps 142-145 are the same as in FIG. 8 related to the processing of the FAC target at the beginning of the TC frame 120.
[0065] To obtain local FAC synthesis and check whether the change in the coding mode involved by selecting the TC coding mode of frame 120 is an optimal choice, when applied in the encoder to calculate the resulting reconstruction, the processing in FIGS. 8 and 9 is executed completely from left to right. In the decoder, the processing in FIGS. 8 and 9 is only applied from the middle to the right. That is, the encoded and quantized transform coefficients transmitted by the processing device Q 143 are decoded to form the input of the IMDCT. See, for example, FIGS. 10 and 11. FIG. 10 is equal to the right side of FIG. 8, while FIG. 11 is equal to the right side of FIG. 9. The transition handler 60 of FIG. 3 can be realized according to the specific embodiment outlined here, according to FIGS. 10 and 11. That is, the transition handler 60 can re-convert the transform coefficient information in the FAC data 34 in the current frame 14b to generate the first FAC synthesis signal 146 in the case of a transition from the ACELP time segment sub-part to the FD time segment or the TCX sub-part, or the second FAC synthesis signal 149 when transitioning from the FD time segment or the TCX sub-part of the time segment to the ACELP time segment sub-part.
[0066] Furthermore, if the FAC data 34 is such that the parser 20 simply derives the existence of the FAC data 34 from the syntax portion 24, it can be associated with this type of transition occurring within the current time segment. On the other hand, it should be noted that the parser 20 needs to utilize the syntax portion 26 to determine whether the FAC data 34 exists for this type of transition at the leading edge of the current time segment 16b when there is no previous frame.
[0067] FIG. 12 shows how a complete synthesized or reconstructed signal for the current frame 120 can be obtained by using the FAC synthesis signals of FIGS. 8 to 11 and applying the reverse steps of FIG. 6. Further, it should be noted that the steps shown in FIG. 12 here are also executed by the encoder to check whether the encoding mode for the current frame leads to the best optimization, for example, in terms of rate / distortion points. In FIG. 12, it is assumed that the ACELP frame 122 to the left of the marker LPC1 has already been synthesized or reconstructed by the module 58 of FIG. 3 or the like and leads to the ACELP synthesis signal of line 2 of FIG. 12 having the reference numeral 110 thereto up to the marker LPC1. Since the FAC correction is also used at the end of the TC frame, it is also assumed that the frame 124 after the marker LPC2 is an ACELP frame. Next, the following steps are executed to generate the signal synthesized or reconstructed in the TC frame 120 between the markers LPC1 and LPC2 of FIG. 12. These steps are also shown in FIGS. 13 and 14. FIG. 13 shows the steps executed by the transition handler 60 to process the transition from the TC-encoded segment or segment sub-part to the ACELP-encoded segment sub-part, while FIG. 14 shows the operation of the transition handler for the reverse transition.
[0068] 1. One step is to decode the MDCT - encoded TC frame and position the thus - obtained time - domain signal between markers LPC1 and LPC2 as shown in line 2 of FIG. 12. The decoding is performed by module 54 or module 56 such that the decoded TC frame includes window function processing and time - domain aliasing effects and includes an inverse MDCT, for example, for TDAC reconversion. In other words, the currently decoded segment or time - segment sub - portion indicated by index k in FIGS. 13 and 14 can be the time - segment 16b which is an ACELP - encoded time - segment sub - portion 92b as shown in FIG. 13, or an FD - encoded or TCX - encoded sub - portion 92a as shown in FIG. 14. In the case of FIG. 13, the previously processed frame is thus a TC - encoded segment or time - segment sub - portion, and in the case of FIG. 14, the previously processed time - segment is an ACELP - encoded sub - portion. The reconstructed or synthesized signal as an output by modules 54 - 58 is partially subject to aliasing effects. This also applies to signal segments 78 / 108.
[0069] 2. Another step in the processing of the transition handler 60 is the generation of the FAC composite signals according to FIG. 10 in the case of FIG. 14 and according to FIG. 11 in the case of FIG. 13. That is, the transition handler 60 can perform the reconversion 191 on the conversion coefficients in the FAC data 34 respectively to obtain the FAC composite signals 146 and 149. The FAC composite signals 146 and 149 are then subject to the aliasing effect and are positioned at the beginning and end of the TC-encoded segments shown in the time segments 78 / 108. In the case of FIG. 13, for example, as also shown in line 1 of FIG. 12, the transition handler 60 positions the FAC composite signal 149 at the end of the TC-encoded frame k-1. In the case of FIG. 14, as also shown in line 1 of FIG. 12, the transition handler 60 positions the FAC composite signal 146 at the beginning of the TC-encoded frame k. Furthermore, it should be noted that the frame k is the currently decoded frame and the frame k-1 is the previously decoded frame.
[0070] 3. As far as the situation of FIG. 14 where the coding mode change occurs at the beginning of the current TC frame k is concerned, from the ACELP frame k-1 before the TC frame k, the windowed and folded (inverted) ACELP composite signal 130, and the windowed zero input response of the LPC1 synthesis filter, or ZIR, that is, the signal 132, are positioned so as to be aligned with the reconverted signal segments 78 / 108 that are subject to aliasing. This contribution is shown in line 3 of FIG. 12. As shown in FIG. 14 and as already explained above, the transition handler 60 continues the LPC synthesis filtering of the previous CELP subframe across the boundary line at the beginning of the current time segment k and obtains the anti-aliasing signal 132 by windowing the continuity of the signals in the current signal k in both steps indicated by the reference numerals 190 and 192 in FIG. 14. To obtain the anti-aliasing signal 130, the transition handler 60 also window-processes the reconstructed signal segment 110 of the previous CELP frame in step 194 and uses this windowed and time-reversed signal as the signal 130.
[0071] 4. The contributions of lines 1, 2, and 3 in FIG. 12, and the contributions 78 / 108, 132, 130, and 146 in FIG. 14, and the contributions 78 / 108, 149, and 196 in FIG. 13 are added together by the transition handler 60 at the aligned positions described above, and as shown in line 4 of FIG. 12, a synthesized or reconstructed audio signal for the current frame k is formed in the original region. It should be noted that the processing in FIGS. 13 and 14 is such that the time-domain aliasing and windowing effects are eliminated at the beginning and end of the frame, and the potential discontinuity at the frame boundary near the marker LPC1 is smoothed and perceptually masked by the filter 1 / W(z) in FIG. 12 to generate a synthesized or reconstructed signal 198 in the TC frame.
[0072] Thus, FIG. 13 leads to forward aliasing cancellation at the end of the previous TC-encoded segment in relation to the current processing of the CELP-encoded frame k. As indicated by 196, the last reconstructed audio signal is the aliasing that is not reconstructed across the boundary between segment k-1 and segment k. The processing in FIG. 14 leads to forward aliasing cancellation at the beginning of the current TC-encoded segment k, as indicated by the reference numeral 198 showing the signal reconstructed across the boundary between segment k and segment k-1. The aliasing remaining at the trailing end of the current segment k is eliminated by the TDAC if the subsequent segment is a TC-encoded segment, or by the FAC according to FIG. 13 if the subsequent segment is an ACELP-encoded segment. FIG. 13 refers to this latter possibility by assigning the reference numeral 198 to the signal segment of time segment k-1.
[0073] The following refers to how the second syntax portion 26 can be implemented for a particular possibility.
[0074] For example, in order to handle the occurrence of lost frames, the syntax part 26 can be embodied as a 2-bit field prev_mode that explicitly sends, within the current frame 14b, a signal of the coding mode applied in the previous frame 14a according to the following table. JPEG0007693891000001.jpg45120
[0075] In other words, this 2-bit field may be called prev_mode and can thus indicate the coding mode of the previous frame 14a. That is, in the case of the example just mentioned, four different states are distinguished. 1) The previous frame 14a is an LPD frame and its last sub-frame is an ACELP sub-frame. 2) The previous frame 14a is an LPD frame and its last sub-frame is a TCX-coded sub-frame. 3) The previous frame is an FD frame using a long transform window, 4) The previous frame is an FD frame using a short transform window.
[0076] The potential use of different window lengths in the FD coding mode has already been mentioned above with respect to the description of FIG. 3. Of course, the syntax part 26 can have only three different states, and the FD coding mode can operate with a certain window length, thereby combining the last two options 3 and 4 of the options listed above.
[0077] In any case, based on the 2-bit field outlined above, the parser 20 can determine whether the FAC data for the transition between the current time segment and the previous time segment 16a is within the current frame 14a. As will be outlined in more detail below, the parser 20 and the reconstructor 22 can determine, based on prev_mode, whether the previous frame 14a was an FD frame using a long window (FD_long), or whether the previous frame was an FD frame using a short window (FD_short), and whether the current frame 14b (if the current frame is an LPD frame) inherits an FD frame or an LPD frame, which are differentiations that are necessary according to the following embodiments respectively to correctly analyze the data stream and reconstruct the information signal.
[0078] Thus, according to the aforementioned possibility of using a 2-bit identifier as the syntax part 26, each frame 16a - 16c is supplied with an additional 2-bit identifier in addition to the syntax part 24 that determines the coding mode of the current frame, which is a sub-framing structure in the case of the FD or LPD coding mode and the LPD coding mode.
[0079] Regarding all of the above embodiments, it must be stated that other internal frame dependencies need to be avoided as well. For example, the decoder of FIG. 1 would be SBR-capable. In that case, instead of analyzing this type of crossover frequency, which has an SBR header transmitted not very frequently in the data stream 12, all frames 16a - 16c from each SBR extension data can be analyzed by the parser 20. Other inter-frame dependencies can be removed as well.
[0080] Regarding all of the above embodiments, it is worth noting that the parser 20 is configured to buffer at least the currently decoded frame 14b in the buffer by passing all of the frames 14a to 14c through this buffer in a first in first out (FIFO) manner. In buffering, the parser 20 can perform frame removal from this buffer in units of frames 14a to 14c. That is, the filling and removal of the buffer of the parser 20 can be performed in units of frames 14a to 14c in order to comply with the constraints imposed by, for example, simply one, or a plurality of, frames of the maximum size at a time, the maximum available buffer space for receiving frames.
[0081] Another signaling possibility for syntax portion 26 with reduced bit consumption is described next. According to this variant, different structural configurations of syntax portion 26 are used. In the previously described embodiment, syntax portion 26 was a 2-bit field transmitted in all frames 14a - 14c of the encoded USAC data stream. Regarding the FD portion, since it is only important for the decoder to know whether it needs to read FAC data from the bitstream when the previous frame 14a is lost, these 2 bits can be divided into two 1-bit flags, one of which is signaled as fac_data_present in all frames 14a - 14c. This bit can be introduced into the single_channel_element and channel_pair_element structures accordingly, as shown in the tables of FIGS. 15 and 16. FIGS. 15 and 16 can be regarded as the upper structure definition of the syntax of frame 14 according to this embodiment. Here, the function "function_name(…)" calls a subroutine, and the syntax element names written in bold indicate reading each syntax element from the data stream. In other words, the marked or hatched portions of FIGS. 15 and 16 indicate that each frame 14a - 14c is supplied with the flag fac_data_present according to this embodiment. Reference numeral 199 indicates these portions.
[0082] The other 1-bit flag prev_frame_was_lpd is only transmitted in the current frame next if it is encoded using the LPD portion of USAC, indicating whether the previous frame was also encoded using the LPD path of USAC. This is shown in the table of FIG. 17.
[0083] The table of FIG. 17 shows a part of the information 28 of FIG. 1 when the current frame 14b is an LPD frame. As shown at 200, each LPD frame is supplied with a flag prev_frame_was_lpd. This information is used to parse the syntax of the current LPD frame. It can be derived from FIG. 18 that the content and position of the FAC data 34 of the LPD frame depend on the transition between the TCX coding mode and the CELP coding mode or the transition at the front end of the current LPD frame from the FD coding mode to the CELP coding mode. In particular, when the currently decoded frame 14b is an LPD frame immediately following the FD frame 14a and fac_data_present indicates that (since the leading subframe is an ACELP subframe) FAC data is present in the current LPD frame, the FAC data is then read at 202 at the end of the LPD frame syntax using the FAC data 34 that includes a gain coefficient fac_gain as shown at 204 in FIG. 18. For this gain coefficient, the contribution 149 in FIG. 13 adjusts the gain.
[0084] However, when the current frame is an LPD frame having a previous frame that is also an LPD frame, i.e., when the transition between the TCX and CELP subframes occurs between the current frame and the previous frame, the FAC data is read at 206 without a gain adjustment option, i.e., without the FAC data 34 that includes the FAC gain syntax element fac_gain. Further, when the current frame is an LPD frame and the previous frame is an FD frame, the position of the FAC data read at 206 is different from the position where the FAC data is read at 202. The read position 202 occurs at the end of the current LPD frame, while the reading of the FAC data at 206 occurs before reading the specific data of the subframe, i.e., the ACELP or TCX data that depends on the mode of the subframe of the subframe structure at 208 and 210, respectively.
[0085] In the example of FIGS. 15 to 18, the LPC information 104 (FIG. 5) is read after the data specific to sub-frames such as 212 and 90a and 90b (compare FIG. 5).
[0086] For completeness only, the syntax structure of the LPD frame according to FIG. 17 is further described with respect to potentially additionally included FAC data in the LPD frame in order to supply FAC information regarding the transition between the TCX and ACELP sub-frames within the current LPD-coded time segment. In particular, according to the embodiments of FIGS. 15 to 18, the LPD sub-frame structure is restricted to subdivide the current LPD-coded time segment in quarter units only by allocating one quarter of these to either TCX or ACELP. The exact LPD structure is determined by the syntax element lpd_mode read at 214. While the ACELP frame is restricted to a quarter length, the first, second, third, and fourth quarters can all form a TCX sub-frame. The TCX sub-frame also extends over the entire LPD-coded time segment, in which case the number of sub-frames is simply one. The while loop in FIG. 17 steps through the quarters of the currently LPD-coded time segment and whenever the current quarter k is at the start of a new sub-frame within the currently LPD-coded time segment, the sub-frame immediately preceding the currently starting / decoding LPD frame transmits the FAC data supplied at 216, which is in the other mode, i.e., TCX mode if the current sub-frame is in ACELP mode and ACELP mode if the current sub-frame is in TCX mode.
[0087] For completeness only, FIG. 19 shows a possible syntax structure of the FD frame according to the embodiments of FIGS. 15 to 18. It can be seen that the FAC data is read at the end of the FD frame by a determination as to whether there is FAC data 34 that is only involved in the fac_data_present flag. In comparison, the syntax analysis of fac_data34 in the case of the LPD frame as shown in FIG. 17 requires knowing about the flag prev_frame_was_lpd for correct syntax analysis.
[0088] Thus, the 1-bit flag prev_frame_was_lpd is only transmitted when the current frame is encoded using the LPD part of USAC and indicates whether the previous frame was encoded using the LPD path of the USAC codec (see the syntax of lpd_channel_stream() in FIG. 17).
[0089] Regarding the embodiments of FIGS. 15 to 19, further syntax elements can be transmitted at 220 so that the FAC data is read at 202 at the front end of the current LPD frame for the transition from the FD frame to the ACELP subframe, i.e., when the current frame is an LPD frame and the previous frame is an FD frame (by the first frame of the current LPD frame which is an ACELP frame). This additional syntax element read at 220 can indicate whether the previous FD frame 14a is FD_long or FD_short. Depending on this syntax element, the FAC data 202 can be affected. For example, the length of the synthesized signal 149 can be affected according to the length of the window used to transform the previous LPD frame. Summarizing the embodiments of FIGS. 15 and 19 and applying the features mentioned therein to the embodiments described with respect to FIGS. 1 to 14, the following can be applied to the latter embodiments, individually or in combination.
[0090] 1) The FAC data 34 referred to in the previous figure enables forward aliasing cancellation that occurs at the transition between the previous frame 14a and the current frame 14b, i.e., between the corresponding time segments 16a and 16b. This mainly means that FAC data exists in the current frame 14b. However, additional FAC data may be present. However, this additional FAC data handles the transition between the TCX-encoded subframe and the CELP-encoded subframe located internally in the current frame 14b when it is in the LPD mode. The presence or absence of this additional FAC data is independent of the syntax part 26. In FIG. 17, this additional FAC data is read at 216. Its presence or absence depends only on the lpd_mode read at 214. The latter syntax element is part of the syntax part 24 that reveals the encoding mode of the current frame instead. Together with the core_mode read at 230 and 232 shown in FIGS. 15 and 16, lpd_mode corresponds to the syntax part 24.
[0091] 2) Further, the syntax part 26 can consist of one or more syntax elements as described above. The flag FAC_data_present indicates whether there is fac_data for the boundary between the previous frame and the current frame. This flag exists in the LPD frame as well as in the FD frame. In the said embodiment, a further flag called prev_frame_was_lpd is transmitted to the LPD frame only for indicating whether the previous frame 14a was in the LPD mode. In other words, this second flag included in the syntax part 26 indicates whether the previous frame 14a was an FD frame. The parser 20 predicts and reads this flag only when the current frame is an LPD frame. In FIG. 17, this flag is read at 200. Depending on this flag, the parser 20 can predict that the FAC data contains the gain value fac_gain and thus can read it from the current frame. The gain value is used by the reconstructor to set the gain of the FAC synthesis signal for the FAC at the transition between the current and the previous time segments. In the embodiments of FIGS. 15 to 19, this syntax element is read at 204 depending on the dependence on a second flag that is obvious from comparing the situations leading to the reads 206 and 202 respectively. Alternatively, or in addition, prev_frame_was_lpd can control the position where the parser 20 predicts and reads the FAC data. In the embodiments of FIGS. 15 to 19, these positions were 206 or 202. Further, the second syntax part 26 can further include a further flag when the current frame is an LPD frame, the subframe at its head is an ACELP frame, and the previous frame is an FD frame, for indicating whether the previous FD frame used a long conversion window or a short conversion window for encoding. This latter flag can be read at 220 in the case of the foregoing embodiments of FIGS. 15 to 19. This knowledge about the FD conversion length can be used respectively to determine the length of the FAC synthesis signal and the size of the FAC data 38.With this method, the FAC data can be sized to fit the overlap length of the window of the previous FD frame so that a better compromise between coding quality and coding rate can be achieved.
[0092] 3) By dividing the second syntax portion 26 into the three flags described above, when the current frame is an FD frame, only one flag or bit indicating that it is the second syntax portion 26 needs to be transmitted. When the current frame is an LPD frame and the previous frame is also an LPD frame, only two flags or bits need to be transmitted. The third flag needs to be transmitted only in the case of a transition from an FD frame to the current LPD frame. Alternatively, as described above, the second syntax portion 26 is transmitted for each frame and indicates the mode of the frame preceding this frame within the range required by the parser to determine whether the FAC data 38 needs to be read from the current frame, and if so, where the FAC composite signal is from and how long it is. That is, the specific embodiments of FIGS. 15 to 19 can be easily diverted to embodiments using the 2-bit identifier to execute the second syntax portion 26. Instead of FAC_data_present in FIGS. 15 and 16, a 2-bit identifier is transmitted. The flags 200 and 220 do not need to be transmitted. Instead, the content of fac_data_present in the if statement connected to 206 and 218 can be derived from the 2-bit identifier by the parser 20. The following table can be accessed by the decoder to utilize the 2-bit identifier. JPEG0007693891000002.jpg68120
[0093] When the FD frame uses only one possible length, the syntax portion 26 can simply have only three different possible values.
[0094] As the embodiments of FIGS. 20 to 22 are referred to for the description of the embodiments, a syntax structure that is slightly different but very similar to that described above with respect to FIGS. 15 to 19 is shown in FIGS. 20 to 22 using the same reference numerals as those used with respect to FIGS. 15 to 19.
[0095] It should be noted that for the embodiments described with respect to FIG. 3 and the like, any transform coding scheme having aliasing fitness other than MDCT can be used in association with the TCX frame. Furthermore, transform coding schemes such as FFT can also be used without aliasing in the LPD mode, that is, without FAC for sub-frame transitions within the LPD frame, and thus without the need to transmit FAC data for sub-frame boundaries between LPD boundaries. The FAC data is then included only for any transitions from FD to LPD and vice versa.
[0096] Regarding the embodiments described below with respect to FIG. 1, note that they are directed to cases where, arranged such that an additional syntax portion 26 is defined in the first syntax portion of the previous frame, i.e., uniquely depending on the comparison between the coding mode of the current frame and the coding mode of the previous frame, as a result, in all of the foregoing embodiments, a decoder or parser can uniquely predict the content of the second syntax portion of the current frame by using or comparing these frames, i.e., the first syntax portions of the previous frame and the current frame. That is, in the absence of frame loss, the decoder or parser was able to derive from the transition between frames whether the FAC data is present in the current frame. When a frame is lost, a second syntax portion such as the flag fac_data_present bit explicitly conveys that information. However, according to other embodiments, the encoder uses this explicit signalability provided by the second syntax portion 26 such that the syntax portion 26 optimally, i.e., by the decisions made there, which are typically executed on a per-frame basis, for example, from FD / TCX, i.e., the TC coding mode, to ACELP, i.e., the time-domain coding mode, or vice versa, applies the reverse coding set such that the transition between the current frame and the previous frame indicates the absence of FAC in the syntax portion of the current frame. The decoder is then implemented to operate strictly according to the syntax portion 26, thereby effectively disabling or suppressing the FAC data transmission in an encoder indicating this stop, simply by setting, for example, fac_data_present = 0. A scenario where this would be a preferred option is in very low bitrate coding where the resulting aliasing artifacts are acceptable compared to the overall audio quality, while the additional FAC data may be too costly in terms of bits.
[0097] Although several aspects have been described in relation to an apparatus, it will be apparent that these aspects also indicate a description of a corresponding method. Here, a block or device corresponds to a method step or the function of a method step. Similarly, aspects described in relation to a method step also indicate a description of a corresponding block or item, or the function of a corresponding apparatus. Some or all of the method steps can be performed (or by using) by a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.
[0098] The encoded audio signal of the present invention can be stored in a digital storage medium or transmitted to a transmission medium such as, for example, a wireless transmission medium or a wired transmission medium such as the Internet.
[0099] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or in software. The embodiment has an electronically readable control signal stored thereon that cooperates (or can cooperate) with a programmable computer system such that each method is performed, and can be implemented using a digital storage medium such as a floppy (registered trademark) disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory. Therefore, the digital storage medium can be computer-readable.
[0100] Some embodiments according to the present invention include a data carrier having an electronically readable control signal that can cooperate with a programmable computer system in which one of the methods described herein is performed.
[0101] Generally, embodiments of the present invention can be realized as a computer program product having program code. And when the computer program product operates on a computer, the program code functions to execute one of the methods. The program code can be stored, for example, in a machine-readable carrier.
[0102] Other embodiments include a computer program for executing one of the methods described in the present specification and stored in a machine-readable carrier.
[0103] Therefore, in other words, an embodiment of the method of the present invention is a computer program having program code for executing one of the methods described in the present specification when the computer program operates on a computer.
[0104] Therefore, a further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) that contains a computer program for executing one of the methods described in the present specification and on which the computer program is recorded. The data carrier, digital storage medium or recording medium is generally tangible and / or non-transitory.
[0105] Therefore, a further embodiment of the method of the present invention is a sequence of data streams or signals that represents a computer program for executing one of the methods described in the present specification. The sequence of data streams or signals can be configured to be transferred, for example, via a data communication connection, such as via the Internet.
[0106] Further embodiments include processing means, such as a computer or programmable logic device, configured or adapted to execute one of the methods described in the present specification.
[0107] A further embodiment includes a computer having installed thereon a computer program for performing one of the methods described herein in the present specification.
[0108] A further embodiment according to the present invention includes an apparatus or system configured to transfer (e.g., electronically or optically) to a receiver a computer program for performing one of the methods described herein in the present specification. The receiver can be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system can include, for example, a file server for transferring the computer program to the receiver.
[0109] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein in the present specification.
[0110] In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein in the present specification. Generally, the method is preferably performed by no hardware device.
[0111] The above-described embodiments are merely illustrative of the principles of the present invention. Modifications and details of the apparatus described herein in the present specification will be apparent to other persons skilled in the art. Accordingly, it is intended to be limited only by the scope of the impending claims and not by the specific details shown as the description and illustration of the embodiments herein in the present specification.
Claims
1. A decoder (10) for decoding a data stream (12) each comprising a sequence of frames in which a time segment of an information signal (18) is encoded, comprising: a parser (20) configured to parse the data stream (12), the parser being configured to read a first syntactic portion (24) and a second syntactic portion from a current frame (14b) when parsing the data stream (12); a reconstructor (22) configured to reconstruct the current time segment (16b) of the information signal (18) based on information (28) obtained from the current frame by the analysis, by performing a first selection from among a time domain aliasing cancellation transform decoding mode and a time domain decoding mode based on the first syntax part (24) to obtain a first selected decoding mode, and by performing a reconstruction of the current time segment (16b) of the information signal (18) associated with the current frame (14b) using the first selected decoding mode; Including, The parser (20) is configured to perform a second selection between a first operation and a second operation in dependence on the second syntactic portion when parsing the data stream (12) to obtain a second selected operation, the parser and the reconstructor being configured to: if the first operation is the second selected operation, reading forward aliasing cancellation data (34) from the current frame (14b) and using the forward aliasing cancellation data (34) to perform forward aliasing cancellation at a boundary between the current time segment (16b) and a previous time segment (16a) of a previous frame (14a); configured to not retrieve forward aliasing cancellation data (34) from the current frame (14b) during said parsing of said data stream if said second operation is the second selected operation; the first syntax part and the second syntax part are included in each frame, the first syntax part (24) associating each of the frames from which it is read with a first frame type or a second frame type, and, if each of the frames is of the second frame type, associating subframes of a subdivision of each of the frames, which consists of several subframes, with a respective one of the first subframe type and the second subframe type; The reconstructor comprises: performing, for each frame of the first frame type, a spectrally variable dequantization (70) of transform coefficient information in each frame of the first frame type based on scale factor information in each frame of the first frame type, and a retransform on the dequantized transform coefficient information to obtain a retransformed signal segment (78) extending in time throughout and beyond the time segment associated with each frame of the first frame type; deriving a first forward aliasing cancellation composite signal from the forward aliasing cancellation data (34) and adding the first forward aliasing cancellation composite signal to the retransformed signal segment (78) in the previous time segment to reconstruct the information signal (18) across the boundary between the previous frame and the current frame (14a, 14b), if the previous frame is of the first frame type or the previous frame is of the second frame type with a last subframe of the first subframe type and the current frame (14b) is of the second frame type with a first subframe of the second subframe type. A decoder (10) configured as follows.
2. A method for decoding a data stream (12) comprising a sequence of frames, each frame being an encoded time segment of an information signal (18), the method comprising the steps of: a parsing step of parsing the data stream (12), the parsing step comprising the steps of reading a first syntactic portion (24) and a second syntactic portion from a current frame (14b); performing a first selection between a time domain aliasing cancellation transform decoding mode and a time domain decoding mode in dependence on the first syntax part (24) to obtain a first selected decoding mode, and performing a reconstruction of the current time segment (16b) of the information signal (18) associated with the current frame (14b) using the first selected decoding mode, thereby reconstructing the current time segment of the information signal (18) based on information obtained from the current frame (14b) by the analyzing step. Including, A second selection between the first operation and the second operation is performed upon analyzing the data stream (12), where: if the first operation is the second selected operation, forward aliasing cancellation data (34) is read from the current frame (14b); performing forward aliasing cancellation at a boundary between the current time segment (16b) and a previous time segment (16a) of a previous frame (14a) using the forward aliasing cancellation data (34); if the second operation is the second selected operation, parsing the data stream does not include retrieving forward aliasing cancellation data (34) from the current frame (14b); the first syntax part and the second syntax part are included in each frame, the first syntax part (24) associating each of the frames from which it is read with a first frame type or a second frame type, and, if each of the frames is of the second frame type, associating subframes of a subdivision of each of the frames, which consists of several subframes, with a respective one of the first subframe type and the second subframe type; The method comprises: performing, for each frame of the first frame type, a spectrally variable dequantization (70) of transform coefficient information in each frame of the first frame type based on scale factor information in each frame of the first frame type and a retransform on the dequantized transform coefficient information to obtain a retransformed signal segment (78) extending in time through and beyond the time segment associated with each frame of the first frame type; if the previous frame is of the first frame type or the previous frame is of the second frame type whose last subframe is of the first subframe type and the current frame (14b) is of the second frame type whose first subframe is of the second subframe type, deriving a first forward aliasing cancellation synthesis signal from the forward aliasing cancellation data (34); and applying the first forward aliasing cancellation synthesis signal to the retransformed signal segment (78) in the previous time segment to reconstruct the information signal (18) across the boundary between the previous frame and the current frame (14a, 14b); A method comprising:
3. 3. A computer program having a program code for performing the method according to claim 2, when the computer program runs on a computer.
Citation Information
Patent Citations
Forward time-domain aliasing cancellation with application in weighted or original signal domain
WO2010148516A1