Time-varying time-frequency tiling using non-uniform orthogonal filter banks based on mdct analysis / synthesis and tdar
By combining cascaded overlapped critical sampling transform and time domain aliasing reduction technology with MDCT and TDAR, the problem of insufficient flexibility of the TDAR filter bank when the signal characteristics change is solved, and time-varying time-frequency tiling and coding efficiency are improved.
Patent Information
- Application Number
- CN202080060582.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-28
- Filing Date
- 2020-08-25
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2040-08-25
AI Technical Summary
The existing TDAR filter bank lacks flexibility when the input signal characteristics change and cannot achieve time-varying adaptive time-frequency tiling, which makes window switching difficult.
By introducing cascaded overlapped critical sampling transform and time domain aliasing reduction technology, combining MDCT analysis and TDAR, using non-uniform orthogonal filter banks for time-varying time-frequency tiling, using STDAR parameters and MDCT length parameters for inter-frame switching, the equalization and aliasing reduction of time-frequency tiling are achieved.
The compactness of the impulse response and the improvement of the coding efficiency of the non-uniform filter bank when the input signal characteristics change are achieved, the data volume requirement is reduced, and the flexibility and quality of signal processing are improved.
Smart Images

Figure CN114503196B_ABST
Abstract
Description
Technical Field
[0001] Embodiments relate to audio processors / methods for processing an audio signal to obtain a subband representation of the audio signal. Further embodiments relate to audio processors / methods for processing a subband representation of an audio signal to obtain an audio signal. Some embodiments relate to time-varying time-frequency tiling using a non-uniform orthogonal filter bank based on MDCT (MDCT = Modified Discrete Cosine Transform) analysis / synthesis and TDAR (TDAR = Time Domain Aliasing Reduction). Background Art
[0002] It has been previously shown that it is possible to design non-uniform orthogonal filter banks using subband merging [1], [2], [3], and that by introducing a post-processing step called time-domain aliasing reduction (TDAR), compact impulse responses are possible [4]. Furthermore, the use of TDAR filter banks in audio coding has been shown to yield higher coding efficiency and / or improved perceptual quality than window switching [5].
[0003] However, a major drawback of TDAR is that it requires two adjacent frames to use the same time-frequency tiling. This limits the flexibility of the filter bank when time-varying adaptive time-frequency tiling is required, as TDAR must be temporarily disabled to switch from one tiling to another. This switching is typically required when the input signal characteristics change, i.e. when transients are encountered. In the uniform MDCT, this is achieved using window switching [6].
[0004] It is therefore an object of the present invention to improve the compactness of the impulse response of a non-uniform filter bank even when the input signal characteristics vary. Summary of the Invention
[0005] This object is achieved by the independent claims.
[0006] Advantageous embodiments are set forth in the dependent claims.
[0007] An embodiment provides an audio processor for processing an audio signal to obtain a subband representation of the audio signal. The audio processor includes a cascaded overlapped critical sampling transform stage configured to perform the cascaded overlapped critical sampling transform on at least two partially overlapping blocks of samples of the audio signal to obtain a subband sample set based on a first block of samples of the audio signal and to obtain a subband sample set based on a second block of samples of the audio signal. Furthermore, the audio processor comprises: a first time-frequency transform stage configured to, in a case where the subband sample set based on the first sample block represents a different region in the time-frequency plane than the subband sample set based on the second sample block [e.g., time-frequency plane representations of the first sample block and the second sample block], identify one or more subband sample sets from the subband sample set based on the first sample block, and identify one or more subband sample sets from the subband sample set based on the second sample block, the identified one or more subband sample sets combined representing the same region in the time-frequency plane, and perform time-frequency transform on the one or more subband sample sets identified from the subband sample set based on the first sample block and / or the one or more subband sample sets identified from the subband sample set based on the second sample block to obtain one or more time-frequency transformed subband samples, each of the time-frequency transformed subband samples representing the same region in the time-frequency plane as a corresponding one of the identified one or more subband samples or the time-frequency transformed versions of the one or more subband samples. Furthermore, the audio processor comprises a time-domain aliasing reduction stage configured to perform a weighted combination of two corresponding subband sample sets or time-frequency transformed versions thereof to obtain an aliasing-reduced subband representation of the audio signal (102), wherein one of the two corresponding subband sample sets or the time-frequency transformed version thereof is obtained based on a first sample block of the audio signal (102) and the other of the two corresponding subband sample sets or the time-frequency transformed version thereof is obtained based on a second sample block of the audio signal.
[0008] In an embodiment, the time-frequency transform performed by the time-frequency transform stage is an overlapped critical sampling transform.
[0009] In an embodiment, the time-frequency transform performed by the time-frequency transform stage on the identified one or more subband sample sets in the subband sample set based on the second sample block and / or the identified one or more subband sample sets in the subband sample set based on the second sample block corresponds to the transform described by the following formula:
[0010]
[0011] where S(m) describes the transform, where m describes the index of a block of samples of the audio signal, where T0…T K The sub-band samples of the corresponding identified one or more sub-band sample sets are described.
[0012] For example, the time-frequency transform stage may be configured to perform time-frequency transform on one or more identified subband sample sets in the subband sample set based on the second sample block and / or one or more identified subband sample sets in the subband sample set based on the second sample block based on the above formula.
[0013] In an embodiment, the cascaded overlapped critical sampling transform stages are configured to process a first set of binary bins obtained based on a first sample block of the audio signal and a second set of binary bins obtained based on a second sample block of the audio signal using a second overlapped critical sampling transform stage of the cascaded overlapped critical sampling transform stages, wherein the second overlapped critical sampling transform stage is configured to perform a first overlapped critical sampling transform on the first set of binary bins and a second overlapped critical sampling transform on the second set of binary bins depending on a signal characteristic of the audio signal (for example, when the signal characteristic of the audio signal changes), one or more of the first critical sampling transforms having a different length than that of the second critical sampling transform.
[0014] In an embodiment, the time-frequency transform stage is configured to: identify one or more subband sample sets from the subband sample sets based on the first sample block and identify one or more subband sample sets from the subband sample sets based on the second sample block, if one or more of the first critically sampled transforms have a different length [e.g., a merging factor] compared to the second critically sampled transform, the identified one or more subband sample sets representing the same time-frequency portion of the audio signal.
[0015] In an embodiment, the audio processor comprises a second time-frequency transform stage configured to time-frequency transform the aliasing reduced subband representation of the audio signal, wherein the time-frequency transform applied by the second time-frequency transform stage is opposite to the time-frequency transform applied by the first time-frequency transform stage.
[0016] In an embodiment, the time-domain aliasing reduction performed by the time-domain aliasing reduction stage corresponds to the transform described by the following formula:
[0017]
[0018] where R(z,m) describes the transform, where z describes the frame index in the z-domain, where m describes the index of the sample block of the audio signal, where F′0…F′ K A modified version of the NxN overlapped critically sampled transform pre-permutation / folding matrix is described.
[0019] In an embodiment, the audio processor is configured to provide a bitstream comprising a STDAR parameter indicating whether the length of the identified one or more subband sample sets corresponding to the first sample block or the second sample block is used in a time domain aliasing reduction stage to obtain a corresponding aliasing-reduced subband representation of the audio signal, or wherein the audio processor is configured to provide a bitstream comprising an MDCT length parameter [e.g. a merging factor [MF] parameter] indicating the length of the subband sample set.
[0020] In an embodiment, the audio processor is configured to perform joint channel coding.
[0021] In an embodiment, the audio processor is configured to perform M / S or MCT as joint channel processing.
[0022] In an embodiment, the audio processor is configured to provide a bitstream comprising at least one STDAR parameter indicating the lengths of one or more time-frequency transformed subband samples corresponding to the first block of samples and one or more time-frequency transformed subband samples corresponding to the second block of samples for use in a time-domain aliasing reduction stage to obtain a corresponding aliasing-reduced subband representation of the audio signal or an encoded version thereof [e.g. an entropy or differentially encoded version thereof].
[0023] In an embodiment, the cascaded overlapped critical sampling transform stages include a first overlapped critical sampling transform stage configured to perform an overlapped critical sampling transform on a first block of samples and a second block of samples of at least two partially overlapping blocks of samples of an audio signal to obtain a first set of binary bins for the first block of samples and a second set of binary bins for the second block of samples.
[0024] In an embodiment, the cascaded overlapped critical sampling transform stages further comprise a second overlapped critical sampling transform stage configured to perform overlapped critical sampling transform on segments of the first set of binary bins and to perform overlapped critical sampling transform on segments of the second set of binary bins to obtain a set of subband samples for the first set of binary bins and a set of subband samples for the second set of binary bins, wherein each segment is associated with a subband of the audio signal.
[0025] Another embodiment provides an audio processor for processing a subband representation of an audio signal to obtain the audio signal, the subband representation of the audio signal comprising an aliasing-reduced sample set. The audio processor comprises: a second inverse time-frequency transform stage configured to time-frequency transform one or more of the aliasing-reduced subband sample sets corresponding to a second sample block of the audio signal and / or one or more of the aliasing-reduced subband sample sets corresponding to the second sample block of the audio signal to obtain one or more time-frequency transformed aliasing-reduced subband samples, each of the time-frequency transformed aliasing-reduced subband samples representing the same region in the time-frequency plane as a corresponding one of the one or more aliasing-reduced subband samples or the time-frequency transformed versions of the one or more aliasing-reduced subband samples corresponding to another sample block of the audio signal. Furthermore, the audio processor comprises: an inverse time-domain aliasing reduction stage configured to perform a weighted combination of the corresponding aliasing-reduced subband sample sets or the time-frequency transformed versions thereof to obtain the aliased subband representation. Furthermore, the audio processor comprises a first inverse time-frequency transform stage configured to perform a time-frequency transform on the aliased subband representation to obtain a subband sample set corresponding to a first sample block of the audio signal and a subband sample set corresponding to a second sample block of the audio signal, wherein the time-frequency transform applied by the first inverse time-frequency transform stage is the inverse of the time-frequency transform applied by the second inverse time-frequency transform stage. Furthermore, the audio processor comprises a cascaded inverse overlapped critically sampled transform stage configured to perform a cascaded inverse overlapped critically sampled transform on the sample set to obtain a sample set associated with the sample block of the audio signal.
[0026] Another embodiment provides a method for processing an audio signal to obtain a subband representation of the audio signal. The method comprises the steps of performing a cascaded overlapped critical sampling transform on at least two partially overlapping blocks of samples of the audio signal to obtain a subband sample set based on a first block of samples of the audio signal and a subband sample set based on a second block of samples of the audio signal. The method further comprises the steps of identifying one or more subband sample sets from the subband sample set based on the first block of samples, if the subband sample set based on the first block of samples represents a different region in a time-frequency plane than the subband sample set based on the second block of samples, and identifying one or more subband sample sets from the subband sample set based on the second block of samples, the identified one or more subband sample sets being combined to represent the same region in the time-frequency plane. The method further comprises the steps of: performing a time-frequency transform on the identified one or more subband sample sets of the subband sample sets based on the first sample block and / or the identified one or more subband sample sets of the subband sample sets based on the second sample block to obtain one or more time-frequency transformed subband samples, each of the time-frequency transformed subband samples representing the same region in the time-frequency plane as a corresponding one of the identified one or more subband samples or the time-frequency transformed versions of the one or more subband samples. The method further comprises the steps of: performing a weighted combination of two corresponding subband sample sets or the time-frequency transformed versions thereof to obtain an aliasing-reduced subband representation of the audio signal, wherein one of the two corresponding subband sample sets or the time-frequency transformed version thereof is obtained based on the first sample block of the audio signal and the other of the two corresponding subband sample sets or the time-frequency transformed version thereof is obtained based on the second sample block of the audio signal.
[0027] Further embodiments provide a method of processing a subband representation of an audio signal, the subband representation of the audio signal comprising an aliasing- reduced set of samples. The method comprises the steps of performing a time- frequency transform on one or more of the aliasing-reduced sets of subband samples corresponding to a second block of samples of the audio signal and / or on a time- frequency transformed version of one or more of the aliasing-reduced sets of subband samples corresponding to a second block of samples of the audio signal, to obtain one or more time- frequency transformed aliasing-reduced subband samples each representing a same area in a time- frequency plane as a corresponding one of the one or more aliasing-reduced subband samples or the time- frequency transformed version of the one or more aliasing-reduced subband samples corresponding to a further block of samples of the audio signal. Further, the method comprises the step of performing a weighted combination of the corresponding sets of aliasing-reduced subband samples or the time- frequency transformed version thereof to obtain an aliased subband representation. Further, the method comprises the step of performing a time- frequency transform on the aliased subband representation to obtain a set of subband samples corresponding to a first block of samples of the audio signal and a set of subband samples corresponding to a second block of samples of the audio signal, wherein the time- frequency transform applied by the first inverse time- frequency transform stage is inverse to the time- frequency transform applied by the second inverse time- frequency transform stage. Further, the method comprises the step of performing a concatenated inverse overlap-critical sampling transform on the sets of samples to obtain a set of samples associated with the blocks of samples of the audio signal.
[0028] According to the concept of the present application, by introducing another symmetric subband merging / subband splitting step, which equalizes the time- frequency tilings of the two frames, it is allowed to perform time-domain aliasing reduction between two frames of different time- frequency tilings. After equalizing the tilings, time-domain aliasing reduction can be applied and the original tilings can be reconstructed.
[0029] Embodiments provide switching time-domain aliasing reduction (STDAR) filter banks with single-sided or double-sided STDAR.
[0030] In embodiments, the STDAR parameters can be derived from the MDCT length parameters, e.g. the merging factor (MF) parameters. For example, in case of using single-sided STDAR, 1 bit can be transmitted for each merging factor. This bit can signal whether the merging factor of frame m or frame m-1 is used for STDAR. Alternatively, the transform can always be performed towards the higher merging factor. In this case, the bit can be omitted.
[0031] In embodiments, joint channel processing can be performed, e.g. M / S or multi-channel coding tool (MCT)
[10] . For example, some or all channels can be transformed to the same TDAR layout based on bilateral STDAR and jointly processed. There is probably less likelihood for varying factors such as 2, 8, 1, 2, 16, 32 than for uniform factors such as 4, 4, 8, 8, 16, 16. This correlation can be exploited to reduce the amount of data required, e.g. by means of differential coding.
[0032] In embodiments, fewer merge factors can be transmitted, wherein the omitted merge factors can be derived or interpolated from neighboring merge factors. For example, if the merge factors are actually uniform as described in the previous paragraph, then all merge factors can be interpolated based on a few merge factors.
[0033] In embodiments, the bilateral STDAR factors can be signaled in the bitstream. For example, some bits in the bitstream are required to signal the STDAR factors describing the current frame restriction. These bits can be entropy coded. In addition, these bits can be coded with respect to each other.
[0034] A further embodiment provides an audio processor for processing an audio signal to obtain a subband representation of the audio signal. The audio processor comprises a cascaded overlap critical sampling transform stage and a time domain aliasing reduction stage. The cascaded overlap critical sampling transform stage is configured to perform a cascaded overlap critical sampling transform on at least two partially overlapping blocks of samples of the audio signal to obtain a set of subband samples based on a first block of samples of the audio signal and to obtain a corresponding set of subband samples based on a second block of samples of the audio signal. The time domain aliasing reduction stage is configured to perform a weighted combination of the two corresponding sets of subband samples to obtain an aliasing-reduced subband representation of the audio signal, wherein one of the two corresponding sets of subband samples is obtained based on the first block of samples of the audio signal and the other of the two corresponding sets of subband samples is obtained based on the second block of samples of the audio signal.
[0035] A further embodiment provides an audio processor for processing a subband representation of an audio signal to obtain the audio signal. The audio processor comprises an inverse time domain aliasing reduction stage and a cascaded inverse overlap critical sampling transform stage. The inverse time domain aliasing reduction stage is configured to perform a weighted (and shifted) combination of two corresponding aliasing-reduced subband representations of the audio signal (of different partially overlapping blocks of samples) to obtain an aliased subband representation, wherein the aliased subband representation is a set of subband samples. The cascaded inverse overlap critical sampling transform stage is configured to perform a cascaded inverse overlap critical sampling transform on the set of subband samples to obtain a set of samples associated with a block of samples of the audio signal.
[0036] According to the concepts of the present invention, an additional post-processing stage is added to the overlapped critically sampled transform (e.g., MDCT) pipeline, which includes another overlapped critically sampled transform (e.g., MDCT) along the frequency axis and time-domain aliasing reduction along the time axis of each subband. This allows the extraction of arbitrary frequency scales from the overlapped critically sampled transform (e.g., MDCT) spectrogram and improves the temporal compactness of the impulse response without introducing additional redundancy and only introducing the reduced overlapped critically sampled transform frame delay.
[0037] Another embodiment provides a method of processing an audio signal to obtain a sub-band representation of the audio signal. The method comprises:
[0038] - performing a cascaded overlapped critical sampling transform on at least two partially overlapping blocks of samples of the audio signal to obtain a set of subband samples based on a first block of samples of the audio signal and to obtain a corresponding set of subband samples based on a second block of samples of the audio signal; and
[0039] - performing a weighted combination of two corresponding subband sample sets to obtain an aliasing-reduced subband representation of the audio signal, wherein one of the two corresponding subband sample sets is obtained based on a first block of samples of the audio signal and the other of the two corresponding subband sample sets is obtained based on a second block of samples of the audio signal.
[0040] Another embodiment provides a method for processing a subband representation of an audio signal to obtain an audio signal. The method comprises:
[0041] - performing a weighted (and shifted) combination of two corresponding aliasing-reduced subband representations (of different, partially overlapping blocks of samples) of the audio signal to obtain an aliased subband representation, wherein the aliased subband representation is a set of subband samples; and
[0042] - performing a concatenated inverse lapped critical sampling transform on the sets of subband samples to obtain sets of samples associated with a block of samples of the audio signal.
[0043] Advantageous embodiments are set forth in the dependent claims.
[0044] Subsequently, advantageous embodiments of an audio processor for processing an audio signal to obtain a sub-band representation of the audio signal are described.
[0045] In an embodiment, the cascaded overlapped critically sampled transform stages may be cascaded MDCT (MDCT=Modified Discrete Cosine Transform), MDST (MDST=Modified Discrete Sine Transform) or MLT (MLT=Modulated Overlapped Transform) stages.
[0046] In an embodiment, the cascaded overlapped critical sampling transform stages may include a first overlapped critical sampling transform stage, which is configured to perform an overlapped critical sampling transform on a first sample block and a second sample block of at least two partially overlapping sample blocks of the audio signal to obtain a first set of binary bins for the first sample block and a second set of binary bins (overlapped critical sampling coefficients) for the second sample block.
[0047] The first overlapped critical sampling transform stage may be a first MDCT, MDST or MLT stage.
[0048] The cascaded overlapped critical sampling transform stage may further include a second overlapped critical sampling transform stage, which is configured to perform overlapped critical sampling transform on segments (appropriate subsets) of the first binary bin set and perform overlapped critical sampling transform on segments (appropriate subsets) of the second binary bin set to obtain a subband sample set for the first binary bin set and a subband sample set for the second binary bin set, wherein each segment is associated with a subband of the audio signal.
[0049] The second overlapped critical sampling transform stage may be a second MDCT, MDST or MLT stage.
[0050] Thus, the first overlapped critically sampled transform stage and the second overlapped critically sampled transform stage may be of the same type, ie one of an MDCT, MDST or MLT stage.
[0051] In an embodiment, the second overlapped critical sampling transform stage may be configured to perform an overlapped critical sampling transform on at least two partially overlapping segments (appropriate subsets) of the first binary bin set, and to perform an overlapped critical sampling transform on at least two partially overlapping segments (appropriate subsets) of the second binary bin set, to obtain at least two subband sample sets for the first binary bin set and at least two subband sample sets for the second binary bin set, wherein each segment is associated with a subband of the audio signal.
[0052] Thus, the first subband sample set may be a result of a first overlapped critical sampling transform based on a first segment of the first binary bin set, wherein the second subband sample set may be a result of a second overlapped critical sampling transform based on a second segment of the first binary bin set, wherein the third subband sample set may be a result of a third overlapped critical sampling transform based on a first segment of the second binary bin set, wherein the fourth subband sample set may be a result of a fourth overlapped critical sampling transform based on a second segment of the second binary bin set. The time domain aliasing reduction stage may be configured to perform a weighted combination of the first subband sample set and the third subband sample set to obtain a first aliasing-reduced subband representation of the audio signal, and to perform a weighted combination of the second subband sample set and the fourth subband sample set to obtain a second aliasing-reduced subband representation of the audio signal.
[0053] In an embodiment, the cascaded overlapping critical sampling transform stage can be configured to segment the set of bins obtained based on the first block of samples using at least two window functions, and obtain at least two sets of subband samples based on the segmented set of bins corresponding to the first block of samples, wherein the cascaded overlapping critical sampling transform stage can be configured to segment the set of bins obtained based on the second block of samples using at least two window functions, and obtain at least two sets of subband samples based on the segmented set of bins corresponding to the second block of samples, wherein the at least two window functions comprise different window widths.
[0054] In an embodiment, the cascaded overlapping critical sampling transform stage can be configured to segment the set of bins obtained based on the first block of samples using at least two window functions, and obtain at least two sets of subband samples based on the segmented set of bins corresponding to the first block of samples, wherein the cascaded overlapping critical sampling transform stage can be configured to segment the set of bins obtained based on the second block of samples using at least two window functions, and obtain at least two sets of subband samples based on the segmented set of bins corresponding to the second block of samples, wherein the filter slopes of the window functions corresponding to adjacent sets of subband samples are symmetrical.
[0055] In an embodiment, the cascaded overlapping critical sampling transform stage can be configured to segment the samples of the audio signal into a first block of samples and a second block of samples using a first window function, wherein the overlapping critical sampling transform stage can be configured to segment the set of bins obtained based on the first block of samples and the set of bins obtained based on the second block of samples using a second window function to obtain corresponding sets of subband samples, wherein the first window function and the second window function comprise different window widths.
[0056] In an embodiment, the cascaded overlapping critical sampling transform stage can be configured to segment the samples of the audio signal into a first block of samples and a second block of samples using a first window function, wherein the overlapping critical sampling transform stage can be configured to segment the set of bins obtained based on the first block of samples and the set of bins obtained based on the second block of samples using a second window function to obtain corresponding sets of subband samples, wherein the window width of the first window function and the window width of the second window function are different from each other by a factor different from a power of two.
[0057] Subsequently, advantageous embodiments of an audio processor for processing a subband representation of an audio signal to obtain an audio signal are described.
[0058] In an embodiment, the inverse cascaded overlapping critical sampling transform stage can be an inverse cascaded MDCT (MDCT = Modified Discrete Cosine Transform), MDST (MDST = Modified Discrete Sine Transform) or MLT (MLT = Modulated Lapped Transform) stage.
[0059] In an embodiment, the cascaded inverse overlapped critically sampled transform stages may comprise a first inverse overlapped critically sampled transform stage configured to perform an inverse overlapped critically sampled transform on a set of subband samples to obtain a set of binary bins associated with a given subband of the audio signal.
[0060] The first inverse overlapped critical sampling transform stage may be a first inverse MDCT, MDST or MLT stage.
[0061] In an embodiment, the cascaded inverse overlapped critically sampled transform stages may comprise a first overlap and add stage configured to perform a concatenation of binary bins associated with a plurality of subbands of the audio signal, comprising a weighted combination of a binary bin associated with a given subband of the audio signal and a binary bin associated with another subband of the audio signal to obtain a binary bin associated with a block of samples of the audio signal.
[0062] In an embodiment, the cascaded inverse overlapped critical sampling transform stages may include a second inverse overlapped critical sampling transform stage configured to perform an inverse overlapped critical sampling transform on a set of binary bins associated with a block of samples of the audio signal to obtain a set of samples associated with the block of samples of the audio signal.
[0063] The second inverse overlapped critical sampling transform stage may be a second inverse MDCT, MDST or MLT stage.
[0064] Thus, the first inverse overlapped critically sampled transform stage and the second inverse overlapped critically sampled transform stage may be of the same type, ie one of an inverse MDCT, MDST or MLT stage.
[0065] In an embodiment, the cascaded inverse overlapped critically sampled transform stages may comprise a second overlap and add stage configured to overlap and add a set of samples associated with a block of samples of the audio signal and another set of samples associated with another block of samples of the audio signal to obtain the audio signal, the block of samples of the audio signal and the another block of samples partially overlapping. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Herein, embodiments of the present invention are described with reference to the accompanying drawings.
[0067] Figure 1 shows a schematic block diagram of an audio processor configured to process an audio signal to obtain a sub-band representation of the audio signal according to an embodiment;
[0068] Figure 2 shows a schematic block diagram of an audio processor configured to process an audio signal to obtain a sub-band representation of the audio signal according to another embodiment;
[0069] Figure 3 shows a schematic block diagram of an audio processor configured to process an audio signal to obtain a sub-band representation of the audio signal according to another embodiment;
[0070] Figure 4 shows a schematic block diagram of an audio processor for processing a sub-band representation of an audio signal to obtain an audio signal according to an embodiment;
[0071] Figure 5 shows a schematic block diagram of an audio processor for processing a sub-band representation of an audio signal to obtain an audio signal according to another embodiment;
[0072] Figure 6 shows a schematic block diagram of an audio processor for processing a sub-band representation of an audio signal to obtain an audio signal according to another embodiment;
[0073] Figure 7 Examples of sub-band samples are shown graphically (top) and their spread over time and frequency (bottom);
[0074] Figure 8 The spectral and temporal uncertainties obtained by several different transformations are shown graphically;
[0075] Figure 9 A comparison of two exemplary impulse responses generated by subband combining with and without TDAR, simple MDCT short blocks, and Hadamard matrix subband combining is shown graphically;
[0076] Figure 10 shows a flow chart of a method for processing an audio signal to obtain a sub-band representation of the audio signal according to an embodiment;
[0077] Figure 11 shows a flow chart of a method for processing a sub-band representation of an audio signal to obtain an audio signal according to an embodiment;
[0078] Figure 12 shows a schematic block diagram of an audio encoder according to an embodiment;
[0079] Figure 13 shows a schematic block diagram of an audio decoder according to an embodiment;
[0080] Figure 14 shows a schematic block diagram of an audio analyzer according to an embodiment;
[0081] Figure 15 shows a schematic block diagram of an audio processor configured to process an audio signal to obtain a sub-band representation of the audio signal according to another embodiment;
[0082] Figure 16 shows a schematic representation of the time-frequency transform performed by the time-frequency transform stage in the time-frequency plane;
[0083] Figure 17 shows a schematic block diagram of an audio processor configured to process an audio signal to obtain a sub-band representation of the audio signal according to another embodiment;
[0084] Figure 18 shows a schematic block diagram of an audio processor for processing a sub-band representation of an audio signal to obtain an audio signal according to another embodiment;
[0085] Figure 19 shows a schematic representation of the STDAR operation in the time-frequency plane;
[0086] Figure 20 Graphs show example impulse responses for two frames with binning factors of 8 and 16 before (top) and after (bottom) STDAR;
[0087] Figure 21 Graphing impulse response and frequency response compactness for upper matching;
[0088] Figure 22 Graphing impulse response and frequency response compactness for downmatching;
[0089] Figure 23 A flow chart showing a method for processing an audio signal to obtain a sub-band representation of the audio signal according to another embodiment; and
[0090] Figure 24 A flow chart of a method for processing a subband representation of an audio signal to obtain an audio signal according to another embodiment is shown, the subband representation of the audio signal comprising an aliasing-reduced sample set. DETAILED DESCRIPTION
[0091] In the following description, the same or equivalent elements or elements having the same or equivalent functions are denoted by the same or equivalent reference numerals.
[0092] In the following description, a number of details are set forth to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention. Furthermore, unless otherwise specifically indicated, the features of the different embodiments described below may be combined with one another.
[0093] First, in Section 1, a non-uniform orthogonal filter bank based on the cascade of two MDCTs and time-domain aliasing reduction (TDAR) is described, which enables compact impulse responses in both time and frequency [1]. Then, in Section 2, switched time-domain aliasing reduction (STDAR) is described, which allows TDAR between two frames with different time-frequency tiles. This is achieved by introducing another symmetric subband merging / subband splitting step that equalizes the time-frequency tiles of the two frames. After equalizing the tiles, conventional TDAR is applied and the original tiles are reconstructed.
[0094] 1. Non-uniform orthogonal filter banks based on cascading two MDCTs and time domain aliasing reduction (TDAR)
[0095] Figure 1 A schematic block diagram of an audio processor 100 configured to process an audio signal 102 to obtain a sub-band representation of the audio signal according to an embodiment is shown. The audio processor 100 comprises a cascaded overlapped critically sampled transform (LCST) stage 104 and a time domain aliasing reduction (TDAR) stage 106.
[0096] The cascaded overlapped critical sampling transform stage 104 is configured to perform a cascaded overlapped critical sampling transform on at least two partially overlapping blocks of samples 108_1 and 108_2 of the audio signal 102 to obtain a subband sample set 110_1,1 based on a first block of samples 108_1 (of the at least two overlapping blocks of samples 108_1 and 108_2) of the audio signal 102, and to obtain a corresponding subband sample set 110_2,1 based on a second block of samples 108_2 (of the at least two overlapping blocks of samples 108_1 and 108_2) of the audio signal 102.
[0097] The time-domain aliasing reduction stage 104 is configured to perform a weighted combination of two corresponding subband sample sets 110_1,1 and 110_2,1 (i.e., subband samples corresponding to the same subband), one of which is obtained based on the first block of samples 108_1 of the audio signal 102 and the other of which is obtained based on the second block of samples 108_2 of the audio signal, to obtain an aliasing-reduced subband representation 112_1 of the audio signal 102.
[0098] In an embodiment, the cascaded overlapped critical sampling transform stage 104 may include at least two cascaded overlapped critical sampling transform stages, or in other words, two overlapped critical sampling transform stages are connected in a cascade manner.
[0099] The cascaded overlapped critically sampled transform stages may be cascaded MDCT (MDCT=Modified Discrete Cosine Transform) stages.The cascaded MDCT stages may comprise at least two MDCT stages.
[0100] Naturally, the cascaded overlapped critically sampled transform stages can also be a cascaded MDST (MDST=Modified Discrete Sine Transform) or MLT (MLT=Modulated Lapped Transform) stages, comprising at least two MDST or MLT stages, respectively.
[0101] The two corresponding subband sample sets 110_1,1 and 110_2,1 may be subband samples corresponding to the same subband (ie, frequency band).
[0102] Figure 2 Shown is a schematic block diagram of an audio processor 100 configured to process an audio signal 102 to obtain a sub-band representation of the audio signal according to another embodiment.
[0103] like Figure 2 As shown, the cascaded overlapped critical sampling transform stage 104 may include a first overlapped critical sampling transform stage 120, which is configured to: perform an overlapped critical sampling transform on at least two partially overlapping sample blocks 108_1 and 108_2 of the audio signal 102 consisting of (2M) samples (x i-1 (n), 0≤n≤2M-1) and the first sample block 108_1 consisting of (2M) samples (x i The overlapped critical sampling transform is performed on the second sample block 108_2 composed of (n), 0≤n≤2M-1) to obtain a binary bin (LCST coefficient) (X i-1 (k), 0≤k≤M-1) and the first binary bin set 124_1 consisting of (M) binary bins (LCST coefficients) (X i (k), 0≤k≤M-1) constitutes a second binary bin set 124_2.
[0104] The cascaded overlapped critical sampling transform stage 104 may include a second overlapped critical sampling transform stage 126 configured to: v,i-1 (k)) performs overlapped critical sampling transform and performs overlapped critical sampling transform on the segments 128_2,1 (appropriate subset) of the second set of binary bins 124_2 (X v,i (k) Perform overlapped critical sampling transform to obtain subband samples for the first binary bin set 124_1 Set 110_1,1 and subband samples for the second binary bin set 124_2 A set 110_2,1, wherein each segment is associated with a subband of the audio signal 102.
[0105] Figure 31 shows a schematic block diagram of an audio processor 100 configured to process an audio signal 102 to obtain a sub-band representation of the audio signal according to another embodiment. Figure 3 A diagram of the analysis filter bank is shown in FIG. Thus, an appropriate window function is assumed. Note that for simplicity, Figure 3 Indicates (only) the first half of the sub-band frame (y[m], 0 <= m <N / 2)的处理(即,仅等式(6)的第一行)。
[0106] like Figure 3 As shown, the first overlapped critical sampling transform stage 120 can be configured to: i-1 (n), 0≤n≤2M-1) performs a first overlapped critical sampling transform 122_1 (eg, MDCTi-1) on the first sample block 108_1 to obtain a binary bin (LCST coefficient) (X i-1 (k), 0≤k≤M-1) composed of the first binary bin set 124_1; and for the (2M) samples (x i (n), 0≤n≤2M-1) performs a second overlapped critical sampling transform 122_2 (eg, MDCT i) on the second sample block 108_2 to obtain a binary bin (LCST coefficient) (X i (k), 0≤k≤M-1) constitutes a second binary bin set 124_2.
[0107] In detail, the second overlapped critical sampling transform stage 126 may be configured to: perform a multi-segment multiplication of at least two partially overlapping segments 128_1,1 and 128_1,2 (appropriate subsets) of the first set of binary bins 124_1 (X v,i-1 (k)) Perform an overlapped critical sampling transform and perform an overlapped critical sampling transform on at least two partially overlapping segments 128_2,1 and 128_2,2 (appropriate subsets) of the second set of binary bins (X v,i (k) Perform overlapped critical sampling transform to obtain at least two subband samples for the first binary bin set 124_1 The sets 110_1,1 and 110_1,2 and at least two subband samples for the second binary bin set 124_2 Sets 110_2,1 and 110_2,2, wherein each segment is associated with a subband of the audio signal.
[0108] For example, the first subband sample set 110_1,1 may be a result of performing a first overlapped critical sampling transform 132_1,1 based on the first segment 132_1,1 of the first binary bin set 124_1, wherein the second subband sample set 110_1,2 may be a result of performing a second overlapped critical sampling transform 132_1,2 based on the second segment 128_1,2 of the first binary bin set 124_1, wherein the third subband sample set 110_2,1 may be a result of performing a third overlapped critical sampling transform 132_2,1 based on the first segment 128_2,1 of the second binary bin set 124_2, wherein the fourth subband sample set 110_2,2 may be a result of performing a fourth overlapped critical sampling transform 132_2,2 based on the second segment 128_2,2 of the second binary bin set 124_2.
[0109] Thus, the time domain aliasing reduction stage 106 may be configured to perform a weighted combination of the first subband sample set 110_1,1 and the third subband sample set 110_2,1 to obtain a first aliasing-reduced subband representation 112_1 (y 1,i [m1]), wherein the time domain aliasing reduction stage 106 may be configured to perform a weighted combination of the second subband sample set 110_1,2 and the fourth subband sample set 110_2,2 to obtain a second aliasing-reduced subband representation 112_2 (y 2,i [m2]).
[0110] Figure 4 A schematic block diagram of an audio processor 200 according to an embodiment is shown for processing a subband representation of an audio signal to obtain the audio signal 102. The audio processor 200 comprises an inverse time domain aliasing reduction (TDAR) stage 202 and a cascaded inverse lapped critically sampled transform (LCST) stage 204.
[0111] The inverse time-domain aliasing reduction stage 202 is configured to perform aliasing reduction on two corresponding sub-band representations 112_1 and 112_2 (y v,i (m), y v,i-1 (m)) to obtain the aliased subband representation 110_1 The aliased subband is represented by the subband sample set 110_1.
[0112] The cascaded inverse overlapped critically sampled transform stage 204 is configured to perform a cascaded inverse overlapped critically sampled transform on the set of subband samples 110_1 to obtain a set of samples associated with the block of samples 108_1 of the audio signal 102 .
[0113] Figure 5A schematic block diagram of an audio processor 200 according to another embodiment for processing a subband representation of an audio signal to obtain the audio signal 102 is shown. The cascaded inverse overlapped critically sampled transform stage 204 may include a first inverse overlapped critically sampled transform (LCST) stage 208 and a first overlap and add stage 210.
[0114] The first inverse overlapped critical sampling transform stage 208 may be configured to perform an inverse overlapped critical sampling transform on the set of subband samples 110_1,1 to obtain a set of binary bins 128_1,1 associated with a given subband of the audio signal.
[0115] The first overlap and add stage 210 may be configured to perform a concatenation of binary bins associated with a plurality of subbands of the audio signal, including the binary bin 128_1,1 associated with a given subband (v) of the audio signal 102. and a set of binary bins 128_1,2 associated with another subband (v-1) of the audio signal 102 to obtain a set of binary bins 124_1 associated with the block of samples 108_1 of the audio signal 102.
[0116] like Figure 5 As shown, the cascaded inverse overlapped critically sampled transform stage 204 may include a second inverse overlapped critically sampled transform (LCST) stage 212, which is configured to perform an inverse overlapped critically sampled transform on the binary bin set 124_1 associated with the sample block 108_1 of the audio signal 102 to obtain a sample set 206_1,1 associated with the sample block 108_1 of the audio signal 102.
[0117] Furthermore, the cascaded inverse overlapped critically sampled transform stage 204 may include a second overlap and add stage 214 configured to overlap and add a set of samples 206_1,1 associated with the block of samples 108_1 of the audio signal 102 and another set of samples 206_2,1 associated with another block of samples 108_2 of the audio signal to obtain the audio signal 102, wherein the block of samples 108_1 and the another block of samples 108_2 of the audio signal 102 partially overlap.
[0118] Figure 6 1 shows a schematic block diagram of an audio processor 200 for processing a subband representation of an audio signal to obtain the audio signal 102 according to another embodiment. In other words, Figure 6 A diagram of the synthesis filter bank is shown in FIG. Thus, an appropriate window function is assumed. Note that for simplicity, Figure 6Indicates (only) the first half of the sub-band frame (y[m], 0 <= m <N / 2)的处理(即,仅等式(6)的第一行)。
[0119] As described above, the audio processor 200 includes the inverse time-domain aliasing reduction stage 202 and the inverse cascaded overlap critical sampling stages 204 , which includes the first inverse overlap critical sampling stage 208 and the second inverse overlap critical sampling stage 212 .
[0120] The inverse time domain reduction stage 104 is configured to perform a first aliasing reduction on the sub-band representation y 1,i-1 [ m1 ] and the second aliasing-reduced subband representation y 1,i [ m1 ] is a first weighted and shifted combination 220_1 to obtain a first aliased subband representation 110_1, 1 Wherein the aliased subband representation is a subband sample set; and the subband representation y where the third aliasing reduction is performed is 2,i-1 [ m1 ] and the fourth aliasing-reduced subband representation y 2,i [ m1 ] is a second weighted and shifted combination 220_2 to obtain a second aliased subband representation 110_2, 1 The aliased subband representation is a subband sample set.
[0121] The first inverse overlap critical sampling transform stage 208 is configured to: A first inverse overlapped critical sampling transform 222_1 is performed to obtain a set of binary bins 128_1, 1 associated with a given subband of the audio signal. And for the second subband sample set 110_2,1 Perform the second inverse overlap critical sampling transform 222_2 to obtain the audio signal The set of binary bins 128_2,1 associated with a given subband.
[0122] The second inverse overlap critical sampling transform stage 212 is configured to perform an inverse overlap critical sampling transform on the overlapped and added binary bins obtained by overlapping and adding the binary bins 128_1,1 and 128_2,1 provided by the first inverse overlap critical sampling transform stage 208 to obtain the sample block 108_2.
[0123] Subsequently, it was described Figures 1 to 6, wherein it is exemplarily assumed that the cascaded overlap critical sampling transform stage 104 is an MDCT stage (i.e., the first overlap critical sampling transform stage 120 and the second overlap critical sampling transform stage 126 are MDCT stages), and the inverse cascaded overlap critical sampling transform stage 204 is an inverse cascaded MDCT stage (i.e., the first inverse overlap critical sampling transform stage 120 and the second inverse overlap critical sampling transform stage 126 are inverse MDCT stages). Naturally, the following description is also applicable to other embodiments of the cascaded overlap critical sampling transform stage 104 and the inverse overlap critical sampling transform stage 204, such as to cascaded MDST or MLT stages or inverse cascaded MDST or MLT stages.
[0124] Thus, the described embodiments can be applied to MDCT spectrum sequences of finite length and use MDCT and time-domain aliasing reduction (TDAR) as the subband merging operation. The resulting non-uniform filter bank is overlapping, orthogonal, and allows subband widths k = 2n, where n∈N. Due to TDAR, more compact subband impulse responses can be achieved in both time and spectrum.
[0125] Subsequently, embodiments of the filter bank are described.
[0126] The filter bank implementation builds directly on the common overlapped MDCT transform scheme: the original transform remains unchanged with overlapping and windowing.
[0127] Without loss of generality, the following notation assumes an orthogonal MDCT transform, eg, where the analysis and synthesis windows are the same.
[0128] x i (n)=x(n+iM) 0≤n≤2M (1)
[0129]
[0130] Where k(k,n,M) is the MDCT transform kernel and h(n) is the appropriate analysis window
[0131]
[0132] Then, the output X of this transformation i (k) The segments have their own width N υ The u subbands are transformed again using MDCT. This results in filter banks that overlap in both time and spectrum.
[0133] For simpler notation in this paper, a common merging factor N is used for all subbands, however any valid MDCT window switching / ordering can be used to achieve the desired time-frequency resolution. More on resolution design follows.
[0134] X v,i (k) = X i (k+vN) 0≤k<2N (4)
[0135]
[0136] Where w(k) is a suitable analysis window, which is usually of a different size than h(n) and may also be of a different window type. Since the embodiment applies the window in the frequency domain, it is worth noting that the time and frequency selectivity of the window are swapped.
[0137] For correct boundary handling, an additional offset of N / 2 can be introduced in equation (4) in conjunction with half of the rectangular start / stop window at the boundary. Again, for simpler notation, this offset is not considered here.
[0138] Output Has the corresponding bandwidth The coefficients of time resolution proportional to the bandwidth have respective lengths N v A list of v vectors.
[0139] However, these vectors contain aliasing from the original MDCT transform and therefore show poor temporal compactness. TDAR can help compensate for this aliasing.
[0140] The samples used for TDAR are taken from two adjacent subband sample blocks v in the current MDCT frame i and the previous MDCT frame i- 1. The result is that aliasing is reduced in the second half of the previous frame and the first half of the second frame.
[0141] For 0≤m <N / 2,
[0142]
[0143] in
[0144]
[0145] The TDAR coefficient a can be designed v (m), b v (m), c v (m) and d v (m) to minimize residual aliasing. A simple estimation method based on the synthesis window g(n) is introduced below.
[0146] Note also that if A is nonsingular, then operations (6) and (8) correspond to a biorthogonal system. In addition, if g(n) = h(n) and v(k) = w(k), for example, both MDCTs are orthogonal, and the matrix A is orthogonal, then the entire pipeline constitutes an orthogonal transform.
[0147] To compute the inverse transform, a first inverse TDAR is performed,
[0148]
[0149] Then, inverse MDCT and time domain aliasing cancellation (TDAC, although here aliasing cancellation is performed along the frequency axis) must be performed to remove the aliasing produced in Equation 5
[0150]
[0151]
[0152] X i (k+vN)=X v,i (k) (11)
[0153] Finally, the initial MDCT in Equation 2 is inverted and TDAC is performed again
[0154]
[0155]
[0156] x(n+iM)=x i (n) (14)
[0157] Subsequently, the time-frequency resolution design constraints are described. Although any desired time-frequency resolution is possible, some constraints on the design of the resulting window function must be observed to ensure reversibility. Specifically, the slopes of two adjacent subbands can be symmetrical so that equation (6) satisfies the Princen-Bradley condition [J. Princen, A. Johnson, and A. Bradley, "Subband / transform coding using filter bank designs based on timedomain aliasing cancellation", in Acoustics, Speech, and Signal Processing, IEEE International Conference on ICASSP'87., April 1987, Vol. 12, pp. 2161–2164]. The window switching scheme originally designed to combat the pre-echo effect, introduced in [B. Edler, "Codierung von Audiosignalen mitüberlappender Transformation und adaptiven Fensterfunktionen", Frequenz, Vol. 43, pp. 252–256, September 1989], can be applied here. See [Olivier Derrien, Thibaud Necciari, and Peter Balazs, “A quasi-orthogonal, invertible, and perceptually relevant time-frequency transform for audio coding,” in EUSIPCO, Nice, France, August 2015].
[0158] Secondly, the sum of all second MDCT transform lengths must add up to the total length of the provided MDCT coefficients. Frequency bands can be selected so that they are not transformed using a unit step window with zeros at the desired coefficients. However, care must be taken to ensure symmetry between adjacent windows [B. Edler, "Codierung von Audiosignalen mitüberlappender Transformation und adaptiven Fensterfunktionen", Frequenz, vol. 43, pp. 252–256, September 1989]. The resulting transform will produce zeros in these frequency bands, so the original coefficients can be used directly.
[0159] As a possible time-frequency resolution, the scale factor bands from most modern audio coders can be used directly.
[0160] Subsequently, the time-domain aliasing reduction (TDAR) coefficient computation is described.
[0161] Following the above time resolution, each subband sample corresponds to M / N v original samples or to an interval of size N v times the size of an original sample.
[0162] Furthermore, the amount of aliasing in each subband sample depends on the amount of aliasing in the interval it represents. When aliasing is weighted with the analysis window h(n), using an approximation of the synthesis window at the subband sample interval is considered to be a good first estimate of the TDAR coefficients.
[0163] Experiments have shown that two very simple coefficient computation schemes allow for good initial values with improved time and spectral compactness. Both methods are based on an assumed synthesis window g v (m) of length 2N v .
[0164] 1) For parametric windows like sine windows or Kaiser Bessel derived windows, a simple and short window of the same type can be defined.
[0165] 2) For parametric windows without closed representation and tabulated windows, the window can simply be cut into 2N v equal parts, allowing to use the average of each part to obtain the coefficients:
[0166]
[0167] Taking the MDCT boundary conditions and aliasing mirroring into account, the TDAR coefficients
[0168] a v (m) = g v (N / 2 + m) (16)
[0169] b v (m) = -g v (N / 2 - 1 - m) (17)
[0170] c v (m) = g v (3N / 2 + m) (18)
[0171] d v (m) = g v (3N / 2 - 1 - m) (19)
[0172] Or in the case of orthogonal transforms
[0173] a v (m) = d v (m) = g v (N / 2+m) (20)
[0174]
[0175] Regardless of the coefficient approximation chosen, as long as A is non-singular, the perfect reconstruction of the entire filter bank is preserved. In addition, suboptimal coefficient selection will only affect the subband signal y v,i (m), but does not affect the residual aliasing in the signal x(n) combined by the inverse filter.
[0176] Figure 7 Examples of subband samples are shown in the diagram (top) and their sample spread over time and frequency (bottom). Compared to the bottom sample, the annotated samples have a wider bandwidth but a shorter time spread. The analysis window (bottom) has a full resolution of one coefficient per original time sample. Therefore, for each time region of the subband sample (m=256:::384), the TDAR coefficients must be approximated (annotated by the dots).
[0177] Subsequently, the (simulation) results are described.
[0178] Figure 8 The spectral and temporal uncertainties obtained by several different transforms are shown, as shown in [Frederic Bimbot, Ewen Camberlein, and Pierrick Philippe, "Adaptive filter banks using fixed size mdct and subband merging for audio coding - comparison with the mpeg aac filterbanks", in Audio Engineering Society Convention 121, October 2016].
[0179] It can be seen that the Hadamard-based transform provides a tightly bounded time-frequency tradeoff. For increasing bin sizes, the additional time resolution comes at a disproportionately high cost in terms of spectral uncertainty.
[0180] In other words, Figure 8A comparison of the spectral and temporal energy compactness of different transforms is shown. The inline labels indicate the frame length for MDCT, the partitioning factor for Heisenberg partitioning, and the combining factors for all others.
[0181] However, similar to the simple uniform MDCT, subband merging with TDAR exhibits a linear tradeoff between temporal and spectral uncertainty. While slightly higher than the simple uniform MDCT, the product of the two is constant. For this analysis, the sinusoidal analysis window and the Kaiser-Bessel-derived subband merging window showed the most compact results and were therefore selected.
[0182] However, for the merging factor N v = 2, the use of TDAR appears to reduce temporal and spectral compactness. We attribute this to the fact that the coefficient calculation scheme presented in Section II-B is too simple and does not properly approximate the value of the steep window function slope. Numerical optimization schemes will be presented in subsequent publications.
[0183] Gear center of gravity and squared effective length using impulse response x[n] Calculate these compactness values as defined in [Athanasios Papoulis, Signal analysis, Electrical and electronic engineering series. McGraw-Hill, New York, San Francisco, Paris, 1977.]
[0184]
[0185]
[0186] The average of all impulse responses for each individual filter bank is shown.
[0187] Figure 9 A comparison of two exemplary impulse responses generated by subband combining with and without TDAR, simple MDCT short blocks, and Hadamard matrix subband combining as proposed in [OA Niamut and R. Heusdens, “Flexible frequency decompositions for cosine-modulated filterbanks,” in Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP'03). 2003 IEEE International Conference, April 2003, Vol. 5, pp. V–449–52].
[0188] The poor time compactness of the Hadamard matrix merging transform is clearly visible. It is also clear that most of the aliasing artifacts in the subband are significantly reduced by TDAR.
[0189] In other words, Figure 9 An exemplary impulse response of a merged subband filter comprising 8 of the 1024 original binary bins using the method without TDAR / with TDAR proposed in this article is shown, which is proposed in [OA Niamut and R. Heusdens, "Subband merging in cosine-modulated filter banks", SignalProcessing Letters, IEEE, Vol. 10, No. 4, pp. 111–114, April 2003] and uses a shorter MDCT frame length of 256 samples.
[0190] Figure 10 A flow chart of a method 300 for processing an audio signal to obtain a subband representation of the audio signal is shown. The method 300 comprises a step 302 of performing a cascaded overlapped critical sampling transform on at least two partially overlapping blocks of samples of the audio signal to obtain a subband sample set based on a first block of samples of the audio signal and a corresponding subband sample set based on a second block of samples of the audio signal. Furthermore, the method 300 comprises a step 304 of performing a weighted combination of the two corresponding subband sample sets to obtain an alias-reduced subband representation of the audio signal, wherein one of the two corresponding subband sample sets is obtained based on the first block of samples of the audio signal and the other of the two corresponding subband sample sets is obtained based on the second block of samples of the audio signal.
[0191] Figure 11 A flow chart of a method 400 for processing a subband representation of an audio signal to obtain an audio signal is shown. The method 400 comprises a step 402 of performing a weighted (and shifted) combination of two corresponding alias-reduced subband representations (in different, partially overlapping blocks of samples) of the audio signal to obtain an aliased subband representation, wherein the aliased subband representation is a set of subband samples. Furthermore, the method 400 comprises a step 404 of performing a concatenated inverse overlapped critically sampled transform on the set of subband samples to obtain a set of samples associated with the block of samples of the audio signal.
[0192] Figure 12A schematic block diagram of an audio encoder 150 according to an embodiment is shown. The audio encoder 150 comprises: an audio processor (100) as described above; an encoder 152 configured to encode an aliasing-reduced subband representation of an audio signal to obtain an encoded aliasing-reduced subband representation of the audio signal; and a bitstream shaper 154 configured to form a bitstream 156 from the encoded aliasing-reduced subband representation of the audio signal.
[0193] Figure 13 : A schematic block diagram of an audio decoder 250 according to an embodiment is shown. The audio decoder 250 comprises: a bitstream parser 252 configured to parse the bitstream 154 to obtain an encoded aliasing-reduced subband representation; a decoder 254 configured to decode the encoded aliasing-reduced subband representation to obtain an aliasing-reduced subband representation of an audio signal; and the audio processor 200 as described above.
[0194] Figure 14 : A schematic block diagram of an audio analyzer 180 according to an embodiment is shown. The audio analyzer 180 comprises: the audio processor 100 as described above; an information extractor 182 configured to analyze the aliasing reduced subband representation to provide information describing the audio signal.
[0195] Embodiments provide time domain aliasing reduction (TDAR) in subbands of a non-uniform orthogonal modified discrete cosine transform (MDCT) filter bank.
[0196] An embodiment adds an additional post-processing step to the widely used MDCT transform pipeline, which itself consists of only another overlapped MDCT transform along the frequency axis and time domain aliasing reduction (TDAR) along the time axis for each subband, allowing the extraction of arbitrary frequency scales from the MDCT spectrogram and improving the temporal compactness of the impulse response, while introducing no additional redundancy and only one MDCT frame delay.
[0197] 2. Time-varying time-frequency tiling using non-uniform orthogonal filter banks based on MDCT analysis / synthesis and TDAR
[0198] Figure 15 A schematic block diagram of an audio processor 100 configured to process an audio signal to obtain a sub-band representation of the audio signal according to another embodiment is shown. The audio processor 100 includes a cascaded overlapped critical sampling transform (LCST) stage 104 and a time domain aliasing reduction (TDAR) stage 106, both of which are described in detail in Section 1 above.
[0199] The cascaded overlapped critical sampling transform stage 104 includes a first overlapped critical sampling transform (LCST) stage 120, which is configured to perform LCST (e.g., MDCT) 122_1 and 122_2 on the first sample block 108_1 and the second sample block 108_2, respectively, to obtain a first set of binary bins 124_1 for the first sample block 108_1 and a second set of binary bins 124_2 for the second sample block 108_2. Furthermore, the cascaded overlapped critical sampling transform stage 104 includes a second overlapped critical sampling transform (LCST) stage 126 configured to perform an LCST (e.g., MDCT) 132_1,1-132_1,2 on the piecewise binary bin sets 128_1,1-128_1,2 of the first binary bin set 124_1 and to perform an LCST (e.g., MDCT) 132_2,1-132_2,2 on the piecewise binary bin sets 128_2,1-128_2,2 of the second binary bin set 124_1 to obtain a subband sample set 110_1,1-110_1,2 based on the first sample block 108_1 and a subband sample set 110_2,1-110_2,2 based on the second sample block 108_1.
[0200] As already pointed out in the introductory part, the time domain aliasing reduction (TDAR) stage 106 can only apply time domain aliasing reduction (TDAR) if the same time-frequency tiling is used for the first block of samples 108_1 and the second block of samples 108_2, i.e. if the set of subband samples 110_1,1-110_1,2 based on the first block of samples 108_1 represents the same area in the time-frequency plane as the set of subband samples 110_2,1-110_2,2 based on the second block of samples 108_2.
[0201] However, if the signal characteristics of the input signal change, the LCST (e.g., MDCT) 132_1, 1-132_1, 2 used to process the set of segmented binary bins 128_1, 1-128_1, 2 based on the first block of samples 108_1 may have a different frame length (e.g., merging factor) than the LCST (e.g., MDCT) 132_2, 1-132_2, 2 used to process the set of segmented binary bins 128_2, 1-128_2, 2 based on the second block of samples 108_2.
[0202] In this case, the subband sample sets 110_1,1-110_1,2 based on the first sample block 108_1 represent different regions in the time-frequency plane compared to the subband sample sets 110_2,1-110_2,2 based on the second sample block 108_2, that is, if the first subband sample set 110_1,1 represents a different region in the time-frequency plane compared to the third subband sample set 110_2,1, and the second subband sample set 110_1,2 represents a different region in the time-frequency plane compared to the fourth subband sample set 110_2,1, then time domain aliasing reduction (TDAR) cannot be directly applied.
[0203] In order to overcome this limitation, the audio processor 100 further includes a first time-frequency transform stage 105, which is configured to: identify one or more subband sample sets from the subband sample sets 110_1, 1-110_1, 2 based on the first sample block 108_1, and identify one or more subband sample sets from the subband sample sets 110_2, 1-110_2, 2 based on the second sample block 108_2, in the case where the subband sample sets 110_1, 1-110_1, 2 based on the first sample block 108_1 represent different regions in the time-frequency plane compared to the subband sample sets 110_2, 1-110_2, 2 based on the second sample block 108_2; The method further comprises: performing time-frequency transformation on the one or more subband sample sets identified in the subband sample sets 110_2, 1-110_2, 2 based on the second sample block 108_2 and / or the one or more subband sample sets identified in the subband sample sets 110_2, 1-110_2, 2 based on the second sample block 108_2 to obtain one or more time-frequency transformed subband samples, each of the time-frequency transformed subband samples representing the same region in the time-frequency plane as a corresponding one of the one or more identified subband samples or the time-frequency transformed versions of the one or more subband samples.
[0204] Thereafter, the time domain aliasing reduction stage 106 may apply time domain aliasing reduction (TDAR), i.e., by performing a weighted combination of two corresponding subband sample sets or time-frequency transformed versions thereof, wherein one subband sample set or a time-frequency transformed version thereof is obtained based on the first sample block 108_1 of the audio signal 102 and the other subband sample set or a time-frequency transformed version thereof is obtained based on the second sample block 108_2 of the audio signal, to obtain an aliasing-reduced subband representation of the audio signal 102.
[0205] In an embodiment, the first time-frequency transform stage 105 may be configured to perform a time-frequency transform on the identified one or more subband sample sets of the subband sample sets 110_2, 1 - 110_2, 2 based on the first sample block 108_1 or the identified one or more subband sample sets of the subband sample sets 110_2, 1 - 110_2, 2 based on the second sample block 108_2 to obtain one or more time-frequency transformed subband samples, each of the time-frequency transformed subband samples representing the same region in the time-frequency plane as a corresponding one of the one or more identified subband samples.
[0206] In this case, the time-domain aliasing reduction stage 106 may be configured to perform a weighted combination of a time-frequency transformed subband sample set and a corresponding (non-time-frequency transformed) subband sample set, wherein one subband sample set is obtained based on a first sample block 108_1 of the audio signal 102 and the other subband sample set is obtained based on a second sample block 108_2 of the audio signal. This is referred to herein as single-sided STDAR.
[0207] Naturally, the first time-frequency transform stage 105 may also be configured to perform time-frequency transform on the identified one or more subband sample sets in the subband sample sets 110_2, 1 - 110_2, 2 based on the first sample block 108_1 and the identified one or more subband sample sets in the subband sample sets 110_2, 1 - 110_2, 2 based on the second sample block 108_2, so as to obtain one or more time-frequency transformed subband samples, each of which represents the same region in the time-frequency plane as a corresponding one of the time-frequency transformed versions of the other identified one or more subband samples.
[0208] In this case, the time-domain aliasing reduction stage 106 may be configured to perform a weighted combination of two corresponding time-frequency transformed subband sample sets, one of which is obtained based on a first sample block 108_1 of the audio signal 102 and the other of which is obtained based on a second sample block 108_2 of the audio signal. This is referred to herein as bilateral STDAR.
[0209] Figure 16 A schematic representation of the time-frequency transform performed by the time-frequency transform stage 105 in the time-frequency plane is shown.
[0210] like Figure 16As indicated in diagrams 170_1 and 170_2 in FIG, the first subband sample set 110_1,1 corresponding to the first sample block 108_1 and the third subband sample set 110_2,1 corresponding to the second sample block 108_2 represent different regions 194_1,1 and 194_2,1 in the time-frequency plane, such that the time domain aliasing reduction stage 106 will not be able to apply time domain aliasing reduction (TDAR) to the first subband sample set 110_1,1 and the third subband sample set 110_2,1.
[0211] Similarly, the second subband sample set 110_1,2 corresponding to the first block of samples 108_1 and the fourth subband sample set 110_2,2 corresponding to the second block of samples 108_2 represent different regions 194_1,2 and 194_2,2 in the time-frequency plane, such that the time domain aliasing reduction stage 106 will not be able to apply time domain aliasing reduction (TDAR) to the second subband sample set 110_1,2 and the fourth subband sample set 110_2,2.
[0212] However, the combination of the first subband sample set 110_1,1 and the second subband sample set 110_1,2 and the combination of the third subband sample set 110_2,1 and the fourth subband sample set 110_2,2 represent the same region 196 in the time-frequency plane.
[0213] Therefore, the time-frequency transform stage 105 may perform time-frequency transform on the first subband sample set 110_1,1 and the second subband sample set 110_1,2, or perform time-frequency transform on the third subband sample set 110_2,1 and the fourth subband sample set 110_2,2, to obtain time-frequency transformed subband sample sets, each of which represents the same region in the time-frequency plane as a corresponding one of the other subband sample sets.
[0214] exist Figure 16 In the example, it is assumed that the time-frequency transform stage 105 performs time-frequency transform on the first subband sample set 110_1,1 and the second subband sample set 110_1,2 to obtain a first time-frequency transformed subband sample set 110_1,1′ and a second time-frequency transformed subband sample set 110_1,2′.
[0215] like Figure 16 As indicated in diagrams 170_3 and 170_4 in the figure, the first time-frequency transformed subband sample set 110_1,1′ and the third subband sample set 110_2,1 represent the same area 194_1,1′ and 194_2,1 in the time-frequency plane, so that time domain aliasing reduction (TDAR) can be applied to the first time-frequency transformed subband sample set 110_1,1′ and the third subband sample set 110_2,1.
[0216] Similarly, the second time-frequency transformed subband sample set 110_1,2' and the fourth subband sample set 110_2,2 represent the same regions 194_1,2' and 194_2,3 in the time-frequency plane, such that time domain aliasing reduction (TDAR) can be applied to the second time-frequency transformed subband sample set 110_1,2' and the fourth subband sample set 110_2,2.
[0217] Although in the above description it is assumed that the first time-frequency transform stage 105 is configured to perform a time-frequency transform on the first subband sample set 110_1,1 and the second subband sample set 110_1,2 corresponding to the first sample block 108_1, in an embodiment, the first time-frequency transform stage 105 can also be configured to perform a time-frequency transform on the first subband sample set 110_1,1 and the second subband sample set 110_1,2 corresponding to the first sample block 108_1 and on the third subband sample set 110_2,1 and the fourth subband sample set 110_2,2 corresponding to the second sample block 108_2. Figure 16
[0218] Figure 17 A schematic block diagram of an audio processor 100 configured to process an audio signal to obtain a subband representation of the audio signal is shown according to another embodiment.
[0219] As shown in Figure 17 , the audio processor 100 can further comprise a second time-frequency transform stage 107 configured to perform a time-frequency transform on the aliasing-reduced subband representation of the audio signal, wherein the time-frequency transform applied by the second time-frequency transform stage is inverse to the time-frequency transform applied by the first time-frequency transform stage.
[0220] Figure 18 A schematic block diagram of an audio processor 200 for processing a subband representation of an audio signal to obtain the audio signal is shown according to another embodiment.
[0221] The audio processor 200 comprises a second inverse time-frequency transform stage 201 which is inverse to the second time-frequency transform stage 107 of the audio processor 100 shown in Figure 17 . In detail, the second inverse time-frequency transform stage can be configured to perform a time-frequency transform on one or more of the aliasing-reduced subband sample sets corresponding to the second sample block of the audio signal and / or on one or more time- frequency transformed aliasing-reduced subband samples of the aliasing-reduced subband sample sets corresponding to the second sample block of the audio signal to obtain one or more time- frequency transformed aliasing-reduced subband samples each representing the same region in the time-frequency plane having the same length as a corresponding one of the one or more aliasing-reduced subband samples corresponding to another sample block of the audio signal or a time- frequency transformed version of the one or more aliasing-reduced subband samples.
[0222] Furthermore, the audio processor 200 comprises an inverse time domain aliasing reduction (ITDAR) stage 202 configured to perform weighted combining of corresponding aliased reduced subband sample sets or time-frequency transformed versions thereof to obtain aliased subband representations.
[0223] Furthermore, the audio processor 200 comprises a first inverse time-frequency transform stage 203 configured to perform a time-frequency transform on the aliased subband representations to obtain a subband sample set 110_1,1-110_1,2 corresponding to a first sample block 108_1 of the audio signal and a subband sample set 110_2,1-110_2,2 corresponding to a second sample block 108_1 of the audio signal, wherein the time-frequency transform applied by the first inverse time-frequency transform stage 203 is the opposite of the time-frequency transform applied by the second inverse time-frequency transform stage 201.
[0224] Furthermore, the audio processor 200 comprises a cascaded inverse overlapped critically sampled transform stage 204 configured to perform a cascaded inverse overlapped critically sampled transform on the sample sets 110_1,1-110_2,2 to obtain a sample set 206_1,1 associated with the sample block of the audio signal 102.
[0225] The embodiments of the present invention are described in further detail below.
[0226] 2.1 Time Domain Aliasing Reduction
[0227] When representing the lapped transform using polyphase representation, the frame index can be expressed in the z domain, where z -1 Refer to the previous frame [7]. In this representation, the MDCT analysis can be expressed as
[0228]
[0229] where D is the N×N DCT-IV matrix and F(z) is the N×N MDCT pre-permutation / folding matrix [7].
[0230] Then the subband merge M and TDARR(z) become another pair of block diagonal transform matrices
[0231]
[0232]
[0233] Where T k is a suitable transform matrix (in some embodiments an overlapped MDCT), F′(z) k is an improved and smaller variant of F(z)[4]. Contains the submatrix Tk and F′(z) k A vector of the size It is called sub-band layout. The overall analysis becomes
[0234]
[0235] For simplicity, we only analyze the special case of uniform tiling in M and R(z), i.e., Where c∈{1,2,4,8,16,32}, it is easy to see that the embodiments are not limited to these.
[0236] 2.2 Switching time domain aliasing reduction
[0237] Since STDAR will be applied between two different transform frames, in an embodiment, the subband merging matrix M, the TDAR matrix R(z), and the subband layout is expanded to the time-varying representations M(m), R(z,m) and where m is the frame index [8].
[0238]
[0239]
[0240] Of course, STDAR can also be extended to time-varying matrices F(z,m) and D(m), but this case will not be considered here.
[0241] If the tiling of two frames m and m-1 is different, that is,
[0242]
[0243] Then an additional transformation matrix S(m) can be designed that temporarily transforms the time-frequency tile of frame m to match the tile of frame m-1 (backward matching). Figure 19 An overview of STDAR operation can be seen in .
[0244] In detail, Figure 19 FIG shows a schematic representation of the STDAR operation in the time-frequency plane. Figure 19As shown, the subband sample set 110_1,1-110_1,4 corresponding to the first sample block 108_1 (frame m-1) and the subband sample set 110_2,1-110_2,4 corresponding to the second sample block 108_2 (frame m) represent different regions in the time-frequency plane. Therefore, the subband sample set 110_1,1-110_1,4 corresponding to the first sample block 108_1 (frame m-1) can be time-frequency transformed to obtain the time-frequency transformed subband sample set 110_1,1'-110_1,4' corresponding to the first sample block 108_1 (frame m-1), each of which represents the same region in the time-frequency plane as the corresponding one of the subband sample sets 110_2,1-110_2,4 corresponding to the second sample block 108_2 (frame m), so that it can be shown as follows Figure 19 TDAR(R(z,m)) is applied as shown. Afterwards, an inverse time-frequency transform may be applied to obtain an aliasing-reduced subband sample set 112_1,1-112_1,4 corresponding to the first block of samples 108_1 (frame m-1) and an aliasing-reduced subband sample set 112_2,1-112_2,4 corresponding to the second block of samples 108_2 (frame m).
[0245] In other words, Figure 19 TDAR using forward up-matching is shown. The time-frequency tile of the relevant half of frame m-1 is altered to match the time-frequency tile of the relevant half of frame m, after which TDAR can be applied and the original tile reconstructed. The tile of frame m is unchanged, as indicated by the identity matrix I.
[0246] Naturally, we can also transform frame m-1 to match the time-frequency tile of frame m (forward matching). In this case, we can consider S(m-1) instead of S(m). Both forward matching and backward matching are symmetric, so we only consider one of the two operations.
[0247] If, through this operation, the temporal resolution is improved by the subband merging step, it is referred to as up-matching in this paper. If the temporal resolution is reduced by the subband splitting step, it is referred to as down-matching in this paper. Both up-matching and down-matching are evaluated in this paper.
[0248] This matrix S(m) is also block diagonal, however with κ≠K, and will be applied before TDAR and then inverted afterwards.
[0249]
[0250] Therefore, the analysis becomes:
[0251]
[0252] Naturally, only half of each frame is affected by the TDAR between the two frames, so only half of the corresponding frame needs to be transformed. Therefore, half of S(m) can be chosen as the identity matrix.
[0253] 2.3 Additional considerations
[0254] Obviously, the order of the impulse responses (ie, row order) of each transform matrix needs to match the order of its neighboring matrices.
[0255] In the case of conventional TDAR, no special consideration is required because the order of two adjacent identical frames is always equal. However, depending on the choice of parameters, when STDAR is introduced, the input ordering of STDARS(m) may be incompatible with the output ordering of the subband merge M. In this case, two or more coefficients that are not adjacent in memory are jointly transformed and therefore need to be realigned before the operation.
[0256] Furthermore, the output ordering of STDARS(m) is generally incompatible with the input ordering of TDARR(z,m) as originally defined. Again, this is because the coefficients of a subband are not adjacent in memory.
[0257] Both reordering and unordering can be represented as additional permutation matrices P and P -1 , which are introduced into the transformation pipeline at the appropriate location.
[0258] The order of the coefficients in these matrices depends on the operation, memory layout, and transformations used. Therefore, no general solution can be provided here.
[0259] All introduced matrices are orthogonal, so the overall transformation remains orthogonal.
[0260] 2.4 Evaluation
[0261] In the evaluation, DCT-IV and DCT-II are considered for T(m) in S(m), both used without overlap. An input frame length of N=1024 is chosen as an example. Thus, the system is analyzed for different switching ratios r(m), which is the ratio of the merging factors between two frames, i.e.
[0262]
[0263] As in the analysis of TDAR, investigations have focused on the shape, and in particular on the compactness of the impulse response and frequency response of the entire transform [4], [9].
[0264] 2.5 Results
[0265] DCT-II produces the best results, so the following focuses on this transform. Forward matching and backward matching are symmetric and produce the same results, so only the results of forward matching are described.
[0266] Figure 20 Example impulse responses for two frames with binning factors of 8 and 16 are shown graphically before (top) and after (bottom) STDAR.
[0267] In other words, Figure 20 Two example impulse responses for two frames with different time-frequency tiling, before and after STDAR, are shown. The impulse responses exhibit different widths because their binning factors are different—c(m-1)=8 and c(m)=16. After STDAR, the aliasing is significantly reduced, but some residual aliasing is still visible.
[0268] Figure 21 The impulse response and frequency response compactness for the upper matching are shown in the graph. The labels in the lines indicate the frame length for uniform MDCT, the merging factor for TDAR and the merging factor for STDAR for frames m-1 and m. Thus, in Figure 21 , a first curve 500 represents TDAR, a second curve 502 represents no TDAR, a third curve 504 represents STDAR with c(m)=4, a fourth curve 506 represents STDAR with c(m)=8, a fifth curve 508 represents STDAR with c(m)=16, a sixth curve 510 represents STDAR with c(m)=32, a seventh curve 512 represents MDCT, and an eighth curve 514 represents the Heisenberg boundary.
[0269] Figure 22 The impulse response and frequency response compactness for the down-matching are shown in the graph. The labels in the lines indicate the frame length for uniform MDCT, the merging factor for TDAR, and the merging factor for STDAR for frames m-1 and m. Thus, in Figure 21 , a first curve 500 represents TDAR, a second curve 502 represents no TDAR, a third curve 504 represents STDAR with c(m)=4, a fourth curve 506 represents STDAR with c(m)=8, a fifth curve 508 represents STDAR with c(m)=16, a sixth curve 510 represents STDAR with c(m)=32, a seventh curve 512 represents MDCT, and an eighth curve 514 represents the Heisenberg boundary.
[0270] Therefore, in Figure 21 and Figure 22 , are the average impulse response compactness of various filter banks for up-matching and down-matching, respectively. and frequency response compactness [3], [9]. For baseline comparison, uniform MDCT is shown using curves 512, 500, and 502, as well as subband merging with and without TDAR [3], [4]. STDAR filter banks are shown using curves 504, 506, 508, and 510. Each line represents all filter banks with the same merging factor c. The labels within the lines for each data point indicate the merging factor for frames m-1 and m.
[0271] exist Figure 21 In
[15] , frame m-1 is transformed to match the tile of frame m. It can be seen that the temporal compactness of frame m is improved without losing spectral compactness. For the compactness of frame m-1, it can be seen that it improves for all binning factors c>2, while there is a fallback for binning factor c=2. This fallback is expected since the original TDAR with c=2 already leads to a deterioration in impulse response compactness [4].
[0272] exist Figure 22 A similar situation can be seen in . Again, frame m-1 is transformed to match the tile of frame m. In this case, the temporal compactness of frame m-1 is improved without loss in spectral compactness. Again, the merging factor c = 2 is still problematic.
[0273] Overall, it can be clearly seen that STDAR reduces the impulse response width by reducing aliasing for combining factors c > 2. Among all combining factors, the smallest switching factor r has the best compactness.
[0274] 2.6 Further Examples
[0275] Although the above embodiments primarily refer to single-sided STDAR, where the STDAR operation changes the time-frequency tiling of only one of the two frames to match the other, it should be noted that the present invention is not limited to such embodiments. Instead, embodiments may also employ double-sided STDAR, where the STDAR operation changes the time-frequency tiling of both frames to ultimately match each other. Such a system can be used to improve system compactness for very high switching ratios, i.e., instead of changing one frame from one extreme tiling to another (32 / 2 → 2 / 2), both frames can be changed to an intermediate base tiling of 32 / 2 → 8 / 8.
[0276] Furthermore, it is possible to numerically optimize the coefficients in R(z,m) and S(m) as long as orthogonality is not violated. This can improve the performance of STDAR at lower merging factors c or higher switching ratios r.
[0277] Time Domain Aliasing Reduction (TDAR) is a method for improving the compactness of the impulse response of the non-uniform orthogonal modified cosine transform (MDCT). Traditionally, TDAR can only be performed between frames of the same time-frequency tile, but the embodiments described herein overcome this limitation. The embodiments enable the use of TDAR between two consecutive frames with different time-frequency tile by introducing another subband merging or subband splitting step. Furthermore, the embodiments allow more flexible and adaptive filter bank tile while still maintaining a compact impulse response, which are two properties required for efficient perceptual audio coding.
[0278] An embodiment provides a method for applying time domain aliasing reduction (TDAR) between two frames of different time-frequency tiles. Previously, it was not possible to perform TDAR between such frames, which resulted in suboptimal impulse response compactness when the time-frequency tiles had to be adaptively changed.
[0279] An embodiment introduces another subband merging / subband splitting step to allow matching of the time-frequency tiles of the two frames before applying TDAR.After TDAR, the original time-frequency tiles can be reconstructed.
[0280] The embodiment provides two scenarios. The first is up-matching, in which the temporal resolution of one frame is increased to match the temporal resolution of another frame. The second is down-matching, which is the opposite situation.
[0281] Figure 23A flow chart of a method 320 for processing an audio signal to obtain a subband representation of the audio signal is shown. The method comprises step 322 of performing a cascaded overlapped critical sampling transform on at least two partially overlapping blocks of samples of the audio signal to obtain a subband sample set based on a first block of samples of the audio signal and a subband sample set based on a second block of samples of the audio signal. Furthermore, the method 320 comprises step 324 of identifying one or more subband sample sets from the subband sample set based on the first block of samples, and identifying one or more subband sample sets from the subband sample set based on the second block of samples, if the subband sample set based on the first block of samples represents a different region in a time-frequency plane than the subband sample set based on the second block of samples, the identified one or more subband sample sets being combined to represent the same region in the time-frequency plane. Furthermore, the method 320 comprises step 326 of performing a time-frequency transform on the identified one or more subband sample sets of the subband sample sets based on the first sample block and / or the identified one or more subband sample sets of the subband sample sets based on the second sample block to obtain one or more time-frequency transformed subband samples, wherein each of the time-frequency transformed subband samples represents the same region in the time-frequency plane as a corresponding one of the identified one or more subband samples or the time-frequency transformed versions of the one or more subband samples. Furthermore, the method 320 comprises step 328 of performing a weighted combination of two corresponding subband sample sets or the time-frequency transformed versions thereof to obtain an aliasing-reduced subband representation of the audio signal, wherein one of the two corresponding subband sample sets or the time-frequency transformed versions thereof is obtained based on the first sample block of the audio signal and the other of the two corresponding subband sample sets or the time-frequency transformed versions thereof is obtained based on the second sample block of the audio signal.
[0282] Figure 24A flow chart of a method 420 for processing a subband representation of an audio signal comprising an alias-reduced sample set is shown. The method 420 comprises a step 422 of performing a time-frequency transform on one or more of the alias-reduced subband sample sets corresponding to a second block of samples of the audio signal and / or one or more of the alias-reduced subband sample sets corresponding to the second block of samples of the audio signal to obtain one or more time-frequency transformed alias-reduced subband samples, each of the time-frequency transformed alias-reduced subband samples representing the same region in the time-frequency plane as a corresponding one of the one or more alias-reduced subband samples or time-frequency transformed versions of the alias-reduced subband samples corresponding to another block of samples of the audio signal. Furthermore, the method 420 comprises a step 424 of performing a weighted combination of the corresponding alias-reduced subband sample sets or time-frequency transformed versions thereof to obtain the aliased subband representation. Furthermore, the method 420 includes step 426 of performing a time-frequency transform on the aliased subband representation to obtain a subband sample set corresponding to a first sample block of the audio signal and a subband sample set corresponding to a second sample block of the audio signal, wherein the time-frequency transform applied by the first inverse time-frequency transform stage is the inverse of the time-frequency transform applied by the second inverse time-frequency transform stage. Furthermore, the method 420 includes step 428 of performing a cascaded inverse overlapped critical sampling transform on the sample set to obtain a sample set associated with the sample block of the audio signal.
[0283] Further embodiments are described below. Thus, the following embodiments can be combined with the above embodiments.
[0284] Embodiment 1: An audio processor (100) for processing an audio signal (102) to obtain a subband representation of the audio signal (102), the audio processor (100) comprising: a cascaded overlapped critical sampling transform stage (104) configured to perform a cascaded overlapped critical sampling transform on at least two partially overlapping sample blocks (108_1; 108_2) of the audio signal (102) to obtain a subband sample set (110_1, 1) based on a first sample block (108_1) of the audio signal (102), and a subband sample set (110_1, 1) based on a second sample block (108_2) of the audio signal (102). 2) to obtain a corresponding subband sample set (110_2, 1); and a time-domain aliasing reduction stage (106) configured to perform a weighted combination of the two corresponding subband sample sets (110_1, 1; 110_1, 2) to obtain an aliasing-reduced subband representation (112_1) of the audio signal (102), wherein one of the two corresponding subband sample sets is obtained based on a first sample block (108_1) of the audio signal (102) and the other of the two corresponding subband sample sets is obtained based on a second sample block (108_2) of the audio signal.
[0285] Embodiment 2: The audio processor (100) according to embodiment 1, wherein the cascaded overlapped critical sampling transform stage (104) comprises: a first overlapped critical sampling transform stage (120) configured to perform an overlapped critical sampling transform on a first sample block (108_1) and a second sample block (108_2) of at least two partially overlapping sample blocks (108_1; 108_2) of the audio signal (102) to obtain a first binary bin set (124_1) for the first sample block (108_1) and a second binary bin set (124_2) for the second sample block (108_2).
[0286] Embodiment 3: The audio processor (100) according to embodiment 2, wherein the cascaded overlapped critical sampling transform stage (104) further comprises: a second overlapped critical sampling transform stage (126) configured to perform overlapped critical sampling transform on the segments (128_1, 1) of the first binary bin set (124_1) and to perform overlapped critical sampling transform on the segments (128_2, 1) of the second binary bin set (124_2) to obtain a subband sample set (110_1, 1) for the first binary bin set and a subband sample set (110_2, 1) for the second binary bin set, wherein each segment is associated with a subband of the audio signal (102).
[0287] Embodiment 4: The audio processor (100) according to embodiment 3, wherein the first subband sample set (110_1, 1) is a result of performing a first overlapping critical sampling transform (132_1, 1) based on the first segment (128_1, 1) of the first binary bin set (124_1), wherein the second subband sample set (110_1, 2) is a result of performing a second overlapping critical sampling transform (132_1, 2) based on the second segment (128_1, 2) of the first binary bin set (124_1), wherein the third subband sample set (110_2, 1) is a result of performing a third overlapping critical sampling transform (132_2, 1) based on the first segment (128_2, 1) of the second binary bin set (128_2, ... first subband sample set (110_1, 2) is a result of performing a second overlapping critical sampling transform (132_2, 1) based on the first segment (128_2, 1) of the second binary bin set (128_2, 1), wherein the second subband sample set (110_1, 2) is a result of performing a second overlapping critical sampling transform (132_2, 1) based on the first segment (128_2, 1) of the second binary bin set (128_2, 1), wherein the third subband sample set (110_2, 1) is a result of performing a third overlapping critical sampling transform (132_2, 1) based on the first segment (128_2, 1) of the second binary bin set (128_2, 1), wherein the third subband sample set (110_2, 1 The four subband sample sets (110_2, 2) are the result of a fourth overlapped critical sampling transform (132_2, 2) based on the second segment (128_2, 2) of the second binary bin set (128_2, 1); and wherein the time-domain aliasing reduction stage (106) is configured to perform a weighted combination of the first subband sample set (110_1, 1) and the third subband sample set (110_2, 1) to obtain a first alias-reduced subband representation (112_1) of the audio signal; wherein the time-domain aliasing reduction stage (106) is configured to perform a weighted combination of the second subband sample set (110_1, 2) and the fourth subband sample set (110_2, 2) to obtain a second alias-reduced subband representation (112_2) of the audio signal.
[0288] Embodiment 5: The audio processor (100) according to one of embodiments 1 to 4, wherein the cascaded overlapped critical sampling transform stage (104) is configured to segment the binary bin set (124_1) obtained based on the first sample block (108_1) using at least two window functions, and obtain at least two segmented subband sample sets (128_1, 1; 128_1, 2) based on the segmented binary bin set corresponding to the first sample block (108_1); wherein the cascaded overlapped critical sampling transform stage (104) is configured to segment the binary bin set (124_2) obtained based on the second sample block (108_2) using at least two window functions, and obtain at least two segmented subband sample sets (128_2, 1; 128_2, 2) based on the segmented binary bin set corresponding to the second sample block (108_2); and wherein the at least two window functions include different window widths.
[0289] Embodiment 6: The audio processor (100) according to one of embodiments 1 to 5, wherein the cascaded overlapped critical sampling transform stage (104) is configured to segment the binary bin set (124_1) obtained based on the first sample block (108_1) using at least two window functions, and obtain at least two segmented subband sample sets (128_1, 1; 128_1, 2) based on the segmented binary bin set corresponding to the first sample block (108_1); wherein the cascaded overlapped critical sampling transform stage (104) is configured to segment the binary bin set (124_2) obtained based on the second sample block (108_2) using at least two window functions, and obtain at least two subband sample sets (128_2, 1; 128_2, 2) based on the segmented binary bin set corresponding to the second sample block (108_2); and wherein the filter slopes of the window functions corresponding to adjacent subband sample sets are symmetrical.
[0290] Embodiment 7: An audio processor (100) according to one of embodiments 1 to 6, wherein the cascaded overlapped critical sampling transform stage (104) is configured to segment samples of the audio signal into a first sample block (108_1) and a second sample block (108_2) using a first window function; wherein the overlapped critical sampling transform stage (104) is configured to segment a binary bin set (124_1) obtained based on the first sample block (108_1) and a binary bin set (124_2) obtained based on the second sample block using a second window function to obtain corresponding subband samples; and wherein the first window function and the second window function include different window widths.
[0291] Embodiment 8: An audio processor (100) according to one of embodiments 1 to 6, wherein the cascaded overlapped critical sampling transform stage (104) is configured to segment samples of the audio signal into a first sample block (108_1) and a second sample block (108_2) using a first window function; wherein the overlapped critical sampling transform stage (104) is configured to segment a binary bin set (124_1) obtained based on the first sample block (108_1) and a binary bin set (124_2) obtained based on the second sample block (108_2) using a second window function to obtain corresponding subband samples; and wherein a window width of the first window function and a window width of the second window function are different from each other, wherein the window width of the first window function and the window width of the second window function are different from each other by a factor that is different from a power of two.
[0292] Embodiment 9: The audio processor (100) according to one of embodiments 1 to 8, wherein the time domain aliasing reduction stage (106) is configured to perform a weighted combination of two corresponding subband sample sets according to the following equation:
[0293] For 0≤m <N / 2
[0294]
[0295] in
[0296]
[0297] To obtain the aliasing-reduced subband representation of the audio signal, where y v,i (m) is the first aliasing-reduced subband representation of the audio signal, y v,i-1 (N-1-m) is the second alias-cancelled subband representation of the audio signal, is a set of subband samples based on a second block of samples of the audio signal, is a subband sample set based on the first sample block of the audio signal, a v (m) is..., b v (m) is..., c v (m) is... and d v (m) is...
[0298] Embodiment 10: An audio processor (200) for processing subband representations of an audio signal to obtain an audio signal (102), the audio processor (200) comprising: an inverse time-domain aliasing reduction stage (202) configured to perform a weighted combination of two corresponding aliasing-reduced subband representations of the audio signal (102) to obtain an aliased subband representation, wherein the aliased subband representation is a subband sample set (110_1, 1); and a cascaded inverse overlapped critical sampling transform stage (204) configured to perform a cascaded inverse overlapped critical sampling transform on the subband sample set (110_1, 1) to obtain a sample set (206_1, 1) associated with a sample block of the audio signal (102).
[0299] Embodiment 11: An audio processor (200) according to embodiment 10, wherein the cascaded inverse overlapped critical sampling transform stage (204) comprises: a first inverse overlapped critical sampling transform stage (208) configured to perform an inverse overlapped critical sampling transform on a subband sample set (110_1, 1) to obtain a binary bin set (128_1, 1) associated with a given subband of the audio signal; and a first overlap and add stage (210) configured to perform a concatenation of the binary bin sets associated with a plurality of subbands of the audio signal, comprising a weighted combination of the binary bin set (128_1, 1) associated with a given subband of the audio signal (102) and a binary bin set (128_1, 2) associated with another subband of the audio signal (102) to obtain a binary bin set (124_1) associated with a sample block of the audio signal (102).
[0300] Embodiment 12: The audio processor (200) according to embodiment 11, wherein the cascaded inverse overlapped critical sampling transform stage (204) comprises: a second inverse overlapped critical sampling transform stage (212) configured to perform an inverse overlapped critical sampling transform on a binary bin set (124_1) associated with a sample block of the audio signal (102) to obtain a sample set associated with the sample block of the audio signal (102).
[0301] Embodiment 13: The audio processor (200) according to embodiment 12, wherein the cascaded inverse overlapped critical sampling transform stage (204) comprises: a second overlap and add stage (214) configured to overlap and add a set of samples (206_1, 1) associated with a block of samples of the audio signal (102) and another set of samples (206_2, 1) associated with another block of samples of the audio signal (102) to obtain the audio signal (102), wherein the block of samples of the audio signal (102) and the another block of samples partially overlap.
[0302] Embodiment 14: The audio processor (200) according to one of embodiments 10 to 13, wherein the inverse time-domain aliasing reduction stage (202) is configured to perform a weighted combination of two corresponding aliasing-reduced subband samples of the audio signal (102) based on the following equation
[0303] For 0≤m <N / 2
[0304]
[0305] in
[0306]
[0307] The aliased subband representation is obtained, where y v,i (m) is the first aliasing-reduced subband representation of the audio signal, y v,i-1 (N-1-m) is the second alias-cancelled subband representation of the audio signal, is a set of subband samples based on a second block of samples of the audio signal, is a subband sample set based on the first sample block of the audio signal, a v (m) is..., b v (m) is..., c v (m) is... and d v (m) is...
[0308] Embodiment 15: An audio encoder comprising: an audio processor (100) according to one of embodiments 1 to 9; an encoder configured to encode an aliasing-reduced subband representation of an audio signal to obtain an encoded aliasing-reduced subband representation of the audio signal; and a bitstream former configured to form a bitstream based on the encoded aliasing-reduced subband representation of the audio signal.
[0309] Embodiment 16: An audio decoder comprising: a bitstream parser configured to parse a bitstream to obtain an encoded aliasing-reduced subband representation; a decoder configured to decode the encoded aliasing-reduced subband representation to obtain an aliasing-reduced subband representation of an audio signal; and an audio processor (200) according to one of embodiments 10 to 14.
[0310] Embodiment 17: An audio analyzer comprising: an audio processor (100) according to one of embodiments 1 to 9; and an information extractor configured to analyze the aliasing-reduced subband representation to provide information describing the audio signal.
[0311] Embodiment 18: A method (300) for processing an audio signal to obtain a subband representation of the audio signal, the method comprising: performing (302) a cascaded overlapped critical sampling transform on at least two partially overlapping sample blocks of the audio signal to obtain a subband sample set based on a first sample block of the audio signal and to obtain corresponding subband samples based on a second sample block of the audio signal; and performing (304) a weighted combination of the two corresponding subband sample sets to obtain an aliasing-reduced subband representation of the audio signal, wherein one of the two corresponding subband sample sets is obtained based on the first sample block of the audio signal and the other of the two corresponding subband sample sets is obtained based on the second sample block of the audio signal.
[0312] Embodiment 19: A method (400) for processing a subband representation of an audio signal to obtain an audio signal, the method comprising: performing (402) a weighted combination of two corresponding aliasing-reduced subband representations of the audio signal to obtain an aliased subband representation, wherein the aliased subband representation is a subband sample set; and performing (404) a cascaded inverse overlapped critical sampling transform on the subband sample set to obtain a sample set associated with a sample block of the audio signal.
[0313] Embodiment 20: A computer program for performing the method according to one of embodiments 18 and 19.
[0314] Although some aspects have been described in the context of an apparatus, it will be clear that these aspects also represent descriptions of corresponding methods, wherein a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent descriptions of the features of the corresponding block or item or corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device (such as a microprocessor, a programmable computer, or an electronic circuit). In some embodiments, one or more of the most important method steps may be performed by such a device.
[0315] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or in software. Implementation may be performed using a digital storage medium (e.g., a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) having stored thereon electronically readable control signals that cooperate (or are capable of cooperating) with a programmable computer system to perform the corresponding method. Thus, the digital storage medium may be computer-readable.
[0316] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0317] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer.The program code may, for example, be stored on a machine-readable carrier.
[0318] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0319] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0320] A further embodiment of the inventive method is therefore a data carrier (or a digital storage medium or a computer-readable medium) having recorded thereon the computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory.
[0321] Therefore, another embodiment of the method of the present invention is a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can, for example, be configured to be transmitted via a data communication connection (for example, via the Internet).
[0322] A further embodiment comprises a processing means, such as a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0323] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0324] Another embodiment according to the present invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, a mobile device, a storage device, etc. For example, the apparatus or system may include a file server for transmitting the computer program to the receiver.
[0325] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can collaborate with a microprocessor to perform one of the methods described herein. Typically, the method is preferably performed by any hardware device.
[0326] The devices described herein may be implemented using hardware devices, or using computers, or using a combination of hardware devices and computers.
[0327] The apparatus described herein or any component of an apparatus described herein may be implemented at least partially in hardware and / or software.
[0328] The methods described herein may be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0329] Any component of a method described herein or an apparatus described herein may be performed at least in part by hardware and / or software.
[0330] The above-described embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. Accordingly, it is intended that the present invention be limited solely by the scope of the appended patent claims and not by the specific details presented in the description and explanation of the embodiments herein.
[0331] References
[0332] [1] HSMalvar, "Biorthogonal and nonuniform lapped transforms for transform coding with reduced blocking and ringing artifacts," IEEETransactions on Signal Processing, vol.46, no.4, pp.1043–1053, Apr.1998.
[0333] [2]OANiamut and R.Heusdens, "Subband merging in cosine-modulated filter banks," IEEE Signal Processing Letters, vol.10, no.4, pp.111–114, Apr.2003.
[0334] [3]Frederic Bimbot,Ewen Camberlein,and Pierrick Philippe,“AdaptiveFilter Banks using Fixed Size MDCT and Subband Merging for Audio Coding-Comparison with the MPEG AAC Filter Banks,”in Audio Engineering SocietyConvention 121.Oct.2006,Audio Engineering Society.
[0335] [4]N.Werner and B.Edler,“Nonuniform Orthogonal Filterbanks Based onMDCT Analysis / Synthesis and Time-Domain Aliasing Reduction,”IEEE SignalProcessing Letters,vol.24,no.5,pp.589–593,May 2017.
[0336] [5]Nils Werner and Bernd Edler,“Perceptual Audio Coding with AdaptiveNon-Uniform Time / Frequency Tilings using Subband Merging and Time DomainAliasing Reduction,”in 2019IEEE International Conference on Acoustics,Speechand Signal Processing,2019.
[0337] [6]B.Edler,“Codierung von Audiosignalen mit¨uberlappenderTransformation und adaptiven Fensterfunktionen,”Frequenz,vol.43,pp.252–256,Sept.1989.
[0338] [7]G.D.T.Schuller and M.J.T.Smith,“New framework for modulatedperfect reconstruction filter banks,”IEEE Transactions on Signal Processing,vol.44,no.8,pp.1941–1954,Aug.1996.
[0339] [8]Gerald Schuller,“Time-Varying Filter Banks With Variable SystemDelay,”in In IEEE International Conference on Acoustics,Speech,and SignalProecessing(ICASSP,1997,pp.21–24.
[0340] [9]Carl Taswell,“Empirical Tests for Evaluation of MultirateFilterBank Parameters,”in Wavelets in Signal and Image Analysis,MaxA.Viergever,Arthur A.Petrosian,and Franc,ois G.Meyer,Eds.,vol.19,pp.111–139.SpringerNetherlands,Dordrecht,2001.
[0341]
[10] F.Schuh,S.Dick,R.Füg,C.R.Helmrich,N.Rettelbach,andT.Schwegler,“Efficient Multichannel Audio Tranform Coding with LowDelay and Complexity.”Audio Engineering Society,Sep.2016.[Online].Available:http: / / www.aes.org / e-lib / browse.cfm?elib=18464。
Claims
1. An audio processor (100) for processing an audio signal (102) to obtain a subband representation of the audio signal (102), the audio processor (100) comprising: a cascaded overlapped critical sampling transform stage (104) configured to perform a cascaded overlapped critical sampling transform on at least two partially overlapping blocks of samples (108_1; 108_2) of the audio signal (102) to obtain a subband sample set (110_1, 1; 110_1, 2) based on a first block of samples (108_1) of the audio signal (102) and to obtain a subband sample set (110_2, 1; 110_2, 2) based on a second block of samples (108_2) of the audio signal (102); A first time-frequency transform stage (105) is configured to, in the case where the subband sample set (110_1, 1; 110_1, 2) based on the first sample block (108_1) represents a different region in the time-frequency plane than the subband sample set (110_2, 1; 110_2, 2) based on the second sample block (108_2), identify one or more subband sample sets from the subband sample set (110_1, 1; 110_1, 2) based on the first sample block (108_1), and identify one or more subband sample sets from the subband sample set (110_2, 1; 110_2, 2) based on the second sample block (108_2), the identified one or more subband sample sets being the same as the subband sample set (110_2, 1; 110_2, 2) based on the second sample block (108_2). The subband sample sets are combined to represent a same region in the time-frequency plane, and time-frequency transform is performed on the identified one or more subband sample sets in the subband sample set (110_2, 1; 110_2, 2) based on the first sample block (108_1) and / or the identified one or more subband sample sets in the subband sample set (110_2, 1; 110_2, 2) based on the second sample block (108_2) to obtain one or more time-frequency transformed subband samples, each of the time-frequency transformed subband samples representing a same region in the time-frequency plane as compared with a corresponding one of the identified one or more subband samples or a time-frequency transformed version of the identified one or more subband samples; as well as A time-domain aliasing reduction stage (106) is configured to perform a weighted combination of two corresponding subband sample sets or time-frequency transformed versions thereof to obtain aliasing-reduced subband representations (112_1-112_2) of the audio signal (102), wherein one of the two corresponding subband sample sets or the time-frequency transformed version thereof is obtained based on a first sample block (108_1) of the audio signal (102) and the other of the two corresponding subband sample sets or the time-frequency transformed version thereof is obtained based on a second sample block (108_2) of the audio signal.
2. The audio processor (100) according to claim 1, in, The time-frequency transform performed by the first time-frequency transform stage is an overlapped critical sampling transform.
3. The audio processor (100) according to claim 1, in, The time-frequency transform performed by the time-frequency transform stage on the identified one or more subband sample sets in the subband sample set (110_2, 1; 110_2, 2) based on the second sample block (108_2) and / or the identified one or more subband sample sets in the subband sample set (110_2, 1; 110_2, 2) based on the second sample block (108_2) corresponds to the transform described by the following formula: in The transformation is described, where describes an index of a block of samples of the audio signal, where The sub-band samples of the corresponding identified one or more sub-band sample sets are described.
4. The audio processor (100) according to claim 1, in, The cascaded overlapped critical sampling transform stages (104) are configured to: process a first set of binary bins (124_1) obtained based on a first block of samples (108_1) of the audio signal and a second set of binary bins (124_2) obtained based on a second block of samples (124_2) of the audio signal using a second overlapped critical sampling transform stage (126) of the cascaded overlapped critical sampling transform stages (104), The second overlapped critical sampling transform stage (126) is configured to: perform a first overlapped critical sampling transform on a first set of binary bins (124_1) to obtain a set of subband samples (110_1, 1; 110_1, 2) based on the first sample block (108_1), and perform a second overlapped critical sampling transform on a second set of binary bins (124_2) to obtain a set of subband samples (110_2, 1; 110_2, 2) based on the second sample block (108_2), depending on signal characteristics of the audio signal, one or more of the first critical sampling transforms having a different length than the second critical sampling transform.
5. The audio processor (100) according to claim 4, in, The first time-frequency transform stage is configured to: identify one or more subband sample sets from the subband sample sets (110_1, 1; 110_1, 2) based on the first sample block (108_1) and identify one or more subband sample sets from the subband sample sets (110_2, 1; 110_2, 2) based on the second sample block (108_2) when one or more of the first critically sampled transforms have different lengths compared to the second critically sampled transform, the identified one or more subband sample sets representing the same area in the time-frequency plane of the audio signal.
6. The audio processor (100) according to claim 1, in, The audio processor (100) comprises a second time-frequency transform stage configured to perform a time-frequency transform on an alias-reduced subband representation (112_1) of the audio signal (102), The time-frequency transform applied by the second time-frequency transform stage is opposite to the time-frequency transform applied by the first time-frequency transform stage.
7. The audio processor (100) according to claim 1, in, The time-domain aliasing reduction performed by the time-domain aliasing reduction stage corresponds to the transform described by the following formula: in The transformation is described, where describes the frame index in the z-domain, where describes an index of a block of samples of the audio signal, where Described A modified version of the overlapped critical sampling transform pre-permutation / folding matrix.
8. The audio processor (100) according to claim 1, in, The audio processor (100) is configured to provide a bitstream comprising a STDAR parameter indicating the length of the identified one or more subband sample sets corresponding to the first block of samples or the second block of samples for use in a time domain aliasing reduction stage to obtain a corresponding aliasing-reduced subband representation (112_1) of the audio signal (102), or The audio processor (100) is configured to provide a bitstream including an MDCT length parameter, wherein the MDCT length parameter indicates the length of a subband sample set (110_1, 1; 110_1, 2; 110_2, 1; 110_2, 2).
9. The audio processor (100) according to claim 1, in, The audio processor (100) is configured to perform joint channel coding.
10. The audio processor (100) according to claim 9, in, The audio processor (100) is configured to perform M / S or MCT as joint channel processing.
11. The audio processor (100) according to claim 1, in, The audio processor (100) is configured to provide a bitstream comprising at least one STDAR parameter indicating the lengths of one or more time-frequency transformed subband samples corresponding to the first block of samples and one or more time-frequency transformed subband samples corresponding to the second block of samples for use in a time domain aliasing reduction stage to obtain a corresponding aliasing-reduced subband representation (112_1) of the audio signal (102) or an encoded version thereof.
12. The audio processor (100) according to claim 1, in, The cascaded overlapped critical sampling transform stage (104) comprises a first overlapped critical sampling transform stage (120) configured to perform an overlapped critical sampling transform on a first sample block (108_1) and a second sample block (108_2) of at least two partially overlapping sample blocks (108_1, 108_2) of the audio signal (102) to obtain a first binary bin set (124_1) for the first sample block (108_1) and a second binary bin set (124_2) for the second sample block (108_2).
13. The audio processor (100) according to claim 12, in, The cascaded overlapped critical sampling transform stage (104) further comprises a second overlapped critical sampling transform stage (126) configured to perform an overlapped critical sampling transform on segments (128_1, 1) of the first set of binary bins (124_1) and to perform an overlapped critical sampling transform on segments (128_2, 1) of the second set of binary bins (124_2) to obtain a set of subband samples (110_1, 1) for the first set of binary bins and a set of subband samples (110_2, 1) for the second set of binary bins, wherein each segment is associated with a subband of the audio signal (102).
14. An audio processor (200) for processing a subband representation of an audio signal to obtain the audio signal (102), the subband representation of the audio signal comprising a set of subband samples with aliasing reduction, the audio processor (200) comprising: a second inverse time-frequency transform stage configured to time-frequency transform one or more of the sets of aliasing-reduced subband samples corresponding to the first block of samples of the audio signal and / or one or more of the sets of aliasing-reduced subband samples corresponding to the second block of samples of the audio signal to obtain one or more time-frequency transformed aliasing-reduced subband samples, each of the time-frequency transformed aliasing-reduced subband samples representing a same region in the time-frequency plane as a corresponding one of the one or more aliasing-reduced subband samples or the time-frequency transformed versions of the one or more aliasing-reduced subband samples corresponding to the other block of samples of the first block of samples and the second block of samples of the audio signal, an inverse time-domain aliasing reduction stage (202) configured to perform weighted combining of corresponding alias-reduced subband sample sets or time-frequency transformed versions thereof to obtain aliased subband representations, a first inverse time-frequency transform stage configured to perform a time-frequency transform on the aliased subband representation to obtain a subband sample set (110_1, 1; 110_1, 2) corresponding to a first sample block (108_1) of the audio signal and a subband sample set (110_2, 1; 110_2, 2) corresponding to a second sample block (108_1) of the audio signal, wherein the time-frequency transform applied by the first inverse time-frequency transform stage is opposite to the time-frequency transform applied by the second inverse time-frequency transform stage, A cascaded inverse overlapped critical sampling transform stage (204) is configured to perform a cascaded inverse overlapped critical sampling transform on a set of samples (110_1,1; 110_,2; 110_2,1; 110_2,2) to obtain a set of samples (206_1,1) associated with a block of samples of the audio signal (102).
15. A method (320) for processing an audio signal to obtain a subband representation of the audio signal, the method comprising: at least two partially overlapping blocks of samples (108_1; 108_2) performing a cascaded overlapped critical sampling transform to obtain a subband sample set (110_1, 1; 110_1, 2) based on a first sample block (108_1) of the audio signal (102), and to obtain a subband sample set (110_2, 1; 110_2, 2) based on a second sample block (108_2) of the audio signal (102); In a case where a subband sample set (110_1, 1; 110_1, 2) based on the first sample block (108_1) represents a different region in a time-frequency plane than a subband sample set (110_2, 1; 110_2, 2) based on the second sample block (108_2), one or more subband sample sets are identified (324) from the subband sample set (110_1, 1; 110_1, 2) based on the first sample block (108_1), and one or more subband sample sets are identified from the subband sample set (110_2, 1; 110_2, 2) based on the second sample block, the identified one or more subband sample sets being combined to represent the same region in the time-frequency plane, performing a time-frequency transform on the identified one or more subband sample sets in the subband sample set (110_2, 1; 110_2, 2) based on the first sample block (108_1) and / or the identified one or more subband sample sets in the subband sample set (110_2, 1; 110_2, 2) based on the second sample block (108_2) to obtain one or more time-frequency transformed subband samples, each of the time-frequency transformed subband samples representing a same region in the time-frequency plane as a corresponding one of the identified one or more subband samples or the time-frequency transformed versions of the one or more subband samples; as well as A weighted combination of two corresponding subband sample sets or time-frequency transformed versions thereof is performed to obtain an aliasing-reduced subband representation (112_1; 112_2) of the audio signal (102), wherein one of the two corresponding subband sample sets or the time-frequency transformed version thereof is obtained based on a first sample block (108_1) of the audio signal (102) and the other of the two corresponding subband sample sets or the time-frequency transformed version thereof is obtained based on a second sample block (108_2) of the audio signal.
16. A method (420) for processing a subband representation of an audio signal to obtain the audio signal, the subband representation of the audio signal comprising a set of aliasing-reduced subband samples, the method comprising: performing a time-frequency transform on one or more sets of aliasing-reduced subband samples of the sets of aliasing-reduced subband samples corresponding to the first block of samples of the audio signal and / or one or more sets of aliasing-reduced subband samples of the sets of aliasing-reduced subband samples corresponding to the second block of samples of the audio signal to obtain one or more time-frequency transformed aliasing-reduced subband samples, each of the time-frequency transformed aliasing-reduced subband samples representing a same region in a time-frequency plane as a corresponding one of the one or more aliasing-reduced subband samples or time-frequency transformed versions of the one or more aliasing-reduced subband samples corresponding to the other block of samples of the first block of samples and the second block of samples of the audio signal, performing a weighted combination of corresponding alias-reduced subband sample sets or time-frequency transformed versions thereof to obtain aliased subband representations, performing a time-frequency transform on the aliased subband representation to obtain a subband sample set (110_1, 1; 110_1, 2) corresponding to a first sample block (108_1) of the audio signal and a subband sample set (110_2, 1; 110_2, 2) corresponding to a second sample block (108_1) of the audio signal, wherein the time-frequency transform performed on one or more of the aliasing-reduced subband sample sets corresponding to the first sample block of the audio signal or one or more of the aliasing-reduced subband sample sets corresponding to the second sample block of the audio signal is opposite to the time-frequency transform performed on the aliased subband representation, A concatenated inverse overlapped critical sampling transform is performed on the sample sets (110_1, 1; 110_, 2; 110_2, 1; 110_2, 2) to obtain a sample set (206_1, 1) associated with a sample block of the audio signal (102).
17. A computer readable medium having stored thereon a computer program operable to perform the method according to claim 15 or claim 16.
Citation Information
Patent Citations
Voice encoding device, voice decoding device, voice encoding method, voice decoding method, voice encoding program, and voice decoding program.
BR122012021668A2
Time domain aliasing reduction for non-uniform filterbanks which use spectral analysis followed by partial synthesis
CN109863555A