Harmonic conversion

By employing frequency oversampling and bi-orthogonal windows, the method addresses the challenge of achieving high-frequency resolution and transient response in harmonic conversion, resulting in improved signal quality and reduced complexity.

JP7855043B2Active Publication Date: 2026-05-07DOLBY INTERNATIONAL AB
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2024-09-24
Publication Date
2026-05-07

Smart Images

  • Figure 0007855043000023
    Figure 0007855043000023
  • Figure 0007855043000024
    Figure 0007855043000024
  • Figure 0007855043000025
    Figure 0007855043000025
Patent Text Reader

Abstract

To relate to transposing signals in time and / or frequency and in particular to coding of audio signals.SOLUTION: More particular, the present invention relates to high frequency reconstruction (HFR) methods including a frequency domain harmonic transposer. A method and system for generating a transposed output signal from an input signal using a transposition factor T is described. The system comprises an analysis window of length La, extracting a frame of the input signal, and an analysis transformation unit of order M transforming the samples into M complex coefficients. The system further comprises a nonlinear processing unit altering the phase of the complex coefficients by using the transposition factor T, a synthesis transformation unit of order M transforming the altered coefficients into M altered samples, and a synthesis window of length Ls, generating a frame of the output signal.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the encoding of audio signals, particularly to the conversion of signals in frequency and / or the expansion / compression of signals in time. In other words, the present invention relates to the modification of time scales and / or frequency scales. More specifically, the present invention relates to high-frequency reconstruction (HFR) including a frequency-domain harmonic transposer. [Background technology]

[0002] High-frequency (HFR) technologies, such as Spectral Band Replication (SBR), can significantly improve the encoding efficiency of traditional perceptual audio codecs. Combined with MPEG-4 Advanced Audio Coding (AAC), HFR technology creates a highly efficient audio codec. It is already in use within the XM Satellite Radio system and Digital Radio Mondiale, and is standardized within organizations such as 3GPP® and the DVD Forum. The combination of AAC and SBR is called aacPlus, which is part of the MPEG-4 standard, where it is referred to as the High Efficiency AAC Profile. Generally, HFR technology can be combined with any perceptual audio codec in a backward-compatible manner, thus offering the potential to upgrade established broadcast systems such as MPEG-2 Layer 2 used in the Eureka DAB system. The HFR conversion method, when combined with an audio codec, can enable wide bandwidth audio at ultra-low bitrates.

[0003] The fundamental idea behind HRF is the observation that there is usually a strong correlation between the characteristics of a signal in the high-frequency range and those of the same signal in the low-frequency range. Therefore, a good approximation for representing the original input high-frequency range of a signal can be achieved by converting the signal from the low-frequency range to the high-frequency range.

[0004] The concept of conversion was established in WO98 / 57436 as a method for regenerating high-frequency bands from lower-frequency bands of audio signals. Using this concept in acoustic coding and / or speech coding results in substantial bitrate savings. While the following refers to acoustic coding (audio coding), it should be noted that the methods and systems described are equally applicable to speech coding and unified speech and audio coding.

[0005] In HFR-based audio coding systems, low-bandwidth signals are presented to the core waveform encoder, while higher frequencies are regenerated on the decoder side using the conversion of the low-bandwidth signals and additional sub-information. This sub-information is typically encoded at a very low bitrate and describes the target spectral shape. Due to the narrow bandwidth and low bitrate of the core coded signal, it becomes increasingly important to regenerate or synthesize the high-band, i.e., the high-frequency range of the audio signal, with perceptually pleasing characteristics.

[0006] Conventional techniques include several methods for reconstructing harmonic frequencies, such as using harmonic transposition or time stretching. One method is based on a phase vocoder, which operates on the principle of performing frequency analysis with sufficiently high frequency resolution. Signal modification is performed in the frequency domain before the signal is resynthesized. Signal modification may be time stretching or transposition.

[0007] One of the underlying problems with these methods is the conflicting constraints between the intended high-frequency resolution for obtaining high-quality conversion for steady-state sounds and the system's temporal response to transient or percussive sounds. In other words, while high-frequency resolution is beneficial for the conversion of steady-state signals, such high-frequency resolution typically requires a large window size, which becomes detrimental when dealing with the transient portion of a signal. One approach to address this problem is to adaptively vary the converter window as a function of the input signal characteristics, for example, by using window switching. Typically, for the steady portion of a signal, a long window is useful to achieve high frequency resolution. On the other hand, for the transient portion of a signal, a short window is used to implement a good transient response of the converter, i.e., good temporal resolution. However, this approach has the drawback that signal analysis measures, such as transient detection, must be incorporated into the conversion system. Such signal analysis measures often involve decision steps that trigger the switching of signal processing, such as a decision on the presence of a transient signal. Furthermore, such measures typically affect the reliability of the system and can introduce signal artifacts when switching signal processing, for example, when switching the window size.

[0008] The present invention solves the aforementioned problems concerning the transient performance of harmonic conversion without the need for window switching. Furthermore, improved harmonic conversion is achieved without adding significant complexity. [Prior art documents] [Patent Documents]

[0009] [Patent Document 1] EP0940015B1 / WO98 / 57436 [Overview of the Initiative] [Problems that the invention aims to solve]

[0010] This invention relates to improved transient performance for harmonic conversion and to various improvements over known methods for harmonic conversion. Furthermore, the invention explains how to minimize the added complexity while maintaining the proposed improvements. In particular, the present invention may have at least one of the following aspects: • Oversample the frequency by a factor that is a function of the conversion factor at the operating point of the converter; • Appropriate selection of combinations of decomposition windows and composite windows; and • To ensure the time alignment of such signals when different converted signals are combined. [Means for solving the problem]

[0011] According to one aspect of the present invention, a system is described for generating an output signal converted from an input signal using a conversion factor T. The converted output signal may be a time-stretched and / or frequency-shifted version of the input signal. With respect to the input signal, the converted output signal may be temporally stretched by the conversion factor T. Alternatively, the frequency components of the converted output signal may be shifted upward by the conversion factor T.

[0012] The system may include a decomposition window of length L from which L sample values ​​of an input signal are extracted. Typically, the L sample values ​​of the input signal are sample values ​​of the input signal in the time domain, such as an audio signal. The extracted L sample values ​​are referred to as frames of the input signal. The system further has a decomposition transform unit of order M = F × L that transforms the L time-domain sample values ​​into M complex coefficients, where F is the frequency oversampling factor. The M complex coefficients are typically coefficients in the frequency domain. The decomposition transform may be a Fourier transform, fast Fourier transform, discrete Fourier transform, wavelet transform, or a decomposition stage of a (possibly modulated) filter bank. The oversampling factor F is based on or a function of the conversion factor T.

[0013] The oversampling operation may also be described as zero-padding of the decomposition window by an additional (F-1) × L zeros. It can also be seen as choosing a decomposition transformation size M that is a factor F times larger than the size of the decomposition window.

[0014] The system may also have a nonlinear processing unit that modifies the phase of the complex coefficients by using a conversion factor T. The phase modification may involve multiplying the phase of the complex coefficients by the conversion factor T. Furthermore, the system may have a composite transform unit of order M that transforms the modified coefficients into M modified sample values, and a composite window of length L for generating the output signal. The composite transform may be an inverse Fourier transform, inverse fast Fourier transform, inverse discrete Fourier transform, inverse wavelet transform, or (possibly) a composite stage of a modulated filter bank. Typically, the decomposition and composite transforms relate to each other, for example, when the conversion factor T=1, to achieve a complete reconstruction of the input signal.

[0015] According to another aspect of the present invention, the oversampling factor F is proportional to the conversion factor T. In particular, the oversampling factor F may be greater than or equal to (T+1) / 2. This selection of the oversampling factor F ensures that unwanted signal artifacts that may be caused by the conversion, such as pre-echoes and post-echoes, are blocked by the synthesis window.

[0016] In a more general form, it should be noted that the length of the analysis window may be La and the length of the synthesis window may be Ls. In such cases, it may be beneficial to select the order M of the transformation unit based on the transformation order T, i.e., as a function of the transformation order T. Furthermore, it may be beneficial to select M to be greater than the average length of the analysis and synthesis windows, i.e., greater than (La+Ls) / 2. In one embodiment, the difference between the order M of the transformation unit and the average window length is proportional to (T-1). In a further embodiment, M is selected to be greater than or equal to (TLa+Ls) / 2. It should be noted that the case where the lengths of the analysis and synthesis windows are equal, i.e., La=Ls=L, is a special case of the general case described above. For the general case, the oversampling factor is

number

[0017] In other words, the decomposition window may extract or isolate L or more generally La sample values ​​of the input signal, for example, by multiplying a set of L sample values ​​of the input signal by a non-zero window coefficient. Such a set of L sample values ​​may be referred to as an input signal frame or a frame of the input signal. The decomposition stride unit shifts the decomposition window along the input signal, thereby selecting different frames of the input signal; that is, it generates a sequence of frames of the input signal. The sample value distance between the sequence of frames is given by the decomposition stride. Similarly, the synthesis stride unit shifts the synthesis window and / or frames of the output signal; that is, it generates a sequence of shifted frames of the output signal. The sample value distance between the sequence of frames of the output signal is given by the synthesis stride. The output signal may be determined by superimposing the sequence of frames of the output signal and adding together the sample values ​​that coincide in time.

[0018] According to a further aspect of the present invention, the combined stride is T times the decomposed stride. In such a case, the output signal corresponds to the input signal stretched in time by a conversion factor T. In other words, by selecting the combined stride to be T times larger than the decomposed stride, a time shift or stretching of the output signal relative to the input signal can be obtained. This time shift is of order T.

[0019] In other words, the system described above may also be described as follows: a suite or sequence of M sets of complex coefficients may be determined from an input signal using a decomposition window unit, a decomposition transform unit, and a decomposition stride unit having a decomposition stride Sa. The decomposition stride defines the number of sample values ​​the decomposition window moves forward along the input signal (how many sample values ​​it moves). Since the elapsed time between two consecutive sample values ​​is given by the sampling rate, the decomposition stride also defines the elapsed time between two frames of the input signal. Consequently, the elapsed time between two consecutive sets of M complex coefficients is also given by the decomposition stride Sa.

[0020] After passing through a non-linear processing unit where the phase of the complex coefficients can be changed, for example, by multiplying by a conversion factor T, a suite or sequence of M sets of complex coefficients may be reconverted into the time domain. Each set of M changed complex coefficients may be converted into M changed sample values using a synthesis conversion unit. In a subsequent overlapping and adding operation involving a synthesis stride unit with a synthesis window unit and a synthesis stride Ss, a suite of sets of M changed sample values may be overlapped and added to form an output signal. In this overlapping and adding operation, successive sets of M changed sample values may be shifted relative to each other by only Ss sample values, then multiplied by the synthesis window, and then added to produce an output signal. As a result, if the synthesis stride Ss is T times the decomposition stride Sa, the signal may be time-expanded by a factor T.

[0021] According to a further aspect of the present invention, the synthesis window is derived from the decomposition window and the synthesis stride. In particular, the synthesis window may be given by the following formula.

[0022] [Number] Here, v s (n) is the synthesis window, v a (n) is the decomposition window, and Δt is the synthesis stride Ss. The decomposition window and / or the synthesis window may be a Gaussian window, a cosine window, a Hamming window, a Hann window, a rectangular window, a Bartlett window, a Blackman window, or one of the functions v(n) = sin{(π / L)(n + 0.5)} with 0 ≦ n < L. Here, if the lengths of the decomposition window and the synthesis window are different, L may be La or Ls, respectively.

[0023] According to another aspect of the present invention, the system further includes a stenosis unit that performs rate conversion of the output signal by, for example, conversion order T, thereby producing a converted output signal. By selecting the combined stride to be T times the decomposed stride, a time-stretched output signal can be obtained as outlined above. If the sampling rate of the time-stretched signal is increased by a factor of T, or if the time-stretched signal is downsampled by a factor of T, a converted output signal corresponding to the input signal frequency-shifted by the conversion factor T can be produced. The downsampling operation may have a step of selecting only a subset of the sample values ​​of the output signal. Typically, only the T-th sample value of the output signal is retained. Alternatively, the sampling rate may be increased by a factor of T, i.e., the sampling rate is interpreted as T times higher. In other words, resampling or sampling rate conversion means that the sampling rate is changed to a higher or lower value. Downsampling means a rate conversion to a lower value.

[0024] According to a further aspect of the present invention, the system may generate a second output signal from an input signal. The system may have a second nonlinear processing unit that modifies the phase of the complex coefficients by using a second conversion factor T2, and a second composite stride unit that shifts the composite window and / or the frame of the second output signal by a second composite stride. The phase modification may include multiplying the phase by a factor T2. The frame of the second output signal may be generated from the frame of the input signal by modifying the phase of the complex coefficients using a second conversion factor, converting the second modified coefficients into M second modified sample values, and applying a composite window. The second output signal may be generated in a superposition unit by applying a second composite stride to the sequence of frames of the second output signal.

[0025] The second output signal may be condensed in a second condensation unit, for example, by performing a rate conversion of the second output signal using a second conversion order T2. This results in a second converted output signal. In summary, the first converted output signal can be generated using a first conversion factor T, and the second converted output signal can be generated using a second conversion factor T2. These two converted output signals may then be merged in a combination unit to produce a converted output signal as a whole. The merging operation may include adding the two converted output signals together. The generation and combination of such multiple converted output signals can be useful in obtaining a good approximation of the high-frequency signal components to be synthesized. It should be noted that any number of converted output signals may be generated using multiple conversion orders. These multiple converted output signals may then be merged, for example, added together in a combination unit to produce a whole converted output signal.

[0026] It may be beneficial for the combination unit to weight the first and second converted output signals prior to merging. The weighting may be performed such that the energy or energy per bandwidth of the first and second converted output signals corresponds to the energy or energy per bandwidth of the input signal, respectively.

[0027] According to a further aspect of the present invention, the system may have an alignment unit that applies a time offset to the first and second converted output signals before they enter the combination unit. Such a time offset may include the time-domain shift of the two converted output signals relative to each other. The time offset may be a function of the conversion order and / or the window length. In particular, the time offset is (T-2)L / 4 It may be decided as such.

[0028] According to another aspect of the present invention, the above-described conversion system may be incorporated into a system for decoding a received multimedia signal, including an audio signal. The decoding system may have a conversion unit corresponding to the system outlined above, where the input signal is typically the low-frequency components of an audio signal, and the output signal is the high-frequency components of an audio signal. In other words, the input signal is typically a low-pass signal with a certain bandwidth, and the output signal is typically a band-pass signal with a higher bandwidth. Furthermore, it may have a core decoder for decoding the low-frequency components of the audio signal from the received bitstream. Such a core decoder may be based on an encoding scheme such as Dolby E, Dolby Digital, or AAC. In particular, such a decoding system may be a set-top box for decoding a received multimedia signal, including an audio signal and other signals such as video.

[0029] It should be noted that the present invention also describes a method for converting an input signal using a conversion factor T. This method corresponds to the system outlined above and may include any combination of the aspects described above. It may include the steps of extracting sample values ​​of the input signal using a decomposition window of length L and selecting an oversampling factor F as a function of the conversion factor T. Furthermore, it may include the steps of converting L sample values ​​from the time domain to the frequency domain to produce F × L complex coefficients and changing the phase of the complex coefficients using the conversion factor T. In a further step, the method may convert the F × L modified complex coefficients to the time domain to produce F × L modified sample values ​​and generate an output signal using a composite window of length L. It should also be noted that the method may be adapted to general lengths of the decomposition and composite windows, i.e., general La and Ls as outlined above.

[0030] According to a further aspect of the present invention, the method may include a step of shifting the resolution window by a resolution stride of Sa sample values ​​along the input signal, and / or shifting the composite window and / or the frame of the output signal by a composite stride of Ss sample values. By selecting that the composite stride is T times the resolution stride, the output signal may be time-stretched by a factor of T relative to the input signal. A transformed output signal may be obtained by performing an additional step of rate-transforming the output signal by a conversion order T. Such a transformed output signal may include frequency components that are shifted up by a factor of T relative to the corresponding frequency components of the input signal.

[0031] The method may further include steps for generating a second output signal. This may be implemented by changing the phase of the complex coefficients by using a second conversion factor T2. The second output signal may be generated using the second conversion factor T2 and the second composite stride by shifting the composite window and / or the frame of the second output signal by a second composite stride. The second converted output signal may be generated by performing a rate transform on the second output signal by a second conversion order T2. Finally, by merging the first and second converted output signals, a merged or overall converted output signal can be obtained that includes high-frequency signal components generated by two or more transformations with different conversion factors.

[0032] According to another aspect of the present invention, the present invention describes a software program adapted for execution on a processor and for performing the method steps of the present invention when executed on a computing device. The present invention also describes a storage medium having a software program adapted for execution on a processor and for performing the method steps of the present invention when executed on a computing device. Furthermore, the present invention describes a computer program product comprising executable instructions for performing the method of the present invention when executed on a computer.

[0033] In a further aspect, another method and system for converting an input signal using a conversion factor T is described. This method and system may be used standalone or in combination with the methods and systems outlined above. Any of the features outlined in this paper may be applied to this method / system, and vice versa.

[0034] The method may include a step of extracting a frame of sample values ​​of the input signal using a decomposition window of length L. The frame of the input signal may then be converted from the time domain to the frequency domain to produce M complex coefficients. The phase of the complex coefficients may be modified using a conversion factor T, and the M modified complex coefficients may be converted to the time domain to produce M modified sample values. Finally, the frame of the output signal may be generated using a synthesis window of length L. The method and system may use different decomposition and synthesis windows. The decomposition and synthesis windows may differ in their shape, length, the number of coefficients defining the window, and / or the value of the coefficients defining the window. Doing so may provide additional degrees of freedom in the selection of the decomposition and synthesis windows and may reduce or eliminate aliasing of the converted output signal.

[0035] From another perspective, the decomposition window and the composite window are bi-orthogonal to each other. Composite window v s (n) may also be given by the following equation.

[0036]

number

[0037]

number

[0038] In one further aspect, the decomposition window may be selected such that its z-transform has dual zeros on the unit circle. Preferably, the z-transform of the decomposition window has only dual zeros on the unit circle. For example, the decomposition window may be a squared sine window. In another example, a decomposition window of length L may be determined by convolving two sine windows of length L to produce a squared sine window of length 2L-1. In one further step, zeros may be appended to the squared sine window to produce a base window of length 2L. Finally, the base window may be resampled using linear interpolation to produce an even symmetric window of length L as the decomposition window.

[0039] The methods and systems described in this paper may be implemented as software, firmware, and / or hardware. Certain components may be implemented, for example, as software running on a digital signal processor or microprocessor. Other components may be implemented, for example, as hardware and / or application-specific integrated circuits. Signals encountered in the methods and systems described may be stored on media such as random-access memory or optical storage media. These signals may be transmitted over networks such as radio networks, satellite networks, wireless networks, or wired networks, such as the Internet. Typical devices using the methods and systems described in this paper are set-top boxes or other customer premises equipment that decodes audio signals. On the encoding side, the methods and systems may be used in broadcast stations, for example, in video or television headend systems.

[0040] It should be noted that the embodiments and aspects of the present invention described herein may be combined in any way. In particular, it should be noted that the aspects outlined for the system are also applicable to the corresponding methods encompassed by the present invention. Furthermore, it should be noted that the disclosure of the present invention covers combinations of claims other than those explicitly given by reference in the dependent claims. That is, the claims and their technical features can be combined in any order and in any form.

[0041] The present invention will now be described with reference to the accompanying drawings, using examples that do not limit the scope or spirit of the invention. [Brief explanation of the drawing]

[0042] [Figure 1] This figure shows the Dirac at a specific position appearing in the separation window and combination window of a harmonic converter. [Figure 2] This figure shows the Dirac at different positions appearing in the separation window and combination window of a harmonic converter. [Figure 3] This figure shows the Dirac for the position shown in Figure 2, which appears according to the present invention. [Figure 4] This diagram shows the operation of an HFR-enhanced audio decoder. [Figure 5] This diagram shows the operation of a harmonic converter that uses several orders. [Figure 6] This diagram shows the operation of a frequency domain (FD) harmonic converter. [Figure 7] This diagram shows a series of decomposition and synthesis windows. [Figure 8] This diagram shows the breakdown window and the composite window at different stride lengths. [Figure 9] This figure shows the effect of resampling on the combined stride of the window. [Figure 10]This figure shows an embodiment of an encoder using the improved harmonic conversion method outlined in this paper. [Figure 11] This figure shows an embodiment of a decoder using the improved harmonic conversion method outlined in this paper. [Figure 12] This figure shows an embodiment of the conversion unit shown in Figures 10 and 11. [Modes for carrying out the invention]

[0043] The embodiments described below merely illustrate the principles of the present invention for improved harmonic conversion. It will be understood that modifications and variations to the configurations and details described herein will be obvious to those skilled in the art. Accordingly, the present invention is intended to be limited only by the appended claims and not by the specific details presented by the description and explanation of the embodiments herein.

[0044] The principle of harmonic conversion in the frequency domain and the proposed improvements taught by this invention are outlined below. The key element of harmonic conversion is time stretching by an integer conversion factor T, which preserves the frequency of the sine wave. In other words, harmonic conversion is based on time stretching of the fundamental signal by a factor of T. The time stretching is performed so as to preserve the frequency of the sine wave that constitutes the input signal. Such time stretching can be performed using a phase vocoder. A phase vocoder has a resolution window v a (n) and composite window v s It is based on a frequency-domain representation established by a DFT filter bank windowed using (n). Such a decomposition / composition transform is also called a short-time Fourier transform (STFT).

[0045] The Short-Time Fourier Transform (STFT) is performed on a time-domain input signal to obtain a series of overlapping spectral frames. To minimize possible sideband effects, an appropriate decomposition / combination window should be selected, such as a Gaussian window, cosine window, Hamming window, Hann window, rectangular window, Bartlett window, or Blackman window. The time delay at which each spectral frame is picked from the input signal is called the hop size or stride. The STFT of the input signal is called the decomposition stage, which leads to the frequency-domain representation of the input signal. The frequency-domain representation includes multiple subband signals, where each subband signal represents a specific frequency component of the input signal.

[0046] Next, the frequency domain representation of the input signal can be processed in a desired manner. For the purpose of time stretching of the input signal, each subband signal may be time stretched, for example, by delaying the sample values ​​of the subband signals. This may be achieved by using a composite hop size larger than the decomposed hop size. The time domain signal may be reconstructed by performing an inverse (fast) Fourier transform on all frames and then sequentially accumulating the frames. This operation of the synthesis stage is called superimposed summation. The resulting output signal is a time-stretched version of the input signal, containing the same frequency components as the input signal. In other words, the resulting output signal has the same spectral composition as the input signal, but is slower than the input signal, i.e., its progression is time-stretched.

[0047] Next, conversion to higher frequencies can be achieved in a subsequent process, or in an integrated manner, through downsampling of the extended signal. As a result, the converted signal has the time length of the initial signal, but has frequency components shifted upward by a predetermined conversion factor.

[0048] Mathematically, a phase vocoder can be described as follows: An input signal x(t) is sampled at a sampling rate R to produce a discrete input signal x(n). During the decomposition stage, a specific decomposition time t is performed for a series of values ​​k.a k The STFT is determined for the input signal x(n) at. The decomposition time is preferably t a k = kΔt a is uniformly selected through. Here, Δt a is the decomposition hop factor or decomposition stride. At each of these decomposition times t a k a Fourier transform is calculated for the windowed portion of the original signal x(n). Here, the decomposition window v a (t) is centered on t a k . That is, v a (t - t a k ). This windowed portion of the input signal x(n) is called a frame. The result is the STFT representation of the input signal x(n) and can be expressed as follows.

[0049]

Equation

[0050] The synthesis stage can typically be performed at synthesis times t s k = kΔt s distributed uniformly according to. Here, Δt s k where. Here, Δt sis the composite Hop factor or composite stride. At each of these composite times, the short-time signal y k (n) is the synthesis time t s k In this case, X(t a k ,Ω m The STFT subband signal Y(t) may be the same as ) s k ,Ω m It is obtained by inverse Fourier transforming ). However, typically the STFT subband signal is modified, for example, by time stretching and / or phase modulation and / or amplitude modulation, thereby decomposing the subband signal X(t a k ,Ω m ) is the composite subband signal Y(t s k ,Ω m This is different from the above. In one preferred embodiment, the STFT subband signal is phase-modulated, i.e., the phase of the STFT subband signal is corrected. Short-term combined signal y k (n) It can be expressed as follows:

[0051]

number

[0052]

number

[0053] The following outlines the implementation of time stretching in the frequency domain. A suitable starting point for describing the various aspects of a time stretcher is to consider the case where T=1, i.e., the conversion factor T is equal to 1 and no stretching occurs. The resolution time stride Δt of the DFT filter bank. a and combined time stride Δt s They are equal, i.e., Δt a =Δt s Assuming = Δt, the combined effect of decomposition and subsequent composition is a function of period Δt.

number

number

[0054] For conversion factors T > 1, i.e., greater than 1, time extension is Δt for the composite stride. s Maintaining =Δt while stride Δt aThis can be obtained by performing decomposition with =Δt / T. In other words, the time extension due to factor T can be obtained by applying a Hop factor or stride in a decomposition window that is T times smaller than the Hop factor or stride in the synthesis stage. As can be seen from the formula above, using a synthesis stride that is T times larger than the decomposition stride results in a short-term synthesis signal y k In the superposition summation operation, (n) is shifted by an interval that is T times larger. This ultimately leads to a time extension of the output signal y(n).

[0055] It should be noted that time extension due to factor T may also involve phase multiplication by factor T between decomposition and synthesis. In other words, time extension due to factor T includes phase multiplication by factor T of the subband signal.

[0056] The following outlines how the above time-stretching operation can be transitioned during harmonic transposition. Pitch-scale modification, or harmonic transposition, can be obtained by performing a sample rate transformation on the time-stretched output signal y(n). To perform harmonic transposition by factor T, the output signal y(n), which is a time-stretched version of the input signal x(n) by factor T, may be obtained using the phase vocoding method described above. The harmonic transposition may then be obtained by downsampling the output signal y(n) by factor T, or by converting the sampling rate from R to TR. In other words, instead of interpreting the output signal y(n) as having the same sampling rate as the input signal x(n) but with a duration T times longer, the output signal y(n) may be interpreted as having the same duration but a sampling rate T times longer. Then, the subsequent downsampling by T may be interpreted as making the output sampling rate equal to the input sampling rate so that the signals can ultimately be added together.

[0057] Assume the input signal x(n) is a sine wave, and the symmetric resolution window va Assuming (n), the above time stretching method based on the phase vocoder works perfectly for odd T values, producing a time-stretched version of the input signal x(n) with the same frequency. Combined with subsequent downsampling, a sine wave y(n) with a frequency T times that of the input signal x(n) is obtained.

[0058] For even numbers T, the time stretching / harmonic transformation method outlined above becomes a more approximate approximation. (Decomposition window v) a This is because the negative side lobes of the frequency response of (n) are reproduced with different fidelity by phase multiplication. Negative side lobes typically stem from the fact that most practical windows (or prototype filters) have a number of discrete zeros that lie on the unit circle, resulting in a 180-degree phase shift. When multiplying the phase angle with an even conversion factor, the phase shift is typically converted to 0 degrees (or rather, a multiple of 360), depending on the conversion factor used. In other words, when using an even conversion factor, the phase shift disappears. This typically leads to aliasing in the converted output signal y(n). A particularly undesirable scenario can occur when a sine wave lies at a frequency corresponding to the top of the first side lobe of the decomposition filter. Depending on the blocking of this lobe in the magnitude response, the aliasing can become more or less audible in the output signal. For even factors T, reducing the overall stride Δt typically improves the performance of the time stretcher at the cost of higher computational complexity.

[0059] Patent Document 1, titled "Improving Source Coding Using Spectral Band Replication," which is incorporated here by reference, describes a method for avoiding aliasing that occurs from harmonic converters when using even-numbered conversion factors. This method, called relative phase locking, evaluates the relative phase difference between adjacent channels and determines whether the sine wave is phase-inverted in any of the channels. Detection is performed using equation (32) of Patent Document 1. Channels detected as phase-inverted are corrected after their phase angle is multiplied by the actual conversion factor.

[0060] The following describes a novel method for avoiding aliasing when using even and / or odd conversion factors T. Unlike the relative phase-locking method described in the patent literature, this method does not require phase angle detection and correction. The novel solution to the above problem utilizes non-identical decomposition and synthesis transformation windows. In the case of perfect reconstruction (PR), this corresponds to a biorthogonal transformation / filter bank rather than an orthogonal transformation / filter bank.

[0061] A certain disassembly window v a Given (n), to obtain a biorthogonal transformation, we need the composite window v s (n)

number

number

[0062] However, in what follows, another sequence w(n) is introduced. w(n) is an indicator of how much the synthesis window v s (n) deviates from the decomposition window v a (n), that is, how different the biorthogonal transform is from the orthogonal transform. The sequence w(n) is w(n) = v s (n) / v a (n) for 0 ≦ n < L given by

[0063] Then, the condition for perfect reconstruction is

Eq.

[0064]

Eq.

[0065]

Eq.

[0066] To obtain a decomposition / synthesis window pair that suppresses aliasing for even conversion factors, some embodiments are outlined below. According to a first embodiment, the window or prototype filter is made long enough to attenuate the level of the first side lobe in the frequency response below a certain "aliasing" level. The decomposition window stride Δt a is in this case only a (small) fraction of the window length L. This typically leads to, for example, smearing of transient components in impact signals.

[0067] According to a second embodiment, the decomposition window v a (n) is chosen to have dual zeros on the unit circle. The phase response resulting from the dual zeros is a 360-degree phase shift. These phase shifts are retained when the phase angle is multiplied by the conversion factor, regardless of whether the conversion factor is odd or even. When an appropriate and smooth decomposition filter v a (n) with dual zeros on the unit circle is obtained, the synthesis window is obtained from the equations outlined above.

[0068] In an example of the second embodiment, the decomposition filter / window v a (n) is a "squared sine window", i.e., the sine window v(n)=sin{(π / L)(n + 0.5)} 0≦n<L is

Number

[0069] Overall, we have outlined how pairs of separation and synthesis windows can be selected so that aliasing in the converted output signal can be avoided or significantly reduced. This method is particularly important when using even conversion factors.

[0070] Another aspect to consider in the context of vocoder-based harmonic converters is phase unwrapping. While careful attention is needed regarding the phase unwrapping problem in general-purpose phase vocoders, it should be noted that harmonic converters have unambiguously defined phase behavior when an integer conversion factor T is used. Thus, in preferred embodiments, the conversion order T is an integer value. Otherwise, a phase unwrapping technique can be applied. Here, phase unwrapping is the process of estimating the instantaneous frequency of a nearby sine wave in each channel using the phase increment between two consecutive frames.

[0071] Another aspect to consider when dealing with the conversion of acoustic and / or speech signals is the handling of steady-state and / or transient signal sections. Typically, the frequency resolution of the DFT filter bank needs to be relatively high in order to convert steady-state acoustic signals without intermodulation artifacts, and therefore the window is long compared to the input signal x(n), especially the transient components in the acoustic and / or speech signal. As a result, the converter has a poor transient response. However, as will be discussed below, this problem can be solved by modifying the window design, conversion size, and time stride parameters. Thus, unlike many current techniques for improving the transient response of phase vocoders, the proposed solution does not rely on any signal-adaptive operation such as transient component detection.

[0072] The following outlines the harmonic conversion of transient signals using a vocoder. As a starting point, we consider the discrete-time Dirac pulse at time t=t0, which is the prototype transient signal.

number

[0073]

number

[0074] This demonstrates that the phase multiplication operation of the decomposed subband signals by factor T leads to the desired time shift of the Dirac pulse, i.e., the transient input signal. It should be noted that for more realistic transient signals with two or more non-zero sample values, further operation of time stretching of the decomposed subband signals by factor T should be performed. In other words, different hop sizes should be used on the decomposition and synthesis sides.

[0075] However, it should be noted that the above considerations pertain to decomposition / composition stages using infinite-length decomposition and composition windows. In fact, theoretical converters with infinite-duration windows give the correct extension of the Dirac pulse δ(t-t0). For finite-duration windowed decompositions, the situation becomes complicated by the fact that each decomposition block should be interpreted as a periodic interval of a periodic signal with a period equal to the size of the DFT.

[0076] This is shown in Figure 1. Figure 1 shows the decomposition and synthesis of the Dirac pulse δ(t-t0). The upper part of Figure 1 shows the input to the decomposition stage 110, and the lower part of Figure 1 shows the output of the synthesis stage 120. The upper and lower graphs represent the time domain. The stylized decomposition window 111 and synthesis window 121 are drawn as triangular (Bartlett) windows. The input pulse δ(t-t0) 112 at time t=t0 is drawn as a vertical arrow in the upper graph 110. The DFT transform block is assumed to be of size M=L. That is, the size of the DFT transform is chosen to be equal to the size of the window. Phase multiplication of the subband signal by factor T yields the DFT decomposition of the Dirac pulse δ(t-Tt0) at t=Tt0, which is divided into a Dirac pulse train with period L. This is due to the finite length of the window and Fourier transform applied. The segmented pulse trains with period L are shown by the dashed arrows 123 and 124 in the graph below.

[0077] In a real-world system where the decomposition window and the synthesis window are of finite length, the pulse train actually contains only a few pulses (depending on the transform factor). One main pulse, i.e., the desired term, and a few pre-pulses and a few post-pulses, i.e., the undesired terms. The pre-pulses and post-pulses occur because the DFT is periodic (period L). When a pulse is located within the decomposition window and is folded (wrapped) when the complex phase is multiplied by T (i.e., the pulse is shifted beyond the end of the window and returns to the beginning), the undesired pulses appear. The undesired pulses may or may not have the same polarity as the input pulse, depending on their position in the decomposition window and the transform factor.

[0078] This can be seen mathematically when converting the Dirac pulse δ(t - t0) located in the interval -L / 2 ≤ t0 < L / 2 using a DFT of length L centered at t = 0.

[0079]

Number

Number

[0080] In the example of Fig. 1, the synthesis windowing uses the finite window v s (n) 121. The finite synthesis window 121 picks up the desired pulse δ(t - Tt0) at t = Tt0 depicted as the solid arrow 122 and cancels the other contributions shown as the dashed arrows 123, 124.

[0081] As the decomposition and synthesis stages move along the time axis according to the Hop factor or time stride Δt, the pulse δ(t-t0) 112 will have a different position relative to the center of each decomposition window 111. As outlined above, the action to achieve time stretching is to move the pulse 112 by a factor of T relative to the center of the window. As long as this position is within window 121, this time stretching action ensures that when all contributions are added together, they result in a single time-stretched synthesis pulse δ(t-Tt0) at t=Tt0.

[0082] However, a problem arises in the situation shown in Figure 2. Here, pulse δ(t-t0) 212 moves further outward towards the edge of the DFT block. Figure 2 shows a decomposition / combination configuration 200 similar to that in Figure 1. The upper graph 210 shows the input to the decomposition stage and the decomposition window 211, and the lower graph 220 shows the output of the combination stage and the combination window 221. When the input Dirac pulse 212 is time-stretched by factor T, the time-stretched Dirac pulse 222, i.e., δ(t-Tt0), falls outside the combination window 221. At the same time, another Dirac pulse 224 in the pulse train, i.e., δ(t-Tt0+L) at time t=Tt0-L, is picked up by the combination window. In other words, the input Dirac pulse 212 is not delayed to a time T times later, but rather advanced to a time earlier than the input Dirac pulse 212. The final effect on the audio signal is the generation of a pre-echo at a time t=Tt0-L that is L-(T-1)t0 earlier than the input Dirac pulse 212, on a time distance scaled by a longer converter window.

[0083] The principle of the solution proposed by the present invention is described with reference to Figure 3. Figure 3 shows a decomposition / combination scenario 300 similar to that in Figure 2. The upper graph 310 shows the input to the decomposition stage with the decomposition window 311, and the lower graph 320 shows the output to the combination stage with the combination window 321. The basic idea of ​​the present invention is to adapt the DFT size to avoid pre-echoes. This can be achieved by setting the DFT size M so that unwanted Dirac pulse images are not picked up by the combination window from the resulting pulse train. The size of the DFT transform 301 is increased to M = FL, where L is the length of the window function 302 and factor F is the frequency domain oversampling factor. In other words, the size of the DFT transform 301 is selected to be larger than the window size 302. In particular, the size of the DFT transform 301 may be selected to be larger than the window size 302 of the combination window. Due to the increased length 301 of the DFT transform, the period of the pulse train containing the Dirac pulses 322, 324 is FL. By selecting a sufficiently large value for F, that is, by selecting a sufficiently large frequency-domain oversampling factor, unwanted contributions to pulse stretching can be eliminated. This is shown in Figure 3. The Dirac pulse 324 at time t=Tt0-FL is outside the blending window 321. Therefore, the Dirac pulse 324 is not picked up by the blending window 321, and as a result, pre-echoes can be avoided.

[0084] It should be noted that in one preferred embodiment, the composite window and the decomposition window have equal “normal” lengths. However, when implicit resampling of the output signal is used by discarding or inserting sample values ​​in the frequency band of the conversion or filter bank, the composite window size is typically different from the decomposition size, depending on the resampling or conversion factor.

[0085] The minimum value of F, i.e., the minimum frequency-domain oversampling factor, can be deduced from Fig. 3. The condition for not picking up an undesired Dirac pulse image can be formulated as follows: for any input pulse δ(t - t0) at position t = t0 < L / 2, i.e., for any input pulse contained within the decomposition window 311, the undesired image δ(t - Tt0 + FL) at time t = Tt0 - FL must be located to the left of the left end of the synthesis window at t = -L / 2. Equivalently, the condition T(L / 2) - FL ≤ -L / 2 must be satisfied. This leads to the rule F ≥ (T + 1) / 2 (3) as follows.

[0086] As can be seen from formula (3), the minimum frequency-domain oversampling factor F is a function of the conversion / time stretching factor T. More specifically, the minimum frequency-domain oversampling factor F is proportional to the conversion / time stretching factor T.

[0087] By repeating the above train of thought for the case where the decomposition and synthesis windows have different lengths, a more general formula can be obtained. Let L A and L S be the lengths of the decomposition window and the synthesis window respectively, and let M be the DFT size used. Then, the rule for extending formula (3) is M ≥ (TL A + L S ) / 2 (4) as follows.

[0088] That this rule is actually an extension of (3) can be verified by substituting M = FL and L A = L S - L into (4) and dividing both sides of the resulting formula by L.

[0089] The above analysis is performed on a somewhat special model of transient signals, i.e., Dirac pulses. However, the idea can be extended to show that when using the above time stretching method, an input signal with a nearly flat spectral envelope and zero outside the time interval [a,b] is stretched to a small output signal outside the interval [Ta,Tb]. Furthermore, the disappearance of pre-echoes in the stretched signal when the above rules for selecting an appropriate frequency-domain oversampling factor are respected can also be checked by examining the spectrograms of actual acoustic and / or speech signals. More quantitative analysis reveals that pre-echoes are mitigated even when using a frequency-domain oversampling factor slightly less than the value imposed by the conditions of formula (3). This is typical of the window function v s This is due to the fact that (n) is small near the edges, thereby attenuating unwanted pre-echoes located near the edges of the window function.

[0090] In summary, the present invention teaches a novel method for improving the transient response of a frequency response harmonic converter or time expander by introducing an oversampled transform such that the oversampling amount is a function of a selected transform factor.

[0091] The following describes in more detail the application of harmonic conversion in audio decoders based on the present invention. A common use case for harmonic converters is in acoustic / speech codec systems that utilize so-called bandwidth expansion or high-frequency regeneration (HFR). Although acoustic coding [audio coding] is mentioned, it should be noted that the methods and systems described are equally applicable to both speech coding and unified speech and audio coding.

[0092] In such an HFR system, a converter may be used to generate high-frequency signal components from low-frequency signal components provided by a so-called core decoder. The envelope of the high-frequency components may be shaped in time and frequency based on sub-information transmitted in the bitstream.

[0093] Figure 4 illustrates the operation of an HFR-enhanced audio decoder. The core audio decoder 401 outputs a low-bandwidth audio signal, which is input to an upsampler 404. The upsampler 404 may be required to generate the final audio output contribution at the desired full sampling rate. Such upsampling is required for dual-rate systems where the bandwidth-limited core audio codec operates at half the external audio sampling rate, while the HFR portion is processed at the full sampling frequency. Consequently, for single-rate systems, this upsampler 404 is omitted. The low-bandwidth output of 401 is also sent to a converter or conversion unit 402 that outputs a converted signal, i.e., a signal containing the desired high-frequency range. This converted signal may be shaped in time and frequency by an envelope tuner 403. The final audio output is the sum of the low-bandwidth core signal and the envelope-tuned converted signal.

[0094] As outlined in the context of Figure 4, the core decoder output signal may be upsampled by factor 2 as a preprocessing step in the conversion unit 402. The conversion by factor T, in the case of time stretching, produces a signal with T times the length of the unconverted signal. Downsampling or rate conversion of the time-stretched signal is then performed to achieve the desired pitch-shifting or frequency transposition to a frequency T times higher. As described above, this operation may be achieved through the use of different decomposition strides and combined strides in the phase vocoder.

[0095] The overall conversion order can be obtained in various ways. The first possibility is to upsample the decoder output signal by factor 2 at the input of the converter, as noted above. In such a case, the time-stretched signal must be downsampled by factor T in order to obtain the desired output signal that has been frequency-converted by factor T. The second possibility is to omit the above preprocessing step and directly perform the time-stretching operation on the output signal of the core decoder. In such a case, in order to retain the global upsampling factor 2 and achieve frequency conversion by factor T, the converted signal must be downsampled by factor T / 2. In other words, when downsampling the output signal of converter 402 by T / 2 instead of T, upsampling of the core decoder signal may be omitted. However, it should be noted that the core signal still needs to be upsampled in the upsampler 404 before it is combined with the converted signal.

[0096] It should also be noted that converter 402 may use several different integer conversion factors to generate high-frequency components. This is shown in Figure 5. Figure 5 shows the operation of harmonic converter 501, corresponding to converter 402 in Figure 4, with several converters of different conversion orders or conversion factors T. The signals to be converted are, respectively, conversion orders T=2,3,...,T max Individual converters 501-2, 501-3, ..., 501-T max It is passed to the bank. Typically, the conversion order is T. max =3 is sufficient for most audio coding applications. Different converters 501-2, 501-3, ..., 501-T maxThe contributions are summed in 502 to give a combined converter output. In the first embodiment, this summing operation may include adding up the individual contributions. In another embodiment, the contributions are weighted with different weights so that the effect of adding multiple contributions to a certain frequency is mitigated. For example, a third-order contribution may be added with a lower gain than a second-order contribution. Finally, the summing unit 502 may selectively add these contributions depending on the output frequency. For example, a second-order conversion may be used for a first lower target frequency unit, and a third-order conversion may be used for a second higher target frequency unit.

[0097] Figure 6 shows the operation of one of the individual blocks of 501, i.e., a harmonic converter such as one of the converters 501-T with conversion order T. The decomposition stride unit 601 selects a series of frames of the input signal to be converted. These frames are superimposed, for example, multiplied, with the decomposition window in the decomposition window unit 602. Note that the operation of selecting frames of the input signal and multiplying the sample values ​​of the input signal by the decomposition window function may be performed in unique steps, for example, by using a window function that is shifted along the input signal by the decomposition stride. In the decomposition transform unit 603, the windowed frames of the input signal are transformed into the frequency domain. The decomposition transform unit 603 may perform a DFT, for example. The size of the DFT is selected to be F times larger than the size L of the decomposition window, thereby generating M = F × L complex frequency domain coefficients. These complex coefficients are modified in the nonlinear processing unit 604, for example, by multiplying their phase by a conversion factor T. The complex coefficients of the sequence of complex frequency domain signals, i.e., the sequence of frames of the input signal, may be viewed as subband signals. The combination of the disassembly stride unit 601, the disassembly window unit 602, and the disassembly conversion unit 603 may be viewed as a combined disassembly stage or disassembly filter bank.

[0098] The modified coefficients or modified subband signals are converted back into the time domain using the synthesis conversion unit 605. For each set of converted complex coefficients, this gives a frame of modified sample values, i.e., a set of M modified sample values. Using the synthesis window unit 606, L sample values ​​may be extracted from each set of modified sample values, thereby giving a frame of the output signal. Overall, a sequence of frames of the output signal may be generated for a sequence of frames of the input signal. The frames of this sequence are shifted from each other by a composite stride in the synthesis stride unit 607. The composite stride may be T times greater than the decomposition stride. The output signal is generated in the superposition summation unit 608, where the shifted frames of the output signal are superimposed and sample values ​​at the same time are added together. By passing through the above system, the input signal may be time-stretched by a factor T. That is, the output signal may be a time-stretched version of the input signal.

[0099] Finally, the output signal may be temporally compressed using a condensation unit 609. The condensation unit 609 may perform a sampling rate transformation of order T. That is, the sampling rate of the output signal may be increased by factor T while keeping the number of sample values ​​constant. This gives a transformed output signal that has the same temporal length as the input signal but has frequency components shifted up by factor T relative to the input signal. The combination unit 609 may also perform a downsampling operation by factor T. That is, only the T-th sample value may be retained and the other sample values ​​may be discarded. This downsampling operation may be achieved by a low-pass filter operation. If the overall sampling rate remains constant, the transformed output signal has frequency components that are shifted up by factor T relative to the frequency components of the input signal.

[0100] It should be noted that the stenosis unit 609 may perform a combination of rate conversion and downsampling. For example, the sampling rate may be increased by factor 2. At the same time, the signal may be downsampled by factor T / 2. Overall, such a combination of rate conversion and downsampling also leads to an output signal which is a harmonic conversion of the input signal by factor T. In general, it may be said that the stenosis unit 609 performs a combination of rate conversion and / or downsampling to give a harmonic conversion by conversion order T. This is particularly useful when performing harmonic conversion of the low-bandwidth output of the core audio decoder 401. As outlined above, such a low-bandwidth output may be downsampled by factor 2 in the encoder and therefore may require upsampling in the upsampling unit 404 before being merged with the reconstructed high-frequency components. Nevertheless, performing harmonic conversion in the conversion unit 402 using a “not upsampled” low-bandwidth output can be useful to reduce computational complexity. In such cases, the stenosis unit 609 of the conversion unit 402 may perform a rate conversion of order 2, thereby implicitly performing an upsampling operation that requires high-frequency components. As a result, the converted output signal of order T is downsampled by factor T / 2 in the stenosis unit 609.

[0101] In the case of multiple parallel converters of different conversion orders as shown in Figure 5, some conversion or filter bank operations are different for converters 501-2, 501-3, ..., 501-T max The filter bank operation may be shared between them. Sharing of the filter bank operation may preferably be done with respect to decomposition to obtain a more effective implementation of the conversion unit 402. A preferred method for resampling the outputs from different converters may be to discard the DFT bins or subband channels before the synthesis stage. Thus, the resampling filter may be omitted, and the computational cost may be reduced when performing smaller inverse DFT / synthesis filter banks.

[0102] As mentioned above, the resolution window may be common to signals with different conversion factors. When using a common resolution window, an example of a window 700 stride applied to a low-band signal is shown in Figure 7. Figure 7 shows the resolution Hop factor or resolution time stride Δt a This shows the strides of the disassembly windows 701, 702, 703, and 704, which are displaced relative to each other.

[0103] Figure 8(a) shows an example of a window stride applied to a low-band signal, such as the output signal of a core decoder. The stride of a decomposition window of length L moved for each decomposition transform is Δt a This is expressed as follows. Each such decomposition transform and the windowed portion of the input signal are also called a frame. The decomposition transform transforms the frame, consisting of input sample values, into a set of complex FFT coefficients. After the decomposition transform, the complex FFT coefficients may be converted from Cartesian coordinates to polar coordinates. The suite of FFT coefficients for the subsequent frames constitutes the decomposed subband signal. The conversion factors used are T=2,3,…,T max For each of these, the phase angle of the FFT coefficient is multiplied by the respective conversion factor T, converted back to Cartesian coordinates, and then converted back.

[0104] Therefore, for each conversion factor T, there is a different set of complex FFT coefficients that represent a specific frame. In other words, conversion factors T = 2, 3, ..., T max For each of these, and for each frame, a separate set of FFT coefficients is determined. As a result, for each conversion order T, the composite subband signal Y(t) s k ,Ω m A different set of ) is generated.

[0105] In the synthesis stage, the combined stride Δt of the synthesis window. sThis is determined as a function of the conversion order T used in each converter. As outlined above, the time stretching operation also includes the time stretching of subband signals, i.e., the time stretching of a suite of frames. This operation is decomposed by the factor T, stride Δt a The synthetic Hop factor or synthetic stride Δt is being increased. s This can be done by selecting [this option]. As a result, the combined stride Δt for a converter of order T is [this option]. sT is Δt sT =TΔt a This is given by [formula]. Figures 8(b) and (c) show the combined stride Δt of the combined window for conversion factors T=2 and T=3, respectively. sT This shows that Δt s2 =2Δt a Δt s3 =3Δt a That is the case.

[0106] Figure 8 also shows the reference time t "stretched" by factors T=2 and T=3 in Figures 8(b) and (c), respectively, compared to Figure 8(a). r This also indicates that, however, in the output, this reference time t r The two conversion factors need to be aligned. To align the outputs, the third-order conversion signal, i.e., Figure 8(c), needs to be downsampled or rate-transformed by factor 3 / 2. This downsampling leads to harmonic conversion with respect to the second-order conversion signal. Figure 9 shows the effect of this resampling on the combined stride of the window for T=3. If the decomposed signals are the output signals of the core decoder that have not been upsampled, then the signal in Figure 8(b) is effectively frequency-transformed by factor 2, and the signal in Figure 8(c) is effectively frequency-transformed by factor 3.

[0107] The following discussion addresses the aspect of time alignment of conversion sequences with different conversion factors when using a common resolution window. In other words, it addresses the aspect of aligning the output signals of frequency converters using different conversion orders. When using the method outlined above, the Dirac function δ(t-t0) is time-stretched, i.e., moved along the time axis, by an amount of time given by the applied conversion factor T. To convert the time-stretching operation into a frequency-shifting operation, decimation or downsampling is performed using the same conversion factor T. When such decimation by the conversion factor or conversion order T is performed on the time-stretched Dirac function δ(t-Tt0), the downsampled Dirac pulse is time-aligned with respect to the zero reference time 710 at the center of the first resolution window 701. This is shown in Figure 7.

[0108] However, when using different conversion orders T, decimation leads to different offsets with respect to the zero reference unless the zero reference is aligned with the "zero" time of the input signal. As a result, time offset adjustment of the decimated converted signals must be performed before they can be summed in the summing unit 502. As an example, consider a first converter of order T=3 and a second converter of order T=4. Furthermore, assume that the output signal of the core decoder is not upsampled. Then the converter decimates the third-order time-stretched signal by factor 3 / 2 and the fourth-order time-stretched signal by factor 2. The second-order time-stretched signal, i.e., T=2, is interpreted at the end as having a higher sampling frequency than the input signal, i.e., twice as high a sampling frequency, effectively pitch-shifting the output signal by factor 2.

[0109] It can be shown that in order to align the converted and downsampled signals, a time offset of (T-2)L / 4 must be added to the converted signal before decimation. That is, for third-order and fourth-order conversions, offsets of L / 4 and L / 2, respectively, must be applied. To verify this with a concrete example, let the zero reference for a second-order time-stretched signal correspond to the time or sample value L / 2, i.e., the zero reference 710 in Figure 7. This is because decimation is not used. For a third-order time-stretched signal, the reference shifts to (L / 2)(2 / 3)=L / 3 due to downsampling by factor 3 / 2. If the time offset according to the above rule is added before decimation, the reference shifts to ((L / 2)+(L / 4))(2 / 3)=L / 2. This means that the reference of the downsampled converted signal is aligned with the zero reference 710. Similarly, for a fourth-order conversion without offset, the zero reference corresponds to (L / 2)(1 / 2)=L / 4, but when using the proposed offset, the reference shifts to ((L / 2)+(L / 2))(1 / 2)=L / 2. This, too, aligns with the second-order zero reference 710, i.e., the zero reference for a converted signal using T=2.

[0110] Another aspect to consider when using multiple conversion orders simultaneously concerns the gain applied to conversion sequences of different conversion factors. In other words, we may address the aspect of combining the output signals of converters of different conversion orders. When selecting the gain of the converted signal, there are two principles, which can be considered under different theoretical approaches. In one option, the converted signal is energy-conserving, meaning that the total energy in the low-band signal that constitutes the high-band signal that is subsequently converted and multiplied by T is conserved. In this case, the energy per bandwidth should be reduced by the conversion factor T, because the signal is stretched by the same amount T at frequency. However, a sine wave with energy in an infinitesimally small bandwidth retains its energy after conversion. This is due to the fact that, just as the Dirac pulse is moved in time by the converter during time stretching, i.e., the duration of the pulse is not changed by the time stretching operation, the sine wave is moved in frequency when converted, i.e., the duration at frequency (i.e., bandwidth) is not changed by the frequency conversion operation. In other words, even if the energy per unit bandwidth decreases by a factor of T, a sine wave still has all its energy at a single point in frequency, and the energy at each point is conserved.

[0111] Another option when selecting the gain of the converted signal is to preserve the energy per bandwidth after conversion. In this case, broadband white noise and transient signals exhibit a flat frequency response after conversion, while the energy of sinusoidal waves increases by factor T.

[0112] A further aspect of the present invention is the selection of decomposition and synthesis phase vocoder windows when using a common decomposition window. Decomposition and synthesis phase vocoder windows, i.e., v a (n) and v s It is beneficial to carefully select (n). A composite window v is used to allow for complete reconstruction. s (n) should not only follow formula 2 above, but also the decomposition window v a(n) should also have sufficient sidelobe level blocking. Otherwise, undesirable "aliasing" terms will typically be heard as interference with the principal term for a frequency-varying sine wave. Such undesirable "aliasing" terms can also appear for stationary sine waves in the case of even conversion factors, as described above. The present invention proposes the use of a sinusoidal window for a good sidelobe blocking ratio. Thus, the resolution window is v a (n) = sin{(π / L)(n+0.5)} 0≦n <L (4) It is proposed that this be done.

[0113] Synthetic hop size Δt s If is not a divisor of the decomposition window length L, i.e., if the decomposition window length L cannot be divided by the composite hop size, then composite window v s (n) is the decomposition window v a It is either identical to (n) or given by formula (2) above. For example, L=1024, Δt s If =384, then 1024 / 384 = 2.667 is not an integer. It should be noted that, as outlined above, it is also possible to select a pair of biorthogonal decomposition and composite windows. This can be useful for reducing aliasing in the output signal, especially when using even conversion order T.

[0114] In the following, Figures 10 and 11 show exemplary encoder 1000 and exemplary decoder 1100 for Integrated Speech Acoustic Coding (USAC), respectively. The general structure of the USAC encoder 1000 and decoder 1100 is described as follows: First, there may be common pre-processing / post-processing consisting of an MPEG Surround (MPEGS) function unit for handling stereo or multi-channel processing and enhanced spectral band replication (eSBR) units 1001 and 1101 for handling the parametric representation of higher audio frequencies in the input signal. The eSBR may utilize the harmonic conversion method outlined in this paper. There are two branches, one consisting of a modified Advanced Audio Coding (AAC) toolpath and the other consisting of a linear predictive coding (LP or LPC domain) based path. The latter features a frequency-domain or time-domain representation of the LPC residual. All transmitted spectra for both AAC and LPC may be represented in the MDCT domain and then quantized and arithmetic coded. The time-domain representation may use the ACELP excitation coding scheme.

[0115] The improved spectral band replication (eSBR) unit 1001 of the encoder 1000 may have the high-frequency reconstruction system outlined in this paper. In some embodiments, the eSBR unit 1001 may have the conversion unit outlined in the context of Figures 4, 5, and 6. Encoded data related to harmonic conversion, such as the conversion order used, the amount of frequency-domain oversampling required, or the gain used, may be derived in the encoder 1000, merged with other encoded information in a bitstream multiplexer, and transferred to the corresponding decoder 1100 as an encoded audio stream.

[0116] The decoder 1100 shown in Figure 11 also has an improved spectral bandwidth replication (eSBR) unit 1101. This eSBR unit 1101 receives an encoded audio bitstream or encoded signal from encoder 1000, generates high-frequency components or high bands of the signal using the method outlined herein, and merges them with the decoded low-frequency components or low bands to produce a decoded signal. The eSBR unit 1101 may have various components outlined herein. In particular, it may have a conversion unit outlined in the context of Figures 4, 5, and 6. The eSBR unit 1101 may use information about the high-frequency components provided by encoder 1000 via the bitstream to perform high-frequency reconstruction. Such information may include the spectral envelope of the original high-frequency components, the conversion order used, the amount of frequency-domain oversampling required, or the gain used, for generating the high-frequency components of the composite subband signal and thus the decoded signal.

[0117] Furthermore, Figures 10 and 11 show the following possible additional components of the USAC encoder / decoder:

[0118] • Bitstream payload demultiplexer tool. This separates the bitstream payload into parts for each tool, providing each tool with the bitstream payload information relevant to that tool.

[0119] • A noiseless scale factor decoding tool. This tool receives information from a bitstream payload demultiplexer, parses that information, and decodes the Huffman and DPCM encoded scale factors.

[0120] • Spectrum noiseless decoding tool. This tool receives information from a bitstream payload demultiplexer, parses that information, decodes the arithmetic encoded data, and reconstructs the quantized spectrum.

[0121] • Inverse quantization tool. This takes quantized values ​​for a spectrum and converts the integer values ​​into an unscaled, reconstructed spectrum. This quantizer is preferably a compression / decompression quantizer, and its compression / decompression factor depends on the chosen core coding mode.

[0122] • Noise filling tool. This is used to fill spectral gaps in a decoded spectrum. These spectral gaps appear when spectral values ​​are quantized to zero, for example, due to strong bit demand constraints in an encoder.

[0123] • Rescaling tool. This converts the integer representation of the scaling factor to its actual value and multiplies the unscaled, inversely quantized spectrum by the associated scaling factor.

[0124] M / S tools as described in ISO / IEC 14496-3.

[0125] • Temporal noise shaping (TNS) tools as described in ISO / IEC 14496-3.

[0126] • Filter bank / block switching tool. This applies the inverse of the frequency mapping performed in the encoder. For the filter bank tool, the inverse modified discrete cosine transform (IMDCT) is preferably used.

[0127] • Time-Distorted Filter Bank / Block Switching Tool. This replaces the normal filter bank / block switching tool when time distortion mode is enabled. The filter bank is preferably the same as the normal filter bank (IMDCT), and furthermore, windowed time-domain samples are mapped from the distorted time domain to the linear time domain by time-varying resampling.

[0128] • MPEG Surround (MPEGS) tool. This tool generates multiple signals from one or more input signals by applying a sophisticated upmix procedure to the input signals, which are controlled by appropriate spatial parameters. In the context of USAC, MPEGS is preferably used to encode multi-channel signals by transmitting parametric subinformation along with the downmixed signals that are transmitted.

[0129] • Signal classifier tool. This tool analyzes the original input signal and then generates control information that triggers the selection of various encoding modes. The analysis of the input signal is typically implementation-dependent and attempts to select the optimal core encoding mode for a given input signal frame. The output of the signal classifier may optionally also be used to influence the behavior of other tools, such as MPEG surround, enhanced SBR, and time distortion filter banks.

[0130] • LPC filter tool. This generates a time-domain signal from an excitation-domain signal by filtering the reconstructed excitation signal through a linear predictive synthesis filter.

[0131] • ACELP tool. This provides an efficient method for representing time-domain excitation signals by combining long-term predictors (adaptive codewords) with pulse-like sequences (innovation codewords).

[0132] Figure 12 shows an embodiment of the eSBR unit shown in Figures 10 and 11. The eSBR unit 1200 is described below in the context of a decoder, and the input to the eSBR unit 1200 is the low-frequency component of a signal, also known as the low-band.

[0133] In Figure 12, the low-frequency component 1213 is input to the QMF filter bank to generate QMF frequency bands. These QMF frequency bands should not be confused with the decomposed subbands outlined in this paper. The QMF frequency bands are used to manipulate and merge the low-frequency and high-frequency components of a signal in the frequency domain, not the time domain. The low-frequency component 1214 is input to the conversion unit 1204, which corresponds to the system for high-frequency reconstruction outlined in this paper. The conversion unit 1204 generates a high-frequency component 1212, also known as the high band of the signal, which is converted to the frequency domain by the QMF filter bank 1203. Both the QMF-converted low-frequency component and the QMF-converted high-frequency component are input to the manipulation and merging unit 1205. This unit 1205 may also perform envelope tuning of the high-frequency component to combine the tuned high-frequency and low-frequency components. The combined output signal is converted back to the time domain by the inverse QMF filter bank 1201.

[0134] Typically, the QMF filter bank 1202 has 32 QMF frequency bands. In such a case, the low-frequency component 1213 has a bandwidth f s It has / 4. Here, f s / 2 is the sampling frequency of signal 1213. The high-frequency component 1212 is the bandwidth f. s It has a / 2 and is filtered through QMF bank 1203 which has 64 QMF frequency bands.

[0135] This paper has outlined a method for harmonic conversion. This harmonic conversion method is particularly suitable for the conversion of transient signals. The method involves a combination of frequency-domain oversampling and harmonic conversion using a vocoder. The conversion operation depends on a combination of the decomposition window, decomposition window stride, conversion size, combination window, and combination window stride, as well as the phase adjustment of the decomposed signal. By using this method, undesirable effects such as pre-echo and post-echo can be avoided. Furthermore, the method does not use signal analysis measures such as transient signal detection, which typically introduce signal distortion due to discontinuities in signal processing. Moreover, the proposed method has reduced computational complexity. The harmonic conversion method based on the present invention can be further improved by appropriate selection of the decomposition / combination window, gain value, and / or time alignment.

[0136] Several aspects are described below. [Aspect 1] A system that generates an output signal from an input signal using a conversion factor T: A decomposition window unit that applies a decomposition window of length La, thereby extracting the frame of the input signal; • A decomposition transformation unit of order M that transforms sample values ​​into M complex coefficients; A nonlinear processing unit that changes the phase of the complex coefficient by using a conversion factor T; A composite transformation unit of order M that transforms the modified coefficient into M modified sample values; The system includes a composite window unit that applies a composite window of length Ls to the M modified sample values ​​to generate a frame of the output signal, M is based on the conversion factor T. system. [Aspect 2] The system according to embodiment 1, wherein the difference between M and the average length of the decomposition window and the composite window is proportional to (T-1). [Aspect 3] The system according to embodiment 2, wherein M is (TLa + Ls) / 2 or greater. [Aspect 4] The decomposition transform unit performs one of the following: Fourier transform, fast Fourier transform, discrete Fourier transform, or wavelet transform; The synthesis unit performs the corresponding inverse transformation. A system as described in any one of the descriptions in 1 to 3. [Aspect 5] A decomposition stride unit that shifts the decomposition window by a decomposition stride of Sa sample values ​​in accordance with the input signal; A composite stride unit that shifts a series of frames of the output signal by a composite stride equal to Ss of sample values; The system further comprises a superposition summing unit that superimposes and adds a series of shifted frames from the composite stride unit to generate the output signal. A system as described in any one of the descriptions in 1 to 4. [Aspect 6] The combined stride is T times the decomposed stride; The output signal corresponds to the input signal stretched in time by the conversion factor T. The system described in aspect 5. [Aspect 7] The system according to any one of embodiment 5 or 6, wherein the composite window is derived from the decomposition window and the decomposition stride. [Aspect 8] The aforementioned composite window is official

number

number

number

Claims

1. An audio signal processing device that converts an input audio signal with a conversion factor T to generate an output audio signal, wherein the audio signal processing device: The steps include: extracting frames of L time-domain sample values ​​of the input audio signal using a decomposition window of length L with the function v(n) = sin((π / L)(n + 0.5)) and 0 ≤ n < L; A step of converting the L time-domain samples into M complex frequency-domain coefficients, the step of converting the L time-domain samples into M complex frequency-domain coefficients, the step of determining a frequency-domain oversampling factor F, and determining M according to L and F; A step of changing the phase of the complex frequency domain coefficient using the conversion factor T; The steps include: converting the modified frequency-domain coefficients into M modified time-domain samples; It has one or more components that perform the steps of generating L frames of time-domain output sample values ​​of the output audio signal from the M modified time-domain sample values ​​using a composite window, M = F * L, where F is the frequency-domain oversampling factor determined in response to the frequency-domain oversampling information received in the encoded bitstream. The frame of L time-domain output sample values ​​of the output audio signal includes multiple high-frequency components that are not present in the frame of L time-domain sample values ​​of the input audio signal, at least one of the high-frequency components is generated using a conversion factor T, and at least one of the other high-frequency components is generated using a second conversion factor T 2 It is generated using, and T is T 2 not equal to, Audio signal processing unit.

2. The audio signal processing apparatus according to claim 1, wherein the oversampling factor F is (T+1) / 2 or greater, and the conversion factor T is an integer greater than 1.

3. The audio signal processing apparatus according to claim 1, wherein the phase change includes multiplying the phase by a conversion factor T.

4. The audio signal processing apparatus according to claim 1, wherein the decomposition window has a length L, along with zero padding of an additional (F-1)*L zeros.

5. The aforementioned one or more components further: The steps include: shifting the decomposition window by the decomposition stride along the input audio signal to generate a series of frames of the input audio signal; The first step is to shift a series of frames of L time-domain output sample values ​​by the combined stride; The process involves performing the steps of generating the output audio signal by superimposing and summing a series of shifted frames of L time-domain output sample values. The audio signal processing device according to claim 1.

6. The audio signal processing apparatus according to claim 5, wherein one or more of the components further increase the sampling rate of the output audio signal by a conversion factor T to produce a converted output audio signal.

7. The audio signal processing apparatus according to claim 6, wherein the combined stride is T times the decomposed stride.

8. A method performed by an audio signal processing device, which converts an input audio signal with a conversion factor T to generate an output audio signal, wherein the method is: The steps include: extracting frames of L time-domain sample values ​​of the input audio signal using a decomposition window of length L with the function v(n) = sin((π / L)(n + 0.5)) and 0 ≤ n < L; A step of converting the L time-domain samples into M complex frequency-domain coefficients, the step of converting the L time-domain samples into M complex frequency-domain coefficients, the step of determining a frequency-domain oversampling factor F, and determining M according to L and F; A step of changing the phase of the complex frequency domain coefficient using the conversion factor T; The steps include: converting the modified frequency-domain coefficients into M modified time-domain samples; The process includes the step of generating a frame of L time-domain output sample values ​​of the output audio signal from the M modified time-domain sample values ​​using a composite window, M = F * L, where F is a frequency-domain oversampling factor determined in response to the frequency-domain oversampling information received in the encoded bitstream. The frame of L time-domain output sample values ​​of the output audio signal includes multiple high-frequency components that are not present in the frame of L time-domain sample values ​​of the input audio signal, at least one of the high-frequency components is generated using a conversion factor T, and at least one of the other high-frequency components is generated using a second conversion factor T 2 It is generated using, and T is T 2 not equal to, method.

9. The method according to claim 8, wherein converting the L time-domain sample values ​​into M complex frequency-domain coefficients is performed by performing one of the Fourier transform, fast Fourier transform, discrete Fourier transform, or wavelet transform.

10. The method according to claim 8, wherein the oversampling factor F is (T+1) / 2 or greater, and the conversion factor T is an integer greater than 1.

11. The method according to claim 8, wherein the input audio signal includes low-frequency components of the audio signal.

12. A non-temporary computer-readable medium having instructions for execution by an audio signal processing device, wherein, when executed by the audio signal processing device, the instructions cause the audio signal processing device to perform the method described in claim 8.

13. A computer program for causing a computer to perform the method described in claim 8.

Citation Information

Patent Citations

  • Source coding enhancement using spectral-band replication

    EP0940015B1

  • Enhancing Primitive Coding Using Spectral Band Duplication

    JP2001521648A

  • Partitioned fast convolution in time and frequency domain

    JP2008020913A

  • Low-delay transform coding using weighting windows

    WO2008081144A2

  • Device and method for a bandwidth extension of an audio signal

    WO2009095169A1