Device and method for processing audio signal using harmonic post filter

The harmonic post-filter addresses inter-harmonic noise in transform-based audio codecs by utilizing a transfer function with multi-tap filters based on pitch lag components, improving audio quality without pre-filters.

JP2025131828APending Publication Date: 2025-09-09FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2025099285
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2014-07-28
Filing Date
2025-06-13
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Transform-based audio codecs introduce inter-harmonic noise, especially at low bit rates, which significantly degrades the subjective quality of harmonic audio signals due to poor frequency resolution and selectivity, and existing solutions based on prediction techniques fail to adequately address this issue.

Method used

A harmonic post-filter is applied in the decoder using a transfer function with a numerator and denominator, where the denominator includes a multi-tap filter dependent on the integer and fractional parts of the pitch lag, effectively reducing inter-harmonic noise without requiring pre-filters.

Benefits of technology

The post-filter significantly improves the subjective quality of audio signals by accurately removing inter-harmonic noise, enhancing the performance of transform-based audio codecs, especially at low bit rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025131828000001_ABST
    Figure 2025131828000001_ABST
Patent Text Reader

Abstract

To provide an audio processing device, a method and a system, which use a harmonic post filter.SOLUTION: A device for processing an audio signal associated with pitch lag information and gain information includes: a domain transformer (100) for transforming representation of a first domain of the audio signal into representation of a second domain of the audio signal; and a harmonic post filter (104) based on a transfer function including a numerator including a gain value indicated by gain information and a denominator including a multi-tap filter depending on an integer portion of a pitch lag and a fraction portion of the pitch lag, which are indicated by pitch lag information, for filtering representation of the second domain of the audio signal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to audio processing, and in particular to audio processing using harmonic postfilters. [Background technology]

[0002] Transform-based audio codecs generally introduce inter-harmonic noise when processing harmonic audio signals, especially at low bit rates.

[0003] This effect is exacerbated when transform-based audio codecs operate at low delay due to the poor frequency resolution and / or selectivity brought about by shorter transform sizes and / or poor window frequency response.

[0004] This inter-harmonic noise is generally perceived as a very objectionable artifact and significantly reduces the performance of transform-based audio codecs when subjectively evaluated against highly tonal audio material.

[0005] Several solutions exist to improve the subjective quality of transform-based audio codecs for harmonic audio signals, all of which are based on prediction-based techniques, either in the transform domain or in the time domain.

[0006] An example of the transform domain approach is as follows.

[0007] ·[1]H.Fuchs “Improving MPEG Audio Coding by Backward Adaptive Linear Stereo Prediction” (99th AES Convention, New York 1995, Preprint 4086) ·[2] L. Yin, M. Suonio, M. Vaananen “A New Backward Predictor for MPEG Audio Coding” (103rd AES Convention, New York 1997, Preprint 4521) ·[3]Juha Ojanpera, Mauri Vaananen, Lin Yin "Long Term Predictor for Transform Domain Perceptual Audio Coding" (107th AES Convention, New York 1999, Preprint 5036) An example of a time domain approach is as follows.

[0008] [4] Philip J. Wilson, Harprit Chhatwal, "Adaptive transform coder having long term predictor," U.S. Patent No. 5,012,517, April 30, 1991. ·[5] Jeongook Song, Chang-Heon Lee, Hyen-O Oh, Hong-Goo Kang “Harmonic Enhancement in Low Bitrate Audio Coding Using and Efficient Long-Term Predictor” (EURASIP Journal on Advances in Signal Processing 2010) [6] Juin-Hwey Chen, “Pitch-based pre-filtering and post-filtering for compression of audio signals,” U.S. Patent No. 8,738,385, May 27, 2014. [Prior art documents] [Patent documents]

[0009] [Patent Document 1] U.S. Patent No. 5,012,517 [Patent Document 2] U.S. Patent No. 8,738,385 [Non-patent literature]

[0010] [Non-Patent Document 1] H.Fuchs "Improving MPEG Audio Coding by Backward Adaptive Linear Stereo Prediction" (99th AES Convention, New York 1995, Preprint 4086) [Non-patent document 2] L. Yin, M. Suonio, M. Vaananen "A New Backward Predictor for MPEG Audio Coding" (103rd AES Convention, New York 1997, Preprint 4521) [Non-patent document 3] Juha Ojanpera, Mauri Vaananen, Lin Yin "Long Term Predictor for Transform Domain Perceptual Audio Coding" (107th AES Convention, New York 1999, Preprint 5036) [Non-patent document 4] Jeongook Song, Chang-Heon Lee, Hyen-O Oh, Hong-Goo Kang "Harmonic Enhancement in Low Bitrate Audio Coding Using and Efficient Long-Term Predictor" (EURASIP Journal on Advances in Signal Processing 2010) Summary of the Invention [Problem to be solved by the invention]

[0011] The object of the present invention is to provide an improved concept for processing audio signals. [Means for solving the problem]

[0012] This object is achieved by an apparatus for processing an audio signal according to claim 1, a method for processing an audio signal according to claim 12, a system according to claim 13, a method for operating a system according to claim 17 or a computer program according to claim 18.

[0013] The present invention is based on the finding that the subjective quality of an audio signal can be significantly improved by using a harmonic post-filter having a transfer function including a numerator and a denominator, the numerator of the transfer function including a gain value indicated by the transmitted gain information, and the denominator including a multi-tap filter that depends on the integer part of the pitch lag indicated by the pitch lag information and the fractional part of the pitch lag.

[0014] Thus, it is possible to remove interharmonic noise, which is introduced as an artifact by typical domain-modified audio decoders. This harmonic postfilter is particularly useful in that it relies on transmitted information, i.e., pitch gain and pitch lag, which are available in the decoder anyway. This information is received from the corresponding encoder via the decoder input signal. Furthermore, due to the fact that not only the integer part of the pitch lag is taken into account, but also the fractional part of the pitch lag, the postfiltering is particularly accurate. The fractional part of the pitch lag can be introduced into the postfilter via a multi-tap filter, in particular, with filter coefficients that depend on the fractional part of the pitch lag. This filter can be implemented as an FIR filter, or any other filter, such as an IIR filter or a different filter implementation. Any domain modification, such as a time-frequency modification, an LPC-time modification, a time-LPC modification, or a frequency-time modification, can be advantageously improved by the inventive postfilter concept. However, preferably, the domain modification is a frequency-time domain modification.

[0015] Thus, embodiments of the present invention reduce the inter-harmonic noise introduced by transform audio codecs that are based on long-term predictors operating in the time domain. In contrast to

[04] -[6], in which both a pre-filter before transform coding and a post-filter after transform decoding are used, the present invention preferably applies only a post-filter.

[0016] Furthermore, it has been found that the prefilters used in

[04] -[6] tend to introduce instabilities into the input signal provided to the transform coder. These instabilities are due to frame-to-frame changes in gain and / or pitch lag. Transform coders have trouble encoding such instabilities, especially at low bit rates, and sometimes introduce even more noise into the decoded signal compared to the situation without any pre- or post-filter.

[0017] Preferably, the present invention does not utilize any pre-filters, thus avoiding the problems associated with pre-filters entirely.

[0018] Furthermore, the present invention relies on a postfilter applied to the decoded signal after transform coding, which is based on a long-term prediction filter that takes into account the integer and fractional parts of the pitch lag, reducing the inter-harmonic noise introduced by the transform audio codec.

[0019] For better robustness, the post-filter parameters pitch lag and pitch gain are estimated at the encoder side and transmitted in the bitstream, however, in other embodiments, the pitch lag and pitch gain can also be estimated at the decoder side based on a decoded audio signal obtained by an audio decoder comprising a frequency-to-time transformer for converting a frequency representation of the audio signal into a time-domain representation of the audio signal.

[0020] In a preferred embodiment, the numerator further comprises a multi-tap filter for the zero fractional part of the pitch lag to compensate for the spectral tilt introduced by the multi-tap filter in the denominator that depends on the fractional part of the pitch lag.

[0021] Preferably, the post-filter is configured to suppress an amount of energy between harmonics within a frame, the amount of energy suppressed being less than 20% of the total energy of the time domain representation within the frame.

[0022] In a further embodiment, the denominator comprises a product between a multi-tap filter and a gain value.

[0023] In a further embodiment, the filter numerator further includes a product of the first scalar value and the second scalar value, and the denominator includes only the second scalar value, not the first scalar value. These scalar values ​​are set to predetermined values ​​greater than 0 and less than 1, and the second scalar value is smaller than the first scalar value. Thus, it is possible to very efficiently set the generally undesirable energy removal characteristics and, in addition, the filter strength, i.e., how strongly the filter attenuates inter-harmonic artifacts in the transform domain decoder output signal.

[0024] In a preferred embodiment the device further comprises a filter controller for setting the at least second scalar value depending on the bit rate, such that a higher value is set for a lower bit rate and vice versa.

[0025] Furthermore, the filter controller is configured to signal-adaptively select a corresponding multi-tap filter depending on the fractional part of the pitch lag, so as to set the harmonic post-filter in a signal-adaptive manner, i.e., depending on an actually given fractional part value of the pitch lag.

[0026] Preferred embodiments of the present invention will now be discussed in the context of the accompanying drawings. [Brief explanation of the drawings]

[0027] [Figure 1] 1 illustrates an embodiment of an apparatus of the present invention for processing an audio signal; [Figure 2] FIG. 1 illustrates a preferred embodiment of a harmonic post-filter represented as a transfer function in the z-domain. [Figure 3] FIG. 10 shows a further preferred embodiment of the harmonic post-filter represented as a transfer function in the z-domain. [Figure 4] 2 shows a preferred embodiment of an encoder for generating an encoded signal to be decoded by the transform domain audio decoder shown in FIG. 1; [Figure 5] FIG. 1 shows a preferred implementation of a multi-tap filter as an FIR filter controlled by a filter controller. [Figure 6] FIG. 10 illustrates the cooperation between the filter controller and the memory pre-stored tap weights depending on the fractional part. [Figure 7a] FIG. 1 illustrates the frequency response of a filter with a zero alpha value. [Figure 7b] FIG. 10 illustrates the frequency response of a preferred harmonic post-filter with an α value equal to 1. [Figure 7c] FIG. 10 illustrates the frequency response of a preferred harmonic postfilter with an α value of 0.8. [Figure 8a] FIG. 1 illustrates a preferred embodiment of a harmonic post-filter with a β value equal to 0.4. [Figure 8b] FIG. 10 illustrates the frequency response of a harmonic postfilter with a β value of 0.2. DETAILED DESCRIPTION OF THE INVENTION

[0028] 1 shows an apparatus for processing an audio signal having associated pitch lag and gain information. This gain information can be transmitted to a decoder 100 via a decoder input 102 that receives the encoded signal, or alternatively, this information can be calculated within the decoder itself when this information is not available. However, for more robust operation, it is preferable to calculate the pitch lag and pitch gain information at the encoder side.

[0029] The decoder 100 includes, for example, a frequency-to-time converter for converting a frequency-time representation of an audio signal into a time-domain representation of the audio signal. The decoder is therefore not a pure time-domain speech codec, but rather a pure transform-domain decoder or a mixed transform-domain decoder, or any other coder operating in a domain different from the time domain. Furthermore, the second domain is preferably the time domain.

[0030] The apparatus further comprises a harmonic post-filter 104 for filtering the time-domain representation of the audio signal, the harmonic post-filter being based on a transfer function including a numerator and a denominator, in particular the numerator including a gain value indicated by the gain information and the denominator including the integer part of the pitch lag indicated by the pitch lag information, and importantly further including a multi-tap filter dependent on the fractional part of the pitch lag.

[0031] A preferred embodiment of this harmonic postfilter, with transfer function H(z), is shown in FIG. 2. This filter receives a decoder output signal 106 and subjects it to a postfiltering operation to obtain a postfiltered output signal 108. This postfiltered output signal can be output as a processed signal or can be further processed by any procedure to remove any discontinuities introduced by the postfiltering operation, which, of course, is signal-dependent, i.e., can vary from frame to frame. This discontinuity removal operation can be any of the known discontinuity removal operations, such as crossfading, which means that the previous frame is faded out while the new frame is simultaneously faded in, and preferably the fading characteristics are such that the fading coefficient is 1 throughout the crossfading operation. However, other discontinuity removal methods, such as low-pass filtering or LPC filtering, can also be applied.

[0032] 1 further comprises a multi-tap filter information storage device 112 and a filter controller 114. In particular, the filter controller 114 receives side information 116 from the decoder 100, which side information includes, for example, pitch gain information g as well as pitch lag information, i.e., the integer part of the pitch lag T int and the fractional part of the pitch lag, T frThis information sets the per-frame harmonic post-filter, and in addition, the multi-tap filter information B(z,T fr ) can be used by the filter control unit 114 to select the scalar values ​​α, β for a particular encoder and / or decoder setting, especially with respect to bit rate and sampling rate.

[0033] 2 shows a pole / zero representation of a filter transfer function H(z) in the z-domain, as is known in the art. Of course, there are many other representations of harmonic postfilters, all of which are filter representations and can be converted to this type of pole / zero representation in the z-domain. Thus, the present invention is applicable to each filter that can be described in any way by a transfer function such as those illustrated herein.

[0034] FIG. 3 shows a preferred embodiment of the harmonic postfilter, also written as a transfer function in pole / zero notation in the z-domain.

[0035] This filter can be written as follows:

number

[0036] B(z,0) in the fraction of H(z) is B(z,T fr ) is used to compensate for the tilt introduced by

[0037] β is used to control the strength of the post-filter. β equal to 1 produces the full effect, suppressing the maximum possible amount of energy between harmonics. β equal to 0 disables the post-filter. Typically, a very low value is used to avoid suppressing too much energy between harmonics. This value may also depend on the bitrate, with higher values ​​at lower bitrates, for example 0.4 at low bitrates and 0.2 at high bitrates.

[0038] α is used to add a slight slope to the frequency response of H(z) to compensate for the slight loss of energy at low frequencies. The value of α is typically chosen to be close to 1, e.g., 0.8.

[0039] B(z,T fr An example of B(z,T fr The order and coefficients of ) may also depend on the bit rate and the output sampling rate. For each combination of bit rate and output sampling rate, a different frequency response can be designed and adjusted.

[0040] In particular, it has been found that even values ​​of α between 0.6 and less than 1.0 are useful, and in addition, values ​​of β between 0.1 and 0.5 have also proven useful.

[0041] Furthermore, a multi-tap filter can have a variable number of taps. For a particular embodiment, one tap is z +1 It has been found that four taps, where , are sufficient. However, smaller filters with only two taps, or even larger filters with five or more taps, may be useful for certain implementations.

[0042] FIG. 6 shows a preferred embodiment of the filter B(z) for different fractional values ​​of pitch lag, specifically for a quarter pitch lag resolution. For this embodiment, four different filter descriptions are shown for a multi-tap filter in the denominator of the harmonic post-filter transfer function. However, it has been found that the filter coefficients do not necessarily have to exactly represent the values ​​shown in FIG. 6, and that a constant variation of ±0.05 may also be useful in other embodiments.

[0043] In particular, as shown in Figure 1, the tap weights shown in Figure 6 are stored in memory 112 for multi-tap filter information. Filter controller 114 receives fractional part T from line 116 in Figure 1. fr and, in response to this value, addresses memory 112 to extract specific filter information for a specific fractional portion of the pitch lag via extraction line 200. This information is then forwarded to harmonic postfilter 104 via output line 202 so that the harmonic postfilter can be accurately configured. A specific implementation of a multi-tap FIR filter is shown in FIG. 5. Weight indications w1 through w4 correspond to the notation in FIG. 6, and filter controller 114 applies the corresponding weights for a specific audio frame in response to the actual fractional portion of the pitch lag. Other sections, such as delay sections 501, 502, and 503 and combiner 505, may be implemented as shown. It is emphasized that delay value 501 is a negative delay value in z-notation, since FIR filter representations with negative delay values ​​in addition to positive delay values ​​such as 503 and 504 have proven particularly useful in this context.

[0044] Subsequently, a preferred encoder embodiment having specific functional blocks and operating without any pre-filter is shown in Figure 4. The filter section shown in Figure 4 comprises a pitch estimator 402, a pitch refiner 404, a fractional part estimator 406, a transient detector 408, a gain estimator 410, and a gain quantizer 412. Information provided by decision bits generated by the gain quantizer 412, the fractional part estimator 406, the pitch refiner 404, and the transient detector 408 is input to an encoded signal former 414. The encoded signal former provides an encoded signal 102, which is then input to the decoder 100 shown in Figure 1. The encoded signal 102 includes additional signal information not shown in Figure 4.

[0045] Next, the function of pitch estimator 402 will be described.

[0046] One pitch lag (integer part + fractional part) is estimated per frame (frame size is e.g. 20 ms). This is done in three steps to reduce complexity and improve estimation accuracy.

[0047] A pitch analysis algorithm that generates a smooth pitch evolution contour is used (e.g., open-loop pitch analysis as described in Rec. ITU-T G.718, sec. 6.6). This analysis is typically performed on a subframe-by-subframe basis (the subframe size is, for example, 10 ms) and produces one pitch lag estimate per subframe. Note that these pitch lag estimates do not have any fractional parts and are typically estimated for a downsampled signal (the sampling rate is, for example, 6400 Hz). The signal used can be any audio signal, for example, an LPC-weighted audio signal as described in Rec. ITU-T G.718, sec. 6.5.

[0048] The pitch refiner operates as follows.

[0049] The final integer part of the pitch lag is estimated for the audio signal x[n], typically performed at a core coder sampling rate higher than the sampling rate of the downsampled signal used in a (e.g., 12.8 kHz, 16 kHz, 32 kHz...). The signal x[n] can be any audio signal, for example, an LPC-weighted audio signal.

[0050] The integer part of the pitch lag is then the lag d that maximizes the autocorrelation function m and

number

[0051] (Number 3) T-δ1≦d≦T+δ2 The fractional part estimator 406 operates as follows.

[0052] The fractional part is found by interpolating the autocorrelation function C(d) calculated in step 2.b. and selecting the fractional pitch lag that maximizes the interpolated autocorrelation function. The interpolation can be performed using a low-pass FIR filter such as that described in Rec. ITU-T G.718, sec. 6.6.7.

[0053] The transient detector 408 shown in FIG. 4 is configured to generate a decision bit.

[0054] If the input audio signal does not contain any harmonic components, no parameters are coded in the bitstream. Only one bit is transmitted so that the decoder knows whether it has to decode the post-filter parameters or not. The decision is made based on several parameters:

[0055] a. Normalized correlation at integer pitch lags estimated in step 1.b.

number

[0056] If (norm.corr(curr.)*norm.corr.(prev.))>0.25, the current frame contains some harmonic content (bit=1).

[0057] b. Features computed by the transient detector (e.g., temporal flatness measure, maximum energy change) to avoid post-filtering for signals containing transients, e.g.

[0058] If (tempFlatness>3.5 or maxEnergychange>3.5), set bit=0, otherwise do not send any parameters.

[0059] Additionally, the gain estimator 410 calculates the gain to be input to the gain quantizer 412 .

[0060] The gain is typically estimated for an input audio signal at the core coder sampling rate, but it can also be any audio signal, such as an LPC-weighted audio signal, which is denoted y[n] and may or may not be the same as x[n].

[0061] First, we obtain the prediction y[n] by filtering y[n] with the following filter: P [n] is required.

number

[0062] An example of B(z) when the pitch lag resolution is 1 / 4 is as follows:

number

number

[0063] Finally, the gain is quantized, eg, with 2 bits, using, eg, uniform quantization.

[0064] If the gain is quantized to 0, then no parameter is coded in the bitstream, resulting in only one decision bit (bit=0).

[0065] As already outlined, a postfilter is applied to the output audio signal after the transform decoder. The postfilter processes the signal frame by frame with the same frame size as used at the encoder side, such as 20 ms. As shown, it is based on a long-term prediction filter H(z) whose parameters are estimated at the encoder side and determined from parameters decoded from the bitstream. This information includes the decision bit, pitch lag, and gain. If the decision bit is 0, the pitch lag and gain are not decoded, are assumed to be 0, and are not written to the bitstream at all.

[0066] As discussed, if the filter parameters differ from one frame to the next, a discontinuity may be introduced at the boundary between the two frames. To avoid the discontinuity, a discontinuity remover such as a crossfader or any other implementation for that purpose is applied.

[0067] Additionally, several different methods for configuring the harmonic postfilter are shown in Figures 7a-8b. These plots show frequency-domain transfer functions. The horizontal axis is related to normalized frequency 1, and the vertical axis is the magnitude of the filter response in dB. In all illustrations except Figure 7b, it is emphasized that the filter introduces low-frequency amplification, i.e., a constant positive dB magnitude value.

[0068] In particular, Figure 7a shows a transfer function implementing the filter of Figure 3 with constant parameter values ​​as shown above. Additionally, the α value, i.e., the first scalar value, is set to 0. Figure 7b shows a similar situation, but here the α value is equal to 1. The other parameters are identical to Figure 7a.

[0069] FIG. 7c shows a further embodiment where α is equal to 0.8, which has a slight tilt and a boost at lower frequencies. Again, FIG. 7 has the same other parameters as FIG. 7a. It becomes clear that α equal to 1 eliminates the tilt and all harmonic frequencies have a gain of 1. The drawback of this configuration is the loss of energy at interharmonic frequencies. Therefore, a value of α equal to 0.8, as in FIG. 7c, is preferred. This value adds a slight tilt compared to the situation in FIG. 7b where α is equal to 1. This slight tilt is preferably used to compensate for the loss of energy at interharmonic frequencies.

[0070] 8a and 8b show filter settings for an α value equal to 0.8 and different β values, i.e., a β value of 0.4 in Fig. 8a and a β value of 0.2 in Fig. 8b. It is clear that a β value of 0.4 has a stronger post-filtering effect compared to a β value of 0.2, and therefore, at lower bit rates, a β value of 0.4 is used to remove inter-harmonic noise introduced by such low bit rates.

[0071] On the other hand, β equal to 0.2 has a less strong effect of suppressing inter-harmonic energy, and therefore this β value is preferred for higher bit rates due to the fact that there is not as much inter-harmonic noise at such higher bit rates.

[0072] While some aspects are described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, with a block or device corresponding to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, any one or more of the most important method steps may be performed by such an apparatus.

[0073] The transmitted or encoded signals of the present invention can be stored on a digital storage medium or can be transmitted over a transmission medium, such as a wireless or wired transmission medium, such as the Internet.

[0074] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementations can be implemented using digital storage media, such as floppy disks, DVDs, Blu-ray®, CDs, ROMs, PROMs, and EPROMs, EEPROMs, or flash memories, on which electronically readable control signals are stored, which cooperate (or can cooperate) with a programmable computer system so that the respective methods are implemented. Hence, the digital storage media may be computer-readable.

[0075] Some embodiments according to the present invention include a data carrier having electronically readable control signals capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0076] Generally, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product runs on a computer. The program code may, for example, be stored on a machine-readable carrier.

[0077] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0078] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0079] A further embodiment of the inventive method is therefore a data carrier (or a non-transitory storage medium such as a digital storage medium or a computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium or storage medium is generally tangible and / or non-transitory.

[0080] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, the data stream or the sequence of signals being, for example, adapted to be transmitted via a data communication connection, for example over the Internet.

[0081] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0082] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0083] Further embodiments according to the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.

[0084] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0085] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented herein by way of description and explanation of the embodiments.

Claims

1. 1. An apparatus for processing an audio signal having associated pitch lag information and gain information, comprising: a domain transformer (100) for transforming a representation of a first domain of the audio signal into a representation of a second domain of the audio signal; a harmonic postfilter (104) for filtering the representation of the second region of the audio signal, the postfilter being based on a transfer function including a numerator and a denominator, the numerator including a gain value indicated by the gain information, and the denominator including a multi-tap filter dependent on an integer part of a pitch lag indicated by the pitch lag information and a fractional part of the pitch lag; An apparatus comprising:

2. The apparatus of claim 1 , wherein the transfer function of the postfilter includes, in the numerator, an additional multi-tap FIR filter for a zero fractional portion of the pitch lag.

3. The apparatus of claim 1 or 2, wherein the denominator comprises a product between the multi-tap filter and the gain value.

4. 4. The apparatus of claim 1, wherein the numerator further comprises a product of a first scalar value and a second scalar value, the denominator comprises the second scalar value but not the first scalar value, the first scalar value and the second scalar value are predetermined and have values ​​greater than 0 and less than 0, and the second scalar value is less than the first scalar value.

5. 5. The apparatus of claim 4, wherein a filter controller (114) is configured to set the second scalar value depending on a bit rate, whereby the frequency-to-time converter (100) is operated such that the second scalar value is set to a first value when the bit rate has a first value, and the second scalar value is set to a second value when the bit rate has a second value, the second value of the bit rate being smaller than the first value of the bit rate, and the second value of the second scalar value being larger than the first value of the second scalar value.

6. 6. The apparatus of claim 4, wherein the first scalar value is set between 0.6 and 1.0, and the second scalar value is set between 0.1 and 0.

5.

7. The post-filter has the transfer function H(z) in pole-zero representation according to the following equation: [Equation 1] where α is the first scalar value, β is the second scalar value, B(z,0) is the multi-tap filter for zero fractional partial pitch lag, and B(z,T fr ) is a multi-tap filter that depends on the fractional part of the pitch lag, and T int is the integer part of the pitch lag, and T fr 7. The apparatus of claim 1, wherein: ≡(x,y) is the fractional part of the pitch lag; g is the gain value indicated by the gain information; and z is a variable in the z-plane.

8. The apparatus of any one of claims 1 to 7, wherein the multi-tap filter is a finite impulse response (FIR) filter and has at least three taps.

9. the multi-tap filter in the denominator includes four taps, the first tap being between 0.0 and 0.1, the second tap being between 0.2 and 0.3, the third tap being between 0.5 and 0.6, and the fourth tap being between 0.2 and 0.3 for a zero fractional part; the multi-tap filter includes, for a first fractional portion, four filter taps, the first tap being between 0.0 and 0.1, the second tap being between 0.3 and 0.4, the third tap being between 0.45 and 0.55, and the fourth tap being between 0.1 and 0.2; the multi-tap filter includes four filter taps for a second fractional portion, a first tap between 0.0 and 0.1, a second tap between 0.35 and 0.45, a third tap between 0.35 and 0.45, and a fourth tap between 0.0 and 0.1; the multi-tap filter includes, for a third fractional portion, four filter taps, the first tap being between 0.1 and 0.2, the second tap being between 0.45 and 0.55, the third tap being between 0.3 and 0.4, and the fourth tap being between 0.0 and 0.1; The apparatus of any one of claims 1 to 8, wherein the third fractional portion is greater than the second fractional portion, and the second fractional portion is greater than the first fractional portion.

10. the post-filter is configured to have a negative spectral tilt to compensate for the loss of energy by the harmonic post-filter; or 10. The apparatus of claim 1, wherein the post-filter is configured to suppress an amount of energy between harmonics within a frame, the amount of energy suppressed being less than 20% of the total energy of the time domain representation within the frame.

11. the domain transformer is a frequency-to-time transformer, the first domain being the frequency domain and the second domain being the time domain; or The apparatus of any one of claims 1 to 10, wherein the domain transformer is an LPC residual-to-time transformer, the first domain being the LPC residual domain and the second domain being the time domain.

12. 1. A method for processing an audio signal having associated pitch lag information and gain information, comprising: Transforming (100) a frequency representation of the audio signal into a time domain representation of the audio signal; filtering the time-domain representation of the audio signal with a harmonic postfilter (104), the postfilter being based on a transfer function including a numerator and a denominator, the numerator including a gain value indicated by the gain information, and the denominator including a multi-tap filter dependent on an integer portion of a pitch lag indicated by the pitch lag information and a fractional portion of the pitch lag; A method comprising:

13. 1. A system for processing an audio signal, comprising an encoder for encoding an audio signal, and a decoder comprising a processor, the processor comprising: a domain transformer (100) for transforming a frequency representation of the audio signal into a time domain representation of the audio signal; a harmonic postfilter (104) for filtering the time domain representation of the audio signal, a harmonic post-filter (104) based on a transfer function including a numerator and a denominator, the numerator including a gain value indicated by gain information, and the denominator including a multi-tap filter depending on an integer part of a pitch lag indicated by pitch lag information and a fractional part of the pitch lag; A system comprising:

14. 14. The system of claim 13, wherein the encoder comprises a pitch lag calculator (402, 404, 406) for calculating integer and fractional parts of the pitch lag, a gain calculator (410, 412) for calculating the gain value, and an encoded signal former (414) for generating an encoded signal (102) including the pitch lag information and the gain information.

15. 1. A method for processing an audio signal, including a method for encoding and decoding an audio signal, comprising: Transforming (100) a frequency representation of the audio signal into a time domain representation of the audio signal; filtering the time-domain representation of the audio signal using a harmonic postfilter (104), the postfilter being based on a transfer function including a numerator and a denominator, the numerator including a gain value indicated by gain information, and the denominator including a multi-tap filter dependent on an integer portion of a pitch lag indicated by pitch lag information and a fractional portion of the pitch lag; A method comprising:

16. A computer program for carrying out the method of claim 12 or claim 15 when running on a computer or processor.

Citation Information

Patent Citations

  • Voice synthesizing method

    JP1998214100A

  • Long-period post-filter

    JP2004302257A

  • Frequency-selective pitch enhancement method and device for synthetic speech

    JP2005528647A

  • Encoding device, decoding device, encoding method and decoding method

    JP2011175278A

  • Encoding method, encoding device, program, and recording medium

    JP2013120225A