Time-Domain Super-Wideband Bandwidth Extension for Crosstalk Scenarios

JP2025504131A5Pending Publication Date: 2026-02-04VOICEAGE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024546171
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-03
Filing Date
2023-01-27
Publication Date
2026-02-04

AI Technical Summary

Benefits of technology

【0016】 クロストーク音声信号を符号化/復号する際の励振信号の時間領域帯域幅拡張のための方法およびデバイスの、前述の内容、ならびに他の目的、利点、および特徴は、ほんの一例として与えられるその例示的な実施形態についての以下の非制限的な説明を、添付の図面を参照して読めば、より明らかとなろう。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for time-domain bandwidth extension of an excitation signal when decoding a crosstalk speech signal includes the steps of decoding a high-band mixing factor received in a bitstream, and using the high-band mixing factor to mix a low-band excitation signal and a random noise excitation signal to generate a time-domain extended excitation signal. A method for time-domain bandwidth extension of an excitation signal when encoding a crosstalk speech signal includes the steps of calculating a high-band residual signal using a speech signal, calculating a time envelope of the high-band residual signal, calculating a high-band voicing factor based on the time envelope of the high-band residual signal, calculating a high-band mixing factor that can be used to mix the low-band excitation signal and the random noise excitation signal to generate a time-domain extended excitation signal, and estimating gain / shape parameters using the high-band voicing factor.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a method and device for time-domain bandwidth extension of an excitation signal when encoding / decoding a crosstalk audio signal. [Background technology]

[0002] In this disclosure and the accompanying claims: - The term "crosstalk" is generally intended to refer to an audio period in which a first audio element is overlaid on a second audio element, for example, but not limited to, a speech period when a first person speaks over a second person. - The term "low band" is intended to refer to a low frequency range. In this disclosure, the frequency ranges 0 kHz to 6.4 kHz and 0 kHz to 8 kHz are given as examples of "low band", but the frequency boundaries of the low band frequency range can obviously be modified / adapted to suit the bit rate of a codec and / or to achieve specific goals such as compliance with application-related constraints, system-related constraints, network-related constraints, and design / business-related constraints. - The term "high band" is intended to refer to a high frequency range. In this disclosure, the frequency ranges 6.4 kHz to 14 kHz and 8 kHz to 16 kHz are given as examples of "high band", but the frequency boundaries of the high band frequency range can obviously be modified / adapted to suit the bit rate of a codec and / or to achieve specific goals such as compliance with application-related constraints, system-related constraints, network-related constraints, and design / business-related constraints.

[0003] In many conventional applications, there are often situations where one person speaks over another person. As mentioned above in this specification, such situations are often referred to as "crosstalk". Crosstalk speech periods can be problematic in modern speech encoding / decoding systems. Since conventional speech encoding techniques are primarily designed and optimized for single-talk content (only one person speaking), the quality of crosstalk speech can be seriously affected by the encoding / decoding operations. As an example, the most serious problem with crosstalk speech encoding / decoding in the 3GPP EVS codec (Reference [1], the entire contents of which are incorporated herein by reference) is the occasional presence of "rattling noise". "Rattling" is a strong and unpleasant sound generated at frequencies between 8 kHz and 14 kHz, which falls within the high-band frequency range example defined above in this specification.

[0004] At low bit rates in the 3GPP EVS codec, the high-band frequency content is coded / decoded using the Super Wideband Bandwidth Extension (SWB TBE) tool described in reference [1]. Due to the limited number of bits available for the SWB TBE tool, the high-band excitation signal in the high-band frequency range is not directly coded. Instead, the low-band excitation signal in the low-band frequency range is calculated using an ACELP (Algebraic Code Excited Linear Prediction) coder (reference [2], the entire contents of which are incorporated herein by reference), then upsampled and extended to 14 kHz or 16 kHz depending on the high-band frequency range, and used as a replacement for the high-band excitation signal. If there is a mismatch between the low-band excitation signal and the high-band excitation signal, the synthesized speech may sound different compared to the original speech. When the low-band excitation signal is voiced but the high-band excitation signal is unvoiced, the synthesized speech is perceived as a rattle as defined above. The rattle problem in crosstalk content is illustrated in the spectrum plot of FIG. 1.

[0005] The plot in FIG. 1 shows the power spectrum P versus frequency f of an exemplary crosstalk speech in which two speakers pronounce different types of speech. The speech from the first speaker (speaker 1) contains mainly voiced content, while the speech from the second speaker (speaker 2) contains unvoiced intervals. Assuming a monophonic capture device, such as a smartphone or an omnidirectional microphone, the speech from the two speakers is mixed together in the capture device. As a result, the spectral content of the input speech signal recognized by the encoder will resemble a superset of the two spectra. A similar situation also occurs in multi-channel capture devices, such as stereo microphones and ambisonic microphones. If the encoder includes a downmixing module, the resulting monophonic input signal may contain different types of speech that are clearly distinguishable in the spectral domain. [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] 3GPP TS 26.445, "EVS Codec Detailed Algorithmic Description", 3GPP Technical Specification (Release 12), 2014, Sections 5.2.6.1 and 6.1.3.1 [Non-Patent Document 2] Bessette, B., Lefebvre, R., Salami, R. et al.,"Techniques for high-quality ACELP coding of wideband speech",Int. Conference EUROSPEECH 2001 Scandinavia, 7th European Conference on Speech Communication and Technology, 2nd INTERSPEECH Event,Aalborg, Denmark,September 3-7, 2001 Summary of the Invention [Means for solving the problem]

[0007] The present disclosure relates to the following aspects:

[0008] - a method for time-domain bandwidth extension of an excitation signal when decoding a crosstalk audio signal, the method comprising the steps of decoding a high-band mixing factor received in a bitstream, and mixing a low-band excitation signal and a random noise excitation signal using the high-band mixing factor to generate a time-domain bandwidth extended excitation signal.

[0009] - a method for time-domain bandwidth extension of an excitation signal when encoding a crosstalk speech signal, the method comprising the steps of: (a) calculating a high-band residual signal using the speech signal; and (b) calculating a time envelope of the high-band residual signal; and calculating a high-band voicing factor based on the time envelope of the high-band residual signal.

[0010] - A method for time-domain bandwidth extension of an excitation signal when encoding a crosstalk audio signal, the method comprising a step of calculating a high-band mixing factor that can be used to mix a low-band excitation signal and a random noise excitation signal to generate a time-domain bandwidth extended excitation signal.

[0011] - a method for time-domain bandwidth extension of an excitation signal when encoding a crosstalk speech signal, the method comprising the steps of: (a) calculating a high-band residual signal using the speech signal; (b) calculating a time envelope of the high-band residual signal; calculating a high-band voicing factor based on the time envelope of the high-band residual signal; calculating a high-band mixing factor that can be used to mix a low-band excitation signal and a random noise excitation signal to generate a time-domain bandwidth extended excitation signal; and estimating gain / shape parameters using the high-band voicing factor.

[0012] - a device for time-domain bandwidth extension of an excitation signal when decoding a crosstalk audio signal, the device comprising: a decoder of a high-band mixing factor received in a bitstream; and a mixer of a low-band excitation signal and a random noise excitation signal using the high-band mixing factor for generating a time-domain bandwidth extended excitation signal.

[0013] - a device for time domain bandwidth extension of an excitation signal when encoding a crosstalk speech signal, the device comprising: (a) a calculator for calculating a highband residual signal using the speech signal; (b) a calculator for calculating a time envelope of the highband residual signal; and a calculator for calculating a highband voicing factor based on the time envelope of the highband residual signal.

[0014] - a device for time-domain bandwidth extension of an excitation signal when encoding a crosstalk audio signal, the device comprising: a calculator for calculating a high-band mixing factor that can be used to mix a low-band excitation signal and a random noise excitation signal to generate a time-domain bandwidth extended excitation signal.

[0015] - a device for time-domain bandwidth extension of an excitation signal when encoding a crosstalk speech signal, the device comprising: a calculator for (a) calculating a high-band residual signal using a speech signal and (b) calculating a time envelope of the high-band residual signal; a calculator for calculating a high-band voicing factor based on the time envelope of the high-band residual signal; a calculator for calculating a high-band mixing factor that can be used to mix a low-band excitation signal and a random noise excitation signal to generate a time-domain bandwidth extended excitation signal; and an estimator for estimating gain / shape parameters using the high-band voicing factor.

[0016] The foregoing, as well as other objects, advantages and features of the method and device for time-domain bandwidth extension of an excitation signal when encoding / decoding a crosstalk audio signal, will become more apparent from the following non-limiting description of illustrative embodiments thereof, given by way of example only, and with reference to the accompanying drawings, in which: [Brief description of the drawings]

[0017] [Figure 1] 1 is a plot showing the power spectrum P (dB) versus frequency f (kHz) of an example crosstalk speech in which two speakers (speaker 1 and speaker 2) pronounce different types of speech (voiced and unvoiced). [Diagram 2] 1 is a schematic block diagram illustrating simultaneously a highband voicing factor calculation / calculator in a method and device for time-domain bandwidth extension of an excitation signal when encoding a crosstalk speech signal; [Diagram 3] 4 is a graph illustrating how the temporal envelope of a highband residual signal is determined. [Figure 4] 13 is a graph illustrating interpolation of segmental normalization factors calculated using average values ​​of successive intervals of the downsampled temporal envelope of the highband residual signal; [Diagram 5] 1 is a schematic block diagram showing simultaneously a decoder-side calculation / calculator of a time-domain bandwidth extended excitation signal in a method and device for time-domain bandwidth extension of an excitation signal; [Figure 6] 1 is a schematic block diagram showing simultaneously an encoder-side calculation / calculation of a high-band mixing factor formed / represented by a quantized normalized gain in a method and device for time-domain bandwidth extension of an excitation signal; [Figure 7] 1 is a schematic block diagram of a gain shape estimation / estimator in a method and device for time-domain bandwidth extension of an excitation signal; [Figure 8] 13 is a graph showing interpolated subframe gains. [Figure 9] FIG. 2 is a simplified block diagram of an exemplary configuration of hardware components forming a method and device for time-domain bandwidth extension of an excitation signal when encoding / decoding a crosstalk audio signal. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] The following description relates to a technique for encoding / decoding crosstalk audio signals. In this disclosure, the basis for this encoding / decoding technique is the SWB TBE tool of the 3GPP EVS codec described in reference [1]. However, it should be noted that this technique can be used with other encoding / decoding techniques.

[0019] More specifically, this disclosure proposes a set of modifications to the SWB TBE tool, the purpose of which is to improve the quality of the synthesized crosstalk audio signal, such as the crosstalk speech signal, in particular but not exclusively to eliminate the rattle sound defined above. The set of modifications concerns the time domain bandwidth extension of the excitation signal and is distributed in one or more of the following three areas: - Calculation of the high-band voicing factor in the coder using the time envelope of the high-band residual signal. In the SWB TBE tool, the high-band corresponds to SHB (Super Higher-Band). - Calculation of the highband mixing factor for the highband excitation signal in the encoder and decoder. - Improved estimation of gain / shape parameters and frame gain in the encoder and decoder.

[0020] The calculation of the highband voicing factor according to the present disclosure uses a highband autocorrelation function itself calculated from the time envelope of, for example, a downsampled domain of the highband residual signal. The highband voicing factor is used in the encoder to replace the so-called voice factor, which is obtained from the lowband voicing parameters in the SWB TBE tool.

[0021] The calculation of the high-band mixing factor according to the present disclosure replaces the corresponding method in the SWB TBE tool. The high-band mixing factor determines the ratio of the low-band excitation signal (e.g., from the ACELP core) and the random noise (which can also be defined as "white noise") excitation signal to generate the time-domain bandwidth-extended excitation signal. In the disclosed implementation, the high-band mixing factor is calculated using, for example, MSE (Mean Square Error) minimization between the time envelope of the random noise excitation signal in the downsampled domain and the time envelope of the low-band excitation signal. The quantization of the high-band mixing factor can be performed by the existing quantizer of the SWB TBE tool. Adding the quantized high-band mixing factor to the SWB TBE bitstream results in a slight increase in the bit rate. The mixing operation is performed in both the encoder and the decoder. Other characteristics of the mixing operation may include rescaling the random noise excitation signal at the beginning of each frame and interpolating the high-band mixing factor to ensure a smooth transition between the current frame and the previous frame.

[0022] The estimation of gain / shape parameters according to the present disclosure involves post-processing of gain / shape parameters using adaptive smoothing of the unquantized gain / shape parameters (at the encoder), with weighting between the original and the interpolated gain / shape parameters. The quantization of the gain / shape parameters can be performed by the existing quantizer in the SWB TBE tool. The adaptive smoothing is applied twice: first to the unquantized gain / shape parameters (at the encoder), and then to the quantized gain / shape parameters (at both the encoder and the decoder). At the encoder, adaptive attenuation is applied to the unquantized frame gains. The adaptive attenuation is based on the MSE excess error, which is a by-product of the SHB voicing parameter calculation in the SWB TBE tool.

[0023] FIG. 2 is a schematic block diagram illustrating simultaneously a method 200 for time-domain bandwidth extension of an excitation signal in encoding a crosstalk speech signal and a calculation / calculator of a high-band voicing factor in a device 250 .

[0024] 1. Low-band excitation signal Referring to FIG. 2, an input audio signal s for the 3GPP EVS codec is inp (n) is, for example, the following relationship (1)

number

[0025] The method 200 includes a downsampling operation 201, and the device 250 includes a downsampler 251 for performing the operation 201. The downsampler 251 downsamples an input audio signal s inp 2. The ACELP encoder 253 downsamples the input audio signal (n) from 32 kHz to 12.8 kHz or 16 kHz depending on the bit rate of the encoder. For example, the input audio signal in the 3GPP EVS codec is downsampled to 12.8 kHz for all bit rates up to 24.4 kbps, and to 16 kHz otherwise. The resulting signal is the lowband signal 202. The lowband signal 202 is encoded using an ACELP encoder 253 in an ACELP encoding operation 203.

[0026] The method 200 includes an ACELP encoding operation 203, while the device 250 comprises an ACELP encoder 253 of the 3GPP EVS codec for performing the ACELP encoding. The ACELP encoder 253 generates two types of excitation signals: an adaptive codebook excitation signal 204 and a fixed codebook excitation signal 205, as described in reference [1].

[0027] In the method 200 and device 250, the SWB TBE tool in the 3GPP EVS codec comprises a corresponding generator 257 for performing a low-band excitation signal generation operation 207 and generating a low-band excitation signal 208. The generator 257 uses the two excitation signals 204 and 205 as input, mixes them together and applies a nonlinear transformation to generate a mixed signal with spectral inversion, which is further processed in the SWB TBE tool, resulting in the low-band excitation signal 208 of Fig. 2. Details about the low-band excitation signal generation can be found in reference [1], in particular in section 5.2.6.1 for SWB TBE encoding and in section 6.1.3.1 for SWB TBE decoding.

[0028] As a non-limiting example, a low-band excitation signal 208 with spectral inversion may be sampled at 16 kHz and may be expressed as follows:

number

[0029] 2. High-bandwidth target signal Referring to FIG. 2, the highband target signal 210 is essentially the input audio signal s inp(n) containing spectral components in the frequency range of 6.4 kHz to 14 kHz or 8 kHz to 16 kHz depending on the bit rate of the codec. The highband target signal 210 is always sampled at 16 kHz regardless of the bit rate of the codec and its spectral content is inverted. Thus, the first frequency bin of the highband target spectrum corresponds to the last frequency bin of the spectrum and vice versa. In the method 200 and device 250, the highband target signal 210 can be generated using a QMF (quadrature mirror filter) analysis operation 209, for example performed by a QMF analysis filter bank 259 of the 3GPP EVS codec described in reference [1]. Alternatively, the highband target signal 210 can be generated by a QMF ... inp The highband target signal 210 may also be generated by bandpass filtering (n), shifting the filtered version in the frequency domain, inverting its spectral content as described above, and finally downsampling it from 32 kHz to 16 kHz. In this disclosure, the use of QMF processing is assumed, and the highband target signal 210 may be generated, for example, by the following relationship (3):

number

[0030] Following processing in the QMF filter bank 259, the method 200 includes an operation 211 of estimating high-band filter coefficients 212, and the device 250 includes an estimator 261 for performing operation 211. The estimator 261 estimates high-band LP (Linear Prediction) filter coefficients 212 from the high-band target signal 210 in four consecutive subframes per frame, each subframe having a length of 80 samples. The estimator 261 calculates the high-band LP filter coefficients 212 using the Levinson-Durbin algorithm as described in Reference [1]. The high-band LP filter coefficients 212 are calculated according to the following relationship (4):

number

number

[0031] The method 200 includes an operation 213 of generating a highband residual signal 214, and the device 250 includes a generator 263 of a highband residual signal for performing the operation 213. The generator 263 generates the highband residual signal 214 by filtering the highband target signal 210 from the QMF analysis filter bank 259 with the highband LP filter (LP filter coefficients 212) from the estimator 261. The highband residual signal 214 may be, for example, expressed as a function of the following relationship (5):

number

[0032] The first P samples of the high-band residual signal 214 are calculated using the high-band target signal 210 from the previous frame. This is done by using the s HB (-k), denoted by negative indexes in k=1,...,P. Negative indexes refer to samples of the highband target signal 210 at the end of the previous frame.

[0033] 3. Highband autocorrelation function and voicing factor Section 3 (highband autocorrelation function) concerns the characteristics of the coder.

[0034] The highband residual signal 214 calculated by generator 263 using relation 5 is used to calculate the highband autocorrelation function and the highband voicing factor. The highband autocorrelation function is not calculated directly for the highband residual signal 214. To calculate the highband autocorrelation function directly would require significant computational resources. Furthermore, the dynamics of the highband residual signal 214 are typically small, and the spectral inversion process often smears the difference between voiced and unvoiced speech signals. To avoid these problems, the highband autocorrelation function is estimated for the time envelope of, for example, a downsampled region of the highband residual signal 214.

[0035] The method 200 includes an operation 215 of calculating a temporal envelope of the highband residual signal 214, and the device 250 includes a calculator 265 for performing the operation 215. TD To calculate (n) 216, calculator 265 processes highband residual signal 214 through a sliding moving average (MA) filter with M=20 taps in an exemplary implementation. The time envelope calculation may be performed, for example, using the following relationship (6):

number

number

[0036] The time envelope R of the highband residual signal 214 TD The operation 215 of computing (n) 216 is illustrated in FIG.

[0037] The method 200 includes a time envelope downsampling operation 217, and the device 250 includes a downsampler 267 for performing the operation 217. The downsampler 267 may be implemented, for example, by using the following relationship (8):

number

[0038] The method 200 includes an average value calculation operation 219, and the device 250 includes a calculator 269 for performing the operation 219. The calculator 269 calculates the downsampled temporal envelope R 4kHz (n)218 is divided into four consecutive intervals, and the downsampled time envelope R 4kHz (n) 218, the average value 220 in each interval, for example, the following relationship (9)

number

[0039] Calculator 269 limits all average values ​​to a maximum value of 1.0.

[0040] The method 200 includes a normalization factor calculation operation 221, and the device 250 includes a calculator 271 for performing the operation 221. The calculator 271 uses the average value 220 of the downsampled time envelope to calculate, for each interval k, an interval normalization factor, for example, according to the following relationship (10):

number

[0041] Calculator 271 then linearly interpolates the interval normalization factor from relationship (10) within the entire interval of the current frame to obtain an interpolated normalization factor 222, for example, according to the following relationship (11):

number

[0042] This interpolation process, performed by operation 221 and calculator 271, is illustrated in FIG.

[0043] In relation (11), the term η -1 refers to the last interval normalization factor in the previous frame. Thus, the term η -1 is updated with η3 after the interpolation process in each frame.

[0044] The method 200 includes a normalization operation 223 of the downsampled temporal envelope, and the device 250 comprises a normalizer 273 for performing the operation 223. The normalizer 273 normalizes the downsampled temporal envelope R 4kHz (n) 218 ​​using the interpolated normalization factor γ(n) 222, for example, the following relationship (12):

number

[0045] Then, the normalizer 273 calculates the value R of the relation (12). γ From (n), the global mean value of the normalized envelope

number

number

[0046] It is useful to estimate the slope of the time envelope of the highband residual signal. To that end, the method 200 includes a time envelope slope estimation operation 225, and the device 250 comprises an estimator 275 for performing the operation 225. The time envelope slope estimation is performed using the Linear Least Squares (LLS) method, taking the interval mean value calculated in relation (9) as

number

number

[0047] According to the LLS method, the goal is to find, for all k = 0, ..., 3,

number

number

[0048] Optimal slope a LLS (slope 226) is calculated by the estimator 275 according to the relationship (16)

number

[0049] The method 200 includes a highband autocorrelation function calculation operation 227, and the device 250 includes a calculator 277 for performing the operation 227. The calculator 277 calculates the highband autocorrelation function X corr 228 based on the normalized time envelope, e.g., relation (17)

number

number

number

[0050] In the case of mode switching, the factor before the summation term in relation (17) is 1 / E f because the normalized temporal envelope R norm (n)The energy of 224 is unknown.

[0051] The method 200 includes a highband voicing factor calculation operation 229 , and the device 250 comprises a calculator 279 for performing the operation 229 .

[0052] The voicing of the highband residual signal is performed using the highband autocorrelation function X corr Variance σ of 228 corr The calculator 279 has a variance σ corr For example, the following relation (19)

number

[0053] Voicing parameter ν mult In order to improve the discriminability (voiced / unvoiced decision) of

number

[0054] Then, the calculator 279 calculates, for example, the following relationship (21):

number

[0055] 4. Excitation Mixing Factor FIG. 5 is a schematic block diagram illustrating simultaneously the decoder-side time-domain bandwidth extended excitation signal calculation / calculator in the method 200 and device 250. As shown in FIG.

[0056] Section 4 (Excitation Mixing Factors) concerns the characteristics of both the encoder and the decoder.

[0057] The SWB TBE tool in the 3GPP EVS codec uses the low-band excitation signal 208 (FIG. 2) described in section 1 (low-band excitation signal) to predict the high-band residual signal 214 (FIG. 2) described in section 2 (high-band target signal). At low bit rates below 24.4 kbps in the EVS codec, the SWB TBE tool uses 19 bits to code the spectral envelope and energy of the predicted high-band residual signal. At a frame length of 20 ms, this results in a bit rate of 0.95 kbps. At bit rates higher than 24.4 kbps, the SWB TBE tool uses 32 bits to code the spectral envelope and energy of the predicted high-band residual signal. At a frame length of 20 ms, this results in a bit rate of 1.6 kbps. At both bit rates (0.95 and 1.6 kbps) of the SWB TBE tool, no bits are used to code the highband residual signal 214 or the highband target signal 210 .

[0058] Referring to FIG. 5, the method 200 includes a pseudorandom noise generation operation 501, and the device 250 comprises a pseudorandom noise generator 551 for performing the operation 501.

[0059] The pseudorandom noise generator 551 generates a uniformly distributed random noise excitation signal 502. For example, the pseudorandom number generator of the 3GPP EVS codec described in reference [1] can be used as the pseudorandom noise generator 551. The random noise excitation signal w rand 502 is the next relation (22)

number

[0060] Random noise excitation signal w rand 502 is a function with zero mean and non-zero variance σ rand = 1.14e + 11. Note that the variance is only an estimate and represents the average value over 100 frames.

[0061] The method 200 includes providing a lowband excitation signal l LB (n) 208 , and a power calculator 553 for performing the operation 503 .

[0062] The power calculator 503 calculates the power of the low-band excitation signal l transmitted from the encoder. LB (n) The power 504 of 208, for example, in the following relationship (23)

number

[0063] The method 200 includes an operation 505 of normalizing the power of the random noise excitation signal 502 , and a power normalizer 555 for performing the operation 505 .

[0064] The power normalizer 555 may, for example, use the following relationship (24):

number

[0065] Although the true variance of the random noise excitation signal 502 varies from frame to frame, the exact value is not needed for power normalization. Instead, to save computational resources, an approximation of the variance defined above is used in the above relation (24).

[0066] The method 200 includes providing a lowband excitation signal l LB (n) 208 is the power-normalized random noise excitation signal w white (n) 506 and an operation 507 of mixing, and a mixer 557 for performing operation 507.

[0067] Mixer 557 receives the low-band excitation signal l LB (n) 208 is the power-normalized random noise excitation signal w white(n) 506 using high-band mixing factors, as described later in this disclosure, to generate a time-domain bandwidth extended excitation signal 508.

[0068] FIG. 6 is a schematic block diagram showing simultaneously an encoder-side calculation / calculation of a high-band mixing factor formed / represented by a quantized normalized gain in a method and device for time-domain bandwidth extension of an excitation signal.

[0069] Referring to FIG. 6, in the encoder: The method 200 comprises: white (n) 506, an operation 602 of calculating the time envelope of the low-band excitation signal l LB (n) 208, a mean squared error (MSE) minimization operation 601, and a gain quantization operation 607; The device 250 comprises a time envelope calculator 652 for performing operation 602, a time envelope calculator 654 for performing operation 604, an MSE minimizer 651 for performing operation 601, and a gain quantizer 657 for performing operation 607.

[0070] As shown in Figure 6, to save computational resources, the optimal gain is calculated based on the time envelope of the signal in the downsampled domain using a mean square error (MSE) minimization process.

number

[0071] Calculator 652 calculates the power-normalized random noise excitation signal w (also calculated at the encoder side as shown in FIG. 5 and the corresponding description). white (n) Downsampled time envelope W of 506 4kHz(n) 606 is calculated using the same algorithms described in section 3 (Highband Autocorrelation Function and Voicing Factor) for calculating the time envelope of the highband residual signal 214 (operation 215 and calculator 265 in FIG. 2) and for downsampling that time envelope (operation 217 and downsampler 267 in FIG. 2). The downsampling factor used is, for example, 4. The downsampled time envelope of the power-normalized random noise excitation signal is calculated using the following relationship (25):

number

[0072] Similarly, the calculator 654 calculates the low-band excitation signal l LB (n)208, 4kHz downsampled time envelope L 4kHz (n) 605 is calculated again using the same algorithm as described in Section 3 (Highband Autocorrelation Function and Voicing Factor). LB The downsampled temporal envelope 605 of (n) 208 can be expressed as:

number

[0073] The objective of the MSE minimization operation 601 is to (a) minimize the combined temporal envelope (L 4kHz (n), W 4kHz (n)) and (b) the high-band residual signal r HB (n)214 Time envelope R 4kHz (n) The optimal gain pair that minimizes the energy of the error between

number

number

[0074] To that end, the MSE minimizer 651 solves a system of linear equations. Methods for solving this can be found in the scientific literature. For example,

number

number

number

[0075] Then, the MSE minimizer 651 calculates the minimum MSE error energy (excess error) as, for example, the following relation (30):

number

[0076] For further processing, a gain quantizer 657 determines the optimal gain

number

number

[0077] The consequence / benefit of the rescaling of relation (31) is that instead of two parameters, we have a normalized gain gwn The idea is that only one parameter, L, needs to be coded and transmitted in the bitstream from the encoder to the decoder. Scaling the gains using relation (31) therefore reduces bit consumption and simplifies the quantization process. On the other hand, the combined temporal envelope (L 4kHz (n) and W 4kHz The energy of (n) is the time envelope R 4kHz (n) does not match the energy of (n). This is not a problem because the SWB TBE tool uses subframe and global gains, which contain information about the energy of the high-band residual signal. The computation of subframe and global gains is described in Section 6 (Gain / Shape Estimation) of this disclosure.

[0078] The gain quantizer 657 calculates the normalized gain g wn between a maximum threshold of 1.0 and a minimum threshold of 0.0. Gain quantizer 657 limits the normalized gain g wn For example, the following relation (32)

number

[0079] Referring again to FIG. 5, the method 200 includes a mixed factor decoding operation 509 at the decoder, and the device 250 comprises a mixed factor decoder 559 for performing the operation 509 .

[0080] The mixed factor decoder 559 receives the index idx g 610, the decoded gain can be calculated, for example, by the following relationship (33):

number

[0081] The decoded gain from relation (33) is the highband mixing factor f mix Form 510.

[0082] For example, a low-band excitation signal sampled at 16 kHz LB (n) 208 and a normalized random noise excitation signal w sampled at, for example, 16 kHz. white (n) 506 are mixed together in mixer 557. However, the low-band excitation signal l LB (n)208 energy and random noise excitation signal w rand The energies of both 502 vary from frame to frame. The high-band mixing factor f mix 510 to generate the low-band excitation signal LB (n)208 and random noise excitation signal w rand 502 directly, the energy fluctuations may eventually result in audible artifacts at the frame boundaries. To ensure a smooth transition, the random noise excitation signal w rand The energy of 502 is linearly interpolated between the previous and current frames in generator 551. This is done by applying the random noise excitation signal w rand 502 for the first half of the current frame,

number

number

[0083] To further smooth the transition between the previous and current frames, the decoder 559 also uses a highband blending factor f mix 510 is linearly interpolated. This can be done, for example, by using the following relationship:

number

number

[0084] Low-band excitation signal LB (n)208 and random noise excitation signal w white (n) 506 is finally mixed by mixer 557 using, for example, relation (36), to obtain the time-domain bandwidth extended excitation signal u(n) 508.

number

[0085] 5. High-frequency synthesis (LP synthesis) In the relation (4), the high-band input signal s HB (n) High-band LP filter coefficients a calculated using LP analysis j HB The (n) 212 are converted to LSF parameters and quantized in the SWB TBE tool's encoder. At a bit rate of 0.95 kbps, the SWB TBE encoder quantizes the LSF index using 8 bits. At a bit rate of 1.6 kbps, the SWB TBE encoder quantizes the LSF index using 21 bits.

[0086] Referring again to FIG. 5, in the decoder: To decode the quantized LSF index, the method 200 comprises a decoding operation 511 and the device 250 comprises a corresponding decoder 561 . To convert the decoded LSF indices 512 into high-band LP filter coefficients 514 , the method 200 includes a conversion operation 513 , and the device 250 comprises a corresponding converter 563 .

[0087] The decoded highband LP filter coefficients 514 are

number

number

[0088] The method 200 includes a filtering operation 515, in which the device 250 includes a corresponding synthesis filter 565, and the decoded high-band LP filter coefficients 514 are used to filter the mixed time-domain bandwidth extended excitation signal 508 of the relationship (36), for example, using the following relationship (38), to obtain the LP filtered high-band signal y HB The result is 516.

number

[0089] 6. Gain / shape estimation (Figure 7) Gain / shape parameter smoothing is applied at both the encoder and decoder. Adaptive attenuation of frame gain is applied only at the encoder.

[0090] High-bandwidth target signals HBThe spectral shape of the signal (n) 210 is encoded using the quantized LSF coefficients. Referring to FIG. 7, the SWB TBE tool is used to encode the high-band target signal s as described in [1]. HB It also comprises an estimation operation 701 / estimator 751 for estimating the temporal subframe gains 702 of (n) 210. The estimator 751 normalizes the estimated temporal subframe gains to unit energy.

[0091] The normalized estimated temporal subframe gain 702 from the estimator 751 is given by the relationship (39):

number

[0092] Normalized estimated temporal subframe gain g k To determine the temporal tilt 704 of 702 using linear least squares (LLS) interpolation, the method 200 includes a calculation operation 703, and the device 250 includes a corresponding calculator 753. As shown in Figure 8, this interpolation process can be performed by fitting a linear curve 801 to the true subframe gains 702 in four consecutive subframes (subframes 0-3 in Figure 8) and calculating its slope.

[0093] The linear curve 801 constructed using the LLS interpolation method is expressed as follows:

number

number

[0094] By extending the relationship (41), the estimated temporal subframe gain g k 702 time gradient g tilt It is possible to express the time gradient g tilt 704 is in fact the optimal slope of the linear curve c LLS The time gradient g is equal to tilt is calculated in the calculator 753 according to the following relationship (42)

number

[0095] For example, if the following condition is true:

number

number

[0096] In that case, the temporal subframe gain g k The smoothing of 702 is performed by a smoother 755, for example, according to the following relationship (44):

number

number

[0097] Smoothed Temporal Subframe Gain

number

number

number

[0098] After the quantization operation 707, the quantized temporal subframe gains

number

number

[0099] Interpolated quantized temporal subframe gain

number

number

number

[0100] Then, the quantized temporal subframe gain

number

number

[0101] For that purpose, the quantized temporal subframe gains

number

number

number

[0102] The method 200 includes a frame gain estimation operation 715, and the device 250 includes a corresponding frame gain estimator 765. The SWB TBE tool uses the frame gain to control the global energy of the synthesized highband speech signal. The frame gain is calculated by: (a) the LP filtered highband signal y HB At 516, the smoothed quantized temporal subframe gains from the relation (49) are

number

number

number

[0103] The details of the frame gain estimation operation 715 are given in reference [1]. The estimated frame gain parameters are f (See 716).

[0104] The method 200 includes an operation 717 of calculating a composite highband signal 718, and the device 250 includes a calculator 767 for performing the operation 717. The calculator 767 calculates the estimated frame gain g f 716 can be modified under some specific conditions. For example, the frame gain g f Relationship (51)

number

number

[0105] Under some specific conditions, the frame gain g f Further modifications to are given in reference [1].

[0106] Calculator 767 then quantizes the modified frame gain using the frame gain quantizer of the SWB TBE tools encoder of reference [1].

[0107] Finally, the calculator 767 calculates the synthesized highband speech signal 718, e.g., according to the following relationship (53):

number

[0108] 7. Exemplary Configurations of Hardware Components FIG. 9 is a simplified block diagram of an exemplary configuration of hardware components forming the above-described method 200 and device 250 (hereinafter “method 200 and device 250”) for time-domain bandwidth extension of an excitation signal when encoding / decoding a crosstalk audio signal.

[0109] The method 200 and device 250 can be implemented as part of a mobile terminal, as part of a portable media player, or in any similar device. The device 250 (identified in FIG. 9 as 900) includes an input 902, an output 904, a processor 906, and a memory 908.

[0110] The input unit 902 is configured to receive an input signal. The output unit 904 is configured to provide a time-domain bandwidth-extended excitation signal. The input unit 902 and the output unit 904 can be implemented in a common module, for example a serial input / output device.

[0111] The processor 906 is operatively connected to the input 902, the output 904, and the memory 908. The processor 906 is implemented as one or more processors for executing code instructions that facilitate the functions of the various operations and elements of the method 200 and device 250 described above, as illustrated in the accompanying figures and / or described in this disclosure.

[0112] The memory 908 may comprise a non-transitory memory for storing code instructions executable by the processor 906, and in particular may comprise a processor-readable memory that comprises / stores non-transitory instructions that, when executed, cause the processor to perform the operations and elements of the method 200 and device 250. The memory 908 may also comprise a random access memory or buffer for storing intermediate processing data from various functions performed by the processor 906.

[0113] Those skilled in the art will recognize that the description of the method 200 and device 250 is merely illustrative and not limiting in any way. Other embodiments will readily occur to those of skill in the art having the benefit of this disclosure. Moreover, the disclosed method 200 and device 250 can be customized to provide valuable solutions to current needs and problems in encoding and decoding audio.

[0114] For clarity, not all routine features of implementations of method 200 and device 250 have been shown and described. Of course, it will be understood that in the development of any such actual implementation of method 200 and device 250, numerous implementation-specific decisions may need to be made to achieve the developer's particular goals, such as compliance with application-, system-, network-, and business-related constraints, and that these particular goals will vary from implementation to implementation and developer to developer. Moreover, it will be understood that the development effort may be complex and time-consuming, but will nevertheless be a routine engineering undertaking for those skilled in the art of speech processing having the benefit of this disclosure.

[0115] According to the present disclosure, the elements, processing operations, and / or data structures described herein can be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. In addition, those skilled in the art will recognize that less general-purpose devices, such as hardware devices, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc., can also be used. When a method including a series of operations and sub-operations is performed by a processor, computer, or machine, and the operations and sub-operations can be stored as a series of non-transitory code instructions readable by a processor, computer, or machine, the operations and sub-operations can be stored on a tangible and / or non-transitory medium.

[0116] The processing operations and elements of the method 200 and device 250 described herein may include software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.

[0117] In the method 200 and device 250, various process operations and sub-operations may be performed in various orders, and some of the process operations and sub-operations may be optional.

[0118] Although the present disclosure has been described above by its non-limiting and exemplary embodiments, these embodiments can be freely modified within the scope of the appended claims without departing from the spirit and essence of the present disclosure.

[0119] 8.References This disclosure refers to the following references, the contents of which are incorporated herein by reference in their entireties: [1] 3GPP TS 26.445, "EVS Codec Detailed Algorithmic Description," 3GPP Technical Specification (Release 12) (2014) - Sections 5.2.6.1 and 6.1.3.1 [2] Bessette, B., Lefebvre, R., Salami, R. et al. "Techniques for high-quality ACELP coding of wideband speech". Int. Conference EUROSPEECH 2001 Scandinavia, 7th European Conference on Speech Communication and Technology, 2nd INTERSPEECH Event, Aalborg, Denmark, September 3-7, 2001 [Explanation of symbols]

[0120] 200 ways 202 Low-bandwidth signal 204 Adaptive Codebook Excitation Signal 205 Fixed codebook excitation signal 208 Low-band excitation signal LB (n) 210 High Bandwidth Target SignalsHB (n) 212 High-band LP (linear prediction) filter coefficient a j HB (n) 214 High-band residual signal, High-band residual signal r HB (n) 216 Time envelope R TD (n) 218 Downsampled Time Envelope R 4kHz (n) 220 Mean value of downsampled time envelope 222 Interpolated normalization factor γ(n) 224 Normalized Time Envelope R norm (n) 226 Incline 228 High-band autocorrelation function X corr 230 Highband Voicing Factor ν HB , the voicing parameter ν HB 250 devices 251 Down Sampler 253 ACELP encoder 257 Generator 259 QMF (Quadrature Mirror Filter) Analysis Filter Bank 261 Estimator 263 Generator 265 Calculator 267 Down Sampler 269 ​​Calculator 271 Calculator 273 Normalizer 275 Estimator 277 Calculator 279 Calculator 502 Random noise excitation signal w rand 504 Power 506 Power-normalized random noise excitation signal w white (n) 508 Time-domain bandwidth-extended excitation signal u(n) 510 Highband mixing factor f mix 512 Decoded LSF Index 514 Decoded High-Band LP Filter Coefficients 516 LP filtered high-band signal y HB 551 Pseudorandom Noise Generator 553 Power Calculator 555 Power Normalizer 557 Mixer 559 Mixed Factor Decoder 561 Decoder 563 Converter 565 Synthesis Filter 605 Downsampled Time Envelope L 4kHz (n) 606 Downsampled Time Envelope W 4kHz (n) 610 index idx g 651 MSE Minimizer 652 Time envelope calculator 654 Time Envelope Calculator 657 Gain quantizer 702 Normalized estimated temporal subframe gain g k 704 temporal slope g tilt 706 Smoothed Temporal Subframe Gain 708 Quantized Temporal Subframe Gain 710 Interpolated Quantized Temporal Subframe Gain 714 Smoothed Quantized Temporal Subframe Gain 716 Estimated frame gain g f 718 Synthetic Highband Signal, Synthesized Highband Speech Signal 751 Estimator 753 Calculator 755 Smoother 757 Gain shape quantizer 759 Interpolator 761 Slope Calculator 763 Smoother 765 Frame Gain Estimator 767 Calculator 801 Linear curve 902 Input section 904 Output section 906 Processor 908 Memory

Claims

1. 1. A method for time domain bandwidth extension of an excitation signal when decoding a crosstalk audio signal, comprising: decoding a highband mixing factor received in the bitstream; mixing the low-band excitation signal with the random noise excitation signal using the high-band mixing factor to generate a time-domain bandwidth-extended excitation signal; A method comprising:

2. 2. The method of claim 1, wherein decoding the high-band mixing factor comprises decoding a quantized normalized gain received in the bitstream and calculating the high-band mixing factor using the decoded quantized normalized gain.

3. 2. The method of claim 1, further comprising interpolating the energy of the random noise excitation signal between a previous frame and a current frame of the speech signal to smooth a transition between the previous frame and the current frame.

4. 4. The method of claim 3, comprising scaling the random noise excitation signal within a portion of the current frame to interpolate the energy of the random noise excitation signal.

5. 2. The method of claim 1, further comprising interpolating the highband mixing factors between a previous frame and a current frame of the speech signal to ensure a smooth transition between the previous frame and the current frame.

6. The method of claim 1 , comprising estimating quantized gain / shape parameters.

7. 1. A method for time domain bandwidth extension of an excitation signal when encoding a crosstalk audio signal, comprising: Calculating a high-band mixing factor that can be used to mix the low-band excitation signal with the random noise excitation signal to generate a time-domain bandwidth-extended excitation signal. A method comprising:

8. A method for generating a high-band residual signal comprising the steps of: (a) calculating a high-band residual signal using the speech signal; and (b) calculating a time envelope of the high-band residual signal. calculating a highband voicing factor based on the temporal envelope of the highband residual signal; estimating gain / shape parameters using said highband voicing factors; 8. The method of claim 7, comprising:

9. 9. The method of claim 8, wherein the step of calculating the highband voicing factor comprises: (a) calculating a highband autocorrelation function based on the time envelope; and (b) calculating the highband voicing factor using the highband autocorrelation function.

10. Calculating the highband voicing factor comprises: downsampling the temporal envelope of the highband residual signal by a given factor; dividing the downsampled temporal envelope into a plurality of intervals and calculating a mean value for each interval of the downsampled temporal envelope; performing piecewise normalization of the downsampled temporal envelope of the highband residual signal; Including, 9. The method of claim 8, wherein the interval-wise normalization of the downsampled temporal envelope comprises: (a) calculating interval normalization factors from the calculated average values; (b) interpolating the interval normalization factors within a current frame; and (c) normalizing the downsampled temporal envelope using the interpolated interval normalization factors.

11. The method of claim 8 , comprising calculating the slope of the temporal envelope of the highband residual signal based on a linear least squares method.

12. The method of claim 7 , wherein calculating the high-band mix factor comprises calculating and quantizing a gain from which the high-band mix factor is derived.

13. Calculating the highband blending factor comprises: generating the random noise excitation signal; mixing the low-band excitation signal with the random noise excitation signal and minimizing the mean square error between the mixed excitation signal and a high-band residual signal calculated from the speech signal; calculating a time envelope of the random noise excitation signal, calculating a time envelope of the low-band excitation signal, and finding a gain for each of the time envelope of the random noise excitation signal and the time envelope of the low-band excitation signal using a mean square error minimization process; 13. The method of claim 12, comprising:

14. wherein calculating the high-band mix factor comprises scaling the respective gains of the time envelope of the random noise excitation signal and the time envelope of the low-band excitation signal; 14. The method of claim 13, wherein scaling the respective gains comprises obtaining a single gain parameter, and wherein calculating the highband mix factor comprises quantizing the single gain parameter to obtain a quantized gain from which the highband mix factor is derived.

15. The gain / shape parameters are: - the spectral shape of the high-band target signal, - a subframe gain of the highband target signal; - Frame Gain parameter 9. The method of claim 8, selected from the group consisting of:

16. estimating the gain / shape parameters includes calculating the temporal slope of the gain / shape parameters; The method of claim 8 , wherein calculating the temporal tilt comprises interpolating the gain / shape parameters.

17. the step of estimating the gain / shape parameters includes smoothing the gain / shape parameters using adaptive weighting parameters calculated using the highband voicing factors; 9. The method of claim 8, wherein the method includes smoothing the gain / shape parameter using the adaptive weight parameter in response to a given condition involving the highband voicing factor.

18. estimating the gain / shape parameters quantizing the smoothed gain / shape parameters; Interpolating and smoothing the quantized gain / shape parameters. Including, 18. The method of claim 17, wherein smoothing the quantized gain / shape parameters is performed using averaging of interpolated quantized gain / shape parameters.

19. The method of claim 8 , wherein the step of estimating the gain / shape parameters comprises adaptively attenuating frame gain parameters using an MSE excess error.

20. 1. A device for time domain bandwidth extension of an excitation signal when decoding a crosstalk audio signal, comprising: a decoder for the highband mixing factors received in the bitstream; a mixer of the low-band excitation signal and the random noise excitation signal using said high-band mixing factor to generate a time-domain bandwidth-extended excitation signal; A device comprising:

21. 21. The device of claim 20, wherein the decoder of the highband mixing factor decodes quantized normalized gains received in the bitstream and calculates the highband mixing factor using the decoded quantized normalized gains.

22. 21. The device of claim 20, comprising: a generator of the random noise excitation signal, the generator interpolating energy of the random noise excitation signal between a previous frame and a current frame of the speech signal to smooth a transition between the previous frame and the current frame.

23. 23. The device of claim 22, wherein the generator of the random noise excitation signal scales the random noise excitation signal within a portion of the current frame to interpolate the energy of the random noise excitation signal.

24. 21. The device of claim 20, wherein the decoder of the highband mixing factor interpolates the highband mixing factor between the previous frame and the current frame of the audio signal to ensure a smooth transition between the previous frame and the current frame.

25. 21. The device of claim 20, comprising a quantized gain / shape parameter estimator.

26. 1. A device for time domain bandwidth extension of an excitation signal when encoding a crosstalk audio signal, comprising: A calculator of a high-band mixing factor that can be used to mix a low-band excitation signal with a random noise excitation signal to generate a time-domain bandwidth-extended excitation signal. A device comprising:

27. A calculator of (a) a highband residual signal using said audio signal, and (b) a time envelope of said highband residual signal. a highband voicing factor calculator based on the temporal envelope of the highband residual signal; a gain / shape parameter estimator using the high-band voicing factors; 27. The device of claim 26, comprising:

28. 28. The device of claim 27, wherein the calculator of the highband voicing factor calculates a highband autocorrelation function based on the temporal envelope and calculates the highband voicing factor using the highband autocorrelation function.

29. the calculator of the highband voicing factor comprises: a downsampler of the temporal envelope of the highband residual signal by a given factor; a divider of the downsampled time envelope into multiple intervals; a calculator of the mean value of each interval of the downsampled time envelope; a piecewise normalizer of the downsampled temporal envelope of the highband residual signal; Equipped with 28. The device of claim 27, wherein the interval normalizer (a) calculates interval normalization factors from the calculated average values, (b) interpolates the interval normalization factors within a current frame, and (c) normalizes the downsampled temporal envelope using the interpolated interval normalization factors.

30. 27. The device of claim 26, wherein the calculator of the highband mix factor calculates and quantizes gains that form the highband mix factor.

31. The calculator of the highband blend factor comprises: a generator of the random noise excitation signal; mixing the low-band excitation signal with the random noise excitation signal and minimizing the mean square error between the mixed excitation signal and a high-band residual signal calculated from the speech signal; 31. The device of claim 30, comprising a calculator of a time envelope of the random noise excitation signal and a calculator of a time envelope of the low-band excitation signal, wherein a mean square error minimization process is used to find a gain for each of the time envelope of the random noise excitation signal and the time envelope of the low-band excitation signal.

32. the high-band mix factor calculator scales the respective gains of the time envelope of the random noise excitation signal and the time envelope of the low-band excitation signal; 32. The device of claim 31 , wherein the high-band mixing factor calculator calculates a single gain parameter and quantizes the single gain parameter to obtain a quantized gain that forms the high-band mixing factor, to scale the respective gains of the temporal envelope of the random noise excitation signal and the temporal envelope of the low-band excitation signal.

33. The gain / shape parameters are: - the spectral shape of the high-band target signal, - a subframe gain of the highband target signal; - Frame Gain parameter 28. The device of claim 27, selected from the group consisting of:

34. 34. The device of claim 33, wherein the gain / shape parameters include subframe gains of the highband target signal, the estimator of the gain / shape parameters comprises a calculator of temporal slopes of the subframe gains, and the calculator of temporal slopes comprises an interpolator of the subframe gains.

35. the gain / shape parameters include subframe gains of the highband target signal, and the estimator of the gain / shape parameters comprises a smoother of the subframe gains using adaptive weighting parameters; 34. The device of claim 33, wherein the smoother of the subframe gains calculates the adaptive weighting parameters using the highband voicing factors and smooths the gain / shape parameters using the adaptive weighting parameters in response to a given condition involving the highband voicing factors.

36. the estimator of the gain / shape parameters is a quantizer for the subframe gain; an interpolator of the quantized subframe gains; a subframe gain smoother; Equipped with 36. The device of claim 35, wherein the subframe gain smoother smooths the quantized gain / shape parameters using averaging of interpolated quantized gain / shape parameters.

37. 34. The device of claim 33, wherein the gain / shape parameters include frame gain parameters of the high-band target signal, and the estimator of the gain / shape parameters performs adaptive attenuation of the frame gain parameters using an MSE excess error.