Parametric audio coding per integral band

The joint encoder and decoder system addresses high computational complexity and inefficient bit rate distribution in parametric audio coding by using a band-wise parametric coder within a waveform-preserving coder, achieving efficient coding of subbands quantized to zero with high spectral resolution and adaptive bit distribution.

JP2026042037APending Publication Date: 2026-03-10FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing parametric audio coding methods suffer from high computational complexity, low spectral resolution, and inefficient bit rate distribution, particularly when dealing with tonal and non-tonal components, and lack adaptive spectral envelope propagation.

Method used

A joint encoder and decoder system that employs a band-wise parametric coder within a waveform-preserving coder, allowing efficient coding of subbands quantized to zero, with adaptive bit distribution and high spectral resolution, using a quantizer and per-band parametric coder to generate parametric representations of subbands.

Benefits of technology

Achieves efficient parametric coding with high spectral resolution and adaptive bit rate distribution, reducing computational complexity and improving signal quality by selectively coding subbands quantized to zero.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026042037000001_ABST
    Figure 2026042037000001_ABST
Patent Text Reader

Abstract

An encoder, decoder, method and program for a spectral representation of an audio signal divided into multiple subbands are provided. Spectral representation X MR is composed of frequency bins or frequency coefficients, and at least one subband includes two or more frequency bins, and the encoder 1000 generates a spectral representation X of the audio signal divided into a plurality of subbands. MR quantized representation of X Q and a quantizer that generates the quantized representation X Q Spectral expression according to X MR , and the coded parametric representation zfl is a spectral representation X MR and a spectral representation X in at least two different subbands, MR and a per-band parametric coder, where there are parameters describing
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Embodiments of the present invention refer to an encoder and a decoder. Further embodiments refer to methods for encoding and decoding, and corresponding computer programs. Generally, embodiments of the present invention are in the field of parametric coders per integration band. [Background technology]

[0002] Modern low-bit-rate audio and speech coders typically use some kind of parametric coding for at least part of their spectral bandwidth, which is either separate from a waveform-preserving coder (in which case it is called a core coder with bandwidth extension) or very simple (e.g., noise filling). In the prior art, several approaches are already known in the field of parametric coders. In [1], comfort noise of amplitude derived from the transmitted noise fill-in level is inserted into the subvectors rounded to zero. In [2], the noise level calculation and noise substitution detection in the encoder are Detecting and marking spectral bands that can be reproduced perceptually equivalently at the decoder by noise substitution. For example, measures of tonality or spectral flatness may be checked for this purpose. Calculate and quantize the average quantization error (which may be calculated over multiple scale factor bands or over all scale factor bands that are not quantized to zero), and Calculate scale factors for bands quantized to zero so that the noise introduced (by the decoder) matches the original energy. In [2], noise is introduced into the spectral lines quantized to zero starting from the "noise filling onset line", the amplitude of the introduced noise depends on the average quantization error, and the introduced noise is scaled by a scale factor for each band.

[0003] In [3], noise filling in frequency domain coders is proposed, where zero quantization lines are replaced with random noise shaped according to the tonality and the position of non-zero quantization lines, and the level of the inserted noise is set based on the global noise level. In [4], noise-like components are detected in the encoder based on the coder frequency band. Spectral coefficients in the scale factor band containing the noise-like components are omitted from quantization / coding, and only the noise substitution flag and the total power of the substitution band are transmitted. At the decoder, a random vector with the desired total power is inserted for the substitution spectral coefficient. [5] proposes a bandwidth expansion method operating in the time domain that avoids harmonicity. The harmonicity of the decoded signal is ensured by calculating the autocorrelation function of the amplitude spectrum, which is obtained from the decoded time-domain signal. By using the autocorrelation, the estimation of F0 is avoided. The LF part of the analytical signal is generated by a Hilbert transform and amplified by a modulator, resulting in bandwidth expansion. Envelope shaping and noise addition are performed by SBR.

[0004] In [6], the complete core band is copied to the HF domain, then shifted so that the highest harmonic of the core coincides with the lowest harmonic of the replicated spectrum. Finally, the spectral envelope is restored. The frequency shift, also called the modulation frequency, is calculated based on f0, which can be calculated at the encoder side using the full spectrum, or at the decoder side using only the core band. The proposal also utilizes a steep bandpass filter in the MDCT to separate the LF and HF bands. In [7-14], a semi-parametric coding technique called Intelligent Gap Filling (IGF) was proposed. This technique fills spectral holes in the high-frequency region using a synthetic HF generated from low-frequency content and post-processing with parametric side information consisting of the HF spectral and temporal envelopes. The IGF range is determined by user-defined IGF start and stop frequencies. Waveforms deemed to need to be coded in a waveform-preserving manner by the core coder, such as prominent tones, may be located above the IGF start frequency. The encoder codes the spectral envelope within the IGF range and then quantizes the MDCT spectrum. The decoder uses conventional noise filling below the IGF start frequency. A tabular, user-defined division of the spectral bandwidth is used with possible signal-adaptive selection of source partitions (tiles) and post-processing of tiles (e.g., crossfading) to reduce issues related to tones at tile boundaries.

[0005] In

[11] , an automatic selection of source-target tile mapping and whitening level in IGF is proposed based on a psychoacoustic model. In

[15] , an encoder finds extreme coefficients in a spectrum, modifies the extreme coefficients or their neighboring coefficients, and generates side information, resulting in pseudo coefficients represented by the modified spectrum and side information. The pseudo coefficients are determined in the decoded spectrum and set to predefined values ​​in the spectrum to obtain a modified spectrum. A time-domain signal is generated by an oscillator controlled by the spectral position and the value of the pseudo coefficients. The generated time-domain signal is mixed with a time-domain signal obtained from the modified spectrum. In

[16] , pseudo coefficients are determined in the decoded spectrum and replaced by a fixed tone pattern or a frequency sweep pattern. In

[17]

[18] , the quantizer uses a deadband adapted according to the input signal characteristics. The deadband ensures that low-level, potentially noisy, spectral coefficients are quantized to zero.

[0006] Below, the shortcomings of the prior art are explained, and the analysis of the prior art and identification of the shortcomings are part of the present invention. In the prior art, simple noise filling is simply integrated into the core coder [1][2][3][4], the core coder is a waveform-preserving quantizer for spectral lines, or there is a distinction between the core coder and bandwidth extension [1][5][6][7-14]. Even if IGF [7-14] allows preservation of spectral lines across the entire bandwidth, it requires a spectral analyzer operating before the spectral domain encoder, so it is not possible to parametrically select which parts of the spectrum to code depending on the results of the spectral domain encoder. The PNS in [4] determines which subbands to zero out based solely on their tonality before quantization, and uses only random noise for subband replacement.

[15] only considers parametric coding of a single tonal component. Before the quantizer, a decision is made as to which spectral lines to parametrically code, using only a simple maximum value. The result of the quantizer is not used to determine which spectral lines to parametrically code. Non-zero pseudo-coefficients must be coded within the spectrum, and coding non-zero coefficients is almost always more expensive than coding zero coefficients. In addition to coding the pseudo-coefficients, side information is required to distinguish them from waveform-preserving spectral coefficients. Therefore, a lot of information needs to be conveyed to generate a signal with many tonal components. The method also does not propose any solution for the non-tonal parts of the signal. Additionally, the computational complexity of generating a signal with many tonal components that is parametrically coded is very high.

[0007] In

[16] , the high computational complexity compared to

[15] is reduced by using spectral patterns instead of time-domain generators. Furthermore, only predetermined patterns or their modifications are used to replace the pseudo-coefficients, thus requiring a lot of storage or limiting the range of possible tones that can be generated. Other drawbacks from

[15] remain in

[16] . Noise filling in [1][2][3] and similar methods provides replacement of spectral lines quantized to zero, but the spectral resolution is very low, typically using only a single level for the entire bandwidth. The IGF has a predefined subband division and the spectral envelope is propagated for the complete IGF range without the possibility to adaptively propagate the spectral envelope only for some subbands. In [5], only the properties of the autocorrelation of the amplitude spectrum and a predefined constant are used to select the offset used in the modulator: only one offset is found for the entire spectral bandwidth. In [6], only one modulation frequency for the whole bandwidth is used for frequency shifting, and the modulation frequency is calculated based on the fundamental frequency only.

[0008] In

[11] , only predefined source tiles below the IGF start frequency are used to satisfy the IGF target range, which is above the start frequency. The tile selection is determined by adaptive coding and therefore needs to be coded in the bitstream. The proposed brute-force approach has high computational complexity. In IGF, the source tiles are obtained below the IGF start frequency, and therefore waveform-preserving core-coded salient tones located above the IGF start frequency are not used. There is also no mention of using synthesized low-frequency content and waveform-preserving core-coded salient tones located above the IGF start frequency as source tiles. This indicates that the IGF is a tool added to the core coder, not an integral part of it. Methods using deadbands

[17]

[18] attempt to estimate the range of values ​​of spectral coefficients that should be set to zero. Because they do not use the actual output of quantization, the estimation is prone to errors. Summary of the Invention [Problem to be solved by the invention]

[0009] It is an object of the present invention to provide a concept for efficient coding, in particular for efficient parametric coding. [Means for solving the problem]

[0010] This object is solved by the subject matter of the independent claims. One embodiment provides a spectral representation of an audio signal (X MR ) and provides an encoder for encoding the spectral representation (X MR ) consists of frequency bins or frequency coefficients, and at least one subband contains two or more frequency bins. The encoder includes a quantizer and a band-wise parametric coder. The quantizer generates a spectral representation (X MR ) quantized representation (X Q ) The per-band parametric coder is configured to generate a quantized representation (X Q ), for example, for each band, the spectral representation (X MR ), where the coded parametric representation (zfl) consists of parameters describing the energy in the subbands, or coded versions of parameters describing the energy in the subbands, and there are corresponding parameters describing the energy in at least two different subbands, and therefore at least two different subbands. Note that the at least two subbands may belong to multiple subbands. One aspect of the present invention is based on the finding that an audio signal or a spectral representation of an audio signal divided into several subbands can be coded efficiently band-wise (band-wise may mean band / subband-wise). According to an embodiment, the concept allows restricting parametric coding only in subbands quantized to zero by the quantizer (used to quantize the spectrum). This concept allows efficient joint coding of spectral and band-wise parameters, resulting in a high spectral resolution for the parametric coding, although lower than the spectral resolution of the spectral coder. The resulting coder is defined as an integrated band-wise parametric coding entity within a waveform-preserving coder. According to an embodiment, the band-wise parametric coder together with the spectral coder encodes the spectral representation (X MR ) together to obtain a coded version of the joint coder. This joint coder concept has the advantage that the bit rate distribution between the two coders can be done jointly.

[0011] According to a further embodiment, at least one subband is quantized to zero. For example, a parametric coder identifies which subbands are zero and codes (only) a representation of the subbands that are zero. According to an embodiment, at least two subbands may have different parameters. According to an embodiment, the spectral representation is perceptually flattened. This can be done, for example, by using a spectral shaper configured to provide a perceptually flattened spectral representation from the spectral representation based on a spectral shape obtained from the coded spectral shape. It should be noted that the perceptually flattened spectral representation is divided into subbands with a different or higher frequency resolution than the coded spectral shape. According to a further embodiment, the encoder may further comprise a time-spectral transformer, such as an MDCT transformer, configured to transform the audio signal having the sampling rate into a spectral representation. Starting from said extension, the per-band parametric coder is configured to provide a parametric representation of a perceptually flattened spectral representation or a derivative of a spectrally flattened spectral representation, the parametric representation may depend on an optimal quantization step and may consist of parameters describing the energy within a subband, the quantized spectrum being zero, so that at least two subbands have different parameters or at least one parameter is limited to only one subband.

[0012] According to a further embodiment, the spectral representation is used to determine an optimal quantization step. For example, the encoder can be enhanced by the use of a so-called rate-distortion loop configured to determine the quantization step. This allows said rate-distortion loop to determine or estimate the optimal quantization step used above. This may be done in such a way that said loop performs several (at least two) iteration steps, the quantization step being adapted depending on one or more previous quantization steps. To code the representation of the quantized spectrum, the encoder may further comprise a lossless spectral coder. According to a further embodiment, the encoder comprises a spectral coder and / or a spectral coder decision entity configured to provide a decision whether joint coding of the coded representation of the quantized spectrum and the coded representation of the parametric spectrum satisfies the constraint that the total number of bits for joint coding is below a predetermined threshold. This is particularly sensible when both the coded representation of the quantized spectrum and the coded representation of the parametric spectrum are based on a variable number of bits (optional feature) that depends on the spectral representation or on the derivative of the perceptually flattened spectral representation and the quantization step. According to a further embodiment, both the parametric coder per band and the spectral coder form a joint coder that allows interaction, e.g., to take into account parameters used for both the variable number of bits or the quantization step.

[0013] According to a further embodiment, the encoder further comprises a modifier configured to adaptively set at least subbands in the quantization step to zero depending on the quantized spectrum and / or the content of the subbands in the spectral representation. According to a further embodiment, the per-band parametric coder comprises two stages, wherein a first of the two stages of the per-band parametric coder is configured to provide individual parametric representations of subbands above a certain frequency, and a second of the two stages provides an additional average parametric representation for subbands above a certain frequency, e.g. based on the parametric representations of the (individual) subbands, and the individual parametric representations are zero, for subbands below a certain frequency.

[0014] According to an embodiment, the encoder comprises the following method: -Spectral representation of an audio signal divided into multiple subbandsX MR quantized representation of X Q generating a -Quantized representation Q Spectral expression according to X MR providing a coded parametric representation zfl of the spectral representation X in the subband, MR and a spectral representation X MR There are parameters describing the steps and The present invention may be implemented by a method for encoding an audio signal, including: Here, there are at least two sub-bands that are different, and therefore the parameters describing the energy in the at least two sub-bands are different.

[0015] Another embodiment provides a decoder. The decoder includes a spectral domain decoder and a per-band parametric decoder. The spectral domain decoder is configured to generate a decoded spectrum or a dequantized (and decoded) spectrum based on an encoded audio signal, where the decoded spectrum is divided into subbands. Optionally, the spectral domain decoder is used to decode / dequantize information related to the quantization step. The per-band parametric decoder is configured to identify a zero subband in the decoded and / or dequantized spectrum and decode a parametric representation of the zero subband based on the encoded audio signal. Note that the parametric representation includes parameters describing the subband, such as the energy within the subband, and there are parameters describing at least two different subbands, and thus at least two different subbands, and the identification can be performed based only on the decoded and dequantized spectrum or on the spectrum processed by the spectral domain decoder without the dequantization step, referred to as the decoded spectrum. Additionally or alternatively, the coded parametric representation is coded using a variable number of bits, and / or the number of bits used to represent the coded parametric representation depends on the spectral representation of the audio signal. In other words, this means that the decoder is configured to generate a decoded output from the jointly coded spectral and per-band parameters.

[0016] Another embodiment provides another decoder having the following entities: a spectral domain decoder, a per-band parametric decoder combined with a per-band spectrum generator, a synthesizer, and a spectrum-to-time converter. The spectral domain decoder and the per-band parametric decoder may be defined as above, or another parametric decoder such as from IGF (see [7-14]) may be used. The per-band spectrum generator is configured to generate a per-band generated spectrum according to a parametric representation of the zero subband. The synthesizer is configured to provide a per-band synthesized spectrum, which may include a synthesis of the per-band generated spectrum and a decoded spectrum, or a synthesis of the per-band generated spectrum, or a synthesis of the predicted spectrum and a decoded spectrum. The spectrum-to-time converter is configured to convert the per-band synthesized spectrum or a derivative thereof (e.g., a reshaped spectrum reshaped by SNS or TNS, or alternatively by using an LP predictor) into a time representation.

[0017] The per-band parametric decoder, according to an embodiment, generates a parametric representation of the zero subband (E) based on the encoded audio signal using a quantization step. B ) According to a further embodiment, the decoder comprises a spectral shaper configured to provide a reshaped spectrum from the composite spectrum per band or a derivative of the composite spectrum per band. For example, the spectral shaper can use a spectral shape obtained from a coded spectral shape of a different or lower frequency resolution than the sub-band decomposition. According to a further embodiment, the parametric representation is composed of parameters describing the energy in the zero subband, such that at least two subbands have different parameters or at least one parameter is limited to only one subband. Note that the zero subband is defined by the decoded and / or dequantized spectral output of the spectral decoder.

[0018] According to another embodiment, a per-band parametric spectrum generator may be provided together with the decoder or independently. The parametric spectrum generator is configured to generate a generated spectrum to be added to the decoded and dequantized spectrum or to the combination of the predicted spectrum and the decoded spectrum. It should be noted that the step of adding to the decoded and dequantized spectrum is performed, for example, when there is no LTP in the system. Here, the generated spectrum (X G ) may be obtained band-wise from the source spectrum, which is - the second predicted spectrum (X NP ),or -Random noise spectrum (X N ),or - the already generated part of the generation spectrum, or - a combination of one of the above It is one of them. The decoder may be implemented by a method. A method for decoding an audio signal comprises: - the spectrum (X) decoded and dequantized from the coded representation of the spectrum (spect) D ), generating the decoded and dequantized spectrum (X D ) is divided into subbands; - the decoded and dequantized spectrum (X D ) and based on the coded parametric representation (zfl), identify the zero subband (E B ) and Contains Parametric Expression (E BIt should be noted that the coded parametric representation (zfl) contains parameters describing the subbands, and there are at least two different subbands, and thus there are parameters describing at least two different subbands, and / or the coded parametric representation (zfl) is coded using a variable number of bits, and / or the number of bits used to represent the coded parametric representation (zfl) depends on the coded representation of the spectrum (spect).

[0019] Alternatively, the method may be - the decoded and dequantized spectrum (X D ), generating the decoded and dequantized spectrum (X D ) is divided into subbands; - based on the coded audio signal, the decoded and dequantized spectrum (X D ) and determine the parametric representation of the zero subband (E B ) and -Parametric representation of the zero subband (E B ) generating a generated spectrum for each band according to -Synthetic spectrum for each band (X CT ), providing a composite spectrum per band (X CT ) is the generated spectrum for each band and the decoded and dequantized spectrum (X D ) or the synthesis of the generated spectrum for each band, and the predicted spectrum (X PS ) and the decoded and dequantized spectrum (X D ) synthesis (X DT ), and -Synthetic spectrum for each band (X CT ) or the composite spectrum per band (X CT ) to convert its derivative to a time representation. Includes.

[0020] The generator described above may be implemented by a method for generating a generated spectrum to be added to a decoded and dequantized spectrum or to a combination of a predicted spectrum and a decoded spectrum, the generated spectrum being obtained band-by-band from a source spectrum, the source spectrum being a second predicted spectrum, or - random noise spectrum, or - the already generated part of the generation spectrum, or - a combination of one of the above It is one of them. Note that the source spectrum can be derived from any of the listed possibilities.

[0021] According to an embodiment, the source spectrum is weighted based on the energy parameter of the zero subband. According to a further embodiment, the selection of the source spectrum for a subband depends on the subband position, tonality information, the power spectrum estimate, the energy parameter, pitch information, and / or time information. The tonality information is φ H and / or the pitch information may be Note that the time information may be TIFF2026042037000002.tif6150 and / or the time information may be information on whether TNS is active or not. According to an embodiment, the source spectrum is weighted based on the energy parameter of the zero band. It should be noted that all of the methods described above may be implemented using computer programs. Embodiments of the present invention will now be described with reference to the enclosed figures. [Brief explanation of the drawings]

[0022] [Figure 1a] 1 is a schematic diagram of a basic implementation of an encoder with a parametric coder per band according to one embodiment; [Figure 1b]FIG. 10 is a schematic diagram of another implementation of an encoder with a parametric coder per band according to one embodiment; [Figure 1c] FIG. 2 is a schematic diagram of a decoder implementation according to one embodiment. [Figure 2a] 2 is a schematic block diagram illustrating an encoder according to one embodiment and a decoder according to another embodiment; [Figure 2b] FIG. 2b is a schematic block diagram illustrating an excerpt from FIG. 2a according to one embodiment. [Figure 2c] FIG. 2b is a schematic block diagram of the extract of FIG. 2a including a decoder according to another embodiment; [Figure 3] 4 is a schematic block diagram of a signal encoder for a residual signal according to an embodiment and a decoder according to another embodiment; [Figure 4] FIG. 10 is a schematic block diagram of a decoder including the zero-filling principle according to a further embodiment; [Figure 5] 1 is a schematic diagram for explaining the principle of determining a pitch contour (see block gap pitch contour) according to an embodiment; [Figure 6] FIG. 10 is a schematic block diagram of a pulse extractor that uses information about the pitch contour according to a further embodiment; [Figure 7] FIG. 10 is a schematic block diagram of a pulse extractor using pitch contour as additional information according to an alternative embodiment. [Figure 8] FIG. 10 is a schematic block diagram of a pulse coder according to a further embodiment; [Figure 9a] 1 is a schematic diagram illustrating the principle of spectrally flattening a pulse according to an embodiment; [Figure 9b] 1 is a schematic diagram illustrating the principle of spectrally flattening a pulse according to an embodiment; [Figure 10] FIG. 10 is a schematic block diagram of a pulse coder according to a further embodiment; [Figure 11a] 1 is a schematic diagram illustrating the principle of determining a prediction residual signal starting from a flattened original pulse waveform; [Figure 11b]1 is a schematic diagram illustrating the principle of determining a prediction residual signal starting from a flattened original pulse waveform; [Figure 12] FIG. 10 is a schematic block diagram of a pulse coder according to a further embodiment; [Figure 13] FIG. 2 is a schematic diagram showing a residual signal and coded pulses for explaining an embodiment. [Figure 14] FIG. 10 is a schematic block diagram of a pulse decoder according to a further embodiment; [Figure 15] FIG. 10 is a schematic block diagram of a pulse decoder according to a further embodiment; [Figure 16] 1 is a schematic flowchart illustrating the principle of estimating an optimal quantization step (i.e., step size) using block IBPC according to an embodiment. [Figure 17a] FIG. 1 is a schematic diagram for explaining the principle of long-term prediction according to an embodiment; [Figure 17b] FIG. 1 is a schematic diagram for explaining the principle of long-term prediction according to an embodiment; [Figure 17c] FIG. 1 is a schematic diagram for explaining the principle of long-term prediction according to an embodiment; [Figure 17d] FIG. 1 is a schematic diagram for explaining the principle of long-term prediction according to an embodiment; [Figure 18a] FIG. 10 is a schematic diagram for explaining the principle of harmonic post-filtering according to a further embodiment; [Figure 18b] FIG. 10 is a schematic diagram for explaining the principle of harmonic post-filtering according to a further embodiment; [Figure 18c] FIG. 10 is a schematic diagram for explaining the principle of harmonic post-filtering according to a further embodiment; [Figure 18d] FIG. 10 is a schematic diagram for explaining the principle of harmonic post-filtering according to a further embodiment; DETAILED DESCRIPTION OF THE INVENTION

[0023] Hereinafter, embodiments of the present invention will be described subsequently with reference to the enclosed figures, in which objects having the same or similar functions are given the same reference numbers, and the descriptions thereof are mutually applicable and interchangeable. 1a shows an encoder 1000 comprising a quantizer 1030, a per-band parametric coder 1010, and an optional (lossless) spectral coder 1020. Before describing the per-band parametric coder 1010, the surroundings of the per-band parametric coder will be described. Around the parametric coder 1010, the encoder 1000 comprises several optional elements.

[0024] According to an embodiment, the parametric coder 1010 is combined with a spectral coder or a lossless spectral coder 1020 to form a joint coder 1010 plus 1020. The signal processed by the joint coder 1010 plus 1020 is provided to a quantizer 1030, which quantizes a spectral representation X of the audio signal divided into a number of subbands. MR is used as input. The quantizer 1030 calculates the X MR to obtain a quantized representation X of the spectral representation X of the audio signal (divided into multiple subbands). Q Optionally, the quantizer may be configured to provide a quantized spectrum of the perceptually flattened spectral representation or a derivative of the perceptually flattened spectral representation. The quantization may depend on an optimal quantization step according to a further embodiment that is iteratively determined (see FIG. 16 ). Both coders 1010 and 1020 use the quantized representation X Q , i.e., the signal X preprocessed by a quantizer 1030 and an optional corrector (not shown in FIG. 1a, but shown as 156m in FIG. 3). MR The parametric coder 1010 receives X Q Check which subbands in X are zero Q for subbands that are zero in X MRWith respect to the modifier, it should be noted that the modifier provides the quantized and modified audio signal to the joint coder 1010 plus 1020 (as shown in FIG. 3). For example, the modifier can set different subbands to zero, as described with respect to FIG. 16 (in FIG. 16 the modifier is marked with 302).

[0025] According to an embodiment, the coded parametric representation (zfl) uses a variable number of bits. For example, the number of bits used to represent the coded parametric representation (zfl) may vary depending on the spectral representation (X MR ) depends on According to an embodiment, the coded representation (spect) uses a variable number of bits or the number of bits used to represent the coded representation (spect) varies depending on the spectral representation (X MR ) Note that the coded representation (spect) may be obtained by a lossless spectral coder. According to an embodiment, the (total) number of bits required to represent the coded parametric representation (zfl) and the coded representation (spect) may be below a predetermined limit. According to an embodiment, the parameters are the quantized representation (X Q ) is zero (i.e., X Q describes the energy only within the subbands where all frequency bins of X are zero. Other parametric representations of the zero subband may be used. This is called the "quantized representation (X Q ) may be a specification that depends on According to an embodiment, the per-band parametric coder 1010 is configured to provide a parametric description of the subbands quantized to zero. The parametric representation is based on the optimal quantization step (step size in FIG. 16 and step size in FIG. 3). The quantized spectrum may depend on the parametric coding scheme (see TIFF2026042037000003.tif6150) and may consist of parameters describing the energy in subbands where the quantized spectrum is zero, so that at least two subbands have different parameters or at least one parameter is limited to only one subband. The lossless spectral coder 1020 is configured to provide a coded representation of the (quantized) spectrum. This joint coding 1010 plus 1020 is highly efficient, in particular allowing the high spectral resolution of the parametric coding 1010, but lower than the spectral resolution of the spectral coder 1020. The above approach further allows restricting parametric coding only within subbands quantized to zero by the quantizer used to quantize the spectrum. Due to the use of a modifier, it is further possible to provide an adaptive method of distributing bits between the per-band parametric coder 1010 and the spectral coder 1020, each coder taking into account the bit demands of the other and allowing the realization of bit rate limits. According to a further embodiment, the encoder 1000 may comprise an entity such as a splitter (not shown) configured to split the spectral representation of the audio signal into said subbands. Optionally or additionally, the encoder 1000 may comprise in the upstream path a TD-to-FD transformer (not shown), such as an MDCT transformer (see entity 152, MDCT or equivalent) configured to provide a spectral representation based on the time-domain audio signal. A further optional element is a temporal noise shaping (TNS E、 154 in Figure 2a), as well as the spectral shaper SNS / temporal noise shaping TNS. E Signal X MS , X MT , and X PS is an entity 155 that synthesizes the above. At the output of the audio signals 1010 plus 1020, a bitstream multiplexer (not shown) may be located, the purpose of which is to combine the parametrically coded bitstream and the spectrally coded bitstream per band.

[0026] According to an embodiment, the output of the MDCT 152 has length L M X M For example, for an input sampling rate of 48 kHz and an exemplary frame length of 20 ms, L M is equal to 960. The codec may operate at other sampling rates and / or other frame lengths. X M All other spectra derived from:X MS , X MT , X MR , X Q , X D , X DT , X CT , X CS , X C , X P , X PS , X N , X NP , X S are also the same length L M The spectrum may be of any order, but in some cases, only a portion of the spectrum may be needed and used. The spectrum is made up of spectral coefficients, also known as spectral bins or frequency bins. In the case of an MDCT spectrum, the spectral coefficients may have positive and negative values. Each spectral coefficient is said to cover a bandwidth. For a sampling rate of 48 kHz and a frame length of 20 ms, the spectral coefficients cover a bandwidth of 25 Hz. The spectral coefficients may range, for example, from 0 to L M May be indexed down to -1.

[0027] Social Media E and social media D The SNS scale factor used in (see Figure 2a) is N SBThe energy may be obtained from the energy within 64 frequency subbands (often called bands), where the energy is obtained from the spectrum divided into the frequency subbands. For example, the subband boundaries, expressed in Hz, are: 0, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2050, 2200, 2350, 2500, 2650, 2800, 2950, ​​3100, 3200, 3300, 3400, 3500, 3650, 3700, 3800, 3950, 4000, 4100, 4200, 4300, 4400, 4500, 4650, 4700, 4800, 4950, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5950, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6950, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7 The subbands may be set to 00, 3500, 3700, 3900, 4100, 4350, 4600, 4850, 5100, 5400, 5700, 6000, 6300, 6650, 7000, 7350, 7750, 8150, 8600, 9100, 9650, 10250, 10850, 11500, 12150, 12800, 13450, 14150, 15000, 16000, or 24000. SB They may be indexed down to -1. In this example, the 0th subband (0-50 Hz) contains 2 spectral coefficients, the same as subbands 1-11, subband 62 contains 40 spectral coefficients, and subband 63 contains 320 coefficients.

[0028] N SB The energy in the = 64 frequency subbands may be downsampled to 16 coded values, denoted "sns" (see Figures 2a, 2b, and 2c). The 16 decoded values ​​obtained from "sns" are interpolated to SNS scale factors; for example, there may be 32, 64, or 128 scale factors. For further details on obtaining SNS, the reader is referred to [21-25]. In iBPC, "zfl decoding", and / or "zero filling" blocks, the spectrum is TIFF2026042037000004.tif6150 may be divided into subbands Bi, where subband i is Starting at TIFF2026042037000005.tif6150. The same 64 subband boundaries used for the energy to obtain the SNS scale factors may be used, but any other number of subbands and any other subband boundaries may be used, independent of SNS. To emphasize, the same subband division principles as in SNS may be used, but the subband division in the iBPC, "zfl decoding", and / or "zero filling" blocks may be different from those in SNS, as well as from SNS. E and independent of the SNS blocks. In the subband division example above, TIFF2026042037000006.tif6150 and TIFF2026042037000007.tif6150, TIFF2026042037000008.tif6150 and TIFF2026042037000009.tif6150,…, TIFF2026042037000010.tif6150 and TIFF2026042037000011.tif6150, TIFF2026042037000012.tif6150 and The file is TIFF2026042037000013.tif6150. In another example, the iBPC may include an SNS at the input of the time-to-frequency transformer (e.g., at the input of 152). E is replaced by the LP analysis filter, and at the output of the frequency-to-time converter (e.g., at the output of 161) SNS D is sometimes used in codecs where it replaces the LP synthesis filter. According to a further embodiment, the per-band parametric coder 1010 is integrated into the rate-distortion loop (see FIG. 16) thanks to an efficient modification of the quantized spectrum, as illustrated by FIG. 1b.

[0029] FIG. 1b shows a portion of a rate-distortion loop 1001. The portion of the rate-distortion loop 1001 includes a quantizer 1030, parametric and spectral coders 1010-1020 per joint band, a bit counter 1050, and a recoder 1055. The recoder 1055 is configured to record spectral and band-specific parameters (e.g., as shown in detail in FIG. 16). For example, the bit counter 1050 can estimate / calculate / recalculate the bits required for coding a spectral line to arrive at an efficient way to store the bits required for coding. In other words, an estimation of the maximum number of bits required for coding may be performed instead of the actual coding. This helps to perform efficient coding with a limited bit budget. Note that FIG. 1b shows a portion of FIG. 16. Here, 1030 corresponds to 301, 1010 + 1020 corresponds to 303, 1050 is 304, and 1055 corresponds to the "recoder." Therefore, according to an embodiment, the rate distortion loop includes a bit counter 1050 configured to estimate or calculate the bits used for coding and / or the spectral representation (X MR ), for example, spectral parameters and band parameters.

[0030] Note that although the same blocks are used in Figures 1a and 1b to indicate that the blocks have the same functionality, the entities in Figure 1a (part of Figure 3) are different from the entities in Figure 1b (part of Figure 16). The decoder 1200 will now be described with reference to Figure 1c. Figure 1c shows a decoder for decoding an audio signal. It comprises a spectral domain decoder 1230, with a per-band parameter decoder 1210 arranged in a processing path with a per-band spectrum generator 1220, which uses the output of the spectral decoder 1230. Both decoders have outputs to a synthesizer 1240, with a spectro-temporal converter 1250 arranged at the output of the synthesizer 1240.

[0031] The spectral domain decoder 1230 (which may comprise an inverse quantizer combined with the decoder) generates the inverse quantized spectrum (X D ), where the dequantized spectrum is divided into subbands. The per-band parameter decoder 1210 identifies zero subbands, i.e. subbands consisting only of zeros, in the dequantized spectrum and decodes the energy parameters of the zero subbands, which are defined by the dequantized spectral output of the spectral decoder. For this purpose, a quantized representation (X Q ) may be used, since which subbands have parametric representation depends on the decoded spectrum obtained from spect. Note that the output of 1230 used as input to 1220 can have information about the decoded spectrum or its derivatives, such as information about the dequantized spectrum, since both the decoded spectrum and the dequantized spectrum can have the same zero subbands. The decoded spectrum obtained from spect may contain the same information as the input to 1010+1020 in Figure 1a. Quantization Step TIFF2026042037000014.tif6150 is the dequantized spectrum (X D ) The location of the zero subband in the decoded and / or dequantized spectrum may be determined by the quantization step This may be determined independently of TIFF2026042037000015.tif6150.

[0032] Starting from this, the band-wise spectrum generator 1220 generates a band-wise generated spectrum X in response to the parametric representation of the zero subband. G The synthesizer 1240 provides a band-by-band synthesized spectrum X CT For example, the synthetic spectrum X CT In this case, the following synthesis is possible: -Generated spectrum for each band X G and the decoded spectrum X D ,or -Generated spectrum for each band X G and the combined predicted and decoded spectrum X DT . In other words, the interaction of entities 1220, 1230 with entity 1240 can be described as follows: The per-band parametric spectrum generator 1220 generates a generated spectrum X G Provide the generated spectrum X G is obtained band-wise from the source spectrum, and the source spectrum is converted to a second predicted spectrum X NP , or the random noise spectrum X N , or the already generated part of the generated spectrum, or a combination thereof. CT is X G Note that this may include

[0033] X CT The already generated part of X G, which may be used to generate the source spectrum. The source spectrum may be weighted based on the energy parameter of the zero subband. The selection of the source spectrum for a subband may be based on band position, tonality, a power spectrum estimate, an energy parameter, a pitch parameter, and temporal information. This method obtains the selection of the subband to be parametrically coded based on the decoded spectrum, thus avoiding additional side information in the bitstream. Another aspect of this adaptive method is that, for each subband, it is determined which source spectrum to use to replace the zeros in the subband is provided to the decoder 1200, thus avoiding additional side information in the bitstream and allowing multiple possibilities for source spectrum selection. The output of the combiner 1240 is optionally combined with a TNS or SNS to obtain a so-called reshaped spectrum. D (not shown). Based on the output of the synthesizer 1240 or based on this reshaped spectrum, an optional spectrum-to-time converter 1250 outputs a time representation. According to further embodiments, the decoder 1200 may comprise a spectral shaper for providing a reshaped spectrum from the band-wise synthesized spectrum or from a derivative of the band-wise synthesized spectrum.

[0034] According to further embodiments, the encoder may comprise a spectral coder decision entity for providing a decision whether joint coding of the quantized spectral coded representation and the parametric zero sub-band coded representation satisfies a constraint that the total number of bits of the joint coding is below a predetermined threshold, where both the quantized spectral coded representation and the parametric zero sub-band coded representation may use a perceptually flattened spectral representation or a variable number of bits and / or quantization step that depends on a derivative of the perceptually flattened spectral representation. As described above, the per-band parametric spectrum generator and synthesizer 1240 may be implemented as follows: The per-band parametric spectrum generator provides a generated spectrum for each band and adds it to the decoded spectrum or to a combination of the predicted spectrum and the decoded spectrum. The generated spectrum is obtained for each band from the source spectrum, and the source spectrum is a second predicted spectrum, a random noise spectrum, an already generated portion of the generated spectrum, or a combination thereof. The source spectrum may be weighted based on an energy parameter for the zero band. The use of the already generated portion of the generated spectrum provides a combination of any two separate portions of the decoded spectrum, thus providing a harmonic or tonal source spectrum that is not available by using only one portion of the decoded spectrum. The combination of the second predicted spectrum and the source spectrum is another advantage for creating a harmonic or tonal source spectrum that is not available by using only the decoded spectrum.

[0035] FIG. 2 a shows the encoder 101 combined with a decoder 201 . The main entities of the encoder 101 are marked by the reference numerals 110, 130 and 150. Entity 110 performs pulse extraction and the pulse p is encoded using entity 132 for pulse coding. The signal encoder 150 is implemented by several entities 152, 153, 154, 155, 156, 157, 158, 159, 160, and 161. These entities 152-161 form the main path of the encoder 150, and in parallel, additional entities 162, 163, 164, 165, and 166 may be deployed. Entity 162 (zfl decoder) informatively connects entity 156 (iBPC) with entity 158 for zero filling. Entity 165 (TNS acquisition) informs entity 153 (SNS E) informatively connects entity 152 with entities 154, 158, and 159. Entity 166 (SNS Acquisition) informatively connects entity 152 with entities 153, 163, and 160. Entity 158 performs zero-filling and may comprise combiner 158c, which will be described in the context of FIG. 4. Note that there may be implementations in which entities 153 and 160 are not present, e.g., systems with LP analysis filtering of the MDCT input and LP synthesis filtering of the IMDCT output. These entities 153 and 160 are therefore optional.

[0036] Entities 163 and 164 represent the predicted spectrum X P and / or perceptually flattened prediction X PS , the pitch contour and coding residual y from entity 180 C The functions and interactions of the different entities are described below. Before describing the functionality of encoder 101, and in particular the functionality of encoder 150, a short description of decoder 210 is given. Decoder 210 may comprise entities 157, 162, 163, 164, 158, 159, 160, 161, as well as encoder-specific entities 214 (HPF), 23 (signal combiner), and 22 (for decoding and recovering pulse portions comprising the recovered pulse waveform).

[0037] The encoding function is explained below. The pulse extractor 110 extracts the input audio signal PCM I The STFT of y is taken and the nonlinear amplitude and phase spectrograms of the STFT are used to find and extract pulses, each of which has a waveform with high-pass characteristics. M is obtained by removing the pulses from the input audio signal. The pulses are coded by pulse coding 132 and the coded pulses CP are transmitted to the decoder 201. Pulse residual signal y M is the length L M XM The window is selected from among three windows as in

[19] . In the example below, the longest window is 30 ms long with a 10 ms overlap, but any other window and overlap length may be used. X M The spectral envelope of X MS Get SNS E 153. Optionally, temporal noise shaping TNS E 154 is applied to flatten the time envelope in at least a portion of the spectrum, and X MT Generate a part of the spectrum (X M or X MS or X MT ) at least one tonality flag φ H may be estimated and communicated to the decoder 201 / 210. Optionally, a long-term prediction LTP 164 following the pitch contour 180 may estimate the spectrum X predicted from past decoded samples. P is used to construct the perceptually flattened prediction X PS is X MT is subtracted in the MDCT domain from the LTP residual X MR The pitch contour 180 is obtained for frames with high average harmonics and communicated to the decoder 201 / 210. The pitch contour 180 and harmonics are used to drive many parts of the codec. The average harmonics may be calculated for each frame.

[0038] Figure 2b shows an excerpt from Figure 2a focusing on encoder 101' with entities 180, 110, 152, 153, 153, 155, 156', 165, 166, and 132. Note that 156 in Figure 2a is a kind of composition of 156' in Figure 2b and 156'' in Figure 2c. Note that entity 163 (in Figures 2a, 2c) may be the same as or equivalent to 153 and is the inverse of 160. According to an embodiment, the encoder divides the input signal into frames and outputs, for example for each frame, at least one or more of the following parameters: -Pitch contour -MDCT window selection, 2 bits -LTP parameters -Coded Pulse -sns, i.e., coded information for spectrum shaping via SNS -tns, i.e., coding information for time shaping via TNS - global gain gQo, i.e. the global quantization step size for the MDCT codec -spect, which consists of the entropy-coded quantized MDCT spectrum -zfl, consisting of a quantized parametrically coded zero section X PS comes from LTP, which is also used in the encoder, but LTP is only shown in the decoder (see Figures 2a and 2c).

[0039] Figure 2c shows an excerpt from Figure 2a focusing on encoder 201' comprising entities 156'', 162, 163, 164, 158, 159, 160, 161, 214, 23, and 22 described in the context of Figure 2a. Regarding LTP 164: Essentially, LTP is part of the decoder (except for the HPF, "waveform building" and their outputs) which may also be used / required in the encoder (as part of the inner decoder). In implementations without LTP, no inner decoder is needed within the encoder. X emitted by entity 155 MR The coding (of the residual from LTP) is done in an integral band-wise parameter coder (iBPC) as described with respect to FIG.

[0040] Figure 3 shows entity iBPC 156, which may have sub-entities 156q, 156m, 156pc, 156sc, and 156mu. Figure 1a shows a portion of Figure 3, where 1030 corresponds to 156q, 1010 corresponds to 156pc, and 1020 corresponds to 156sc. At the output of the bitstream multiplexer 156mu, a per-band parametric decoder 162 is located together with a spectral decoder 156sd. The entity 162 receives the signal zfl and the entity 156sd receives the signal spect, both of which have a global gain / step size g Q0 The parametric decoder 162 receives the output X of the spectral decoder 156sd to decode zfl. D Note that the spectral decoder 156sd uses the inputs X of 156pc and 156sc. Alternatively, another signal output from the decoder 156sd may be used. The background is that the spectral decoder 156sd may comprise two parts: a spectral lossless decoder and an inverse quantizer. For example, the output of the spectral lossless decoder may be the decoded spectrum obtained from spect and used as input to the parametric decoder 162. The output of the spectral lossless decoder is the inputs X of 156pc and 156sc. Q The inverse quantizer uses a global gain / step size to extract X from the output of the spectral lossless decoder. D The decoded spectrum and / or the dequantized spectrum X can be derived. D The position of the zero subband in This may be determined independently of TIFF2026042037000016.tif6150.

[0041] X MR is the quantized spectrum X Q quantized and coded, including quantizing and coding the energy of (a part of) the zero values ​​of X Q is X MR is a quantized version of XMR The quantization and coding of It will be held at iBPC156. As part of iBPC, quantization (quantizer 156q) along with adaptive band zeroing 156m, optimizes the optimal quantization step size. Quantized spectrum X based on TIFF2026042037000017.tif6150 Q Generate. iBPC156 is (X Q )spect 156sc and (X Q The coded information is generated by generating coded information consisting of zfl162 (which can represent zero-value energy in a portion of the A zero-filling entity 158 placed at the output of entity 157 is shown by FIG.

[0042] FIG. 4 shows the signal E B and a zero-filling entity 158 that receives the predicted spectrum (X PS ) and the decoded and dequantized spectrum (X D ) synthesis (X DT ) The zero-filling entity 158 may include two sub-entities 158sc and 158sg and a combiner 158c. spect is X MR The dequantized spectrum corresponds to the quantized version of XD (decoded LTP residual, error spectrum). E B is X D It is obtained from zfl taking into account the position of zero values ​​in E B is X Q may be a smoothed version of the zero-value energy in E B can have a different resolution than zfl, preferably a higher resolution resulting from smoothing.

[0043] E B (see 162) and then obtain the perceptually flattened prediction X PS is the optionally decrypted X D is added to X DT is generated. Zero-filled X G is obtained and then "zero-filled" (e.g., using addition 158c) to obtain X DT and zero-filled X G is E B The source spectrum for each band is weighted based on Source spectrum X consisting of TIFF2026042037000018.tif6150 (see 156sc) S Band-wise zero-filling obtained iteratively from It consists of TIFF2026042037000019.tif6150. X CT Zero-filled X G and Spectrum X DT This is a band-by-band synthesis of (158c).

[0044] X S is configured by band (158sg is X G ), X CT is obtained band-by-band starting from the lowest subband. For each subband, the source spectrum is calculated using, for example, the subband position, the tonality flag (toi), and X DT The power spectrum estimated from E B , pitch information (pii), and time information (tei) (see 158sc). X DT The power spectrum estimated from X DT or X D Note that the starting frequency f may be derived from the source spectrum. Alternatively, the selection of the source spectrum may be obtained from the bitstream. ZFstart X up to S The lowest subband in TIFF2026042037000020.tif6150 may be set to 0, which means that XCT is X DT This means that it may be a copy of f ZFStart can be 0, which means that a source spectrum different from zero can also be selected from the beginning of the spectrum. The source spectrum for subband i can be, for example, a random noise or predicted spectrum, or X CT The source spectrum X can be a composite of the already obtained sub-portion of X, random noise, and the predicted spectrum. S Zero-filled X G To obtain E B are weighted based on

[0045] The weighting may be performed, for example, by entity 158sg and may have a higher resolution than the subband decomposition, which may be determined sample by sample to obtain a smoother weighting; TIFF2026042037000021.tif6150 is X DT is added to subband i of X CT Generate subband i of the complete X CT After obtaining the time envelope, the time envelope can be optionally calculated using the TNS D 159 (see Figure 2a) through X MS is modified to match the time envelope of X CS Then, X CS The spectral envelope of X M SNS to match the spectral envelope of D 160 and fixed with X C Generate the time domain signal y C is the output of the IMDCT161. C The IMDCT 161 is obtained from the inverse MDCT, windowing, and overlap-add. C is used to update the LTP buffer 164 (corresponding either to buffer 164 in Figures 2a and 2c or to the combination of 164 + 163) for the next frame.H A harmonic post-filter (HPF) that follows the pitch contour is used to output y C The coded pulses that are composed of the coded pulse waveform are decoded, and the time domain signal y P is constructed. y P is a decoded audio signal (PCM O ) to generate y H Or, y P Yes C and the combination may be used as the input to the HPF, in which case the output of the HPF 214 is the decoded audio signal.

[0046] With reference to FIG. 5, the entity "Get Pitch Contour" 180 is described below. Next, the processing in the block "Get Pitch Contour 180" will be described. The input signal is downsampled from the full sampling rate to a lower sampling rate, for example 8 kHz. The pitch contour is determined by pitch_mid and pitch_end from the current frame and pitch_start equal to pitch_end from the previous frame. A frame is exemplarily shown by FIG. 5. All values ​​used in the pitch contour may be stored as pitch lags with decimal precision. The pitch lag values ​​are determined by a minimum pitch lag dFmin = 2.25 ms (corresponding to 444.4 Hz) and a maximum pitch lag dFmin = 51.3 Hz (corresponding to 51.3 Hz). Fmax = 19.5 ms and d Fmin From dF max The range up to is called the full pitch range. Other ranges of values ​​may also be used. The values ​​of pitch_mid and pitch_end are found in multiple steps. In each step, a pitch search is performed in the domain of the downsampled signal or in the domain of the input signal. The pitch search is performed by the normalized autocorrelation ρ of its input and a delayed version of the input. H [d F ]. Calculate the lag d F is the start of pitch search.FStart and pitch search ends d Fend The pitch search start time is between d and d. Fstart , pitch search completed d Fend , autocorrelation length l ρH , and past pitch candidates d Fpast are the parameters of the pitch search. The pitch search is performed to find the optimal pitch d as a pitch lag with fractional precision. Foptim and the harmonic level ρ obtained from the autocorrelation value at the optimal pitch lag. Hopim and returns. ρ Hoptim The range is 0 to 1, where 0 means no harmonics and 1 means maximum harmonics.

[0047] The location of the absolute maximum in the normalized autocorrelation is the first candidate for the optimal pitch lag. The file is TIFF2026042037000022.tif6150. d Fpast d F1 If it is close to , the second candidate for optimal pitch lag is d F2 d Fpast otherwise, d Fpast The location of the maximum value close to is the second candidate d F2 is. d Fpast d F1 If it is close to , d F1 d F2 is again chosen for , so that no local maxima are searched for. d F1 and d F2 The difference between the normalized autocorrelations in the pitch candidate threshold τ DF If it is greater than d Foptim d F1 (ρ H [d F1 ]-ρ H [d F2 ]>τ dF ⇒d Foptim =d F1 ), otherwise, d Foptim d F2 is set to τ dF is d F1 , dF2 , and d Fpast is chosen adaptively depending on the F1 ≦d Fpast ≦1.25 d F1 If τ dF =0.01, and not d F1 ≦d F2 If τ dF = 0.02, and d F1 >d F2 If τ dF =0.03 (for small pitch changes it is easier to switch to a new maximum position, and for large changes it is easier to switch to a small pitch lag than to a large pitch lag).

[0048] The location of the regions for the pitch search relative to the framing and windowing is shown in Figure 5. For each region, the pitch search is performed with an autocorrelation length l set to the length of the region. ρH First, in the pitch search, d Fpast =pitch_start, d Fstart =d Fmin , and d Fend =d Fmax The pitch lag start_pitch_ds and associated harmonics start_norm_corr_ds are calculated at the lower sampling rate using Fpast =start_pitch_ds, d Fstart =d Fmin , and d Fend =d Fmax The pitch lag avg_pitch_ds and associated harmonics avg_norm_corr_ds are calculated at the lower sampling rate using the d Fpast =avg_pitch_ds, d Fstart =0.3·avg_pitch_ds, and d Fend= 0.7·avg_pitch_ds, the pitch lags mid_pitch_ds and end_pitch_ds and the associated harmonics mid_norm_corr_ds and end_norm_corr_ds are calculated at a lower sampling rate. Fpast =pitch_ds, d Fstart =pitch_ds-Δ Fdown , and d Fend ==pitch_ds+Δ Fdown The pitch lags pitch_mid and pitch_end and the associated harmonics norm_corr_mid and norm_corr_end are calculated at the full sampling rate using Δ Fdown is the ratio of the full sampling rate to the lower sampling rate, and mid_pitch_ds for pitch_ds=pitch_mid and end_pitch_ds for pitch_ds=pitch_end.

[0049] If the average harmonic is below 0.3, or norm_corr_end is below 0.3, or norm_corr_mid is below 0.6, it is notified in a bitstream having a single bit that there is no pitch contour in the current frame. If the average harmonic is above 0.3, the pitch contour is encoded using absolute coding for pitch_end and differential coding for pitch_mid. Pitch_mid is differentially encoded to (pitch_start + pitch_end) / 2 using 3 bits by using a code for the difference from (pitch_start + pitch_end) / 2 among 8 predefined values, which minimizes the autocorrelation within the pitch_mid region. If there is an end of harmonic in the frame, for example, when norm_corr_end < norm_corr_mid / 2, linear extrapolation from pitch_start and pitch_mid is used for pitch_end, and as a result, pitch_mid may be encoded (e.g., norm_corr_mid > 0.6 and norm_corr_end < 0.3). |pitch_mid - pitch_start| ≤ τ HPFconst and |norm_corr_mid - norm_corr_start| ≤ 0.5, and if the expected HPF gain within the regions of pitch_start and pitch_mid is close to 1 and does not change significantly, it is notified in the bitstream that the HPF should use constant parameters.

[0050] According to an embodiment, the pitch contour provides the pitch lag value d Fmax at every sample i within the current window and at least d contour past samples to d contour . The pitch lag of the pitch contour is obtained by linear interpolation of pitch_mid and pitch_end from the current frame, the previous frame, and the second previous frame. Average pitch lag TIFF2026042037000023.tif6150 is calculated per frame as the average of pitch_start, pitch_mid, and pitch_end.

[0051] According to a further embodiment, half-pitch lag compensation is also possible. The LTP buffer 164, available in both the encoder and decoder, detects when the pitch lag of the input signal is d Fmin It is used to check whether the pitch lag of the input signal is below d Fmin The detection of a pitch lag below d is called "half pitch lag detection", and if it is detected, it is said to be "half pitch lag detected". The coded pitch lag values ​​(pitch_mid, pitch_end) are coded and Fmin ~d Fmax From these coded parameters, the pitch contour is derived as defined above. If a half pitch lag is detected, the coded pitch lag value is an integer multiple n of the true pitch lag value. Fcorrection (Equivalently, the input signal pitch is an integer multiple n of the coded pitch. Fcorrection ). To extend the pitch lag range beyond the codable range, the corrected pitch lag values ​​(pitch_mid_corrected, pitch_end_corrected) are used. If the true pitch lag value is within the codable range, the corrected pitch lag values ​​(pitch_mid_corrected, pitch_end_corrected) may be equal to the coded pitch lag values ​​(pitch_mid, pitch_end). Note that in the same way that the pitch contour is derived from the pitch lag values, the corrected pitch lag values ​​may be used to obtain the corrected pitch contour. In other words, this makes it possible to extend the frequency range of the pitch contour outside the frequency range for the coded pitch parameter, and a corrected pitch contour is generated.

[0052] Half-pitch detection assumes that the pitch is constant within the current window, TIFF2026042037000024.tif6150 is executed only if max(|pitch_mid-pitch_start|,|pitch_mid-pitch_end|)<τ Fconst For half-pitch detection, the pitch is considered constant within the current window if n Fmultiple ∈{1,2,…n FMaxcorrection} For each, the pitch search is TIFF2026042037000025.tif6150,d Fpast = TIFF2026042037000026.tif6150, d FStart =d Fpast -3, and d Fend =d Fpast It is performed using +3.

[0053] n Fcorrection is the n that maximizes the normalized correlation returned by the pitch search. Fmultiple It is set to n Fcorrection > 1 and n Fcorrection The normalized correlation returned by the pitch search for is greater than 0.8, and n Fmultiple A half-pitch is considered detected if it exceeds the normalized correlation returned by the pitch search when =1 by 0.02. If a half pitch lag is detected, pitch_mid_corrected and pitch_end_corrected are n Fmultiple =n Fcorrection otherwise pitch_mid_corrected and pitch_end_corrected are set to pitch_mid and pitch_end, respectively. Average corrected pitch lag TIFF2026042037000027.tif6150 is calculated as the average of pitch_start, pitch_mid_corrected, and pitch_end_corrected after correcting the final octave jump. The octave jump correction is calculated by finding the minimum value among pitch_start, pitch_mid_corrected, and pitch_end_corrected, and then calculating the minimum value among pitch_start, pitch_mid_corrected, and pitch_end_corrected for each pitch, (n Fmultiple ∈{1,2,…n FMaxcorrection}) the pitch / n closest to the minimum Fmultiple Then find pitch / n Fmultiple is used in place of the original value in calculating the average.

[0054] Pulse extraction may be described below in the context of Figure 6. Figure 6 shows pulse extractor 110 with entities 111hp, 112, 113c, 113p, 114, and 114m. The first entity at the input is optional high-pass filter 111hp, which outputs a signal to pulse extractor 112 (extracting pulse and statistical data). At the output, two entities 113c and 113p are arranged, which interact together and receive as input the pitch contour from entity 180. The entity for selecting pulses 113c outputs a pulse p directly to another entity 114 for generating a waveform. This is the waveform of the pulse, which can be subtracted from the PCM signal using mixer 114m to generate a residual signal R (the residue after extracting the pulse).

[0055] A maximum of eight pulses per frame are extracted and coded. In other cases, other numbers of maximum pulses may be used. TIFF2026042037000028.tif6150 pulses are retained and used for extraction and predictive coding ( TIFF2026042037000029.tif6150). Another example is Other constraints may be used for TIFF2026042037000030.tif6150. "Pitch Contour Acquisition" 180 is TIFF2026042037000031.tif6150, or TIFF2026042037000032.tif6150 may be used. For frames with low harmonics, TIFF2026042037000033.tif6150 is expected to be zero.

[0056] Time-frequency analysis via a short-time Fourier transform (STFT) is used to find and extract the pulses (see entity 112). In other examples, other time-frequency representations may be used. I The signals may be high-pass filtered (111hp), windowed using a 2 ms long raised sine window with 75% overlap, and transformed to the frequency domain (FD) via a discrete Fourier transform (DFT). Alternatively, high-pass filtering may be performed in the FD (at the output of 112s or 112s). Thus, in each 20 ms frame, there are 40 points per frequency band, each consisting of amplitude and phase. Each frequency band is 500 Hz wide, and the remaining 47 bands may be constructed via symmetric extension, thus resulting in a sampling rate of F S We consider only 49 bands for = 48 kHz. Therefore, there are 49 points at each time instance of the STFT, and 40 49 points in the time-frequency plane of the frame. The STFT hop size is H P =0.0005F S is.

[0057] 7 shows entity 112 in more detail. At 112te, the time envelope is obtained from the log-magnitude spectrogram by integration over the frequency axis, i.e., the log-magnitudes are summed for each time instance of the STFT to obtain one sample of the time envelope. The illustrated entity 112 is a PCM I The spectrogram entity 112s outputs a phase and / or amplitude spectrogram based on the signal. The phase spectrogram is forwarded to the pulse extractor 112pe, and the amplitude spectrogram is further processed. The amplitude spectrogram may be processed using a background remover 112br, and a background estimator 112be is used to estimate the background signal to be removed. Additionally or alternatively, a time envelope determiner 112te and a pulse locator 112pl process the amplitude spectrogram. The entities 112pl and 112te allow for determining pulse positions to be used as inputs to the pulse extractor 112pe and the background estimator 112be. The pulse locator finder 112pl can use pitch contour information. Optionally, some entities, such as the entities 112be and 112te, can use an algorithmic representation of the amplitude spectrogram obtained by the entity 112lo.

[0058] The function is explained below. The smoothed time envelope is filtered using a short symmetric FIR filter (e.g., F S = 48 kHz). The normalized autocorrelation of the time envelope is calculated: TIFF2026042037000034.tif15150 TIFF2026042037000035.tif13150 where, TIFF2026042037000036.tif6150 is the time envelope after removing the mean. The exact delay ( TIFF2026042037000037.tif6150) is estimated using a three-point Lagrange polynomial that forms a peak in the normalized autocorrelation.

[0059] The expected mean pulse distance may be estimated from the normalized autocorrelation of the time envelope and the average pitch lag within the frame. TIFF2026042037000038.tif23150Here, for frames with low harmonics, TIFF2026042037000039.tif6150 is set to 13, which corresponds to 6.5 milliseconds. The pulse locations are local peaks in the smoothed time envelope with the requirement that the peaks be above their perimeter. The perimeter is defined as a low-pass filtered version of the time envelope using a simple moving average filter with adaptive length, where the length of the filter is determined by the expected average pulse distance ( TIFF2026042037000040.tif6150). The exact pulse position ( TIFF2026042037000041.tif6150) is estimated using a three-point Lagrange polynomial that forms a peak in the smoothed time envelope. TIFF2026042037000042.tif6150) are exact positions rounded to STFT time instances, so the distance between pulse center positions is a multiple of 0.5 milliseconds. Each pulse can be thought of as extending two time instances to the left and two time instances to the right from its center position. Other numbers of time instances may also be used. A maximum of 8 pulses per 20 milliseconds are found, and if more pulses are detected, the smaller pulses are ignored. The number of pulses found is It is written as TIFF2026042037000043.tif6150.

[0060] The i-th pulse is It is written as TIFF2026042037000044.tif6150. The average pulse distance is It is defined as follows. TIFF2026042037000045.tif18150 Amplitudes are enhanced based on pulse position so that the enhanced STFT, also known as the enhanced spectrogram, consists of only the pulse. The pulse background is estimated as a linear interpolation of the left and right backgrounds, which are the average of the third through fifth time instances from the center position. The background is estimated in the logarithmic amplitude domain at 112be and removed by subtracting it in the linear amplitude domain at 112br. Amplitudes in the enhanced STFT are linearly scaled. The phase is not corrected. All amplitudes at time instances not belonging to a pulse are set to zero. The starting frequency of the pulses is proportional to the inverse of the average pulse distance (between nearby pulse waveforms) within a frame, but is limited to between 750 Hz and 7250 Hz. TIFF2026042037000046.tif11150

[0061] Start frequency ( TIFF2026042037000047.tif6150) is expressed as an index of the STFT band. The change in start frequency in successive pulses is limited to 500 Hz (one STFT band). The amplitude of the enhanced STFT below the start frequency is set to zero at 112 pe. The waveform of each pulse is obtained from the enhanced STFT at 112 pe. The pulse waveform is non-zero for 4 ms around its center, and the pulse length is TIFF2026042037000048.tif6150 (The sampling rate of the pulse waveform is the sampling rate of the input signal F S (Equal to ). TIFF2026042037000049.tif6150 represents the waveform of the ith pulse. Each pulse Pi is located at the center position TIFF2026042037000050.tif6150 and pulse waveform The pulse extractor 112pe is uniquely determined by the center position TIFF2026042037000052.tif6150 and pulse waveform TIFF2026042037000053.tif6150. The pulses are aligned to the STFT grid. Alternatively, the pulses may not be aligned to the STFT grid and / or the exact pulse position ( TIFF2026042037000054.tif6150) The pulse may be determined instead of TIFF2026042037000055.tif6150.

[0062] The features are calculated for each pulse: Percentage of local energy in the pulse TIFF2026042037000056.tif6150 Percentage of frame energy in the pulse TIFF2026042037000057.tif6150·Percentage of the band where the pulse energy exceeds half of the local energy- TIFF2026042037000058.tif6150 Between each pulse pair (pulse in the current frame and pulse from the previous frame) TIFF2026042037000059.tif6) Correlation between the 150 last coded pulses TIFF2026042037000060.tif6150 and distance TIFF2026042037000061.tif6150 Pitch lag at the exact position of the pulse TIFF2026042037000062.tif6150 Local energy is calculated from 11 time instances around the pulse center in the original STFT. All energy is calculated only above the starting frequency. Distance between pulse pairs TIFF2026042037000063.tif6150 is obtained from the position of maximum cross-correlation between pulses TIFF2026042037000064.tif7150. The cross-correlation is windowed with a rectangular window 2 ms long and normalized by the norm of the pulse (also windowed with a rectangular window of 2 ms). The pulse correlation is the maximum of the normalized cross-correlation. TIFF2026042037000065.tif18150 TIFF2026042037000066.tif20150 TIFF2026042037000067.tif22150 TIFF2026042037000068.tif8150 TIFF2026042037000069.tif9150 The value of TIFF2026042037000070.tif7150 is in the range of 0 to 1.

[0063] The error between pitch and pulse distance is calculated as follows: TIFF2026042037000071.tif15150Multiple of pulse distance( TIFF2026042037000072.tif6150) takes into account the pitch estimation error. TIFF2026042037000073.tif6150) solves missing pulses resulting from imperfections in the pulse train when the pulses in the train are distorted or when there are transient signals that do not belong to the pulse train that prevent the detection of pulses that belong to the train.

[0064] Probability that the i-th and j-th pulses belong to the pulse train (see p. 113): TIFF2026042037000074.tif51145The probability of a pulse having a relationship only with previously coded pulses is defined as follows: TIFF2026042037000075.tif11150 Pulse probability ( TIFF2026042037000076.tif6150) is found repeatedly (see entity 113p): 1. All pulse probabilities ( TIFF2026042037000077.tif6150) is set to 1

[0065] 2. In the time sequence of pulses, it is still possible ( TIFF2026042037000078.tif6150) For each pulse: a. The probability of a pulse belonging to a pulse train in the current frame is calculated: TIFF2026042037000079.tif16150 b. The initial probability that it is truly a pulse is: TIFF2026042037000080.tif6150c. The probability increases for pulses with energy in many bands greater than half the local energy. TIFF2026042037000081.tif6150d. The probability is limited by the time envelope correlation and the proportion of local energy within the pulse. TIFF2026042037000082.tif6150e.If the pulse probability is below the threshold, the probability is set to zero and is not considered further. TIFF2026042037000083.tif11150 3. Set to at least one zero in the current iteration As long as TIFF2026042037000084.tif6150 exists, or all Step 2 is repeated until TIFF2026042037000085.tif6150 is set to zero. At the end of this procedure, a value equal to 1 will be generated. TIFF2026042037000086.tif6150 TIFF2026042037000087.tif6There are 150 true pulses. All true pulses and only true pulses constitute the pulse part P, coded as CP. TIFF2026042037000088.tif6 Of the 150 pulses, the last three pulses are TIFF2026042037000089.tif6150 and TIFF2026042037000090.tif6150 are stored in memory to calculate. If there are fewer than three true pulses in the current frame, some pulses already in memory are kept. A maximum of three pulses in total are kept in memory. There may be other limits on the number of pulses kept in memory, for example, two or four. After three pulses are in memory, the memory remains full and the oldest pulse in memory is replaced by the newly found pulse. In other words, the number of past pulses kept in memory TIFF2026042037000091.tif6150 is TIFF2026042037000092.tif increases to 6150 and then remains at 3.

[0066] Pulse coding (encoder side, see entity 132) will now be described with respect to FIG. FIG. 8 illustrates a pulse coder 132 comprising entities 132fs, 132c, and 132pc in the main path, where entity 132as is arranged to determine and provide a spectral envelope as input to entity 132fs, which is configured to perform spectral flattening. Within main paths 132fs, 132c, and 132pc, pulse P is coded to determine a coded spectrally flattened pulse. The coding performed by entity 132pc is performed on the spectrally flattened pulse. The coded pulse CP of FIGS. 2a-2c is composed of a coded spectrally flattened pulse and a pulse spectral envelope. Coding of multiple pulses is described in detail with reference to FIG. 10. The pulses are coded using the following parameters: Number of pulses in a frame TIFF2026042037000093.tif6150 Position within the frame TIFF2026042037000094.tif6150 Pulse start frequency TIFF2026042037000095.tif6150 Pulse spectrum envelope Predicted Gain TIFF2026042037000096.tif6150, and If TIFF2026042037000097.tif6150 is not zero o Index of the forecast source TIFF2026042037000098.tif6150o predicted offset TIFF2026042037000099.tif6150 Innovation Gain TIFF2026042037000100.tif6150 · An innovation consisting of up to four impulses, each coded by its position and sign

[0067] A single coded pulse is determined by the following parameters: Pulse start frequency TIFF2026042037000101.tif6150 Pulse Spectral Envelope Predicted Gain TIFF2026042037000102.tif6150, and If TIFF2026042037000103.tif6150 is not zero ○ Forecast source index TIFF2026042037000104.tif6150 ○ Predicted offset TIFF2026042037000105.tif6150 Innovation Gain TIFF2026042037000106.tif6150 · An innovation consisting of up to four impulses, each coded by its position and sign A waveform presenting a single coded pulse can be constructed from the parameters that determine the single coded pulse, and the coded pulse waveform can then be said to be determined by the parameters of the single coded pulse. The pulse numbers are Huffman coded. First pulse position TIFF2026042037000107.tif6150 is coded absolutely using Huffman coding. For subsequent pulses, the position delta TIFF2026042037000108.tif6150 is Huffman coded. There are different Huffman codes depending on the number of pulses in the frame and the position of the first pulse.

[0068] First Pulse Start Frequency TIFF2026042037000109.tif6150 is coded absolutely using Huffman coding. The starting frequency of subsequent pulses is differentially coded. If a difference of zero exists, all subsequent differences are also zero, so the number of non-zero differences is coded. Because all differences have the same sign, the sign of the difference can be coded with one bit per frame. In most cases, the absolute difference is at most 1, so if the maximum absolute difference is 1 or greater, a single bit is used for coding. Finally, all non-zero absolute differences need to be coded only if the maximum absolute difference is 1 or greater, and they are unary coded. Spectral flattening, performed for example using the STFT (see entity 132 fs in Figure 8), is illustrated by Figures 9a and 9b, where Figure 9a shows the original pulse waveform compared to the flattened version in Figure 9b. Note that spectral flattening may alternatively be performed, for example by a filter in the time domain.

[0069] All pulses within a frame may use the same spectral envelope (see entity 132as) consisting of eight bands. The band boundary frequencies are 1 kHz, 1.5 kHz, 2.5 kHz, 3.5 kHz, 4.5 kHz, 6 kHz, 8.5 kHz, 11.5 kHz, and 16 kHz. Spectral content above 16 kHz is not explicitly coded. In other examples, other band boundaries may be used. The spectral envelope at each time instance of a pulse is obtained by summing the amplitudes within the envelope bands, where the pulse consists of five time instances. The envelope is averaged over all pulses in the frame. Points between pulses in the time-frequency plane are not considered. The values ​​are compressed using the fourth root and the envelope is vector quantized. The vector quantizer has two stages, with the second stage divided into two halves. TIFF2026042037000110.tif6150 and For a frame with TIFF2026042037000111.tif6150, and TIFF2026042037000112.tif6150 and There are different codebooks for the value TIFF2026042037000113.tif6150. Different codebooks require different numbers of bits.

[0070] The quantized envelope may be smoothed using linear interpolation. The spectrogram of the pulse is flattened using the smoothed envelope (see entity 132fs). Flattening is achieved by dividing the amplitude by the envelope (received from entity 132as), which is equivalent to subtraction in the log-magnitude domain. The phase value is not altered. Alternatively, the filter processor may be configured to spectrally flatten the amplitude or pulse STFT by filtering the pulse waveform in the time domain. Spectrally flattened pulse waveform TIFF2026042037000114.tif6150 is obtained from the STFT via inverse DFT, windowing, and overlap add at 132c.

[0071] 10 shows an entity 132pc for encoding a single spectrally flattened pulse waveform of a plurality of spectrally flattened pulse waveforms. Each single encoded pulse waveform is output as a coded pulse signal. From another perspective, the entity 132pc for encoding a single pulse in FIG. 10 is the same as the entity 132pc configured to encode a pulse waveform as shown in FIG. 8, but is used several times to encode several pulse waveforms. Entity 132pc in FIG. 10 comprises a pulse coder 132spc, a constructor for flattened pulse waveforms 132cpw, and a memory 132m, arranged as a kind of feedback loop. The constructor 132cpw has the same function as 220cpw, and the memory 132m has the same function as 229 in FIG. 14. Each single / current pulse is coded by entity 132spc based on a flattened pulse waveform taking into account past pulses. Information about past pulses is provided by memory 132m. Note that the past pulses coded by 132pc are fed via pulse waveform constructor 132cpw and memory 132m. This allows prediction. The result of using such a prediction technique is illustrated in FIG. 11, where FIG. 11a shows the flattened original pulse together with the prediction, resulting in the prediction residual signal in FIG. 11b.

[0072] According to an embodiment, Among the 6150 pulses and the already quantized pulses from the current frame, the most similar previously quantized pulse is found. TIFF2026042037000116.tif6150 is used to select the most similar pulse. If the correlation difference is below 0.05, the closer pulse is selected. The most similar previous pulse is the source of the prediction. TIFF2026042037000117.tif6150 and its index for the currently coded pulse TIFF2026042037000118.tif6150 is used in pulse coding. Up to four relative prediction source indices. TIFF2026042037000119.tif6150 is grouped and Huffman coded. The grouping and Huffman coding are TIFF2026042037000120.tif6150 and TIFF2026042037000121.tif6150 mosquito Depends on TIFF2026042037000122.tif6150.

[0073] The offset for maximum correlation is the pulse prediction offset TIFF2026042037000123.tif6150. It is coded absolutely, differentially, or relative to the estimate, which is the exact position of the pulse. It is calculated from the pitch lag in TIFF2026042037000124.tif6150. The number of bits required for each coding type is calculated and the one with the least bits is selected. prediction For scaling TIFF2026042037000125.tif6150, use the gain to maximize the SNR. TIFF2026042037000126.tif6150 is used. The prediction gain is non-uniformly quantized with 3-4 bits. If the energy of the prediction residual is not at least 5% less than the energy of the pulse, prediction is not used. TIFF2026042037000127.tif6150 is set to zero.

[0074] The prediction residual is quantized using a maximum of four impulses. In other cases, other maximum numbers of impulses may be used. The quantized residual, which consists of impulses, is the innovation. This is called TIFF2026042037000128.tif6150. This is shown in Figure 12. To save bits, the number of impulses is reduced by one for each pulse predicted from a pulse in this frame. In other words, if the prediction gain is zero or if the source of the prediction is a pulse from the previous frame, four impulses are quantized; otherwise, the number of impulses is reduced compared to the predicted source. Figure 12 shows a processing path used as process block 132spc in Figure 10. The processing path makes it possible to determine the coded pulse and may include three entities 132bp, 132qi, 132ce. The first entity 132bp for finding the best prediction uses past pulses and pulse waveforms to determine iSOURCE, shift, GP', and the prediction residual. The quantization impulse entity 132gi quantizes the prediction residual and outputs GI' and impulse. The entity 132ce is configured to calculate and apply correction coefficients. All this information, together with the pulse shape, is received by the entity for correcting the energy 132ce to output a coded impulse. According to an embodiment, the following algorithm may be used:

[0075] The following algorithm is used for impulse detection and coding: 1. Absolute pulse waveform using full-wave rectification TIFF2026042037000129.tif6150 is constructed. TIFF2026042037000130.tif61502. A vector containing the number of impulses at each position TIFF2026042037000131.tif6150 is initialized with zeros. TIFF2026042037000132.tif61503. The position of the maximum value of TIFF2026042037000133.tif6150 is found. TIFF2026042037000134.tif91504. A vector containing the number of impulses and the position of the maximum value found TIFF2026042037000135.tif6150 increases by one. TIFF2026042037000136.tif6150

[0076] 5. The maximum value of TIFF2026042037000137.tif6150 is reduced. TIFF2026042037000138.tif121506. Steps 3-5 are repeated until the required number of impulses is found, where the number of pulses is Equals TIFF2026042037000139.tif8150. Note that impulses may have the same location. Pulse locations are ordered by their distance from the pulse center. The location of the first impulse is coded absolutely. The locations of subsequent impulses are coded differentially, with probabilities depending on the location of the previous impulse. Huffman coding is used for impulse locations. The code of each impulse is also coded. If multiple impulses share the same location, the code is coded only once.

[0077] The four detected and scaled impulses 15i of the residual signal 15r are shown in Figure 13. In particular, the impulses represented by the lines TIFF2026042037000140.tif8150 will change the gain accordingly, e.g. impulse + / - 1 It may be scaled by multiplying TIFF2026042037000141.tif6150. Gain that maximizes SNR TIFF2026042037000142.tif6150 is an innovation made up of impulses Used to scale TIFF2026042037000143.tif6150. The innovation gain is the number of pulses TIFF2026042037000144.tif6150, non-uniformly quantized between 2 and 4 bits depending on the image.

[0078] Then, a first estimate for the quantization of the flattened pulse waveform is TIFF2026042037000145.tif6150 is TIFF2026042037000146.tif8150, where TIFF2026042037000147.tif6150 indicates quantization. The gain is found by maximizing the SNR, so The energy of TIFF2026042037000148.tif6150 is the original target To compensate for the energy reduction, a correction factor c g is calculated. TIFF2026042037000150.tif16150The final gain is: TIFF2026042037000151.tif15150 TIFF2026042037000152.tif6150

[0079] The memory for prediction is the quantized flattened pulse waveform. Updated with TIFF2026042037000153.tif6150. TIFF2026042037000154.tif8150 TIFF2026042037000155.tif6150 At the end of coding the quantized flattened pulse waveforms are kept in memory for prediction in subsequent frames.

[0080] A technique for recovering the pulse will now be described with reference to FIG. FIG. 14 shows entity 220 for reconstructing a single pulse waveform. The technique described below for reconstructing a single pulse waveform is performed multiple times for multiple pulse waveforms. Multiple pulse waveforms are used by entity 22′ of FIG. 15 to reconstruct a waveform containing multiple pulses. From another perspective, entity 220 processes a signal composed of multiple coded pulses and multiple pulse spectral envelopes, and for each coded pulse and associated pulse spectral envelope, outputs a single reconstructed pulse waveform, resulting in a signal composed of multiple reconstructed pulse waveforms at the output of entity 220. Entity 220 includes several sub-entities, such as entity 220cpw for constructing a spectrally flattened pulse waveform, entity 224 for generating a pulse spectrogram (phase and amplitude spectrogram) of the spectrally flattened pulse waveform, and entity 226 for spectrally shaping the pulse amplitude spectrogram. This entity 226 uses the amplitude spectrogram and the pulse spectral envelope. The output of entity 226 is supplied to a converter for converting the pulse spectrogram into a waveform marked by reference numeral 228. This entity 228 receives the phase spectrogram and the spectrally shaped pulse amplitude spectrogram to reconstruct the pulse waveform. Note that entity 220cpw (configured to construct a spectrally flattened pulse waveform) receives a signal describing the coded pulse at its input. Constructor 220cpw includes a kind of feedback loop including update memory 229. This allows the pulse waveform to be constructed taking into account past pulses. Here, previously constructed pulse waveforms are fed back so that past pulses can be used by entity 220cpw to construct the next pulse waveform. The function of this pulse restorer 220 will be described below. It should be noted that only the quantized flattened pulse waveform (also called decoded flattened pulse waveform or coded flattened pulse waveform) exists on the decoder side, and since the original pulse waveform does not exist on the decoder side, we use "flattened pulse waveform" to name the quantized flattened pulse waveform on the decoder side, and "pulse waveform" to name the quantized pulse waveform (also called decoded pulse waveform or coded pulse waveform or decoded pulse waveform).

[0081] To restore the pulses at the decoder side 220, the gain ( TIFF2026042037000156.tif6150 and TIFF2026042037000157.tif6150), Impulse / Innovation, Forecast Source ( TIFF2026042037000158.tif6150), and offset ( After decoding the TIFF2026042037000159.tif6150, a quantized flattened pulse waveform is constructed (see entity 220cpw). The prediction memory 229 is updated similarly to the encoder in entity 132m. An STFT (see entity 224) is then taken for each pulse waveform. For example, the same 2 ms long raised sine window with 75% overlap is used as in the pulse extraction. The amplitude of the STFT is reshaped using the decoded and smoothed spectral envelope to obtain the pulse start frequency. The STFT is zeroed below TIFF2026042037000160.tif6150. A simple multiplication of the amplitude with the envelope is used to shape the STFT (see entity 226). The phase is not modified. The reconstructed waveform of the pulse is obtained from the STFT via an inverse DFT, windowing, and overlap-add (see entity 228). Alternatively, the envelope can be shaped via an FIR filter, avoiding the STFT.

[0082] Figure 15 shows the waveform y P 2a, 2c, which is used as the last entity in the waveform builder 22 of 2a or 2c. The restored pulse waveform is TIFF2026042037000161.tif6150 and insert zeros between the pulses of entity 22' in Figure 15. The concatenated waveform is added to the decoded signal (see 23 in Figure 2a or 2c or 114m in Figure 6). Similarly, the original pulse waveform TIFF2026042037000162.tif6150 are concatenated (see 114 in Figure 6) and subtracted from the input of the MDCT-based codec (see Figure 6). The restored pulse waveform is The waveforms are concatenated based on TIFF2026042037000163.tif6150, and zeros are inserted between the pulses. The concatenated waveform is added to the decoded signal. Similarly, the original pulse waveform TIFF2026042037000164.tif6150 are concatenated and subtracted from the input of the MDCT-based codec. The reconstructed pulse waveform is not a perfect representation of the original pulse. Therefore, removing the reconstructed pulse waveform from the input leaves some of the transient portion of the signal. Because transient signals cannot be well represented by the MDCT codec, there is noise that spreads across the entire frame, reducing the benefit of coding the pulses separately. For this reason, the original pulse is removed from the input.

[0083] According to an embodiment, the HF tonality flag φ H may be defined as follows: Normalized correlation ρ HF is the sample in the current window and TIFF2026042037000165.tif6150 (or TIFF2026042037000166.tif6150) with delayed version MHF is calculated for, where y MHF is the pulse residual signal y M As an example, a high-pass filter with a crossover frequency of about 6 kHz may be used. For each MDCT frequency bin above a specified frequency, it is determined whether the frequency bin is tonal or noise-like, as per 5.3.3.2.5 of

[20] . The total number of tonal frequency bins, n HFTonalCurr is calculated in the current frame, and the total number of smoothed tonal frequencies is n HFTonal =0.5 n HFTonal +n HFTonalCurr It is calculated as: HF tonality flag φ His set to 1 when TNS is inactive, pitch contours are present, and tonality is present in high frequencies, and ρ HF >0 or n HFTonal If >1, it is present at high frequencies.

[0084] The iBPC method is explained with reference to Figure 16. Next, the optimal quantization step size The process of obtaining TIFF2026042037000167.tif6150 is described. The process can be an integral part of the block iBPC. The entity 300 in FIG. 16 is X MR Based on Note that the output is TIFF2026042037000168.tif6150. On another device, you might want to use X MR and TIFF2026042037000169.tif6150 may be used (see Figure 3 for details).

[0085] 16 shows a flowchart of a technique for estimating the step size. The process starts at i=0, and then four steps are performed: quantization, adaptive band zeroing, jointly determining per-band parameters and spectrum, and determining whether the spectrum is codable. These steps are marked by reference numerals 301 to 304. If the spectrum is codable, the step size is decreased (see step 307) and the next iteration ++i is performed, see reference numeral 308. This is performed as long as i is not equal to the maximum iteration (see decision step 309). If the maximum iteration is reached, the step size is output. If the maximum iteration is not reached, the next iteration is performed.

[0086] If the spectrum is not codable, a process is applied comprising steps 311 and 312 together with a verification step (spectrum is now codable) 313. Then the step size is increased (see 314) before starting the next iteration (see step 308). Spectrum X, whose spectral envelope is perceptually flattened MR is a single quantization step size g across the entire coded bandwidth. Q The coded spectral bandwidth is then scalar quantized using a scalar quantizer, e.g., entropy coded with a context-based arithmetic coder to generate the coded spectrum. It is divided into sub-bands Bi of TIFF2026042037000170.tif6150. Optimal quantization step size, also known as global gain TIFF2026042037000171.tif6150 is found repeatedly as described.

[0087] At each iteration, the spectrum X MR teeth The image is quantized in block quantization 301 to generate TIFF2026042037000172.tif6150. In block "adaptive band zeroing" 302, the ratio of the energy of the zero-quantized line to the original energy is calculated for the subband Bi The energy ratio is calculated at the adaptive threshold If it exceeds TIFF2026042037000173.tif6150, Entire subbands in TIFF2026042037000174.tif6150 are set to zero. Threshold TIFF2026042037000175.tif6150 is the tonality flag φ H and flags Calculated based on TIFF2026042037000176.tif6150, flag TIFF2026042037000177.tif6150 indicates whether the subband was zeroed in the previous frame. TIFF2026042037000178.tif12150

[0088] For each zeroed subband, a flag TIFF2026042037000179.tif6150 is set to 1. At the end of processing, the current frame TIFF2026042037000180.tif6150 Alternatively, there may be more than one tonality flag and a mapping from multiple tonality flags to tonality for each subband, with a tonality value for each subband TIFF2026042037000182.tif6150 is generated. The values ​​in TIFF2026042037000183.tif6150 are, for example, the set of values ​​{0.25, 0.5, 0.75} Alternatively, the energy of the zero quantization line and the original energy and content TIFF2026042037000184.tif6150 and X BR Based on Other decisions may be used to determine whether to zero out the entire subband i of TIFF2026042037000185.tif6150.

[0089] The frequency range in which adaptive band nulling is used is the frequency range at which the ABZStart , for example, limited above 7000 Hz, so long as the lowest subband is zeroed, adaptive band nulling can be performed at a specific frequency f ABZMin , for example, extended downwards to 700Hz. f is completely zero EZ exceed Individual zero-filling levels (individual zfl) of the subbands of TIFF2026042037000186.tif6150, where f EZ is explicitly coded, and f is quantized to zero. EZ All zero subbands below f EZ One zero-filling level (zfl small) is coded due to quantization in block quantization, even if it is not explicitly set to 0 by adaptive band zeroing. Subbands in TIFF2026042037000187.tif6150 can be entirely zero. The zero-filling level (individual zfl and zfl small ZFL) and The number of bits required for entropy coding of the spectral lines of TIFF2026042037000188.tif6150 is calculated (e.g., by a band-wise parametric coder). Furthermore, the number of spectral lines N that can be explicitly coded with the available bit budget is calculated. Q is found.

[0090] N Q is an integral part of the coded spect and is used in the decoder to find how many bits are used to code the spectral lines. Other methods for finding the number of bits to code the spectral lines may be used, such as using a special EOF character. N Q exceed The lines of TIFF2026042037000189.tif6150 are set to zero and the number of bits required is recalculated. For the calculation of the bits required to code a spectral line, the bits required to code the line starting from the bottom are calculated. The recalculation of the bits required to code a spectral line is carried out for each n ≤ N Q This calculation is only needed once, as efficiency is gained by storing the number of bits required to code n lines for . At each iteration, if the number of required bits exceeds the available bits, the global gain g Q decreases (307), otherwise g Qis increased (314). At each iteration, the rate of change of the global gain is adapted. The same rate of change adaptation as the rate-distortion loop from EVS

[20] may be used to iteratively modify the global gain. At the end of the iteration process, the optimal quantization step size TIFF2026042037000190.tif6150, for example, generates optimal coding of spectra using criteria from EVS. Q is equal to X Q corresponds to Equals TIFF2026042037000191.tif6150. Instead of the actual coding, an estimate of the maximum number of bits required for coding may be used. The output of the iterative process is the optimal quantization step size TIFF2026042037000192.tif6150, and the output may also include the coded spect and coded noise filling level (zfl), as they are usually already available, to avoid iterative processing in retrieving them again.

[0091] Zero filling is explained in detail below. According to an embodiment, starting from an example of how to select a source spectrum, the block "zero filling" is then described. To create the zero filling, the following parameters are adaptively found: Optimal long copy-up distance TIFF2026042037000193.tif6150 Minimum Copy-Up Distance TIFF2026042037000194.tif6150 · Minimum copy-up source start TIFF2026042037000195.tif6150·Copy-up distance shift Δ C Source spectrum is X CT If the already obtained bottom part of TIFF2026042037000196.tif6150 determines the optimal distance. For example, the value of TIFF2026042037000197.tif6150 is the minimum value set in the index corresponding to 5600Hz. TIFF2026042037000198.tif6150 and the maximum value set in the index corresponding to 6225Hz, for example TIFF2026042037000199.tif6150. Other values ​​are constrained. May be used with TIFF2026042037000200.tif6150.

[0092] Distance between harmonics TIFF2026042037000201.tif6150 is the average pitch lag Average pitch lag calculated from tif2026042037000202.tif6150 TIFF2026042037000203.tif6150 is either decoded from the bitstream or inferred from parameters (e.g., pitch contour) from the bitstream, or TIFF2026042037000204.tif6150 is X DT or (e.g., X DT The distance between harmonics may be obtained by analyzing its derivative (from the time domain signal obtained using TIFF2026042037000205.tif6150 is not necessarily an integer. If the file is TIFF2026042037000206.tif6150 TIFF2026042037000207.tif6150 is set to zero, which is a way of signaling that there is no significant pitch lag. The value of TIFF2026042037000208.tif6150 is the minimum optimal copy-up distance Harmonic distance greater than TIFF2026042037000209.tif6150 This is the smallest multiple of TIFF2026042037000210.tif6150. If TIFF2026042037000211.tif3350 is zero, TIFF2026042037000212.tif6150 is not used.

[0093] The starting TNS spectral line plus TNS order is i T which may be an index corresponding to, for example, 1000 Hz. If TNS is inactive in the frame, TIFF2026042037000213.tif6150 TIFF2026042037000214.tif7150. If TNS is active, TIFF2026042037000215.tif6150 is i T , and if the HF is tonal (e.g., Φ H is 1), Lower bounded by TIFF2026042037000216.tif7150.

[0094] Amplitude spectrum Z C is the decrypted spect, X DT It is estimated from TIFF2026042037000217.tif18150The normalized correlation of the estimated amplitude spectrum is calculated. TIFF2026042037000218.tif22150 Correlation length L C is set to the maximum allowed by the available spectrum, optionally limited to some value (e.g., a length equivalent to 5000 Hz).

[0095] Basically, copy-up source TIFF2026042037000219.tif6150 and destination TIFF2026042037000220.tif6150, where 0 ≤ m <L C is

[0096] n ( TIFF2026042037000221.tif6150) Select TIFF2026042037000222.tif6150, where ρ C has the first peak, and ρ C above the average of, i.e., TIFF2026042037000223.tif7150 and TIFF2026042037000224.tif11150 and all About TIFF2026042037000225.tif6150 TIFF2026042037000226.tif6150 is not satisfied. TIFF2026042037000227.tif6150~ TIFF2026042037000228.tif6150 is the absolute maximum value in the range TIFF2026042037000229.tif6150 can be selected. Where a long copy-up distance is expected, TIFF2026042037000230.tif6150~ Any other value in the range TIFF2026042037000231.tif6150 It may be selected as TIFF2026042037000232.tif6150.

[0097] If TNS is active, You can select TIFF2026042037000233.tif6150. If TNS is inactive, TIFF2026042037000234.tif7150, where TIFF2026042037000235.tif6150 is the normalized correlation, TIFF2026042037000236.tif6150 is the best fit distance in the previous frame. TIFF2026042037000237.tif6150 indicates whether there was a tonality change in the previous frame. TIFF2026042037000238.tif6150 is TIFF2026042037000239.tif6150, TIFF2026042037000240.tif6150, or Returns either TIFF2026042037000241.tif6150. The decision on which value to return in TIFF2026042037000242.tif6150 is mainly based on the value TIFF2026042037000243.tif7150, TIFF2026042037000244.tif7150, and Based on TIFF2026042037000245.tif6150. Flags TIFF2026042037000246.tif6150 is true, TIFF2026042037000247.tif7150 or If TIFF2026042037000248.tif7150 is valid, TIFF2026042037000249.tif6150 will be ignored. TIFF2026042037000250.tif6150 and The value TIFF2026042037000251.tif6150 is used in rare cases.

[0098] In one example, TIFF2026042037000252.tif6150 may be defined by the following decisions: · TIFF2026042037000253.tif7150 is at least For TIFF2026042037000254.tif6150 TIFF2026042037000255.tif7150 and at least For TIFF2026042037000256.tif6150 If the file size is larger than TIFF2026042037000257.tif6150 TIFF2026042037000258.tif6150 is returned, where: TIFF2026042037000259.tif6150 and TIFF2026042037000260.tif6150 is, TIFF2026042037000261.tif7150 and TIFF2026042037000262.tif7150 is an adaptive threshold proportional to the TIFF2026042037000263.tif7150 may be required to exceed some absolute threshold, e.g. 0.5, If not, TIFF2026042037000264.tif7150 is at least the threshold, e.g., 0.2 If the file size is larger than TIFF2026042037000265.tif6150 TIFF2026042037000266.tif6150 is returned, If not, TIFF2026042037000267.tif6150 is set, If it is TIFF2026042037000268.tif7150 TIFF2026042037000269.tif6150 is returned, If not, TIFF2026042037000270.tif6150 is set, If the value of TIFF2026042037000271.tif6150 is valid, i.e., if there is significant pitch lag, TIFF2026042037000272.tif6150 is returned, If not, TIFF2026042037000273.tif6150 is small, for example, below 0.1, The value of TIFF2026042037000274.tif6150 is valid, i.e., there is significant pitch lag and the pitch lag change from the previous frame is small. TIFF2026042037000275.tif6150 is returned, If not, TIFF2026042037000276.tif6150 is returned.

[0099] Flags TIFF2026042037000277.tif6150 is displayed when TNS is active or TIFF2026042037000278.tif6150 and set to true if the tonality is low. H is false, or Low if TIFF2026042037000279.tif6150 is 0. TIFF2026042037000280.tif6150 is a value smaller than 1, for example 0.7. The value set in TIFF2026042037000281.tif6150 will be used in the next frame. Between the previous frame and the current frame Percent change of TIFF2026042037000282.tif6150 TIFF2026042037000283.tif6150 is also calculated. Optimal copy-up distance TIFF2026042037000284.tif6150 TIFF2026042037000285.tif6150 is not equal to Unless it's TIFF2026042037000286.tif6150 ( TIFF2026042037000287.tif6150 is a predetermined threshold), copy-up distance shift Δ C teeth TIFF2026042037000288.tif6150, in which case, Δ C is set to the same value as the previous frame and remains constant across successive frames. TIFF2026042037000289.tif6150 is the time between the previous frame and the current frame TIFF2026042037000290.tif6150 is a measure of change (e.g., percent change). TIFF2026042037000291.tif6150 If the percent change is TIFF2026042037000292.tif6150, TIFF2026042037000293.tif6150 may be set to 0.1 for example. If TNS is active in the frame, Δ C is not used.

[0100] Minimum copy-up source start TIFF2026042037000294.tif6150 can be set to iT, for example, if TNS is active, and optionally, if HF is tonal. If bounded by TIFF2026042037000295.tif7150 or if TNS is not active in the current frame, e.g. It can be set to TIFF2026042037000296.tif6150. Minimum Copy-Up Distance TIFF2026042037000297.tif6150, for example, if TNS is inactive TIFF2026042037000298.tif6150. If TNS is active, TIFF2026042037000299.tif6150, for example, if HF is not tonal TIFF2026042037000300.tif6150, or TIFF2026042037000301.tif6150 is, for example, when HF is tonal It is set to TIFF2026042037000302.tif15150.

[0101] For example, the initial condition is Random noise spectrum X using TIFF2026042037000303.tif11150 N teeth TIFF2026042037000304.tif6150, where the short function truncates the result to 16 bits. Any other random noise generator and initial conditions may be used. Then, the random noise spectrum X N is X D is set to zero at non-zero values ​​of , and optionally, N The part is X D is windowed to reduce random noise near the locations of non-zero values ​​of X CT of Length starting from TIFF2026042037000305.tif6150 For each subband Bi in TIFF2026042037000306.tif6150, The source spectrum for TIFF2026042037000307.tif6150 can be found here. The subband decomposition can be the same as the subband decomposition used to code the zfl, but can also be different, higher, or lower. For example, if the TNS is not active and the HF is not tonal, the random noise spectrum X N is used as the source spectrum for all subbands. In another example, X N The other source is empty, or the smallest copy-up destination: It is used as the source spectrum for several subbands starting from the bottom of TIFF2026042037000308.tif6150.

[0102] In another example, if the TNS is not active and the HF is tonal, the predicted spectrum X NP teeth, Starting from the bottom of TIFF2026042037000309.tif6150, E B E B may be used as the source for subbands at least 12 dB above , and the predicted spectrum is obtained from a previously decoded spectrum or from a signal obtained from a previously decoded spectrum (e.g., from a decoded TD signal). In cases not included in the above examples, the distance d C teeth, TIFF2026042037000310.tif6150 or TIFF2026042037000311.tif6150 and A mixture of TIFF2026042037000312.tif6150, Starting from TIFF2026042037000313.tif6150 TIFF2026042037000314.tif6150, where TIFF2026042037000315.tif6150. In one example, if the TNS is active but only starting at higher frequencies (e.g., 4500 Hz), and the HF is not tonal, TIFF2026042037000316.tif6150 and TIFF2026042037000317.tif6150 is a mixture of TIFF2026042037000318.tif6150 may be used as the source spectrum. TIFF2026042037000319.tif6150 or a spectrum consisting of zeros only may be used as a source. If it is TIFF2026042037000320.tif6150, d C teeth TIFF2026042037000321.tif6150. If TNS is active, a positive integer n may be found, resulting in TIFF2026042037000322.tif10150 and d C teeth, TIFF2026042037000323.tif10150, for example, may be set to the smallest such integer n. If TNS is not active, another positive integer n may be found, resulting in TIFF2026042037000324.tif6150 and d C teeth, TIFF2026042037000325.tif6150, for example, may be set to the smallest such integer n.

[0103] In another example, the starting frequency f ZFStart X up to S The lowest subband in TIFF2026042037000326.tif6150 may be set to 0, the lowest subband X CT is X DT This means that it may be a copy of Next, in the block "zero filling" B An example of weighting the source spectra is given based on: E B One example of smoothing is TIFF2026042037000327.tif6150 may be obtained from zfl, each TIFF2026042037000328.tif6150 is E B Then, TIFF2026042037000329.tif6150 is smoothed: TIFF2026042037000330.tif9150 and TIFF2026042037000331.tif9150. Scaling Factor TIFF2026042037000332.tif6150 is calculated for each subband Bi according to the source spectrum. TIFF2026042037000333.tif18150

[0104] Additionally, scaling is performed by a factor calculated as Limited by TIFF2026042037000334.tif6150: TIFF2026042037000335.tif11150 Source spectral band TIFF2026042037000336.tif6150( TIFF2026042037000337.tif6150) is split into two halves, each half is scaled, and the first half is TIFF2026042037000338.tif6150, and the second half is The file is TIFF2026042037000339.tif6150.

[0105] In the above explanation, TIFF2026042037000340.tif6150 Derived using TIFF2026042037000341.tif6150, TIFF2026042037000342.tif6150 Derived using TIFF2026042037000343.tif6150, TIFF2026042037000344.tif6150 and TIFF2026042037000345.tif6150 TIFF2026042037000346.tif6150 and Note that this is derived using TIFF2026042037000347.tif6150. TIFF2026042037000348.tif6150 TIFF2026042037000349.tif6150 and TIFF2026042037000350.tif6150 and Derived using TIFF2026042037000351.tif6150. This description is It was used only to clarify the use of TIFF2026042037000352.tif6150. According to a further embodiment, E B teeth TIFF2026042037000353.tif6150 may be used to derive the above formula, which can be written in different ways. TIFF2026042037000354.tif18150

[0106] E B but Even in this further embodiment, which can be derived using TIFF2026042037000355.tif6150, TIFF2026042037000356.tif6150 and The value of TIFF2026042037000357.tif6150 can be the same as in the previous example. Scaled Source Spectral Bands TIFF2026042037000358.tif6150, where the scaled source spectral bands are TIFF2026042037000359.tif6150 is To get TIFF2026042037000360.tif6150 TIFF2026042037000361.tif6150 is added to.

[0107] Next, an example of quantizing the energy of a zero quantization line (as part of an iBPC) is given. X QZ X by setting non-zero quantization lines to zero MR For example, X N In the same way as Q The values ​​at the positions of the non-zero quantization lines of are set to zero, and the zero parts between the non-zero quantization lines are MR Windowed in X QZ is generated. Energy per band i for the zero line ( TIFF2026042037000362.tif6150) is X QZ : Calculated from TIFF2026042037000363.tif18150.

[0108] TIFF2026042037000364.tif6150, for example, is quantized using a step size of 1 / 8 and limited to 6 / 8. TIFF2026042037000365.tif6150 is quantized to exactly zero EZ is coded only as individual zfl for the subbands above, where f EZ is 3000Hz. In addition, one energy level TIFF2026042037000366.tif6150 is f EZ Zero subbands below f EZ All from the zero subbands above tif2026042037000367.tif6150, where TIFF2026042037000368.tif6150 is quantized to zero, zero subband means the complete subband is quantized to zero. TIFF2026042037000369.tif6150 is quantized with a step size of 1 / 16 and bounded to 3 / 16. The energy of individual zero lines within the non-zero subbands is estimated (e.g., by the decoder) and is not explicitly coded. The value of TIFF2026042037000370.tif6150 is obtained from zfl on the decoder side, and The value of TIFF2026042037000371.tif6150 is This corresponds to the quantized value of TIFF2026042037000372.tif6150. Therefore, E consisting of TIFF2026042037000373.tif6150 B The value of is the optimal quantization step It may be coded according to TIFF2026042037000374.tif6150, which is taken as input by a parametric coder 156pc. Another example is shown in Figure 3, where the optimal quantization step Other quantization step sizes specific to the parametric coder may be used, independent of TIFF2026042037000376.tif6150. In yet another example, a non-uniform scalar or vector quantizer may be used to code zfl. However, in the presented example, X MR Quantization from 0 to 0 is the optimal quantization step. TIFF2026042037000377.tif6150 depends on the optimal quantization step It is advantageous to use TIFF2026042037000378.tif6150.

[0109] Long-Term Prediction (LTP) Next, the block LTP is explained. The time domain signal y C is used as input to the LTP, where y is obtained as the output of the IMDCT from X. The IMDCT consists of the inverse MDCT, windowing, and overlap-add. C The overlapping and non-overlapping parts on the left side of are stored in the LTP buffer. The LTP buffer is used in subsequent frames in the LTP to generate prediction signals for the entire window of the MDCT, as shown in Figure 17a. If a shorter overlap, for example half an overlap, is used for the right overlap in the current window, the non-overlapped portion "Overlap Difference" is also saved in the LTP buffer. Therefore, the sample at position "Overlap Difference" (see Figure 17b) is also saved in the LTP buffer together with the sample at the position between the two vertical lines before "Overlap Difference". The non-overlapped portion "Overlap Difference" is not present in the decoder output in the current frame, but only in subsequent frames (see Figures 17b and 17c). If a shorter overlap is used for the left overlap within the current window, the entire non-overlapping portion up to the start of the current window is used as part of the LTP buffer for generating the prediction signal.

[0110] The prediction signal for the entire MDCT window is generated from the LTP buffer. The time interval of the window length is the hop size L updateF0 =L subF0 / 2 with length L sub The overlap length is L. Other hop sizes and relationships between subinterval lengths and hop sizes may be used. updateF0 -L subF0 It can be the following: L subF0 is chosen so that significant pitch changes are not expected within the subinterval. updateF0 teeth, It is closest to TIFF2026042037000379.tif6150, TIFF2026042037000380.tif6150 is an integer not greater than LsubF0 is 2L updateF0 In another example, the frame length or window length is set to L updateF0 It may be further required that the function be divisible by Below, an example of a calculation means (1030) configured to derive subinterval parameters from coded pitch parameters depending on the position of the subinterval within an interval associated with a frame of the coded audio signal is also given, as well as an example where the parameters are derived from the coded pitch parameters and the subinterval position within an interval associated with a frame of the coded audio signal. subCenter The pitch lag at d is obtained from the pitch contour. In the first step, the subinterval pitch lag d subF0 is the pitch lag d at the center of the subinterval contour [i subCenter ]. The distance of the subinterval end to the window start (i subCenter +L subF0 / 2) is d subF0 As long as it is greater than d subF0 is the left position d of the subinterval center subF0 lag from the pitch contour at i subCenter +L subF0 / 2 <d subF0 Until subF0 =d subF0 +dcontour[i subCenter -d subF0 ]. The distance of the subinterval end to the window start (i subCenter +L subF0 / 2) is sometimes called the subinterval end.

[0111] In each subinterval, the predicted signal is fed to the LTP buffer and the transfer function H LTP (z) where TIFF2026042037000381.tif6150 where T int d subF0 The integer part of, i.e. TIFF2026042037000382.tif6150, T fr d subF0 The fractional part of T fr =d subF0 -T int and B(z,T fr ) is a fractional delay filter. B(z,T fr ) may have low-pass characteristics (or it may not emphasize high frequencies). The predicted signals are then cross-faded in the overlapping regions of the subintervals.

[0112] Alternatively, the predicted signal can be expressed as a function of the transfer function H LTP2(z) and an LTP buffer used as the initial output of the filter, where The file is TIFF2026042037000383.tif11150. B(z,T fr ) example: TIFF2026042037000384.tif9150 TIFF2026042037000385.tif9150 TIFF2026042037000386.tif9150 TIFF2026042037000387.tif9150 In this example, fr is typically rounded to the nearest value from a list of values, and a filter B is predefined for each value in the list. The predicted signal XP' is X M windowed with the same window used to generate X P is transformed via MDCT to obtain

[0113] Below, an example is given of a means for modifying the predicted spectrum, or a derivative of the predicted spectrum, depending on parameters derived from the coded pitch parameters. P X P Harmonics of at least n Fsafeguard The amplitudes of distant MDCT coefficients are set to zero (or multiplied by a positive factor less than 1), where n Fsafeguard is, for example, 10. Alternatively, other windows than rectangular windows may be used to reduce the amplitude between harmonics. X P The harmonics of TIFF2026042037000388.tif6150, where L M is X P It is long, TIFF2026042037000389.tif6150 is the average corrected pitch lag. The harmonic positions are TIFF2026042037000390.tif6150. This removes noise between harmonics, especially when half-pitch lag is detected. The spectral envelope of X is PS To get this, for example, SNS E Via X M is perceptually flattened in the same way as

[0114] An example is given below in which the number of predictable harmonics is determined based on the coded pitch parameters: PS , X MS , and Using TIFF2026042037000391.tif6150, the number of predictable harmonics, n LTP is determined. n LTP is coded and transmitted to the decoder. LTP harmonics may be predicted, e.g., N LTP =8. X PS and X MS is the length TIFF2026042037000392.tif6150 N LTP Each band is divided into Starting with TIFF2026042037000393.tif6150, The file is TIFF2026042037000394.tif6150. n LTP for all n ≤ n LTP About X MS -X PS and X MS The ratio of the energies of LTP For example, τ=0.7. If no such n exists, then n=0 and LTP is not active in the current frame. Whether LTP is active or not is indicated by a flag. X PS and X MS Instead of X P and X M may be used. X PS and X MS Instead of XP S and X MT Alternatively, the number of predictable harmonics can be determined by the pitch contour d contour It may be determined based on the

[0115] When LTP is active, X excluding the zeroth coefficient PS First TIFF2026042037000395.tif6150 coefficients are X MT is subtracted from X MR Generate the zeroth and The coefficient exceeding TIFF2026042037000396.tif6150 is X MT From X MR will be copied to In the quantization process, X Q is X MR is obtained from, X is coded as spect, and X D It is obtained from spect by decrypting

[0116] Below, the predicted spectrum (X Pat least a part of ) or a derivative of the predicted spectrum (X PS a part of ) is combined with the error spectrum (X D ) to give an example of a synthesizer (157) configured to. When LTP is active, excluding the zero - th coefficient, X PS the first of TIFF2026042037000397.tif6150 coefficients are added to X D to generate X DT . The zero - th and TIFF2026042037000398.tif6150 coefficients greater than are copied from X D to X DT .

[0117] <{ Hereinafter, optional features of harmonic post - filtering are described. The time - domain signal y C is obtained from X C as the output of the IMDCT, where the IMDCT is composed of an inverse MDCT, windowing, and overlap - add. To reduce the noise between harmonics and output y C , a harmonic post - filter (HPF) following the pitch contour is applied to y H . Instead of y C , the synthesis of y C constructed from the decoded pulse waveform and the time - domain signal y may be used as the input to the HPF. As shown by Figure 18a. The HPF input to the current frame k is y c [n] (0 ≤ n < N). Past output samples y H [n] (-d HPFmax ≤ n < 0, where d HPFmax is at least the maximum pitch lag) are also available. The time - aliased part of the right overlap region of the inverse MDCT output may be included. N​IMDCT look-ahead samples are also available. An example is shown in which the time interval over which the HPF is applied is equal to the current frame, although different intervals may be used. The positions of the HPF current input / output, HPF past output, and IMDCT look-ahead relative to the MDCT / IMDCT window are shown by Figure 18a, which also shows the overlap portions that can be added as usual to produce the overlap-add.

[0118] If the HPF is signaled in the bitstream that it should use constant parameters, smoothing is used at the beginning of the current frame, followed by the HPF with constant parameters for the remainder of the frame. Alternatively, to determine whether constant parameters should be used, y C A pitch analysis may be performed on the region where smoothing is used. The length of the region where smoothing is used may depend on the pitch parameters. If no constant parameters are signaled, the HPF input is k,update =L k / 2 with length L k Other hop sizes may be used. The overlap length is L k,update -L k It can be the following: L k is chosen so that significant pitch changes are not expected within the subinterval. k,update is the integer closest to but not greater than pitch_mid / 2, and L k is 2L k,update Instead of pitch_mid, some other value can be used, e.g. the average of pitch_mid and pitch_start, or y C A value obtained from a pitch analysis for , or the minimum pitch lag expected within an interval, for example for a signal with varying pitch, may be used. Alternatively, a fixed number of subintervals may be selected. In another example, if the frame length is L, k,update It may be further required that it be divisible by (see Figure 18b). The number of subintervals in the current interval k is K k and in the previous interval k-1, k-1 In the next interval k+1, k+1 In the example shown in Figure 18b, K k =6 and K k-1 =4.

[0119] In other examples, the current (time) interval may be divided into a non-integer number of subintervals and / or the lengths of the subintervals may vary within the current interval, as shown below, as illustrated by Figures 18c and 18d. Current section k(1≦l≦K k For each subinterval l in k,l is found using a pitch search algorithm, which may be the same as or different from the pitch search used to obtain the pitch contour. The pitch search of the subinterval l reduces the complexity of the search across the subintervals and / or finds the value p k,l To increase the stability of p, values ​​derived from the coded pitch lags (pitch_mid, pitch_end) can be used, e.g., the values ​​derived from the coded pitch lags can be values ​​of the pitch contour. In other examples, the complexity of the search over the subintervals and / or the value p k,l Instead of coded pitch lag, y C In another example, when searching for a subinterval pitch lag, it is assumed that the intermediate output of the harmonic post-filtering for the previous subinterval is available and is used in the pitch search (including the subinterval of the previous interval). N aheadThe (potentially time-aliased) look-ahead samples may also be used to find the pitch within subintervals that cross interval / frame boundaries, or, for example, if look-ahead is not available, a delay may be introduced in the decoder to look ahead to the last subinterval within the interval. Alternatively, values ​​derived from the coded pitch lag (pitch_mid, pitch_end) may be used to find the pitch within the subinterval. May be used for TIFF2026042037000399.tif6150.

[0120] For harmonic post-filtering, a gain-adaptive harmonic post-filter may be used. In the example, the HPF has the transfer function: TIFF2026042037000400.tif11150, where B(z,T fr ) is a fractional delay filter. The selections are independent, so B(z,T fr ) may be the same as the fractional delay filters used in LTP, or they may be different. In HPF, B(z,T fr ) also acts as a low-pass (or tilt filter to de-emphasize high frequencies).

[0121] Transfer functions H(z) and B(z,T fr ) as a coefficient of b j (T fr An example of a difference equation for a gain adaptive harmonic postfilter with Instead of a low-pass filter with fractional delay, a discriminant filter may be used, B(z,T fr )=1 and different expressions: The result is TIFF2026042037000402.tif6150. The parameter g is the optimal gain. It models the amplitude variation of the signal (modulation) and is signal adaptive. The parameter h is the harmonic level. It controls the desired increase in signal harmonics and is signal adaptive. The parameter β also controls the increase in signal harmonics and can be constant or dependent on the sampling rate and bit rate. The parameter β can also be equal to 1. The value of the product βh should be between 0 and 1, with 0 resulting in no change in harmonics and 1 resulting in the maximum increase in harmonics. In practice, βh < 0.75 is common. The feedforward portion of the harmonic post-filter (i.e., 1-αβhB(z,0)) functions as a high-pass (or tilt filter that de-emphasizes low frequencies). The parameter α determines the strength of the high-pass filtering (or in other words, it controls the de-emphasis slope) and has a value between 0 and 1. The parameter α may be constant or depend on the sampling rate and bit rate. In embodiments, values ​​between 0.5 and 1 are preferred.

[0122] For each subinterval, the optimal payoff g k,l and harmonic level h k,l is found, or in some cases it can be derived from other parameters. Given B(z,T fr ), we define a function to shift / filter the signal as follows: TIFF2026042037000403.tif14150 TIFF2026042037000404.tif6150 TIFF2026042037000405.tif6150By these definitions, TIFF2026042037000406.tif6150 is a (sub)interval l with length L. TIFF2026042037000407.tif6150 signals C represents TIFF2026042037000408.tif6150 has B(z,0) C represents the filtering of y -p is the y for p (possibly small) samplesH represents a shift of

[0123] A signal y in a (sub)interval l with length L and shift p C and y H The normalized correlation of TIFF2026042037000409.tif6150 is defined as follows: TIFF2026042037000410.tif16150 An alternative definition of TIFF2026042037000411.tif6150 could be: TIFF2026042037000412.tif17150 TIFF2026042037000413.tif6150

[0124] An alternative definition is TIFF2026042037000414.tif6150 is n <T INT y in the past subinterval for H Represents. In the above definition, the fourth order B(z,T fr ) is used. Any other order may be used which requires a change of range for j. B(z,T fr )=1, which can be used when only integer shifts are considered. TIFF2026042037000415.tif6150 and The resulting file is TIFF2026042037000416.tif6150. The normalized correlation thus defined allows the calculation of the fractional shift p.

[0125] The parameters normcorr, l, and L define the window for the normalized correlation. In the above definition, a rectangular window is used. Any other type of window (e.g., Hann, cosine) may be used instead, which is given by w[n]. TIFF2026042037000417.tif6150 and This can be done by multiplying TIFF2026042037000418.tif6150, where w[n] represents the window. To obtain the normalized correlation for a subinterval, set l to the interval number and L to the length of the subinterval. The output in TIFF2026042037000419.tif6150 represents the ZIR of the gain-adaptive harmonic postfilter H(z) for subframe l, with β=h=g=1 and TIFF2026042037000420.tif6150 and T fr =pT int is.

[0126] Optimal gain g k,l models the amplitude variation (modulation) within subframe l. It may be calculated, for example, as the correlation of the predicted signal with the low-pass input divided by the energy of the predicted signal. TIFF2026042037000421.tif15150 Another example shows the optimal gain g k,l may be calculated as the energy of the low-pass input divided by the energy of the predicted signal. TIFF2026042037000422.tif15150 Harmonic level h k,l controls the desired increase in signal harmonics and can be calculated, for example, as the square of the normalized correlation. TIFF2026042037000423.tif6150Usually, the normalized correlation of the subinterval is already available from the pitch search in the subinterval. Harmonic level h k,l may also be modified depending on the LTP and / or depending on the decoded spectral characteristics. For example, TIFF2026042037000424.tif6150 can be set, where h modLTP is a value between 0 and 1 and is proportional to the number of harmonics predicted by LTP, and h modTilt is a value between 0 and 1, and X C In one example, n LTPIf h is 0 modLTP =0.5, otherwise The file is TIFF2026042037000425.tif17153. X C The slope of may be the ratio of the energy of the first 7 spectral coefficients to the energy of the following 43 coefficients.

[0127] Once the parameters for subinterval l are calculated, an intermediate harmonic post-filtering output can be generated for the portion of subinterval j that does not overlap with subinterval l+1, which is used in finding the parameters for subsequent subintervals, as described above. The subintervals overlap and a smoothing operation between the two filter parameters is used. The smoothing described in [3] may be used.

[0128] A preferred embodiment is described below. According to an embodiment, there is provided an apparatus for encoding an audio signal, the apparatus comprising the following entities: a time-spectral transformer (MDCT) for converting an audio signal having a sampling rate into a spectral representation; a spectral shaper (SNS) for providing a perceptually flattened spectral representation from a spectral representation, the perceptually flattened spectral representation being divided into subbands of a different (higher) frequency resolution than the spectral shaper; a rate-distortion loop for finding the optimal quantization step; a quantizer for providing a quantized spectrum of the perceptually flattened spectral representation or a derivative of the perceptually flattened spectral representation in response to an optimal quantization step; a lossless spectral coder for providing a coded representation of the quantized spectrum; a band-wise parametric coder for providing a perceptually flattened spectral representation or a parametric representation of a derivative of the perceptually flattened spectral representation, the parametric representation being dependent on an optimal quantization step and consisting of parameters describing the energy in subbands where the quantized spectrum is zero, so that at least two subbands have different parameters or at least one parameter is limited to only one subband; Equipped with.

[0129] Another embodiment provides an apparatus for encoding an audio signal, the apparatus comprising the following entities on the other hand: a time-spectral transformer (MDCT) for converting an audio signal having a sampling rate into a spectral representation; a spectral shaper (SNS) for providing a perceptually flattened spectral representation from a spectral representation, the perceptually flattened spectral representation being divided into subbands of a different (higher) frequency resolution than the spectral shaper; a rate-distortion loop for finding an optimal quantization step, the rate-distortion loop providing a quantization step and selecting the optimal quantization step in response to the quantization step at each loop iteration; a quantizer for providing a quantized spectrum of the perceptually flattened spectrum or a derivative of the perceptually flattened spectral representation in response to an optimal quantization step; a per-band parametric coder for providing a perceptually flattened spectral representation or a parametric representation of a derivative of the perceptually flattened spectral representation, the parametric representation depending on an optimal quantization step and consisting of parameters describing the energy in subbands where the quantized spectrum is zero; and a spectral coder decision for providing a decision as to whether joint coding of the coded representation of the quantized spectrum and the coded representation of the parametric zero subband representation satisfies a constraint that the total number of bits for joint coding is below a predetermined limit; Equipped with Both the quantized spectral coding representation and the parametric zero subband coding representation require a variable number of bits depending on the perceptually flattened spectral representation or the derivative of the perceptually flattened spectral representation.

[0130] According to an embodiment, both devices may be enhanced by a modifier that adaptively sets at least subbands in the quantized spectrum to zero depending on the content of the subbands in the quantized spectrum and the perceptually flattened spectral representation. Here, a two-step per-band parametric coder may be used, which is configured to provide, for subbands whose quantized spectrum is zero (such that at least two subbands have different parametric representations), a perceptually flattened spectral representation or a parametric representation of a derivative of the perceptually flattened spectral representation depending on the quantization step; In the first of the two steps, the band-wise parametric coder calculates the frequency f EZ Providing individual parametric representations for the higher subbands, In the second step, the frequencies f where the individual parametric expressions are zero are EZ upper subbands, and f EZ It provides additional average parametric representations for the lower subbands.

[0131] Another embodiment provides an apparatus for decoding an encoded audio signal, the apparatus for decoding comprising the following entities: a spectral domain audio decoder for generating a decoded spectrum in response to a quantization step, the decoded spectrum being divided into subbands; a per-band parametric decoder that identifies a zero sub-band consisting of only zeros in the decoded spectrum and decodes a parametric representation of the zero sub-band using a quantization step, the parametric representation consisting of parameters that describe the energy in the zero sub-band, such that at least two sub-bands have different parameters or at least one parameter is limited to only one sub-band; a per-band generator providing a spectrum generated per band according to a parametric representation of the zero subband; The composite spectrum for each band is The generated spectrum and decoded spectrum for each band, or Combining the spectrum generated for each band, the predicted spectrum, and the decoded spectrum a combiner that provides the a spectral shaper (SNS) for providing a reshaped spectrum from a composite spectrum per band or a derivative of the composite spectrum per band, the spectral shaper having a different (lower) frequency resolution than the subband division; a spectrum-to-time converter for converting the reshaped spectrum into a time representation; Equipped with.

[0132] Another embodiment is a generated spectrum combined with a decoded spectrum, or Combining predicted and decoded spectra a band-wise parametric spectrum generator that provides The product spectrum is obtained band by band from the source spectrum, and the source spectrum is Zero spectrum, or a second predicted spectrum, or Random noise spectrum, or Combining the decoded spectrum (and predicted spectrum) with the already generated part, Their synthesis It is one of the In at least some cases, the source is a combination of already generated parts and the decoded spectrum (and predicted spectrum). It should be noted that, according to a further embodiment, the source spectrum is weighted based on the energy parameter of the zero subband. The selection of the source spectrum for a subband depends on the subband position, the power spectrum estimate, the energy parameter, the pitch information, and the time information. According to an embodiment, the spectral representation (X MR ) is the number of parameters that describe the quantized representation (X Q ) may depend on

[0133] In yet another embodiment, the subbands (i.e., subband boundaries) for iBPC, zfl decoding, and zero filling are determined by the following equation: D and / or X Q Note that the position of the zero spectral coefficients in While some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0134] The encoded audio signal of the present invention may be stored on a digital storage medium or may be transmitted over a transmission medium, such as a wireless transmission medium or a wired transmission medium, such as the Internet. Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementation can be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, storing electronically readable control signals, which cooperate (or can cooperate) with a programmable computer system to execute the respective methods. Thus, the digital storage medium can be computer-readable. Some embodiments according to the present invention include a data carrier having electronically readable control signals capable of cooperating with a programmable computer system to perform one of the methods described herein. Generally, embodiments of the present invention may be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer. The program code may, for example, be stored on a machine-readable carrier.

[0135] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. A further embodiment of the inventive method is therefore a data carrier (or digital storage medium or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals may for example be adapted to be transmitted via a data communication connection, for example via the Internet.

[0136] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein. Further embodiments according to the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver. In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0137] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the forthcoming claims and not by the specific details presented as descriptions and explanations of the embodiments herein.

[0138] References [1] 3GPP, Technical Specification Group Services and System Aspects, Audio codec processing functions, Extended Adaptive Multi-Rate-Wideband (AMR-WB+) codec, Transcoding functions (Release 16), no. 26.290. 3GPP (registered trademark), 2020. [2] N. Rettelbach, B. Grill, G. Fuchs, S. Geyrsberger, M. Multrus, H. Popp, J. Herre, S. Wabnik, G. Schuller, and J. Hirschfeld, “Audio Encoder, Audio Decoder, Methods For Encoding And Decoding An Audio Signal, Audio Stream And Computer Program”, PCT / EP2009 / 0046022009 [3] S. Disch, M. Gayer, C. Helmrich, G. Markovic, and M. Luis Valero, “Noise Filling Concept”, PCT / EP2014 / 0516302014

[0139] [4] J. Herre and D. Schultz, “Extending the MPEG-4 AAC Codec by Perceptual Noise Substitution” in Audio Engineering Society Convention 104, 1998. [5] F. Nagel, S. Disch, and S. Wilde, "A continuous modulated single sideband bandwidth extension," in 2010 IEEE International Conference on Acoustics, Speech and Signal Processing, 2010, pp. 357–360. [6] C. Neukam, F. Nagel, G. Schuller, and M. Schnabel, "A MDCT-based harmonic spectral bandwidth extension method," in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, 2013, pp. 566–570.

[0140] [7] S. Disch, R. Geiger, C. Helmrich, F. Nagel, C. Neukam, K. Schmidt, and M. Fischer, “Apparatus, Method And Computer Program For Decoding An Encoded Audio Signal”, PCT / EP2014 / 0651182013. [8] S. Disch, F. Nagel, R. Geiger, BNThoshkahna, K. Schmidt, S. Bayer, C. Neukam, B. Edler, and C. Helmrich, “Apparatus And Method For Encoding Or Decoding An Audio Signal With Intelligent Gap Filling In The Spectral Domain”, PCT / EP2014 / 0651232013 [9] S. Disch, F. Nagel, R. Geiger, B. N. Thoshkahna, K. Schmidt, S. Bayer, C. Neukam, B. Edler, and C. Helmrich, "Apparatus And Method For Encoding And Decoding An Encoded Audio Signal Using Temporal Noise / Patch Shaping", PCT / EP2014 / 065123 2013

[10] S. Disch, A. Niedermeier, C. R. Helmrich, C. Neukam, K. Schmidt, R. Geiger, J. Lecomte, F. Ghido, F. Nagel, and B. Edler, "Intelligent Gap Filling in Perceptual Transform Coding of Audio", 2016

[0141]

[11] S. Disch, S. van de Par, A. Niedermeier, E. Burdiel Perez, A. Berasategui Ceberio, and B. Edler, "Improved Psychoacoustic Model for Efficient Perceptual Audio Codecs" in Audio Engineering Society Convention 145, 2018

[12] C. R. Helmrich, A. Niedermeier, S. Disch, and F. Ghido, "Spectral envelope reconstruction via IGF for audio transform coding" in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pages 389 - 393

[13] C.Neukam、S.Disch、F.Nagel、A.Niedermeier、K.SchmidtおよびBNThoshkahna、「Apparatus And Method For Decoding And Encoding An Audio Signal Using Adaptive Spectral Tile Selection」、PCT / EP2014 / 0651162013

[14] A.Niedermeier、C.Ertel、R.Geiger、F.Ghido、およびC.Helmrich、「Apparatus And Method For Decoding Or Encoding An Audio Signal Using Energy Information Values ​​For A Reconstruction Band」、PCT / EP2014 / 0651102013

[0142]

[15] S.Disch、B.Schubert、R.Geiger、およびM.Dietz、Apparatus And Method For Audio Encoding And Decoding Employing Sinusoidal Substitution」、PCT / EP2012 / 0767462012

[16] S.Disch、B.Schubert、R.Geiger、B.Edler、およびM.Dietz、Apparatus And Method For Efficient Synthesis Of Sinusoids And Sweeps By Employing Spectral Patterns」、PCT / EP2013 / 0695922013

[17] M.Dietz、G.Fuchs、C.Helmrich、およびG.Markovic、、Low-Complexity Tonality-Adaptive Audio Signal Quantization」、PCT / EP2014 / 0516242014

[18] M. Oger, S. Ragot, and M. Antonini, “Model-based deadzone optimization for stack-run audio coding with uniform scalar quantization,” in 2008 IEEE International Conference on Acoustics, Speech and Signal Processing, 2008, pp. 4761–4764.

[0143]

[19] C. Helmrich, J. Lecomte, G. Markovic, M. Schnell, B. Edler, and S. Reuschl, “Apparatus And Method For Encoding Or Decoding An Audio Signal Using A Transient-Location Dependent Overlap”, PCT / EP2014 / 053293, 2014.

[20] 3rd Generation Partnership Project, Technical Specification Group Services and System Aspects, Codec for Enhanced Voice Services (EVS), Detailed algorithmic description, no. 26.445. 3GPP (registered trademark), 2019.

[21] G. Markovic, E. Ravelli, M. Dietz, and B. Grill, “Signal Filtering”, PCT / EP2018 / 080837, 2018.

[22] E. Ravelli, M. Schnell, C. Benndorf, M. Lutzky, and M. Dietz, Apparatus And Method For Encoding And Decoding An Audio Signal Using Downsampling Or Interpolation Of Scale Parameters, US Patent PCT / EP2017 / 078921

[0144]

[23] E. Ravelli, M. Schnell, C. Benndorf, M. Lutzky, M. Dietz, and S. Korse, Apparatus And Method For Encoding And Decoding An Audio Signal Using Downsampling Or Interpolation Of Scale Parameters, U.S. Patent PCT / EP2018 / 0801372018

[24] Low Complexity Communication Codec.Bluetooth, 2020

[25] Digital Enhanced Cordless Telecommunications(DECT), Low Complexity Communication Codec plus(LC3plus), no.103 634.ETSI, 2019

Claims

1. A spectral representation of an audio signal divided into multiple subbands (X MR ), wherein the spectral representation (X MR ) is composed of frequency bins or frequency coefficients, and at least one subband includes more than one frequency bin, and the encoder (1000) The spectral representation (X MR ) quantized representation (X Q a quantizer (1030) configured to generate The quantized representation (X Q ) according to the spectral representation (X MR a per-band parametric coder (1010) configured to provide a coded parametric representation (zfl) of the spectral representation (X) in the subband, MR ) or the spectral representation (X MR ) and the spectral representation (X MR ) and a per-band parametric coder (1010) in which there are parameters describing An encoder (1000) comprising:

2. At least one subband of the plurality of subbands is quantized to zero, or the quantized representation (X Q ) a spectral representation (X MR ) is zero, and / or The per-band parametric coder (1010) converts the quantized representation (X Q ) and / or the per-band parametric coder (1010) determines the at least one sub-band of the plurality of sub-bands in the quantized representation (X Q ) and / or coding the at least one subband of the plurality of subbands quantized to zero within the parameters describe the energy within the subbands, or the parameters describe the energy within the subbands quantized to zero; The encoder (1000) of claim 1.

3. The coded parametric representation (zfl) uses a variable number of bits or the number of bits used to represent the coded parametric representation (zfl) varies depending on the spectral representation (Xfl) of the audio signal. MR ) and / or The coded representation (spect) uses a variable number of bits or the number of bits used to represent the coded representation (spect) varies depending on the spectral representation (X MR ) and / or The coded representation (spect) uses entropy coding with a variable number of bits, and / or the number of bits required for the entropy coding of the zero-filling level is calculated, and / or the number of bits used to represent the coded parametric representation (zfk) and the coded representation (spect) is below a predetermined threshold; An encoder (1000) according to claim 1 or 2.

4. The quantized representation (X Q and / or a spectral coder (1020) configured to generate a coded representation (spect) of The per-band parametric coder (1010) forms a joint coder together with a spectral coder (1020) and / or the per-band parametric coder (1010) together with a spectral coder (1020) encodes the spectral representation (X MR ) together to obtain a coded version of An encoder (1000) according to any one of claims 1 to 3.

5. The encoder (1000, 101, 101') of any one of claims 1 to 4, further comprising a time-spectral transformer or an MDCT transformer configured to transform an audio signal having a sampling rate into said spectral representation to obtain said spectral representation.

6. The spectral representation is perceptually flattened, and / or the encoder (1000, 101, 101′) and / or further comprising a spectral shaper configured to provide a perceptually flattened spectral representation from said spectral representation. The perceptually flattened spectral representation is divided into subbands of a different or higher frequency resolution than the coding spectral shape used for spectral flattening, and / or the encoder (1000, 101, 101′) means for processing the input signal of the time-spectral transformer or the MDCT transformer with an LP filter to spectrally flatten the audio signal; An encoder (1000, 101, 101') according to any one of claims 1 to 5.

7. further comprising a rate-distortion loop configured to determine or estimate an optimal quantization step; and / or and / or a rate-distortion loop configured to perform at least two iteration steps or at least two iteration steps for two quantization steps; a rate-distortion loop configured to adapt a quantization step depending on a previous quantization step or to adapt the quantization step depending on a previous quantization step to determine an optimal quantization step; An encoder (1000, 1001) according to any one of claims 1 to 6.

8. The rate distortion loop may include a bit counter (1050) configured to estimate bits used for coding, and / or the spectral representation (X MR 8. The encoder (1000, 101, 101') of claim 7, comprising a recoder (1055) configured to recode the parameters describing the .

9. The spectral representation (X MR ) is expressed as: Q 9. An encoder (1000, 101, 101') according to any one of claims 1 to 8, which relies on

10. The quantized representation (X Q and / or the coded representation of the quantized spectrum and the coded representation of the parametric representation are both based on a variable number of bits and a quantization step that depend on the spectral representation or on a derivative of the perceptually flattened spectral representation. An encoder (1000) according to any one of claims 4 to 9.

11. the quantized spectrum and / or the spectral representation of the audio signal (X MR 11. The encoder (1000, 300) of claim 1, further comprising a modifier (156m, 302) configured to adaptively set at least subbands in the quantized spectrum to zero depending on the content of the subbands in the quantized spectrum.

12. The parameters describe the energy within the subbands, and the per-band parametric coder (1010) comprises two stages, and in the first of the two stages, the per-band parametric coder (1010) performs a code at a certain frequency (f EZ ), wherein the second of the two stages is configured to provide an individual parametric representation of the subbands above the frequency (f EZ ), wherein the individual parameter representations are zero, and the frequency (f EZ 12. The encoder (1000) of claim 1, for sub-bands below 12.

13. A decoder (1200) for decoding an encoded audio signal, said encoded audio signal consisting of at least a coded representation of the spectrum (spect) and a coded parametric representation (zfl), said coded audio signal being subjected to a quantization step and wherein the decoder (1200) The coded representation of the spectrum (spect) and the quantization step The decoded and dequantized spectrum (X D a spectral domain decoder (1230, 156sd) configured to generate the decoded and dequantized spectrum (X D a spectral domain decoder (1230, 156sd) in which the sigma-based signal is split into subbands; The decoded spectrum or the decoded and dequantized spectrum (X D ) and identifying a zero subband in the coded parametric representation (zfl), and B a per-band parametric decoder (1210, 162) configured to decode the Equipped with The parametric representation (E B ) contains parameters describing a subband, and there are at least two different subbands, and thus parameters in at least two different subbands, and / or the coded parametric representation (zfl) is expressed using a variable number of bits, and / or the number of bits used to represent the coded parametric representation (zfl) depends on the coded representation of the spectrum (spect), Decoder (1200).

14. A decoder (1200) for decoding an encoded audio signal, wherein the encoded audio signal is subjected to a quantization step and wherein the decoder (1200) The decoded and dequantized spectrum (X D a spectral domain decoder (1230, 156sd) configured to generate the decoded and dequantized spectrum (X D a spectral domain decoder (1230, 156sd) in which the sigma-based signal is split into subbands; The decoded spectrum or the decoded and dequantized spectrum (X D ) and based on the encoded audio signal, generate a parametric representation (E B a per-band parametric decoder (1210, 162) configured to decode the The parametric representation of the zero subband (E B a per-band spectrum generator (1220, 158sg) configured to generate a per-band generated spectrum in response to the The composite spectrum for each band (X CT ), wherein the combined spectrum for each band (X CT ) is the generated spectrum for each band and the decoded and dequantized spectrum (X D ) or the generated spectrum for each band and the predicted spectrum (X PS ) and the decoded and dequantized spectrum (X D Synthesis of (X DT a combiner (1240, 158c) including combining The composite spectrum (X CT ) or the composite spectrum (X CT a spectral-to-time converter (1250, 161) configured to convert the derivative of A decoder (1200) comprising:

15. The composite spectrum (X CT ) is a reshaped spectrum (X ) that has been reshaped by using a spectrum shaper (SNS) and / or a noise shaper (TNS). C ) and / or said decoder (1200) comprises: further comprising means configured to obtain a time domain signal from the output of the spectrum-time converter and / or means configured to spectrally shape the time domain signal (derived from the output of the spectrum-time converter) by processing with an LP filter; The decoder (1200) of claim 14.

16. The per-band parametric decoder (1210, 162) generates a parametric representation (E) of the zero subband based on the encoded audio signal using a quantization step. B ) and / or The parametric representation (E B ) comprises parameters describing the energy in the subbands, and there are at least two different subbands, and thus there are parameters describing the energy in at least two different subbands; and / or The parametric representation (E B ) includes parameters describing the energy within the subbands, and / or The energy of individual zero lines within the non-zero subbands is estimated and not explicitly coded, and / or the zero subband is defined by the decoded spectrum or the decoded and dequantized spectrum output from the spectral decoder (1200); and / or the coded parametric representation (zfl) is coded using a variable number of bits, and / or the number of bits used to represent the coded parametric representation (zfl) depends on the coded representation of spectrum (spect), and / or The parametric representation (E B ) depends on the coded representation of the spectrum (spect), A decoder (1200) according to claim 13, 14 or 15.

17. The parametric representation of the zero subband (E B ) is the quantization step or decrypted accordingly the parametric representation depends on the coded representation of the spectrum (spect), A decoder (1200) according to claim 13, 14, 15 or 16.

18. The per-band parametric decoder (1210, 162) uses information from the output of the spectral domain decoder (1230, 156sd) or the decoded and dequantized spectrum (X D ) to determine the parametric representation (E B 18. The decoder (1200) of claim 13, 14, 15, 16 or 17, configured to decode a

19. The spectral shaper uses the spectral shape obtained from the coded spectral shape to generate the composite spectrum for each band (X CT ) or the composite spectrum (X CT 15. The decoder (1200) of claim 14, configured to spectrally shape the derivative of (a) (b), wherein the coding spectral shape uses a different or lower frequency resolution than the subband decomposition.

20. The decoded and dequantized spectrum (X D ) or the predicted spectrum and the decoded and dequantized spectrum (X DT ) to be added to the synthesis of the resulting spectrum (X G ), wherein the generated spectrum (X G ) is obtained band-by-band from the source spectrum, and the source spectrum is - second predicted spectrum (X NP ),or - Random noise spectrum (X N ),or - an already generated part of the generated spectrum, or - the decoded and dequantized spectrum (X DT ) or the predicted spectrum and the decoded and dequantized spectrum (X DT ) or - a combination of one or two of the above One of the A decoder (1200) according to any one of claims 13 to 19.

21. The decoded and dequantized spectrum (X D ) or the predicted spectrum and the decoded and dequantized spectrum (X DT ) to be added to the synthesis of the resulting spectrum (X G a per-band parametric spectrum generator (158sg) configured to generate the generated spectrum (X G ) is obtained band-by-band from the source spectrum, and the source spectrum is - second predicted spectrum (X NP ),or - Random noise spectrum (X N ),or - the product spectrum (X G ), or - the decoded and dequantized spectrum (X DT ) or the predicted spectrum and the decoded and dequantized spectrum (X DT ) or - a combination of one or two of the above One of the Parametric spectrum generator per band (158sg).

22. 20. The decoder (1200) of any one of claims 13 to 19, wherein the source spectrum is weighted based on an energy parameter of the zero subband.

23. The source spectrum is the zero-band energy parameter (E B 22. The parametric spectrum generator (158sg) for each band according to claim 20 or 21, wherein the parametric spectrum generator (158sg) is weighted based on the

24. The selection of the source spectrum (158sc) for a subband is determined by the subband position, the tonality information (toi), the power spectrum estimate (Z C ), energy parameter (E B 21. The decoder (1200) of any one of claims 13 to 20, which relies on at least one of the following information: pitch information (pii), and / or time information (tei).

25. The selection of the source spectrum (158sc) for a subband is determined by the subband position, the tonality information (toi), the power spectrum estimate (Z C ), energy parameter (E B 24. A parametric spectrum generator (158sg) per band according to claim 21 or 23, which depends on at least one of the following: pitch information (pii), and / or time information (tei).

26. The tonality information is φ H and / or the pitch information is and / or time information is the information whether TNS is active or not.

27. The tonality information is φ H and / or the pitch information is and / or time information is the information whether TNS is active or not.

28. A spectral representation of an audio signal divided into multiple subbands (X MR ), comprising the steps of: MR ) is comprised of frequency bins or frequency coefficients, and at least one subband includes more than one frequency bin, and the method further comprises: The spectral representation of the audio signal divided into a plurality of subbands (X MR ) quantized representation (X Q ) The quantized representation (X Q ) according to the spectral representation (X MR ), wherein said coded parametric representation (zfl) is a coded parametric representation of said spectral representation (X MR ) or the spectral representation (X MR ) and the spectral representation (X MR ) there are parameters describing the steps and A method comprising:

29. A method for decoding an encoded audio signal, said encoded audio signal consisting of at least a coded representation of the spectrum (spect) and a coded parametric representation (zfl), said encoded audio signal being subjected to a quantization step wherein the method further comprises: The coded representation of the spectrum (spect) and the quantization step The decoded and dequantized spectrum (X D ), generating the decoded and dequantized spectrum (X D ) is divided into subbands; The decoded spectrum or the decoded and dequantized spectrum (X D ) and identifying a zero subband in the coded parametric representation (zfl), and B ) and Including, The parametric representation (E B ) contains parameters describing a subband, and there are at least two different subbands, and thus parameters in at least two different subbands, and / or the coded parametric representation (zfl) is expressed using a variable number of bits, and / or the number of bits used to represent the coded parametric representation (zfl) depends on the coded representation of the spectrum (spect), method.

30. 1. A method for decoding an encoded audio signal, said method comprising: The decoded and dequantized spectrum (X D ), generating the decoded and dequantized spectrum (X D ) is divided into subbands; The decoded spectrum or the decoded and dequantized spectrum (X D ) and generating a parametric representation (E B ) and The parametric representation of the zero subband (E B ) generating a generated spectrum for each band according to the The composite spectrum for each band (X CT ), wherein the band-by-band composite spectrum (X CT ) is the generated spectrum for each band and the decoded and dequantized spectrum (X D ) ), or a combination of the generated spectrum for each band, and a predicted spectrum (X PS ) and the decoded and dequantized spectrum (X D Synthesis of (X DT ) and The composite spectrum (X CT ) or the composite spectrum (X CT ) to convert its derivative to a time representation; A method comprising:

31. A method for generating a generated spectrum per band, comprising: D ) or the predicted spectrum and the decoded and dequantized spectrum (X DT ) to be added to the synthesis of the resulting spectrum (X G ), wherein the generated spectrum (X G ) is obtained band-by-band from the source spectrum, and the source spectrum is - second predicted spectrum (X NP ),or - Random noise spectrum (X N ),or - the product spectrum (X G ), or - combination of at least two of the above The method is one of the above.

32. A computer readable digital storage medium storing a computer program having a program code for performing the method of any one of claims 28 to 31 when the program is run on a computer.