Audio coding with distortion weight conversion

By controlling distortion in the gammatone domain and converting it to the transform domain, the encoder optimizes quantization step sizes for improved noise shaping and speech quality in audio encoding.

WO2025226573A1PCT designated stage Publication Date: 2025-10-30DOLBY LABORATORIES LICENSING CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/025549
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-19
Filing Date
2025-04-21
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing audio encoding methods struggle with implementing perceptual modeling in the codec transform domain, particularly in controlling noise shape and quantization distortion, which can be challenging and inefficient.

Method used

An encoder is configured to control distortion in the analysis filter bank domain, such as the gammatone domain, and convert it back to the transform domain using weighted distortion functions to optimize quantization step sizes across time and frequency, ensuring perceptual-based distortion control.

Benefits of technology

This approach achieves smoother noise shaping, better transient signal handling, and improved speech quality by optimizing the distribution of distortion and bit allocation, resulting in a more efficient audio encoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000003_0001
    Figure IMGF000003_0001
  • Figure IMGF000007_0001
    Figure IMGF000007_0001
  • Figure IMGF000008_0001
    Figure IMGF000008_0001
Patent Text Reader

Abstract

Encoding of an audio signal comprising transforming the audio signal using a frame-based frequency transform, applying an analysis filter bank to the audio signal, the filter bank domain having B frequency bands, defining a set of B weight functions wherein each weight function is related to a perception based distortion of the audio signal in one frequency band of the filter bank domain, converting the set of B weight functions into a set of bin weights in the frame-based transform domain, using the bin weights υα to calculate quantization step sizes configured to distribute the mean encoding bitrate across time and frequency; and quantizing the transform domain representation of the audio signal using the quantization step sizes. With this approach, a perception-based distortion function is defined for use in audio coding in a frame-based transform domain, via weighting defined in an analysis filter bank domain.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AUDIO CODING WITH DISTORTION WEIGHT CONVERSION CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from U.S. Provisional Application No.63 / 638,198, filed on April 24, 2024, and European Patent Application number 24183208.8, filed on June 19, 2024, each of which are incorporated by reference herein in their entirety. TECHNICAL FIELD

[0002] The present invention relates to audio encoding, and in particular to encoding where quantization of individual frames is scaled based on a weighted distortion. BACKGROUND

[0003] Today, most audio encoding involves a frequency transform into a transform domain representation, for example using the Modified Discrete Cosine Transform (MDCT). The transform representation typically includes one (real or complex) coefficient for each time- frequency bin, and in the next stage these coefficients are quantized to reduce the required bitrate.

[0004] It is known to apply a perceptually motivated control of the noise shape in time and frequency incurred by the coding scheme (e.g. quantization). For instance, the model discussed in WO 2021 / 113416 may be used to define a required Signal to Noise Ratio (SNR) for each band of a Modified Discrete Cosine Transform (MDCT) domain of the codec.

[0005] For more advanced perceptual modelling, a gammatone filter bank domain may be used, and an encoding scheme based on distortion analysis in the gammatone domain was proposed in “A new perceptual model for audio coding based in spectro-temporal masking”, by van de Par et al, Audio Engineering Society, 2008. SUMMARY

[0006] The approach formulated by van de Par et al can be difficult to implement, and it is therefore desirable to provide an improved approach to performing perceptual modelling in one domain and applying such modelling to the codec transform domain to provide an improved control of the noise shape (e.g. quantization distortion).

[0007] It is an object of the present disclosure to provide such an improved approach. Aspects of this approach are defined by the appended claims.

[0008] As discussed above, an encoder quantizes transform coefficients, thereby introducing distortion in the transform domain, such as the MDCT domain. According to aspects of the present invention, an encoder is configured to control this distortion in the currency of a weighted distortion defined in an analysis filter bank domain, such as a gammatone domain. This weighted distortion is converted back to a weighted mean distortion in the transform domain.

[0009] According to one aspect of the invention, the above object is achieved by a method for encoding an audio signal, s(t), having a duration, I, into an encoded audio signal, the method comprising transforming the audio signal using a frame-based frequency transform to obtain a transform domain representation of the audio signal, applying an analysis filter bank to the audio signal to obtain a filter bank domain representation of the audio signal, the filter bank domain having B frequency bands, defining a set of B weight functions, wb(t), wherein each weight function is related to a perception based distortion of the audio signal in one frequency band, b, of the filter bank domain, converting the set of B weight functions, wb(t), into a set of bin weights υα in the frame-based transform domain, where each bin weight with a specific time-frequency bin of the frame-based transform domain, using weights υαto calculate quantization step sizes ΔJ configured to distribute a total number of bits, R, across time and frequency; and quantizing the transform domain representation of the audio signal using the quantization step sizes.

[0010] According to another aspect, the above object is achieved by an encoder for encoding an audio signal, s(t), having a duration, I, into an encoded audio signal, comprising a transform block for transforming the audio signal using a frame-based frequency transform to obtain a transform domain representation of the audio signal, an filter bank application block for applying an analysis filter bank to the audio signal to obtain a filter bank domain representation of the audio signal, the filter bank domain having B frequency bands, a weight determination block for defining a set of B weight functions, wb(t), wherein each weight function is related to a perception based distortion of the audio signal in one frequency band, b, of the filter bank domain, a weight conversion block for converting the set of B weight functions, wb(t), into a set of bin weights υα in the frame-based transform domain, where each bin weight is associated with a specific time-frequency bin of the frame-based transform domain, a scaling block for calculating quantization step sizes ΔJ based on the bin weights υα which step sizes are configured to distribute a total number of bits, R, across time and frequency; and a quantization block for quantizing the transform domain representation of the audio signal using the quantization step sizes ΔJ.

[0011] According to these aspects, a perception-based distortion function is defined for use in audio coding in a frame-based transform domain, via weighting defined in an analysis filter bank domain. Put differently, a perceptually motivated distortion function is defined by weight functions in the analysis filter bank domain for use in audio coding in a frame-based transform domain (the “codec domain”). This is achieved by efficient conversion of the weighting in the analysis domain into the transform domain of the audio coder. The converted weights can be used to reach a given bit budget (or mean bitrate) over the considered signal segment.

[0012] The weights (i.e. the distortion function) relate to a “perception based” distortion of the audio signal. By “perception based” is intended to mean a distortion as it is perceived by a (generic) listener. A band in which a distortion is perceived strongly will be associated with a larger weight. A perception based distortion may be determined using some sort of model of the human auditory system, e.g., calibrated by the results of psychoacoustic experiments.

[0013] Compared to audio coding based on transform domain perceptual noise shaping, the observed benefits of the present disclosure are in the smoothness of resulting noise shaping, in better transient signal handling, and improved speech quality.

[0014] The weighted distortion is defined over the complete signal, i.e. a signal time segment which is significantly longer than the frame length (block size, transform length) on which the frame-based frequency transform operates. This allows unequal distribution contributions across time and frequency. By a proper scaling according to this disclosure, the rate optimization pushes the result in a somewhat uniform distribution of weighted distortion, which results in a controllable noise shaping.

[0015] By “frame-based” is intended to indicate that the encoding transform is performed on one “frame” (block, window) of samples at a time. The number of (frequency) samples in the transform domain representation of a frame may be less than the number of (time) samples in the frame. For example, the MDCT transform reduces the number of samples by a factor two (2N time samples result in N frequency samples). This may be compensated by taking time steps (“strides”) shorter than the transform length. For example, the stride may be 50% of the frame. This is referred to as a “lapped” transform.

[0016] Encoding according to this aspect enables a satisfactory control of a tradeoff between distortion and total number of bits (or mean bitrate over the signal duration). For example, the total number of bits can be controlled (e.g., minimized) given a limit on a weighted distortion based on the bin weights. Or, the weighted distortion can be controlled (e.g., minimized) given a maximum available mean bitrate. Other satisfactory combinations of distortion and number of bits may also be contemplated.

[0017] The analysis filter bank domain introduces no artificial framing in the distortion measure, and is therefore advantageous for distortion analysis. Also, the analysis filter bank typically has a significantly higher time resolution, thereby allowing a physiologically plausible definition of distortion.

[0018] In some implementations, each weight function is equal to a constant divided by a masking function, where the masking function represents a level at which a distortion will be masked by the audio signal to make the distortion unperceivable by a human listener. The sum of all constants is preferably equal to the number of bands.

[0019] In a simple case the constant is equal to one, but alternatively the constants may be different for different bands. Such a set of constants may provide a calibration of the distortion weights to achieve a desired stationary noise shaping.

[0020] In some applications, the transform domain is grouped into a set of “scale-factor” bands based on the frequency resolution of the human ear, and the quantization step size may then be set equal for all bins in a group of bins belonging to the same scale-factor band. In this case, a variance of a quantization error for a specific bin group of bins having a same quantization step size may be assumed to be proportional to an inverse of a group bin weight equal to an average of the bin weights in the group. This makes it possible to apply a simplified distortion target conversion to the transform domain of the codec, where each specific step size is set to be proportional to the inverse of the group weight. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The present invention will be described in more detail with reference to the appended drawings, showing currently preferred embodiments of the invention.

[0022] Figure 1 is a schematic block diagram of an encoder system according to an embodiment of the present invention.

[0023] Figure 2 is a flow chart of a method according to an embodiment of the present invention.

[0024] Figure 3 is a spectral representation of impulse responses of a gammatone filter bank.

[0025] Figure 4 shows a time-frequency representation of a signal in a gammatone filter bank domain.

[0026] Figure 5 shows a time-frequency representation of a signal in an MDCT domain.

[0027] Figure 6 is a graph showing test results comparing a signal encoded with the method in figure 2 and a signal encoded with conventional encoding.

[0028] Figure 7 illustrates a schematic block diagram of an example device or architecture that may be used to implement embodiments of the invention. DETAILED DESCRIPTION

[0029] Figure 1 shows schematically a variable bitrate (VBR) encoder 10 according to an example implementation of the invention.

[0030] The upper part of the block diagram includes a transform block 11 of an input audio signal 12 into a frame-based transform domain representation 13 of the signal. The input audio signal 12 has a duration I which may be in the order of seconds. Most significantly, it is considerably (e.g. an order of magnitude) longer than the frame length on which the transform block operates.

[0031] The transform block 11 operates on one frame of the input signal 12 at a time, and each frame is transformed into a set of frequency bins. Each bin (also referred to as a time- frequency bin) has a coefficient (value) cα and is associated with a specific waveform uα in the transform domain filter bank. The frames typically overlap, such that a next frame includes a portion of a preceding frame. The advancement of one frame compared to the previous frame is sometimes referred to as “stride” and may be 50% of the frame length. The length of the frame (also referred to as window length, or block size) can be variable, and depend on the dynamics of the signal. In the encoder in figure 1, the frame length is controlled by a transform control block 14.

[0032] In the illustrated example the transform block 11 implements a Modified Discrete Cosine Transform (MDCT). The frame length (block size, window length) of an MDCT transform is typically 768 or 1024 samples, and the stride is typically 50%, i.e.384 samples or 512 samples. A frame of 768 time samples is transformed into 384 bins, while a frame of 1024 time samples is transformed into 512 bins. The reduced number of samples is compensated for by the 50% stride, such that the total number of frequency samples will be the same as the number of time samples. This is referred to as a “lapped” transform.

[0033] The transform 11 is followed by a quantization block 15, configured to quantize the bin coefficients, cα. It is this quantization that introduces distortion that the present disclosure is intended to control.

[0034] The lower part of the block diagram includes an analysis filter bank application block 16. Contrary to the transform applied in the transform block 11, which operates on one signal frame at a time (as discussed above for the case of MDCT), the analysis filter bank is not frame-based, i.e. it operates on the entire input signal, one sample at a time. The analysis filterbank includes a set of frequency transforms each associated with a particular frequency band ^ ∈^. In the illustrated example, the analysis filter bank is a Gammatone filter bank, which will bediscussed in more detail below.

[0035] The output of the filter bank analysis is used in a consecutive weight definition block 17 to determine a set of weight functions or curves wb(t), one weight for each band b of the analysis domain.

[0036] The weights are provided to a conversion block 18, which is configured to convert the weights wb(t) into a set of bin weights vα, one bin weight for each bin of the transform domain representation. The conversion receives information about the transform length (frame length) that has been applied in the transform block 11 from the transform control block 14.

[0037] Finally, the lower branch includes a scaling block 19, configured to determine an appropriate step size for the quantization of each bin coefficient cα. The scaling in block 19 is governed by application specific control parameters, such as target or maximum available number of bits (mean bitrate), maximum acceptable distortion, etc. The scaling block 19 is connected to provide the calculated step sizes to the quantization block 15.

[0038] In one embodiment, the scaling block 19 is configured to calculate the quantization step sizes by controlling (e.g., minimizing) the mean encoding bitrate given a pre-defined limit on a weighted distortion measure based on the bin weights. In another embodiment, the scaling block is configured to calculate the quantization step sizes by controlling (e.g., minimizing) a weighted distortion measure based on the bin weights given a maximum available number of bits.

[0039] Figure 2 shows an example encoding method for encoding an audio signal, s(t), having a duration, I, with a variable encoding bitrate.

[0040] In step S11 (‘transform audio’), the audio signal is transformed using a frame-based frequency transform to obtain a transform domain representation of the audio signal. The transform includes a set of bins (also referred to as time-frequency bins), where each bin has a coefficient (value) cαand is associated with a specific waveform uαin the transform domain filter bank.

[0041] In step S12 (‘apply filter bank to audio’), an analysis filter bank including B filters each associated with a frequency band b, is applied to the audio signal to obtain a filter bank domain representation of the audio signal. It is noted that steps S11 and S12 are independent of each other and may be performed in any order.

[0042] In step S13 (‘define weight functions’), a set of B weight functions, wb(t), are defined, wherein each weight function is related to a perception based the audio signal in one frequency band, b, of the filter bank domain. The perception based distortion may be assessed or estimated using a model of the human auditory system. Specifically, as will be discussed below, the weights may be determines based on a masking level representing a level at which a distortion will be masked by the original signal and therefore not perceived.

[0043] In step S14 (‘convert weights’), the set of weight functions, wb(t), are converted into a set of bin weights υα in the frame-based transform domain, where each bin weight is associated with a specific time-frequency bin of the frame-based transform domain.

[0044] In step S15 (‘calculate quantization steps’) the bin weights are used to by scaling block 15 to calculate quantization step sizes Δ configured to distribute the a total number of available bits, R, or mean bitrate, across time and frequency. In many practical implementations, individual frequency lines of the transform domain are grouped in bands (scale-factor bands) based on the frequency resolution of the human ear, and one single step size (quantization level) is applied to all bins in a group of bins belonging to the same scale-factor band. In the following, the common step size shared by all bins in a group J is referred to as ΔJ.

[0045] Finally, in step S16 (‘quanitize coefficients’), the transform domain representation is quantized using the quantization step sizes ΔJ.

[0046] In the following, additional details of a possible encoding process performed in the encoder 10 will be discussed in more detail. Distortion definition

[0047] Consider a bank of 4thorder gammatone filters gb which have impulse responses ofthe form ^^^^^ = ^^^^^^^ cos^2^^^^^, ^ ≥ 0. For each center frequency ^^, the parameters ^and ^ are adjusted to realize the Equivalent Rectangular Bandwidth (ERB) and unit magnitude response at that frequency. The collection of center frequencies can be chosen with a 0.5 ERBspacing covering the frequency range from 20 Hz to the coded bandwidth. An example of ^ =62 resulting filter frequency responses for a coded bandwidth of 6750 Hz is given in figure 3.

[0048] A coding distortion signal ℎ^^^ can be defined as a difference between an original audio signal and a coded version of the same audio signal. A weighted distortion in the gammatone domain over the considered time duration ^ of length|^|is then defined by: 1 ( ^1^ ^ "^1^ where the weights %^^^^ in each gammatone band ^ are nonnegative functions of time. For convenience of presentation, a continuous time variable ^ is assumed. It is understood that integrals have to be replaced by sums over time samples of discrete time signals in the implementations. The symbol ‘*’ is used throughout this disclosure to denote time domain convolution.

[0049] The weights %^^^^ are based on a perceptual model, so that distortion contributions which are more easily perceived by a listener should have a higher weight. They may be determined using a signal specific masking level mb(t) representing a level at which a distortionwill be masked by the original signal. In other words, a signal distortion contribution ^^^ ∗ ℎ^$can be expected to be masked if it is smaller than +^^^^. This masking level will depend on both a hearing threshold in quiet, the original signal, and a perceptual model of the human ear. Anappropriate distortion weight definition can then be %^^^^ = 1 / +^^^^, leading to ^^ℎ^ ≤ 1 inEq. (1) for masked distortions.

[0050] In one embodiment, a starting point can be the analysis of the original input signal .to achieve an excitation pattern such as^^^^^ = max2^ ∗ ^^^ ∗ .^$^^^ , 3$^ 4 , ^2^where ^ is the impulse response of a low pass filter and 3^is the linear value of the threshold in quiet at the center frequency of gammatone band ^. A perceptually motivated masking model (such as the one discussed in International Publication No. WO2021 / 113416 titled “A Psychoacoustic Model for Audio Processing”, hereby incorporated by reference in its entirety) can then be used to convert the excitation pattern ^^^^^ into a masking level +^^^^ with the properties discussed above.

[0051] The linear filtering with ^ in Eq. (2) can be replaced by a nonlinear smoothingoperation on ^^^ ∗ .^$ aimed to align the temporal post-masking properties of the model withexperimental data. An example of this is to use envelope creation tools based on two different first order filters with different time constants for attack and decay, where the decay time constant depends on the gammatone filter index ^. Alignment correction

[0052] An improvement of the modelling of time domain masking aspects can be obtainedby replacing ^^^ ∗ ℎ^$ in Eq. (1) with a smoothed version of this quantity performed in a similarfashion as for the creation of the excitation pattern ^^^^^ from ^^^ ∗ .^$.

[0053] In the case of a linear envelope, where the smoothed version is defined by^ ∗ ^^^ ∗ ℎ^$, a convenient method taught by this disclosure to perform this alignment correctionis to filter the distortion weights %^^^^ by the time inverse ^^−^^ of the smoothing filter, while keeping Eq (1) unchanged.

[0054] In the case of a nonlinear envelope, the corresponding modification of the distortion weights involves the transpose of the gradient of the envelope operator, but a simpler alternative is to use the time inverse of the first order filter defined by the attack time constant. Weight conversion

[0055] As the transform control block 14 provides the applied frame lengths to the weight conversion block 18, the MDCT domain is given on the considered time interval ^. In other words, the distribution of time-frequency bins α is known.

[0056] The distortion signal can then be written ℎ = ∑^ &^7^ , where 7^ is the MDCTsynthesis waveform for a given bin ^, and &^is the quantization error for that bin. From a standard assumption that quantization errors are uncorrelated with zero mean, one arrives at the weighted distortion estimate ^^ℎ^ ≈ ! 9$^:^ , ^3^^where 9^$is the variance of the quantization error and :^are MDCT domain weights :1( ^= ! " ^^^ ∗ 7^^$^^^ %^^^^ &^ = ^^7^^ ^4^

[0057] In other bank domain bands, wherein each band term includes an integrated weighted contribution over the duration I, wherein the weighted contribution of a specific band includes a pre-computed factor and the weight function wb(t) of the band. The pre-computed factor includes a time- convolution between an analysis filter associated with the band and a transform waveform associated with the bin weight.

[0058] A relief in computational complexity is obtained by observing that ^^^ ∗ 7^^$depends only on the analysis filter bank (here, gammatone filter bank) and transform waveforms (here, MDCT waveforms). The convolution can therefore be approximated to high numerical precision by time shifts of a precomputed table of waveforms with length depending on transform size and the decay of the gammatone filters. Additionally, a downsampling of the summation implementing the integral of eq. (4) is possible when the weights are sufficiently slowly varying for an interpolated version of these to be a good approximation. Application to coding

[0059] The main free variables of MDCT domain coding are the quantization step sizes ∆ . The variance of the quantization error 9^$can be estimated based on the usual ∆$ / 12 rule forscalar quantization. Upon adaption to a companding quantizer with exponent > ≥ 1, one can use>$ $9^ = min A^^, B12 ∆ |^^|$^$$ $ BC . ^5^

[0060] Here, ^^is the estimate of the weightedstep size Δ = ΔH is then^H = ! 9$^:^ . ^6^

[0061] Estimates of the an observation that the Huffman are a Laplace distribution model, and that the codebook search can be modelled by an optimization of the Laplace scale parameter, leads to the following bit cost approximation *1 ^ B= F ^ . ^7^Rate distortion

[0062] The resulting estimates of total distortion ^ and total bit cost I are obtained by summing ^Hand IHover all (step size) subsets F under consideration. By neglecting details such that Huffman coding sections can be different from scale factor bands and disregarding the coding cost of the step sizes, one can then use gradient descent methods to control (e.g., pursue minimization of) the MDCT coding mean bitrate estimate I / |^| , given an upper bound on thedistortion ^ ≤ ^P.

[0063] The performance of this approach is illustrated in figures 4-6.

[0064] Figure 4 shows a conventional approach, where an audio signal 41 is transformed into an MDCT domain by banded MDCT analysis in block 42. The MDCT signal representation is represented by a time-frequency-diagram 43. The representation is used by block 44 to determine appropriate step sized for quantizing the coefficients in the representation 43.

[0065] Figure 5 shows an approach according to an implementation of the method in figure 2. Here, the audio signal 41 is subject to a gammatone analysis in block 52. The output from the gammatone analysis is represented as a time-frequency diagram 53. The output is used by block 54 to define a weighted distortion, which is converted by block 55 into the MDCT domain before being used in block 56 to determine step sizes.

[0066] Figure 6 shows MUSHRA test results for different categories of 48 kb / s mono signals (11 listeners). In figure 6, series 61 relates to conventional encoding according to figure 4, while series 62 relates to encoding according to figure 5. Consequences of high-rate theory and simplified step size control

[0067] It is well-known from rate distortion theory in the high-rate limit that the mean over coded dimensions of the resulting weighted distortion is constant. This result can also be recovered directly from explicit minimization using the high-rate versions of Eqs. (5) and (7) where the min and max operations are omitted and only the terms involving the step size are kept. Inthis case, there is a constant Q such that^H = Q |F|, ^8^where |F| is the dimension of a subset J of bins sharing the same step size Δ = ΔH .

[0068] With > = 1 in Eq. (5) one recovers a constant 9$ $ $^ = 9H = ∆H / 12 for ^ ∈ F. Asimplified interface to step size control is then obtained by choosing 9H$proportional to 1 / :H: 9$ SH =: , ^9^where : is a “group bin weight” defined subset J: 1:H =|F| ! :^ . ^10^^∈H

[0069] One can verify that Eq. (10) is consistent with the equidistribution of mean distortion as defined by Eq. (8). The step size for each group will now be determined by S∆$= 12 $H 9H = 12 .In other words, a simplified step size adjustment can now be achieved by selection of the constant of proportionality S. The scaling block 19 can then be configured to set each specific step size Δj based on the group weight and the constant E.

[0070] In this case, there is also an additional relief in computational complexity for weight conversion since the sum over F of Eq. (10) can be moved inside in Eq. (4) so precomputedtables are required only for the mean values of ^^^ ∗ 7^^$ over ^ ∈ F.1( 1:U =^|^| ! " !^^^ ∗ 7^^$^^^ % ^^^ &^^4′^ '|F| ^Calibration to a

[0071] The selection of distortion weights according to %^^^^ = 1 / +^^^^ in Eq. (1)results in a distortion ^^ℎ^ ≤ 1 if ^^^ ∗ ℎ^$ is bounded by +^^^^. In this case, ^^ℎ^ ≤ 1 alsoholds for any redistribution of the weights across gammatone bands %^^^^ = ^^ / +^^^^ with∑^ ^^ = ^. This offers an opportunity to select ^^ such that rate-distortion optimization based ondistortion approximates a desired noise shaping under the assumption of time of both the noise and the weights. The calibration of ^^can be obtained according to two methods. A heuristic based on gammatone filter bank bandwidth

[0072] To apply the high-rate theoretical prediction of optimized distortion distribution directly to the weighted distortion of eq. (1) one needs a heuristic for estimating coded dimensionality per band ^. This invention proposes to use the bandwidth BW^^^^ of thegammatone filter ^ . Then the result is tha Y^ t there is a constant Q such that1" ^^^ ∗ ℎ^$^^^ %^^^^ &^ = QY ∙ BW^^^^ ^11^'For a constant weight contribution ^^^ ∗ ℎ^$ of +^, this leads to ^^ = QY ∙ BW^^^^. Combined with the requirement∑^ ^^ = ^ this then leads toBW^^^^^^ = ^^ . ^12^ Calibration in the case of simplified

[0073] In the case where the step size is chosen inversely proportional to the mean of MDCT domain distortion weights as in Eqs. (9-10), the time mean value >^of the signaldistortion contribution ^^^ ∗ ℎ^$ resulting from a stationary weight %^ = ^^ / +^ can beestimated to \>^ = S^ ! F ^,H^ , ^13^where the MDCT band to ]\^,H = ! " ^^^ ∗ 7^^$^^^ &^. ^14^^∈H. ^]

[0074] Achieving >^ = ,but good approximations can sufficiently fast as a function of the difference in center frequencies of the gammatone bands andthe set of bins F. For calibration purposes, aiming at the existence of a single example with >^ =+^ is therefore worthwhile. In the case described below, this is obtained for a spectrally flat(white) noise shape.

[0075] If the gammatone squared magnitude frequency responses sum up to a constant and the MDCT synthesis waveforms are normalized to unit energy and cover the whole frequencyspectrum, one finds that ∑^ \^,H = |F| ∑ [ ‖^[‖$ and ∑H \^,H = ‖^^‖$ = _] $^] ^^^^^ &^ . Theseidentities can be applied in Eq. (13) with S = 1 to constants ‖^ $^‖^^ = ^ , ^15^it follows that >^ = +^ for any +^

[0076] In the assumed case where the gammatone squared magnitude frequency responses sum up to a constant, (as is approximately the case in Error! Reference source not found.3), there is a good agreement between the heuristic rule of Eq. (12) and the calibration of Eq. (15). Implementation Details

[0077] Figure 7 shows a schematic block diagram of an example electronic device or architecture 200 (e.g., an apparatus 200) suitable for implementing example embodiments of the present disclosure. The architecture 200 may form part of a mobile device such as a smartphone 2, but may also be a stand-alone piece of equipment. Architecture 200 may embody, but is not limited to, the system as described in relation to figure 1.

[0078] As shown, the architecture 200 includes central processing unit (CPU) 201 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 202 or a program loaded from, for example, storage unit 208 to random access memory (RAM) 203. The CPU 201 may be, for example, an electronic processor 201, which may include one or more processor cores, and in some examples the processor 201 may be multiple processors. In RAM 203, the data required when CPU 201 performs the various processes is also stored, as required. CPU 201, ROM 202 and RAM 203 are connected to one another via bus 204. Input / output (I / O) interface 205 is also connected to bus 204.

[0079] The following components are connected to I / O interface 205: input unit 206, that may include a keyboard, a mouse, or the like; output unit 207 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 208 including a hard disk, or another suitable storage device; and communication unit 209 which may include a network interface card such as a network card (e.g., wired or wireless).

[0080] In some embodiments, communication unit 209 is configured to communicate with other devices (e.g., via a network). Drive 210 is also connected to I / O interface 205, as required. Removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 210, so that a computer program read therefrom is installed into storage unit 208, as required. A person skilled in the art would understand that although apparatus 200 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.

[0081] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 209, and / or installed from the removable medium 211, as shown in figure 7.

[0082] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the various steps of figure 2 can be executed by control circuitry (e.g., CPU 201 in combination with other components of figure 7), thus, the control circuitry may be performing the actions described in this disclosure.

[0083] Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, a processor and / or other computing device(s), which may include control circuitry. While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in, as non- limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0084] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0085] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to one or more processors of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by one or more processors of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.

[0086] The one or more processors may operate as a standalone device or may be connected, e.g., networked to other processor(s). Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.

[0087] The software may be distributed on computer readable media, which may comprise computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to a person skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, physical (non-transitory) storage media in various forms, such as ROM, PROM, EPROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, it is well known to the skilled person that communication media (transitory) typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.

[0088] The implementation of the technologies disclosed in the figures are merely illustrative examples, and the invention is not so limited. For example, the illustrated partitions such as blocks in Figure 1 are merely illustrative logical partitions for ease of discussion, where such partitions may be split into additional partitions, combined into fewer partitions, supplemented with additional partitions, or reduced by eliminating partitions, without departing from the spirit of the present invention. For the illustrated flow chart of Figure 2, the partitions of the operational steps, which may be also referred to as functions, steps, operations, processes, or acts, may be combined into fewer steps or split into additional steps, where steps may be reordered or eliminated, in whole or in part, without departing from the spirit of this disclosure.

[0089] Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the disclosure discussions utilizing terms such as “processing”, “computing”, “calculating”, “determining”, “analyzing” or the like, may refer to the function, action, steps and / or processes of a computer hardware or computing system, or similar electronic computing devices, that manipulate and / or transform data represented as physical, such as electronic, quantities into other data similarly represented as physical quantities.

[0090] It should be appreciated that in the above description of example embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment. Thus, the claims following the Detailed Description are hereby expressly incorporated into this Detailed Description, with each claim standing on its own as a separate embodiment of this invention. Furthermore, while some embodiments described herein include some, but not other, features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and form different embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0091] Furthermore, some of the embodiments are described herein as a method or combination of elements of a method that can be implemented by a processor of a computer system or by other means of carrying out the function. Thus, a processor with instructions for carrying out such a method or element of a method forms a means for carrying out the method or element of a method. Note that when the method includes several elements, e.g., several steps, no ordering of such elements is implied, unless specifically stated. Furthermore, an element described herein of an apparatus embodiment is an example of a means for carrying out the function performed by the element for the purpose of carrying out the embodiments of the invention. In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the invention may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.

[0092] The person skilled in the art realizes that the present invention by no means is limited to the preferred embodiments described above. On the contrary, many modifications and variations are possible within the scope of the appended claims. For example, other transforms than MDCT may be used in the encoder, and other filter banks than gammatone may be used in the distortion analysis.

[0093] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs): EEE 1. An encoding method for encoding an audio signal, s(t), having a duration, I, into an encoded audio signal, the method comprising: transforming the audio signal using a frame-based frequency transform to obtain a transform domain representation of the audio signal; applying an analysis filter bank to the audio signal to obtain a filter bank domain representation of the audio signal, the filter bank domain having B frequency bands; defining a set of B weight functions, wb(t), wherein each weight function is related to a perception based distortion of the audio signal in one frequency band, b, of the filter bank domain; converting the set of B weight functions, wb(t), into a set of bin weights υαin said frame-based transform domain, where each bin weight is associated with a specific time-frequency bin of the frame-based transform domain; using the bin weights υαto calculate quantization step sizes ΔJconfigured to distribute an total number of bits, R, across time and frequency; and quantizing the transform domain representation of the audio signal using the quantization step sizes ΔJ. EEE 2. The method according to EEE 1, wherein the quantization step sizes are calculated by controlling the total number of bits, R, given a pre-defined limit on a weighted distortion measure based on the bin weights. EEE 3. The method according to EEE 1, wherein the quantization step sizes are calculated by controlling a weighted distortion measure based on the bin weights given a maximum available number of bits. EEE 4. The method according to any one of the preceding EEEs, wherein the frame-based frequency transform is configured to transform a first number of time samples in a frame to a second number of frequency samples, wherein the second number is smaller than the first number. EEE 5. The method according to any one of the preceding EEEs, wherein the frame-based transform domain has a first time-resolution and the analysis filter bank domain has a second time-resolution, wherein the second time-resolution is higher than the first time-resolution. EEE 6. The method according to any one of the preceding EEEs, wherein each weight function wb(t) is equal to a constant cbdivided by a masking function mb(t), which masking function mb(t) represents a level at which a distortion will be masked by the audio signal to make the distortion unperceivable by a human listener. EEE 7. The method according to EEE 6, wherein a sum of all constants cb is equal to the number of frequency bands, B. EEE 8. The method according to EEE 7, wherein each constant cb is determined as ^^= ^`a^bc^∑d `a^bd^ , where B is the number of frequency bands, BW(gb) is the bandwidth of filter bank∑[ BW^^[^ is a sum of all bandwidths of all filter bank filters. EEE The method according to EEE 8, wherein each constant cb is determined as ^^= ^‖bc‖edbd‖ wher ‖ ‖$∑ ‖ e e B is the number of frequency bands, ^^ is the sum of squared samples of theresponse of filter bank filter gb, and ∑ [ ‖^[‖$ is a sum of all squared impulse responsesof all filter bank filters. EEE 10. The method according to any one of the preceding EEEs, wherein each bin weight υα is formed as a sum of band terms over all filter bank domain bands, wherein each band term includes an integrated weighted contribution over the duration I, wherein the weighted contribution of a specific band includes a pre-computed factor and the weight function wb(t) of the band. EEE 11. The method according to EEE 9, wherein the pre-computed factor includes a time- convolution between an analysis filter associated with the band and a transform waveform associated with the bin weight. EEE 12. The method according to any one of EEEs 1-9, wherein the transform domain is divided into a set of scale-factor bands, and wherein all bins in a group J of bins belonging to a same scale-factor band have a common quantization step size ΔJ. EEE 13. The method according to EEE 12, wherein a variance 9H$of a quantization error for all bins in a group J of bins having a same quantization step size ΔJ is set to be proportional to an inverse of a group weight, : $ fH, according to: 9H = gh , where E is a constant of proportionality, and the group weight :His defined as an average of all bin weights υα associated with step size ΔJ::H = *|H| ∑ ^∈H :^ .The method according to EEE 13, wherein each bin group weight υj is formed as over all filter bank domain bands, wherein each term includes an integrated weighted contribution over the duration I, wherein the weighted contribution of a specific band includes a pre-computed factor and the weight function wb(t) of the band. EEE 15. The method according to EEE 14, wherein the pre-computed factor includes a sum of bin terms over all bins in the bin group, where each bin term includes a time-convolution between an analysis filter associated with the band and a transform waveform associated with the bin. EEE 17. An encoder for encoding an audio signal, s(t), having a duration, I, into an encoded audio signal, comprising: a transform block (11) for transforming the audio signal using a frame- based frequency transform to obtain a transform domain representation of the audio signal; an filter bank application block (16) for applying an analysis filter bank to the audio signal to obtain a filter bank domain representation of the audio signal, the filter bank domain having B frequency bands; a weight determination block (17) for defining a set of B weight functions, wb(t), wherein each weight function is related to a perception based distortion of the audio signal in one frequency band, b, of the filter bank domain; a weight conversion block (18) for converting the set of B weight functions, wb(t), into a set of bin weights υα in said frame-based transform domain, where each bin weight is associated with a specific time-frequency bin of the frame-based transform domain; a scaling block (19) for calculating quantization step sizes ΔJ based on the bin weights υα which step sizes are configured to distribute a total number of bits, R, across time and frequency; and a quantization block (15) for quantizing the transform domain representation of the audio signal using the quantization step sizes ΔJ. EEE 18. The encoder according to EEE 17, wherein the scaling block is configured to calculate the quantization step sizes by controlling the total number of bits, R, given a pre- defined limit on a weighted distortion measure based on the bin weights. EEE 19. The encoder according to EEE 17, wherein the scaling block is configured to calculate the quantization step sizes by controlling a weighted distortion measure based on the bin weights given a maximum available number of bits. EEE 20. The encoder according to any one of EEE 17-19, wherein the frame-based frequency transform is configured to transform a first number of time samples in a frame to a second number of frequency samples, wherein the second number is smaller than the first number. EEE 21. The encoder according to any one of EEEs 17-20, wherein the frame-based transform domain has a first time-resolution and the analysis filter bank domain has a second time- resolution, wherein the second time-resolution is higher than the first time-resolution. EEE 22. The encoder according to any one of EEEs 17-21, wherein each weight function wb(t) is equal to a constant cb divided by a masking function mb(t), which masking function mb(t) represents a level at which a distortion will be masked by the audio signal to make the distortion unperceivable by a human listener. EEE 23. The encoder according to EEE 22, wherein a sum of all constants cb is equal to the number of frequency bands, B. EEE 24. The encoder according to EEE 23, wherein each constant cb is determined as ^^= ^`a^bc^∑ , where B is the number of frequency bands, BW(gbd `a^bd^ ) is the bandwidth of filter bankfilter gb, and ∑[ BW^^[^ is a sum of all bandwidths of all filter bank filters.EEE The encoder according to EEE 23, wherein each constant cb is determined as ^^= ^bce ‖‖where B is the number of frequency band ‖ ‖$∑ d ‖bd‖e s, ^^ is the sum of squared samples of theof filter bank filter gb, and ∑ [ ‖^[‖$ is a sum of all squared impulse responses of all filter bank filters. EEE 26. The encoder according to any one of EEEs 17-25, wherein the weight conversion block is configured to form each bin weight υα as a sum of band terms over all filter bank domain bands, wherein each band term includes an integrated weighted contribution over the duration I, wherein the weighted contribution of a specific band includes a pre-computed factor and the weight function wb(t) of the band. EEE 27. The encoder according to EEE 26, wherein the pre-computed factor includes a time-convolution between an analysis filter associated with the band and a transform waveform associated with the bin weight. EEE 28. The encoder according to any one of EEEs 17-25, wherein the transform domain is divided into a set of scale-factor bands, and wherein all bins in a group J of bins belonging to a same scale-factor band have a common quantization step size ΔJ. EEE 29. The encoder according to EEE 28, wherein a variance 9H$of a quantization error for all bins in a group J of bins having a same quantization step size ΔJ is set to be proportional to an inverse of a group weight, :H, according to: 9$H = fgh , where E is a constant of proportionality,and the group weight :His defined as an average of all bin weights υα associated with step size ΔJ: :*H = ∑ ^∈H : . EEE 30. The encoder according to EEE 29, wherein the scaling block is configured to set the square ∆H$of the specific step size to be proportional to the inverse of the group weight. EEE 31. The encoder according to EEE 29 or 30, wherein the weight conversion block is configured to form each bin group weight υj as a sum of terms over all filter bank domain bands, wherein each term includes an integrated weighted contribution over the duration I, wherein the weighted contribution of a specific band includes a pre-computed factor and the weight function wb(t) of the band. EEE 32. The encoder according to EEE 31, wherein the pre-computed factor includes a sum of bin terms over all bins in the bin group, where each bin term includes a time-convolution between an analysis filter associated with the band and a transform waveform associated with the bin. EEE 33. An apparatus comprising a processor and a memory coupled to the processor, wherein the processor is adapted to cause the apparatus to carry out the method according to any one of claims 1 - 16. EEE 34. A computer-readable storage medium storing a program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1 - 16.

Claims

CLAIMS 1. An encoding method for encoding an audio signal, s(t), having a duration, I, into an encoded audio signal, the method comprising: transforming the audio signal using a frame-based frequency transform to obtain a transform domain representation of the audio signal; applying an analysis filter bank to the audio signal to obtain a filter bank domain representation of the audio signal, the filter bank domain having B frequency bands; defining a set of B weight functions, wb(t), wherein each weight function is related to a perception based distortion of the audio signal in one frequency band, b, of the filter bank domain; converting the set of B weight functions, wb(t), into a set of bin weights υαin said frame- based transform domain, where each bin weight is associated with a specific time-frequency bin of the frame-based transform domain; using the bin weights υα to calculate quantization step sizes ΔJ configured to distribute an total number of bits, R, across time and frequency; and quantizing the transform domain representation of the audio signal using the quantization step sizes ΔJ.

2. The method according to claim 1, wherein the quantization step sizes are calculated by controlling the total number of bits, R, given a pre-defined limit on a weighted distortion measure based on the bin weights.

3. The method according to claim 1, wherein the quantization step sizes are calculated by controlling a weighted distortion measure based on the bin weights given a maximum available number of bits.

4. The method according to any one of the preceding claims, wherein the frame-based frequency transform is configured to transform a first number of time samples in a frame to a second number of frequency samples, wherein the second number is smaller than the first number.

5. The method according to any one of the preceding claims, wherein the frame-based transform domain has a first time-resolution and the analysis filter bank domain has a second time-resolution, wherein the second time-resolution is higher than the first time-resolution.

6. The method according to any one of the preceding claims, wherein each weight function wb(t) is equal to a constant cbdivided by a masking function mb(t), which masking function mb(t) represents a level at which a distortion will be masked by the audio signal to make the distortion unperceivable by a human listener.

7. The method according to claim 6, wherein a sum of all constants cb is equal to the number of frequency bands, B.

8. The method according to claim 7, wherein each constant cb is determined as BW^^ ^^^ = ^ ^∑ ,[ BW^^[^where B is the number ofis the bandwidth of filter bank filtergb, and ∑[ BW^^[^ is a sum of all bandwidths of all filter bank filters.The method according to claim 8, wherein each constant cb is determined as ‖^ $^‖^^ = ^where B is the number ofis the sum of squared samples of theimpulse response of filter bank filter gb, and ∑ [ ‖^[‖$ is a sum of all squared impulse responsesof all filter bank filters.

10. The method according to any one of the preceding claims, wherein each bin weight υα is formed as a sum of band terms over all filter bank domain bands, wherein each band term includes an integrated weighted contribution over the duration I, wherein the weighted contribution of a specific band includes a pre-computed factor and the weight function wb(t) of the band.

11. The method according to claim 10, wherein the pre-computed factor includes a time-convolution between an analysis filter associated with the band and a transform waveform associated with the bin weight.

12. The method according to any one of claims 1-9, wherein the transform domain is divided into a set of scale-factor bands, and wherein all bins in a group J of bins belonging to a same scale-factor band have a common quantization step size ΔJ.

13. The method according to claim 12, wherein a variance 9H$of a quantization error for all bins in a group J of bins having a same quantization step size ΔJ is set to be proportional to an inverse of a group weight, :H, according to: $S9H =: ,Hwhere E is a constant ofthe group weight :His defined as an average of all bin weights υαassociated with step size ΔJ: :1H =|F| ! :^ .^∈H14. The method according to claim 13, wherein the square ∆H$of the specific step size is selected to be proportional to the inverse of the group weight 15. The method according to claim 13 or 14, wherein each bin group weight υjis formed as a sum of terms over all filter bank domain bands, wherein each term includes an integrated weighted contribution over the duration I, wherein the weighted contribution of a specific band includes a pre-computed factor and the weight function wb(t) of the band.

16. The method according to claim 15, wherein the pre-computed factor includes a sum of bin terms over all bins in the bin group, where each bin term includes a time-convolution between an analysis filter associated with the band and a transform waveform associated with the bin.

17. An encoder for encoding an audio signal, s(t), having a duration, I, into an encoded audio signal, comprising: a transform block (11) for transforming the audio signal using a frame-based frequency transform to obtain a transform domain representation of the audio signal; an filter bank application block (16) for applying an analysis filter bank to the audio signal to obtain a filter bank domain representation of the audio signal, the filter bank domain having B frequency bands; a weight determination block (17) for defining a set of B weight functions, wb(t), wherein each weight function is related to a perception based distortion of the audio signal in one frequency band, b, of the filter bank domain;a weight conversion block (18) for converting the set of B weight functions, wb(t), into a set of bin weights υαin said frame-based transform domain, where each bin weight is associated with a specific time-frequency bin of the frame-based transform domain; a scaling block (19) for calculating quantization step sizes ΔJ based on the bin weights υα which step sizes are configured to distribute a total number of bits, R, across time and frequency; and a quantization block (15) for quantizing the transform domain representation of the audio signal using the quantization step sizes ΔJ.

18. The encoder according to claim 17, wherein the scaling block is configured to calculate the quantization step sizes by controlling the total number of bits, R, given a pre- defined limit on a weighted distortion measure based on the bin weights.

19. The encoder according to claim 17, wherein the scaling block is configured to calculate the quantization step sizes by controlling a weighted distortion measure based on the bin weights given a maximum available number of bits.

20. The encoder according to any one of claims 17-19, wherein the frame-based frequency transform is configured to transform a first number of time samples in a frame to a second number of frequency samples, wherein the second number is smaller than the first number.

21. The encoder according to any one of claims 17-20, wherein the frame-based transform domain has a first time-resolution and the analysis filter bank domain has a second time-resolution, wherein the second time-resolution is higher than the first time-resolution.

22. The encoder according to any one of claims 17-21, wherein each weight function wb(t) is equal to a constant cbdivided by a masking function mb(t), which masking function mb(t) represents a level at which a distortion will be masked by the audio signal to make the distortion unperceivable by a human listener.

23. The encoder according to claim 22, wherein a sum of all constants cb is equal to the number of frequency bands, B.

24. The encoder according to claim 23, wherein each constant cb is determined as BW^^ ^^ = ^^ ^ ,where B is the number of frequency bands, BW(gb) is the bandwidth of filter bank filter gb, and∑[ BW^^[^ is a sum of all bandwidths of all filter bank filters.

25. The encoder according to claim 23, wherein each constant cbis determined as ‖^ ‖$^ = ^ ^^ ∑ [ ‖^[‖$where B is the number of frequency sum of squared samples of the impulseresponse of filter bank filter g , and ∑b[ all squared impulse responses of allfilter bank filters.

26. The encoder according to any one of claims 17-25, wherein the weight conversion block is configured to form each bin weight υαas a sum of band terms over all filter bank domain bands, wherein each band term includes an integrated weighted contribution over the duration I, wherein the weighted contribution of a specific band includes a pre-computed factor and the weight function wb(t) of the band.

27. The encoder according to claim 26, wherein the pre-computed factor includes a time-convolution between an analysis filter associated with the band and a transform waveform associated with the bin weight.

28. The encoder according to any one of claims 17-25, wherein the transform domain is divided into a set of scale-factor bands, and wherein all bins in a group J of bins belonging to a same scale-factor band have a common quantization step size ΔJ.

29. The encoder according to claim 28, wherein a variance 9H$of a quantization error for all bins in a group J of bins having a same quantization step size ΔJis set to be proportional to an inverse of a group weight, :H, according to: $S9H =: ,Hwhere E is a constant of proportionality, and the group weight :His defined as an average of all bin weights υαassociated with step size ΔJ: 1= ! :^ .

30. The encoder according to claim 29, wherein the scaling block is configured to set the square ∆H$of the specific step size to be proportional to the inverse of the group weight.

31. The encoder according to claim 29 or 30, wherein the weight conversion block is configured to form each bin group weight υjas a sum of terms over all filter bank domain bands, wherein each term includes an integrated weighted contribution over the duration I, wherein the weighted contribution of a specific band includes a pre-computed factor and the weight function wb(t) of the band.

32. The encoder according to claim 31, wherein the pre-computed factor includes a sum of bin terms over all bins in the bin group, where each bin term includes a time-convolution between an analysis filter associated with the band and a transform waveform associated with the bin.

33. An apparatus comprising a processor and a memory coupled to the processor, wherein the processor is adapted to cause the apparatus to carry out the method according to any one of claims 1 - 16.

34. A computer-readable storage medium storing a program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1 - 16.

Citation Information

Patent Citations

  • A psychoacoustic model for audio processing

    WO2021113416A1

  • Quality improvement techniques in an audio encoder

    US20090326962A1