Level-dependent channel bit distribution
The method addresses the challenge of bit allocation in multi-channel audio codecs by using a level-dependent psychoacoustic model to adjust perceptual entropy, resulting in improved coding efficiency and maintained audio quality for immersive audio content.
Patent Information
- Application Number
- PCT/EP2024/083738
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-01
- Filing Date
- 2024-11-27
- Publication Date
- 2025-06-05
AI Technical Summary
Existing multi-channel audio codecs face challenges in efficiently allocating bits across channel elements while maintaining reasonable audio quality, especially in immersive audio content.
A method for determining bit allocation for multiple channel elements in audio content using a level-dependent psychoacoustic model, which calculates perceptual entropy for each channel element and adjusts it based on target bits and aggregated perceptual entropy, ensuring that louder channels are coded with higher precision.
This approach achieves higher coding efficiency for multi-channel immersive audio content by distributing bits in a perceptually meaningful way, maintaining audio quality while adhering to bitrate constraints.
Smart Images

Figure EP2024083738_05062025_PF_FP_ABST
Abstract
Description
LEVEL-DEPENDENT CHANNEL BIT DISTRIBUTION Cross-Reference to Related Applications
[0001] This application claims the benefit of priority from U.S. Provisional Application Ser. No.63 / 604,361, filed on 30 November 2023 and European Patent Application 24160941.1 filed 1 March 2024 which are incorporated herein by reference. Technical Field
[0002] The present disclosure relates to techniques for multi-channel audio coding. In particular, the present disclosure describes methods and apparatus for determining a bit allocation for a plurality of channel elements of audio content. Background
[0003] Bitrate control is an essential part of any multi-channel audio codec and typically determines how the available bits should be spent across (time) frames and / or across frequency bands. Low memory footprint immersive audio codecs especially may require more efficient allocation of bits so as to satisfy both bitrate constraints and reasonable audio quality.
[0004] Thus, there is a need for improved techniques for bit allocation in multi- channel audio content, especially immersive audio content. There is a particular need for such techniques that do not degrade perceived audio quality. Summary
[0005] In view of this need, the present disclosure provides methods of determining a bit allocation for a plurality of channel elements of audio content as well as corresponding apparatus, programs, and computer-readable storage media, having the features of the respective independent claims.
[0006] An aspect of the present disclosure relates to a method of determining a bit allocation for a plurality of channel elements of audio content. The audio content may be frame-based, channel-based, or object-based, for example. The method may include, for a frame (e.g., time frame) of the audio content, determining a measure of perceptual entropy for each of the plurality of channel elements in the frame. Here, a channel element may relate to asingle channel (including channels of audio objects) or a pair of channels. The measure of perceptual entropy of a channel element may be based on or correspond to a minimum number of bits for transparent quality coding of the channel element, for example. The method may further include adjusting the determined measures of perceptual entropy based on a number of target bits for the frame and a measure of aggregated perceptual entropy of the plurality of channel elements in the frame. For example, the adjusted measure of perceptual entropy of a channel element may correspond to a number of bits to be allocated to that channel element. The method may yet further include performing bit allocation for the plurality of channel elements based on the adjusted measures of perceptual entropy. The method may further include determining the bit allocation for the plurality of channel elements based on the adjusted measures of perceptual entropy.
[0007] Configured as described above, the proposed method can distribute bits to audio channels (audio channel elements) in an audio frame in a perceptually meaningful way based on a level-dependent psychoacoustic model. For example, audio channels with higher energy (i.e., louder channels) may be coded at higher precision (e.g., with more bits) than those with lower energy (i.e., less loud or fainter channels). This intuitive approach is backed- up and guided by the level-dependent psychoacoustic model. In consequence, the proposed method can achieve higher coding efficiency especially for multi-channel immersive audio content.
[0008] In some embodiments, the measure of perceptual entropy for a given channel element may be positively correlated with the energy of the given channel element. Further, the amount of bits to be allocated to a given audio channel element may be positively correlated with the measure of perceptual entropy. Thereby, louder channels may be coded (e.g., quantized) with higher precision, whereas less loud channels may be coded with less precision.
[0009] In some embodiments, the measure of aggregated perceptual entropy of the plurality of channel elements in the frame may be based on (or correspond to) a sum of the determined measures of perceptual entropy for the plurality of channel elements in the frame. In some embodiments, determining the measure of perceptual entropy for a given channel element may include determining, for each of a plurality of frequency bands, an energy per band of the given channel element in the respective frequency band based on a frequency- domain representation of the given channel element in the frame. Without intended limitation, the frequency representation may relate to a Modified Discrete Cosine Transform (MDCT) of the channel element, or to a complex modulated filterbank transform, such as a QuadratureMirror Filter (QMF) transform or Discrete Fourier Transform (DFT), for example. The frequency bands may be scale factor bands, for example. The energy per band may be adjusted based on playback level changes (known at the encoder), for example for Low Frequency Enhancement (LFE). Determining the measure of perceptual entropy may further include determining, for each of the plurality of frequency bands, a number of bits per band for the respective frequency band based on the energy per band of the given channel element. The number of bits per band may be the minimum number of bits required for transparent quality coding of the channel element in the respective frequency band, for example. Determining the measure of perceptual entropy may yet further include determining the measure of perceptual entropy based on the numbers of bits per band for the plurality of frequency bands.
[0010] In some embodiments, the measure of perceptual entropy for a given channel element may be based on a sum, over the plurality of frequency bands, of the numbers of bits per band for the given channel element.
[0011] In some embodiments, determining the number of bits per band may be further based on a threshold in quiet for the respective frequency band. As can be understood and appreciated by the skilled person, the “threshold in quiet”, for example in the general technical context of an audio codec, may be used to generally represent the sound signal energy per frequency band above which a sound signal (at this frequency band) just becomes audible for an average human listener. In other words, sound signals at or below the threshold in quiet may be considered as irrelevant and receive no or fewer bits from a bit distribution algorithm. Determining the number of bits per band may be further based on a number of perceptually relevant bins per band of the frequency representation of the channel element for the respective frequency band. In some embodiments, determining the number of bits per band may be further based on an estimate of tonality for the respective frequency band.
[0012] In some embodiments, determining the number of bits per band may be based on a level-dependent psychoacoustic model.
[0013] In some embodiments, the number of bits per band may be set to at least a predetermined minimum number of bits per band. In other words, the method may ensure that at least a predetermined minimum number of bits are considered for each frequency band, regardless of whether the number of bits per band is lower than this predetermined minimum number.
[0014] In some embodiments, the number of bits per band ^^,^for channel element ^ and frequency band ^ may be given by ^ ^^,^^,^ = max ^^^^^,^ log^^ ^max ^^^ , 1.0^^ , ^^^^ ^,!^is a threshold in quiet for frequency band ^, ^^is a band-dependent number, ^^^^is a predetermined minimum number of bits per band, and ^^,^is a number of perceptually relevant bins of the frequency representation of the given channel element in frequency band ^. For example, ^^may be a number dependent on a measure of tonality in that band.
[0015] In some embodiments, adjusting the perceptual entropy for a given channel element may be further based on a number of perceptually relevant bins of the frequency representation of the given channel element. The bins of the frequency representation of the channel element may correspond to MDCT lines, for example.
[0016] In some embodiments, adjusting the perceptual entropy for a given channel element may be further based on a total number of perceptually relevant bins of the frequency representations of the plurality of channel elements.
[0017] In the simplest form, the number of perceptually relevant bins may be set to the total number of bins in the frequency band.
[0018] In some embodiments, the adjusted measure "^#of perceptual entropy for a channel element ^ may be given by " #^ = max$" #^ − &^^ , " ^^^ ', where " ^ is the measureof perceptual entropy for channel element ^, & is a constant based on a total number of perceptually relevant bins of the frequency representations of the plurality of channel elements in the frame, a total perceptual entropy of the frame, and the number of target bits for the frame, ^^is a number of perceptually relevant bins of the frequency representation of channel element ^, and "^#^^is a minimum number (minimum value) for the adjusted measure of perceptual entropy. The minimum number "^#^^ for the adjusted measure of perceptual entropy may be 0, for example.
[0019] In some embodiments, the plurality of channel elements in the frame may be processed sequentially. Further, in some implementations, the plurality of channel elements may be processed independently of each other.
[0020] In some embodiments, the method may sequentially operate on frames (e.g., time frames) of the audio content. Then, the method may further include buffering the adjusted measures of perceptual entropy between frames. Additionally or alternatively, the method may further include smoothing the adjusted measures of perceptual entropy between frames using a low pass filter. The low-pass filter may be a first order recursive filter, forexample. The low-pass filtering over time frames may be applied, for example, only for declining adjusted measures of perceptual entropy.
[0021] In some embodiments, the method may further include encoding the frame of the audio content based on the adjusted measures of perceptual entropy and the number of target bits. The encoding may involve quantizing of frequency domain coefficients of channel elements, for example.
[0022] In some embodiments, a number of bits ^^()^,*to be allocated per channelelement ^ in frame + of the audio content may be given by ^^()^,* = ,^,*^^()*, where ,^,* isa ratio based on the adjusted measure of perceptual entropy for channel element ^ in frame + and ^^()*is the number of target bits for frame +. In some embodiments, the method may further include encoding a subsequent frame of the audio content based on the adjusted measures of perceptual entropy determined for the frame and the number of target bits for the subsequent frame. The encoding may involve quantizing of frequency domain coefficients of channel elements, for example.
[0023] In some embodiments, a number of bits ^^()^,*to be allocated per channelelement ^ in frame + of the audio content may be given by ^^()^,* = ,^,*-^^^()*, where,^,*-^is a ratio based on the adjusted measure of perceptual entropy for channel element ^ inframe + − 1 and ^^()* is the number of target bits for frame +.
[0024] When using this coding mode, each audio channel element of the subsequent frame may be encoded at once, before the perceptual entropies of the remaining audio channel elements in the subsequent frame are determined. Thereby, the audio channels in the subsequent frame can be sequentially and independently processed, leading to significant savings in memory footprint.
[0025] According to another aspect, an apparatus for determining a bit allocation for a plurality of channel elements of audio content is provided. The apparatus may include a processor and a memory coupled to the processor and storing instructions for the processor. The processor may be configured to perform all steps of the methods according to the preceding aspect and its embodiments.
[0026] According to another aspect, a computer program is described. The computer program may comprise executable instructions for performing the methods or method steps outlined throughout the present disclosure when executed by a computing device. According to yet another aspect, a computer-readable storage medium is described. The storage medium may store a computer program adapted for execution on a processor and forperforming the methods or method steps outlined throughout the present disclosure when carried out on the processor.
[0027] It should be noted that the methods and systems including its preferred embodiments as outlined in the present disclosure may be used stand-alone or in combination with the other methods and systems disclosed in this document. Furthermore, all aspects of the methods and systems outlined in the present disclosure may be arbitrarily combined. In particular, the features of the claims may be combined with one another in an arbitrary manner.
[0028] It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus, and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) (and, e.g., their steps) are understood to likewise apply to the corresponding apparatus (and, e.g., their blocks, stages, units), and vice versa. Brief Description of the Drawings
[0029] The invention is explained below in an exemplary manner with reference to the accompanying drawings, wherein
[0030] Fig.1 is a block diagram schematically illustrating a framework for memory- efficient level-dependent channel bit allocation according to embodiments of the disclosure;
[0031] Fig.2 is a flowchart illustrating an example method of determining a bit allocation for a plurality of channel elements of audio content according to embodiments of the disclosure;
[0032] Fig.3 is a flowchart illustrating details of a step of the method of Fig.2 according to embodiments of the disclosure;
[0033] Fig.4 and Fig.5 are flowcharts illustrating different modes of encoding audio frames according to embodiments of the disclosure; and
[0034] Fig.6 is a block diagram of an example of an apparatus for performing methods according to embodiments of the disclosure. Detailed Description
[0035] In general terms, the present disclosure relates to techniques (e.g., methods and apparatus, such as encoders) for allocating bits to audio channels (or audio channel elementsin general) based on a level-dependent psychoacoustic model in a perceptually meaningful way. The present disclosure further relates to a memory-efficient implementation thereof. Implementations of the disclosed techniques relate to channel bit distribution based on a level-dependent psychoacoustic model for low-delay immersive audio codecs. One non- limiting example of a suitable psychoacoustic model that may be used for the present purposes is given in U.S. Patent Application, Publication No.2022 / 0415334 A1, “A Psychoacoustic Model for Audio Processing,” published December 29, 2022, which is hereby incorporated in its entirety.
[0036] Bitrate control is an essential part of any multi-channel audio codec and typically determines how the available bits are spent across (time) frames and / or across frequency bands. Bit allocation over time frames may also simply be static in case of constant bitrate requirements.
[0037] The present disclosure describes a perceptually motivated approach for spatial bit allocation, i.e., across audio channels, using a level-dependent psychoacoustical model. In other words, the disclosure relates to distributing bits in an audio frame in a perceptually meaningful way based on a level-dependent psychoacoustic model. The underlying idea is that audio channels with higher energy (i.e., louder channels) shall be coded at higher precision (e.g., with more bits) than those with lower energy (i.e., less loud or fainter channels). This intuitive approach is backed-up and guided by the level-dependent psychoacoustic model. Notably, the psychoacoustic model used for channel bit distribution may differ from the psychoacoustic model(s) that is / are used for bit allocation over frequency bands and / or over time frames.
[0038] In techniques according to the present disclosure, the desired number of frame bits (i.e., a number of target bits for the frame, or frame bit budget) is given. Then the frame bits are distributed to audio channels. Finally, the channel bits may be allocated over frequency bands in accordance with known techniques, for example techniques given in 3GPP TS 26.403.
[0039] In the following, an audio encoder is considered that encodes channel elements, which may be either single channels (including channels of audio objects) or pairs of channels. For this type of audio encoder, the number of target bits per channel element is estimated. This may be done based on energy estimates that are input to a psychoacoustic model. These energy estimates may be either the initial audio input energies or, in case of channel pairs with joint stereo coding, the energies of the mid and side signal, respectively, for example.
[0040] For such audio encoder, the present disclosure foresees two different coding (or quantization) modes, namely a first mode that seeks to improve memory efficiency and in which knowledge on a given audio frame is used for encoding channel elements of a subsequent (e.g., next) audio frame (memory efficient mode or chain mode), and a second mode in which the encoder holds data associated with all channel elements of a given frame at the same time for encoding the channel elements of the given frame (sequential mode). In the context of the present disclosure, immersive audio content is understood to relate to multi-channel audio content, where channels are associated to fixed or flexible loudspeaker layouts including non-horizontal positions or audio objects defined in at least three dimensions. Immersive audio content may thus relate to spatial multi-channel and object audio content.
[0041] Next, examples of basic processing steps / blocks of techniques according to the present disclosure will be described in more detail with reference to the block diagram of Fig.1.
[0042] In this framework 100, a memory efficient encoder may sequentially operate on single channel elements (e.g., channels or channel pairs) and mainly store the data associated with the currently processed channel element. While Fig.1 relates to this memory- efficient implementation (memory efficient mode, chain mode), it is understood that this is without intended limitation and that the channel elements may also be processed simultaneously, etc., in some implementations.
[0043] The incoming audio 105 is organized in time frames and transformed into the frequency domain using a time-frequency transform at time-frequency transform block 110. Frames + are sequentially processed. Without intended limitation, the time-frequency transform may be a MDCT or a complex modulated filterbank transform, such as a Quadrature Mirror Filter (QMF) transform or Discrete Fourier Transform (DFT), for example. While frequent reference may be made to MDCT in the following, this is understood to be without intended limitation.
[0044] Based on the resulting frequency-domain representation 115 (e.g., MDCTcoefficients) of a given channel element ^ = 1, … , / in the current audio frame +, energies(energies per band) are computed in frequency bands at block 120. For example, the frequency bands may correspond to the scale factor bands that determine the quantization precision.
[0045] In case when a channel pair is processed, the energies in the scale factor bands associated with the two channels, or channels derived therefrom, e.g., by (joint) stereo codingsuch as mid / side transform, are computed. Then, the energies in the scale factor bands associated with both channels may be added to obtain energies in the scale factor bands associated with the pair of channels, for example.
[0046] The energies together with the threshold in quiet 135 are input to block 130 implementing a level-dependent psychoacoustic model. This psychoacoustic model determines the Signal-to-Mask (SMR) ratio in every frequency band. From the SMR, the number of required bits to encode at just transparent audio quality, in the following referred to as Perceptual Entropy (PE), is computed. Transparent quality means that despite present distortions, no difference to the uncoded audio can be perceived.
[0047] At this, when computing the PE, a minimum number of bits per band may be assigned. Further, relevant playback level changes (known to the encoder), for example a 10 dB boost for the Low Frequency Enhancement (LFE) channel, may be considered in the PE computation.
[0048] The determined PEs are subsequently stored in a buffer 140 which is persistent between audio frames.
[0049] In the memory efficient mode, target bits 162 for the current channel element are computed at block 160 based on the (previous frame’s) PE, for example based on the previous frame’s PE ratio 155 (e.g., ratio with respect to a total of previous frame’s PE) for the current channel element multiplied by the current number of bits per frame 165. In some implementations, the computation of target bits may not only be based on the previous frame’s PE (or PE ratio), but also on the PE (or PE ratio) of one or more frames preceding the previous frame. For example, the PE (or PE ratio) used for the computation may be obtained by averaging over a number of frames preceding the current frame, or by any other suitable means of combining the PE (or PE ratio) of these frames for the current channel element (such as by smoothing, filtering, weighting, etc.).
[0050] The target bits for the current channel element are then distributed over frequency bands at bit allocation block 170, and frequency-domain coefficients (e.g., MDCT coefficients) are quantized and encoded accordingly at quantization / encoding block 180. Encoded versions of the frequency-domain coefficients, in suitable format, are then output to bitstream 190.
[0051] Once the last channel element (^ == / ) in a frame has been processed, thestored PEs are adjusted at PE adjustment block 150 such that the aggregated (e.g., total) PE equals the number of bits per frame 152. When adjusting the PEs, the number of relevantfrequency-domain lines (e.g., MDCT lines) per channel element may be considered. Then, PE ratios may be computed by dividing the adjusted PEs by the number of bits per frame. When not using the memory efficient mode, bit allocation for each channel element may be based on the adjusted PE or PE ratio for that channel element in the current frame. In this case, bit allocation and quantization / encoding can only be performed after all channel elements in the current frame have been processed, since the aggregated PE is required for adjusting the channel elements’ PEs.
[0052] Fig.2 is a flowchart illustrating an example of a method 200 of determining a bit allocation for a plurality of channel elements of audio content in accordance with the above. Method 200 includes steps S210 through S230.
[0053] The audio content may be frame-based, channel-based, or object-based, for example.
[0054] At step S210, for a frame (e.g., time frame) of the audio content, a measure of perceptual entropy is determined for each of the plurality of channel elements in the frame. For example, the measure of PE for a given channel element may be positively correlated with the energy of the given channel element in the frame.
[0055] At step S220, the determined measures of perceptual entropy are adjusted based on a number of target bits for the frame and a measure of aggregated perceptual entropy of the plurality of channel elements in the frame.
[0056] Therein, the measure of aggregated perceptual entropy of the plurality of channel elements in the frame may be based on a sum of the determined measures of perceptual entropy for the plurality of channel elements in the frame, for example. At step S230, bit allocation for the plurality of channel elements is performed based on the adjusted measures of perceptual entropy.
[0057] This may relate to determining the bit allocation for the plurality of channel elements based on the adjusted measures of perceptual entropy, for output to a quantization and / or encoding stage. For example, the adjusted measure of perceptual entropy of a channel element may correspond to a number of bits to be allocated to that channel element. Fig.3 is a flowchart illustrating an example implementation 300 of step S210 (i.e., determining the measure of perceptual entropy) in method 200 for a given channel element. Example implementation 300 comprises steps S310 through S330.
[0058] At step S310, for each of a plurality of frequency bands, an energy^per band of the given channel element in the respective frequency band is determined based on afrequency-domain representation of the given channel element in the frame. The aforementioned frequency representation may relate to an MDCT of the channel element, for example. The frequency bands may be scale factor bands, for example. For example, energy^in band ^ may be computed from the frequency-domain representation (e.g., MDCT coefficients) 0^of audio channel (element) ^ as ^4= 1where 2^and 2^are the(e.g., bins / lines corresponding to lower and upper limits of the band ^).
[0059] The energy^per band may be adjusted based on playback level changes (known at the encoder), for example for the Low Frequency Enhancement (LFE) channel, which may be typically played back with a 10 dB boost, or for audio objects with associated gains.
[0060] At step S320, for each of the plurality of frequency bands ^, a number of bits per band ^^,^for the respective frequency band ^ is determined based on the energy per band of the given channel element ^. This number of bits per band ^^,^may be the minimum number of bits per band required for transparent quality coding of the channel element ^ in the respective frequency band ^, for example.
[0061] In general, the determination of the number of bits per band may be based on a level-dependent psychoacoustic model. Examples are given below.
[0062] In addition to the energy per band^,^, the number of bits per band ^^,^may be further based on a threshold in quiet !^for the respective frequency band ^. As mentioned earlier, as can also be understood and appreciated by the skilled person, the “threshold in quiet”, for example in the general technical context of an audio codec, may be used to generally represent the sound signal energy per frequency band above which a sound signal (at this frequency band) just becomes audible for an average human listener. In other words, sound signals at or below the threshold in quiet may be considered as irrelevant and receive no or fewer bits from a bit distribution algorithm.
[0063] Also, the number of bits per band ^^,^may be further based on a number of perceptually relevant bins (e.g., MDCT lines) ^^,^per band of the frequency representation of the channel element ^ for the respective frequency band ^.In line with the above, the number of bits per band ^^,^may be given, for example, by ^= max ^ ^ log ^,^^,^ ^ ^ ^,^ ^^ 9max :! , 1.0;< , ^^^^ ^^where^,^is the energy per band of channel element ^ inper Eq. (1). Further, ^^,^is a number of perceptually relevant bins (e.g., MDCT lines) of the frequency representation of the given channel element ^ in frequency band ^. ^^is a band- dependent number and ^^^^is a predetermined minimum number of bits per band.
[0064] In Eq. (2), the estimate of bits per line (e.g., the log-term multiplied by ^^) is multiplied by the number of relevant bins (e.g., lines) ^^,^in the band to compute the level- dependent PE for this band.
[0065] In one non-limiting example, the number of perceptually relevant bins (e.g., relevant MDCT lines) ^^,^may be given by ^,^^ = = > 1.0: ^^ − ^^ + 1^(3) where ^^and ^^are bin numbers of an upper limit bin and a lower limit bin, respectively, ofthe frequency band, so that ^^ − ^^ + 1 gives the total number of bins in the frequency band.In alternative implementations, the number of perceptually relevant bins ^^,^may be given by the number of bins for which the (magnitude of the) corresponding frequency-domain coefficient (e.g., MDCT coefficient) is greater than a threshold, for example.
[0066] Band-dependent number (e.g., band-dependent constant) ^^may be a value dependent on a measure of tonality in that band. Typically, ^^is in the order of ^^∼10.0 ^.3IJ = 3.IJ . Here, the factor 10.0 is used to calculate the sound level above threshold inquiet in dB. The required Signal-to-Mask ratio (SMR) at just transparent quality according to the level dependent psycho-acoustic model is calculated by multiplying the level by 0.25. Notably, this is a simplification of a model where this factor varies from approximately 0.25 for non-tonal content to 0.35 for tonal content. The SMR is higher for tonal content than for noisy content at the same level. If tonality information per frequency band is available, the factor may be applied accordingly. The division by 3 gives an approximation of the number of bits per MDCT line needed to achieve the SMR.
[0067] In some implementations, ^^may be identical for all frequency bands, leading to a constant ^. In this case, the number of bits per band ^^,^for transparent quality coding can be estimated as ^= max ^^ ,^ lo ^,^^,^ : ^ g^^ :max :! , 1.0;; , ^^^^;^By virtue ofadjustment) a predetermined minimum number of bits per band. At step S330, the measure of perceptual entropy is determined (e.g., estimated) based on the numbers of bits per band for the plurality of frequency bands. For example, the perceptual entropy "^of an audio element ^ may be estimated as K "^ = 1where L is the total number of
[0068] Accordingly, the aforementioned measure of perceptual entropy "^for a given channel element ^ may be based, for example, on a sum, over the plurality of frequencybands ^ = 1, … , L, of the numbers of bits per band ^^,^ for the given channel element ^. Or,in other words, the measure of perceptual entropy "^of a channel element ^ may for example be based on or correspond to a number of bits for transparent quality coding of the (full-band) channel element.
[0069] Adjustment of the determined PEs at step S220 in method 200 may be performed as follows. In general, adjusting the (measure of) perceptual entropy for a given channel element may be based to the number of target bits and the measure of aggregated (e.g., total) PE, as described above.
[0070] The PE of an audio frame, which is an example of the aggregated (e.g., total) perceptual entropy of the plurality of channel elements in the audio frame, may be computed asM "= 1 " ^where / is the number of elements.
[0071] In addition, adjusting the (measure of) perceptual entropy for a given channel element may be further based on a number of perceptually relevant bins of the frequency representation of the given channel element. These bins of the frequency representation of the channel element may correspond to MDCT lines, for example. Adjusting the perceptual entropy for the given channel element may be further based on a total number of perceptually relevant bins of the frequency representations of the plurality of channel elements.
[0072] The number of relevant (e.g., perceptually relevant) bins (e.g., MDCT lines) of the frequency representation of a channel element ^ may be for example computed as K ^^ = 1with the number of perceptually
[0073] The total number of (perceptually relevant) bins (e.g., number of relevant MDCT lines) in an audio frame may then be computed for example as M ^= 1 ^^^5^ (8)
[0074] With this, the PE adjustment per bin (e.g., per MDCT line) & may be based on a total number of perceptually relevant bins (e.g., MDCT lines) of the frequency representations of the plurality of channel elements in the frame, the total perceptual entropy " of the frame, and the desired number ^^() of bits per frame (e.g., number of target bits for the frame). For example, the PE adjustment per bin, &, may be computed as $'& = " − ^^()^(9)
[0075] With this, the adjusted measure "^#of perceptual entropy for a channel element ^ may be given by" #^ = max$" #^ − &^^, " ^^^ '(10) where "^is the measure of perceptual entropy for channel element ^, & is the PE adjustment per bin, ^^is the number of perceptually relevant bins of the frequency representation of channel element ^, and "^#^^is a minimum number (minimum value) for the adjusted measure of perceptual entropy. The minimum number "^#^^for the adjusted measure of perceptual entropy may be 0, for example, leading to an adjusted "^#for channel element ^ given by "#^ = max $" ^ − &^^, 0'(11)
[0076] The adjusted (measures of) PE, "^′ may then be used for bit allocation, for example for determining a bit allocation, for the plurality of channel elements at step S230 in method 200.
[0077] For example, the bit allocation may be based on PE ratios ,^for the plurality of channel elements.
[0078] The PE ratio ,^for channel element ^ may be computed as, for example "#, = ^^ " #where 0 ≤ ,^ ≤ 1In some implementations, the level dependent ratio ,^that is used to determine the target number of bit per channel (element) may be blended with a static channel bit distribution to avoid zero target bits for certain channels. This may be achieved by modifying the level dependent ratio for example via ,′^ = ,^ R + $1 − R')^(13) where )^are the static channel bit distributions with M 1)^ = 1.0
[0079] Without intended limitation, a typical value for constant R may be 0.85, for example.
[0080] Next, different modes (encoding modes) for encoding a sequence of audio frames will be described with reference to Fig.4 and Fig.5. Of these, Fig.4 shows a flowchart of a method 400 that encodes audio frames in a sequential mode, one by one and independently of each other, and Fig.5 shows a flowchart of a method 500 that encodes audio frames in a memory efficient mode (chain mode) where bit allocation for a given audio frame is performed based on information derived for the previous audio frame. In both modes, the encoding sequentially operates on the frames of the audio content.
[0081] Especially in the memory efficient mode, the encoding may further comprise buffering the adjusted measures of perceptual entropy between frames. Further, again especially in the memory efficient mode, the encoding may further comprise smoothing the adjusted measures of perceptual entropy between frames using a low pass filter. This low-pass filter may be a first order recursive filter, for example.
[0082] Method 400 in Fig.4, relating to the sequential mode, comprises steps S410 and S420.
[0083] At step S410, a bit allocation for a plurality of channel elements in a frame + of audio content is determined. This may be done as described above in the framework of Fig.1, and / or using method 200 of Fig.2.
[0084] At step S420, the frame + of the audio content is encoded (e.g., quantized) based on the adjusted measures of perceptual entropy and the number of target bits. This involves distributing the bits available for the frame in accordance with the adjusted measures of perceptual entropy.
[0085] For example, the target number of bits ^^()^,*per channel element ^ in frame + may be given by ^^()^,* = ,′^,*^^()*(15)
[0086] Alternatively, instead of using the modified level dependent ratio ,^#,*, the unmodified level dependent ratio ,^,*(i.e., without using the modification of Eq. (13)) may be used for this purpose.
[0087] Method 500 in Fig.5, relating to the memory efficient mode, comprises steps S510 and S520.
[0088] At step S510, a bit allocation for a plurality of channel elements in a frame + −1 of audio content is determined. This may be done as described above in the framework of Fig.1, and / or using method 200 of Fig.2.
[0089] At step S520, a subsequent frame + of the audio content is encoded based on the adjusted measures of perceptual entropy determined for the frame and the number of target bits for the subsequent frame. This involves distributing the bits available for the subsequent frame + in accordance with the adjusted measures of perceptual entropy determined for the frame +-1.
[0090] For example, the target number of bits ^^()^,*per channel element ^ in frame + may be given by ^^()^,* = ,′^,*-^^^()*(16)
[0091] Alternatively, instead of using the modified level dependent ratio ,^#,*-^, the unmodified level dependent ratio ,^,*-^(i.e., without using the modification of Eq. (13)) may be used for this purpose.
[0092] It is further noted that while the target number of bits ^^()^,*per channel element ^ in frame + may be based on more than one preceding frame, such as based on apredefined number ^ of frames (i.e., frames + − 1, + − 2, … , + − ^'. The (modified) leveldependent ratios ,$#'^,*-^ , ,$#'^,*-3 , … , ,$#'^,*-T of these frames, or alternatively, the (adjusted)measures ofof channel element bit estimates) of these frames, may be combined in any suitable manner, for example by averaging, smoothing, filtering, weighting, etc., for determining the target number of bits ^^()^,*per channel element ^ in frame +, or alternatively, a corresponding level dependent ratio.
[0093] Advantageously, when using this mode, the plurality of channel elements in the frame may be processed and encoded sequentially, since information on the bit allocation to the channel elements can be derived from the number of target bits and information retained from the previous frame (notably, the adjusted measures of PE or PE ratios), without knowledge of the PEs of the other channel elements in the frame. This allows to significantly reduce the memory footprint of encoding.Apparatus for Implementing Methods According to the Disclosure
[0094] Finally, the present disclosure likewise relates to an apparatus (e.g., computer- implemented apparatus) for performing methods and techniques described throughout the present disclosure. For example, this apparatus may relate to an audio encoder.
[0095] Fig.6 shows an example of such apparatus 600. In particular, apparatus 600 comprises a processor 610 and a memory 620 coupled to the processor 610. The memory 620 may store instructions for the processor 610. The processor 610 may also receive, among others, suitable input data 630 (e.g., audio content, target bitrate, etc.), depending on use cases and / or implementations. The processor 610 may be adapted to carry out or implement the methods / techniques described throughout the present disclosure (e.g., method 200 of Fig.2, method 300 of Fig.3, method 400 of Fig.4, and / or method 500 of Fig.5) and to generate corresponding output data 640 (e.g., bitrate allocations, quantized or encoded audio frames, etc.), depending on use cases and / or implementations.
[0096] The present disclosure likewise relates to corresponding computer programs, computer program products, and computer-readable storage media storing such computer programs or computer program products. Summary
[0097] One aspect of the present disclosure seeks to distribute bits in an audio frame to audio channels in a perceptually meaningful way. This aspect relates to the use of a level- dependent psychoacoustic model for channel bit distribution by computing channel perceptual entropies, adjusting the determined channel perceptual entropies given the target bitrate, and deriving channel perceptual entropy ratios (e.g., relative to total perceptual entropy or relative to a number of target bits). The channel bits may then be computed based on the ratios and the frame bit budget (e.g., number of target bits).
[0098] Some embodiments use a level-dependent psychoacoustic model operating on frequency bands to determine the perceptual entropy (PE) for all audio channels / audio objects as a basis for channel bit allocation. Some embodiments adjust the estimated PE such that the total PE corresponds to the target frame size.
[0099] Some embodiments compute PE ratios by dividing adjusted PEs by the number of frame bits and computing channel bits by multiplying PE ratios by the number of frame bits where the PE ratio may be the ratio from the previous frame for memory efficient implementation.
[0100] Some embodiments compute the level dependent PE by considering known relevant playback level changes.
[0101] Another aspect relates to a low memory implementation of the aforementioned bit distribution in encoding. In this implementation, channel perceptual entropy ratios are calculated after the last channel element in a frame has been processed, and the channel perceptual entropy ratios are stored (e.g., buffered) for the next frame. In this configuration, the channel bits for a given frame are computed based on the channel perceptual entropy ratios of the previous frame, allowing channel elements in the given frame to be encoded at once, without need for first processing all the channel elements in the given frame. Interpretation
[0102] Aspects of the methods and apparatus / systems described herein may be implemented in an appropriate computer-based audio processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files. Portions of the audio system may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0103] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (non-transitory), non- volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.
[0104] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium)executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, the apparatus (e.g., encoders) described above can include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.
[0105] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements. Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Enumerated Example Embodiments
[0106] Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.
[0107] EEE1. A method of determining a bit allocation for a plurality of channel elements of audio content, the method comprising: for a frame of the audio content, determining a measure of perceptual entropy for each of the plurality of channel elements in the frame; adjusting the determined measures of perceptual entropy based on a number of target bits for the frame and a measure of aggregated perceptual entropy of the plurality of channel elements in the frame; and performing bit allocation for the plurality of channel elements based on the adjusted measures of perceptual entropy.
[0108] EEE2. The method according to EEE1, wherein the measure of perceptual entropy for a given channel element is positively correlated with the energy of the given channel element.
[0109] EEE3. The method according to EEE1 or EEE2, wherein the measure of aggregated perceptual entropy of the plurality of channel elements in the frame is based on a sum of the determined measures of perceptual entropy for the plurality of channel elements in the frame.
[0110] EEE4. The method according to any one of EEE1 to EEE3, wherein determining the measure of perceptual entropy for a given channel element comprises: determining, for each of a plurality of frequency bands, an energy per band of the given channel element in the respective frequency band based on a frequency-domain representation of the given channel element in the frame; determining, for each of the plurality of frequency bands, a number of bits per band for the respective frequency band based on the energy per band of the given channel element; and determining the measure of perceptual entropy based on the numbers of bits per band for the plurality of frequency bands.
[0111] EEE5. The method according to EEE4, wherein the measure of perceptual entropy for a given channel element is based on a sum, over the plurality of frequency bands, of the numbers of bits per band for the given channel element.
[0112] EEE6. The method according to EEE4 or EEE5, wherein determining the number of bits per band is further based on a threshold in quiet for the respective frequency band.
[0113] EEE7. The method according to any one of EEE4 to EEE6, wherein determining the number of bits per band is further based on an estimate of tonality for the respective frequency band.
[0114] EEE8. The method according to any one of EEE4 to EEE7, wherein determining the number of bits per band is based on a level-dependent psychoacoustic model.
[0115] EEE9. The method according to any one of EEE4 to EEE8, wherein the number of bits per band is set to at least a predetermined minimum number of bits per band.
[0116] EEE10. The method according to any one of EEE4 to EEE9, wherein the number of bits per band ^^,^for channel element ^ and frequency band ^ is given by ^= max ^ ^ lo ^,^^,^ ^ ^ ^,^ g^^ 9max :! , 1.0;< , ^^^^ ^^where^,^is the energy per band of channel element ^ in frequency band ^, !^is a threshold in quiet for frequency band ^, ^^is a band-dependent constant, ^^^^is a predetermined minimum number of bits per band, and ^^,^is a number of perceptually relevant bins of the frequency representation of the given channel element for frequency band ^.
[0117] EEE11. The method according to any one of EEE1 to EEE10, wherein adjusting the perceptual entropy for a given channel element is further based on a number of perceptually relevant bins of the frequency representation of the given channel element.
[0118] EEE12. The method according to any one of EEE1 to EEE11, wherein adjusting the perceptual entropy for a given channel element is further based on a total number of perceptually relevant bins of the frequency representations of the plurality of channel elements.
[0119] EEE13. The method according to any one of EEE1 to EEE12, wherein the adjusted measure "^#of perceptual entropy for a channel element ^ is given by "#^ = max$" ^ − &^ #^, " ^^^ 'where "^is the measure of perceptual entropy for channel element ^, & is a constant based on a total number of perceptually relevant bins of the frequency representations of the plurality of channel elements in the frame, a total perceptual entropy of the frame, and the number of target bits for the frame, ^^is a number of perceptually relevant bins of the frequency representation of channel element ^, and "^#^^is a minimum number for the adjusted measure of perceptual entropy.
[0120] EEE14. The method according to any one of EEE1 to EEE13, wherein the plurality of channel elements in the frame are processed sequentially.
[0121] EEE15. The method according to any one of EEE1 to EEE14, wherein the method sequentially operates on frames of the audio content, and wherein the method further comprises: buffering the adjusted measures of perceptual entropy between frames.
[0122] EEE16. The method according to any one of EEE1 to EEE15, wherein the method sequentially operates on frames of the audio content, and wherein the method further comprises: smoothing the adjusted measures of perceptual entropy between frames using a low pass filter.
[0123] EEE17. The method according to any one of EEE1 to EEE16, further comprising:encoding the frame of the audio content based on the adjusted measures of perceptual entropy and the number of target bits.
[0124] EEE18. The method according to EEE17, wherein a number of bits ^^()^,*to be allocated per channel element ^ in frame + of the audio content is given by ^^()^,* = ,^,*^^()*where ,^,*is a ratio based on the adjusted measure of perceptual entropy for channel element ^ in frame + and ^^()*is the number of target bits for frame +.
[0125] EEE19. The method according to any one of EEE1 to EEE16, further comprising: encoding a subsequent frame of the audio content based on the adjusted measures of perceptual entropy determined for the frame and the number of target bits for the subsequent frame.
[0126] EEE20. The method according to EEE19, wherein a number of bits ^^()^,*to be allocated per channel element ^ in frame + of the audio content is given by ^^()^,* = ,^,*-^^^()*where ,^,*-^is a ratio based on the adjusted measure of perceptual entropy for channelelement ^ in frame + − 1 and ^^()* is the number of target bits for frame +.
[0127] EEE21. An apparatus, comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of EEE1 to EEE20.
[0128] EEE22. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE1 to EEE20.
[0129] EEE23. A computer-readable storage medium storing the program according to EEE22.
Claims
CLAIMS 1. A method of determining a bit allocation for a plurality of channel elements of audio content, the method comprising: for a frame of the audio content, determining a measure of perceptual entropy for each of the plurality of channel elements in the frame; adjusting the determined measures of perceptual entropy based on a number of target bits for the frame and a measure of aggregated perceptual entropy of the plurality of channel elements in the frame; and performing bit allocation for the plurality of channel elements based on the adjusted measures of perceptual entropy.
2. The method according to claim 1, wherein the measure of perceptual entropy for a given channel element is positively correlated with the energy of the given channel element.
3. The method according to claim 1 or 2, wherein the measure of aggregated perceptual entropy of the plurality of channel elements in the frame is based on a sum of the determined measures of perceptual entropy for the plurality of channel elements in the frame.
4. The method according to any one of claims 1 to 3, wherein determining the measure of perceptual entropy for a given channel element comprises: determining, for each of a plurality of frequency bands, an energy per band of the given channel element in the respective frequency band based on a frequency-domain representation of the given channel element in the frame; determining, for each of the plurality of frequency bands, a number of bits per band for the respective frequency band based on the energy per band of the given channel element; and determining the measure of perceptual entropy based on the numbers of bits per band for the plurality of frequency bands.
5. The method according to claim 4, wherein the measure of perceptual entropy for a given channel element is based on a sum, over the plurality of frequency bands, of the numbers of bits per band for the given channel element.
6. The method according to claim 4 or 5, wherein determining the number of bits per band is further based on a threshold in quiet for the respective frequency band.
7. The method according to any one of claims 4 to 6, wherein determining the number of bits per band is further based on an estimate of tonality for the respective frequency band.
8. The method according to any one of claims 4 to 7, wherein determining the number of bits per band is further based on a level-dependent psychoacoustic model.
9. The method according to any one of claims 4 to 6, wherein the number of bits per band is set to at least a predetermined minimum number of bits per band.
10. The method according to any one of claims 4 to 9, wherein the number of bits per band ^^,^for channel element ^ and frequency band ^ is given by ^= max ^ ^ log 9ma ^,^^,^ ^ ^ ^,^ ^^ x :! , 1.0;< , ^^^^ ^^where^,^is the energy per band of channel element ^ in frequency band ^, !^is a threshold in quiet for frequency band ^, ^^is a band-dependent number, ^^^^is a predetermined minimum number of bits per band, and ^^,^is a number of perceptually relevant bins of the frequency representation of the given channel element for frequency band ^.
11. The method according to any one of claims 1 to 10, wherein adjusting the perceptual entropy for a given channel element is further based on a number of perceptually relevant bins of the frequency representation of the given channel element.
12. The method according to any one of claims 1 to 10, wherein adjusting the perceptual entropy for a given channel element is further based on a total number of perceptually relevant bins of the frequency representations of the plurality of channel elements.
13. The method according to any one of claims 1 to 12, wherein the adjusted measure "^#of perceptual entropy for a channel element ^ is given by "#^ = max$" #^ − &^^, " ^^^ 'where "^is the measure of perceptual entropy for channel element ^, & is a constant based on a total number of perceptually relevant bins of the frequency representations of the plurality of channel elements in the frame, a total perceptual entropy of the frame, and the number of target bits for the frame, ^^is a number of perceptually relevant bins of the frequency representation of channel element ^, and "^#^^is a minimum number for the adjusted measure of perceptual entropy.
14. The method according to any one of claims 1 to 13. wherein the plurality of channel elements in the frame are processed sequentially.
15. The method according to any one of claims 1 to 14, wherein the method sequentially operates on frames of the audio content, and wherein the method further comprises: buffering the adjusted measures of perceptual entropy between frames.
16. The method according to any one of claims 1 to 15, wherein the method sequentially operates on frames of the audio content, and wherein the method further comprises: smoothing the adjusted measures of perceptual entropy between frames using a low pass filter.
17. The method according to any one of claims 1 to 16, further comprising: encoding the frame of the audio content based on the adjusted measures of perceptual entropy and the number of target bits.
18. The method according to claim 11, wherein a number of bits ^^()^,*to be allocated per channel element ^ in frame + of the audio content is given by ^^()^,* = ,^,*^^()*where ,^,*is a ratio based on the adjusted measure of perceptual entropy for channel element ^ in frame + and ^^()*is the number of target bits for frame +.
19. The method according to any one of claims 1 to 10, further comprising: encoding a subsequent frame of the audio content based on the adjusted measures of perceptual entropy determined for the frame and the number of target bits for the subsequent frame.
20. The method according to claim 19, wherein a number of bits ^^()^,*to be allocated per channel element ^ in frame + of the audio content is given by ^^()^,* = ,^,*-^^^()*where ,^,*-^is a ratio based on the adjusted measure of perceptual entropy for channelelement ^ in frame + − 1 and ^^()* is the number of target bits for frame +.
21. An apparatus, comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of claims 1 to 20.
22. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1 to 20.
23. A computer-readable storage medium storing the program according to claim 22.
Citation Information
Patent Citations
A psychoacoustic model for audio processing
US20220415334A1
Audio encoding apparatus
EP2202724A1
Encoding method and apparatus, electronic device, and storage medium
US20230326467A1