AUDIO ENCODER WITH A NUMBER DEPENDENT ON THE SIGNAL AND PRECISION CONTROL, AUDIO DECODER, AND RELATED COMPUTER METHODS AND PROGRAMS

MX431608BActive Publication Date: 2026-02-25FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
MX2021015564
Authority / Receiving Office
MX · MX
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-17
Filing Date
2021-12-14
Publication Date
2026-02-25
Estimated Expiration
2040-06-10

AI Technical Summary

Technical Problem

State-of-the-art audio encoders face challenges in accurately estimating bit consumption due to computational complexity, leading to inefficiencies and quality loss, particularly for highly tonal signals, as they rely on psychoacoustic models that do not account for signal-dependent variations.

Method used

An audio encoder that preprocesses input data based on signal characteristics to reduce the number of data elements in a signal-dependent manner, reallocating bits from an initial encoding stage to a refinement stage, using a two-stage encoding process with entropy and residual coding, and incorporating a signal-adaptive noise floor to enhance precision for tonal signals.

Benefits of technology

This approach improves encoding efficiency and audio quality by optimizing bit allocation based on signal characteristics, ensuring precise encoding of remaining data elements while maintaining computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure MX431608B0
    Figure MX431608B0
  • Figure MX431608B1
    Figure MX431608B1
Patent Text Reader

Abstract

An audio encoder for encoding audio input data (11) comprises: a preprocessor (10) for preprocessing the audio input data (11) to obtain audio data to be encoded; an encoder processor (15) for encoding the audio data to be encoded; and a controller (20) for controlling the encoder processor (15) such that, depending on a first signal characteristic of a first frame of the audio data to be encoded, a number of audio data elements of the audio data to be encoded by means of the encoder processor (15) for the first frame is reduced compared to a second signal characteristic of a second frame, and a first number of information units used to encode the reduced number of audio data elements for the first frame is further improved compared to a second number of information units for the second frame.
Need to check novelty before this filing date? Find Prior Art

Description

AUDIO ENCODER WITH A SIGNAL-DEPENDENT NUMBER AND PRECISION CONTROL, AUDIO DECODER, AND RELATED COMPUTER METHODS AND PROGRAMS FIELD OF INVENTION The present invention relates to audio signal processing and, in particular, to audio encoders / decoders that apply a signal-dependent number and precision control. BACKGROUND OF THE INVENTION Modern transform-based audio encoders apply a series of psychoacoustically motivated processings to a spectral representation of an audio segment (a frame) to obtain a residual spectrum. This residual spectrum is quantized, and the coefficients are encoded using entropic coding. In this process, the quantization step size, which is generally controlled by an overall gain, has a direct impact on the encoder's bit consumption and needs to be selected to satisfy the bit budget, which is usually limited and often fixed. Since the encoder's bit consumption, and particularly that of an arithmetic encoder, is not known exactly before encoding, the calculation of the optimal overall gain can only be done in a closed-loop iteration of quantization and encoding. However, this is not feasible under certain complexity constraints, as arithmetic encoding introduces significant computational complexity. Therefore, state-of-the-art encoders, such as those found in the 3GPP EVS codec, typically feature a bit consumption estimator to derive a first estimate of overall gain, usually operating in the residual signal's power spectrum. Depending on the complexity constraint, this may be followed by a rate loop to refine the first estimate. Using such an estimate alone or in conjunction with very limited correction capabilities reduces complexity, but also reduces accuracy, leading to significant under- or overestimations of bit consumption. Overestimating bit consumption leads to excess bits after the first encoding stage. State-of-the-art encoders utilize these to refine the quantization of the encoded coefficients in a second encoding stage known as residual coding. Residual coding differs fundamentally from the first encoding stage because it operates at bit granularity and therefore does not incorporate any entropic coding. Furthermore, residual coding is generally only applied at frequencies with non-zero quantized values, leaving dead zones that are not further enhanced. On the other hand, underestimating bit consumption inevitably leads to a partial loss of spectral coefficients, generally the higher frequencies. In state-of-the-art encoders, this effect is mitigated by applying noise substitution in the decoder, which is based on the assumption that high-frequency content is generally noisy. In this configuration, it is evident that it is desirable to encode as much of the signal as possible in the first stage of / QQLLn / ZZnZ / q / Yli encoding, which uses entropic encoding and is therefore more efficient than the residual encoding stage. Therefore, one would want to select the overall gain with a bit estimate as close as possible to the available bit budget. While the power spectrum-based estimator works well for most audio content, it can cause problems for highly tonal signals. The first-stage estimate relies primarily on irrelevant sidelobes from the filter bank's frequency decomposition, while important components are lost due to an underestimation of bit consumption. BRIEF DESCRIPTION OF THE INVENTION An objective of the present invention is to provide an improved concept for audio encoding or decoding that is nevertheless efficient and achieves good audio quality. This objective is achieved by means of an audio encoder according to claim 1, a method for encoding audio input data according to claim 33, and an audio decoder according to claim 35, a method for decoding encoded audio data according to claim 41, or a computer program according to claim 42. The present invention is based on the finding that, in order to improve efficiency, particularly with respect to bit rate on the one hand and audio quality on the other, a signal-dependent change is necessary from the typical situation given by psychoacoustic considerations. Typical psychoacoustic models or considerations result in good audio quality at a low bit rate for all signal classes on average, that is, for all audio signal frames regardless of their signal characteristics, when an average result is considered.However, it has been discovered that for certain signal classes or for signals that have certain signal characteristics such as fairly tonal signals, the direct psychoacoustic model or the direct psychoacoustic control of the encoder only gives suboptimal results with respect to audio quality (when the bit rate is kept constant), or with respect to the bit rate (when the audio quality is kept constant). Therefore, in order to address this drawback of typical psychoacoustic considerations, the present invention provides, in the context of an audio encoder with a preprocessor for preprocessing the input audio data to obtain audio data to be encoded, and an encoder processor for encoding the audio data to be encoded, a controller for controlling the encoder processor so that, depending on a certain signal characteristic of the frame, a number of audio data elements is reduced from the audio data to be encoded by the encoder processor compared to typical direct results obtained by means of prior art psychoacoustic considerations. Furthermore, this reduction of the number of audio data elements is done in a signal-dependent manner so that, for a frame with a certain first signal characteristic,The number is reduced more than for another frame with a different signal characteristic than the signal characteristic of the first frame. This reduction in the number of audio data elements can be considered as a reduction in the absolute number or a reduction in the relative number, although this is not decisive. However, it is a characteristic that the units of information "saved" by the intentional reduction in the number of audio data elements are not simply lost, but are used to more accurately encode the remaining number of data elements, that is, the data elements that have not been eliminated by the intentional reduction in the number of audio data elements. According to the invention, the controller for controlling the encoder processor operates such that, depending on the first signal characteristic of a first frame of the audio data to be encoded, a number of audio data elements of the audio data to be encoded by the encoder processor for the first frame is reduced compared to a second signal characteristic of a second frame, and, at the same time, a first number of information units used to encode the reduced number of audio data elements for the first frame is further enhanced compared to a second number of information units for the second frame. In a preferred mode, the reduction is performed such that, for more tonal signal frames, a stronger reduction is carried out, and at the same time, the number of bits for individual lines is increased more compared to a less tonal, i.e., noisier frame. Here, the number is not reduced to such a high degree, and correspondingly, the number of information units used to encode the less tonal audio data elements is not significantly increased. The present invention provides an environment where, in a signal-dependent manner, the generally provided psychoacoustic considerations are violated to a greater or lesser degree. However, this violation is not treated as in normal encoders, where a violation of psychoacoustic considerations occurs, for example, in an emergency situation such as one where, in order to maintain a required bit rate, the highest frequency portions are set to zero. Instead, according to the present invention, such a violation of normal psychoacoustic considerations occurs independently of any emergency situation, and the "saved" information units are used to further refine the "surviving" audio data elements. In preferred embodiments, a two-stage encoder processor is used, which has, as an initial encoding stage, for example, an entropy encoder such as an arithmetic encoder, or a variable-length encoder such as a Huffman encoder. The second encoding stage serves as a refinement stage, and this second encoder is generally implemented in preferred embodiments as a residual encoder or a bit encoder operating at a bit granularity that can, for example, be implemented by adding a certain defined offset for the first value of an information unit or subtracting an offset for the opposite value of the information unit. In one embodiment, this refinement encoder is preferably implemented as a residual encoder by adding an offset for the first bit value and subtracting an offset for the second bit value.In a preferred embodiment, reducing the number of audio data elements results in a situation where the distribution of available bits in a typical fixed-frame-rate scenario is altered such that the initial encoding stage receives a lower bit budget than the refinement encoding stage. Until now, the paradigm was that the initial encoding stage should receive as high a bit budget as possible, regardless of the signal characteristics, since it was believed that an initial encoding stage, acting as an arithmetic encoding stage, has the highest efficiency and therefore encodes much better than a residual encoding stage from an entropy standpoint.However, according to the present invention, this paradigm is overcome, since it has been discovered that for certain signals, such as signals with a higher pitch, the efficiency of an entropy encoder, such as an arithmetic encoder, is not as high as the efficiency obtained by a downstream residual encoder, such as a bit encoder. While it is true that the entropy encoding stage is highly efficient for audio signals on average, the present invention addresses this problem not by averaging, but by reducing the bit budget for the initial encoding stage in a signal-dependent manner, preferably for portions of the pitch signal. In a preferred embodiment, the shift of the bit budget from the initial encoding stage to the refinement encoding stage, based on the signal characteristic of the input data, is done such that at least two refinement information units are available for at least one, and preferably 50%, and even more preferably all, of the data elements that have survived the reduction in the number of data elements. Furthermore, a particularly efficient procedure for calculating these refinement information units on the encoder side and applying them on the decoder side has been found to be an iterative procedure where, in a certain order, such as from a low frequency to a high frequency, the remaining bits of the bit budget for the refinement encoding stage are consumed one after the other.Depending on the number of surviving audio data elements and depending on the number of information units for the refinement encoding stage, the number of iterations can be significantly greater than two, and it has been found that for strongly tonal signal frames, the number of iterations can be four, five, or even more. In a preferred embodiment, the determination of a control value by the controller is done indirectly, that is, without explicit determination of the signal characteristic. For this purpose, the control value is calculated based on manipulated input data, where this manipulated input data is, for example, the input data to be quantized or amplitude-related data derived from the data to be quantized. Although the control value for the encoder processor is determined based on manipulated data, the actual quantization / coding is performed without this manipulation.Thus, the signal-dependent procedure is obtained by determining a manipulation value for manipulation in a signal-dependent manner, where this manipulation influences more or less the reduction obtained from the number of audio data elements, without explicit knowledge of the specific signal characteristic. In another implementation, the direct mode can be applied in which a certain signal characteristic is directly estimated and depending on the result of this signal analysis, a certain reduction of the number of data elements is carried out in order to obtain higher accuracy for the surviving data elements. In a further implementation, a separate procedure can be applied to reduce the audio data elements. In this separate procedure, a certain number of data elements are obtained through quantization controlled by a quantizer, typically psychoacoustically actuated. Based on the input audio signal, the already quantized audio data elements are reduced in number, preferably by eliminating the smallest audio data elements in terms of amplitude, energy, or power. The control for this reduction can again be achieved through direct / explicit signal characteristic determination or through indirect / non-explicit signal control. In a further preferred embodiment, the integrated procedure is applied, in which the variable quantizer is controlled to perform a single quantization based on manipulated data, while simultaneously quantizing the unmanipulated data. A quantizer control value, such as an overall gain, is calculated using signal-dependent manipulated data while quantizing the unmanipulated data. The quantization result is then encoded using all available information units, so that, in the case of two-stage encoding, a generally high number of information units are retained for refinement encoding. These methods provide a solution to the quality loss problem for highly tonal content by modifying the power spectrum used to estimate the encoder's bit consumption by entropy. This modification consists of a signal-adaptive background noise adder that keeps the estimate for ordinary audio content with a flat residual spectrum virtually unchanged while increasing the bit budget estimate for highly tonal content. The effect of this modification is twofold. First, it causes filter bank noise and irrelevant sidelobes of harmonic components, which are masked by background noise, to be quantized as zero. Second, it shifts bits from the first encoding stage to the residual encoding stage.While such a shift is undesirable for most signals, it is entirely efficient for highly tonal signals because the bits are used to increase the quantization accuracy of harmonic components. This means they are used to encode bits of low importance that generally follow a uniform distribution and are therefore efficiently encoded with a binary representation. Furthermore, the procedure is computationally inexpensive, making it a very effective tool for solving the aforementioned problem. BRIEF DESCRIPTION OF THE DRAWINGS The following are preferred embodiments of the present invention with respect to accompanying drawings, wherein: Figure 1 is one type of audio encoder. Figure 2 illustrates a preferred implementation of the encoder processor of Figure 1. Figure 3 illustrates a preferred implementation of a refinement coding stage. Figure 4a illustrates an exemplary frame syntax for a first or second frame with iteration refinement bits. / QQLLn / ZZnZ / q / Yli Figure 4b illustrates a preferred implementation of an audio data element reducer as a variable quantizer. Figure 5 illustrates a preferred implementation of the audio encoder with a spectrum preprocessor. Figure 6 illustrates a preferred modality of an audio decoder with a time post-processor. Figure 7 illustrates an implementation of the audio decoder encoder processor from Figure 6. Figure 8 illustrates a preferred implementation of the refinement decoding stage of Figure 7. Figure 9 illustrates an implementation of an indirect mode for calculating control value. Figure 10 illustrates a preferred implementation of the manipulation value calculator from Figure 9. Figure 11 illustrates a direct mode control value calculation. Figure 12 illustrates an implementation of separate audio data element reduction. Figure 13 illustrates an implementation of integrated audio data element reduction. DETAILED DESCRIPTION OF THE INVENTION Figure 1 illustrates an audio encoder for encoding audio input data 11. The audio encoder comprises a preprocessor 10, an encoder processor 15, and a controller 20. The preprocessor 10 preprocesses the audio input data 11 to obtain per-frame audio data or audio data to be encoded, as illustrated in element 12. The audio data to be encoded is fed into the encoder processor 15 for encoding, and the encoder processor outputs the encoded audio data. The controller 20 is connected, with respect to its input, to the per-frame audio data from the preprocessor, but alternatively, the controller can also be connected to receive the audio input data without any preprocessing.The controller is configured to reduce the number of audio data elements per frame depending on the signal in the frame and, at the same time, the controller increases a number of information units or, preferably, bits for the number of audio data elements depending on the signal in the frame.The controller is configured to control the encoder processor 15 so that, depending on the first signal characteristic of a first frame of the audio data to be encoded, a number of audio data elements of the audio data to be encoded by the encoder processor for the first frame is reduced compared to a second signal characteristic of a second frame, and a number of information units used to encode the reduced number of audio data elements for the first frame is further enhanced compared to a second number of information units for the second frame. Figure 2 illustrates a preferred implementation of the encoder processor. The encoder processor comprises an initial encoding stage 151 and a refinement encoding stage 152. In one implementation, the initial encoding stage comprises an entropy encoder such as an arithmetic or Huffman encoder. In another embodiment, the refinement encoding stage 152 comprises a bit encoder or a residual encoder operating at an information unit or bit granularity.Furthermore, the functionality regarding the reduction of the number of audio data elements is incorporated in Figure 2 by means of the audio data element reducer / QQ L iP / 77P7 / 3 / YI 150 which can be implemented, for example, as a variable quantizer in the integrated reduction mode illustrated in Figure 13 or, alternatively, as a separate element operating on already quantized audio data elements as illustrated in separate reduction mode 902 and, in a further unillustrated mode, the audio data element reducer can also operate on unquantized elements by setting such unquantized elements to zero or by weighting the data elements to be removed with a certain weighting number so that the audio data elements are quantized to zero and thus removed in a subsequently connected quantizer.The audio data element reducer 150 of Figure 2 can operate on non-quantized or quantized data elements in a separate reduction procedure or can be implemented by means of a variable quantizer controlled specifically by a signal-dependent control value as illustrated in the integrated reduction mode of Figure 13. Controller 20 in Figure 1 is configured to reduce the number of audio data elements encoded by the initial encoding stage 151 for the first frame, and the initial encoding stage 151 is configured to encode the reduced number of audio data elements for the first frame using a number of initial frame information units, and the calculated bits / units from the number of initial information units are provided by block 151 as illustrated in Figure 2, element 151. In addition, refinement encoding stage 152 is set up to use a number of remaining first frame information units for refinement encoding for the reduced number of audio data elements for the first frame, and the number of initial first frame information units added to the number of remaining first frame information units results in a predetermined number of information units for the first frame.In particular, refinement encoding stage 152 outputs the number of remaining bits of the first frame and the number of remaining bits of the second frame, and there are at least two refinement bits for at least one, or preferably at least 50%, or even more preferably all of the non-zero audio data elements, i.e., the audio data elements that survive the reduction of audio data elements and are initially encoded by the initial encoding stage 151. Preferably, the default number of information units for the first frame is equal to or very close to the default number of information units for the second frame so that a constant or substantially constant bit rate operation is obtained for the audio encoder. As illustrated in Figure 2, the audio data element reducer 150 reduces audio data elements beyond the psychoacoustically driven number in a signal-dependent manner. Therefore, for a given signal feature, the number is reduced only slightly above the psychoacoustically driven number, and in a frame with a second signal feature, for example, the number is reduced beyond a psychoacoustically driven number. Preferably, the audio data element reducer eliminates data elements with the smallest amplitudes / powers / energies, and this operation is preferably carried out by means of an indirect selection obtained in integrated mode, where the reduction of audio data elements is performed by quantizing certain audio data elements to zero.In one mode, the initial encoding stage only encodes audio data elements that have not been quantized to zero, and the refinement encoding stage 152 only refines audio data elements already processed by the initial encoding stage, i.e., the audio data elements that have not been quantized to zero by the audio data element reducer 150 of Figure 2. In a preferred embodiment, the refinement encoding stage is configured to iteratively assign the number of remaining information units from the first frame to the reduced number of audio data elements in the first frame in at least two sequential iterations. Specifically, the assigned information unit values ​​are calculated for these at least two sequential iterations, and these calculated values ​​are then fed into the encoded output frame in a predetermined order.Specifically, the refinement encoding stage is configured to sequentially assign one unit of information to each audio data element within the reduced number of audio data elements for the first frame, in order from low-frequency information to high-frequency information in the first iteration. Specifically, the audio data elements can be individual spectral values ​​obtained through a time-to-spectral conversion. Alternatively, the audio data elements can be tupiates of two or more spectral lines that are typically adjacent to each other in the spectrum.The calculation of bit values ​​is carried out from a certain starting value with low frequency information to a certain final value with the highest frequency information and, in an additional iteration, the same procedure is carried out, i.e., again processing from low spectral information values / tuples to high spectral information values / tuples.In particular, refinement coding stage 152 is configured to check if a number of information units already allocated is lower than a predetermined number of information units for the first frame less than the initial number of information units of the first frame, and the refinement coding stage is also configured to stop the second iteration in case of a negative verification result, or in case of a positive verification result, to carry out a number of additional iterations until a negative verification result is obtained, where the number of additional iterations is 1, 2 ... preferably, the maximum number of iterations is limited by a two-digit number such as a value between 10 and 30 and preferably 20 iterations.In an alternative approach, a check for a maximum number of iterations can be omitted if the non-zero spectral lines are counted first and the number of residual bits is adjusted accordingly for each iteration or for the entire procedure. Therefore, when there are, for example, 20 surviving spectral tuples and 50 residual bits, one can, without any check during the procedure in the encoder or decoder, determine that the number of iterations is three, and in the third iteration, a refinement bit must be calculated or is available in the bitstream for the first ten spectral lines / tuples. This alternative thus does not require a check during iteration processing, since the information about the number of non-zero or surviving audio elements is known after the initial processing stage in the encoder or decoder. Figure 3 illustrates a preferred implementation of the iterative procedure performed by the refinement coding stage 152 of Figure 2 that is made possible due to the fact that, contrary to the other procedures, the number of refinement bits for a frame has been significantly increased for certain frames due to the corresponding reduction of audio data elements for those certain frames. In step 300, the surviving audio data elements are determined. This determination can be performed automatically by operating on the audio data elements that have already been processed by the initial encoding stage 151 in Figure 2. In step 302, the procedure begins on a predefined audio data element, such as the audio data element with the lowest spectral information. In step 304, the bit values ​​for each audio data element are calculated in a predefined sequence, where this predefined sequence is, for example, the sequence from the low spectral values / tuples to the high spectral values / tuples. The calculation in step 304 is performed using a start offset 305 and under the control 314 that the refinement bits are still available.In element 316, the first iteration refinement information units are output, i.e., a bit pattern indicating one bit for each surviving audio data element where the bit indicates whether to add or subtract an offset, i.e., the start offset 305, or alternatively, whether to add or not add the start offset. In step 306, the offset is reduced using a predetermined rule. This predetermined rule might be, for example, that the offset is halved, meaning the new offset is half the original offset. However, other offset reduction rules besides the 0.5 weighting can also be applied. In step 308, the bit values ​​for each element in the predefined sequence are recalculated, but now in the second iteration. As input in the second iteration, the refined elements from the first iteration, illustrated in 307, are introduced. Therefore, for the calculation in step 314, the refinement represented by the first-iteration refinement information units has already been applied, and under the prerequisite that the refinement bits are still available, as indicated in step 314, the second-iteration refinement information units are calculated and provided in 318. In step 310, the offset is reduced again with a predetermined rule so that it is ready for the third iteration, and the third iteration is again based on the refined elements after the second iteration illustrated in 309 and again under the prerequisite that the refinement bits are still available as indicated in 314, the third iteration refinement information units are calculated and provided in 320. Figure 4a illustrates an exemplary frame syntax with the units or bits of information for the first or second frame. A portion of the bit data for the frame is composed of the initial number of bits, i.e., element 400. Additionally, the first iteration refinement bits 316, the second iteration refinement bits 318, and the third iteration refinement bits 320 are also included in the frame.Specifically, according to the frame syntax, the decoder is positioned to identify which bits in the frame are the initial bit number, which bits are the first, second, or third iteration refinement bits (316, 318, 320), and which bits in the frame are other bits (402) such as supplementary information, which may also include, for example, an encoded representation of an overall gain (gg), which may be calculated, for example, directly by controller 200 or which may be influenced, for example, by the controller via controller output information (21). Within sections 316, 318, and 320, a certain sequence of individual information units is given. This sequence is preferably such that the bits in the bit sequence are applied to the initially decoded audio data elements to be decoded.Since it is not useful, with respect to bit rate requirements, to explicitly specify anything regarding the refinement bits of the first, second, and third iterations, the order of the individual bits in blocks 316, 318, and 320 must be the same as the corresponding order of the surviving audio data elements. In view of this, it is preferable to use the same iteration procedure on the encoder side as illustrated in Figure 3 and on the decoder side as illustrated in Figure 8. No specific bit designation or bit association needs to be specified, at least in blocks 316 through 320. Furthermore, the numbers for the initial number of bits on one hand and the remaining number of bits on the other are merely illustrative. Generally, the initial number of bits, which typically encode the most significant bit portion of the audio data element, such as spectral values ​​or tuplets of spectral values, is greater than the iteration refinement bits, which represent the least significant portion of the remaining audio data elements. Additionally, the initial number of bits (400) is generally determined by an entropy encoder or arithmetic encoder, while the iteration refinement bits are determined using a residual or bit encoder operating at a granularity of information units.Although the refinement coding stage does not perform any entropic or similar coding, the encoding of the least significant bit portion of the audio data elements is nevertheless done more efficiently by means of the refinement coding stage, since one can assume that the least significant bit portion of the audio data elements, such as the spectral values, are distributed equally and, therefore, any entropic coding with a variable-length code or an arithmetic code along with a certain context does not introduce any additional advantage, but on the contrary introduces even additional overhead. In other words, for the least significant bit (MSB) portion of the audio data elements, using an arithmetic encoder would be less efficient than using a bit encoder, since the bit encoder doesn't require any specific bit rate for a given context. The intentional reduction of audio data elements, as induced by the driver, not only improves the accuracy of the dominant spectral lines or line tupias, but also provides a highly efficient encoding operation for refining the MSB portions of these audio data elements represented by the arithmetic or variable-length code. In view of that, several advantages are obtained, for example, by implementing the encoder processor 15 of Figure 1 as illustrated in Figure 2 with the initial encoding stage 151 on the one hand and the refinement encoding stage 152 on the other. An efficient two-stage coding scheme is proposed, comprising a first entropic coding stage and a second residual coding stage based on single-bit (non-entropic) coding. / QQI I Π / 77Π7 / 3 / YI1 The scheme employs a low-complexity global gain estimator that incorporates an energy-based bit consumption estimator for the first encoding stage, which features a signal-adaptive background noise adder. The background noise adder effectively transfers bits from the first encoding stage to the second encoding stage for highly tonal signals, while leaving the estimate unchanged for other signal types. This bit shift from an entropic encoding stage to a non-entropic encoding stage is completely efficient for highly tonal signals. Figure 4b illustrates a preferred implementation of the variable quantizer that can be implemented, for example, to perform controlled reduction of audio data elements, preferably in the integrated reduction mode illustrated with respect to Figure 13. For this purpose, the variable quantizer comprises a weighter 155 that receives the (unmanipulated) audio data to be encoded, illustrated in line 12. This data is also fed into the controller 20, and the controller is configured to calculate an overall gain 21, but based on the unmanipulated data as input to the weighter 155, and using signal-dependent manipulation. The overall gain 21 is applied to the weighter 155, and the output of the weighter is fed into a quantizer core 157 that is based on a fixed quantization step size.The variable quantizer 150 is implemented as a controlled weight where control is achieved using the global gain (gg) 21 and the downstream fixed-step quantizer kernel 157. However, other implementations are also possible, such as a quantizer kernel with a variable step size controlled by an output value of controller 20. Figure 5 illustrates a preferred implementation of the audio encoder and, in particular, a certain implementation of the processor 10 of Figure 1. Preferably, the processor comprises a window generator 13 that generates, from the audio input data 11, a windowed time-domain audio data frame using a certain analysis window, which may be, for example, a cosine window. The time-domain audio data frame is fed into a spectrum converter 14, which may be implemented to perform a modified discrete cosine transform (MDCT) or any other transform such as FFT or MDST, or any other time-domain spectrum conversion. Preferably, the window generator operates with some feed control so that an overlay frame generation is performed.In the case of a 50% overlap, the advance value of the window generator is half the size of the analysis window applied by the window generator 13. A (non-quantized) frame of spectral values ​​provided by the spectrum converter is fed into a spectral processor 15 that is implemented to carry out some kind of spectral processing such as performing a temporal noise shaping operation, a spectral noise shaping operation, or any other operation such as a spectral whitening operation, whereby the modified spectral values ​​generated by the spectral processor have a spectral envelope that is flatter than a spectral envelope of the spectral values ​​before processing by means of the spectral processor 15.The audio data to be encoded (per frame) is forwarded via line 12 to the / QQI I Π / 77Π7 / 3 / YI1 encoder processor 15 and controller 20, where controller 20 provides control information via line 21 to encoder processor 15. The encoder processor provides its data to writer 30, which is implemented, for example, as a bitstream multiplexer, and the encoded frames are provided on line 35. Regarding decoder-side processing, see Figure 6. The bitstream provided by block 30 can be fed, for example, directly to the bitstream reader 40 after some form of storage or transmission. Naturally, any other processing between the encoder and decoder can be performed, such as transmission processing according to a wireless transmission protocol like DECT, Bluetooth, or any other wireless transmission protocol. The data fed into the audio decoder shown in Figure 6 is fed into a bitstream reader 40. The bitstream reader 40 reads the data and forwards it to the encoder processor 50, which is controlled by a controller 60.Specifically, the bitstream reader receives encoded data, where the encoded audio data comprises, for a frame, a number of initial frame information units and a number of remaining frame information units. The encoder processor 50 processes the encoded audio data, and the encoder processor 50 comprises an initial decoding stage and a refinement decoding stage as illustrated in Figure 7 in element 51 for the initial decoding stage and in element 52 for the refinement decoding stage, both of which are controlled by means of controller 60.Controller 60 is configured to control refinement decoding stage 52 to use, when refining initially decoded data elements as provided by the initial decoding stage 51 of Figure 7, at least two information units from the number of information units remaining to refine one and the same initially decoded data element.Additionally, controller 60 is configured to control the encoder processor in such a way that the initial encoding stage uses the initial number of frame information units to obtain decoded data elements initially in line connection blocks 51 and 52 in Figure 7, where, preferably, controller 60 receives an indication of the initial number of frame information units on one side and the number of remaining initial frame information units from the bitstream reader 40 as indicated by the input line in block 60 of Figure 6 or Figure 7. Post-processor 70 processes the refined audio data elements to obtain decoded audio data 80 at the output of post-processor 70. In a preferred implementation for an audio decoder corresponding to the audio encoder of Figure 5, the post-processor 70 comprises, as an input stage, a spectral processor 71 that performs an inverse temporal noise shaping operation, or an inverse spectral noise shaping operation, or an inverse spectral whitening operation, or any other operation that reduces some type of processing applied by the spectral processor 15 of Figure 5. The output of the spectral processor is fed into a time converter 72 that operates to perform a conversion from a spectral domain to a time domain, and preferably, the time converter 72 coincides with the spectral converter 14 of Figure 5.The output of the time converter 72 is fed into a superposition-sum stage 73 which performs a superposition / sum operation for a number of superposition frames, such as at least two superposition frames, in order to obtain the decoded audio data 80. Preferably, the superposition-sum stage 73 applies a synthesis window to the output of the time converter 72, where this synthesis window coincides with the analysis window applied by the analysis window generator 13. Furthermore, the superposition operation carried out by block 73 coincides with the block feed operation performed by the window generator 13 in Figure 5. As illustrated in Figure 4a, the number of remaining frame information units comprises calculated values ​​of information units 316, 318, and 320 for at least two sequential iterations in a predetermined order, where, in the mode shown in Figure 4a, three iterations are even illustrated. Furthermore, controller 60 is configured to control the refinement decoding stage 52 to use, for a first iteration, the calculated values ​​such as block 316 for the first iteration according to the predetermined order, and to use, for a second iteration, the calculated values ​​of block 318 for the second iteration in the predetermined order. Subsequently, a preferred implementation of the refinement decoding stage under the control of controller 60 is illustrated with respect to Figure 8. In step 800, the controller or refinement decoding stage 52 of Figure 7 determines the audio data elements to be refined. These audio data elements are generally all the audio data elements provided by block 51 of Figure 7. As indicated in step 802, a start is performed on a predefined audio data element, such as the lowest spectral information.Using a start offset 805, the first-iteration refinement information units received from the bitstream or controller 16—for example, the data in block 316 of Figure 4a—are applied 804 to each element in a predefined sequence, where the predefined sequence extends from a low to a high spectral value / spectral tuple / spectral information. The results are refined audio data elements after the first iteration, as illustrated by line 807. In step 808, bit values ​​are applied to each element in the predefined sequence, where the bit values ​​come from the second-iteration refinement information units, as illustrated in 818. These bits are received from the bitstream reader or controller 60, depending on the specific implementation. The result of step 808 is the refined elements after the second iteration.Again, in step 810, the offset is reduced in line with the predetermined offset reduction rule that has already been applied in block 806. With the offset reduced, the bit values ​​for each element in the predefined sequence are applied as illustrated in 812 using the third-iteration refinement information units received, for example, from the bitstream or controller 60. The third-iteration refinement information units are described in the bitstream in element 320 of Figure 4a. The result of the procedure in block 812 is the refined elements after the third iteration as indicated in 821. This procedure continues until all iteration refinement bits included in the bitstream for a frame have been processed. This is verified by controller 60 via control line 814, which monitors the availability of remaining refinement bits, preferably for each iteration but at least for the second and third iterations processed in blocks 808 and 812. In each iteration, controller 60 monitors the refinement decoding stage to check if the number of information units already read is less than the number of information units in the remaining frame information units for the frame. If the check result is negative, the controller stops the second iteration. If the check result is positive, the controller performs a number of additional iterations until a negative check result is obtained. The number of additional iterations is at least one.Due to the application of similar procedures on the encoder side, as discussed in Figure 3, and on the decoder side, as described in Figure 8, no specific signaling is required. Rather, multi-iteration refinement processing is carried out in a highly efficient manner without any specific overhead. Alternatively, a check for a maximum number of iterations can be omitted if the non-zero spectral lines are counted first, and the number of remaining bits is adjusted accordingly for each iteration. In the preferred implementation, refinement decoding stage 52 is configured to add a offset to the initially encoded data element when a read information data unit from the number of remaining frame information units has a first value, and to subtract an offset from the initially decoded element when a read information data unit from the number of remaining frame information units has a second value. This offset is, for the first iteration, the start offset 805 in Figure 8.In the second iteration, as illustrated in block 808 in Figure 8, a reduced offset, as generated by block 806, is used to add a reduced offset, or second offset, to the result of the first iteration when a data unit read from the number of remaining information units in the frame has a first value, and to subtract the second offset from the result of the first iteration when the data unit read from the number of remaining information units in the frame has a second value. Generally, the second offset is smaller than the first offset, and it is preferred that the second offset be between 0.4 and 0.6 times the first offset, and more preferably 0.5 times the first offset. In a preferred implementation of the present invention using an indirect mode illustrated in Figure 9, no explicit signal characteristic determination is required. Rather, a manipulation value is calculated, preferably using the modality illustrated in Figure 9. For the indirect mode, the controller 20 is implemented as shown in Figure 9. In particular, the controller comprises a control preprocessor 22, a manipulation value calculator 23, a combiner 24, and a global gain calculator 25, which ultimately calculates a global gain for the audio data element reducer 150 of Figure 2, which is implemented as a variable quantizer illustrated in Figure 4b.Specifically, controller 20 is configured to analyze the audio data of the first frame to determine a first control value for the variable quantizer for the first frame, and to analyze the audio data of the second frame to determine a second control value for the variable quantizer for the second frame, the second control value being different from the first control value. The analysis of the audio data of a frame is carried out by means of the manipulation value calculator 23. Controller 20 is configured to perform manipulation of the audio data of the first frame. In this operation, the control processor 20, illustrated in Figure 9, is not present, and therefore the bypass line for block 22 is active. However, when the manipulation is not performed on the audio data of the first or second frame, but is applied to amplitude-related values ​​derived from the audio data of the first or second frame, the control preprocessor 22 is present, and the bypass line is absent. The actual manipulation is performed by the combiner 24, which combines the manipulation value produced from block 23 with the amplitude-related values ​​derived from the audio data of a given frame. At the output of the combiner 24, manipulated data (preferably energy) exists, and based on this manipulated data, a global gain calculator 25 calculates a global gain or at least a control value for the global gain specified in 404.The global gain calculator 25 has to apply restrictions with respect to an allowed bit budget for the spectrum in such a way as to obtain a certain data rate or a certain number of allowed information units. In the direct mode illustrated in Figure 11, the controller 20 comprises an analyzer 201 for determining the signal characteristic per frame, and the analyzer 208 outputs, for example, quantitative signal characteristic information such as pitch information and controls a control value calculator 202 using this quantitative data, preferably. One procedure for calculating the pitch of a frame is to calculate the Spectral Flatness Measure (SFM) of a frame. Any other pitch determination procedure or any other signal characteristic determination procedure can be carried out by means of block 201, and a translation of a certain signal characteristic value to a certain control value must be performed in order to achieve a desired reduction in the number of audio data elements for a frame.The output of the control value calculator 202 for the direct mode in Figure 11 can be a control value for the encoder processor, such as for the variable quantizer, or alternatively, for the initial encoding stage. When a control value is given to the variable quantizer, integrated reduction mode is performed, whereas when the control value is given to the initial encoding stage, separate reduction is performed. Another implementation of separate reduction would be to remove or influence specifically selected unquantized audio data elements present before the actual quantization so that, by means of a certain quantizer, these influenced audio data elements are quantized to zero and thus eliminated for the purpose of entropic coding and subsequent refinement coding. Although the indirect mode of Figure 9 has been shown in conjunction with the integrated reduction, i.e., the global gain calculator 25 is configured to calculate the variable global gain, the manipulated data produced by the combiner 24 can also be used to directly control the initial encoding stage to remove any determined quantized audio data elements in such a way that the smaller quantized data elements or, alternatively, the control value can also be sent to an unillustrated audio data influence stage that influences the audio data before actual quantization using a variable quantization control value that has been determined without data manipulation and thus generally obeys psychoacoustic rules that are, however, intentionally violated by the procedures of the present invention. As illustrated in Figure 11 for direct mode, the controller is configured to determine the first pitch feature as the first signal feature and to determine a second pitch feature as the second signal feature in such a way that a bit budget for the refinement encoding stage is increased in the case of a first pitch feature compared to the bit budget for the refinement encoding stage in the case of a second pitch feature, where the first pitch feature indicates a higher pitch than the second pitch feature. The present invention does not result in a more accurate quantization than that generally obtained by applying a higher overall gain. Rather, this calculation of overall gain based on signal-dependent manipulated data only results in a bit budget shift from the initial encoding stage, which receives a smaller bit budget, to the refinement decoding stage, which receives a higher bit budget. However, this bit budget shift is signal-dependent and is greater for a higher-pitched portion of the signal. Preferably, the control processor 22 in Figure 9 calculates amplitude-related values ​​as a plurality of power values ​​derived from one or more audio values ​​in the audio data. Specifically, these power values ​​are manipulated by combining an identical manipulation value, as determined by the combiner 24. This identical manipulation value, calculated by the manipulation value calculator 23, is then combined with all the power values ​​from the plurality of power values ​​for a frame. Alternatively, as indicated by the branch line, the values ​​obtained by subtracting slightly different terms from the same magnitude (but preferably with randomized signs) or complex manipulation value, or, more generally, the values ​​obtained as samples from a certain scaled normalized probability distribution using the calculated complex or real magnitude of the manipulation value, are summed to all the audio values ​​from a plurality of audio values ​​included in the frame. The procedure performed by the control processor 22, such as the calculation of a power spectrum and resolution reduction, can be included within the global gain calculator 25.Therefore, preferably, background noise is added either directly to the spectral audio values ​​or alternatively to amplitude-related values ​​derived from the per-frame audio data, i.e., the output of the control preprocessor 22. Preferably, the controller processor calculates a reduced-resolution power spectrum corresponding to the use of exponentiation with an exponent value of 2. Alternatively, however, a different exponent value greater than 1 can be used. For example, an exponent value of 3 would represent loudness rather than power. But other exponent values, such as smaller or larger exponent values, can also be used. In the preferred implementation illustrated in Figure 10, the manipulation value calculator 23 comprises a searcher 26 for searching for a maximum spectral value in a frame and at least one of the calculations for an independent signal contribution indicated by element 27 in Figure 10, or a calculator for calculating one or more moments per frame as illustrated by block 28 in Figure 10. Essentially, either block 26 or block 28 is there to provide a signal-dependent influence on the manipulation value for the frame. Specifically, the searcher 26 is configured to search for a maximum value from a plurality of audio data elements or amplitude-related values, or to search for a maximum value from a plurality of reduced-resolution audio data or a plurality of reduced-resolution amplitude-related values ​​for the corresponding frame.The actual calculation is done by means of block 29 using the output of blocks 26, 27 and 28, where blocks 26, 28 actually represent a signal analysis. Preferably, the independent signal contribution is determined by means of a bit rate for an actual encoder session, a frame duration, or a sampling frequency for an actual encoder session. Furthermore, the calculator 28 for calculating one or more moments per frame is configured to calculate a signal-dependent weighting value derived from at least a first sum of magnitudes of the audio data or reduced-resolution audio data within the frame, the second sum of magnitudes of the audio data or reduced-resolution audio data within the frame multiplied by an index associated with each magnitude, and the quotient of the second sum and the first sum. In a preferred implementation carried out by the global gain calculator 25 in Figure 9, an estimate of the required bits is calculated for each energy value depending on the energy value and the candidate value for the actual control value. The estimated bits required for the energy values ​​and the candidate value for the control value are accumulated, and it is checked whether an accumulated bit estimate for the candidate value for the control value meets an allowed bit consumption criterion, as illustrated, for example, in Figure 9 as the bit budget for the spectrum entered into the global gain calculator 25.If the allowed bit consumption criterion is not met, the candidate value for the control value is modified, and the calculation of the required bit estimate, the accumulation of the required bit rate, and the verification of compliance with the allowed bit consumption criterion are repeated for a modified candidate value for the control value. As soon as such an optimal control value is found, this value is provided on line 404 of Figure 9. The following are examples of preferred modalities. Detailed description of the encoder (for example, Figure 5) Notation We denote by fsla the underlying sampling frequency in Hz, by Nmsla the underlying frame duration in milliseconds, and by br the underlying bit rate in bits per second. Derivation of residual spectrum (e.g., preprocessor 10) The modality operates on a real residual spectrum Xf(k),k = 0..N — 1, which is generally derived by a time-frequency transform such as an MDCT followed by psychoacoustically motivated modifications such as temporal noise shaping (TNS) to remove the temporal structure and spectral noise shaping (SNS) to remove the spectral structure. For audio content with a slowly varying spectral envelope, the envelope of the residual spectrum Xf(k) is therefore flat. Global profit estimate (e.g., Figure 9) / QQ L ίΠ / 77Π7 / 3 / ΥΙ Spectrum quantization is controlled by an overall gain by means of: ÍXr(k)\ Xq(Je) = round I-----I \9glob / The initial global gain estimate (element 22 of Figure 9) derived from the power spectrum X(fc)2 after resolution reduction by a factor of 4, PXip(k) = A / 4 / c)2+ Xf^k + l)2+ Xf(Ak + 2)2+ Xf(4k + 3)zy an adaptive background noise to the signal N(Xf) that is given by: N^Xf) = max|Xy( / c)| * 2~reaBlts~lowBlts. (for example, element 23 of Figure 9). The regBits parameter depends on the bit rate, frame length, and sampling frequency and is calculated as: / QQLLn / ZZnZ / q / Yli regBits = [ ζθθ] + C(Nms>fs) (for example, element 27 of Figure 10) with C(Nms,fs~) as specified in the following table. Nms ' fs 48000 96000 2.5 -6 -6 5 0 0 10 2 5 The lowBlts parameter depends on the center of mass of the absolute values ​​of the residual spectrum and is calculated as: lowBits = (2Nms- min ZNms^ (for example element 28, Figure 10) where N—l Mo= ^IWl k=0 and Nl M1=^k \Xf(k)\ k=0 are moments of the absolute spectrum. The overall profit is estimated as follows: sgind^gsofr 9glob—W28 of the values: E(k) = 10 \ogw(PXLp(k) + + 2-31), (for example, output of combiner 24 in Figure 9). where ggOff is a bit rate and sampling frequency dependent offset. It should be noted that adding the background noise term N^X^ to PXip(¡í) gives the expected result of adding background noise corresponding to the residual spectrum Xf(k\ for example, randomly adding or subtracting the term 0.5 yÍN(Xf) to each spectral line, before calculating the power spectrum. Estimates based on pure power spectrum can already be found, for example, in the 3GPP EVS codec (3GPP TS 26.445, section 5.3.3.2.8.1). In the modes, the background noise N(χ / ) is summed. The background noise is adaptive to the signal in two ways. First, it scales with the maximum amplitude of Xf. Therefore, the impact on the energy of a flat spectrum, where all amplitudes are close to the maximum amplitude, is very small. But for highly tonal signals, where the spectrum, and by extension also the residual spectrum features, exhibits a number of strong peaks, the total energy is significantly increased, which increases the bit estimate in the overall gain calculation as described later. Secondly, background noise is reduced by the lowBits parameter if the spectrum exhibits a low center of mass. In this case, low-frequency content is dominant, so the loss of high-frequency components is probably not as critical as it would be for high-pitched tonal content. The actual estimate of the overall gain is carried out (e.g., block 25 in Figure 9) by means of a low-complexity bisection search as described in the C code below, where nbitSspec denotes the bit budget for encoding the spectrum. The bit consumption estimate (accumulated in the variable tmp) is based on the energy values ​​E(k), taking into account a context dependency on the arithmetic encoder used for stage 1 encoding. fac = 256; 33ind=255, for (iter = 0; iter < 8; iter++) { fac »= 1; 99índ fac, tmp- 0; iszero = 1; for (i = N / 4-1; i >= 0; i-) if (E[¡]*28 / 20 < {ggind+ggoff)) { if (iszero == 0) { tmp += 2.7*28 / 20; }}else { ¡f ^ggmd+ggoff} < e[¡]*28 / 20 - 43*28 / 20) { tmp += 2*E[¡]*28 / 20 - 2\ggind+ggoff) - 36*28 / 20; else { tmp += E[¡]*28 / 20 - (ggínd+ggoff) + 7*28 / 20;} iszero = 0; } if (tmp > nbitSspec*] .4*28 / 20 && iszero == 0) 99ind +=fac,} / QQLLn / ZZnZ / q / Yli Residual coding (for example, Figure 3) Residual coding utilizes the excess bits available after the arithmetic coding of the quantized spectrum Xq(k). Let B denote the number of excess bits and let K denote the number of non-zero coefficients encoded by Xq(k). Furthermore, let khi = 1..K denote the enumeration of these non-zero coefficients from lowest to highest frequency. The residual bits bi(J) (taking values ​​0 and 1) for the coefficient kt are calculated to minimize the error. (ni\ * 2-7-1- X^. j=i / This can be done interactively by testing whether: (n—1\ Χ^ύ ~ * 2-7-1- XfW > 0.(1) J=1 / If (1) is true, then the nth residual b¡(n) for the coefficient ki is set to 0; otherwise, it is set to 1. The calculation of residual bits is carried out by calculating a first residual bit for each k, then a second bit, and so on until all the residual bits are used up or until a maximum number nmax of iterations has been performed. This leaves: nÉ= minnmax residual bits for the coefficient Xatkj. This residual coding scheme improves upon the residual coding scheme applied in the 3GPP EVS codec, which spends at most one bit per non-zero coefficient. The calculation of residual bits with nmax= 20 is illustrated by the following pseudocode, where gg denotes the overall gain. iter = 0; nbits_residual = 0; offset = 0.25; while (nbits_residual < nbits_residual_max && iter < 20) { k = 0; while (k< Ne&& nbits_residual < nbits_residual_max) { f (YJk] != 0) { if (Xz[kj >= XQ[k]'gg) { res_bits[nbits_residual] = 1; Xf[k] -= offset * gg; } else { res_bits[nbits_residual] = 0; Xz[k] += offset * gg; } nbits_residual++; } k++; } iter++; offset / = 2; Decoder description (for example, Figure 6) In the decoder, the entropy-coded spectrum is obtained by entropic decoding. The remaining bits are used to refine this spectrum, as demonstrated by the following pseudocode (see also, for example, Figure 8). ter = n = 0; offset = 0.25; while (ter < 20 && n < nResBits) { k = 0; while (k < Ne&& n < nResBits) { _ f (Xq[k] !=0) { if (resBits[n++] == 0) { Xq[k] -= offset; } else { _ Xq[k] += offset;}} k++; } iter ++; offset / = 2; / QQ L iΠ / 77Π7 / 3 / YI The decoded residual spectrum is given by 9glob Xq(k')· Conclusions: • An efficient two-stage coding scheme is proposed, comprising a first entropic coding stage and a second residual coding stage based on single-bit (non-entropic) coding. • The scheme employs a low-complexity global gain estimator that incorporates an energy-based bit consumption estimator for the first encoding stage, which features a signal-adaptive background noise adder. • The background noise adder effectively transfers bits from the first encoding stage to the second encoding stage for highly tonal signals, while leaving the estimate unchanged for other signal types. It is argued that this bit shift from an entropic encoding stage to a non-entropic encoding stage is completely efficient for highly tonal signals. Figure 12 illustrates a procedure for reducing the number of audio data elements in a signal-dependent manner rather than using separate reduction. In step 901, quantization is performed using unmanipulated information such as the overall gain as calculated from the unmanipulated signal data. For this purpose, the (total) bit budget for the audio data elements is required, and the output of block 901 yields quantized data elements. In block 902, the number of audio data elements is reduced by removing a (controlled) number of the smaller audio data elements, preferably based on a signal-dependent control value.At the output of block 902, one has obtained a reduced number of data elements, and in block 903, the initial encoding stage is applied, and with the bit budget for the remaining bits due to controlled reduction, a refinement encoding stage is applied as illustrated in 904. Alternatively to the procedure in Figure 12, the 902 reduction block can also be performed before actual quantization using a global gain value or, more commonly, a specific quantizer step size determined using unmanipulated audio data. Therefore, this reduction of audio data elements can also be carried out in the unquantized domain by setting certain, preferably small, values ​​to zero or by weighting certain values ​​with weighting factors that ultimately result in zero-quantized values. In the separate reduction implementation, an explicit quantization step is performed on one hand and an explicit reduction step on the other, with control for the specific quantization being implemented without data manipulation. Conversely, Figure 13 illustrates the integrated reduction mode according to one embodiment of the present invention. In block 911, the manipulated information is determined by means of controller 20, such as, for example, the overall gain illustrated in the output of block 25 in Figure 9. In block 912, quantization of the unmanipulated audio data is performed using the manipulated overall gain, or, more generally, the manipulated information calculated in block 911. At the output of the quantization procedure in block 912, a reduced number of audio data elements are obtained, which are initially encoded in block 903 and further refined in block 904. Due to the signal-dependent reduction of the audio data elements, residual bits remain for at least one complete iteration and for at least a portion of a second iteration, and preferably for even more than two iterations.A shift of the bit budget is carried out from the initial encoding stage to the refinement encoding stage according to the present invention and in a signal-dependent manner. The present invention can be implemented in at least four different modes. The determination of the control value can be done in direct mode with explicit signal characteristic determination or in indirect mode without explicit signal characteristic determination but with the addition of signal-dependent background noise to the audio data or to the derived audio data as an example for manipulation. Simultaneously, the reduction of audio data elements is performed in an integrated or separate manner. An indirect determination and integrated reduction, or an indirect generation of the control value and a separate reduction, can also be carried out. Additionally, a direct determination along with an integrated reduction, or a direct determination of the control value along with a separate reduction, can also be performed.For low-efficiency purposes, an indirect determination of the control value is preferred, along with an integrated reduction of audio data elements. It should be mentioned here that all the alternatives or aspects discussed above and all the aspects defined by the independent claims in the following claims may be used individually, that is, without any other alternative or objective besides the alternative, objective, or independent claim being considered. However, in other embodiments, two or more of the alternatives, aspects, or independent claims may be combined with each other, and in still other embodiments, all aspects, alternatives, and all independent claims may be combined with each other. An inventively encoded audio signal can be stored on a digital storage medium or a non-transient storage medium, or it can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet. Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, the aspects described in the context of a method step also represent a description of a corresponding block, element, or feature of a corresponding apparatus. Depending on certain implementation requirements, the embodiments of the invention can be implemented in hardware or software. Implementation can be carried out using a digital storage medium, for example, a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or flash memory, having electronically readable control signals stored therein, which cooperates (or is capable of cooperating) with a programmable computer system in such a way as to carry out the respective method. Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computing system in such a way as to carry out one of the methods described in this document. Generally, the embodiments of the present invention can be implemented as a computer program product with program code, the program code being operative to carry out one of the methods when the computer program product is executed on a computer. The program code can be stored, for example, on a machine-readable medium. Other modalities include the computer program to carry out one of the methods described in this document, stored on a machine-readable carrier or non-transient storage medium. In other words, a modality of the inventive method is, therefore, a computer program that has program code to carry out one of the methods described in this document, when the computer program is executed on a computer. / QQLLn / ZZnZ / q / Yli An additional embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, engraved thereon, the computer program for carrying out one of the methods described in this document. An additional embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for carrying out one of the methods described herein. The data stream or the sequence of signals can be configured, for example, to be transmitted via a data communication connection, such as the internet. An additional modality comprises a processing means, for example a computer, or a programmable logic device, configured or adapted to carry out one of the methods described in this document. An additional modality comprises a computer that has installed on it the computer program to carry out one of the methods described in this document. In some embodiments, a programmable logic device (for example, a field-programmable gate array) can be used to carry out some or all of the functionalities of the methods described in this document. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to carry out one or more of the methods described in this document. Generally, the methods are preferably implemented by means of any hardware device. The embodiments described above are merely illustrative of the principles of the present invention. It is understood that the modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore intended that the scope of the imminent patent claims be limited only by the specific details presented herein for the purpose of describing and explaining the embodiments.

Claims

1. An audio encoder for encoding audio input data (11), comprising: a preprocessor (10) for preprocessing the audio input data (11) to obtain audio data to be encoded; an encoder processor (15) for encoding the audio data to be encoded; and a controller (20) for controlling the encoder processor (15) such that, depending on a first signal characteristic of a first frame of the audio data to be encoded, a number of audio data elements of the audio data to be encoded by means of the encoder processor (15) for the first frame is reduced compared to a second signal characteristic of a second frame, and a first number of information units used to encode the reduced number of audio data elements for the first frame is further improved compared to a second number of information units for the second frame.

2. The audio encoder according to claim 1, wherein the encoder processor (15) comprises an initial encoding stage (151) and a refinement encoding stage (152), wherein the controller (20) is configured to reduce the number of audio data elements encoded by means of the initial encoding stage (151) for the first frame, wherein the initial encoding stage (151) is configured to encode the reduced number of audio data elements for the first frame using a number of initial first-frame information units, and wherein the refinement encoding stage (152) is configured to use a number of remaining first-frame information units for refinement encoding of the reduced number of audio data elements for the first frame,where the number of initial information units of the first frame added to the number of remaining information units of the first frame results in a predetermined number of information units for the first frame.

3. The audio encoder according to claim 2, wherein the controller (20) is configured to reduce the number of audio data elements encoded by means of the initial encoding stage (151) for the second frame to a higher number of audio data elements compared to the first frame, wherein the initial encoding stage (151) is configured to encode the reduced number of audio data elements for the second frame using a number of initial information units from the second frame, the number of initial information units from the second frame being greater than the number of initial information units from the first frame, and wherein the refinement encoding stage (152) is configured to use a number of remaining information units from the second frame for refinement encoding of the reduced number of audio data elements for the second frame.where the number of initial information units of the second frame added to the number of remaining information units of the second frame results in the predetermined number of information units for the first frame.

4. The audio encoder according to any of the preceding claims, wherein the encoder processor (15) comprises an initial encoding stage (151) and a refinement encoding stage (152), wherein the initial encoding stage (151) is configured to encode the reduced number of audio data elements for the first frame using a number of initial first-frame information units, wherein the refinement encoding stage (152) is configured to use a number of remaining first-frame information units for refinement encoding of the reduced number of audio data elements for the first frame, wherein the number of initial first-frame information units added to the number of remaining first-frame information units results in a predetermined number of information units for the first frame,and wherein the controller (20) is configured to control the encoder processor (15) such that the refinement encoding stage (152) performs refinement encoding of at least one of the reduced number of audio data elements of the first frame using at least two information units, or such that the refinement encoding stage (152) performs refinement encoding of more than 50 percent of the reduced number of audio data elements using at least two information units for each audio data element, or wherein the controller (20) is configured to control the encoder processor (15) such that the refinement encoding stage (152) performs refinement encoding of all the audio data elements of the second frame using fewer than two information units,or in such a way that the refinement coding stage carries out refinement coding of less than 50 percent of the reduced number of audio data elements using at least two units of information for each audio data element.

5. The audio encoder according to any of the preceding claims, wherein the encoder processor (15) comprises an initial encoding stage (151) and a refinement encoding stage (152), wherein the initial encoding stage (151) is configured to encode the reduced number of audio data elements for the first frame using a number of initial first-frame information units, wherein the refinement encoding stage (152) is configured to use a number of remaining first-frame information units for refinement encoding of the reduced number of audio data elements for the first frame, wherein the refinement encoding stage (152) is configured to iteratively allocate (300,302) the number of information units remaining from the first frame to the reduced number of audio data elements in at least two sequentially performed iterations, to calculate (304, 308, 312) values ​​of the information units allocated for said at least two sequentially performed iterations and to input (316, 318, 320) the calculated values ​​ / QQI I n / 77O7 / =l / Yl· of the information units for said at least two sequentially performed iterations into an output frame encoded in a predetermined order.

6. The audio encoder according to claim 5, wherein the refinement encoding stage (152) is configured to sequentially calculate (304) one unit of information for each audio data element of the reduced number of audio data elements for the first frame in an order from low-frequency information for the audio data element to high-frequency information for the audio data element in a first iteration, wherein the refinement encoding stage (152) is configured to sequentially calculate (308) one unit of information for each audio data element of the reduced number of audio data elements for the first frame in an order from low-frequency information for the audio data element to high-frequency information for the audio data element in a second iteration,and wherein refinement encoding stage 152 is configured to check (314) if a number of already allocated information units is lower than a predetermined number of information units for the first frame less than the initial number of information units of the first frame and to stop the second iteration in case of a negative check result, or in case of a positive check result, to carry out (312) a number of additional iterations until a negative check result is obtained, the number of additional iterations being at least one, or wherein refinement encoding stage (152) is configured to count a number of non-zero audio elements,and to determine the number of iterations of the number of non-zero audio elements and a predetermined number of information units for the first frame less than the initial number of information units of the first frame.

7. The audio encoder according to any of the preceding claims, wherein the encoder processor (15) comprises an initial encoding stage (151) and a refinement encoding stage (152), wherein the initial encoding stage (151) is configured to encode a number of more significant information units for each audio data element of the reduced number of audio data elements for the first frame using a number of initial first frame information units, the number being greater than one, and wherein the refinement encoding stage (152) is configured to use a number of remaining first frame information units to encode a number of less significant information units for each audio data element of the reduced number of audio data elements for the first frame,the number being greater than one for at least one audio data element of the reduced number of audio data elements for the first frame.

8. The audio encoder according to any of the preceding claims, wherein the first signal feature is a first pitch value, wherein the second signal feature is a second pitch value, and wherein the first pitch value indicates a higher pitch than the second pitch value, and wherein the controller (20) is configured to reduce the number of audio data elements for the first frame to a first number that is less than the number of audio data elements for the second frame, and to increase an average number of information units used to encode each audio data element of the reduced number of audio data elements of the first frame so that it is greater than an average number of information units used to encode each audio data element of the reduced number of audio data elements of the second frame.

9. The audio encoder according to any of the preceding claims, wherein the encoder processor (15) comprises: a variable quantizer (150) for quantizing the audio data of the first frame to obtain quantized audio data for the first frame and for quantizing the audio data of the second frame to obtain quantized audio data for the second frame; an initial encoding stage (151) for encoding the quantized audio data of the first frame or the second frame; a refinement encoding stage (152) for encoding residual data from the first frame and the second frame; wherein the controller (20) is configured to analyze (26, 28) the audio data of the first frame to determine a first control value (21) for the variable quantizer (150) for the first frame and to analyze (26,28) the audio data of the second frame to determine a second control value for the variable quantizer (150) for the second frame, the second control value being different from the first control value (21), and wherein the controller (20) is configured to perform (23, 24) a manipulation of the audio data of the first frame or the second frame or of amplitude-related values ​​derived from the audio data of the first frame or the second frame depending on the audio data to determine the first control value (21) or the second control value (21), and wherein the variable quantizer (150) is configured to quantize the audio data of the first frame or the second frame without the manipulation.

10. The audio encoder according to any of claims 1 to 9, wherein the encoder processor (15) comprises: a variable quantizer (150) for quantizing the audio data of the first frame to obtain quantized audio data for the first frame and for quantizing the audio data of the second frame to obtain quantized audio data for the second frame; an initial encoding stage (151) for encoding the quantized audio data of the first frame or the second frame; a refinement encoding stage (152) for encoding residual data from the first frame and the second frame; wherein the controller (20) is configured to analyze the audio data of the first frame to determine a first control value (21) for the variable quantizer (150),for the initial encoding stage (151) or for an audio data element reducer (150) for the first frame and for analyzing the audio data of the second frame to determine a second control value for the variable quantizer (150), for the initial encoding stage (151) or for an audio data element reducer (150) for the second frame, the second control value being different from the first control value, and wherein the controller (20) is configured (201) to determine a first pitch characteristic as the first signal characteristic to determine the first control value,and a second pitch feature as the second signal feature to determine the second control value so as to increase the bit budget for the refinement encoding stage (152) in the case of a first pitch feature compared to the bit budget for the refinement encoding stage (152) in the case of a second pitch feature, wherein the first pitch feature indicates a higher pitch than the second pitch feature.

11. The audio encoder according to claim 9 or 10, wherein the initial encoding stage (151) is an entropic encoding stage for encoding by entropy, or the refinement encoding stage (152) is a residual or binary encoding stage for encoding residual data from the first frame and the second frame.

12. The audio encoder according to any of claims 9 to 11, wherein the controller (20) is configured to determine the first or second control value such that a first budget of information units for the initial encoding stage (151) is less than or equal to a predefined value, and wherein the controller (20) is configured to derive a second budget of information units for the refinement encoding stage (152) using the first budget of information units and the maximum number of information units for the first or second frame or the predefined value.

13. The audio encoder according to any of claims 9 to 12, wherein the controller (20) is configured to calculate (22) amplitude-related values ​​as a plurality of power values ​​derived from one or more audio values ​​of the audio data and manipulate (24) the power values ​​using a sum of an identical manipulation value to all the power values ​​of the plurality of power values, or wherein the controller (20) is configured to: randomly add or subtract (24) an identical manipulation value from all the audio values ​​of a plurality of audio values ​​included in the frame, or add or subtract values ​​obtained by the same magnitude of manipulation value but preferably with randomized signs,or adding or subtracting values ​​obtained by subtracting slightly different mines of the same magnitude, adding or subtracting values ​​obtained as samples from a scaled normalized probability distribution using the calculated complex or real magnitude of the manipulation value, or wherein the controller (20) is configured to calculate (22) amplitude-related values ​​using / QQ L ίΠ / 77Π7 / 3 / YΙ( an exponentiation of the audio data from the first or second frame or of reduced-resolution audio data from the first or second frame with an exponent value, the exponent value being greater than 1., 14. The audio encoder according to any of claims 9 to 13, wherein the controller (20) is configured to calculate (23) a manipulation value for manipulation using a maximum value (26) of the plurality of audio data or amplitude-related values ​​or using a maximum value of a plurality of reduced-resolution audio data or a plurality of reduced-resolution amplitude-related values ​​for the first or second frame.

15. The audio encoder according to any of claims 9 to 14, wherein the controller (20) is configured to calculate (23) a manipulation value for manipulation using additionally a signal-independent weighting value (27), the signal-independent weighting value depending on at least one of a bit rate for the first or second frame, a frame duration, and a sampling frequency.

16. The audio encoder according to any of claims 9 to 15, wherein the controller (20) is configured to calculate (23, 29) a manipulation value for manipulation using a signal-dependent weighting value derived from at least one of a first sum of magnitudes of the audio data or reduced-resolution audio data within the frame, a second sum of magnitudes of the audio data or reduced-resolution audio data within the frame multiplied by an index associated with each magnitude, and a ratio of the second sum and the first sum.

17. The audio encoder according to any of claims 9 to 16, wherein the controller (20) is configured to calculate (29) the manipulation value for manipulation based on the following equation: N(Xf) = max|Xr(fc)| * 2-re9Bits-i™Bits where k is a frequency index, where Xf(k) is an audio data value for frequency index k before quantization, where max is the maximum function, where regBits is a first signal-independent weighting value, and where lowBits is a second signal-dependent weighting value.

18. The audio encoder according to any of the preceding claims, wherein the processor (10) further comprises: a time-frequency converter (14) for converting audio data in the time domain into spectral values ​​of a frame; and a spectral processor (15) for calculating modified spectral values ​​having a spectral envelope that is flatter than a spectral envelope of the spectral values, wherein the modified spectral values ​​represent the audio data of the first or second frame to be encoded by means of the encoder processor (15).

19. The audio encoder according to claim 18, wherein the spectral processor (15) is configured to perform at least one of a temporal noise shaping operation, a spectral noise shaping operation, and a spectral whitening operation.

20. The audio encoder according to any of claims 9 to 19, wherein the controller (20) is configured to calculate the control value using a plurality of energy values ​​as amplitude-related values ​​for the frame, wherein each energy value is derived (22, 23, 24) from a power value as an amplitude-related value and a signal-dependent manipulation value for manipulation.

21. The audio encoder according to claim 20, wherein the controller (20) is configured to: calculate a required bit estimate for each power value depending on the power value and a candidate value for the control value; accumulate the required bit estimates for the power values ​​and the candidate value for the control value; verify whether an accumulated bit estimate for the candidate value for the control value meets an allowable bit consumption criterion; modify the candidate value for the control value if the allowable bit consumption criterion is not met; and repeat the calculation of the required bit estimate, the accumulation of the required bit rate, and the verification until compliance with the allowable bit consumption criterion is found for a modified candidate value for the control value.

22. The audio encoder according to claim 20 or 21, wherein the controller (20) is configured to calculate the plurality of energy values ​​based on the following equation: E(k) = 101og10( / %p(fc) + N(Xf) + 2-31), wherein E(k) is an energy value for an index k, wherein PXip(k) is a power value for an index k as the amplitude-related value, and wherein N(Xf) is the signal-dependent manipulation value.

23. The audio encoder according to any of claims 9 to 22, wherein the controller (20) is configured to calculate the first or second control value based on an estimate of accumulated information units required for each manipulated audio data value or manipulated amplitude-related value.

24. The audio encoder according to any of claims 9 to 23, wherein the controller (20) is configured to manipulate in such a way that, due to the manipulation, a bit budget for the initial encoding stage (151) is increased or a bit budget for the refinement encoding stage (152) is decreased.

25. The audio encoder according to any of claims 9 to 24, wherein the controller (20) is configured to manipulate in such a way that a manipulation results in a higher residual encoding stage bit budget for a signal having a first tone compared to a signal having a second tone, wherein the second tone is lower than the first tone.

26. The audio encoder according to any of claims 9 to 25, wherein the controller (20) is configured to manipulate in such a way that an energy of the audio data, from which a bit budget for the initial encoding stage (151) is calculated, is increased with respect to the energy of the audio data to be quantized by means of the variable quantizer (150).

27. The audio encoder according to any of the preceding claims, wherein the encoder processor (15) comprises a variable quantizer (150) for quantizing the audio data of the first frame to obtain quantized audio data for the first frame and for quantizing the audio data of the second frame to obtain quantized audio data for the second frame, wherein the controller (20) is configured to calculate an overall gain for either the first or second frame, and wherein the variable quantizer (150) comprises: a weight (155) for ripple with the overall gain; and a quantizer core (157) having a fixed quantization step size.

28. The audio encoder according to any of the preceding claims, wherein the encoder processor (15) comprises an initial encoding stage (151) and a refinement encoding stage (152), wherein the refinement encoding stage (152) is configured to compute refinement bits for quantized audio values ​​in a plurality of iterations, wherein, in each iteration, a refinement bit indicates a different quantity, or wherein a refinement bit in a smaller iteration indicates a higher quantity than a refinement bit in a larger iteration, or wherein the quantity is a fractional quantity that is a fraction of a quantizer step size indicated by the control value.

29. The audio encoder according to any of the preceding claims, wherein the encoder processor (15) comprises a refinement encoding stage (152), wherein the refinement encoding stage (152) is configured (304, 308, 312) to: perform iterative processing having at least two iterations, verify whether a quantized audio value or the quantized audio value together with a first potential quantity associated with a refinement bit for the quantized audio value in a first iteration, added to or subtracted from a second quantity for the second iteration when weighted by means of an overall gain is greater or less than an unquantized audio value, and set a refinement bit for the second iteration depending on a result of the verification.

30. The audio encoder according to any of the preceding claims, wherein the encoder processor (15) comprises a variable quantizer (150) and a refinement encoding stage (152), wherein the refinement encoding stage (152) is configured to compute a refinement bit only for audio values ​​that are not zero-quantized by the variable quantizer (150). / QQI I n / 77O7 / =l / Yl· 31. The audio encoder according to any of the preceding claims, wherein the controller (20) is configured to reduce the impact of manipulation on audio data having a center of mass at a lower frequency, and wherein an initial encoding stage (151) of the encoder processor (15) is configured to remove high-frequency spectral values ​​from the audio data if it is determined that a bit budget for the first or second frame is insufficient to encode the quantized audio data of the frame.

32. The audio encoder according to any of the preceding claims, wherein the controller (20) is configured to perform a bi-section search for each frame individually using manipulated spectral energy values ​​for the first or second frame as manipulated amplitude-related values ​​for the first or second frame.

33. A method for encoding audio input data, comprising: preprocessing the audio input data (11) to obtain audio data to be encoded; encoding the audio data to be encoded; and controlling the encoding such that, depending on a first signal characteristic of a first frame of the audio data to be encoded, a number of audio data elements of the audio data to be encoded for the first frame is reduced compared to a second signal characteristic of a second frame, and a first number of information units used to encode the reduced number of audio data elements for the first frame is further improved compared to a second number of information units for the second frame.

34. The method according to claim 13, wherein the encoding comprises: variably quantizing the audio data of a frame to obtain quantized audio data; entropy-coding the quantized audio data of the frame; and encoding residual data of the frame; wherein the control comprises determining a control value for the variable quantization, the determination comprising: analyzing the audio data of the first or second frame;and performing manipulation of the audio data of the first or second frame or amplitude-related values ​​derived from the audio data of the first or second frame depending on the audio data to determine the control value, wherein the variable quantization quantizes the audio data of the frame without manipulation, or wherein the control comprises determining a first or second pitch characteristic of the audio data and determining the control value so as to increase a bit budget for residual coding in the case of the first pitch characteristic compared to the bit budget for the residual coding stage in the case of the second pitch characteristic, wherein the first pitch characteristic indicates a higher pitch than the second pitch characteristic.

35. A computer-readable means for encoding audio input data comprising the method according to claim 33 or claim 34.