Audio encoder with signal dependent number and precision control, audio decoder and related methods and computer programs

By using a two-stage encoder processor and a signal adaptive noise floor adder, the problem of inaccurate low-bit loss of high-pitched signals is solved, thereby improving coding efficiency and audio quality.

CN114258567BActive Publication Date: 2026-01-09FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080058343.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-17
Filing Date
2020-06-10
Publication Date
2026-01-09
Estimated Expiration
2040-06-10

AI Technical Summary

Technical Problem

Existing audio encoders often fail to accurately estimate bit consumption when processing high-pitched signals, leading to either degraded audio quality or excessive bit consumption, thus failing to effectively utilize the bit budget.

Method used

A two-stage encoder processor is used. The audio data is preprocessed and encoded in a signal characteristic-dependent manner through the preprocessor and encoder processor. The controller adjusts the operation of the encoder processor according to the signal characteristics, reducing the number of audio data items and enhancing the encoding efficiency of information units. The bit consumption is optimized by using a signal adaptive background noise adder.

Benefits of technology

It improves the encoding efficiency of audio encoders under high-tone signals, reduces the inaccuracy of bit consumption, maintains audio quality, and optimizes encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114258567B_ABST
    Figure CN114258567B_ABST
Patent Text Reader

Abstract

An audio encoder for encoding audio input data (11) comprises a pre-processor (10) for pre-processing the audio input data (11) to obtain audio data to be encoded, an encoder processor (15) for encoding the audio data to be encoded, and a controller (20) for controlling the encoder processor such that, depending on a first signal characteristic of a first frame of the audio data to be encoded, the number of audio data items to be encoded by the encoder processor (15) for the first frame is reduced compared to a second signal characteristic of a second frame, and a first number of information units used for encoding the reduced number of audio data items for the first frame is more strongly enhanced compared to a second number of information units used for the second frame.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to audio signal processing, and in particular to audio encoders / decoders applying signal dependent number and precision control. BACKGROUND

[0002] Modern transform-based audio encoders apply a series of psychoacoustic motivated processes to a spectral representation of an audio segment (frame) to obtain a residual spectrum. This residual spectrum is quantized and the coefficients are encoded using an entropy coder.

[0003] In this approach, the quantization step size, typically via a global gain control, has a direct impact on the bit consumption of the entropy coder and needs to be selected in such a way that the usually limited and often fixed bit budget is met. Since the bit consumption of the entropy coder, and in particular of an arithmetic coder, is not exactly known before encoding, computing the optimal global gain can only be done in a closed loop iteration of quantization and encoding. However, this is not feasible under certain complexity constraints, as arithmetic coding has a significant computational complexity.

[0004] State-of-the-art encoders, as can be seen in the 3GPP EVS codec, thus typically feature a bit consumption estimator for deriving a first global gain estimate, which typically operates on the power spectrum of the residual signal. Depending on the complexity constraints, this can be followed by a rate loop to optimize the first estimate. Using this estimate alone or in combination with a very limited correction capability reduces the complexity, but also the accuracy, resulting in a significant under- or overestimation of the bit consumption.

[0005] An overestimation of the bit consumption leads to excess bits after the first encoding stage. State-of-the-art encoders use these excess bits to optimize the quantization of the encoding coefficients in a second encoding stage, referred to as residual coding. Residual coding fundamentally differs from the first encoding stage, as it operates on a bit granularity and thus does not incorporate any entropy coding. In addition, residual coding is typically only applied at frequencies with a non-zero quantized value, leaving a blind zone that is not further improved.

[0006] On the other hand, an underestimation of the bit consumption inevitably leads to a partial loss of spectral coefficients, typically the highest frequencies. In state-of-the-art encoders, this effect is mitigated by applying a noise substitution at the decoder, which is based on the assumption that high frequency content is typically noisy.

[0007] In this setup, it is apparent that as many signals as possible need to be encoded in the first encoding step, which uses entropy coding and is thus more efficient than the residual encoding step. Therefore, one would like to select a global gain with a bit estimate as close as possible to the available bit budget. While the power spectrum based estimator works well for most audio content, it can cause problems with high pitched signals, where the first level estimate is dominated by the uncorrelated side lobes of the filter bank's frequency decomposition, while important components are lost due to underestimation of the bit consumption. SUMMARY

[0008] It is an object of the present application to provide an improved concept for audio encoding or decoding, which is efficient nevertheless and which achieves good audio quality.

[0009] This object is achieved by the audio encoder of technical solution 1, the method of encoding audio input data of technical solution 33 and the audio decoder of technical solution 35, the method of decoding encoded audio data of technical solution 41 or the computer program of technical solution 42.

[0010] The present application is based on the finding that for improving efficiency, especially with respect to one aspect bit rate and another aspect audio quality, a signal dependent change with respect to typical situations given by psychoacoustic considerations is necessary. When expecting average results, typical psychoacoustic models or psychoacoustic considerations are averaged over all signal classes, i.e. over all audio signal frames, without respect to their signal characteristics, to produce good audio quality at low bit rates. However, it has been found that for certain signal classes or for signals with certain signal characteristics, such as almost tonal signals, a direct psychoacoustic control of a simple psychoacoustic model or encoder only produces suboptimal results with respect to audio quality, when bit rate is kept constant, or with respect to bit rate, when audio quality is kept constant.

[0011] Thus, in order to solve this drawback of typical psychoacoustic considerations, in the context of an audio encoder, the present invention provides: a pre-processor for pre-processing audio input data to obtain audio data to be encoded; and an encoder processor for encoding the audio data to be encoded; a controller for controlling the encoder processor such that, depending on a specific signal characteristic of a frame, the number of audio data items of the audio data to be encoded by the encoder processor is reduced compared to a typical simple result obtained by state-of-the-art psychoacoustic considerations. In addition, this reduction of the number of audio data items is done in a signal-dependent manner such that for a frame having a specific first signal characteristic, the number is reduced more compared to another frame having another signal characteristic different from the signal characteristic of the first frame. Although this reduction of the number of audio data items can be seen as a reduction of the absolute number or a reduction of the relative number, this is not deterministic. However, the information units which are "saved" by the intended reduction of the number of audio data items are not simply lost but are used for a more accurate encoding of the remaining number of data items, i.e. the data items which are not eliminated by the intended reduction of the number of audio data items.

[0012] According to the invention, the controller for controlling the encoder processor operates in a manner such that, depending on a first signal characteristic of a first frame of the audio data to be encoded, the number of audio data items of the audio data to be encoded by the encoder processor for the first frame is reduced compared to a second signal characteristic of a second frame and, at the same time, the first number of information units for encoding the reduced number of audio data items for the first frame is more strongly enhanced compared to the second number of information units of the second frame.

[0013] In a preferred embodiment, the reduction is done in a manner such that for more tonal signal frames, a larger reduction is performed and, at the same time, the number of bits of the respective line is enhanced more for frames with lower tonality, i.e. more noisy frames. Here, the number is not reduced by this higher degree and, correspondingly, the number of information units for encoding the lower tonal audio data items is not increased this much.

[0014] The present invention provides a framework in which the psychoacoustic considerations which are typically provided are violated more or less in a signal-dependent manner. However, on the other hand, this violation is not seen as in an ordinary encoder in which a violation of psychoacoustics is done for example in emergency situations such as a situation in which higher frequency parts are set to zero in order to maintain a required bit rate. In fact, according to the present invention, this violation of the ordinary psychoacoustic considerations is done independent of any emergency situation and the "saved" information units are applied for a further optimization of the "remaining" audio data items.

[0015] In a preferred embodiment, a two-stage encoder processor is used, with an entropy encoder, such as an arithmetic encoder, or a variable length encoder, such as a Huffman encoder, as the initial encoding stage. The second encoding stage acts as an optimization stage, and this second encoder is typically implemented as a residual encoder or a bit encoder operating on bit granularity, which can be implemented, for example, by adding a certain defined offset in case of a first value of an information unit or subtracting the offset in case of the opposite value of the information unit. In an embodiment, this optimization encoder is preferably implemented as a residual encoder adding an offset in case of a first bit value and subtracting the offset in case of a second bit value. In a preferred embodiment, the reduction of the number of audio data items changes the situation in which the distribution of available bits in the typical fixed frame rate situation occurs in such a way that the initial encoding stage receives a lower bit budget than the optimization encoding stage. So far, the paradigm was that the initial encoding stage receives as high a bit budget as possible, independent of the signal characteristics, since it was assumed that the initial encoding stage, such as an arithmetic encoding stage, has the highest efficiency and, thus, encodes better from an entropy point of view than the residual encoding stage. However, according to the present invention, this paradigm is removed, since it has been found that for certain signals, such as signals with higher tones, the efficiency of an entropy encoder, such as an arithmetic encoder, is not as high as the efficiency obtained by a subsequently connected residual encoder, such as a bit encoder. However, while the entropy encoding stage is on average highly efficient for audio signals, the present invention now solves this problem by not observing the average, but reducing the bit budget of the initial encoding stage and preferably of the tonal signal portion in a signal dependent way.

[0016] In a preferred embodiment, the bit budget shift from the initial encoding stage to the optimization encoding stage is done in a way that at least two optimization information units are available for all audio data items that are left in the reduction of the number of data items of at least one and preferably 50% and even better. In addition, it has been found that a particularly efficient process for calculating these optimization information units on the encoder side and applying these optimization information units on the decoder side is an iterative process, in which the remaining bits from the bit budget for the optimization encoding stage are consumed sequentially in a certain order, such as from low frequencies to high frequencies. Depending on the number of left audio data items and depending on the number of information units of the optimization encoding stage, the number of iterations can significantly be larger than two, and it has been found that for strong tonal signal frames, the number of iterations can be four, five or even higher.

[0017] In a preferred embodiment, the determination of the control value by the controller is performed in an indirect manner, i.e. without explicit determination of the signal characteristics. For this purpose, the control value is calculated on the basis of manipulated input data, wherein this manipulated input data is for example the input data to be quantized or amplitude-related data derived from the data to be quantized. Although the control value of the encoder processor is determined on the basis of manipulated data, the actual quantization / encoding is performed without this manipulation. In this way, a signal-dependent process is obtained by determining the manipulation value for the manipulation in a signal-dependent manner, wherein this manipulation more or less influences the resulting reduction of the number of audio data items without explicit knowledge of specific signal characteristics.

[0018] In another implementation, a direct mode can be applied, wherein specific signal characteristics are estimated directly and depending on the results of this signal analysis, a specific reduction of the number of data items is performed in order to obtain a higher precision of the remaining data items.

[0019] In yet another implementation, a separate process can be applied for the purpose of reducing the audio data items. In the separate process, a specific number of data items is obtained by means of a quantization controlled by a generally psychoacoustic driven quantizer control and based on the input audio signal, the quantized audio data items are reduced with respect to their number and preferably this reduction is done by eliminating the smallest audio data items with respect to their amplitude, their energy or their power. Again, the control of the reduction can be obtained by direct / explicit signal characteristic determination or by indirect or non-explicit signal control.

[0020] In another preferred embodiment, an integrated process is applied, wherein the variable quantizer is controlled to perform a single quantization but based on manipulated data, while the data which is not manipulated is quantized. The quantizer control value such as a global gain is calculated using the signal-dependent manipulated data, while the data without this manipulation is quantized and the quantization result is encoded using all available units of information, so that in the case of two-stage encoding, the generally large units of information of the optimized encoding stage are preserved.

[0021] Embodiments provide a solution to the problem of quality loss of high pitch content, which is based on a modification of the power spectrum used to estimate the bit consumption of an entropy encoder. While this modification increases the bit budget estimation for high pitch content, it remains common to the estimated signal adaptive noise floor adder of the audio content with an actually unchanged flat residual spectrum. The impact of this modification is twofold. First, it quantizes the uncorrelated side lobes of the filter bank noise and harmonic components to zero, which are covered by the noise floor. Second, it shifts bits from the first encoding stage to the residual encoding stage. While this shift is undesirable for most signals, it is fully effective for high pitch signals, because the bits are used to increase the quantization accuracy of the harmonic components. This means that the shift is used to encode bits with low significance, which usually follow a uniform distribution and are therefore fully efficiently encoded with a binary representation. In addition, the process is computationally cheap, so that it is an extremely effective tool for solving the aforementioned problem. BRIEF DESCRIPTION OF DRAWINGS

[0022] A preferred embodiment of the application is disclosed hereinafter with respect to the accompanying drawings, in which:

[0023] Figure 1 is an embodiment of an audio encoder;

[0024] Figure 2 Explanation Figure 1 is a preferred implementation of an encoder processor of the

[0025] Figure 3 Explanation of a preferred implementation of an optimization encoding stage;

[0026] Figure 4a Explanation of an exemplary frame syntax with a first frame or a second frame of iterative optimized bits;

[0027] Figure 4b Explanation of a preferred implementation of an audio data item reducer like a variable quantizer;

[0028] Figure 5 Explanation of a preferred implementation of an audio encoder with a spectral pre-processor;

[0029] Figure 6 Explanation of a preferred embodiment of an audio decoder with a temporal post-processor;

[0030] Figure 7 Explanation Figure 6 is an implementation of an encoder processor of the audio decoder of

[0031] Figure 8 Explanation Figure 7 is a preferred implementation of an optimization decoding stage;

[0032] Figure 9An implementation of an indirect mode for controlling value computation is illustrated;

[0033] Figure 10 An implementation of a manipulation value calculator is illustrated Figure 9

[0034] Figure 11 An implementation of a direct mode control value computation is illustrated

[0035] Fig. 12 illustrates an implementation of a separate audio data item reduction; and

[0036] Fig. 13 illustrates an implementation of an integrated audio data item reduction. DETAILED DESCRIPTION

[0037] Figure 1 An audio encoder for encoding audio input data 11 is illustrated. The audio encoder comprises a pre-processor 10, an encoder processor 15 and a controller 20. The pre-processor 10 pre-processes the audio input data 11 such that per frame audio data or audio data to be encoded as illustrated at item 12 is obtained. The audio data to be encoded is input into the encoder processor 15 for encoding the audio data to be encoded and the encoder processor outputs encoded audio data. The controller 20 is connected with respect to its input to the per frame audio data of the pre-processor, but alternatively, the controller can also be connected to receive the audio input data without any pre-processing. The controller is configured to reduce the number of audio data items per frame depending on the signal in the frame and at the same time, the controller increases the number of information units, or preferably, bits, for the reduced number of audio data items depending on the signal in the frame. The controller is configured for controlling the encoder processor 15 such that depending on a first signal characteristic of a first frame of audio data to be encoded, the number of audio data items of the audio data to be encoded by the encoder processor for the first frame is reduced compared to a second signal characteristic of a second frame and for encoding the reduced number of audio data items for the first frame with a number of information units is enhanced more compared to the second number of information units of the second frame.

[0038] Figure 2 A preferred implementation of an encoder processor is illustrated. The encoder processor comprises an initial encoding stage 151 and an optimized encoding stage 152. In one implementation, the initial encoding stage comprises an entropy encoder, such as an arithmetic or Huffman encoder. In another embodiment, the optimized encoding stage 152 comprises a bit encoder or residual encoder operating on a bit or information unit granularity. Additionally, the functionality with respect to the reduction of the number of audio data items is in the initial encoding stage 151. Figure 2 ​The audio data item reducer 150 is embodied by means of an audio data item reducer 150 which can be implemented as a variable quantizer in the integrated reduction mode as illustrated in Fig. 13, for example, or alternatively as a separate element operating on the quantized audio data items as illustrated in the separate reduction mode 902, and in another not illustrated embodiment the audio data item reducer can also operate on unquantized elements by setting such unquantized elements to zero or by weighting the data items to be eliminated with a certain weighting number such that such audio data items are quantized to zero and thus eliminated in the subsequently connected quantizer. Figure 2 The audio data item reducer 150 can operate on unquantized or quantized data elements in the separate reduction procedure, or can be implemented by a variable quantizer controlled by a certain signal dependent control value as illustrated in the integrated reduction mode. Figure 13 integrated reduction

[0039] Figure 1 The controller 20 is configured to reduce the number of audio data items encoded by the initial encoding stage 151 for the first frame, and the initial encoding stage 151 is configured to encode the reduced number of audio data items of the first frame using the first frame initial number of information units, and the calculated bits / units of the initial number of information units are output by the block 151 as illustrated in the item 151. Figure 2

[0040] In addition, the optimized encoding stage 152 is configured to use the first frame remaining number of information units for the optimized encoding of the reduced number of audio data items of the first frame, and the addition of the first frame initial number of information units to the first frame remaining number of information units results in a predetermined number of information units of the first frame. In particular, the optimized encoding stage 152 outputs the first frame remaining number of bits and the second frame remaining number of bits, and for at least one or preferably at least 50% or even better all non-zero audio data items, i.e. the audio data items remaining from the reduction of audio data items and initially encoded by the initial encoding stage 151, there are indeed at least two optimized bits.

[0041] Preferably, the predetermined number of information units of the first frame is equal to or rather close to the predetermined number of information units of the second frame, such that a constant or substantially constant bit rate operation of the audio encoder is obtained.

[0042] As illustrated in Fig. 1, the audio encoder 1 comprises an input 10 for receiving an audio signal, a controller 20 for controlling the operation of the audio encoder 1, an initial encoding stage 151 for encoding the audio signal, an optimized encoding stage 152 for encoding the audio signal, a quantizer 153 for quantizing the audio signal, a bit stream output 154 for outputting the quantized audio signal, and a bit stream output 155 for outputting the quantized audio signal. Figure 2 ​​As explained in the foregoing, the audio data item reducer 150 reduces the audio data items below the psychoacoustic driven number in a signal dependent manner. Thus, for a first signal characteristic, the number is only slightly reduced compared to the psychoacoustic driven number, and, for example, in a frame having a second signal characteristic, the number is significantly reduced below the psychoacoustic driven number. Also, preferably, the audio data item reducer eliminates the data items with a minimum amplitude / power / energy, and this operation is preferably performed via an indirect selection obtained in the integration mode, wherein the reduction of the audio data items is performed by quantizing certain audio data items to zero. In an embodiment, the initial encoding stage only encodes the audio data items that have not been quantized to zero, and the optimization encoding stage 152 only optimizes the audio data items that have been processed by the initial encoding stage, i.e. that have not been quantized to zero by the audio data item reducer 150. Figure 2

[0043] ​In a preferred embodiment, the optimization encoding stage is configured to iteratively allocate the remaining number of information units of the first frame to the reduced number of audio data items of the first frame in at least two sequentially performed iterations. In particular, values of allocated information units for the at least two sequentially performed iterations are calculated and the calculated values of information units for the at least two sequentially performed iterations are introduced into the encoded output frame in a predetermined order. In particular, the optimization encoding stage is configured to sequentially allocate information units of each audio data item of the reduced number of audio data items of the first frame in an order from low frequency information of the audio data item to high frequency information of the audio data item in a first iteration. In particular, the audio data items can be respective spectral values obtained by a time / spectral conversion. Alternatively, the audio data items can be tuples of two or more spectral lines which are typically adjacent to each other in the spectrum. Then, a calculation of bit values is performed from a certain start value with low frequency information to a certain end value with highest frequency information and in a further iteration the same procedure is performed, i.e. again a processing from low spectral information values / tuples to high spectral information values / tuples is performed. In particular, the optimization encoding stage 152 is configured to check whether the number of allocated information units is below a predetermined number of information units of the first frame which is less than the initial number of information units of the first frame and the optimization encoding stage is also configured to stop the second iteration in case of a negative check result or to perform a number of further iterations in case of a positive check result until a negative check result is obtained, wherein the number of further iterations is 1, 2,.... Preferably, the maximum number of iterations is limited by a two digit number, such as a value between 10 and 30 and preferably 20 iterations. In an alternative embodiment, the check on the maximum number of iterations can be omitted if the non-zero spectral lines are counted first and the number of residual bits is adjusted accordingly for each iteration or for the whole procedure. Thus, when there are for example 20 remaining spectral tuples and 50 residual bits, the number of iterations can be determined to be three without any check during the procedure in the encoder or decoder and in the third iteration the optimization bits will be calculated or available in the bit stream for the first ten spectral lines / tuples. Thus, this alternative does not require a check during the iteration processing since the information on the number of non-zero or remaining audio items is known after the processing in the initial stage in the encoder or decoder.

[0044] Figure 3 The preferred implementation of the iteration process performed by the optimization encoding stage 152 of Figure 2 The preferred implementation of the iteration process performed by the optimization encoding stage 152 of

[0045] In step 300, the remaining audio data items are determined. This determination can be performed by having alreadyFigure 2 The automatic execution is performed on the audio data items handled by the initial encoding stage 151. In step 302, the start of the procedure takes place at a predefined audio data item, such as the audio data item with the lowest spectral information. In step 304, the bit value of each audio data item in a predefined sequence, e.g. a sequence from the low spectral value / tuple to the high spectral value / tuple, is calculated. The calculation in step 304 is performed using the start offset 305 and the optimization bits still available in control 314. At item 316, a first iteration optimization information unit is output, i.e. a bit pattern indicating one bit for each retained audio data item, wherein the bit indicates whether the offset, i.e. the start offset 305, is to be added or subtracted, or alternatively, whether the start offset is to be added or not.

[0046] In step 306, the offset is reduced with a predetermined rule. This predetermined rule can e.g. be halving the offset, i.e. the new offset is half the original offset. However, other offset reduction rules than the one weighted with 0.5 can also be applied.

[0047] In step 308, the bit value of each item in the predefined sequence is again calculated, but now in the second iteration. As input to the second iteration, the optimized items after the first iteration as explained at 307 are input. Thus, for the calculation in step 314, the optimization represented by the first iteration optimization information unit has been applied, and in the prerequisite that the optimization bits as indicated in step 314 are still available, a second iteration optimization information unit is calculated and output at 318.

[0048] In step 310, the offset is again reduced by preparing a predetermined rule for a third iteration, and the third iteration again relies on the optimized items after the second iteration as explained at 309 and again in the prerequisite that the optimization bits as indicated at 314 are still available, a third iteration optimization information unit is calculated and output at 320.

[0049] Figure 4aAn exemplary frame syntax is illustrated with information units or bits for the first frame or the second frame. A part of the bit data of the frame is constituted by the initial number of bits, i.e. item 400. In addition, the first, second and third iteration refinement bits 316, 318 and 320 are also included in the frame. In particular, according to the frame syntax, the decoder is in a position to identify which bits of the frame are the initial number of bits, which bits are the first, second or third iteration refinement bits 316, 318, 320 and which bits in the frame are any other bits 402, e.g. which can also comprise an encoded representation of a global gain (gg), for example. This any side information, e.g. which can be directly calculated by the controller 200 or which can be influenced by the controller, e.g. by means of the controller output information 21, can be controlled. Within the part 316, 318, 320, a specific sequence of the respective information units is given. This sequence is preferably such that the bits in the sequence of bits are applied to the initially decoded audio data item to be decoded. Since this sequence is not useful for explicitly signaling anything about the first, second and third iteration refinement bits with respect to the bit rate requirements, the order of the respective bits in the blocks 316, 318, 320 should be the same as the corresponding order of the remaining audio data item. In view of the situation, it is preferred to use the same iteration procedure on the encoder side as illustrated in Figure 3 Figure 8 and on the decoder side as illustrated in Figure 8 It is not necessary to signal any specific bit allocation or bit association at least in the blocks 316 to 320.

[0050] In addition, the number of the initial number of bits on the one hand and the remaining number of bits on the other hand is only exemplary. Typically, the initial number of bits, which typically encodes the most significant bit portion of the audio data item, such as a spectral value or a tuple of spectral values, is greater than the iteration refinement bits, which represent the least significant portion of the "remaining" audio data item. In addition, the initial number of bits 400 is typically determined by means of an entropy encoder or an arithmetic encoder, but the iteration refinement bits are determined using a residual or bit encoder, which operates on the information unit granularity. Although the optimization encoding stage does not perform any entropy encoding in all likelihood, nevertheless, the encoding of the least significant bit portion of the audio data item is more efficiently performed by the optimization encoding stage, since the least significant bit portion of the audio data item, such as a spectral value, can be assumed to be evenly distributed and, therefore, any entropy encoding with variable length codes or arithmetic encoding and specific contexts does not introduce any additional advantage, but rather even an additional burden.

[0051] In other words, for the least significant bit part of the audio data items, the use of an arithmetic encoder should be less efficient than the use of a bit encoder, because the bit encoder does not require any bit rate for a specific context. The intended reduction of the audio data items as caused by the controller not only improves the accuracy of the primary spectral lines or line elements, but additionally provides an efficient encoding operation for the purpose of optimizing the MSB part of these audio data items represented by the arithmetic or variable length code.

[0052] In view of this, several and for example the following advantages are obtained by the implementation of the encoder processor 15 by means of an encoding of Figure 2 as explained in Figure 1 the first encoding stage 151 and the second optimization encoding stage 152.

[0053] A highly efficient two-stage encoding scheme is proposed, comprising a first entropy encoding stage and a second residual encoding stage based on single bit (non-entropy) encoding.

[0054] The scheme employs a low complexity global gain estimator incorporating an energy based bit consumption estimator featuring a signal adaptive floor noise adder for the first encoding stage.

[0055] The floor noise adder effectively transfers bits from the first encoding stage to the second encoding stage for high pitch signals, while leaving the estimation for other signal types unchanged. This bit shifting from the entropy encoding stage to the non-entropy encoding stage is sufficiently efficient for high pitch signals.

[0056] Figure 4b A preferred implementation of a variable quantizer is explained, which can for example be implemented to perform the audio data item reduction in the integrated reduction mode as explained with respect to Fig. 13. For this purpose, the variable quantizer comprises a weighter 155 receiving the audio data to be encoded (not manipulated) as explained at line 12. This data is also input into the controller 20, and the controller is configured to compute a global gain 21, but based on the unmanipulated data as input into the weighter 155, and using a manipulation that depends on the signal. The global gain 21 is applied in the weighter 155, and the output of the weighter is input into a quantizer core 157 that depends on a fixed quantization step size. The variable quantizer 150 is implemented as a controlled weighter, with control using the global gain (gg) 21 and the subsequently connected fixed quantization step size quantizer core 157. However, other implementations can also be performed, such as a quantizer core with a variable quantization step size controlled by the output value of the controller 20.

[0057] Figure 5 A preferred implementation of an audio encoder is explained, and in particular, an implementation of Figure 1of the pre-processor 10. Preferably, the pre-processor comprises a windower 13 which generates frames of time-domain audio data windowed using a specific analysis window, which can for example be a cosine window, from the audio input data 11. The frames of time-domain audio data are input into a spectrum converter 14 which can be implemented to perform a modified discrete cosine transform (MDCT) or any other transform such as an FFT or an MDST or any other time-spectral conversion. Preferably, the windower operates with a specific look-ahead control such that overlapping frame generation is performed. In case of 50% overlap, the look-ahead value of the windower is half the size of the analysis window applied by the windower 13. The (unquantized) frames of spectral values output by the spectrum converter are input into a spectral processor 15 which is implemented to perform several spectral processes such as a run-time noise shaping operation, a spectral noise shaping operation or any other operation such as a spectral whitening operation by which the modified spectral values generated by the spectral processor have a flatter spectral envelope than the spectral envelope of the spectral values before processing by the spectral processor 15. The audio data to be encoded (per frame) is forwarded via line 12 into the encoder processor 15 and into the controller 20, wherein the controller 20 provides control information to the encoder processor 15 via line 21. The encoder processor outputs its data to a bitstream writer 30 which is implemented for example as a bitstream multiplexer and outputs the encoded frames on line 35.

[0058] With respect to the decoder side processing, reference is made to Figure 6 The bitstream output by the block 30 can for example be input directly into the bitstream reader 40 after some storage or transmission. Of course, any other processing such as a transmission processing can be performed between the encoder and the decoder according to a wireless transmission protocol such as a DECT protocol or a Bluetooth protocol or any other wireless transmission protocol. The data input into the audio decoder shown in Figure 6 The data input into the audio decoder shown in Figure 7 The data input into the audio decoder shown in Figure 7When the initial decoding stage 51 outputs the initially decoded data item, at least two information units from the remaining number of information units are used to optimize the same initially decoded data item. Additionally, the controller 60 is configured to control the encoder processor such that the initial decoding stage uses the initial number of information units of the frame to optimize the data item. Figure 7 The initially decoded data items are obtained at line connection blocks 51 and 52, wherein preferably, the controller 60 is as follows: Figure 6 or Figure 7 The input lines in block 60 indicate the information units received from bitstream reader 40 regarding the initial number of frames and the initial remaining number of frames. Postprocessor 70 processes the optimized audio data items to obtain decoded audio data 80 at the output of postprocessor 70.

[0059] In corresponding Figure 5 In a preferred embodiment of the audio decoder of the audio encoder, the post-processor 70 includes a spectrum processor 71 as an input stage, which performs inverse time noise shaping, or inverse spectral noise shaping, or inverse spectral whitening, or reduces noise caused by... Figure 5 The spectrum processor 15 is used for any other operation of a certain processing. The output of the spectrum processor is input to a time converter 72, which performs a conversion from the spectral domain to the time domain, and preferably, the time converter 72 is connected to... Figure 5 The spectrum converter 14 is matched. The output of the time converter 72 is input to the overlap-add stage 73, which performs overlap / add operations on multiple overlap frames, such as at least two overlap frames, to obtain decoded audio data 80. Preferably, the overlap-add stage 73 applies a synthesis window to the output of the time converter 72, wherein this synthesis window matches the analysis window applied by the analysis windower 13. In addition, the overlap operation performed by block 73 is matched with the output of the time converter 72. Figure 5 The window adder 13 performs block advance operations to match.

[0060] like Figure 4a As explained, the information unit for the remaining number of frames includes calculated values ​​of information units 316, 318, and 320 for at least two sequential iterations in a predetermined order, wherein... Figure 4a In the embodiment, even three iterations are described. Additionally, controller 60 is configured to control the optimized decoding level 52 to use computed values ​​such as block 316 for the first iteration in a predetermined order, and to use computed values ​​from block 318 for the second iteration in a predetermined order.

[0061] Subsequently, regarding Figure 8 This describes a preferred implementation of the optimized decoding level under the control of controller 60. In step 800, the controller or... Figure 7The optimization decoding level 52 identifies the audio data items to be optimized. These audio data items are typically composed of... Figure 7 All audio data items output by block 51. As indicated in step 802, a start is performed at a predefined audio data item, such as minimum spectral information. Using a start offset 805, a first iterative optimization information unit received from the bitstream or from controller 16 is applied for each item in the predefined sequence, for example, 804. Figure 4a The data in block 316, wherein the predefined sequence extends from low spectral values / spectral tuples / spectral information to high spectral values / spectral tuples / spectral information. The result is an optimized audio data item after the first iteration as illustrated in line 807. In step 808, the bit value of each item in the predefined sequence is applied, wherein the bit value comes from the second iteration optimization information unit as illustrated in 818, and these bits are received from the bit stream reader or controller 60 depending on the specific implementation. The result of step 808 is the optimized item after the second iteration. Similarly, in step 810, the offset is reduced according to the predetermined offset reduction rule applied in block 806. Using the reduced offset, the bit value of each item in the predefined sequence is applied as illustrated in 812, using, for example, a third iteration optimization information unit received from the bit stream or from controller 60. Figure 4a At item 320, the third iteration optimization information unit is written into the bit stream. The result of the process in block 812 is the optimized term after the third iteration, as indicated at item 821.

[0062] This process continues until all iteratively optimized bits included in the bitstream of the frame have been processed. This is checked by controller 60 via control line 814, which preferably controls the remaining availability of optimized bits for each iteration, but at least for the second and third iterations processed in blocks 808, 812. In each iteration, controller 60 controls the optimized decoding stage to check whether the number of read information units is less than the number of information units in the remaining information units of the frame, thereby stopping the second iteration if the check result is negative, or performing multiple further iterations until a negative check result is obtained if the check result is positive. The number of further iterations is at least one. Since similar processes exist in... Figure 3 The encoder side and such discussed in the context Figure 8 The decoder-side applications outlined herein do not require any specific signal notification. In fact, the multi-iterative optimization process is performed efficiently without any specific overhead. In an alternative embodiment, the check for the maximum number of iterations can be omitted if the non-zero spectral lines are counted first, and the number of residual bits is adjusted accordingly for each iteration.

[0063] In a preferred implementation, the optimization decoding stage 52 is configured to add an offset to the initially decoded data item when the read information data unit in the frame remaining number of information units has a first value and to subtract the offset from the initially decoded item when the read information data unit in the frame remaining number of information units has a second value. For the first iteration, this offset is Figure 8 the start offset 805. In a second iteration as explained at 808 in Figure 8 , the reduced offset as generated by block 806 is used to add a reduced or second offset to the result of the first iteration when the read information data unit in the frame remaining number of information units has a first value and to subtract the second offset from the result of the first iteration when the read information data unit in the frame remaining number of information units has a second value. In general, the second offset is lower than the first offset and preferably, the second offset is between 0.4 and 0.6 times the first offset and optimally 0.5 times the first offset.

[0064] In a preferred implementation of the application using the indirect mode as explained in Figure 9 , any explicit signal characteristic determination is not necessary. In fact, the embodiment as explained in Figure 9 is preferably used to calculate the steering value. For the indirect mode, the controller 20 is implemented as indicated in Figure 9 . In particular, the controller comprises a control pre-processor 22, a steering value calculator 23, a combiner 24 and a global gain calculator 25 which in the last calculation is implemented as Figure 4b the variable quantizer as explained in Figure 2 . In particular, the controller 20 is configured to analyze the audio data of a first frame to determine a first control value for the variable quantizer for the first frame and to analyze the audio data of a second frame to determine a second control value for the variable quantizer for the second frame, the second control value being different from the first control value. The analysis of the audio data of the frames is performed by the steering value calculator 23. The controller 20 is configured to perform a steering of the audio data of the first frame. In this operation, there is no control pre-processor 20 as explained in Figure 9 , so the bypass pipeline of block 22 is active.

[0065] However, when manipulation is not performed on the audio data of the first or second frame, but applied to amplitude-related values ​​derived from the audio data of the first or second frame, a control preprocessor 22 exists and no bypass pipeline exists. The actual manipulation is performed by combiner 24, which combines the manipulated value output from block 23 with the amplitude-related value derived from the audio data of a specific frame. At the output of combiner 24, manipulated (preferably energy) data is indeed present, and based on this manipulated data, global gain calculator 25 calculates the global gain indicated at 404, or at least the control value of the global gain. Global gain calculator 25 must impose a limit on the allowed bit budget of the spectrum to obtain a specific data rate or a specific number of information units allowed for the frame.

[0066] exist Figure 11 In the direct mode described herein, controller 20 includes an analyzer 201 for determining the signal characteristics of each frame, and analyzer 208 outputs quantitative signal characteristic information, such as pitch information, and uses this preferred quantitative data to control control value calculator 202. A process for calculating the pitch of a frame is used to calculate the spectral flatness measure (SFM) of the frame. Any other pitch determination process or any other signal characteristic determination process can be executed by block 201, and a conversion from a specific signal characteristic value to a specific control value will be performed to reduce the expected number of audio data items obtained for the frame. Figure 11 The output of the direct-mode control value calculator 202 can be sent to the encoder processor, such as to a variable quantizer, or alternatively to the control value of the initial encoder level. When the control value is given to the variable quantizer, an integrated reduction mode is performed, while when the control value is given to the initial encoder level, a separate reduction is performed. Another implementation of the separate reduction should remove or specifically affect selected unquantized audio data items that exist before actual quantization, such that, by means of a specific quantizer, this affected audio data item is quantized to zero, and thus eliminated for entropy coding and subsequent optimized coding purposes.

[0067] although Figure 9 The indirect mode has been shown along with the integrated reduction, i.e., the global gain calculator 25 is configured to calculate the variable global gain, but the manipulated data output by the combiner 24 can also be used to directly control the initial coding level to remove any particular quantized audio data item, such as the minimum quantized data item, or alternatively, the control value can also be sent to an unspecified audio data influence level that influences the audio data before the actual quantization using the variable quantization control value determined without any data manipulation, and thus generally follows psychoacoustic rules, however, the process of the present invention intentionally violates the psychoacoustic rules.

[0068] As Figure 11 As explained in the direct mode in the

[0069] The present invention does not generate the coarser quantization that would normally be obtained by applying a global gain. Indeed, this computation based on a global gain depending on the manipulated data of the signal only generates a shift of the bit budget from the initial encoding stage receiving a smaller bit budget to the optimized decoding stage receiving a higher bit budget, but this shift of bit budget is done in a signal dependent manner and is greater for higher tonal signal parts.

[0070] Preferably, Figure 9 The control pre-processor 22 computes the amplitude related values as a plurality of power values derived from the one or more audio values of the audio data. In particular, it is these power values manipulated by means of an addition of the same manipulation values using the combiner 24, and the same manipulation values determined by the manipulation value calculator 23 are combined with all of the plurality of power values of the frame.

[0071] Alternatively, as indicated by the bypass pipeline, values of the same magnitude of the manipulation values computed by the block 23, but preferably having a random sign, and / or values obtained by a subtraction of the same magnitude, but preferably having a random sign, from the same magnitude or complex manipulation values, or more generally, values obtained as samples from a certain normalized probability distribution scaled using the computed complex or real magnitude values of the manipulation values, are added to all of the plurality of audio values comprised in the frame. The processes performed by the control pre-processor 22, such as the computation of the power spectrum and the down-sampling, can be comprised within the global gain calculator 25. Thus, preferably, the noise floor is added directly to the spectral audio values or alternatively to the amplitude related values derived from the audio data of each frame, i.e. the output of the control pre-processor 22. Preferably, the controller pre-processor computes a down-sampled power spectrum corresponding to taking the power with an exponent value equal to 2. However, alternatively, a different exponent value higher than 1 can be used. Exemplarily, an exponent value equal to 3 should represent the loudness rather than the power. However, other exponent values such as a smaller or larger exponent value can be used as well.

[0072] In the preferred implementation explained in the Figure 10 In the preferred implementation explained in the Figure 10 Figure 10 ​The searchers 26 are configured to search for a maximum value of the plurality of audio data items or amplitude related values or to search for a maximum value of the plurality of down-sampled audio data or the plurality of down-sampled amplitude related values of the corresponding frame. Using the outputs of the blocks 26, 27 and 28, the actual calculation is performed by the block 29, wherein the blocks 26, 28 actually represent the signal analysis.

[0073] Preferably, the signal independent contribution is determined by means of the bit rate of the actual encoder session, the frame duration or the sampling frequency of the actual encoder session. In addition, the calculator 28 for calculating one or more moments per frame is configured to calculate a signal dependent weighting value derived from a first sum of the magnitudes of the audio data or the down-sampled audio data within the frame, a second sum of the magnitudes of the audio data or the down-sampled audio data within the frame multiplied by the index associated with each magnitude and the quotient of the second sum and the first sum.

[0074] In the preferred implementation performed by the global gain calculator 25 of Figure 9 The required bit estimate per energy value is calculated depending on the energy value and a candidate value of the actual control value. The required bit estimates of the energy values and the candidate values of the control values are accumulated and it is checked whether the accumulated bit estimates of the candidate values of the control values fulfill the allowed bit consumption criterion, like the bit budget introduced to the spectrum into the global gain calculator 25, as explained for example in Figure 9 If the allowed bit consumption criterion is not fulfilled, the candidate value of the control value is modified and the calculation of the required bit estimates, the accumulation of the required bit rates and the check of the allowed bit consumption criterion for the modified candidate value of the control value are repeated. Once this optimal control value is found, this value is output at line 404 of Figure 9

[0075] Subsequently, a preferred embodiment is explained.

[0076] Detailed description of the encoder (e.g. pre-processor 10) Figure 5

[0077] Notations

[0078] By f s represents the potential sampling frequency in Hertz (Hz), by N ms represents the potential frame duration in milliseconds and by br the potential bit rate in bits per second.

[0079] Derivation of the residual spectrum (e.g. pre-processor 10)

[0080] The embodiment depends on the real residual spectrum X f ​​(k), k = 0..N-1, which is usually derived by a time-to-frequency transform like MDCT, followed by psychoacoustic motivated modifications like time noise shaping (TNS) to remove time structure and spectral noise shaping (SNS) to remove spectral structure. Thus, for audio content with a slowly changing spectral envelope, the residual spectrum X f (k) has a flat envelope.

[0081] ■global gain estimate (e.g. Figure 9 )

[0082] The quantization of the spectrum is controlled via a global gain g glob (k) = X

[0083]

[0084] (k) = X 2 (k) is derived from the power spectrum X(k) Figure 9 (k) = X

[0085] PX lp (k) = X f (4k) 2 + X f (4k+1) 2 + X f (4k+2) 2 + X f (4k+3) 2

[0086] and the signal adaptive floor noise N(X f ) is given by

[0087] (e.g. Figure 9 ) is given by

[0088] The parameter regBits depends on the bit rate, the frame duration and the sampling frequency and is calculated as

[0089] (e.g. Figure 10 ) is given by

[0090] where C(N ms , f s ) is specified in the following table.

[0091] N ms \f s ]]> 48000 96000 2.5 -6 -6 5 0 0 10 2 5

[0092] The parameter lowBits depends on the centroid of the absolute values of the residual spectrum and is calculated as

[0093] (e.g. Figure 10 Item 28 of

[0094] where

[0095]

[0096] and

[0097]

[0098] is the matrix of absolute spectra.

[0099] From the values

[0100] E(k) = 10 log 10 (PX lp (k) + N(X f ) + 2 -31 ), (e.g. Figure 9 Output of combiner 24)

[0101] In the form

[0102]

[0103] The global gain is estimated.

[0104] where gg off is a bit rate and sampling frequency dependent offset.

[0105] It should be noted that adding the noise floor term N(X f ) to PX lp (k) before computing the power spectrum provides the expected result of adding the noise floor to the residual spectrum X f (k), e.g., adding the term randomly to each spectral line or subtracting the term.

[0106] A pure power spectrum based estimate can have been found e.g. in the 3GPP EVS codec (3GPP TS 26.445, section 5.3.3.2.8.1). In an embodiment, the addition of the noise floor N(Xx) is done. The noise floor is signal adaptive in two ways.

[0107] First, it is scaled with the maximum amplitude X f . Thus, the impact on the energy of a flat spectrum, where all amplitudes are close to the maximum amplitude, is minimal. But for a high pitch signal, where the residual spectrum is also characterized by an extension in spectrum and multiple strong peaks, the total energy is significantly increased, which increases the bit estimate of the global gain calculation as outlined below.

[0108] Second, if the spectrum exhibits a low centroid, the noise floor is reduced by a parameter lowBits. In this case, the content is mainly low frequency, whereby the loss of high frequency components is likely to be less critical than with high pitch content.

[0109] The actual estimation of the global gain (block 25) is performed by a low complexity binary search as outlined in the C program code below, where nbits' represents the bit budget used for encoding the spectrum. The bit consumption estimate (accumulated in variable tmp) is based on the energy values E(k) taking into account the context dependencies in the arithmetic coder used for stage 1 encoding. Figure 9 spec (k) of the spectrum. The bit consumption estimate (accumulated in variable tmp) is based on the energy values E(k) taking into account the context dependencies in the arithmetic coder used for stage 1 encoding.

[0110]

[0111] ■Residual encoding (e.g. Figure 3 )

[0112] The residual encoding uses the excess bits available after the arithmetic encoding of the spectrum x q (k). Let B represent the number of excess bits, and let K represent the number of encoded non-zero coefficients X q (k). In addition, let k i , i = 1..K represent the row order of these non-zero coefficients from lowest to highest frequency. The residual bits b i (j) (taking values 0 and 1) are computed such that the error i (j) (taking values 0 and 1) are computed such that the error

[0113]

[0114] This can be done in an iterative fashion testing if the following is true

[0115]

[0116] If (1) is true, the n-th residual bit b i (n) for coefficient k i is set to 0, otherwise it is set to 1. The computation of the residual bits proceeds by computing the first residual bit for each k i , and then the second bit, and so on, until all residual bits are exhausted, or a maximum number n max of iterations is performed. This leaves q (n) residual bits for coefficient X i (k

[0117]

[0118] This residual encoding scheme improves the residual encoding scheme applied in the 3GPP EVS codec, which spends at most one bit per non-zero coefficient. ​

[0119] The calculation of the residual bits with n max = 20 is illustrated by the following pseudo code, where gg denotes the global gain:

[0120]

[0121]

[0122] The description of the decoder (e.g. Figure 6 )

[0123] At the decoder, the entropy coded spectrum The residual bits are used to optimize this spectrum as illustrated by the following pseudo code (see also e.g. Figure 8 ).

[0124]

[0125] The decoded residual spectrum is given by the following

[0126]

[0127] • Conclusion:

[0128] • A highly efficient two-stage encoding scheme is proposed, comprising a first entropy coding stage and a second residual coding stage based on single-bit (non-entropy) coding.

[0129] • The scheme employs a low-complexity global gain estimator incorporating an energy-based bit consumption estimator featuring a signal-adaptive noise floor adder for the first coding stage.

[0130] • The noise floor adder effectively transfers bits from the first to the second coding stage for high-pitched signals, while leaving the estimate for other signal types unchanged. This bit shifting from the entropy to the non-entropy coding stage is considered sufficiently efficient for high-pitched signals.

[0131] Figure 12 illustrates a procedure for reducing the number of audio data items in a signal dependent manner using separate reduction. In step 901, quantization is performed without any manipulation using unmanipulated information such as the global gain as computed from the signal data. For this purpose, the (total) bit budget of the audio data items is needed, and at the output of block 901, the quantized data items are obtained. In block 902, the number of audio data items is reduced by eliminating (controlled) amounts of preferably the smallest audio data items based on a signal dependent control value. At the output of block 902, the reduced number of data items is obtained, and in block 903, the initial coding stage is applied, and the optimization coding stage is applied as illustrated in 904 in case of a bit budget of residual bits that are preserved due to the controlled reduction.

[0132] In addition to the procedure in Fig. 12, the reduction block 902 can also be performed using the global gain value or a certain quantizer step size that is usually determined with unmanipulated audio data prior to the actual quantization. Thus, this reduction of the audio data items can also be performed in the domain of unquantized values by setting certain preferably small values to zero or by weighting certain values with a weighting factor, finally resulting in values quantized to zero. In a separate reduction implementation, the explicit quantization step on the one hand and the explicit reduction step on the other hand are performed without any data manipulation in case of performing a control of the certain quantization.

[0133] In contrast thereto, Fig. 13 illustrates an integrated reduction mode according to an embodiment of the present application. In block 911, a manipulated information, such as a global gain at the output of block 25 of Fig. 12, is determined by the controller 20. In block 912, the quantization of the unmanipulated audio data is performed using the manipulated global gain or manipulated information that is usually calculated in block 911. At the output of the quantization procedure of block 912, the reduced number of audio data items is obtained that is initially encoded in block 903 and optimized encoded in block 904. Due to the signal dependent reduction of the audio data items, residual bits are reserved for at least a single full iteration and for at least a part of a second iteration and preferably for even more than two iterations. The shifting of the bit budget from the initial encoding stage to the optimized encoding stage is performed according to the present application and in a signal dependent manner. Figure 9

[0134] The present application can be implemented in at least four different modes. As an example of manipulation, the determination of the control value can be performed in a direct mode with explicit signal characteristic determination or in an indirect mode without explicit signal characteristic determination but with addition of a signal dependent noise floor to the audio data or to the derived audio data. At the same time, the reduction of the audio data items is performed in an integrated manner or in a separate manner. Also, an indirect determination and integrated reduction or an indirect generation of the control value and a separate reduction can be performed. In addition, a direct determination as well as an integrated reduction and a direct determination of the control value and a separate reduction can also be performed. For purposes of efficiency, the indirect determination of the control value and the integrated reduction of the audio data items is preferred.

[0135] It is to be mentioned here that all alternatives or aspects as discussed before or as defined in the independent claims in the following claims can be used accordingly, i.e. without any other alternative or object than the intended alternative, object or independent claim. However, in other embodiments, two or more of the alternatives or the aspects or the independent claims can be combined with each other and in other embodiments, all aspects or alternatives and all independent claims can be combined with each other.

[0136] ​The inventive encoded audio signal can be stored on a digital storage medium or non-transitory storage medium, or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0137] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or apparatus corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.

[0138] Depending on certain implementation requirements, embodiments of the application can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed.

[0139] Some embodiments according to the application comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0140] Generally, embodiments of the present application can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code can for example be stored on a machine readable carrier.

[0141] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0142] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0143] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein.

[0144] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be configured to be transferred via a data communication connection, for example via the Internet.

[0145] Another embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0146] Another embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0147] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0148] The above-described embodiments are merely meant to illustrate the principles of the application. It is appreciated that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. It is intended that only those limitations appearing in the claims define the scope of the application.

Claims

1. An audio encoder for encoding audio input data (11), the audio encoder comprising: a pre-processor (10) for pre-processing the audio input data (11) to obtain audio data to be encoded; an encoder processor (15) for encoding the audio data to be encoded; and a controller (20) for controlling the encoder processor (15) such that, depending on a first signal characteristic of a first frame of the audio data to be encoded, the number of audio data items to be encoded by the encoder processor (15) for the first frame is reduced compared to a second signal characteristic of a second frame, and for a first number of information units used for encoding the reduced number of audio data items for the first frame is increased stronger compared to a second number of information units used for encoding the second frame having the second signal characteristic.

2. The audio encoder according to claim 1, wherein the encoder processor (15) comprises an initial encoding stage (151) and an optimized encoding stage (152), wherein the controller (20) is configured to reduce the number of audio data items to be encoded by the initial encoding stage (151) for the first frame, wherein the initial encoding stage (151) is configured to encode the reduced number of audio data items for the first frame using a first frame initial number of information units, and wherein the optimized encoding stage (152) is configured to optimize encode the reduced number of audio data items for the first frame using a first frame residual number of information units, wherein the first frame initial number of information units added to the first frame residual number of information units results in a predetermined number of information units for the first frame.

3. The audio encoder according to claim 2, wherein the controller (20) is configured to reduce the number of audio data items to be encoded by the initial encoding stage (151) for the second frame to a higher number of audio data items compared to the first frame, wherein the initial encoding stage (151) is configured to encode the reduced number of audio data items for the second frame using a second frame initial number of information units, the second frame initial number of information units being higher than the first frame initial number of information units, and wherein the optimized encoding stage (152) is configured to optimize encode the reduced number of audio data items for the second frame using a second frame residual number of information units, wherein the second frame initial number of information units added to the second frame residual number of information units results in the predetermined number of information units for the first frame.

4. The audio encoder according to claim 1, wherein the encoder processor (15) comprises an initial encoding stage (151) and an optimized encoding stage (152), wherein the initial encoding stage (151) is configured to encode the reduced number of audio data items for the first frame using a first frame initial number of information units, wherein the optimized encoding stage (152) is configured to optimize encode the reduced number of audio data items for the first frame using a first frame residual number of information units, wherein the first frame initial number of information units added to the first frame residual number of information units results in a predetermined number of information units for the first frame. wherein the optimization encoding stage (152) is configured to perform optimization encoding of the reduced number of audio data items of the first frame using a first frame remaining number of information units, wherein the first frame initial number of information units added to the first frame remaining number of information units results in a predetermined number of information units for the first frame, and wherein the controller (20) is configured to control the encoder processor (15) such that the optimization encoding stage (152) performs optimization encoding of at least one of the reduced number of audio data items of the first frame using at least two information units, or such that the optimization encoding stage (152) performs optimization encoding of more than 50 percent of the reduced number of audio data items using at least two information units per audio data item, or wherein the controller (20) is configured to control the encoder processor (15) such that the optimization encoding stage (152) performs optimization encoding of all audio data items of the second frame using less than two information units, or such that the optimization encoding stage (152) performs optimization encoding of less than 50 percent of the reduced number of audio data items using at least two information units per audio data item.

5. The audio encoder of claim 1, wherein the encoder processor (15) comprises an initial encoding stage (151) and an optimization encoding stage (152), wherein the initial encoding stage (151) is configured to encode the reduced number of audio data items of the first frame using a first frame initial number of information units, wherein the optimization encoding stage (152) is configured to perform optimization encoding of the reduced number of audio data items of the first frame using a first frame remaining number of information units, wherein the optimization encoding stage (152) is configured to iteratively allocate (300, 302) the first frame remaining number of information units to the reduced number of audio data items in at least two sequentially performed iterations, to calculate (304, 308, 312) values of the allocated information units for the at least two sequentially performed iterations, and to introduce (316, 318, 320) the calculated values of the information units for the at least two sequentially performed iterations into an encoded output frame in a predetermined order.

6. The audio encoder of claim 5, wherein the optimization encoding stage (152) is configured to sequentially calculate (304) an information unit for each of the reduced number of audio data items of the first frame in a first iteration in an order from low frequency information of the audio data items to high frequency information of the audio data items, wherein the optimization encoding stage (152) is configured to sequentially calculate (308) an information unit for each of the reduced number of audio data items of the first frame in a second iteration in an order from low frequency information of the audio data items to high frequency information of the audio data items, and wherein the optimization encoding stage (152) is configured to sequentially calculate (312) an information unit for each of the reduced number of audio data items of the first frame in a third iteration in an order from low frequency information of the audio data items to high frequency information of the audio data items. wherein the optimization encoding stage (152) is configured to check (314) whether the number of allocated information units is below a predetermined number of information units for the first frame which is less than the first frame initial number of information units, and to stop the second iteration in case of a negative check result, or to perform (312) a number of further iterations in case of a positive check result, the number of further iterations being at least one, or wherein the optimization encoding stage (152) is configured to count a number of non-zero audio terms, and to determine a number of iterations from the number of non-zero audio terms and a predetermined number of information units for the first frame which is less than the first frame initial number of information units.

7. The audio encoder of claim 1, wherein the encoder processor (15) comprises an initial encoding stage (151) and an optimization encoding stage (152), wherein the initial encoding stage (151) is configured to encode a number of most significant information units of each of the reduced number of audio data terms for the first frame using a first frame initial number of information units, the first frame initial number of information units being greater than one, and wherein the optimization encoding stage (152) is configured to encode a number of least significant information units of each of the reduced number of audio data terms for the first frame using a first frame remaining number of information units, the first frame remaining number of information units being greater than one for at least one of the reduced number of audio data terms for the first frame.

8. The audio encoder of claim 1, wherein the first signal characteristic is a first pitch value, wherein the second signal characteristic is a second pitch value, and wherein the first pitch value indicates a higher pitch than the second pitch value, and wherein the controller (20) is configured to reduce a number of audio data terms for the first frame to a first number which is less than a number of audio data terms for the second frame, and to increase an average number of information units used for encoding each of the reduced number of audio data terms of the first frame to a number which is greater than an average number of information units used for encoding each of the reduced number of audio data terms of the second frame.

9. The audio encoder of claim 1, wherein the encoder processor (15) comprises: a variable quantizer (150) for quantizing the audio data of the first frame to obtain quantized audio data for the first frame, and for quantizing the audio data of the second frame to obtain quantized audio data for the second frame; an initial encoding stage (151) for encoding the quantized audio data of the first frame or the second frame; an optimization encoding stage (152) for encoding residual data of the first frame and the second frame; wherein the controller (20) is configured to analyze (26, 28) the audio data of the first frame to determine a first control value (21) for the variable quantizer (150) for the first frame, and to analyze (26, 28) the audio data of the second frame to determine a second control value for the variable quantizer (150) for the second frame, the second control value being different from the first control value (21), and wherein the controller (20) is configured to perform (23, 24) a manipulation of the audio data of the first frame or the second frame or a manipulation of an amplitude related value derived from the audio data of the first frame or the second frame depending on the audio data used for determining the first control value (21) or the second control value, and wherein the variable quantizer (150) is configured to quantize the audio data of the first frame or the second frame without the manipulation.

10. The audio encoder of claim 1, wherein the encoder processor (15) comprises: a variable quantizer (150) for quantizing the audio data of the first frame to obtain quantized audio data for the first frame, and for quantizing the audio data of the second frame to obtain quantized audio data for the second frame; an initial encoding stage (151) for encoding the quantized audio data of the first frame or the second frame; an optimization encoding stage (152) for encoding residual data of the first frame and the second frame; wherein the controller (20) is configured to analyze the audio data of the first frame to determine a first control value (21) for the variable quantizer (150), for the initial encoding stage (151) or for an audio data item reducer (150) for the first frame, and to analyze the audio data of the second frame to determine a second control value for the variable quantizer (150), for the initial encoding stage (151) or for an audio data item reducer (150) for the second frame, the second control value being different from the first control value (21), and wherein the controller (20) is configured (201) to determine a first tonality characteristic as the first signal characteristic to determine the first control value (21), and to determine a second tonality characteristic as the second signal characteristic to determine the second control value, such that a bit budget for the optimization encoding stage (152) is increased in case of the first tonality characteristic compared to a bit budget for the optimization encoding stage (152) in case of the second tonality characteristic, wherein the first tonality characteristic indicates a greater tonality than the second tonality characteristic.

11. Audio encoder in accordance with claim 9, in which the initial encoding stage (151) is an entropy encoding stage for entropy encoding, or in which, the optimization encoding stage (152) is a residual encoding stage or a binary encoding stage for encoding residual data of the first frame and the second frame.

12. The audio encoder of claim 9, wherein the controller (20) is configured to determine the first control value (21) or the second control value such that a first budget for information units of the initial encoding stage (151) is lower than or equal to a predefined value, and wherein the controller (20) is configured to derive a second budget for information units of the optimized encoding stage (152) using the first budget for information units of the first frame or the second frame and a maximum number of information units or the predefined value.

13. The audio encoder of claim 9, wherein the controller (20) is configured to compute (22) the amplitude-dependent value as a plurality of power values derived from one or more audio values of the audio data and to manipulate (24) the power values using a same manipulation value and an addition of the manipulation value to all of the plurality of power values, or wherein the controller (20) is configured to randomly add or subtract (24) a manipulation value to or from all of a plurality of audio values comprised in the frame, or add or subtract a value obtained by a magnitude of the manipulation value, or add or subtract a value obtained by subtracting a term slightly different from the magnitude of the manipulation value, or add or subtract a value obtained as a scaled normalized probability distribution of a sample from a computed complex or real number value using a manipulation value, or wherein the controller (20) is configured to compute (22) the amplitude-dependent value using an exponentiation of the audio data of the first frame or the second frame or of down-sampled audio data of the first frame or the second frame with an exponent value greater than 1.

14. The audio encoder of claim 13, wherein the controller (20) is configured to add or subtract a value obtained by a magnitude of the manipulation value but with a random sign.

15. The audio encoder of claim 9, wherein the controller (20) is configured to compute (23) a manipulation value for the manipulation using a maximum (26) of the audio data for the first frame or the second frame or of the amplitude-dependent values for the first frame or the second frame or using a maximum of a plurality of down-sampled audio data for the first frame and the second frame or a plurality of down-sampled amplitude-dependent values for the first frame or the second frame.

16. The audio encoder of claim 9, wherein the controller (20) is configured to additionally compute (23) a manipulation value for the manipulation using a signal-independent weighting value (27) that depends on at least one of a bit rate, a frame duration and a sampling frequency for the first frame or the second frame.

17. The audio encoder of claim 9, wherein the controller (20) is configured to compute (23, 29) a manipulation value for the manipulation using signal dependent weighting values derived from at least one of a first sum of magnitudes of the audio data or downsampled audio data within the frame, a second sum of magnitudes of the audio data or the downsampled audio data within the frame multiplied by an index associated with each magnitude, and a quotient of the second sum and the first sum.

18. The audio encoder of claim 9, wherein the controller (20) is configured to compute (29) a manipulation value for the manipulation based on the following equation: where k is a frequency index, where X f (k) is an audio data value for the frequency index k prior to quantization, where max is a maximum function, where regBits is a first signal-independent weighting value, and where lowBits is a second signal-dependent weighting value.

19. The audio encoder of claim 1, wherein the pre-processor (10) further comprises: a time-to-frequency converter (14) for converting time domain audio data into spectral values of the frame; and a spectral processor for computing modified spectral values having a spectral envelope that is flatter than a spectral envelope of the spectral values, wherein the modified spectral values represent the audio data items of the first frame or the second frame to be encoded by the encoder processor (15).

20. The audio encoder of claim 19, wherein the spectral processor (15) is configured to perform at least one of a time noise shaping operation, a spectral noise shaping operation, and a spectral whitening operation.

21. The audio encoder of claim 9, wherein the controller (20) is configured to compute the first control value (21) or the second control value using a plurality of energy values as the amplitude dependent values for the frame, wherein each energy value of the plurality of energy values is derived (22, 23, 24) from a power value that is an amplitude dependent value of a plurality of amplitude dependent values for the frame and a signal dependent manipulation value for the manipulation.

22. The audio encoder of claim 21, wherein the controller (20) is configured to compute a required bit estimate for each energy value of the plurality of energy values depending on the energy value and a candidate value for the first control value (21) or the second control value, accumulate the required bit estimates for the energy values of the plurality of energy values and the candidate value for the first control value (21) or the second control value, check whether the accumulated bit estimate for the candidate value for the first control value (21) or the second control value meets an allowed bit consumption criterion, and modify the candidate value for the first control value (21) or the second control value in case the allowed bit consumption criterion is not met and repeat the computation of the required bit estimate, the accumulation of the bit estimates, and the check until a modified candidate value for the first control value (21) or the second control value is found for which the allowed bit consumption criterion is met.

23. The audio encoder of claim 21, wherein the controller (20) is configured to compute the plurality of energy values based on the following equation: ​ E(k) = 10 log 10 (PX lp (k) + N(X f )+ 2 -31 ), where E(k) is an energy value of the plurality of energy values for index k, where PX lp (k) is a power value for index k as the amplitude dependent value, and where N(X f ) is the signal dependent steering value.

24. The audio encoder of claim 9, wherein the controller (20) is configured to compute the first control value (21) or the second control value based on an estimate of the cumulative information units required for each manipulated audio data value or manipulated amplitude related value.

25. The audio encoder of claim 9, wherein the controller (20) is configured to manipulate in such a way that the manipulation increases the bit budget for the initial encoding stage (151) or decreases the bit budget for the optimized encoding stage (152).

26. The audio encoder of claim 9, wherein the controller (20) is configured to manipulate in such a way that the manipulation results in a higher bit budget for the optimized encoding stage (152) for signals having a first pitch compared to signals having a second pitch, wherein the second pitch is lower than the first pitch.

27. The audio encoder of claim 9, wherein the controller (20) is configured to manipulate in such a way that the energy of the audio data used to compute the bit budget for the initial encoding stage (151) is increased relative to the energy of the audio data to be quantized by the variable quantizer (150).

28. The audio encoder of claim 1, wherein the encoder processor (15) comprises a variable quantizer (150) for quantizing the audio data of the first frame to obtain quantized audio data for the first frame and for quantizing the audio data of the second frame to obtain quantized audio data for the second frame, wherein the controller (20) is configured to compute a global gain for the first frame or the second frame, and wherein the variable quantizer (150) comprises: a weighter (155) for weighting the audio data of the first frame or the audio data of the second frame with the global gain; and a quantizer core (157) having a fixed quantization step size.

29. The audio encoder of claim 1, wherein the encoder processor (15) comprises an initial encoding stage (151) and an optimized encoding stage (152), wherein the optimization encoding stage (152) is configured for calculating in a plurality of iterations an optimized number of bits for quantized audio values, wherein wherein the optimization bit indication is different in each iteration, or wherein the optimization bit indication is higher in lower iterations than in higher iterations, or wherein the amount is a fraction of a quantizer step size indicated by the first control value (21) or the second control value.

30. The audio encoder of claim 1, wherein the encoder processor (15) comprises an optimized encoding stage (152), wherein the optimized encoding stage (152) is configured (304, 308, 312) to perform an iterative process having two iterations comprising at least a first iteration and a second iteration, checking whether the quantized audio values in the first iteration or the quantized audio values together with a potential first number of optimization bits associated with the quantized audio values, when added or subtracted by a global gain weighting, are greater or smaller than the unquantized audio values, and setting the number of optimization bits for the second iteration depending on the result of the checking.

31. The audio encoder of claim 1, wherein the encoder processor (15) comprises a variable quantizer (150) and an optimization encoding stage (152), wherein the optimization encoding stage (152) is configured to calculate optimization bits only for audio values that are not quantized to zero by the variable quantizer (150).

32. The audio encoder of claim 1, wherein the controller (20) is configured to reduce the influence of the manipulation for audio data having a centroid at a lower frequency, and wherein an initial encoding stage (151) of the encoder processor (15) is configured to remove high frequency spectral values from the audio data in case the bit budget for the first frame or the second frame is not sufficient for encoding the quantized audio data of the frame.

33. The audio encoder of claim 1, wherein the controller (20) is configured to perform a binary search for each frame using the manipulated spectral energy values for the first frame or the second frame as manipulated amplitude dependent values for the first frame or the second frame, respectively.

34. The audio encoder of claim 1, wherein the first signal characteristic is a first tonality, wherein the second signal characteristic is a second tonality, and wherein the first tonality is greater than the second tonality.

35. A method of encoding audio input data, comprising: preprocessing the audio input data (11) to obtain audio data to be encoded; encoding the audio data to be encoded; and controlling the encoding such that, depending on a first signal characteristic of a first frame of the audio data to be encoded, the number of audio data items of the audio data to be encoded for the first frame is reduced compared to a second signal characteristic of a second frame, and a first number of information units used for encoding the reduced number of audio data items for the first frame is increased more strongly than a second number of information units used for encoding the second frame having the second signal characteristic.

36. The method of claim 35, wherein the encoding comprises: variable quantizing audio data of a frame to obtain quantized audio data; entropy encoding the quantized audio data of the frame; and encoding residual data of the frame; wherein the controlling the encoding comprises determining a control value for variable quantizing audio data of a frame, the determining comprising analyzing the audio data of the first frame or the second frame; and performing a manipulation of the audio data of the first frame or the second frame or of an amplitude-related value derived from the audio data of the first frame or the second frame depending on the audio data used for determining the control value, wherein a variable quantization of the audio data of a frame quantizes the audio data of the frame without the manipulation, or wherein the controlling the encoding comprises determining a first tonality characteristic or a second tonality characteristic of the audio data and determining the control value such that a bit budget for encoding the residual data is increased in case of the first tonality characteristic compared to a bit budget for encoding the residual data in case of the second tonality characteristic, wherein the first tonality characteristic indicates a greater tonality than the second tonality characteristic.

37. A digital storage medium having stored thereon a computer program for performing, when running on a computer or processor, a method according to claim 35.

Citation Information

Patent Citations

  • Hybrid coded audio data streaming apparatus and method

    US20120290306A1

  • Audio encoder for encoding an audio signal, method for encoding an audio signal and computer program under consideration of a detected peak spectral region in an upper frequency band

    US20190156843A1

  • Audio encoders, audio decoders, methods and computer programs adapting an encoding and decoding of least significant bits

    WO2019091576A1