Encoding and decoding audio signal

By employing frame or subframe partitioning with different sampling rates and innovative codebooks in the speech codec, the problem of low coding efficiency at low bit rates is solved, and flexible pulse position positioning and audio bandwidth preservation are achieved.

CN122029602APending Publication Date: 2026-05-12FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2024-07-31
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing speech codecs struggle to adapt to the target bit rate at low bit rates, and the complex adjustment of the sampling rate during encoding leads to low encoding efficiency. They are particularly unusable when the number of samples is not a power of 2, and reducing the sampling rate limits the audio bandwidth of the codec.

Method used

By dividing the audio signal into multiple frames or subframes and encoding and decoding them using different sampling rates, an innovative codebook is used to define the pulse combination set under the second sampling and generate the decoded audio signal under the first sampling, thus achieving flexible encoding of the pulse position and avoiding direct adjustment of the sampling rate.

Benefits of technology

It enables flexible pulse positioning at low bit rates, improves coding efficiency, reduces the complexity of sampling rate adjustment, adapts to different bit rate requirements, and maintains the audio bandwidth of the codec.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122029602A_ABST
    Figure CN122029602A_ABST
Patent Text Reader

Abstract

The present technology relates to encoding and decoding audio signals, such as applications in encoders, decoders, methods for encoding or decoding, and non-transitory storage units for controlling encoding or decoding. For example, the present technology relates to innovative mapping and / or resampling of codebooks. An apparatus for generating a decoded audio signal divided into a plurality of frames or subframes, comprising: a codec signal reader (722) configured to read at least codec information about a prediction coefficient (105 ') and codec information (119) about at least one pulse from a codec signal (124); a signal processor (703) configured to generate a decoded audio signal (702) from at least a decoded version (704) of prediction coefficients and a decoded pulse combination (710) or a processed version thereof, where the decoded audio signal (702) is generated in a first sample that means a first plurality of sampling positions having a first number of sampling positions in one frame or subframe; wherein the apparatus is configured to derive (716) a decoded pulse combination (710) from codec information (119) about at least one pulse and a second sample codebook (118), where the at least one second sample codebook (118) comprises a set of pulse combinations defined under second samples, the second samples means a second plurality of sampling positions having a second number of sampling positions in the frame or subframe, wherein the first sample and the second sample differ at least in that the first plurality of sample locations differ from the second plurality of sample locations.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ‌ Technical Field ‌

[0002] This technology relates to encoding and decoding audio signals, such as in encoders, decoders, methods for encoding or decoding, and non-transitory storage units that control encoding or decoding. For example, this technology relates to the mapping and / or resampling of innovative codebooks.

[0003] Background Technology ‌

[0004] An audio codec (e.g., a speech codec) is known that relies on a codebook (e.g., an innovative codebook) to quantize prediction residuals, such as those from linear prediction (LP) and long-time prediction (LTP). Specifically, for encoding prediction residual signals (e.g., excitation signals), the position, amplitude, and symbol of the pulses can be encoded, and then they can be decoded.

[0005] Despite its widespread adoption, it still faces some challenges.

[0006] For example, in some cases, it would be preferable to further reduce the number of bits in the bitstream.

[0007] Furthermore, it is often difficult to adapt to the target bit rate. During encoding, it is usually desirable to maintain the sampling rate of the input audio signal, which makes changing the bit rate difficult.

[0008] A more detailed discussion will follow.

[0009] In speech coding and decoding using CELP, an innovative codebook is used to quantize the prediction residuals from Linear Prediction (LP) and Long Time Prediction (LTP). Unlike LP coding and decoding (where the spectral envelope is encoded per time frame), the parameters of LTP and the residuals are quantized over multiple parts of a frame (called subframes). In the specific case of ACELP (Algebraic CELP, i.e., CELP using algebra and an innovative codebook), the innovative codebook is defined by algebraic codes that encode the temporal position and symbol of pulses within a given subframe. The parameters of these pulses are optimized during the encoding process using a least-squares algorithm. While the theoretically possible number of positions for a given number of pulses within a subframe is determined only by the subframe length and sampling rate, the algebraic coding process selects pulse configurations from a subset of the cardinality constrained by the available bit budget.

[0010] In existing ACELP implementations (such as in 3GPP EVS), different sampling rates are applied to different bit rates so that the additional available bits can be used for increased temporal resolution. This increased resolution, and consequently more possible pulse positions, comes at the cost of a reduced number of coded pulses. This technique, in particular, provides an algebraic encoding / decoding scheme that allows pulse positioning at a lower bit rate without reducing the total number of pulses by systematically excluding pulse positions.

[0011] For efficient residual coding, it is convenient and common to encode the number of possible pulse positions as a power of 2. If the number of samples per frame is a power of 2, this can be achieved by dividing the frame into an appropriate number of subframes. This encoding / decoding scheme has two drawbacks: First, it cannot be applied if the number of samples is not a power of 2. Second, the bit consumption of the LTP parameters and residual code increases with the number of subframes.

[0012] Innovative CELP codebooks are often highly constrained. For example, in ACELP, each subframe is divided into tracks with staggered positions. As mentioned above, for convenience, complexity, and optimal coding, the number of positions per track is typically the same, a multiple of 2. For instance, for a 64-sample subframe, two tracks with 32 samples each, or four tracks with 16 samples each, can be designed. The codebook is then designed to distribute the pulse budget evenly or almost evenly across the tracks, thus achieving an equal or almost equal number of pulses per track.

[0013] Therefore, for low bit rates, when the number of pulses is limited, the number of tracks must be reduced, which may be impossible or very complex because it does not result in tracks of equal size or multiples of 2. Another more practical solution is to reduce the sampling rate of the speech codec CELP, which automatically reduces the number of possible positions. This is typically used for wideband or ultrawideband speech codecs operating at bit rates below or around 16 kbps, where the baseband CELP encoder operates at only 12.8 kHz. The disadvantage of reducing the internal sampling rate of the baseband codec is that the baseband codec's encoding and decoding audio bandwidth is further limited, and resampling memory and buffers are required when switching to or from a higher bit rate.

[0014] For a subframe with 64 samples, here's an example of the potential positions of each pulse in a 2-pulse algebraic book using two 32-position tracks:

[0015] track pulse Location 1 0 0, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62 2 1 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63

[0016] For a 64-sampled subframe, here are examples of the potential positions of each pulse in a 4-pulse bit algebra book using 4 tracks of 64 positions:

[0017] track pulse Location 1 0 0, 4, 8, 12, 16, 20, 24, 28, 32 36, 40, 44, 48, 52, 56, 60 2 1 1, 5, 9, 13, 17, 21, 25, 29, 33, 37, 41, 45, 49, 53, 57, 61 3 2 2, 6, 10, 14, 18, 22, 26, 30, 34, 38, 42, 46, 50, 54, 58, 62 4 3 3, 7, 11, 15, 19, 23, 27, 31, 35, 39, 43, 47, 51, 55, 59, 63

[0018] Figure 2 provides an example where 64-sampled frames are divided into four staggered 16-sample tracks, each containing 16 samples (circular, cross, diamond, and star).

[0019] Summary of the Invention ‌

[0020] An apparatus for generating a decoded audio signal divided into multiple frames or subframes is disclosed, comprising:

[0021] A codec signal reader is configured to read at least codec information about prediction coefficients and codec information about at least one pulse from a codec signal;

[0022] A signal processor is configured to generate a decoded audio signal from at least a decoded version of the prediction coefficients and a combination of decoded pulses or a processed version thereof, wherein the decoded audio signal is generated with a first sample, the first sample meaning a first plurality of sample positions having a first number of sample positions in a frame or subframe;

[0023] The apparatus is configured to derive a decoded pulse combination from encoding / decoding information about at least one pulse and a second sample codebook, wherein the at least one second sample codebook contains a pulse combination set defined under the second sample, the second sample meaning a second plurality of sample positions having a second number of sample positions in a frame or subframe, wherein the first sample differs from the second sample at least in that the first plurality of sample positions differ from the second plurality of sample positions.

[0024] An apparatus for encoding an audio signal divided into multiple frames or subframes is disclosed, the apparatus comprising:

[0025] The signal processor is configured to determine the prediction coefficients and prediction residual signal in the time domain under a first sample, where the first sample means a first plurality of sampling positions having a first number of sampling positions in a frame or subframe.

[0026] A pulse information encoder is configured to determine encoding / decoding information regarding selected pulse combinations representing a prediction residual signal, the selected pulse combination being an entry in at least one second sample codebook, wherein the at least one second sample codebook contains a set of pulse combinations defined under a second sample, the second sample meaning a second plurality of sample positions having a second number of sample positions in a frame or subframe, wherein the first sample differs from the second sample at least in that the first plurality of sample positions differs from the second plurality of sample positions; and

[0027] The codec signal writer is configured to write at least codec information about the prediction coefficients (104) and codec information about the selected pulse combination.

[0028] A method for decoding an audio signal from an encoded audio signal is disclosed, comprising:

[0029] Read encoding / decoding information about the prediction coefficients and encoding / decoding information about at least one pulse from the encoded / decoded signal; and

[0030] In the first sampling, a decoded audio signal is generated from at least the decoded version of the prediction coefficients and the combination of decoded pulses or their processed version. The first sampling means a first plurality of sampling positions having a first number of sampling positions in a frame or subframe.

[0031] The method includes deriving a decoded pulse combination from encoding / decoding information about at least one pulse and a second sample codebook, wherein at least one second sample codebook contains a pulse combination set defined under a second sample, the second sample meaning a second plurality of sample positions having a second number of sample positions in a frame or subframe, wherein the first sample differs from the second sample at least in that the first plurality of sample positions differ from the second plurality of sample positions.

[0032] An audio encoding method for encoding audio signals is disclosed, comprising:

[0033] The prediction coefficients and prediction residual signal in the time domain are determined under the first sampling. The first sampling means the first plurality of sampling positions with a first number of sampling positions in a frame or subframe.

[0034] Determine encoding and decoding information for selected pulse combinations representing the predicted residual signal, wherein the selected pulse combination is an entry in at least one second sample codebook, wherein the at least one second sample codebook contains a set of pulse combinations defined under a second sampling, wherein the second sampling means a second plurality of sampling positions having a second number of sampling positions in a frame or subframe, wherein the first sampling differs from the second sampling at least in that the first plurality of sampling positions differs from the second plurality of sampling positions; and

[0035] At least the encoding / decoding information about the prediction coefficients and the encoding / decoding information about the selected pulse combination should be written.

[0036] A non-temporary storage unit for storing instructions is disclosed, which causes the processor to perform the method described above when the instructions are executed by the processor.

[0037] Attached Figure Description ‌

[0038] Figure 1 An example of the encoder's operating modes is shown.

[0039] Figure 2 shows the subframes being split into tracks.

[0040] Figure 3aAn example of the first operating step of the encoder is shown.

[0041] Figure 3b It shows in Figure 3a An example of the encoder's second operation step following the first operation step.

[0042] Figure 4 An example of an encoder is shown.

[0043] Figure 5 It shows Figure 4 An example of encoder operation.

[0044] Figure 6 shows Figure 4 An example of encoder operation.

[0045] Figure 7 An example of a decoder is shown.

[0046] Figure 8 It shows Figure 7 Example of decoder operation.

[0047] Figure 9 An example of an encoder is shown.

[0048] Figure 10 An example of a decoder is shown.

[0049] Detailed Implementation ‌

[0050] Figure 4A means 400 (encoder) for encoding an input audio signal 102 (e.g., speech) onto a codec signal 124 (e.g., a bitstream) is shown, for example according to CELP (specifically ACELP, i.e., CELP with an algebraic and innovative codebook). Encoder 400 may be a CELP encoder (e.g., a CELP with an algebraic and innovative codebook). Encoder 400 may include a signal processor 103. Signal processor 103 may determine (e.g., through LP or LPC analysis) prediction coefficients 104 and prediction residual signals (excitation) 110. Prediction coefficients 104 and prediction residual signals 110 may be in the time domain. Prediction coefficients 104 may be encoded, for example, by prediction coefficient encoder 105, as codec information 105' about prediction coefficients 104. Prediction residual signals 110 may be provided to impulse information encoder 116, which may include (or be connected to) a codebook 118 (e.g., an innovative codebook). Codebook 118 may be an algebraic codebook. Codebook 118 may be a priori known (e.g., the correspondence between codebook entries and specific pulse combinations may be a priori known). Codebook 118 may be the same as that used by the decoder (see below). Pulse information encoder 116 may provide encoding / decoding information 119 about selected pulse combinations representing the prediction residual signal 110. Encoding / decoding information 119 about selected pulse combinations may be identified by a corresponding entry in codebook 118 (e.g., a codebook index). Codebook 118 may contain a set of pulse combinations. Therefore, pulse information encoder 116 may select, from among multiple pulse combinations in this set of pulse combinations in codebook 118, the pulse combination that best represents the prediction residual signal 110, for example, minimizing the error (e.g., 162 in FIG. 6, see below) or cost function, and thus may provide encoding / decoding information about selected pulse combinations as encoding / decoding information 119. Encoding / decoding information 105' regarding prediction coefficients 104 (or more generally, any method of encoding a version of prediction coefficients 104) and encoding / decoding information 119 regarding selected pulse combinations (representing prediction residual signal 110) can be provided to the codec signal writer 122. The codec signal writer 122 can output a codec signal 124 (e.g., a bitstream). The codec signal writer 122 may include an entropy encoder (but may omit this). Therefore, the codec signal 124 can be stored and / or transmitted to a receiving device, which may include means for decoding the codec signal 124 (e.g., a decoder, such as...). Figure 7 decoder 700 and Figure 10 The decoder 1000 (see below).

[0051] The input audio signal 102 can be subdivided into multiple consecutive frames and / or subframes (e.g., a single frame may include more than one subframe). The input audio signal 102 can be provided in the form of a first sample. The first sample can mean a first plurality of sample positions (e.g., a set of sample positions) having a first number of sample positions (e.g., a first cardinality of a first set of sample positions) on a frame or subframe. Thus, according to the first sample, each sample of the input audio signal 102 can have a time-domain value at each sample position. For example, each frame or subframe can include the same first plurality of time slots (e.g., 80 time slots of the same time length in 80 sample positions) and have a first number of sample positions (e.g., 80 sample positions). The first sample can be associated, for example, with a first sampling rate (e.g., 80 samples per frame or subframe, for example, 16 kHz in the case of a frame or subframe of 5 milliseconds, i.e., 16,000 samples per second).

[0052] However, codebook 118 can be at a second sampling level (which is why it can be called a second-sample codebook), where the second sampling differs from the first sampling. The second sampling can mean a second number of sampling positions (e.g., a second set of sampling positions) with a second number of sampling positions (e.g., 64 sampling positions) within the same frame or subframe, where the second number differs from the first number of sampling positions (e.g., the cardinality of the second set can be less than that of the first set). The first sampling differs from the second sampling (e.g., the first set of sampling positions can differ from the second set of sampling positions, and / or the first number of sampling positions can differ from the second number of sampling positions; for example, the first sampling can have a first sampling rate higher than the second sampling rate). Therefore, the frame or subframe of the input audio signal 102 and the prediction residual signal 110 can be at the first sampling level (e.g., 80 samples per frame or subframe), while codebook 118 can be at the second sampling level (e.g., codebook 118 can output a pulse combination of 64 sampling positions for the same frame or subframe). By providing encoding / decoding information 119 about selected pulse combinations at the second sampling level, the bit rate is reduced. It is worth noting that the prediction coefficients 104 (or processing version 105') can be at the first sampling, but the prediction residual signal can be at the second sampling. In the example, the first sampling can mean a first sampling rate, and the second sampling can mean a second sampling rate less than the first sampling rate (e.g., the first sampling rate can mean 80 sampling positions per frame or subframe, while the second sampling can mean 64 sampling positions per frame or subframe, or for example, the first sampling rate can mean 16 kHz, while the second sampling rate can mean 12.8 kHz).

[0053] As will be shown later, in the first alternative (e.g., Figure 5As shown in Figure 6, the transition from the first sample to the second sample is achieved by ignoring some sampling positions of the predicted residual signal 110, while according to the second alternative (e.g., Figure 3a and Figure 3b As shown, the conversion from the first sample to the second sample is achieved by resampling the prediction residual signal 110 or the code vector 118'' of the codebook 118 or its processed version (318'').

[0054] Figure 6 shows a more detailed example of the encoder 400 according to the first alternative. Here, the input audio signal 102 is also mathematically represented as s(n) (where n is the sampling position according to the first sample). The input audio signal 102 can be provided, for example, to an analysis block (short-term prediction block) 130, which can be part of a signal processor 103. The LP analysis block 130 can output prediction coefficients 104, which can be provided to a prediction coefficient encoder 105. The LP analysis block 130 can also control an LPC linear prediction codec analysis filter block 132, mathematically represented as 1 / A(z) (which can also be part of the signal processor 103). The LPC linear prediction codec analysis filter block 132 can output a prediction residual signal (excitation) 110. The prediction residual signal (excitation) 110 can be provided to a pulse information encoder 116. Then, the pulse information encoder 116, as indicated in FIG. 6, includes at least some of blocks 118, 134, 136, 140, 144, 148, 152, 155, 160, and 164. The predicted residual signal 110 (also mathematically represented as r(n)) can be provided to the optimization section 172 via line 110a. The task of the optimization section 172 is to find the pulse combination that best represents the predicted residual signal 110. In the optimization section 172, a plurality of candidate excitation signals 150 (mathematically represented as exc(n)) are iteratively evaluated to find the candidate excitation signal that best represents the residual signal 110. As described below, each candidate excitation signal 150 is obtained from two components: the predicted component 146 and the innovative component 158.

[0055] Input 150b (described below, which is the previous codec excitation 150 or its processed version) can be provided to filter / delay block 136, mathematically represented as P(z), for example, P(z) = z -PType. The filter / delay block 136 can be controlled by a long-term prediction parameter 135, which corresponds to the hysteresis P given by the LTP analysis block 134. The LTP analysis block 134 can, for example, obtain a pitch hysteresis 157, through which the long-term prediction parameter 135 of the filter / delay block 136 is controlled. It is worth noting that the pitch hysteresis 157 can be iteratively optimized (and multiple candidate pitch hysteresis 157s can be tried before finding the most suitable pitch hysteresis). It is important to note that the pitch hysteresis can also have a fractional component, in which case the filter / delay P(z) is a filter composed of an integer number of sample delays and interpolation (such as linear interpolation). The input 150b of the filter / delay block 136 is the past encoding / decoding excitation 150, as reconstructed on the encoder and decoder sides. The output 138 of the filter / delay block 136 can be provided to the adaptive codebook 140. The adaptive codebook 140 can output component 142 (predicted signal, or predicted component of candidate excitation signal 150), mathematically represented as p(n). The output 142 (p(b)) of the adaptive codebook 140 can be scaled at scaler 144 by gain 156 to provide the predicted component 146 of candidate excitation signal 150. The gain 156 used to scale the predicted signal p(n) 142 can be obtained iteratively, for example, by iteratively trying multiple gains and evaluating the gain that provides the best result through analytical synthesis 155, which will be discussed below. It can be understood that the adaptive codebook 140 can contain past encoded / decoded excitation vectors (excitation signals) such that the predicted signal 142 represents the prediction of candidate excitation signal 150. Therefore, the excitation signal 150 can be obtained by adding the predicted component 146 (considering past excitations) obtained from the adaptive codebook 140 at adder 148 to the innovative component 158 ​​obtained from the innovative codebook 118. The innovative codebook 118 and... Figure 4The same applies as indicated in the code. The innovation component 158 ​​of the candidate excitation signal is obtained from the innovation codebook 118 and added to the prediction component 146 at adder 148 to obtain the candidate excitation signal 150. It can be understood that during the execution of loop 179, the excitation signal 150 is actually the candidate excitation signal because it is necessary to evaluate the best excitation signal that best approximates the prediction residual signal 110. The candidate excitation signal 150 can be compared with the prediction residual signal 110 obtained from signal processor 103 (e.g., from block 132). Therefore, an error 162 (or more generally, a cost function) mathematically represented as e(n) can be obtained (e.g., e(n) = abs(r(n) - exc(n)), where “abs” represents the absolute value (which can also be written as e(n) = |r(n) - exc(n)|), but can be replaced by any other norm (e.g., e(n) = ||r(n) - exc(n)||), where n is the sampling position based on the first sample. Ideally, the error 162 should be 0, but since this is generally not achievable, a technique for minimizing the error e(n) (162) is used by iteratively searching for candidate stimuli 150 that minimize the error 162. The error 162 can be filtered at a weighted filter block 164 mathematically represented as W(z). Thus, a processed version 166 of the error, mathematically represented as e, is obtained. w (n), and provides it to the analysis synthesis optimization block 155 to evaluate the error e among several other errors obtained in other iterations of loop 179. w(n). After evaluation, the analysis synthesis optimization block 155 can select a specific pulse combination that best represents the predicted residual signal 110 and provide encoding / decoding information 119 about the selected pulse combination to the encoding / decoding signal writer 122. The encoding / decoding information 119 about the selected pulse combination is obtained iteratively or otherwise by searching for pulse combinations that minimize the error 162 (166). For example, it can be obtained by iteration 179. Here it is shown that the analysis synthesis optimization block 155 provides indices (entries) 119' to the codebook 118. The codebook 118 outputs the associated candidate pulse combination 118' for each candidate index 119'. It is worth noting that the candidate pulse combination 118' is in the second sample, but is mapped to the version 118a' in the first sample by the mapper 118a. The pulse combination 118a' in the second sample is scaled at the scaler 152 by the candidate gain 154 (controlled by the gain information output by the analysis synthesis optimization block 155). Therefore, the scaled version 158 of the pulse combination 118' (118a') can be understood as another component (innovative component 158) of the candidate excitation signal 150. Thus, the analysis synthesis optimization block 155 can iteratively provide multiple candidate indices 119' to the codebook 118, thereby iteratively finding the pulse combination that minimizes the error 162 (166) with the predicted residual signal 110 among all evaluated candidate pulse combinations 119'. It should be noted that the pulse combination 118a' in the second sample can be further processed by one or more filters or processors. For example, formant sharpening can be applied based on the version or weighted version of the LPC coefficients. Pitch sharpening can also be applied based on the LTP parameters. Another possibility is to emphasize high frequencies, as taught in several ACELP implementations in EVS.

[0056] It should be noted that the past excitation signal 150b is not necessarily the same as the candidate excitation signal 150, but rather the best excitation signal among the candidate excitation signals 150 obtained from previous frames or subframes.

[0057] In summary, iterative optimization (through iteration 179) allows finding the optimal candidate excitation signal 150 to approximate the prediction of the residual signal 110. Since the optimal candidate excitation signal 150 is associated with a specific candidate excitation prediction component 146 (associated with a specific gain 156 and a specific pitch hysteresis 157) and a specific candidate excitation innovation component 158 ​​(associated with a specific gain 154 and a specific codebook index 119'), the parameters 156, 154, 156, and 119' associated with the optimal approximate excitation signal 150 can be easily encoded in the encoded / decoded signal 124.

[0058] According to the first alternative, the input audio signal 102, the prediction residual signal 110, the excitation prediction component 146, and the excitation innovation component 158, as well as the excitation 150 and the error 162 (and its version 166), are sampled according to a first sampling (e.g., 80 samples per frame or subframe). However, candidate index 119' is sampled at a second sampling (e.g., a second, fewer number of sample positions per frame or subframe), codebook 118 operates at a second sampling, and candidate pulse combination 118' is also sampled at a second sampling. Mapper 118a maps candidate pulse combination 118' from the second sampling (e.g., 64 samples per frame or subframe) to version 118a' of candidate pulse combination 118' at the first sampling (e.g., 80 samples per frame or subframe). Preferably, scaler 152 is sampled at a first sampling (e.g., a higher sampling, e.g., 80 sample positions per frame or subframe).

[0059] This section explains how, according to the first alternative, sampling is reduced from a first sample (e.g., a first plurality of sampling positions per frame or subframe, e.g., the first number could be 80) to a second sample (e.g., a second plurality of sampling positions per frame or subframe, e.g., the second number could be 64). Within each frame or subframe, each sampling position can be numbered: for example, the first sampling position can be represented as 0, the second sampling position (e.g., temporally immediately following the first sampling position) can be represented as 1, the third sampling position (e.g., temporally immediately following the second sampling position) can be represented as 2, the fourth sampling position (e.g., temporally immediately following the third sampling position) can be represented as 3, the fifth sampling position (e.g., temporally immediately following the fourth sampling position) can be represented as 4, and the sixth sampling position (e.g., temporally immediately following the fifth sampling position) can be... The seventh sampling position (e.g., immediately following the sixth sampling position in time) can be represented as 5, the eighth sampling position (e.g., immediately following the seventh sampling position in time) can be represented as 7, the ninth sampling position (e.g., immediately following the eighth sampling position in time) can be represented as 8, the tenth sampling position can be represented as 9, ... the 76th sampling position can be represented as 75, the 77th sampling position can be represented as 76, the 78th sampling position can be represented as 77, the 79th sampling position can be represented as 78, and the 80th sampling position can be represented as 79.

[0060] track pulse Sampling location 1 0,4 0, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75 2 1,5 1, 6, 11, 12, 21, 26, 31, 36, 41, 46, 51, 56, 61, 66, 71, 76 3 2,6 2, 7, 12, 17, 22, 27, 32, 37, 42, 47, 52, 57, 62, 67, 72, 77 4 3,7 3, 8, 13, 18, 23, 28, 33, 38, 43, 48, 53, 58, 63, 68, 73, 78 5 No pulse 4, 9, 14, 19, 24, 29, 34, 39, 44, 49, 54, 59, 64, 69, 74, 79 (excluded in the second sampling)

[0061] Furthermore, as shown in the table above, the first plurality of samples are defined according to a plurality of tracks (e.g., tracks 1, 2, 3, 4, and 5), which can be regularly staggered with each other. For example, track 1 includes a first sampling position 0, a sixth sampling position 5, ... and a 76th sampling position 75; track 2 includes a second sampling position 1, a seventh sampling position 6, ... and a 77th sampling position 76; and track 5 includes a fifth sampling position 4, a tenth sampling position 9, ... and an 80th sampling position 79. Thus, each sampling position of each track immediately precedes the sampling position of the track immediately following it and follows the sampling position of the track immediately preceding it (a sampling of track 5 is followed by a sampling of track 1). For each track, there can be a predefined number of pulses, which are specifically defined. It is noted here that each of tracks 1, 2, 3, and 4 can have two pulses. Depending on a particular aspect, at least one track (in this case, track 5) is an empty track with no pulses at all. Therefore, if the application specifically has two pulses per track, there can only be eight pulses, and only in the sampling positions of tracks 1, 2, 3, and 4, but no sampling position in track 5 is allowed to carry a pulse (more generally, the second plurality of sampling positions can be a proper subset of the first plurality of sampling positions). Therefore, sampling positions 5, 9, 14… and 79 cannot carry any pulses. Tracks 1, 2, 3, 4, and 5 are the first plurality of sampling tracks that uniformly form 80 sampling positions in the first sampling, while tracks 1, 2, 3, and 4 (but excluding track 5) are the second plurality of sampling tracks that uniformly form in the second sampling. Although the tracks forming the first plurality of sampling are considered in signals 102, 160, 150, 146, 158, 162, and 166 (and all their sampling positions are occupied by some value), the excluded tracks (empty tracks) do not form the second plurality of sampling (and their sampling positions do not carry any value, or carry an ignored value).

[0062] Candidate index 119' (iteratively provided to innovation codebook 118 by analysis synthesis optimization block 155) contains only information about the first plurality of sampling positions (e.g., tracks 1, 2, 3, and 4), but not about the fifth track (equivalently, it can be said that codebook 118 ignores track 5, even when track 5 is provided to codebook 118). Therefore, the candidate pulse combination 118' provided by innovation codebook 118 lacks the sampling position of track 5. Essentially, innovation codebook 118 ignores track 5. In the example, mapper 118a can map the second plurality of sampling positions (i.e., tracks 1, 2, 3, and 4) onto the version of the first plurality of sampling positions by adding samples (e.g., zero-value samples) at the sampling positions of empty track 5. Therefore, version 118a' of candidate pulse combination 118' is an upsampled version of candidate pulse combination 118' in the first sample. However, in the example, there is no pulse at the sampling position of empty track 5 because the value of the normally added sampling position is 0. Subsequently, the first sampled version 118a' of the candidate pulse combination 118' is scaled by the candidate gain 154 (based on the gain information provided by the analysis synthesis optimization block 155) to obtain the innovative component 158 ​​of the excitation 150.

[0063] Figure 5 The operation is conceptually illustrated as transferring data from a first plurality of sampling positions (tracks 1, 2, 3, 4, and 5) to a second plurality of sampling positions.

[0064] It should be noted that the second number of sampling positions is preferably a power of 2 (e.g., 2^32). N Where N is a positive integer, e.g., 64), and the difference between the first number of sampling positions (e.g., 80) and the second number of sampling positions (e.g., 64) can also be a power of 2 (e.g., 2^3). M Where M is a positive integer less than N (e.g., 16). More generally, the length of each track is also preferably a power of 2. This property makes it easy to encode and decode pulse positions into binary format, and the resulting encoding and decoding is often mathematically quasi-optimal or even optimal, with low complexity. Such encoding and decoding schemes are often already available in a given system. In the latter case, the present invention enables the reuse of existing encoding and decoding schemes for new advantageous combinations of bit rates and sampling for CELP without redefining or redesigning the pulse position encoding and decoding.

[0065] It is conceivable that simply discarding one or more tracks from the first plurality of samples would result in a deterioration in the approximation of the candidate excitation 150 to the predicted residual signal 110, thereby reducing coding quality. However, experience shows that the quality reduction is not significant, but the bit rate savings are advantageous. This same bit rate savings can be advantageously reinvested to allow for more pulses or coarser encoding and decoding of other codec parameters.

[0066] This section discusses examples of encoders based on the second alternative. Figure 3a and 3b Another example of a device 400 (specifically denoted as 100 here) for encoding audio signal 102 is shown. Elements 130 and 132 (103), and elements 134, 136, 140, and 144 are not shown here, but they can be obtained from Figure 6. In this case, there are two cyclic steps 179a (as shown in Figure 6). Figure 3a (as shown) and 179b (as shown) Figure 3b (As shown). Figure 3a This shows an example of the first step performed by encoder 100, while Figure 3b This shows an example of the second step performed after the first step. Figure 3a As shown, the prediction residual signal 110 (target signal), mathematically represented as r(n), is downsampled at block 310a, thereby transferring from the first sampling (e.g., 80 sampling positions per frame or subframe) to the second sampling (e.g., 64 sampling positions within the same frame or subframe). The downsampling at block 310a thus allows for the acquisition of a downsampled version 310c of the prediction residual signal 110 (mathematically represented as r_2(n)). Here, the candidate excitation signal is represented as 350 and compared with the downsampled version 310c of the prediction residual signal 110. Similar to the first alternative in Figure 6, the candidate excitation signal 350 is obtained by adder 348 from the downsampled version 146' of the excitation prediction component 146 and the candidate innovation component 358a obtained from the innovation codebook 118 (the innovation codebook 118 can be the same as in the first alternative, so we use the same reference numerals). For example, the excitation prediction component 146 can be downsampled at downsampler 346. It should be noted that in the first step ( Figure 3a All components of loop 179a used in the second sampling are in the second sampling.

[0067] It can be seen that the downsampled version 310c of the predicted residual signal 110 can be processed by the weighted filter W_2(z) at block 364a. The input to block 364a can be the result 362a (error) of the comparison between the candidate excitation signal 350 and the downsampled version of the innovation component 358a (at 360a).

[0068] The filtered signal 366a is provided to the analysis synthesis block 355, which is instantiated here as instance 355a (first step instance). Block 355 (instance 355a) defines multiple indices 319' that are cyclically input into the innovation codebook 118 (in the second, lower sample). The innovation codebook 118 cyclically outputs candidate pulse combinations 118' based on its inputs 319'. Optionally, the candidate pulse combinations 118' can be filtered by a filter or a series of filters 318, mathematically represented as S_2(z), to obtain a filtered version 318' of the candidate pulse combinations 118'. The filter or series of filters can be associated with specific frequency shaping of the candidate pulse combinations, such as formant sharpening and / or pitch sharpening. Then, in the first step loop 179a, the filtered version 318' of the candidate pulse combinations 118' is provided to the scaler 352. The output 358a of the scaler 352, operating at the second sampling rate (fs_2), which is a candidate innovative component of the candidate stimulus 350, can be input to the adder 348 and added to the predicted component 146' of the candidate stimulus 350. In this case, the gain used at the scaler 352 is a predefined optimal gain, which is not changed during the iteration of the first loop 179a by the analysis synthesis block 355 (instance 355a). Therefore, the optimal pulse combination 118 that minimizes the error 362a is found (among all combinations 118' or 318'). Thus, the encoding / decoding information 119 regarding the selected pulse combination 118 (informing the pulse combination 118 that allows the best approximation of the predicted residual signal 110 350 to be obtained in the second sampling) can be provided to the encoding / decoding signal writer 122.

[0069] Once the first step is completed (i.e., when the selected pulse combination that minimizes error 362a or 366a is retrieved), the second step can be triggered. Figure 3b Here, the second loop 179b is iterated, where processing is performed at the first sample (e.g., the higher sample associated with the first, higher sampling rate fs_1). It can be seen that the innovation codebook 118 now provides the selected pulse combination 118 (e.g., which can be filtered at filter block 318, mathematically represented as S_2(z), and provides the filtered selected pulse combination 318'). At this time, the selected pulse combination 118 (whether version 118' or version 318) can be upsampled at upsampling block 318b. Therefore, the upsampled version 318b of the selected pulse combination 118 (or its filtered version 318') can be provided to scaler 352. Scaler 352 can provide candidate innovation components 358b of the candidate excitation signal 350. To obtain the candidate excitation 350, at adder 348, the predicted version 146 of the candidate excitation signal 350 is provided at the first sample. Even in this case, through the second step of loop 179b ( Figure 3bAlthough the preferred pulse combination 118””318” has been obtained, only the optimal gain of the innovative component 358b to be provided to the candidate excitation signal 350 is searched. Here, reference numeral 354b indicates the gain control applied to the scaler 352 so as to retrieve (from the candidate gains indicated by control 354b) the gain that allows for the optimal approximation of the predicted residual signal 110. At comparison block 360b, the predicted residual signal 110 (target signal) is compared with the candidate excitation signal 350, and the error 362b This can be evaluated by the analysis synthesis block 355 (as shown in its second instance 355b) (here, a weighted version 366b of the weighted filter 364b is shown to provide the analysis synthesis block 355 with an error 362b). Similarly, even though not shown, the analysis synthesis optimization block 350 (as shown in its second instance 355b) can provide the gain 156 and pitch hysteresis of the predicted component 146 of the candidate excitation signal 350. Therefore, other information 117 can be provided to the codec signal writer 122 (e.g., gain information, pitch hysteresis information, etc.).

[0070] It should be noted that the predicted residual signal (110) can be downsampled in the time domain (at 310a). This is advantageous because the predicted residual signal (110) is already in the time domain. For example, linear phase filtering can be performed on block 310a.

[0071] In an alternative, at block 310a, the predicted residual signal 110 can first be transformed to the frequency domain (e.g., using a time-frequency block transform such as Short-Time Fourier Transform (STFT), Fast Fourier Transform (FFT), Discrete Cosine Transform (DCT), or a similar linear transform), the frequency domain version of the predicted residual signal 110 can be downsampled in the frequency domain (e.g., using a block transform with no overlap between adjacent blocks and / or using a block transform with spectral truncation, or more specifically using spectral truncation or constant scaling), and then the downsampled frequency domain version of the predicted residual signal 110 can be transformed back to the time domain (e.g., using an inverse time-frequency block transform such as Inverse Short-Time Fourier Transform (ISTFT), Inverse Fast Fourier Transform (IFFT), Inverse Discrete Cosine Transform (IDCT), or a similar inverse linear transform).

[0072] Generally, a linear phase filter can be used in the time domain (e.g., in...). Figure 3a (in) downsamples the predicted residual signal (110) or its processed version (e.g., at 310a), and / or upsamples a selected pulse combination (e.g., 118") or its processed version (e.g., 318") (e.g., at 310a). Figure 3b (At point 318b).

[0073] Yes (in) Figure 3a (in) will predict the residual signal (e.g., 110) or its processed version, or (in) Figure 3bThe selected pulse combination (e.g., 318') or its processed version (e.g., 318") is transformed to the frequency domain, and the predicted residual signal (e.g., 110) or its processed version is downsampled in the frequency domain (e.g., at 310a), and / or (e.g., at...) Figure 3b In the frequency domain, the selected pulse combination (e.g., 118) or its processed version (e.g., 318) is upsampled (e.g., in...). Figure 3b (At point 318b).

[0074] Time-frequency block transforms (such as Short-Time Fourier Transform (STFT), Fast Fourier Transform (FFT), Discrete Cosine Transform (DCT), or similar linear transforms) can be used in the frequency domain (e.g., in...). Figure 3a (in) downsample (310a) the predicted residual signal (110) or its processed version, and / or (e.g. in) Figure 3b (in) upsampling of selected pulse combinations (e.g., 118) or their processed versions (e.g., 318) (e.g., in) Figure 3b (At point 318b).

[0075] Block transforms with no overlap between adjacent blocks can be used in the frequency domain (e.g., in...). Figure 3a (in) downsample (310a) the predicted residual signal (110) or its processed version, and / or (e.g. in) Figure 3b (in) upsampling of selected pulse combinations (e.g., 118) or their processed versions (e.g., 318) (e.g., in) Figure 3b (At point 318b).

[0076] Block transforms that employ zero-padding of the spectrum can be used in the frequency domain (e.g., in...). Figure 3a (in) downsample (310a) the predicted residual signal (110) or its processed version, and / or (e.g. in) Figure 3b (in) upsampling of selected pulse combinations (e.g., 118) or their processed versions (e.g., 318) (e.g., in) Figure 3b (At point 318b).

[0077] Constant scaling can be used in the frequency domain (e.g., in the frequency domain). Figure 3a (in) downsample (310a) the predicted residual signal (110) or its processed version, and / or (e.g. in) Figure 3b (in) upsampling of selected pulse combinations (e.g., 118) or their processed versions (e.g., 318) (e.g., in) Figure 3b (at 318b) to upsample the selected pulse combination or its processed version in the frequency domain.

[0078] It should be noted that, although in Figure 2, Figure 5In the first alternative to Figure 6, the second sampling is obtained by subtracting one or more interleaved tracks from the first plurality of samples, but Figure 3a and Figure 3b The second alternative is to predict the actual downsampling of the residual signal 110 from the first sample to the second sample to search for the optimal pulse combination, but the gain and pitch hysteresis can be searched at the first sample. However, it can be seen that the codebook 118 remains at the second sampling rate.

[0079] It is important to note that, in Figure 3b In this process, the sequence of blocks 118, 318, and 318b can be skipped: for example, a second codebook (not shown) can be used to convert each pulse combination 118' to an upsampled pulse 318b", and the upsampled pulse 318" can be input to scaler 352, or filtering at 318 and upsampling at 318b can be performed actively.

[0080] A less promising alternative could be to avoid Figure 3b The second step, but (in Figure 3a (in the middle) directly upsample the candidate pulse combination 118' upstream of the scaler 352, and in Figure 3a In the same loop 179a, the gain 354b to be applied to the scaler 352 is found simultaneously in the first sample. While this solution is feasible, it is less preferred because... Figure 3a and Figure 3b The two-step technique greatly reduces complexity.

[0081] In summary, encoder 100 (400) according to the first or second alternative scheme can encode the audio signal as follows:

[0082] • Encoding and decoding information 119 regarding a selected pulse combination (e.g., 118) enables the decoder to reconstruct the selected pulse combination from the encoding and decoding information 119; and / or

[0083] • Other information 117, such as including at least one of the following:

[0084] Regarding the gain information for gain 154, this gain will be applied to the selected pulse combination (e.g., 118") once the decoder has reconstructed the selected pulse combination.

[0085] Information regarding the gain 156 to be applied to the excitation prediction components (142, 146) (the excitation prediction components will be estimated, for example, using an adaptive codebook); and

[0086] The pitch lag information for pitch lag 157 enables the decoder to perform LTP synthesis.

[0087] It can be seen from the two alternatives mentioned above that:

[0088] • First alternative ( Figure 5 As shown in Figure 6), the pulse combinations are preferably (but not strictly required) searched iteratively in the following manner: In each iteration:

[0089] Each candidate pulse combination is generated using the second sampling, 118'.

[0090] Subsequently, in the same iteration, candidate pulse combination 118' is converted to version 118a' in the first sample;

[0091] In the same iteration, but in the first sample, candidate gains 154 and 156 and candidate pitch lag 157 are used;

[0092] Evaluate the error 162 (166) of the candidate excitation 150 in the first sample of this iteration;

[0093] Repeat the new iteration, changing the candidate pulse combination 118' in the first sample and the different candidate gains 154, 156 and candidate pitch hysteresis 157 in the second sample, until the best approximate excitation 150 is obtained;

[0094] Information 119 regarding the optimal candidate pulse combination is encoded as the selected pulse combination, and other information 117 (including optimal candidate gains 154, 156 and optimal pitch lag 157) is encoded.

[0095] • Second alternative ( Figure 3a and Figure 3b In this process, the preferred (but not strictly required) method is to iteratively search for pulse combinations in the following manner:

[0096] Execute the first loop ( Figure 3a The first step in the process involves identifying, along multiple iterations 179a (using predefined fixed values ​​for gain and pitch hysteresis), the optimal candidate pulse combination 118” that allows the generation of the optimal approximate candidate excitation signal 350 in the second sample;

[0097] Execute the second loop ( Figure 3b In the second step, along multiple iterations 179b (using predefined fixed values ​​for gain and pitch hysteresis), an upsampled version 318b of the best candidate pulse combination 118” (in the first sample) is used, while searching for different gains 354b, 156 and pitch hysteresis 157 to find those candidate excitation signals 350 that allow for the best approximation of the predicted residual signal 110 in the first sample.

[0098] Information 119 regarding the optimal candidate pulse combination 118” is encoded as the selected pulse combination, and other information 117 (including optimal candidate gain 154, 156 and optimal pitch lag 157) is encoded.

[0099] Figure 7 An example of a device 700 (decoder) for generating a decoded audio signal 702 from a codec signal 124 (e.g., a bitstream) is shown, for example, according to CELP (especially ACELP, e.g., a CELP with an algebraic and innovative codebook). Device 700 can be a CELP decoder, which is, for example, an ACELP. The codec signal 124 is used with... Figure 4 The encoded and decoded signal 124 is represented by the same digital representation, because it is assumed that decoder 700 decodes by encoder 400 (in Figure 3a , Figure 3b , Figure 5 The encoded / decoded signal 124 is generated in (and in any alternative to Figure 6). Nevertheless, it is not strictly required that the encoded / decoded signal 124 be generated by the encoder and decoder 700. The decoder 700 generates the output audio signal 702, making it the most realistic audio representation of the input audio signal 102 as possible.

[0100] Decoder 700 may include a codec signal reader 722 capable of reading the codec signal 124. The codec signal reader 722 may include, for example, an entropy decoder (but this is not strictly required). The codec signal reader can provide codec information 105' about the prediction coefficients. The codec information 105' about the prediction coefficients can be provided to prediction coefficient decoder 705. Prediction coefficient decoder 705 can obtain prediction coefficients 704 from the codec information 105' about the prediction coefficients. Prediction coefficients 704 can be provided to signal processor 703 to generate output audio signal 702.

[0101] The codec signal reader 722 can read the encoded pulse combination 119 from the codec signal 124 (which may be the same as the codec information 119 about the selected pulse combination generated by the pulse information encoder 116 on the encoders 100 and 400, and therefore uses the same reference numerals). The codec signal reader 722 can also read other information 117. The other information 117 may include, for example, other gain information (e.g., the other information 117 may include, for example, information obtained by operating Figure 6 and...). Figure 3b The gain information (154, 156) and / or pitch hysteresis information obtained by the technique (e.g., can also be seen in Figure 6 and Figure 7) Figure 3b The pitch lag information 157 is obtained. The encoded / decoded pulse combination 119 can be encoded / decoded information about at least one pulse (but more commonly about multiple pulses). The encoded / decoded pulse combination 119 can take the form of a codebook index (e.g., Figure 6 and...). Figure 3a The selected codebook index that minimizes the error. The encoded / decoded pulse combination 119 (and optionally, other information 117) can be provided to the pulse information decoder 716. The pulse information decoder 716 can provide a prediction residual signal 710. The pulse information decoder can provide the prediction signal 710 by using an innovation codebook 118. The innovation codebook 118 can be the same as the innovation codebook 118 of device 100 or 400. In the example, the encoded / decoded information 119 about at least one pulse can be an entry in codebook 118, such that codebook 118 provides at least one pulse (or encoded / decoded pulse combination) 118' pre-associated with that entry. Therefore, the pulse information decoder 716 can generate the prediction residual signal 710 from the encoded / decoded information 119 about at least one pulse, for example, by using other information 117 (gain information, pitch hysteresis information, etc.). The prediction residual signal can be provided to the signal processor 703. The signal processor 703 can generate an output audio signal 702 based on the prediction coefficients 704 and the prediction residual signal 710, for example by using LP synthesis technology.

[0102] For example, for Figure 4 The signal to be represented (i.e., from its codec version 124 to audio signal 702), as interpreted by encoder 400, can be subdivided into frames and / or subframes. As described above, each frame or subframe can be further subdivided according to a first sample and a second sample. According to the first sample, a frame or subframe is subdivided into a plurality of adjacent sample positions (the number of which is a first number). According to the second sample, the same frame or subframe is divided according to a second plurality of adjacent sample positions (time slots), the number of which is a second number. The first sample and the second sample described herein are exactly the same as those of the encoder, as they are the same concept. The first sample differs from the second sample (e.g., for the same frame or subframe, the first sample may have more sample positions than the second sample). As described above for encoder 400 or 100, the codec information 119 for a selected combination of pulses is defined using the second sample. This is maintained in the decoding information 710 for at least one pulse under the second sample. Therefore, the pulse information decoder 716 can provide the prediction residual signal 710 under the first sample. It should be noted that, according to many examples, the first sample can be understood as indicating a first sampling rate (e.g., 16 kHz), which is higher than the second sampling rate (the sampling rate of the second sample) (e.g., 12.8 kHz), or in any case, for the same frame or subframe, there are more sampling positions in the first sample than in the second sample. Therefore, the output audio signal 702 (more generally, the prediction coefficients 704 and the prediction residual signal 710) is at the first sample, although the encoding / decoding information 119 regarding at least one pulse is at the first sample.

[0103] As described above, the innovative codebook 118 may include (and output) a set of pulse combinations 710 defined at a second sampling rate, for example at a second sampling rate lower than the first sampling rate used by the rendered output audio signal 702.

[0104] Figure 8 An example of how the prediction coefficients 704 and prediction residual signal 710 obtained from codebook 118 can be processed is shown. For example, this can be performed partly in signal processor 703 and / or partly in pulse information decoder 716. Figure 8 As shown, the output audio signal 702 is obtained from an LTP synthesis filter 830. The LTP synthesis filter 830 can be defined by, for example, prediction coefficients 704 decoded by a prediction coefficient decoder 705. The LTP synthesis filter 830 can be excited by an excitation signal 810. The excitation signal 810 can be obtained at adder block 848 as the sum between a prediction component 856 and an innovation component 858. The innovation component 858 can be obtained from the prediction residual signal 710 (a combination of pulses from codebook 118). The prediction residual signal 710 can be scaled, for example, by a gain (e.g., from gain information 154 in other information 117 written into the codec signal 124). Therefore, the result of scaling at 852 can be the innovation component 858 of the excitation signal 810. The prediction component 846 can be obtained from an LTP synthesis 834, which can include, for example, an adaptive codebook (not shown). The output of the LTP synthesis 834 can be signal 842. LTP synthesis 834 can use pitch hysteresis as usual (e.g., obtained from pitch hysteresis information 157 encoded in additional information 117 written to codec signal 124). Signal 842 can be scaled at scaler 844 by gain (e.g., obtained from gain information 156 written to additional information 117 written to codec signal 124). Thus, the innovative component 846 of the excitation is obtained. Therefore, excitation 810 can be obtained at adder 848 as the sum between components 846 and 858.

[0105] The prediction residual signal output from the innovation codebook 118 can be converted from its second sampled version 710 to its first sampled version 710' using a mapper or resampler (e.g., an upsampler) 818. When using the first alternative (e.g., corresponding to...) Figure 5 In the case of Figure 6), block 818 is a mapper that inserts an empty track interleaved with other tracks (e.g., transferring frames or subframes from a second number of sampling positions (e.g., 64) to a first number of sampling positions (e.g., 80)). When using the first alternative (e.g., corresponding to...), Figure 3a and Figure 3bIn the case of [the above], block 818 can be a resampler (upsampler), which can be similar to upsampler 318b (in any of its embodiments). In any case, the prediction residual signal 710' obtained from the mapper or upsampler 818 is under the first sample (in some examples, similar to other signals 858, 846, 842, 810, and 702), while version 710 of the pulse combination upstream of the mapper or upsampler 818 is under the second sample. It should be understood that the mapper or upsampler 818 can be indistinguishably part of the signal processor 703 or the pulse information decoder 716, and in some examples, so are elements 834, 844, 852, and 848.

[0106] In the second alternative, at decoder 700, resampler 818 (upsampler) can upsample the decoded pulse combination or its processed version to obtain an upsampled decoded pulse combination under the first sample; this can be used to update the adaptive codebook. Figure 8 (Not shown, but can be inserted between the scaling at upsamplers 818 and 852). Therefore, the adaptive codebook can be used under the first sample.

[0107] In any case, upsampler 818 can upsample the decoded pulse combination 710 or its processed version from the second sample to the first sample in the time domain. Alternatively, the decoded pulse combination 710 or its processed version can be upsampled from the second sample to the first sample in the frequency domain (e.g., after a conversion from the time domain to the frequency domain, and in some examples, the frequency-domain upsampled version is subsequently converted back to the time domain, for example). In the latter case, upsampler 818 can upsample the decoded pulse combination 710 or its processed version (from the second sample to the first sample) in the frequency domain using a block transform with no overlap between adjacent blocks or using a block transform and zero-padding of the spectrum.

[0108] Figure 9 An example of encoder 900 is shown, which may include the functionality of encoder 400. Signal processor 103, prediction coefficient 104, and prediction coefficient encoder 105 are not shown here because attention is focused on the evolution of the prediction residual signal from its uncompressed version 110 to its encoded / decoded version 119 or 919. A first operating mode and a second operating mode can be selected (e.g., via selector 920). Pulse information encoder 916 can be selectively instantiated as a first pulse information encoder instance 116b (when the first operating mode is selected) and a second pulse information encoder instance 116a (when the second operating mode is selected). The second operating mode is as follows: Figure 4 , Figure 3a , Figure 3b , Figure 5And any of the operations in Figure 6, therefore their operations will not be described again. In the second operation mode, Figure 4 The role of the pulse information encoder 116 is assumed by the second pulse information encoder instance 116a, which uses Figure 4 The same codebook 118 is operated on under the second sample. Outputs 119 and 117 are the same as... Figure 4 The corresponding outputs 119 and 117 are the same. However, the second operating mode (and the second pulse information encoder instance 116a) is selectively deactivated. In fact, the first pulse information encoder instance 116b, which allows processing according to the first operating mode, can be selected via selector 920. In the second mode, there is no resampling, nor is there a conversion from the first sample to the second sample or vice versa. In short, in the first operating mode, a codebook 918 (e.g., an innovation codebook and / or algebraic codebook) is used, which allows for the provision of encoding / decoding information about selected pulse combinations, but under the first sample. Therefore, in the first operating mode, the encoding / decoding information 919 (encoding / decoding pulse position information) about selected pulse combinations is not like... Figure 4 Obtained that way, but like Figure 1 That is, (e.g., according to conventional CELP). Generally, the length of the codec information 919 can be greater than the length of the codec information 119 in the second operating mode. Therefore, the codec information 919 obtained in the first operating mode can generally be understood to have better quality, although more bits are required for encoding. On the other hand, the quality of the codec information 119 according to the second operating mode is slightly reduced, but fewer bits are required. The selection by the selector 920 can be based on information about the target block size or the instantaneous bit rate (in Figure 9 (Represented by the number 921 in the text). It is worth noting that the sampling of the prediction residual signal 110 remains unchanged (in the first sampling), but in the second operating mode, the codec information 119 is provided in the second sampling (saving bits). In the first operating mode, the codec information 919 is in the second sampling, improving the quality.

[0109] The operation of the first pulse information encoder example 116b (operating under the first sample) is as follows: Figure 1 As shown, this can therefore be done according to the CELP encoder. By comparison... Figure 1 As can be seen from Figure 6 (which represents the first instance 116b) and Figure 6 (which also represents the second instance 116a), in Figure 1 There is no mapper 118a, and the optimized analysis synthesis 955 does not provide simplified format entries to the innovation codebook 918 (e.g., no empty tracks, but all tracks are used). The innovation codebook 918, in the first sample, and Figure 1The part not based on the second sampling is not included. Of course, even in the second operating mode, based on the second alternative (e.g., meaning downsampling, as...),... Figure 3a and Figure 3b In the case shown), the first pulse information encoder instance 116b still works with... Figure 1 same.

[0110] It can be seen that, Figure 1 It basically has the same elements as Figure 6, but with some exceptions: Figure 1 The analysis synthesis optimization block 955 does not provide candidate index 119' under the second sampling, but provides candidate index 919' under the first sampling (i.e., the same as input signal 102 and signals 110 and 166). Figure 1 The innovative codebook 918 is in the first sampling (instead of the innovative codebook 118 in Figure 6 in the second sampling); Figure 1 The candidate pulse combination 918' is obtained under the first sample (instead of under the second sample as in the similar sample combination 118' in Figure 6). The gain 954 applied to the candidate sample combination 918' is obtained using the first sample, and the gain 956 to be applied to the prediction signal 142 is also obtained using the first sample, as is the pitch hysteresis 957. Furthermore, and taking into account the absence of mapper 118a, Figure 1 The operation is the same as that in Figure 6.

[0111] Basically, when operating in the first operating mode (using the first pulse information encoder instance 116b), no track is empty, and Figure 5 Track 5 was also taken into account in the innovative codebook 918. Therefore, in the first operating mode, it is possible to encode more pulses, thus achieving better quality. Nevertheless, the choice between the first and second operating modes at 920 allows for better adaptation to the target block size and / or instantaneous bit rate.

[0112] It should also be noted that the choice between the first and second operating modes at 920 reduces the negative impact of transients, and the predicted residual signal 110 is always provided with the same first sample.

[0113] Figure 10 illustrates decoder 1000, which corresponds to encoder 900 in Figure 9. Pulse information decoder 1016 includes a first pulse information decoder instance 717b (selectable when a first operating mode is selected) and a second pulse information decoder instance 716a (selectable when a second operating mode is selected). Decoder 1000 can operate completely as follows in the second operating mode: Figure 7The decoder 700 operates as described above, but in its first operating mode it can operate without using the second sample at all, and in some examples it can be the same as or similar to a conventional CELP decoder. Here, the codec signal 124 is read by the codec signal reader 722, and at 1020 it can be selected between the first and second operating modes. In the second operating mode, the second pulse information decoder instance 716a is provided with codec information 119 and 117 (while the first pulse information decoder instance 716b is deactivated), thereby obtaining the prediction residual signals 710, 710' (e.g., as shown in FIG8) by using an innovative codebook under the second sample. However, if the encoded input signal 124 contains codec information 909 regarding a selected pulse combination using the first sample, the first operating mode is activated, and the first pulse information decoder instance 716b is activated (while the second pulse information decoder instance 716a is deactivated). In the first operating mode, the codebook 918 under the first sample is used, similar to a conventional CELP decoder. The prediction residual signal provided by the first pulse information decoder instance 716b is labeled 1010. In any case, the prediction residual signal 710 (710') and 1010 are provided to the signal processor 703 to obtain the output audio signal 702. Basically, the first pulse information decoder instance 716b can operate almost identically to when operating in the second operating mode: except that the second sample codebook 118 is replaced by the first sample codebook 918 and there is no mapper or resampler 818, the operation can be the same as when operating in the second operating mode. Figure 8 The operations are the same as those in [the previous section].

[0114] Discussion of this technology

[0115] ■First Alternative Solution / Aspect

[0116] The first alternative to this technology (e.g., Figure 5 As shown in Figure 6), the problem of encoding the optimal number of pulse positions is addressed by introducing interleaved positions not defined in codebook 118. These interleaved positions, organized as empty tracks, are not defined in codebook 118, allowing for greater freedom in customizing codebook 118 to have a size that is a power of 2 and / or the desired number of pulses. In other words, possible positions in a frame or subframe are excluded from codebook 118, such that the number of remaining positions is a power of 2. The codebook 118 thus defined and the associated code vector (e.g., 119') are then mapped to the samples used in encoder 400 by inserting one or more empty tracks corresponding to positions not defined in codebook 118. The constrained codebook 118 is used to locate pulses during pulse search, while the mapped codebook / code vector is used to evaluate performance during optimization (e.g., iterations along loop 179 in Figure 6), thus taking into account the additional constraints in codebook construction.

[0117] For 80 sampled subframes, an example of the potential location of a single pulse in an 8-pulse algebraic book with 5 tracks and 16 locations:

[0118] track pulse Location 1 0,4 0, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75 2 1,5 1, 6, 11, 12, 21, 26, 31, 36, 41, 46, 51, 56, 61, 66, 71, 76 3 2,6 2, 7, 12, 17, 22, 27, 32, 37, 42, 47, 52, 57, 62, 67, 72, 77 4 3,7 3, 8, 13, 18, 23, 28, 33, 38, 43, 48, 53, 58, 63, 68, 73, 78 5 No pulse 4, 9, 14, 19, 24, 29, 34, 39, 44, 49, 54, 59, 64, 69, 74, 79

[0119] In the example above, with a budget of 8 pulses per subframe, the pulse distribution is uneven, resulting in the dropping of one track. This is one way to resample possible positions by discarding every 5th position.

[0120] ■Second alternative / aspect ( Figure 3a and Figure 3b )

[0121] Another approach to maintaining small innovations in the codebook 118 at high sampling rates is to decimate (or more generally, reduce, e.g., downsample) the reachable positions of the codebook (as done in the first aspect) and resample (upsample) them using conventional signal processing resampling techniques. In this sense, the number of positions defined by the codebook 118 can be reduced, and the number of pulses can be increased given a bit budget. Similar to the first aspect of this technique, pulses can be positioned in the reduced positions defined by the codebook 118, and optimization can be performed after resampling (upsampling) the code vectors to be evaluated. However, this can incur high complexity overhead because resampling is costly, especially if resampling is performed for every candidate code vector to be evaluated. Alternatively, and in a preferred embodiment, the optimization process, or a portion thereof, involves resampling (downsampling) the target signal 110 and the impulse response required for optimization at the sampling rate of the codebook 118. Figure 3a To perform (e.g., in 310a) Figure 3a The first step). Then, only the optimal code vector 118” obtained in this way is resampled (upsampled). Figure 3b Samples from 318b in the codec are used for subsequent processing (e.g., Figure 3b (The second step).

[0122] For example, resampling at 318b and / or 310a can be accomplished using a linear filter with low-pass characteristics. However, a disadvantage of linear filtering is that it introduces a delay. A disadvantage of delay-free linear filters (such as IIR filters) is that they have a non-linear phase, which is problematic for pulse-like signals. In a preferred embodiment, frequency-domain resampling involving circular convolution is used. Codebook 118 can preferably be resampled in the frequency domain at 318b, with or without some zero padding, to reduce the number of possible locations of the pulse to be located.

[0123] Example:

[0124] Encoder:

[0125] Target signal -> FFT 80 samples -> Truncation [scaling] -> iFFT 64 samples -> Search in an innovative codebook of 4 tracks, 16 samples per track

[0126] Decoder:

[0127] Innovative code vector with 4 tracks, 16 samples per track -> FFT 64 samples -> Resampling with zero addition and / or spectral copying (e.g., copy / mirror [scaling]) -> FFT 80 samples

[0128] The proposed technique combines the relatively small number of pulses in the innovative codebook 118 with the relatively high sampling rate of the speech codec (CELP).

[0129] Key aspects include: a speech codec that uses linear prediction to operate at a first sampling rate, wherein at least one prediction residual is encoded by locating the following:

[0130] Given a pulse at a possible location,

[0131] The possible locations are subsets or resampled versions of the possible locations at the first sampling rate.

[0132] - A subset of positions at the first sampling rate can be obtained by skipping the regular positions at the first sampling rate.

[0133] - Possible locations are obtained by resampling the codebook vector from the first sampling rate to the given sampling rate.

[0134] The main non-limiting examples are primarily about enhanced CELP (Code-Excited Linear Prediction), an efficient speech codec scheme for efficiently compressing and transmitting speech signals while maintaining fairly high quality.

[0135] The raw speech 102 given as input can be represented as a combination of linear prediction and excitation modeling, i.e., linear prediction residual encoding / decoding. The speech signal can be divided into short frames and / or subframes, for example, ranging from 5 to 20 milliseconds. Within each frame or subframe, CELP performs analysis and encoding to extract the parameters required for synthesis at the receiver end. The CELP encoding process includes the following steps:

[0136] Preprocessing

[0137] The input speech signal is divided into frames and / or subframes and ultimately resampled and high-pass filtered to remove DC bias. The signal may also be pre-emphasized at high frequencies to compensate for the natural negative frequency tilt (i.e., there is currently much more energy in the low frequencies than in the high frequencies), which prevents accurate analysis of high-frequency content via linear prediction. The transition is then smoothed by keeping the filter memory updated, allowing each frame or subframe to be processed individually. Some analyses, such as short-term prediction analysis, i.e., linear prediction codec analysis (LPC analysis), use windowing to obtain better and more accurate analysis.

[0138] LP (C) Analysis (Short-Term) and Quantification

[0139] Using an autocorrelation method with an analysis window of approximately 30 ms, one or two LPC analyses and quantizations are performed per frame or subframe. Depending on the latency constraints, the window can be symmetric or asymmetric and has a look-ahead time relative to the frame or subframe currently being considered for encoding / decoding. Typically, a look-ahead time of 5 ms and 8.75 ms is acceptable for voice communication. The Levinson-Durbin recursion can compute the optimal LPC prediction parameters based on the calculated autocorrelation function. The resulting prediction coefficients can then be efficiently encoded / decoded by vector quantization of the corresponding linear spectral frequencies.

[0140] LPC (Analysis) Filter

[0141] Short-term prediction is performed using an LPC analysis filter 1 / A(z)(132) with quantization coefficients also available at the decoder. The residuals are then modeled using LTP and a fixed codebook 118.

[0142] LTP analysis (134)

[0143] Long-term prediction primarily relies on pitch lag estimation, which directly corresponds to the fundamental frequency f0 or principal periodicity of signal 102. It can be estimated using the autocorrelation function, taking into account a longer order / lag than LPC at a given sampling rate to cover the expected pitch range. This will be used for long-term prediction in subsequent processing.

[0144] Long-term forecast

[0145] Long-Term Prediction (LTP) can be used via an adaptive codebook 140. The adaptive codebook 140 can contain excitation vectors from past encoding and decoding for each frame or subframe. The adaptive codebook 140 can be derived from the long-term prediction parameter 135 and the pitch lag 157, which can be considered as an index of the adaptive codebook 140. The LTP can then be applied in reverse, i.e., for synchronization between the encoder and decoder sides.

[0146] Innovative Encoding and Decoding – Pulse Search in a Fixed Codebook

[0147] The residuals of the two predicted LPCs and LTPs can then be modeled using a fixed (innovative) codebook 118. The fixed codebook 118 may contain non-adaptive excitation vectors (i.e., fixed) used to model the predicted residuals. The selected code vectors, also known as innovators, will excite the LPC synthesis filter 830 on the decoder side, in addition to the adaptive codebook contribution 142. For complexity and memory reasons, the fixed codebook 118 typically contains algebraic codes, which may contain a small number of non-zero pulses with a predefined set of interleaved potential positions (tracks). The amplitude and position of the pulses in the code vectors can be derived from their indices using algebraic rules, requiring little or no memory storage, unlike the lookup tables used in classical random vector quantization. It is this fixed codebook 118 that is the main subject of this technique. In practice, the fixed codebook 118 is typically designed for the sampling rate of the signal received by the CELP codec. At very low bit rates, only a small number of pulses can be located if all positions at the input sampling rate are considered. Therefore, CELP designed for very low bit rates requires reducing its sampling rate, for example, modeling wideband signals using 12.8 kHz instead of 16 kHz. This has the negative impact of reducing the audio bandwidth for encoding and decoding or requiring the use of complementary band extension modules, which is suboptimal globally and structurally complex. The goal of this technique is to maintain a higher input sampling rate, such as 16 kHz, and to design a specific fixed codebook to achieve very low bit rates.

[0148] The codebook structure can be based on the positions of interleaved tracks (e.g., in...). Figure 5 (As in the example in Figure 6). For example, for a 5-millisecond subframe at 16 kHz, the 80 positions in the code vector are divided into 5 equally sized, staggered tracks, each with 16 positions. Different codebooks with different rates (e.g., 118 and 918) can be constructed by placing a certain number of signed pulses in the tracks, with each track having 1 to a maximum of 8 or 10 pulses depending on the available bit budget.

[0149] To achieve even lower bit rates, current technologies propose two solutions:

[0150] - Relax (e.g.) Figure 5 (As shown in Figure 6) Each track is constrained to have at least one pulse, and tracks are allowed to be without pulses. Omitting a track means decimating (or otherwise reducing) the locations reachable from the codebook, and allows encoding and decoding to be more efficient because the number of possible locations is reduced, i.e., the number of bits required. In this way, a greater number of pulses can be retained compared to conventional methods, at the cost of reduced flexibility in pulse positioning, which is often a better trade-off at very low bit rates.

[0151] - Second Solution ( Figure 3a and Figure 3b This includes performing appropriate downsampling only on the fixed codebook contribution, such as involving low-pass filtering. This can be achieved using time filters (such as linear-phase FIR filters) or other linear interpolations. To reduce potential delays, there are preferred techniques such as using resampling in the frequency domain, using rectangular windows, truncation (for downsampling), and zero-padding (for upsampling) in the frequency domain, which correspond to resampling using circular convolutions. The use of circular convolutions associated with rectangular windows is also advantageous, especially in the LPC and LTP residual domains where the signal is highly whitened.

[0152] On the decoder side, the CELP decoding process reverses the encoding step to reconstruct the speech signal 102 as the output signal 702. The decoder 700 can use the received bitstream 124 to synthesize the speech signal 702 by applying inverse linear prediction, reconstructing the excitation signal 810, and finally combining them to obtain the reconstructed speech 702.

[0153] Remapping of fixed codebook 118 (e.g., Figure 5 (and Figure 6)

[0154] More details are needed here.

[0155] The pulse information for encoding and decoding and codebook 118 are defined under the second sample, wherein the possible positions of the pulses are shared with the first positions of the samples defined under the first sample used by the rest of the codec. A mapping (e.g., mapper 118a of Figure 6) is then used to insert empty tracks into the pulses for encoding and decoding. Figure 5 (and Figure 6).

[0156] Resampling of a fixed codebook (e.g., Figure 3a and Figure 3b )

[0157] Figure 3a and Figure 3b As shown, CELP gain optimization ( Figure 3b The second step, loop 179b), can use the upsampled encoding / decoding pulse 318b under the first sample (fs_1), although the pulse search is done under the second sample (fs_2). Figure 3a The first step, loop 179a).

[0158] Figure 3a and Figure 3b The second aspect relates to this technology. Figure 3a: Optimization of pulse search using the optimal gain completed under the second sampling fs_2. Figure 3b: Gain quantization completed after obtaining the selected pulse combination, optionally using filter S_2(z) to shape the selected pulse combination, and upsampling the encoded and finally processed pulse combination from fs_2 to fs_1.

[0159] Notes / definitions that may apply to some examples

[0160] Sampling rate = Number of samples taken over a duration (e.g., within 1 second or within one frame / subframe).

[0161] Sampling = Sampling rate + Sampling location (location is part of the discretization attribute)

[0162] Further examples

[0163] Typically, an example can be implemented as a computer program product having program instructions that, when the computer program product is run on a computer, are operable to perform one of the methods. The program instructions may, for example, be stored on a machine-readable medium.

[0164] Other examples include a computer program stored on a machine-readable medium for performing one of the methods described herein. In other words, examples of methods are therefore computer programs with program instructions that, when run on a computer, are used to perform one of the methods described herein.

[0165] Therefore, another example of the method is a data carrier medium (or digital storage medium, or computer-readable medium) on which a computer program for performing one of the methods described herein is recorded. The data carrier medium, digital storage medium, or recording medium is a tangible and / or non-transitory signal, rather than an intangible and transient one.

[0166] Therefore, another example of a method is a data stream or signal sequence representing a computer program used to perform one of the methods described herein. The data stream or signal sequence can be transmitted, for example, via a data communication connection (e.g., via the Internet).

[0167] Another example includes a processing device, such as a computer or programmable logic device, for performing one of the methods described herein. Another example includes a computer equipped with a computer program for performing one of the methods described herein.

[0168] Another example includes an apparatus or system for transmitting (e.g., electronically or optically) a computer program used to perform one of the methods described herein to a receiver. The receiver may be, for example, a computer, mobile device, storage device, etc. The apparatus or system may include, for example, a file server for transmitting the computer program to the receiver.

[0169] In some examples, programmable logic devices (e.g., field-programmable gate arrays) may be used to perform some or all of the functions of the methods described herein. In some examples, a field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Typically, the methods can be executed by any suitable hardware device.

[0170] The examples described above are merely illustrative of the principles discussed herein. It should be understood that modifications and variations to the arrangements and details described herein will be readily apparent. Therefore, the intent is limited by the scope of the claims, and not by the specific details presented herein through the description and interpretation of the examples.

[0171] In the following description, even if they appear in different figures, the same or equivalent elements or elements with the same or equivalent functions are represented by the same or equivalent reference numerals.

[0172] References

[0173] [1]3GPP, ETSI TS (1)26.441, “EVS Codec: General Overview,” ver. 12,rel. 12, Oct. 2014.[2]3GPP, ETSI TS (1)26.445, “EVS Codec: Detailedalgorithmic description,” May 2022.

Claims

1. An apparatus for generating a decoded audio signal divided into multiple frames or subframes, comprising: The codec signal reader (722) is configured to read at least the codec information about the prediction coefficients (105') and the codec information about at least one pulse (119) from the codec signal (124). The signal processor (703) is configured to generate a decoded audio signal (702) from at least a decoded version (704) of the prediction coefficients and a combination of decoded pulses (710) or a processed version thereof, wherein the decoded audio signal (702) is generated with a first sample, the first sample meaning a first plurality of sample positions having a first number of sample positions in a frame or subframe; The device is configured to derive (716) the decoded pulse combination (710) from the encoding / decoding information (119) and the second sample codebook (118) regarding at least one pulse, wherein at least one second sample codebook (118) comprises: a set of pulse combinations defined under a second sample, the second sample meaning a second plurality of sample positions having a second number of sample positions in the frame or subframe, wherein the first sample differs from the second sample at least in that the first plurality of sample positions differ from the second plurality of sample positions.

2. The apparatus of claim 1 is configured to use the decoded pulse combination (710) or its processed version (710') to excite the synthesized filter (830) derived from the prediction coefficients (704).

3. The apparatus according to any one of the preceding claims, wherein the codec signal reader (703) is configured to read codec information (117) regarding long-term prediction, the long-term prediction being related to a prediction delay (157) and / or at least one long-term prediction gain (156), and wherein the apparatus is configured to generate the decoded audio signal (702) based on a long-term prediction (834) using the prediction delay (157) and / or the at least one long-term prediction gain (156).

4. The apparatus according to any one of the preceding claims, wherein the second number of sampling positions is different from the first number of sampling positions.

5. The apparatus according to any one of the preceding claims, wherein the second number of sampling positions is less than the first number of sampling positions.

6. The apparatus according to any one of the preceding claims is configured such that the first plurality of sampling positions and the second plurality of sampling positions within the frame or subframe are defined by a first plurality of tracks and a second plurality of tracks, the first plurality of tracks and the second plurality of tracks being regularly interleaved with each other, wherein the second plurality of sampling positions are defined by at least one fewer track than the first plurality of sampling positions.

7. The apparatus of claim 6, configured to process the decoded pulse combination (710) by inserting at least one empty track with zero-value samples, the zero-value samples being regularly interleaved in the second plurality of tracks, thereby obtaining a resampled decoded pulse combination under the first sample, the resampled decoded pulse combination being defined at the first plurality of sampling positions.

8. The apparatus of claim 6 or 7, wherein, apart from the at least one empty track, the first plurality of tracks at the first plurality of sampling positions and the second plurality of tracks at the second plurality of sampling positions have correspondingly identical sampling positions.

9. The apparatus according to any one of claims 6-8, configured to define the at least one empty track as having a number of sampling positions that are powers of 2 (e.g., 16).

10. The apparatus according to any one of claims 6-9, wherein the second number of sampling positions is a power of 2.

11. The apparatus according to any one of claims 6-10, wherein the second plurality of sampling positions are mapped to the first plurality of sampling positions by adding at least one empty track to the tracks of the second plurality of tracks.

12. The apparatus according to any one of the preceding claims is configured to resample the decoded pulse combination (710) or a processed version thereof from the second sample (818) to the first sample to obtain a resampled version (710') of the codec pulse combination (710) or a processed version thereof.

13. The apparatus according to any one of the preceding claims further includes a resampler (818) configured to resample the decoded pulse combination (710) from the second sample to the first sample for further processing of the decoding.

14. The apparatus according to any one of the preceding claims further includes a resampler (818) configured to perform upsampling on the decoded pulse combination or a processed version thereof.

15. The apparatus according to any one of the preceding claims, configured to operate between at least a first operating mode and a second operating mode, such that: In the second operating mode, the encoding / decoding pulse position information is determined under the second sampling; and In the first operating mode, the encoding / decoding pulse position information is determined under the first sampling, and the decoding pulse combination is an entry in at least one first sampling codebook, the at least one first sampling codebook containing: a set of pulse combinations defined under the first sampling.

16. The apparatus of claim 15, configured to select between the first operating mode and the second operating mode.

17. The apparatus according to any one of the preceding claims, wherein the second plurality of sampling locations is a proper subset of the first plurality of sampling locations.

18. The apparatus according to any one of the preceding claims is configured to select between at least a first operating mode and a second operating mode based on a target block size or the instantaneous bit rate of the current frame to be encoded.

19. The apparatus according to any one of the preceding claims, wherein the at least one second sampling codebook (118) is or includes an innovative codebook.

20. The apparatus according to any one of the preceding claims, wherein the at least one second sampling codebook (118) is or includes a scalar codebook.

21. The apparatus according to any one of the preceding claims is configured to receive a gain from the encoded / decoded signal and apply the gain to the decoded pulse combination.

22. An apparatus for encoding an audio signal divided into multiple frames or subframes, the apparatus comprising: The signal processor (103) is configured to determine the prediction coefficients (104) and the prediction residual signal (110) in the time domain under a first sampling, the first sampling meaning a first plurality of sampling positions having a first number of sampling positions in a frame or subframe; A pulse information encoder (116) is configured to determine encoding / decoding information (119) for a selected pulse combination representing the predicted residual signal (110), the selected pulse combination being an entry in at least one second sample codebook (118), wherein the at least one second sample codebook (118) comprises: a set of pulse combinations defined under a second sample, the second sample meaning a second plurality of sample positions having a second number of sample positions in the frame or subframe, wherein the first sample differs from the second sample at least in that the first plurality of sample positions differ from the second plurality of sample positions; and The codec signal writer (122) is configured to write at least codec information (105') about the prediction coefficients (104) and codec information about the selected pulse combination.

23. The apparatus of claim 22, wherein the second number of sampling positions is different from the first number of sampling positions.

24. The apparatus of claim 22 or 23, wherein the second number of sampling positions is less than the first number of sampling positions.

25. The apparatus according to any one of claims 22-24, wherein the second plurality of sampling locations is a proper subset of the first plurality of sampling locations.

26. The apparatus according to any one of claims 22-25, configured to define the first plurality of sampling positions and the second plurality of sampling positions within the frame or subframe as a first plurality of tracks and a second plurality of tracks, the first plurality of tracks and the second plurality of tracks being regularly interleaved with each other, wherein the second plurality of sampling positions are defined by at least one fewer track than the first plurality of sampling positions.

27. The apparatus according to any one of claims 22-26, configured to define the first plurality of sampling positions according to a first plurality of tracks that are regularly interleaved with each other, wherein at least one empty track in the first plurality of tracks is ignored by the at least one second sampling codebook (118), such that the second plurality of sampling positions are formed by sampling positions defined by sampling positions in the first plurality of sampling positions that are not in the at least one empty track.

28. The apparatus of claim 26 or 27, configured to define the first plurality of tracks at the first plurality of sampling positions and the second plurality of tracks at the second plurality of sampling positions as having correspondingly identical sampling positions in each track, except for the at least one empty track.

29. The apparatus according to any one of claims 26-28, configured to define the at least one empty track as having a number of sampling positions that are powers of 2.

30. The apparatus according to any one of claims 26-29, wherein the second number of sampling positions is a power of 2.

31. The apparatus according to any one of claims 26-30, wherein the second plurality of sampling positions are mapped (118a) to the first plurality of sampling positions by adding the at least one empty track from the first plurality of tracks to the second plurality of tracks.

32. The apparatus according to any one of claims 26-31, wherein the selected pulse combination defined under the second sample is mapped (118a) to the first sample by adding a zero-value sample at a sampling position defined in an empty track not defined in the second plurality of tracks in the first plurality of tracks.

33. The apparatus according to any one of claims 22-32, configured to downsample the predicted residual signal (110) or a processed version thereof from the first sample (363) to the second sample to obtain a downsampled version (310c) of the predicted residual signal (110) or a processed version thereof, such that the pulse combination is searched within the at least one second sample codebook (118) taking into account the downsampled version (310c) of the predicted residual signal (110) or a processed version thereof.

34. The apparatus according to any one of claims 22-33, further comprising an upsampler (318b) configured to upsample the selected pulse combination (118”, 318’) from the second sample to the first sample for further processing of the encoding.

35. The apparatus of claim 34, wherein the selected pulse combination (318b') of upsampling is used to update the adaptive codebook (140) and / or to determine the encoding / decoding gain.

36. The apparatus according to any one of claims 22-35 further includes a resampler (346) configured to resample additional signal or impulse responses required for a search within the at least one second sample codebook.

37. The apparatus according to any one of claims 22-36, configured to, in a first step, search for a pulse combination (118') in the second sample, and once the selected pulse combination (118') is found, configured to, in a second step, search for the gain of the selected pulse combination (118') in the first sample using an upsampled version (318b) of the selected pulse combination (118').

38. The apparatus according to any one of claims 22-37 is configured to downsample (310a) the predicted residual signal (110) or a processed version thereof in the time domain.

39. The apparatus according to any one of claims 22-38 is configured to downsample (310a) the predicted residual signal (110) or a processed version thereof in the time domain using a linear phase filter, and / or upsample the selected pulse combination or a processed version thereof.

40. The apparatus according to any one of claims 22-39, configured to convert the predicted residual signal or a processed version thereof, or the selected pulse combination or a processed version thereof, to the frequency domain, and to downsample the predicted residual signal or a processed version thereof in the frequency domain, and / or upsample the selected pulse combination or a processed version thereof.

41. The apparatus of claim 40 is configured to downsample (310a) the predicted residual signal (110) or a processed version thereof in the frequency domain by using a time-frequency block transform, such as a short-time Fourier transform (STFT), a fast Fourier transform (FFT), a discrete cosine transform (DCT), or a similar linear transform, and / or upsample the selected pulse combination or a processed version thereof.

42. The apparatus of claim 40 or 41 is configured to downsample (310a) the predicted residual signal or a processed version thereof in the frequency domain using a block transform with no overlap between adjacent blocks, and / or upsample the selected pulse combination or a processed version thereof.

43. The apparatus according to any one of claims 40-42 is configured to downsample the predicted residual signal (110) or a processed version thereof in the frequency domain using a block transform employing spectral truncation, and / or upsample the selected pulse combination or a processed version thereof in the frequency domain using a block transform employing zero-padding of the spectrum.

44. The apparatus of claim 43 is configured to downsample (310a) the predicted residual signal (110) or a processed version thereof in the frequency domain using constant scaling, and to upsample the selected pulse combination or a processed version thereof.

45. The apparatus according to any one of claims 22-44, configured to operate between at least a first operating mode and a second operating mode, such that: In the second operating mode, the encoding / decoding pulse position information (119) is determined using the second sampling; and In the first operating mode, the encoding / decoding pulse position information (919) is determined using the first sample, and the selected pulse combination is an entry in at least one first sample codebook (918), which contains a set of pulse combinations defined under the first sample.

46. ​​The apparatus of claim 45, configured to select between the first operating mode and the second operating mode.

47. The apparatus according to any one of claims 45-46, configured to select between at least a first operating mode and a second operating mode based on the target block size or the instantaneous bit rate of the current frame to be encoded (920).

48. The apparatus according to any one of claims 22-47, wherein the pulse information encoder (116, 316) is configured to compare the predicted residual signal (110) or a processed version thereof with a plurality of candidate signals (150), each of the plurality of candidate signals (150) being obtained from a corresponding codebook index (119', 319'), the pulse information encoder (116) being configured to select a particular codebook index (318') that allows the acquisition of candidate signals that minimize the error (362a) or a processed version thereof (366a) with the predicted residual signal (110) or a processed version thereof (366a) among the plurality of candidate signals (150, 350).

49. The apparatus of claim 48 or 49, wherein the pulse information encoder (316) is configured to select a specific codebook index that allows the acquisition of candidate signals (350) that minimize the error (362a) with respect to a downsampled version (310c) of the predicted residual signal (110) or its processed version among a plurality of candidate signals, wherein the plurality of candidate signals and the downsampled version (736) of the predicted residual signal (110) or its processed version are all under the second sampling.

50. The apparatus of claim 48, wherein the pulse position information encoder is configured to select a specific entry (822) associated with a candidate signal (118) that minimizes the error (362a) with the predicted residual signal (110) or its processed version among the plurality of candidate signals, wherein the predicted residual signal (110) or its processed version is at the first sample and the plurality of candidate signals (839) are at the second sample, the apparatus including an upsampler (840) to convert the selected candidate signal from the second sample to the first sample, thereby performing a comparison (360b) at the first sample.

51. The apparatus of claim 50, wherein the pulse position information encoder is configured to scale (352) the selected candidate signal onto the scaled selected candidate signal so as to compare the candidate signal based on the scaled, upsampled selected candidate signal (358b) with the prediction residual signal (110) or a processed version thereof (360b).

52. The apparatus of claim 51 is configured to scale the selected candidate signal (318b') upsampled by a plurality of candidate gains (354b) in order to select a gain that helps to minimize the error (838) and to encode gain information indicating the selected gain.

53. The apparatus according to any one of claims 22-52, wherein the signal processor is configured to perform a linear prediction codec (LPC) function.

54. The apparatus according to any one of claims 22-53, wherein the signal processor is configured to perform a long-term prediction (LTP) function.

55. The apparatus according to any one of claims 22-54, configured to avoid signaling the at least one second sample codebook in the codec signal and to limit its use to the block size of the current and / or codec mode or any other codec information already present in the block.

56. The apparatus according to any one of claims 22-55 is configured to search for a selected code combination among the code combinations of the at least one second sampled codebook (118) as a code combination that minimizes the error (162, 166) or other cost function between the predicted residual signal (110) and a candidate excitation signal (150) having at least one component obtained from the at least one second sampled codebook (118).

57. The apparatus of claim 56 is configured to perform a search in a loop (179) having multiple iterations by using candidate pulse combinations (118') from the at least one second sample codebook (118) and multiple candidate gains (119') scaled in the same loop for the candidate pulse combinations (118').

58. The apparatus of claim 57, wherein the at least one second sample codebook (118) is configured to output the candidate pulse combination (118') in the second sample and the candidate gain in the first sample, wherein the apparatus further comprises a mapper (118a) to convert the candidate pulse combination (118', 118a') from the second sample to the first sample in the same iteration.

59. The apparatus according to any one of claims 22-58, configured to perform a first step using the second sampling to search for a selected code combination (118') among a plurality of candidate code combinations (118') from the at least one second sample codebook (118) in a first iteration loop (179a) as a candidate code combination that minimizes the error (362a, 366a) or another cost function between the predicted residual signal (110) and a candidate excitation signal (150) having at least one component obtained from the candidate code combination (118') in the second sampling, and It is further configured to perform a second step using the first sample to search for a gain among a plurality of candidate gains in a second iteration loop (179b) using an upsampled version (318b) of the selected code combination (118) in the first sample, the gain minimizing the error or additional cost function between the predicted residual signal (110) and a candidate excitation signal (150) having at least one component obtained from the upsampled version (318b) of the selected code combination (118) scaled by the candidate gain.

60. A method for decoding an audio signal from an encoded / decoded audio signal, comprising: Encoding and decoding information about the prediction coefficients (105') and encoding and decoding information about at least one pulse (119) are read from the encoding and decoding signal (124). and In the first sampling, a decoded audio signal (702) is generated from a decoded version (704) of at least the prediction coefficients and a combination of decoded pulses (710) or a processed version thereof, wherein the first sampling means a first plurality of sampling positions having a first number of sampling positions in a frame or subframe. The method includes deriving (716) the decoded pulse combination (710) from the encoding / decoding information (119) and the second sample codebook (118) regarding at least one pulse, wherein at least one second sample codebook (118) comprises: a set of pulse combinations defined under a second sample, the second sample meaning a second plurality of sample positions having a second number of sample positions in the frame or subframe, wherein the first sample differs from the second sample at least in that the first plurality of sample positions differ from the second plurality of sample positions.

61. An audio encoding method for encoding audio signals, comprising: The prediction coefficients (104) and prediction residual signal (110) in the time domain are determined under the first sampling, where the first sampling means a first plurality of sampling positions having a first number of sampling positions in a frame or subframe; Determine encoding / decoding information (119) for a selected pulse combination representing the predicted residual signal (110), the selected pulse combination being an entry in at least one second sample codebook (118), wherein the at least one second sample codebook (118) comprises: a set of pulse combinations defined under a second sample, the second sample meaning a second plurality of sample positions having a second number of sample positions in a frame or subframe, wherein the first sample differs from the second sample at least in that the first plurality of sample positions differ from the second plurality of sample positions; and Write at least encoding / decoding information (105') about the prediction coefficients (104) and encoding / decoding information about the selected pulse combination.

62. A non-transitory storage unit that stores instructions, when executed by a processor, cause the processor to perform the method according to claim 60 or 61.

63. The apparatus according to any one of claims 1-21, configured to perform upsampling on the decoded pulse combination or a processed version thereof to obtain an upsampled decoded pulse combination or a processed version thereof to update the adaptive codebook.

64. The apparatus according to any one of claims 1-21 or 63, configured to perform upsampling on the decoded pulse combination or a processed version thereof in the time domain.

65. The apparatus according to any one of claims 1-21 or 63, configured to perform upsampling on the decoded pulse combination or a processed version thereof in the frequency domain.

66. The apparatus according to any one of claims 1-21 or 65, configured to perform upsampling on the decoded pulse combination or a processed version thereof in the frequency domain using block transforms that do not overlap between adjacent blocks.

67. The apparatus according to any one of claims 1-21 or any one of claims 65-66, configured to perform upsampling on the decoded pulse combination or a processed version thereof in the frequency domain using block transform and zero-padding of the spectrum.