Fast encoding in multichannel waveform coding with block prediction
Patent Information
- Application Number
- PCT/EP2026/058041
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-24
- Filing Date
- 2026-03-23
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026058041_01102026_PF_FP_ABST
Abstract
Description
[0001] FH250310PEP-2026096817.DOCX 1
[0002] Fast encoding in multichannel waveform coding with block prediction Description
[0003] Embodiments comprise decoders, encoders, methods for decoding, methods for encoding, computer programs and data streams for an efficient fast encoding in waveform codecs. 5 Introductory remarks:
[0004] In the following, different inventive embodiments and aspects will be described.
[0005] Also, some embodiments will be defined by the enclosed claims.
[0006] It should be noted that any embodiments as defined by the aspects can be supplemented by any of the details (features and functionalities).
[0007] 10 Also, the embodiments described can be used individually, and can also be supplemented by any of the features in another section, or by any feature included in the aspects.
[0008] Also, it should be noted that individual aspects described herein can be used individually or in combination. Thus, details can be added to each of said individual aspects without adding details to another one of said aspects.
[0009] 15 Moreover, features and functionalities disclosed herein relating to a method can also be used in an apparatus (configured to perform such functionality). Furthermore, any features and functionalities disclosed herein with respect to an apparatus can also be used in a corresponding method. In other words, the methods disclosed herein can be supplemented by any of the features and functionalities described with respect to the apparatuses.
[0010] 20 Moreover, features and functionalities disclosed herein relating to a method, in particular an encoding method can also be used in a data stream or bitstream (e.g. defining a respective data stream or bitstream element). Furthermore, any features and functionalities disclosed herein with respect to an apparatus can also be used in a corresponding data stream, e.g. as a resulting data stream as providing by said encoder. In other words, the data streams 25 disclosed herein can be supplemented by any of the features and functionalities described with respect to the apparatuses and methods.
[0011] Also, any of the features and functionalities described herein can be implemented in hardware or in software, or using a combination of hardware and software, as will be described in the section “Implementation alternatives”.
[0012] 30 Introduction and Motivation
[0013] Digital waveform codecs, like biophysical (e. g., medical), geophysical (e. g., seismic), or acoustic (e. g., audio) codecs, often employ block prediction methods to improve coding performance. In such codecs, a difference between the per-block / channel waveform signalFH250310PEP-2026096817.DOCX 2
[0014] and a corresponding (usually configurable or parametrizable) prediction signal is formed, and the block / channel-wise residual signal resulting thereform is quantized, entropy coded, and transmitted. At the decoder, the entropy decoded and “dequantized” residual block / channel signal is generated, and the same block prediction signal is added thereon to reconstruct the waveform signal in the initial domain.
[0015] Recently, standardization of a biomedical waveform coder has been started, and the current draft standard includes a block matching predictor, also sometimes referred to as long-term predictor (LTP). This predictor operates in a time domain and is controlled by a temporal offset parameter (or lag) L, which is signalled from encoder to decoder side, via a data stream, for synchronicity. A major drawback of LTP or block matching (BM) prediction is the large search space for finding an “optimal” value for L., i. e., an L resulting in good prediction performance with minimized block residual energy. In other words, the search for “optimal” L may slow down encoding considerably.
[0016] Summary of the invention
[0017] The invention is defined in the independent claims.
[0018] Figures
[0019] Figs. 1, 2A, 2B, 3A, 3B, and 4 show examples according to the invention.
[0020] Fig. 5 shows an example of an encoder according to the invention and of a decoder for decoding the stream encoded by the encoder.
[0021] Fig. 6 shows an operation according to the invention.
[0022] Examples
[0023] The description is extended in the following by the presentation of implementation examples. Before this, however, the description proceeds with a presentation of a possible framework or codec into which the embodiments described above as well as the examples described further below may be built into. Many details described in this framework are, however, optional when being combined with any of the above or subsequently described embodiments. To be more precise, the framework is described with respect to Fig. 5 which shows an encoder for encoding a multi-channel digital signal 14 into a datastream 16 as well as decoder 12 for decoding the multi-channel digital signal 14 from datastream 16. This description of Fig. 5 shall be seen as a presentation of new embodiments of the present application which result when combining any of the embodiments described above or any of the examples described subsequently or even any of the claimed subject matters is combined with the decoder 12 or encoder 10 of Fig. 5 either by adopting all details / functionalities described with respect to Fig. 5 or with leaving-out some of the details / functionalities described with respect to Fig. 5. Sometimes such “optional” features ofFH250310REP-2026096817.DOCX 3
[0024] Fig. 5 are explicitly identified as being optional with respect to the combination of the previously and subsequently described embodiments, but the just-mentioned possible combinations of the previously / subsequently explained embodiments with the description of Fig. 5 shall not be restricted to the these explicitly identified variations of Fig. 5 in terms of leaving-out certain features.
[0025] In Fig. 5, the multi-channel digital signal 14 is illustrated by way of an array of samples with the samples being illustrated as small squares 18. Each line / row corresponds to a certain channel of the multi-channel digital signal 14. Each channel of signal 14 may have associated therewith a respective channel ID and Fig. 5 shows these channels as being ordered according to their channel ID along vertical axis 20 (e.g. channel dimension) which, thus, corresponds to a “source” channel axis 20 (e.g. time dimension). The horizontal axis 22 corresponds to time so that samples 18 forming one column, or being horizontally aligned, are samples belonging to one common time instant. Such set / column of temporally co¬ located samples 18 is illustrated in Fig. 5 at 24.
[0026] Each channel, thus, forms a digital time-varying signal or time / amplitude or time-to-amplitude signal. The multi-channel digital signal m might have been obtained by at least one of Electrocardiography, Electroencephalography, Electromyography or seismic measurement. Differently speaking, the multi-channel digital signal might be a bio-physiological waveform data such as an electroencephalography (EEG) signal, an electrocardiogram (ECG), or an electromyography (EMG) signal, or seismic waveform data. However, each channel / signal might alternatively be another sort of waveform signal data such as scalar media data such as an audio signal and the signal 14 might be a multi-channel audio signal.
[0027] Fig. 5 illustrates the option according to which signal 14 is not coded directly, i.e., in the original domain 26, but in a so-called “coded domain” 28 which might differ from the original domain 26 by one or more of 1) channel transformation, 2) channel permutation and 3) temporal mutual channel alignment. The channel transformation, if applied, transforms, per sample time instant, a set or column 24 of samples from domain 26 to domain 28. Thus, in domain 28, the sample pitch and the time axis is the same as in domain 26, but the meaning of the channels is different, i.e., the “source” channels of domain 26 become transformed channels in domain 28. Accordingly, the vertical axis in Fig. 5 for domain 28 is denoted as 32. Note that the channel transformation might leave the number of channels unchanged so that there is the same number of channels in domain 26 as well as domain 28, but different approaches are also possible. Generally, the channel transformation would aim at reducing redundancy and trying to condense the channels’ energy onto a fewer number of channels in domain 28. As said, the channel transformation is optional. Accordingly, in general terms, the channels in domain 28 are called “coded channels” in order to distinguish them from theFH250310PEP-2026096817.DOCX 4
[0028] “original” or “source” channels of digital signal 14 in domain 26. The permutation is also optional and may be used in combination with, or without, the channel transformation. If used in combination with the channel transformation, the permutation may be performed prior to and / or or subsequent to the channel transformation in order to permute / sort the source 5 channels prior to transformation and the coded channels subsequent to the channel transformation. The channel transformation might be a DCT, DST, FFT or any other transformation. The temporal mutual alignment is also optional and might be seen as a constant temporal alignment between the source channels or the coded channels.
[0029] The module in encoder 10 performing the one or more of channel transformation, channel 10 permutation and temporal mutual alignment is indicated in Fig. 5 as block 34. Side information 36 might be used in order to signal information on one or more of the following: 1) The channel transformation used, 2) information on the permutation(s) among the source channels and / or coded channels and 3) information on the mutual temporal alignment / delays between the source channels or coded channels wherein the temporal mutual alignment 15 might be restricted to full sample precision. A corresponding block 38 in decoder 12 performs the reverse step, i.e., performs one or more of: 1) a channel retransformation, 2) a repermutation of the source channels and / or coded channels and 3) a temporal re-alignment of the source channels or coded channels. Note, that if no channel transformation takes place, the coded channels are, in fact, equal to the source channels except for being temporally 20 mutually aligned or being differently sorted due to permutation. Block 38 might be controlled by the before-mentioned side information 36.
[0030] Thus, the “actual coding" relates to the coded channels in domain 28. In the coded domain 28, the coded channels are depicted in Fig. 5 as lines or rows of samples 40, each extending along time axis 22, the coded channels being depicted one on top of the other along coded 25 channel axis 32 - potentially ordered according to a coded channel ID they have associated therewith - so as to result into an array of samples 40. Again, although Fig. 5 depicts the case that the number of source channels equals the number of coded channels, the number might be different. Further, if channel transformation is used, while there is no longer a clear association between source channels on the one hand and coded channels on the other 30 hand, the temporal association remains: For each temporally co-located samples 24, there is a corresponding temporally co-located set 42 of samples 40 of the coded channels, wherein the set 42 in domain 28 is a column and might be a set of horizontally mutually offset samples in case of, and according to, the mutual temporal alignment, if applied. In case of Fig. 5, it has been assumed that no such temporal alignment took place so that both sets 42 35 and 24 are pure columns in the time / channel representation.rH250310PEP 2026096817. DOCX 5
[0031] The actual coding is done in units of so-called temporal blocks 30. The term “block” or “temporal block” 30 is used so as to denote both a temporal portion of the multi-channel signal in domain 28, i.e., the set of coded channels, as well as a temporal portion of a certain coded channel. That is, for each temporal block 30, each coded channel has a temporal block such as block 140 depicted for some temporal block 30c and same are mutually colocated. The coding is done sequentially along these blocks 140, by following a coding / decoding order, which traverses the blocks 140 temporal block 30 by temporal block 30 with traversing temporally co-located blocks of the coded channels along a channel order corresponding to the order of the coded channels along axis 32. This coding / decoding order is illustrated in Fig. 5 at 60. That is, in case of temporal block 140 being the block currently to be coded / decoded, the previously decoded / encoded temporal blocks include all preceding temporal blocks of all coded channels as well as the temporally co-located temporal blocks of coded channels preceding the coded channel 92 of temporal block 140 in channel order. These previously coded / decoded temporal blocks and their samples are illustrated in Fig. 5 by way of shading. In this regard, note that in Fig. 5, merely one temporal block 140 has been illustrated explicitly in order to reduce the complexity of Fig. 5. Thus, in the specification herein, reference sign 140 is sometimes used to indicate the currently encoded / decoded temporal block or to stand representatively for all temporal blocks. Further, as depicted in Fig. 5, the partitioning of signal 14 into temporal blocks 30 and 140, respectively, might be done in a manner so that these blocks 30 and 140, respectively, are non-overlapping.
[0032] The actual coding in units of the temporal blocks 140 is performed predictively. That is, the encoder 10 comprises a block predictor 62 which predicts the samples of the currently coded temporal block 140, thereby yielding a prediction signal 64, and the prediction residual 66 formed by a subtraction between the actual sample values of temporal block 140 and the predicted samples of prediction signal 64 formed at a subtractor 68 is coded into the datastream 16 by residual coder 70. The residual coding in residual coder 70 may, or may not, involve a coding error by means of quantization. In any case, block predictor 62 uses the reconstructable version as being available by previously coded temporal blocks in order to obtain the prediction signal 64. This reconstructable version 72 might be derived at encoder 10 by means of a residual decoder 74 which reverses, potentially under coding loss, such as quantization, e.g. by means of dequantization, the residual signal 76 as coded into datastream 16, and an adder 78 which sums-up prediction signal 64 and the reconstructable residual signal 80 as obtained by residual decoder 74. To be more precise, let’s call the channel-individual temporal blocks 140 subblocks with temporally collocated subblocks of all channels forming a temporal block 30. Then, the prediction in module 62 or, to be more precise, the prediction at encoder and decoder, is performed in units of the subblocks 140, i.e. subblock wise. The encoder is free to choose different prediction modes for the
[0033] eFH250310PEP-2026096817.DOCX 6
[0034] subblocks within one block 30. As explained in more detail herein, within one block 30, one subblock 140 may be predicted based on one or more subblocks previously - according to the decoding order 60 - en / decoded within this block 30, while another subblock 140 within that block 30 might be coded / decoded based on the previously en / decoded subblock 140 of the same channel (but within the previous block 30). The transform residual en / decoding is then performed subblock wise by use of a one-dimensional transform signaled in the data stream as described hereinbelow.
[0035] The decoder 12 decodes the coded channels from data stream 16 in a corresponding manner, i.e., in units of the temporal blocks 30 or in temporal blocks 140, respectively, and using predictive decoding. To this end, the decoder 12 comprises a residual decoder 82, an adder 84 and a block predictor 86 which correspond to, and are mutually connected in the same manner as, elements 74, 78 and 62 of encoder 10. That is, the residual decoder 82 derives from the residual signal 76 in data stream 16 the reconstructable residual signal 80 for a currently decoded temporal block 140 which is then subject to addition with prediction signal 64 derived by block predictor 86 for temporal block 140 on the basis of the reconstructed version 72 of previously decoded temporal blocks at adder 84. The output of adder 84, thus, yields the reconstructed version 72 of the currently decoded temporal block 140 and becomes part of the pool of already decoded samples of previously decoded temporal blocks when the temporal blocks of the coded channels are, in this manner, traversed along coding / decoding order 60 so as to reconstruct the coded channels in the coded domain 28.
[0036] Note that the above description concentrated on the so-called sample prediction where samples of a current block 140 are predicted based on reconstructed samples of one or more previously decoded blocks, but coding inter dependencies, namely intra-channel and inter-channel coding dependencies may be exploited not only in terms of sample prediction, but also in terms of other coding tools involving, for instance, parameter prediction and / or context derivation.
[0037] In order to enable a high degree of random access capability, some of the temporal blocks 30 may be coded in a random access manner meaning that the coded channels therein are coded independent from previous temporal blocks 30. Imagine, for instance, that temporal blocks 30b and 30e are random access temporal blocks. Then, none of the temporal channel blocks 140 in temporal block 30b as well as 30e would depend on any preceding temporal block 140 and no coding dependency would cross these temporal blocks 30b and 30e, that is no temporal block 140 within any of temporal block 30b-30d would be coded depending on any block 140 temporally preceding temporal block 30b, and no temporal block 140 withinFH250310PEP 2026096817. DOCX 7
[0038] any of temporal block 30e and following would be coded depending on any block 140 temporally preceding temporal block 30e.
[0039] Thus, in other words, coding dependencies are restricted so as to not reach-out beyond the border of a random access temporal block 30b and 30e towards any preceding temporal block 30. Such restriction might also hold for intermediate temporal blocks 30c to 30d between random access temporal blocks 30b and 30e in that same may not depend on any temporal block preceding the leading one among the random access temporal blocks 30b and 30e, here block 30b. Accordingly, leading temporal borders of the random access temporal blocks 30b and 30e are indicated by bold lines in Fig. 5. In a variant, the restriction is not valid for all en / decoding stages. For instance, while the grouping might hold true for prediction, but the residual en / decoding dependencies might cross borders between channel groups. It might be the case, for instance, that for the entropy coding and decoding, all channels are coded jointly, i.e. using a single arithmetic coding engine, but that for the sake of prediction and reconstruction, the channels are grouped as described into independent groups such that, after entropy decoding, each such group can be reconstructed completely independently from each other group. This means that no prediction of sample values or any other information is supported between different channel groups.
[0040] Further, it might be that the coding of the coded channels also interrupts or restricts inter¬ channel dependencies. For example, one or more of the coded channels might be coded as random access coded channels so that same do not use inter-channel dependencies, but merely intra-channel dependencies. The restriction of inter-channel coding dependencies might follow the channel order 32: that is, coding of these random access coded channels and the intermediate coded channels therebetween would be restricted so as to not reach-out beyond such a random access coded channel toward any coded channel preceding that random access coded channel in channel order along axis 32. Two such random access coded channels 88a and 88b and the resulting inter-channel dependency borders are illustrated in Fig. 5. Note that the restriction of inter-channel dependencies might be differently and is illustrated here merely as an example where the definition of, along channel order 32, interspersed random access channels 88a and 88b defines channel groups covering contiguous channels along the channel order 32. Other groups of channels might be defined, which do not necessarily follow the channel order 32, and inter-channel dependencies might be restricted not to render any channel of one group dependent on a channel of any other group, and within each group the inter-channel dependencies may also by restricted or each channel might by coded inter-channel dependent on any previously coded channel within its channel group.
[0041] AFH250310PEP-2026096817.DGCX 8
[0042] The block predictor 62 and 86 of encoder 10 and decoder 12, respectively, operate synchronously, i.e., they generate the same prediction signal 64 based on the previously encoded / decoded samples of previously encoded / decoded temporal blocks 140. On encoder side 10, the prediction for a certain temporal block 140 may be accompanied or determined 5 by one or more prediction parameters. Same might be determined on encoder side based on a rate / distortion optimization. These prediction parameters 90 are coded into data stream 16 and they are decoded from data stream 16 and used by block predictor 86 so as to perform the same prediction.
[0043] It might be that encoder 10 and decoder 12 support more than one prediction mode. For 10 instance, encoder 10 and decoder 12 may support an intra prediction mode (which mode may also be called block-copy mode) according to which the currently encoded / decoded temporal block 140 is predicted based on the reconstructable sample values of previously encoded / decoded temporal blocks of the same coded channel to which the currently encoded / decoded temporal block 140 belongs, which is coded channel 92 in the example of 15 Fig. 5. Additionally or alternatively, encoder 10 and decoder 12 may support an inter¬ prediction mode (which mode may also be called cross-channel prediction mode) according to which the currently encoded / decoded temporal block 140 is predicted based on the reconstructable sample values of previously encoded / decoded temporal blocks of one or more coded channels preceding - in coding order 32 - the coded channel 92 to which the 20 currently encoded / decoded temporal block 140 belongs. Additionally or alternatively, there may be a mixed prediction mode according to which the prediction signal 64 is obtained by both, reconstructed / reconstructable sample values of previously encoded / decoded temporal blocks of coded channel 92 itself as well as reconstructed / reconstructable sample values of one or more coded channels preceding coded channel 92 in channel order along axis 32. 25 Beyond this, there may be temporal blocks 140 which are coded without any prediction at encoder 10 and decoded without any prediction at decoder 12 such as the first temporal blocks 140 in the tiles 94 resulting from mutually separating the temporal blocks by means of the random access borders 96 on the one hand and the random access channel borders 98 on the other hand. This corresponds to the prediction signal 64 being set to zero and this 30 may form an additional mode which could be called bypass mode. Additionally, or alternatively, there may be other modes such as ones deriving a DC predictor or linear function predictor for block 64 based on immediately preceding samples which immediately precede block 140. The prediction parameters 90 may, thus, contain for a currently encoded / decoded temporal block 140 a prediction mode flag or prediction mode indicator 35 indicating the prediction mode to be used for this currently encoded / decoded temporal block 140 and, optionally, one or more parameters parameterizing the prediction mode to be used for this currently encoded / decoded temporal block 140. It might also be that the predictionFH25O31OPtP-2O2bO9681 / .DOCX 9
[0044] parameters are themselves coded predictively from already reconstructed blocks 140. In this prediction process, the laid out random-access capabilities in channel- and temporal¬ direction are, as an example, always maintained, i.e. the mentioned prediction of prediction parameters may never be supported across such a random access segment.
[0045] As mentioned, the aforementioned coding dependencies ought not to cross any of the borders 96 and 98 not only result from the just-described sample prediction capabilities of block predictor 62 and 86, respectively, but may optionally also result from other mechanisms such as parameter prediction according to which parameters such as the aforementioned prediction parameters 90 for a certain temporal block 140 are predicted based on coding parameters conveyed in the data stream 16 for any previous temporal block, or context derivation for context-adaptive entropy coding / decoding any coding parameter such as the prediction parameters 90 or any other side information such as side information 76 and 36 for temporal block 140 based on any coding parameter conveyed in the data stream 16 for any preceding temporal block.
[0046] That is, summarizing, the encoder 10 encodes the multi-channel signal 14 by transferring it into the coded domain 28 and then coding the coded channels into data stream 16 in the just-described block-wise and predictive manner, wherein decoder 12 decodes the coded channels of coded domain 28 from data stream 16 and the corresponding block-wise and predictive manner with then gaining the multi-channel signal 14 in its original form 26 based on the coded channels in coded domain 28 by means of segment 38. As said, the channel transformation is optional and if not used, each sample 40 in the coded domain 28 really corresponds to one sample 18 in the original domain 26. If, further, the temporal mutual alignment is not used, each sample 40 exactly corresponds to a sample 18 in the original domain 26 at exactly the same time instant or, differently speaking, all temporally co-located samples 40 in coded domain 28 remain mutually temporally co-located in the original domain 26.
[0047] It should be noted that the temporal blocks 30 might, other than illustrated in Fig.5, vary in block length rather than being of a constant length as depicted in Fig. 5. For instance, encoder 10 may decide on the length of blocks 30 and signal the block length of blocks 30 (and the corresponding temporal blocks 140 of the coded channels) within data stream 16. Such signaling might be done on block level, such as for each temporal block 30 or, differently speaking for each temporally aligned bundle of blocks 140, so that the encoder may decide on the block size on the fly, or the block length might be signaled in the stream 16 on a larger scope such as for a sequence of blocks or even the whole stream 16.
[0048] As to the residual coder and residual decoder 70 and 82, they may use transform coding / decoding in order to convey the residual signal 76 in data stream 16. That is, theFH250310PEP-2026096817.DOCX 10
[0049] residual signal 80 may be conveyed in data stream 16 in transform or spectral domain by way of transform coefficients in residual signal 76. The transform domain might be a DCT, DST or an FFT. The transform may be non-overlapping, i.e. it may only transform residual signal 80 and its re-transform may only cover residual signal 76 within block 140, and / or may be non-windowed, i.e. the residual signal might be transformed without any transform window used to temporally shape the residual signal 80 before the transform. The transform domain, i.e. the transformation leading from time domain to transform domain which is used by the encoder to transform the prediction residual signal 80 to be coded und the corresponding re-transformation leading from transform domain to time domain which is used by the decoder to derive the prediction residual signal 80, or the transformation, might be selected from a set of available transforms including, for instance, one or more of 1 ) one or more DCTs, 2) one or more DSTs and 3) an identity transform according to which the prediction residual signal 80 is coded into the data stream 14 in time domain directly. The transform may be critically sampled in that the number of transform coefficients resulting from the samples of one block 140 may equal the number of samples of block 140. Again, the samples might be the residual samples or may be, in case of the bypass mode, the channel samples directly.
[0050] The transform coefficients might be encoded by quantization, i.e. they may be quantized with the quantized coefficients then being coded in the datastream 16. Dequantization may occur at decoding. For quantization, either a scalar uniform reconstruction quantizer or a low complexity vector quantizer might be used. In order to determine the quantization indices, the encoder may perform some optimization algorithm such as a rate-distortion optimized scalar quantization, or a trellis quantization with the goal to approximately minimize an approximated Lagrangian rate-distortion cost. At the decoder, the reconstruction process that yiels the transform coefficients may be conducted by multiplying the coded quantization indices with a certain step-size and, in case of the use of a low-complexity vector quantizer, by additionally invoking a state-machine based on the parity of previously decoded quantization indices in order to reconstruct the current quantization index.
[0051] In order to control the quantization noise, the transform coefficients might be subject to noise shaping. Spectral noise shaping may be used to shape the quantization noise spectrally. This may be done by signaling in the data stream spectral-band scale factors, i.e. a scale factor per spectral band, which represent a transfer function of a spectral filter which approximates the spectral envelope of the signal within the current block 140 (or its prediction residual, respectively), or signaling filter coefficients defining a temporal filter having a filter transfer function which approximates the spectral envelope of the signal within the current block 140 (or its prediction residual, respectively). On encoder side, spectral noise shaping may be applied in spectral domain by multiplying an inverse of scale factors,FH250310PEP-2026096817.DOCX 11
[0052] either directly signaled in the data stream or derivable from the filter coefficients by filter-to- factor conversion, with the transform coefficients before quantization. That is, at encoder, the coefficients are shaped by the inverse of the spectral envelope. At decoder side, spectral shaping may be applied in spectral domain by multiplying scale factors, either directly signaled in the data stream or derived from the filter coefficients by filter-to-factor conversion, with the transform coefficients, with then . That is, at decoder, the coefficients are shaped by the spectral envelope before applying retransformation. Additionally or alternatively, temporal noise shaping might be applied. To this end, TNS filter coefficients might be determined and signaled by the encoder. The TNS filter coefficients may represent a transfer function which approximates the temporal envelope of the current block 140 (or its residual signal). The encoder may apply TNS filtering using the filter coefficients by spectrally filtering the possibly spectrally shaped transform coefficients so as to filter them with a transfer function corresponding to an inverse of the temporal envelope. The TNS filter coefficients might be derived by linear prediction analysis of the possibly spectrally shaped transform coefficients so as to derive a linear prediction filter, then used as TNS filter, which minimizes a prediction residual when spectrally applied on the possibly spectrally shaped transform coefficients. At the encoder, the TNS filtered coefficients are then quantized and entropy coded. At decoder side, the inverse takes place: the possibly spectrally shaped transform coefficients are inversely TNS filtered before applying retransformation. Additionally or alternatively, noise filling might be used. The filling may be applied to zero-quantized portions of the spectrum and controlled by the encoder via corresponding noise filling parameters.
[0053] As to the encoding / decoding the block or sequence of quantized transform coefficients of a current block into / from the data stream 16, arithmetic coding, such as context-adaptive binary arithmetic coding, CABAC, may be used. The CABAC encoding / decoding may by performed frame wise. That is, in each channel, the sequence of blocks 140 may be partitioned into immediately consecutive blocks 140, which form frames. This partitioning may be equal among the channels so that, again, a frame denotes both a temporal portion within each channel individually, as well as a temporal portion of the multi-channel signal, i.e. a collection of temporally aligned frames. Within each frame, the sequence of blocks 140 are CABAC en / decoded with once initializing the contexts and resetting the internal CABAC state at the beginning and then updating the contexts’ probabilities during en / decoding the respective frame. That is, blocks 140 are CABAC decodable merely in units of frames. The context initialization might be done independent from previous frames, or depending on the contexts as manifesting itself at the end of, of during, the en / decoding a previous frame. Some deblocking processing might be used to avoid blocking artifacts. If, alternatively, an overlapped transform is used, an overlap-add processing with re-transforms of immediatelyFH250310PEP-2026096817.DOCX 12
[0054] preceding / succeeding temporal blocks of the same coded channel might be used in order to completely reconstruct the current temporal block’s 140 residual signal 76.
[0055] Besides such transform-(residual)-coded blocks there might be temporal blocks 140 which, additionally or alternatively, are coded using, besides the block prediction by block predictor 62 / 86 - which could be called a primary prediction - a secondary sample-wise prediction of the residual samples in residual block 66 such as by predicting a current sample’s residual sample by means of already decoded values of preceding - in sample coding order - residual samples in block 66 or 80, with then correcting same by means of a secondaryprediction-residual sample decoded from the data stream 16. The secondary-prediction- residual samples for such a block may coded into the data stream en block in a transform domain or sample-wise in time domain.
[0056] Note that the afore-mentioned spectral shaping of the residual signal of a block 140 might be seen as a sample wise residual prediction, i.e. the case where filter coefficients are signaled for a block which define a temporal filter having a filter transfer function which approximates the spectral envelope of the residual signal within a current block 140. In sample wise residual prediction, the residual predictor on a current block 140 might either be chosen out of a fixed set of prediction modes, where an index to such a residual prediction mode is signaled in the bit-stream, or the residual prediction mode might be 'signal adaptive’. In the latter case, prediction filter coefficients for the residual predictor are determined at the encoder by solving for example a linear equation, and are then quantized and transmitted to the decoder. At the decoder, the coefficients are inverse quantized and then the sample-wise prediction is conducted with these coefficients. The number of used coefficients may vary per block and might also be signaled in the bit-stream. Additionally, it might optionally (i.e. indicated by some information in the bit-stream) be supported to invoke collocated samples from a previous block for the sample wise residual prediction. Finally, the coefficients of the sample wise residual prediction might be coded predictively, i.e. be predicted from used coefficients of a previous block, where only the differences to the current coefficients are transmitted.
[0057] A final note shall be made with respect to the juxtaposition of frames, blocks 140, channels and channel groups and regarding decoding order. The description above already described the fact that the channels might be grouped into channel group with each channel group being coded independently from each other, meaning that the blocks 140 in a certain channel group are coded without dependencies from channels outside their channel group. The decoding order 60, thus, would traverse the channels channel-group individually, channel group by channel group. Within each channel group, the blocks 140 are traversed as described: all temporally aligned blocks 140 of all channels fist, then proceeding with the next
[0058] IFH250310PEP-2026096817.DOCX 13
[0059] blocks 140 and so forth. A frame may have a sequence of blocks of a channel group encoded thereinto along the mentioned decoding order, such as n temporally consecutive blocks 140 for all channels of a channel group. IF the channel group had m channels, m*n block104 would, thus, be coded into the frame. As mentioned, there might be dependent 5 frames, for which the CABAC contexts are adopted from the preceding frame of the same channel group, i.e. the one having encoded the immediately preceding block 140. For such dependent frames, not only CABAC contexts may be adopted from the preceding frame, but it may also be allowed to allow for prediction from the preceding frame to the dependent frame. Prediction, and possibly also any coding dependencies, towards channels outside the 10 channel group and, within the channel group, towards frames temporally preceding the mostly recently previously en / decoded independent frame would be disallowed. Thus, each tile shown in Fig. 5 by bold lines may represent a sequence of an independent frame flowed by zero, one or more dependent frames.
[0060] As mentioned before, Fig. 5 only represents a possible “framework” into which the previously 15 described embodiments and the embodiments described subsequently may be built into.
[0061] Many modifications may be performed with respect to Fig. 5, and some of these modifications might be mentioned in the subsequent description with respect to certain ones of the subsequently described embodiments, but these modifications shall then be treated as being also applicable with respect to other ones of the subsequently described embodiments.
[0062] 20 The description is now resumed with respect to the announced subsequently described implementation examples and further embodiments where the digital time-varying signal is not restricted to be a channel of a multi-channel signal or to be a multi-channel signal, but where same may only be a single digital scalar signal.
[0063] The present examples
[0064] 25 There is defined the encoder 10 for encoding the digital time-varying signal. The encoder 10 encodes the digital time variant signal in a data stream 16 in temporal blocks 30. The encoder 10 may:
[0065] 1) Predict (62) the current temporal block 30, to obtain a prediction signal 64, from an already encoded (already reconstructed, subjected to encoding, subjected to 30 reconstruction) portion of the signal separated from the current temporal block 30 by a current offset;
[0066] 2) derive the prediction residual signal 66 by comparing (68) the signal with the prediction signal 64;
[0067] 3) encode (70) in the data stream 16, the prediction residual coefficients 76 which 35 represent the prediction residual signal 66 andFH250310PEP-2026096817.DOCX 14
[0068] 4) signal information indicating the current offset (e.g. in the prediction parameters 90).
[0069] The operation of predicting (at block predictor 62) may include at least one of:
[0070] deriving, from at least one previous offset (e.g. L1, L3) used for predicting an already encoded (already predicted) temporal block (e.g. 301, 302), at least one putative offset (e.g., L2', L2”),
[0071] determining the current offset (L2) from the least one putative offset (e.g., L2’, L2”) through coding efficiency measure(s),
[0072] deriving the prediction signal 64 from the current offset (L2).
[0073] Hence, many of the techniques discussed here are for deriving the at least one putative offset (e.g., L2’, L2”) from the previously encoded (predicted) temporal block and for determining the current offset L2 from at the least one putative offset. Efficiency measure(s) may be evaluated, thereby minimizing a cost function indicative of the error in predicting the prediction signal 64.
[0074] There are many techniques for obtaining the putative offset and for obtaining the offset. Figs. 1, 2A, 2B, 3A, 3B and 4 show examples of these techniques. In these examples, diagrams are shown with a time dimension abscissa 20 and a channel dimension 20 (“source channel axis”) in ordinate. In these figures there is shown a current block 30 which is currently to be predicted and an already encoded (already predicted, subjected to prediction, subjected to encoding) portion 310 as well as an already encoded (subjected to prediction, subjected to encoding) block 301. The already encoded block (already predicted, subjected to prediction, subjected to encoding) 301 has been predicted by relying on an offset (lag) L1 from the already encoded portion (already predicted, subjected to prediction, subjected to encoding) 310. Fig. 1, therefore, shows the offset used for the already encoded (already predicted, already subjected to prediction, already subjected to encoding) block. Now, the current block 30 is to be encoded (predicted). The current block 30 is also predicted (encoded) by relying on a lag (offset) L2. The real (actual) lag L2 for encoding (predicting) the current block 30 may be searched in a test area (e.g. interval) 312 derived from L2’. In Fig. 1, the already encoded (already subjected to encoding, prediction) block 301 had been encoded (already subjected to encoding, prediction) through a lag (offset) L1. So, the search of the offset L2 for encoding (predicting) the current block 30 will be in the test area (interval) 312 derived from a putative offset (putative lag) L2’. The putative offset L2’ may in general be the same as the offset L1 used for the already encoded (already subjected to encoding, already subjected to prediction) block 301 or, in some other cases, there can be a different relation (like L2’ being a multiple or a submultiple of L1) or anyway could be derived basedFH250310PEP-2026096817.DOCX 15
[0075] on a predetermined known rule. The test area derived from L2’ can be an interval which contains L2’ (while it could be that L2 -L1). Then, by evaluating the efficiency measure(s), the effective current offset L2 for the current block 30 will be retrieved. Notably, the test area 312 is much less elongated than the length of the already encoded (already predicted) portion 310. Accordingly, a great reduction of the computational power is obtained. It is noted that Fig. 1 is in the case in which the current block 30 and the already encoded (subjected to encoding, predicted, subjected to prediction) block 301 are displaced in different temporal positions. The effective offset L2 can be encoded as parameter in the side information of the datastream 16. The test area (interval) 312 is therefore an area comprising all the candidate offsets and is evaluated based on the efficiency measure(s).
[0076] In Fig. 2A the already encoded block has a lag (offset) L1, and the putative lag L2’ (putative offset) for the current block 30 can be the same of L1 (or could be obtained through a predetermined, a priori known mathematical formula from L1, such as L2’ being a multiple or sub-multiple of L1, or in another predetermined relationship, e.g. known a priori). Fig. 2B shows the same of the example of Fig. 2A, but in this case, a test area 312 determined from L2’ is shown. The test area 312 will permit to search the real (actual) offset (lag) L2 to be used by evaluating efficiency measure(s) only within the test area. As can be seen, the test area 312 is determined from L2’ as, for example, an interval 312 either centered in L2’ or having L2’ as an internal value, or reference etc. It By evaluating only within the test area determined from L2’ like in Fig. 2B, the calculation power waste is greatly reduced, since it is not necessary to calculate the efficiency measure(s) throughout the whole already encoded (already predicted) portion 310. The test area (interval) 312 is therefore an area comprising all the candidate offsets and is evaluated based on the efficiency measure(s).
[0077] The examples of Fig. 1 and Fig. 2B (and also in Fig. 3B and also in other examples shown in the figures) the test area 312 (which in general is a proper subset of the set of possible values of the offset (e.g., the whole already encoded portion)) includes, or is at least determined by, the putative offset L2’. Hence, the offset L2 will be chosen among the multiple candidate offsets based on the coding efficiency measure(s). The test area 312 may be at least one interval, which has interval ends at predefined distances from the putative offset L2’ [e.g. the interval ends are positioned at pre-defined distances may be fixed and / or independent of the values of L2’ and / or any value of the signal to be encoded], [e.g. the interval ends are positioned at distances which are controlled by some values, e.g. by the resolution, e.g.: the finer the temporal and / or channel resolution, the larger the distance in term of samples and / or channels].
[0078] Fig. 3A shows another example in which there is a choice between a first putative lag L2’ and a second putative L2”. In this case, the already encoded blocks 301 and 302 which are takenFH250310PEP-2026096817.DOCX 16
[0079] into account are more than one. Fig. 3A shows that the current block 30 is being predicted based on two colocated blocks in the different channels in the same time coordinate (but it could also be in the same channel at different times, or in the different times or different channels indistinctly). Here, a first already encoded (already predicted, already subjected to 5 prediction, already subjected to encoding) block 302 has been previously predicted (subjected to encoding) from a first lag L3. A second already encoded (already predicted, already subjected to prediction, already subjected to encoding) block 301 has been predicted (and subjected to encoding) through a second lag L1. Therefore, there will be two putative lags (offsets) L2’ and L2”. The putative lag L2’ is obtained from L3 (e.g., L2’ = L3, or L2’ could 10 be in a predefined mathematical relationship with L3, e.g., L2’ could be a multiple or a sub¬ multiple of L3) and the putative lag (offset) L2” can be obtained from L1 (e.g., L2” = L1, or L2’ could be in a predefined mathematical relationship with L3, e.g., L2” could be a multiple or a sub-multiple of L1). Hence, there is more than one putative lag L2’ and L2”. The decision in one example can be between L2’ and L2”, so as to find an offset lag L1 for the current block 15 30. The choice may be made, for example, by evaluating the efficiency measure(s). In another example, shown in Fig. 3B (which is the same as Fig. 3A with the addition of the intervals), from the putative lag L2’ there may be retrieved a first interval (e.g., comprising L2’) 312’ and from L2” there could be identified a second interval 312” (e.g., comprising L2”). The first and second intervals 312' and 312” may be two subsets (proper subsets) of the 20 already encoded portion 310.
[0080] In the examples of Figs. 3A and 3B the chosen offset L1 may be the one which maximizes the efficiency measure(s), similarly to the examples of Figs. 1, 2A, and 2B. It is to be noted that all the examples above have been done with an offset in the time dimension, but the offset can also be in the channel dimension, because it is possible that the offset is in the 25 channel dimension (e.g. vertical in Figs. 1-5).
[0081] In another example there may be a single putative lag which is obtained as a combination (e.g. linear combination, such as an average) between the first lag L3 and the second lag L1, so that the putative lag is e.g. obtained as (L3+L1 ) / 2, or another average (in Fig. 3A or 3B, the uniquely putative lag would be between L2’ and L2”, for example), and even in this case 30 (e.g. in a variant to Fig. 3B) the test area 312 (which would be unique) could be centered (or at least contain) the putative offset.
[0082] In the example of Fig. 4, the block 30 currently predicted is obtained based on one of the sub blocks 303 and 304 (which are proper subsets of the block 30, and which may represent a partition of the block 30, e.g. they may have void mutual intersection and their union may be 35 the whole block 30). In this case, therefore, there may be two putative offsets (e.g. LT from L1 , offset of the first sub block 303, and / or L2’ from L2, offset of the second sub block 303).FH250310PEP-2026096817.DOCX 17
[0083] In practice, instead of determining the current offset L2 based on the complete block 30, each of the two sub blocks 303 and 304 is predicted (encoded), and for each of them a putative lag is obtained. Then, the actual (final, real) lag (offset) L2 will be obtained, e.g. by choosing one of the two previously used lags (e.g. like in Fig. 3A), and / or by analysing two test areas (e.g. like in Fig. 3B), and / or through any combination (e.g. linear combination, e.g. average, either by using the linear combination or average as offset L2 or by determining one ingle interval, e.g. centered or otherwise containing the putative offset). Any of the examples based on Fig. 4 (described or variant) can be based with more than two subblocks.
[0084] In the examples above, the at least one putative offset [e.g. L2’ and / or L2”] may be a scaled value, through a pre-defined [e.g., fixed, independent of L1 ] scaling value, derived from the at least one previous offset [e.g. L1] for the already encoded temporal block [e.g. L2’=m*L1 , where m is a pre-defined constant (e.g. fixed, e.g. independent of L1) value, e.g. a real value, a relative value, a rational value, or a natural value (m may be greater than 0; in some cases, m may be less than 1)].
[0085] Fig. 6 shows an example of operation, e.g. for any of the examples of Figs 1 and 2B. Indeed, it may be possible to decide whether the offset L2 for the currently encoded (predicted) block 30 is to be chosen as the putative offset L2’ itself or shall be chosen according to an alternative technique. Fig. 6 shows an example of operation 600. The step 602 provides for measuring coding efficiency (e.g. obtaining efficiency measure(s). Step 604 includes a decision between using the putative offset L2’ as the current offset L2 or using another solution. In case it is chosen that L2’ can be chosen as the current offset L2 (transition 604A) then at step 606 the putative offset L2’ is used as the current offset L2. If, at step 604, it is chosen that it is preferable to not use L2’ as putative offset L2 (or at least to evaluate other candidates besides L2’), then (after the transition 604B) one of four alternative techniques may be used. One of the following alternative techniques can be chosen:
[0086] 1) to select the test area 312 around the putative offset L2’, like in Fig. 2B.
[0087] 2) to search for a current offset without restricting any error or subset (this may be chosen, for example, in case high performances are required).
[0088] 3) to progressively move away from the putative offset L2’ and end as soon as a candidate offset reaches a minimum value of the coding efficiency measure(s)
[0089] 4) to bypass the prediction.
[0090] The choice between these four options (or from a subset of the four of them) may be signalled (e.g. in the parameters of the side information 90). In some examples, however, only one of the four options is at disposal, and therefore it is not necessary to signal it.FH250310PEP-2026096817.DGCX 18
[0091] There can be several types of prediction, and the particular type of prediction chosen may be signalled in the bitstream (e.g. as side information). Any type of prediction may be used. The present technique
[0092] To address the abovementioned drawback of slow BM (block matching) or LTP (long term prediction or predictor) encoding, the following techniques are proposed:
[0093] • temporal reuse (e.g. Fig. 1) of L-search results for one block 301 and / or 302 in channel c in the next block 30 in channel c; in the next block 30, the encoder 10 may search for an “optimal” L (L2) only within a vicinity (e.g. interval 312, 312’, 312”) of the last L (L1 and / or L3)
[0094] • cross-channel reuse (e.g. Figs. 2A, 2B, 3A, 3B) of L-search results for one block 301 and / or 302 in channel c in the same block or in the same time coordinate (e.g. indicated with 30 in Figs. 2A and 2B) in the next (or generally a further, in coding order) channel c + 1; an encoder may search, for c + 1, for an “optimal” L (L2) only within a vicinity of the “optimal” L (L1 and / or L3) found for c (in a previous encoding step), or may entirely adopt the last L (of c) (e.g. L1 and / or L3) for the coding of said block (in further channel c + 1 ) • cross-block-split reuse (e.g. Fig. 4) of L-search results for a least 2 successive shorter blocks (e.g. sub-blocks) in channel c in the longer block 30 in c spanning the aggregated length of said at least 2 shorter blocks (303, 304), during a search for an optimal block partitioning among said shorter (303, 304) and longer (30) blocks; the encoder 10 may search, for the longer block 30 in c, for an “optimal” L only within a vicinity of the (at least 2) “optimal” L (e.g. L1, L3) found for the shorter blocks (303, 304) in c (in a previous partition-search step), or may readily adopt (one of the at least 2 of) the shorter-block L (e.g. L1 , L3) for the long-block coding.
[0095] In other words, encoder speed-up may be achieved by reusing search information in e.g. all three relevant directions in a block-adaptive multichannel coder: time, channel, and block partitioning direction. Detailed preferred embodiments for each of these three aspects are described in the following.
[0096] temporal reuse (e.g. Fig. 1)
[0097] Assume that the signal has an underlying periodicity in a channel c and a suitable offset L (e.g. L1) was found in a previous block 301. A fast encoder search can look for offset candidates in the neighborhood 312 of L (e.g. L1) as the offset value changes only marginally over time. Additionally, the encoder search can also use multiples of L or fractions (e.g. sub-multiples) of L (and search in their neighborhood 312) as there might be irregularities in the signal at offset L (e.g. L1) in a subsequent block (e.g. 30).FH250310PEP-2026096817.DOCX 19
[0098] cross-channel reuse (e.g. Figs. 2A, 2B, 3A, 3B)
[0099] Given that signals of a similar type are combined in consecutive channels (in a channel group), it can be assumed that the N channels are temporally aligned. Given that a block matching predictor with offset L (e.g. L1 , L3) was already found in channel c with 0 < c < N -1, the offset (e.g. L1 , L3) used for block 301 (or 302) can be used (e.g. for the other block 30) in subsequent channels c + i (in coding order) with 0 < i < N - c. The encoder can use L (e.g. L1 or L3, as L2’ and / or L2”) or multiples of L (e.g. L2’=m*L1 L2”=m*L3 with m natural number) or fractions of L (e.g. L2’=(1 / m)*L1 L2”=(1 / m)*L3, with m natural number) as a starting point for a faster block matching search and search in their neighborhood (e.g. 312, 312’, 312”). Alternatively, the offset L2 (and other parameters such as the filter) can be inferred as the parameter for the block matching predictor in subsequent channels without additional rate for transmitting the offset (and the filter).
[0100] cross-block-split-reuse (e.g. Fig. 4)
[0101] In case that multiple block sizes are allowed (like in Fig. 4), a recursive search is performed. Starting from the largest block size (e.g. the size of block 30), a split in 2 smaller blocks (also called sub-blocks, e.g. 303, 304) may be performed and the combined cost of the smaller blocks (e.g. 303, 304) is calculated before the cost for the larger block 30 is evaluated.
[0102] Therefore, when testing the larger block 30 the information from the smaller blocks (e.g. sub blocks 303, 304) is readily available.
[0103] Given a channel c, it is reasonable to assume that the already searched block matching predictors from the smaller blocks (e.g. 303, 304) in channel c are also fitting candidates for the larger block 30. Assuming the underlying signal is periodic, an optimal L (e.g. L2) could be a multiple or fraction of one of the offsets L (e.g. L1 , L3) from the smaller blocks (e.g. 303, 304) (e.g. L2’=(1 / m’)*L1, L2”=(1 / m”)*L3, with m’ and m” natural numbers (and with possibility of m’=m”), and / or L2’=(1 / m’)*L1, L2’’=(1 / m”)*L3, with m’ and m” natural numbers (and with possibility of m’=m”)). It may also be in the proximity (e.g. test area, interval) of the offset candidates. Additionally, it is also possible to test a combination (e.g. linear combination) of the offsets (e.g. L1, L3, or L2’, L2”) from the smaller blocks (e.g. 303, 304) if multiple hypotheses are allowed.
[0104] Note that reusing the results from the smaller blocks (e.g. sub blocks 303, 304) of a split is not limited to block matching predictors. The cross channel predictor can use also the available information from the smaller blocks (e.g. sub blocks 303, 304) and may use a combination of the prediction hypotheses.
[0105] SummarizationsFH250310PEP-2026096817.DOCX 20
[0106] Hence, it is possible to perform a prediction based on an offset L2 obtained from at least one putative offset L2' and / or L2” which may be the offset L1 and / or L3 used for a block 301, 302, 303, 304 previously subjected to encoding. Information on the offset L2 may therefore be signalled in the parameters 90, and the decoder 12 will perform the decoding by keeping into account the offset L2. The invention is transparent to the decoder 12, which therefore has no knowledge on the particular technique used for quickly obtaining the offset L2.
[0107] The technique may be repeated for a plurality of (e.g. all) blocks.
[0108] Aspects
[0109] Some aspects are discussed below.
[0110] A 1staspect provides an encoder for encoding a digital-time-varying signal, configured to encode the digital-time-varying signal in a data stream (16) in temporal blocks (30) by encoding a current temporal block (30) by:
[0111] predicting (62) the current temporal block (30) from an already encoded portion (310) of the digital-time-varying signal separated from the current temporal block (30) by a current offset (e.g. L2), by:
[0112] deriving, from at least one previous offset (e.g. L1, L3) used for predicting an already encoded (already predicted) temporal block (e.g. 301, 302) [e.g. the already encoded (already predicted) temporal block (e.g. 301, 302) can be in the vicinity of the current temporal block 30, e.g. temporally and / or channel-wise adjacent (or temporally and / or channel-wise co-located) to the current temporal block 30] at least one putative offset [e.g. L2’ and / or L2”] [e.g. the at least one putative offset L2’ and / or L2” may be obtained from the at least one previous offset (e.g. L1 and / or L2) used for predicting the already encoded temporal block (e.g. 301, 302)],
[0113] determining the current offset [L2] from the at least one putative offset (e.g. L1 and / or L2) through coding efficiency measure(s) [e.g. the coding efficiency measure(s) being derived from at least one putative prediction signal (e.g. L2’ and / or L2”), or a processed version thereof, obtained using the at least one putative offset][e.g., the coding efficiency measure(s) may be evaluated between the current temporal block 30 and the already encoded portion 310 for different candidate offsets (which may have been identified from the putative current offset), so as to retrieve the current offset L2 (which may be, for example, the one minimizing a cost function which is, or is associated with, the coding efficiency as measured through the coding efficiency measure(s))], and
[0114] deriving the prediction signal (64) from the current offset [L2];
[0115] deriving a prediction residual signal (66) by comparing the digital-time-varying signal with the prediction signal (64); andFH250310PEP-2026096817.DOCX 21
[0116] encoding (70), in the data stream (16), prediction residual coefficients (76) which represent the prediction residual signal (66) and signaling information indicating the current offset [L2],
[0117] According to a 2ndaspect e.g. when referring back to the 1staspect, the current offset L2 [and / or the at least one previous offset L1 and / or L3 and / or the putative current offset L2’ and / or L2”] may be or may include a temporal offset [e.g. measurable in horizontal in Fig. 5], According to a 3rdaspect e.g. when referring back to the 2ndor 3rdaspect, the offset (e.g. current offset) L2 may be or may include a channel-wise offset [e.g. measurable in vertical in Fig. 5].
[0118] According to a 4thaspect e.g. when referring back to any one of the 1stto 3rdaspects, the encoder 10 may evaluate coding efficiency measure(s) of multiple candidate current offsets [e.g. the candidate current offsets being taken from a particular test area 312, 312”, 312”, e.g. the test area containing the at least one putative offset L2’ and / or L2”, and / or the candidate current offsets being taken from two or more than two previous offsets, e.g. associated with two or more than two already encoded temporal blocks], so as to choose the current offset L2 among the multiple candidate current offsets based on the coding efficiency measure(s) [e.g. the chosen current offset may be the one maximizing the coding efficiency measure(s)][e.g. the candidate current offsets may be, in some examples, putative current offsets, or offsets within at least a test area (312, 312”, 312”) which may be identified by the at least one putative current offset][the coding efficiency measure(s) may be evaluated between the current temporal block 30 and the already encoded portion 310 for different candidate offsets (which may have been identified from the putative current offset L2’ and / or L2”), so as to retrieve the current offset L2 (which may be, for example, the one minimizing a cost function which is, or is associated with, the coding efficiency as measured through the coding efficiency measure(s))].
[0119] According to a 5thaspect e.g. when referring back to the 4thaspect, the encoder 10 may be configured to define a test area (e.g. 312, 312’, 312”) [or more in general a proper subset of the set of possible values of the offset] including (e.g. centered in or containing), or at least delimited by, the at least one putative offset L2’ and / or L2”, so as to choose the current offset L2 among the multiple candidate current offsets based on the coding efficiency measure(s) [e.g. the chosen current offset may be the one maximizing the coding efficiency measure][the test area may be identified from the putative current offset].
[0120] According to a 6thaspect e.g. when referring back to the 5thaspect, the test area (e.g. 312, 312’, 312”) may be or may include an interval [e.g. in the temporal lag dimension and / or in the channel-wise offset dimension] which has interval ends at pre-defined distances [e.g. in the temporal lag dimension and / or in the channel-wise offset dimension] from the at least one putative offset [e.g. L2' and / or L2”] [e.g. the interval ends are positioned at pre-definedFH250310PEP-2026096817.DOCX 22
[0121] distances may be fixed and / or independent of the values of L2’ and / or any value of the signal to be encoded].
[0122] According to a 7thaspect e.g. when referring back to the 5lhor 6thaspect, the test area may be or may include an interval [e.g. in the temporal lag dimension and / or in the channel-wise offset dimension] which has interval ends at pre-defined distances [e.g. in the temporal lag dimension and / or in the channel-wise offset dimension] from the at least one putative offset [e.g. L2’ and / or L2”] [e.g. the interval ends are positioned at distances which are controlled by some values, e.g. by the resolution, e.g.: the finer the temporal and / or channel resolution, the larger the distance in term of samples and / or channels].
[0123] According to an 8thaspect e.g. when referring back to any one of the 1stto 7thaspects, the encoder 10 may be configured to evaluate the coding efficiency measure(s) in correspondence of the at least one putative offset (e.g. L2’ and / or L2”) [e.g. an example of which being in Fig. 6] and, based on the coding efficiency measure(s), decide whether to adopt the at least one putative offset (e.g. L2’, L2”) as current offset or to use an alternative technique [this aspect, as well as its dependent aspects, is particularly adapted to the examples of Figs. 2A-4][e.g. evaluating the coding efficiency measure(s) in correspondence of the at least one putative offset may imply evaluating the coding efficiency measure(s) between the current temporal block and the already encoded portion in correspondence of the at least one putative offset].
[0124] According to a 9thaspect e.g. when referring back to the 8thaspect when depending on at least the 6thaspect, the encoder 10 may adopt the alternative technique of defining the test area 312, 312’, 312” as the interval containing the at least one putative offset (e.g. L2’ and / or L2”), so as to choose the current offset L2 among the multiple candidate current offsets based on the coding efficiency measure(s) within the test area [an example of this alternative technique is referred to with 1 ) in Fig. 6],
[0125] According to a 10thaspect e.g. when referring back to the 8thor 9thaspect, the encoder may be configured to adopt the alternative technique of searching for the current offset L2 without restricting to a test area (e.g. 312, 312’, 312”) defined from the at least one putative offset (e.g. L2’ and / or L2”) [an example of this alternative technique is referred to with 2) in Fig. 6], According to an 11thaspect e.g. when referring back to the 8thor 9thor 10thaspect, the encoder may be configured to adopt the alternative technique of deriving coding efficiency measure(s) of candidate offsets progressively moving away from the at least one putative offset (e.g. L2’ and / or L2”) [e.g., the process may end as soon as a candidate offset reaches a minimum value of the coding efficiency measure(s) (or, said in other terms, as soon as a candidate offset is retrieved having the coding efficiency measure(s) over a predetermined threshold)] [an example of this alternative technique is referred to with 3) in Fig. 6],
[0126] fitFH250310PEP-2026096817.DOCX 23
[0127] According to an 11athaspect e.g. when referring back to the 8thor 9thor 10thor 11thaspect, the encoder may be configured to adopt the alternative technique of skipping the prediction. According to a 12thaspect e.g. when referring back to any one of the 1stto 11 athaspects, the encoder may be configured to evaluate the coding efficiency measure(s) in correspondence to at least a first previous offset [e.g. L1 in Fig. 3A; it may be a first putative current offset] used for predicting a first already encoded temporal block and the coding efficiency measure(s) in correspondence to a second previous offset [e.g. L3 in Fig. 3A; it may be a second putative current offset] used for predicting a second already encoded temporal block, so as to derive the current offset [e.g. by choosing one of the two offsets].
[0128] According to a 13thaspect e.g. when referring back to the 12thaspect, the encoder may be configured to derive the current offset from the one, among the at least the first previous offset and the second previous offset, which maximizes the coding efficiency measure(s). According to a 14thaspect e.g. when referring back to the 12thaspect, the encoder may be configured to derive the current offset as a linear combination between at least two of the at least the first previous offset and the second previous offset.
[0129] According to a 15thaspect e.g. when referring back to any of the 12thto 14thaspect when depending on at least the 6thaspect, the encoder may be configured to define at least a first test area from the first previous offset and a second test area from the second previous offset, so as to choose, as the current offset, the one maximizing the coding efficiency measure(s).
[0130] According to a 16thaspect e.g. when referring back to any of the 1stto 15thaspects, the encoder may be configured to split the current temporal block into at least two temporal sub blocks, to retrieve, for each of the at least two temporal sub blocks, a putative offset (e.g. L2’ and / or L2”), and to adopt, as the current offset, an offset derived from at least one of the at least two temporal sub blocks [an example is in Fig. 4],
[0131] According to a 17thaspect when referring back to the 16thaspect, the encoder may be configured to define the current offset by evaluating the coding efficiency measure(s) in correspondence to the respective putative offsets (e.g. L2’ and / or L2”).
[0132] According to an 18thaspect e.g. when referring back to the 17thaspect, the encoder may be configured to define the current offset by evaluating the coding efficiency measure(s) in correspondence to the respective putative offsets (e.g. L2’ and / or L2") for both the sub blocks and the current temporal block.
[0133] According to a 19thaspect e.g. when referring back to any of the 1stto 18thaspects, t he encoder may be configured to derive the at least one putative offset [e.g. L2’ and / or L2”] as the same as the at least one previous offset.
[0134] According to a 20thaspect e.g. when referring back to any of the 1stto 19thaspects, the encoder may be configured to derive the at least one putative offset [e.g. L2’ and / or L2”] as aFH250310PEP-2026096817.DOCX 24
[0135] scaled value, through a pre-defined [e.g., fixed, independent of L1] scaling value, derived from the at least one previous offset [e.g. L1] for the already encoded (already predicted) temporal block [e.g. L2’=m*L1, where m is a pre-defined constant (e.g. fixed, e.g. independent of L1 ) value, e.g. a real value, a relative value, a rational value, or a natural value (c may be greater than 0; in some cases, m may be less than 1 )].
[0136] According to a 21staspect e.g. when referring back to any of the 1stto 20thaspects, the encoder may be configured to derive the at least one putative offset [e.g. L2’ and / or L2”] as a scaled value, through a pre-defined [e.g., fixed, independent of L1] scaling value, derived from the at least one previous offset [e.g. L1 ] for the already encoded (already predicted) temporal block, where the scaling value is either a natural number greater than 0, or the multiplicative reciprocal of a natural number greater than 0 [e.g. L2’ can be a multiple of L1 (e.g. L2’=m*L1 with m being integer, or more in particular natural), or L2’ can be a submultiple of L1 (e.g. L2’= L1 / m with m being in the form of integer, or more in particular natural)].
[0137] According to a 22ndaspect e.g. when referring back to any of the 1stto 21staspects, the encoder 10 may be configured, in case of retrieving multiple candidate offsets which both satisfy conditions for being the current offset, to define the current temporal block as a multi¬ hypothesis block with multiple current offsets.
[0138] According to a 23rdaspect e.g. when referring back to any of the 1stto 22ndaspects, the coding efficiency measure(s) may include similarity measure(s) between the current temporal block 30 and the already encoded (already predicted) portion 310 separated from the current temporal block 30 by the putative current block L2’ and / or L2” or a candidate current block [e.g. identified from the putative current block]
[0139] According to a 24thaspect e.g. when referring back to any of the 1stto 23rdaspects, the encoder 10 may be configured to evaluate a cost function derived from the coding efficiency measure(s) at the putative current offset L2’ and / or L2” and / or at least one candidate current offset identified from the putative current offset L2’ and / or L2”, so as to choose as the current offset L2 the offset, among the putative and / or the at least one candidate current offset, which minimizes the cost function [the cost function is higher in case e.g. of high dissimilarity, and is low in case of high similarity],
[0140] A 25thaspect provides a method for encoding a digital-time-varying signal, configured to encode the digital-time-varying signal in a data stream (16) in temporal blocks (30) by encoding a current temporal block (30) by:
[0141] predicting (62) the current temporal block (30) from an already encoded portion of the digital-time-varying signal separated from the current temporal block (30) by a current offset, by:FH250310PER-2026096817.DOCX 25
[0142] deriving, from at least one previous offset used for predicting an already encoded temporal block [e.g. the already encoded temporal block can be in the vicinity of the current temporal block, e.g, temporally and / or channel-wise adjacent (or temporally and / or channel-wise co-located) to the current temporal block] at least one putative offset [e.g. L2’ and / or L2”] [e.g. the at least one putative offset may be obtained from the at least one previous offset used for predicting the already encoded temporal block],
[0143] determining the current offset [L2] from the at least one putative offset through coding efficiency measure(s) [e.g. the coding efficiency measure(s) being derived from at least one putative prediction signal, or a processed version thereof, obtained using the at least one putative offset][e.g., the coding efficiency measure(s) may be evaluated between the current temporal block and the already encoded portion for different candidate offsets (which may have been identified from the putative current offset), so as to retrieve the current offset (which may be, for example, the one minimizing a cost function which is, or is associated with, the coding efficiency as measured through the coding efficiency measure(s))], and
[0144] deriving the prediction signal (64) from the current offset [L2];
[0145] deriving a prediction residual signal (66) by comparing the digital-time-varying signal with the prediction signal (64); and
[0146] encoding (70), in the data stream (16), prediction residual coefficients (76) which represent the prediction residual signal (66) and signaling information indicating the current offset [L2],
[0147] A 26thaspect provides a non-transitory storage unit storing instructions which, when executed by a processor, cause the processor to perform the method when referring back to the 25thaspect.
[0148] Implementation alternatives
[0149] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus.
[0150] Depending on certain implementation requirements, embodiments of the present technique can be implemented in hardware or in software. The implementation can be performed usingFH250310PEP-2026096817.DQCX 26
[0151] a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
[0152] Some embodiments according to the present technique comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0153] Generally, embodiments of the present technique can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine-readable carrier.
[0154] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine-readable carrier.
[0155] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0156] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitionary. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.
[0157] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0158] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0159] A further embodiment according to the present technique comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example,FH250310PEP-2026096817.DOCX 27
[0160] be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver. In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
[0161] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0162] The apparatus described herein, or any components of the apparatus described herein, may be implemented at least partially in hardware and / or in software.
[0163] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0164] The methods described herein, or any components of the apparatus described herein, may be performed at least partially by hardware and / or by software.
[0165] The above-described embodiments are merely illustrative for the principles of the present technique. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.
Claims
FH250310PEP-2026096817.DOCX 28Claims1. Encoder (10) for encoding a digital-time-varying signal, configured to encode the digital-time-varying signal in a data stream (16) in temporal blocks (30) by encoding a current temporal block (30) by:predicting (62) the current temporal block (30) from an already encoded portion of the digital-time-varying signal separated from the current temporal block (30) by a current offset (L2), by:deriving, from at least one previous offset (L1, L3) used for predicting an already encoded temporal block (301, 302) at least one putative offset (L2’, L2”), determining the current offset (L2) from the at least one putative offset (L2’, L2”) through coding efficiency measure(s), andderiving the prediction signal (64) from the current offset (L2);deriving a prediction residual signal (66) by comparing the digital-time-varying signal with the prediction signal (64); andencoding (70), in the data stream (16), prediction residual coefficients (76) which represent the prediction residual signal (66) and signaling information indicating the current offset (L2).
2. The encoder of any of the preceding clams, wherein the current offset (L2) is or includes a temporal offset.
3. The encoder of any of the preceding clams, wherein the offset (L2) is or includes a channel-wise offset.
4. The encoder of any of the preceding claims, configured to evaluate coding efficiency measures of multiple candidate current offsets, so as to choose the current offset (L2) among the multiple candidate current offsets based on the coding efficiency measures.
5. The encoder of claim 4, configured to define a test area (312, 312’, 312”) including, or at least delimited by, the at least one putative offset ( L2’ , L2”), so as to choose the current offset (L2) among the multiple candidate current offsets based on the coding efficiency measures.
6. The encoder of claim 5, wherein the test area (312, 312’, 312”) is an interval which has interval ends at pre-defined distances from the at least one putative offset (L2’, L2”).FH250310PEP-2026096817.DOCX 297. The encoder of any of the preceding claims, configured to evaluate a cost function derived from the coding efficiency measure(s) at the putative current offset (L2’) and / or at least one candidate current offset identified from the putative current offset, so as to choose as the current offset (L2) the offset, among the putative and / or the at least one candidate current offset, the one which minimizes the cost function.
8. The encoder of any of the preceding claims, configured to evaluate the coding efficiency measure(s) in correspondence of the at least one putative offset (L2’, L2”) and, based on the coding efficiency measure, decide (604) whether to adopt (604) the at least one putative offset (L2’> L2”) as current offset (L2) or to use an alternative technique (608, 610).
9. The encoder of claim 8 when depending on at least claim 6, configured to adopt (608) the alternative technique of defining the test area (312, 312’, 312”) as the interval containing the at least one putative offset (L2’, L2”), so as to choose the current offset (L2) among the multiple candidate current offsets based on the coding efficiency measures within the test area.
10. The encoder of claim 8, configured to adopt (608) the alternative technique of searching for the current offset (L2) without restricting to a test area defined from the at least one putative offset (L2’, L2”).
11. The encoder of claims 8, configured to adopt (608) the alternative technique of deriving coding efficiency measures of candidate offsets progressively moving away from the at least one putative offset (L2’, L2”).
12. The encoder of claim 8, configured to adopt (610) the alternative technique of skipping the prediction.
13. The encoder of any of the preceding claims, configured to evaluate the coding efficiency measure(s) in correspondence to at least a first previous offset (L3) used for predicting a first already encoded temporal block (302) and the coding efficiency measure in correspondence to a second previous offset (L1) used for predicting a second already encoded temporal block (301 ), so as to derive the current offset (L2).
14. The encoder of claim 13, configured to derive the current offset (L2) from the one, among the at least the first previous offset (L3) and the second previous offset (L1 ), which maximizes the coding efficiency measure(s).4FH250310PEP 2026096817. DOCX 3015. The encoder of claim 13, configured to derive the current offset (L2) as a linear combination between at least two of the at least the first previous offset (L3) and the second previous offset (L1).
16. The encoder of any of claims 13-15 when depending on at least claim 6, configured to define at least a first test area (312’) from the first previous offset and a second test area (312”) from the second previous offset, so as to choose, as the current offset (L2), the one maximizing the coding efficiency measure.
17. The encoder of any of the preceding claims, configured to split the current temporal block (30) into at least two temporal sub blocks (303, 304), to retrieve, for each of the at least two temporal sub blocks (303, 304), a putative offset (L2’, L2”), and to adopt, as the current offset (L2), an offset derived from at least one of the at least two temporal sub blocks.
18. The encoder of claim 17, configured to define the current offset (L2) by evaluating the coding efficiency measures in correspondence to the respective putative offsets (L2’, L2”).
19. The encoder of claim 18, configured to define the current offset (L2) by evaluating the coding efficiency measures in correspondence to the respective putative offsets (L2’, L2”) for both the sub blocks ( L2’, L2”) and the current temporal block (30).
20. The encoder of any of the preceding claims, configured to derive the at least one putative offset (L2’, L2”) as the same as the at least one previous offset (L1 , L3).
21. The encoder of any of the preceding claims, configured to derive the at least one putative offset (L2’;L2”) as a scaled value, through a pre-defined scaling value, derived from the at least one previous offset for the already encoded temporal block.
22. The encoder of any of the preceding claims, configured to derive the at least one putative offset (L21, L2”) as a scaled value, through a pre-defined scaling value, derived from the at least one previous offset for the already encoded temporal block, where the scaling value is either a natural number greater than 0, or the multiplicative reciprocal of a natural number greater than 0.FH250310PEP-2026096817.DOCX 3123. The encoder of any of the preceding claims, configured, in case of retrieving multiple candidate offsets which both satisfy conditions for being the current offset (L2), to define the current temporal block as a multi-hypothesis block with multiple current offsets.5 24. The encoder of any of the preceding claims, wherein the coding efficiency measure(s) include(s) similarity measure(s) between the current temporal block and the already encoded portion separated from the current temporal block by the putative current block or a candidate current block10 25. A method for encoding a digital-time-varying signal, configured to encode the digital- time-varying signal in a data stream (16) in temporal blocks (30) by encoding a current temporal block (30) by:predicting (62) the current temporal block (30) from an already encoded portion of the digital-time-varying signal separated from the current temporal block (30) by a current offset 15 (L2), by:deriving, from at least one previous offset used for predicting an already encoded temporal block at least one putative offset ( L2’ , L2”),determining the current offset (L2) from the at least one putative offset (L2’, L2”) through coding efficiency measure(s), and20 deriving the prediction signal (64) from the current offset (L2);deriving a prediction residual signal (66) by comparing the digital-time-varying signal with the prediction signal (64); andencoding (70), in the data stream (16), prediction residual coefficients (76) which represent the prediction residual signal (66) and signaling information indicating the current 25 offset (L2).
26. A non-transitory storage unit storing instructions which, when executed by a processor, cause the processor to perform the method of claim 25.