Efficient use of block transformation in particular in biomedical waveform coding
Patent Information
- Application Number
- PCT/EP2026/058046
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-24
- Filing Date
- 2026-03-23
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026058046_01102026_PF_FP_ABST
Abstract
Description
[0001] FH250308PEP-2026096718. DOCX 1
[0002] Efficient use of block transformation in particular in biomedical waveform coding Description
[0003] Embodiments comprise decoders, encoders, methods for decoding, methods for encoding, computer programs and data streams for an efficient fast encoding in waveform codecs us¬ ing downsampled block prediction.
[0004] Introductory remarks:
[0005] In the following, different inventive embodiments and aspects will be described.
[0006] Also, further embodiments will be defined by the enclosed aspects.
[0007] It should be noted that any embodiments as defined by the aspects can be supplemented by any of the details (features and functionalities).
[0008] Also, the embodiments described can be used individually, and can also be supplemented by any of the features in another section, or by any feature included in the aspects.
[0009] Also, it should be noted that individual aspects described herein can be used individually or in combination. Thus, details can be added to each of said individual aspects without adding details to another one of said aspects.
[0010] Moreover, features and functionalities disclosed herein relating to a method can also be used in an apparatus (configured to perform such functionality). Furthermore, any features and functionalities disclosed herein with respect to an apparatus can also be used in a corresponding method. In other words, the methods disclosed herein can be supplemented by any of the features and functionalities described with respect to the apparatuses.
[0011] Moreover, features and functionalities disclosed herein relating to a method, in particular an encoding method can also be used in a data stream or bitstream (e.g. defining a respective data stream or bitstream element). Furthermore, any features and functionalities disclosed herein with respect to an apparatus can also be used in a corresponding data stream, e.g. as a resulting data stream as providing by said encoder. In other words, the data streams dis¬ closed herein can be supplemented by any of the features and functionalities described with respect to the apparatuses and methods.
[0012] Also, any of the features and functionalities described herein can be implemented in hard¬ ware or in software, or using a combination of hardware and software, as will be described in the section “Implementation alternatives”.
[0013] Faster encoding in waveform codecs using downsampled block prediction in the prior artFH250308PEP-2026096718. DOCX 2
[0014] Digital waveform codecs, like biophysical (e. g., medical), geophysical (e. g., seismic), or acoustic (e. g., audio) codecs, often employ block prediction methods to improve coding per¬ formance. In such codecs, a difference between the per-block / channel waveform signal and a corresponding (usually configurable or parametrizable) prediction signal is formed, and the block / channel-wise residual signal resulting therefrom is quantized, entropy coded, and transmitted. At the decoder, the entropy decoded and “dequantized” residual block / channel signal is generated, and the same block prediction signal is added thereon to reconstruct the waveform signal in the initial domain. Note that the prediction signal usually represents an already reconstructed part of the waveform signal in the given -or sometimes, a preceding (in coding order)-channel, of equal length as the current waveform signal block to be pre¬ dicted, but offset (relocated) to the current block location.
[0015] Recently, standardization of a biomedical waveform coder has been started, and the current draft standard features cross-channel (CC) and block matching (BM) predictors, the latter also known as long-term predictor (LTP). Both predictors operate in time domain and are controlled by an offset parameter C (channel offset) resp. L (time offset), signalled from en¬ coder to decoder side via a data stream. A major drawback of CC, BM, or LTP prediction is a large search space for finding an “optimal” value for C or L, i.e., one resulting in best prediction performance with minimized block residual energy. In other words, the search for an “optimal” offset value, possibly across hundreds of offset candidates in the channel or time vicin¬ ity of the block, can slow down encoding notably.
[0016] Summary
[0017] According to an aspect, there is provided an encoder for encoding a digital-time-varying signal, configured to encode the digital-time-varying signal into a data stream in temporal blocks by encoding a current temporal block by:
[0018] predicting the current temporal block, to derive a prediction signal,
[0019] deriving a prediction residual signal by comparing the digital-time-varying signal with the prediction signal; and
[0020] encoding, into the data stream, prediction residual coefficients which represent the prediction residual signal,
[0021] configured, in predicting the current temporal block, to:
[0022] downsample an already encoded portion of the digital-time-varying signal from an original sample resolution to a downsampled sampling resolution, to derive a downsampled version of the already encoded portion;FH250308PEP-2026096718. DOCX 3
[0023] determine a similarity between the current temporal block and the already encoded portion for different offset values at the downsampled sample resolution, to ob¬ tain a final offset value for the current temporal block; and
[0024] encode the final offset value for the current temporal block into the data stream.
[0025] Figures
[0026] Figs. 1 and 2 show operations according to examples.
[0027] Fig. 3 shows an example of obtaining the offset value.
[0028] Fig. 4 shows an example of encoder and decoder.
[0029] Examples
[0030] The description is extended in the following by the presentation of implementation examples. Before this, however, the description proceeds with a presentation of a possible framework or codec into which the embodiments described above and below as well as the examples described further below may be built into. Many details described in this framework are, how- ever, optional when being combined with any of the above and below or subsequently de¬ scribed embodiments. To be more precise, the framework is described with respect to Fig. 4 which shows an encoder for encoding a multi-channel digital signal 14 into a datastream 16 as well as decoder 12 for decoding the multi-channel digital signal 14 from datastream 16. This description of Fig. 4 shall be seen as a presentation of new embodiments of the present application which result when combining any of the embodiments described above and be¬ low or any of the examples described subsequently or even any of the claimed subject matters is combined with the decoder 12 or encoder 10 of Fig. 4 either by adopting all de-tails / functionalities described with respect to Fig. 4 or with leaving-out some of the de-tails / functionalities described with respect to Fig. 4. Sometimes such “optional” features of Fig. 4 are explicitly identified as being optional with respect to the combination of the previously and subsequently described embodiments, but the just-mentioned possible combinations of the previously / subsequently explained embodiments with the description of Fig. 4 shall not be restricted to the these explicitly identified variations of Fig. 4 in terms of leaving- out certain features.
[0031] In Fig. 4, the multi-channel digital signal 14 is illustrated by way of an array of samples with the samples being illustrated as small squares 18. Each line / row corresponds to a certain channel of the multi-channel digital signal 14. Each channel of signal 14 may have associ¬ ated therewith a respective channel ID and Fig. 4 shows these channels as being ordered according to their channel ID along vertical axis 20 which, thus, corresponds to a “source” channel axis 20. The horizontal axis 22 corresponds to time so that samples 18 forming one
[0032] AFH250308PEP-2026096718. DOCX 4
[0033] column, or being horizontally aligned, are samples belonging to one common time instant. Such set / column of temporally co-located samples 18 is illustrated in Fig. 4 at 24.
[0034] Each channel, thus, forms a digital time-varying signal or time / amplitude or time-to-amplitude signal. The multi-channel digital signal m might have been obtained by at least one of Electrocardiography, Electroencephalography, Electromyography or seismic measurement. Dif¬ ferently speaking, the multi-channel digital signal might be a bio-physiological waveform data such as an electroencephalography (EEG) signal, an electrocardiogram (ECG), or an elec¬ tromyography (EMG) signal, or seismic waveform data. However, each channel / signal might alternatively be another sort of waveform signal data such as scalar media data such as an audio signal and the signal 14 might be a multi-channel audio signal.
[0035] Fig. 4 illustrates the option according to which signal 14 is not coded directly, i.e., in the origi¬ nal domain 26, but in a so-called “coded domain” 28 which might differ from the original do¬ main 26 by one or more of 1) channel transformation, 2) channel permutation and 3) temporal mutual channel alignment. The channel transformation, if applied, transforms, per sam¬ ple time instant, a set or column 24 of samples from domain 26 to domain 28. Thus, in do¬ main 28, the sample pitch and the time axis is the same as in domain 26, but the meaning of the channels is different, i.e., the “source” channels of domain 26 become transformed channels in domain 28. Accordingly, the vertical axis in Fig. 4 for domain 28 is denoted as 32. Note that the channel transformation might leave the number of channels unchanged so that there is the same number of channels in domain 26 as well as domain 28, but different ap¬ proaches are also possible. Generally, the channel transformation would aim at reducing redundancy and trying to condense the channels’ energy onto a fewer number of channels in domain 28. As said, the channel transformation is optional. Accordingly, in general terms, the channels in domain 28 are called “coded channels” in order to distinguish them from the “original” or “source” channels of digital signal 14 in domain 26. The permutation is also optional and may be used in combination with, or without, the channel transformation. If used in combination with the channel transformation, the permutation may be performed prior to and / or or subsequent to the channel transformation in order to permute / sort the source channels prior to transformation and the coded channels subsequent to the channel transformation. The channel transformation might be a DCT, DST, FFT or any other transformation. The temporal mutual alignment is also optional and might be seen as a constant temporal alignment between the source channels or the coded channels.
[0036] The module in encoder 10 performing the one or more of channel transformation, channel permutation and temporal mutual alignment is indicated in Fig. 4 as block 34. Side information 36 might be used in order to signal information on one or more of the following: 1 ) The channel transformation used, 2) information on the permutation(s) among the sourceFH250308PEP-2026096718. DOCX 5
[0037] channels and / or coded channels and 3) information on the mutual temporal alignment / delays between the source channels or coded channels wherein the temporal mutual alignment might be restricted to full sample precision. A corresponding block 38 in decoder 12 performs the reverse step, i.e., performs one or more of: 1) a channel retransformation, 2) a re-permu-tation of the source channels and / or coded channels and 3) a temporal re-alignment of the source channels or coded channels. Note, that if no channel transformation takes place, the coded channels are, in fact, equal to the source channels except for being temporally mutu¬ ally aligned or being differently sorted due to permutation. Block 38 might be controlled by the before-mentioned side information 36.
[0038] Thus, the “actual coding” relates to the coded channels in domain 28. In the coded domain 28, the coded channels are depicted in Fig. 4 as lines or rows of samples 40, each extending along time axis 22, the coded channels being depicted one on top of the other along coded channel axis 32 - potentially ordered according to a coded channel ID they have associated therewith - so as to result into an array of samples 40. Again, although Fig. 4 depicts the case that the number of source channels equals the number of coded channels, the number might be different. Further, if channel transformation is used, while there is no longer a clear association between source channels on the one hand and coded channels on the other hand, the temporal association remains: For each temporally co-located samples 24, there is a corresponding temporally co-located set 42 of samples 40 of the coded channels, wherein the set 42 in domain 28 is a column and might be a set of horizontally mutually offset sam¬ ples in case of, and according to, the mutual temporal alignment, if applied. In case of Fig. 4, it has been assumed that no such temporal alignment took place so that both sets 42 and 24 are pure columns in the time / channel representation.
[0039] The actual coding is done in units of so-called temporal blocks 30. The term “block” or “tem¬ poral block” 30 is used so as to denote both a temporal portion of the multi-channel signal in domain 28, i.e., the set of coded channels, as well as a temporal portion of a certain coded channel. That is, for each temporal block 30, each coded channel has a temporal block such as block 140 depicted for some temporal block 30c and same are mutually co-located. The coding is done sequentially along these blocks 140, by following a coding / decoding order, which traverses the blocks 140 temporal block 30 by temporal block 30 with traversing temporally co-located blocks of the coded channels along a channel order corresponding to the order of the coded channels along axis 32. This coding / decoding order is illustrated in Fig. 4 at 60. That is, in case of temporal block 140 being the block currently to be coded / decoded, the previously decoded / encoded temporal blocks include all preceding temporal blocks of all coded channels as well as the temporally co-located temporal blocks of coded channels preceding the coded channel 92 of temporal block 140 in channel order. These previously coded / decoded temporal blocks and their samples are illustrated in Fig. 4 by way of shading.FH250308PEP 2026096718. DOCX 6
[0040] In this regard, note that in Fig. 4, merely one temporal block 140 has been illustrated explic¬ itly in order to reduce the complexity of Fig. 4. Thus, in the specification herein, reference sign 140 is sometimes used to indicate the currently encoded / decoded temporal block or to stand representatively for all temporal blocks. Further, as depicted in Fig. 4, the partitioning of signal 14 into temporal blocks 30 and 140, respectively, might be done in a manner so that these blocks 30 and 140, respectively, are non-overlapping.
[0041] The actual coding in units of the temporal blocks 140 is performed predictively. That is, the encoder 10 comprises a block predictor 62 which predicts the samples of the currently coded temporal block 140, thereby yielding a prediction signal 64, and the prediction residual 66 formed by a subtraction between the actual sample values of temporal block 140 and the predicted samples of prediction signal 64 formed at a subtractor 68 is coded into the datastream 16 by residual coder 70. The residual coding in residual coder 70 may, or may not, involve a coding error by means of quantization. In any case, block predictor 62 uses the reconstructable version as being available by previously coded temporal blocks in order to obtain the prediction signal 64. This reconstructable version 72 might be derived at encoder 10 by means of a residual decoder 74 which reverses, potentially under coding loss, such as quantization, e.g. by means of dequantization, the residual signal 76 as coded into datastream 16, and an adder 78 which sums-up prediction signal 64 and the reconstructable residual signal 80 as obtained by residual decoder 74. To be more precise, let’s call the channel-individual temporal blocks 140 subblocks with temporally collocated subblocks of all channels forming a temporal block 30. Then, the prediction in module 62 or, to be more precise, the prediction at encoder and decoder, is performed in units of the subblocks 140, i.e. subblock wise. The encoder is free to choose different prediction modes for the subblocks within one block 30. As explained in more detail herein, within one block 30, one subblock 140 may be predicted based on one or more subblocks previously - according to the decoding order 60 - en / decoded within this block 30, while another subblock 140 within that block 30 might be coded / decoded based on the previously en / decoded subblock 140 of the same channel (but within the previous block 30). The transform residual en / decoding is then per¬ formed subblock wise by use of a one-dimensional transform signaled in the data stream as described hereinbelow.
[0042] The decoder 12 decodes the coded channels from data stream 16 in a corresponding manner, i.e., in units of the temporal blocks 30 or in temporal blocks 140, respectively, and using predictive decoding. To this end, the decoder 12 comprises a residual decoder 82, an adder 84 and a block predictor 86 which correspond to, and are mutually connected in the same manner as, elements 74, 78 and 62 of encoder 10. That is, the residual decoder 82 derives from the residual signal 76 in data stream 16 the reconstructable residual signal 80 for a currently decoded temporal block 140 which is then subject to addition with prediction signal 64FH250308PEP-2026096718. DOCX 7
[0043] derived by block predictor 86 for temporal block 140 on the basis of the reconstructed ver- sion 72 of previously decoded temporal blocks at adder 84. The output of adder 84, thus, yields the reconstructed version 72 of the currently decoded temporal block 140 and be- comes part of the pool of already decoded samples of previously decoded temporal blocks 5 when the temporal blocks of the coded channels are, in this manner, traversed along cod- ing / decoding order 60 so as to reconstruct the coded channels in the coded domain 28. Note that the above and below description concentrated on the so-called sample prediction where samples of a current block 140 are predicted based on reconstructed samples of one or more previously decoded blocks, but coding inter dependencies, namely intra-channel and 10 inter-channel coding dependencies may be exploited not only in terms of sample prediction, but also in terms of other coding tools involving, for instance, parameter prediction and / or context derivation.
[0044] In order to enable a high degree of random access capability, some of the temporal blocks 30 may be coded in a random access manner meaning that the coded channels therein are 15 coded independent from previous temporal blocks 30. Imagine, for instance, that temporal blocks 30b and 30e are random access temporal blocks. Then, none of the temporal channel blocks 140 in temporal block 30b as well as 30e would depend on any preceding temporal block 140 and no coding dependency would cross these temporal blocks 30b and 30e, that is no temporal block 140 within any of temporal block 30b-30d would be coded depending on 20 any block 140 temporally preceding temporal block 30b, and no temporal block 140 within any of temporal block 30e and following would be coded depending on any block 140 tempo¬ rally preceding temporal block 30e.
[0045] Thus, in other words, coding dependencies are restricted so as to not reach-out beyond the border of a random access temporal block 30b and 30e towards any preceding temporal 25 block 30. Such restriction might also hold for intermediate temporal blocks 30c to 30d be¬ tween random access temporal blocks 30b and 30e in that same may not depend on any temporal block preceding the leading one among the random access temporal blocks 30b and 30e, here block 30b. Accordingly, leading temporal borders of the random access temporal blocks 30b and 30e are indicated by bold lines in Fig. 4. In a variant, the restriction is 30 not valid for all en / decoding stages. For instance, while the grouping might hold true for pre¬ diction, but the residual en / decoding dependencies might cross borders between channel groups. It might be the case, for instance, that for the entropy coding and decoding, all channels are coded jointly, i.e. using a single arithmetic coding engine, but that for the sake of prediction and reconstruction, the channels are grouped as described into independent 35 groups such that, after entropy decoding, each such group can be reconstructed completelyFH250308PEP 2026096718. DOCX e
[0046] independently from each other group. This means that no prediction of sample values or any other information is supported between different channel groups.
[0047] Further, it might be that the coding of the coded channels also interrupts or restricts interchannel dependencies. For example, one or more of the coded channels might be coded as random access coded channels so that same do not use inter-channel dependencies, but merely intra-channel dependencies. The restriction of inter-channel coding dependencies might follow the channel order 32: that is, coding of these random access coded channels and the intermediate coded channels therebetween would be restricted so as to not reach- out beyond such a random access coded channel toward any coded channel preceding that random access coded channel in channel order along axis 32. Two such random access coded channels 88a and 88b and the resulting inter-channel dependency borders are illus¬ trated in Fig. 4. Note that the restriction of inter-channel dependencies might be differently and is illustrated here merely as an example where the definition of, along channel order 32, interspersed random access channels 88a and 88b defines channel groups covering contigu¬ ous channels along the channel order 32. Other groups of channels might be defined, which do not necessarily follow the channel order 32, and inter-channel dependencies might be restricted not to render any channel of one group dependent on a channel of any other group, and within each group the inter-channel dependencies may also by restricted or each chan¬ nel might by coded inter-channel dependent on any previously coded channel within its channel group.
[0048] The block predictor 62 and 86 of encoder 10 and decoder 12, respectively, operate synchro¬ nously, i.e., they generate the same prediction signal 64 based on the previously en- coded / decoded samples of previously encoded / decoded temporal blocks 140. On encoder side 10, the prediction for a certain temporal block 140 may be accompanied or determined by one or more prediction parameters. Same might be determined on encoder side based on a rate / distortion optimization. These prediction parameters 90 are coded into data stream 16 and they are decoded from data stream 16 and used by block predictor 86 so as to perform the same prediction.
[0049] It might be that encoder 10 and decoder 12 support more than one prediction mode. For in¬ stance, encoder 10 and decoder 12 may support an intra prediction mode (which mode may also be called block-copy mode) according to which the currently encoded / decoded temporal block 140 is predicted based on the reconstructable sample values of previously en¬ coded / decoded temporal blocks of the same coded channel to which the currently en¬ coded / decoded temporal block 140 belongs, which is coded channel 92 in the example of Fig. 4. Additionally or alternatively, encoder 10 and decoder 12 may support an inter-prediction mode (which mode may also be called cross-channel prediction mode) according to
[0050] AFH250308PEP-2026096718. DOCX 9
[0051] which the currently encoded / decoded temporal block 140 is predicted based on the recon¬ structable sample values of previously encoded / decoded temporal blocks of one or more coded channels preceding - in coding order 32 - the coded channel 92 to which the currently encoded / decoded temporal block 140 belongs. Additionally or alternatively, there may be a mixed prediction mode according to which the prediction signal 64 is obtained by both, re- constructed / reconstructable sample values of previously encoded / decoded temporal blocks of coded channel 92 itself as well as reconstructed / reconstructable sample values of one or more coded channels preceding coded channel 92 in channel order along axis 32. Beyond this, there may be temporal blocks 140 which are coded without any prediction at encoder 10 and decoded without any prediction at decoder 12 such as the first temporal blocks 140 in the tiles 94 resulting from mutually separating the temporal blocks by means of the random access borders 96 on the one hand and the random access channel borders 98 on the other hand. This corresponds to the prediction signal 64 being set to zero and this may form an ad¬ ditional mode which could be called bypass mode. Additionally, or alternatively, there may be other modes such as ones deriving a DC predictor or linear function predictor for block 64 based on immediately preceding samples which immediately precede block 140. The predic¬ tion parameters 90 may, thus, contain for a currently encoded / decoded temporal block 140 a prediction mode flag or prediction mode indicator indicating the prediction mode to be used for this currently encoded / decoded temporal block 140 and, optionally, one or more parame-ters parameterizing the prediction mode to be used for this currently encoded / decoded temporal block 140. It might also be that the prediction parameters are themselves coded predic-tively from already reconstructed blocks 140. In this prediction process, the laid out random¬ access capabilities in channel- and temporal-direction are, as an example, always main¬ tained, i.e. the mentioned prediction of prediction parameters may never be supported across such a random access segment.
[0052] As mentioned, the aforementioned coding dependencies ought not to cross any of the borders 96 and 98 not only result from the just-described sample prediction capabilities of block predictor 62 and 86, respectively, but may optionally also result from other mechanisms such as parameter prediction according to which parameters such as the aforementioned predic-tion parameters 90 for a certain temporal block 140 are predicted based on coding parameters conveyed in the data stream 16 for any previous temporal block, or context derivation for context-adaptive entropy coding / decoding any coding parameter such as the prediction parameters 90 or any other side information such as side information 76 and 36 for temporal block 140 based on any coding parameter conveyed in the data stream 16 for any preceding temporal block.
[0053] That is, summarizing, the encoder 10 encodes the multi-channel signal 14 by transferring it into the coded domain 28 and then coding the coded channels into data stream 16 in theFH250308PEP-2026096718. DOCX 10
[0054] just-described block-wise and predictive manner, wherein decoder 12 decodes the coded channels of coded domain 28 from data stream 16 and the corresponding block-wise and predictive manner with then gaining the multi-channel signal 14 in its original form 26 based on the coded channels in coded domain 28 by means of segment 38. As said, the channel transformation is optional and if not used, each sample 40 in the coded domain 28 really cor¬ responds to one sample 18 in the original domain 26. If, further, the temporal mutual align¬ ment is not used, each sample 40 exactly corresponds to a sample 18 in the original domain 26 at exactly the same time instant or, differently speaking, all temporally co-located samples 40 in coded domain 28 remain mutually temporally co-located in the original domain 26. It should be noted that the temporal blocks 30 might, other than illustrated in Fig.5, vary in block length rather than being of a constant length as depicted in Fig. 4. For instance, encoder 10 may decide on the length of blocks 30 and signal the block length of blocks 30 (and the corresponding temporal blocks 140 of the coded channels) within data stream 16. Such signaling might be done on block level, such as for each temporal block 30 or, differently speaking for each temporally aligned bundle of blocks 140, so that the encoder may decide on the block size on the fly, or the block length might be signaled in the stream 16 on a larger scope such as for a sequence of blocks or even the whole stream 16.
[0055] As to the residual coder and residual decoder 70 and 82, they may use transform coding / de- coding in order to convey the residual signal 76 in data stream 16. That is, the residual signal 80 may be conveyed in data stream 16 in transform or spectral domain by way of transform coefficients in residual signal 76. The transform domain might be a DCT, DST or an FFT. The transform may be non-overlapping, i.e. it may only transform residual signal 80 and its re-transform may only cover residual signal 76 within block 140, and / or may be non-win-dowed, i.e. the residual signal might be transformed without any transform window used to temporally shape the residual signal 80 before the transform. The transform domain, i.e. the transformation leading from time domain to transform domain which is used by the encoder to transform the prediction residual signal 80 to be coded und the corresponding re-transfor- mation leading from transform domain to time domain which is used by the decoder to derive the prediction residual signal 80, or the transformation, might be selected from a set of avail-able transforms including, for instance, one or more of 1 ) one or more DCTs, 2) one or more DSTs and 3) an identity transform according to which the prediction residual signal 80 is coded into the data stream 14 in time domain directly. The transform may be critically sampled in that the number of transform coefficients resulting from the samples of one block 140 may equal the number of samples of block 140. Again, the samples might be the residual samples or may be, in case of the bypass mode, the channel samples directly.FH250308PEP-2026096718. DOCX 11
[0056] The transform coefficients might be encoded by quantization, i.e. they may be quantized with the quantized coefficients then being coded in the datastream 16. Dequantization may occur at decoding. For quantization, either a scalar uniform reconstruction quantizer or a low complexity vector quantizer might be used. In order to determine the quantization indices, the en- coder may perform some optimization algorithm such as a rate-distortion optimized scalar quantization, or a trellis quantization with the goal to approximately minimize an approxi¬ mated Lagrangian rate-distortion cost. At the decoder, the reconstruction process that yiels the transform coefficients may be conducted by multiplying the coded quantization indices with a certain step-size and, in case of the use of a low-complexity vector quantizer, by addi-tionally invoking a state-machine based on the parity of previously decoded quantization indi¬ ces in order to reconstruct the current quantization index.
[0057] In order to control the quantization noise, the transform coefficients might be subject to noise shaping. Spectral noise shaping may be used to shape the quantization noise spectrally. This may be done by signaling in the data stream spectral-band scale factors, i.e. a scale factor per spectral band, which represent a transfer function of a spectral filter which approxi¬ mates the spectral envelope of the signal within the current block 140 (or its prediction resid¬ ual, respectively), or signaling filter coefficients defining a temporal filter having a filter transfer function which approximates the spectral envelope of the signal within the current block 140 (or its prediction residual, respectively). On encoder side, spectral noise shaping may be applied in spectral domain by multiplying an inverse of scale factors, either directly signaled in the data stream or derivable from the filter coefficients by filter-to-factor conversion, with the transform coefficients before quantization. That is, at encoder, the coefficients are shaped by the inverse of the spectral envelope. At decoder side, spectral shaping may be applied in spectral domain by multiplying scale factors, either directly signaled in the data stream or derived from the filter coefficients by filter-to-factor conversion, with the transform coefficients, with then. That is, at decoder, the coefficients are shaped by the spectral enve¬ lope before applying retransformation. Additionally or alternatively, temporal noise shaping might be applied. To this end, TNS filter coefficients might be determined and signaled by the encoder. The TNS filter coefficients may represent a transfer function which approximates the temporal envelope of the current block 140 (or its residual signal). The encoder may ap¬ ply TNS filtering using the filter coefficients by spectrally filtering the possibly spectrally shaped transform coefficients so as to filter them with a transfer function corresponding to an inverse of the temporal envelope. The TNS filter coefficients might be derived by linear prediction analysis of the possibly spectrally shaped transform coefficients so as to derive a lin- ear prediction filter, then used as TNS filter, which minimizes a prediction residual when spectrally applied on the possibly spectrally shaped transform coefficients. At the encoder,FH250308PEP-2026096718. DOCX 12
[0058] the TNS filtered coefficients are then quantized and entropy coded. At decoder side, the in¬ verse takes place: the possibly spectrally shaped transform coefficients are inversely TNS filtered before applying retransformation. Additionally or alternatively, noise filling might be used. The filling may be applied to zero-quantized portions of the spectrum and controlled by the encoder via corresponding noise filling parameters.
[0059] As to the encoding / decoding the block or sequence of quantized transform coefficients of a current block into / from the data stream 16, arithmetic coding, such as context-adaptive bi¬ nary arithmetic coding, CABAC, may be used. The CABAC encoding / decoding may by performed frame wise. That is, in each channel, the sequence of blocks 140 may be partitioned into immediately consecutive blocks 140, which form frames. This partitioning may be equal among the channels so that, again, a frame denotes both a temporal portion within each channel individually, as well as a temporal portion of the multi-channel signal, i.e. a collection of temporally aligned frames. Within each frame, the sequence of blocks 140 are CABAC en / decoded with once initializing the contexts and resetting the internal CABAC state at the beginning and then updating the contexts’ probabilities during en / decoding the respective frame. That is, blocks 140 are CABAC decodable merely in units of frames. The context ini¬ tialization might be done independent from previous frames, or depending on the contexts as manifesting itself at the end of, of during, the en / decoding a previous frame.
[0060] Some deblocking processing might be used to avoid blocking artifacts. If, alternatively, an overlapped transform is used, an overlap-add processing with re-transforms of immediately preceding / succeeding temporal blocks of the same coded channel might be used in order to completely reconstruct the current temporal block’s 140 residual signal 76.
[0061] Besides such transform-(residual)-coded blocks there might be temporal blocks 140 which, additionally or alternatively, are coded using, besides the block prediction by block predictor 62 / 86 - which could be called a primary prediction - a secondary sample-wise prediction of the residual samples in residual block 66 such as by predicting a current sample’s residual sample by means of already decoded values of preceding - in sample coding order - resid¬ ual samples in block 66 or 80, with then correcting same by means of a secondary-predic¬ tion-residual sample decoded from the data stream 16. The secondary-prediction-residual samples for such a block may coded into the data stream en block in a transform domain or sample-wise in time domain.
[0062] Note that the afore-mentioned spectral shaping of the residual signal of a block 140 might be seen as a sample wise residual prediction, i.e. the case where filter coefficients are signaled for a block which define a temporal filter having a filter transfer function which approximates the spectral envelope of the residual signal within a current block 140. In sample wise resid¬ ual prediction, the residual predictor on a current block 140 might either be chosen out of aFH250308PEP-2026096718. DOCX 13
[0063] fixed set of prediction modes, where an index to such a residual prediction mode is signaled in the bit-stream, or the residual prediction mode might be ‘signal adaptive’. In the latter case, prediction filter coefficients for the residual predictor are determined at the encoder by solv¬ ing for example a linear equation, and are then quantized and transmitted to the decoder. At the decoder, the coefficients are inverse quantized and then the sample-wise prediction is conducted with these coefficients. The number of used coefficients may vary per block and might also be signaled in the bit-stream. Additionally, it might optionally (i.e. indicated by some information in the bit-stream) be supported to invoke collocated samples from a previ¬ ous block for the sample wise residual prediction. Finally, the coefficients of the sample wise residual prediction might be coded predictively, i.e. be predicted from used coefficients of a previous block, where only the differences to the current coefficients are transmitted.
[0064] A final note shall be made with respect to the juxtaposition of frames, blocks 140, channels and channel groups and regarding decoding order. The description above and below already described the fact that the channels might be grouped into channel group with each channel group being coded independently from each other, meaning that the blocks 140 in a certain channel group are coded without dependencies from channels outside their channel group. The decoding order 60, thus, would traverse the channels channel-group individually, chan¬ nel group by channel group. Within each channel group, the blocks 140 are traversed as de¬ scribed: all temporally aligned blocks 140 of all channels fist, then proceeding with the next blocks 140 and so forth. A frame may have a sequence of blocks of a channel group en¬ coded thereinto along the mentioned decoding order, such as n temporally consecutive blocks 140 for all channels of a channel group. IF the channel group had m channels, m*n block104 would, thus, be coded into the frame. As mentioned, there might be dependent frames, for which the CABAC contexts are adopted from the preceding frame of the same channel group, i.e. the one having encoded the immediately preceding block 140. For such dependent frames, not only CABAC contexts may be adopted from the preceding frame, but it may also be allowed to allow for prediction from the preceding frame to the dependent frame. Prediction, and possibly also any coding dependencies, towards channels outside the channel group and, within the channel group, towards frames temporally preceding the mostly recently previously en / decoded independent frame would be disallowed. Thus, each tile shown in Fig. 4 by bold lines may represent a sequence of an independent frame flowed by zero, one or more dependent frames.
[0065] As mentioned before, Fig. 4 only represents a possible “framework" into which the previously described embodiments and the embodiments described subsequently may be built into. Many modifications may be performed with respect to Fig. 4, and some of these modifications might be mentioned in the subsequent description with respect to certain ones of the
[0066] / LFH250308PEP-2026096718. DOCX 14
[0067] subsequently described embodiments, but these modifications shall then be treated as being also applicable with respect to other ones of the subsequently described embodiments.
[0068] The description is now resumed with respect to the announced subsequently described im¬ plementation examples and further embodiments where the digital time-varying signal is not restricted to be a channel of a multi-channel signal or to be a multi-channel signal, but where same may only be a single digital scalar signal.
[0069] The present invention
[0070] Fig. 1 shows an example 110. According to the example 110, the encoder (in particular the predictor 62) downsamples an already encoded portion 72 (already subjected to encoding, already subjected to prediction, already subjected to reconstruction) in step 112. In particular, the already encoded portion (already subjected to encoding, already subjected to prediction, already subjected to reconstruction) 72 can be a block or another kind of portion (e.g., a block in another channel). The predictor 62, therefore, through the steps of example 110, creates a prediction pred L)[i] = x[m][sfc- L + i |, i = 0,...,bk- 1. This prediction is based on an offset L, e.g., as shown in Fig. 3. This, however, is obtained in the downsampled resolution (see Fig.
[0071] 3, where here a downsampling factor is k = 2). We note that L = k * L (where stands for multiplication), where L is the offset in the downsampling resolution (downsampled resolution or downsampled sample resolution). In step 114, there is determined the similarity between the current temporal block 30 and samples in the already encoded portion 72. Hence, the cost estimation of a suitable offset L is retrieved. Hence, at step 116, a suitable offset value L is obtained for the current temporal block 30.
[0072] In step 118, there is a conversion of the offset value L to the original sample resolution, to obtain a final offset value L, e.g., through the formula L = k * L. At this point, in step 120, the current temporal block 30 may be predicted from the already encoded portion associated with the final offset value L. This is obtained through the formula pred^L^i] = %[m][s / {- L + i ], i = 0,..., bk- 1. The value pred(L)[i] is referred to 64 in Fig. 4. Then, at step 122 (block 68 in Fig.
[0073] 4) the residual signal 66 is obtained by comparing the prediction signal 64 with the signal to be encoded (in block 30). Then, the residual signal 66 is encoded in the bitstream 16 and the final offset value L (which is equal to k * L) is encoded in the bitstream 16 as side information (e.g. in the parameters 90).
[0074] Fig. 2 shows the variant 111, according to which:
[0075] step 116 of example 110 is substituted by step 117 of variant 111, in which an offset value set is obtained at the downsampled sample resolution; andFH250308PEP-2026096718. DOCX 15
[0076] step 118 of example 110 is substituted by step 119 of variant 111, in which the offset value L is obtained at the original sample resolution from the set of offset values.
[0077] Step 114 is performed like in the example 110. In the downsampled resolution (Fig. 3) the offset value L for the current temporal block 30 is chosen among a plurality of candidates, and the one candidate (or a set of candidates) that minimizes a dissimilarity is the one candidate (or the one set of candidates) that is chosen as offset value L (or that is chosen as set of candidates from which to retrieve the offset value) in the downsampled resolution. Hence, the similarity between the current temporal block 30 and the already encoded portion 72 is evaluated for different offset values (e.g., at the downsampled resolution) to obtain an offset value set [e.g. a test area, or more in general a set of offset values which is less than a global set of offset values (e.g. an unlimited set of offset values)] for the current temporal block 30 at the downsampled resolution (e.g., preselection of step 117 in the variant 111 of Fig. 2), so as to have offset values set which includes a plurality of different offset value candidates. Subse¬ quently, once restricted to different offset value candidate at the original sample resolution, there is determined a similarity between the current sample temporal block 30 and the already encoded portion 72 at the original sample resolution. Hence, there is the possibility of restrict¬ ing the number of candidates of offset values in the downsampled resolution, and then to perform the final evaluation in the original sample resolution, so as to find the offset value L in the original sample resolution.
[0078] The offset values set may be or include at least one offset values interval (e.g., between 71 and 72) so as to subsequently evaluate the final offset value L between the extremities of an interval in the sample resolution between L1 and L2 (with LI = k * LI and 72 = k * 72) corresponding to the already retrieved interval (between 71 and 72) in the downsampled resolution. This is basically an example of steps 117 and 119 of example 111 of Fig. 2. The similarity may be evaluated as a minimization of similarity cost, coding efficiency measurement, distortion measurement and / or correlation measurements. In particular, either the offset 71, or the inter¬ val between 71 and 72 (as in example 111 of Fig. 2) may be determined by evaluating a minimization of a similarity cost, a coding efficiency measurement, a distortion measurement and / or correlation measurements in the subsampling resolution. Hence:
[0079] in example 110 of Fig. 1, 7 is obtained in step 116 in the downsampled sample resolution and is subsequently covered in step 118 (e.g., through the multiplication 7 = k * 7)
[0080] in the variant 111 of Fig. 2, the interval between 71 and 72 may be obtained in step 117, and in the subsequent step 118, at the original sample resolution, the final offset value LFH250308PEP-2026096718. DGCX 16
[0081] is obtained from the value set corresponding to the value set between L1 and L2 (with LI = k * 11 and L2 = k * L2).
[0082] In steps 114, 117, 117, 119, 116, it may be possible to perform operations of minimizing a similarity cost, coding efficiency measurement, distortion measurements, and / or correlation 5 measurements, to obtain L or LI and L2 and / or to retrieve L from L1 and L2.
[0083] As explained above, the downsampling ratio is k which can be a power of 2 or at least is a submultiple of the current block size bk. Hence, it is possible to obtain the offset value L by scaling the offset value by k (i.e., L = k * L or by scaling by k at least one of the offset value indicative of the offset values set [e.g., in case of Fig. 1 and in the case of the offset values set 10 being an interval, and the at least one offset value indicative of the offset values set being, for example, the extremities of the interval, or at least one other offset value indicative of the offset value (for example, there could be one single offset value which is the center of the interval, and the length of the interval would be pre-determined; or, there could be one single offset value which is the center of the interval, and the length of the interval would be provided as a 15 second value indicative of the interval)].
[0084] It is possible to apply in the converting step 118 and / or 119 an interpolation filter.
[0085] The downsampling can be a temporal downsampling. The offset value L may be temporal offset values. The offset values may probably include channel wise offset values.
[0086] 20 In some examples, based on a requirement of having high performance, the choice may be made of avoiding the temporal downsampling and, therefore, to determine similarity between the current temporal block and the radially encoded portion without downsampling (i.e., retrieving L in the downsampling resolution without performing the downsampling).
[0087] 25 Discussion on the present solutions
[0088] To address the above and below mentioned drawback of slow BM, CC or LTP encoding, the following is proposed:
[0089] • since the search for C, L is usually done by calculating, for a given offset candidate, the prediction signal and, thereby, an associated block residual candidate and, thereby, a 30 prediction gain or cost, use temporal downsampling during the calculations of said prediction signal candidate and block residual candidate; i. e., said gain or cost are in a downsampled domain,
[0090] • not using temporal downsampling for encoding configurations requiring high performance,
[0091] A*FH250308PEP-2026096718. DOCX 17
[0092] • using the corresponding non / downsampled BM, LTP, CC prediction signal during actual rate-distortion optimized (RDO) encoding for the given block for encoder-decoder synchronicity purposes when the G, / --search is carried out in a downsampled domain, as pro¬ posed above and below.
[0093] In other words, encoder speedup is achieved by reducing the search space via signal downsampling, such that an encoder performs fewer sample-wise comparisons than when not downsampling. Detailed preferred embodiments for each of these three aspects will be described hereafter.
[0094] The following description is based on a BM predictor, however, the presented concept can also be applied to other predictors such as CC or LTP. Given is the current block with starting position skin channel m and samples x|m]|sk+ i], i = 0,...,bk- 1. When testing a BM predictor with offset L a block prediction is created by shifting the block by L to create pred L)[i] = x[m]|sfc- L + i], 1 = 0,...,bk- 1 from the reconstructed signal of previously encoded samples. Then, the residual
[0095] res(sk, L)[i] = x[m][sk+ i] - pred(L')[i], i = 0,...,bk~ 1 is created and a distortion measure is applied to the residual and / or the rate for the transmitting the offset L is calculated. Based on this cost estimation, a suitable offset L is searched. To reduce the number of tested offset candidates, the following pre search process can be invoked.
[0096] First, a new downsampled signal is calculated from the reconstructed samples x[m]. There¬ fore, a (weighted) sum is calculated every k samples. Alternatively, instead of using all samples, a partial sum or one every k samples can be used. The parameter k is chosen such that the current block size bkis divisible by k, e.g if the block size is a power of 2, then k is also a power of 2. Similarly, the current block of the original signal x[rn][s{+ i], i = 0,...,bk- 1 is downsampled in the same way.
[0097] A given offset L on the downsampled reconstructed samples corresponds to an offset L = k * L on the reconstructed samples. As an estimate for the distortion of the residual of the block prediction with offset L, the same distortion measure as before can be used on the residual between the downsampled original signal and the downsampled reconstructed signal with offset L e.g. sum of squares or sum of absolute differences. Due to the downsampling, the number of calculations per block are reduced by the factor k. Note that methods to adapt the signal to the original signal which are based on the left reconstructed boundary sam- ples(such as adding an offset) are still applicable as long as the size of the left boundary samples is not greater than k.FH250308PEP-2026096718. DOCX 18
[0098] After a subset of candidates has been identified on the downsampled signal, the search can be further refined by switching back to the reference samples. As the offset L on the downsampled signal can only be searched with an accuracy of k, the remaining candidates in the proximity of the offset L with a distance less than k are tested additionally.
[0099] To define a subset of candidates either a fixed maximum number of candidates and / or a maximum cost threshold can be applied. A possible threshold is a multiple of the lowest cost or a multiple of the cost of a constant prediction signal.
[0100] While such a method can greatly decrease the amount of calculations to get a “good” block matching candidate, it may be a feature that can be turned off in a high performance setting to preserve the accuracy. Additionally, when performing refinements to the prediction candi¬ date such as using interpolation filters or applying RDO it is helpful to use the reconstructed signal without downsampling. Using the downsampling only as a speedup at the encoder to get a preselection of candidates also helps to maintain the synchronicity of the encoder and decoder and avoids the overhead at the decoder side to additionally create the downsampled signal.
[0101] In case of a CC prediction the pre selection of “good” candidates can be done in the same manner as for the BM predictor by testing all candidates on a downsampled signal. Then, a final decision on the “best” candidate should be made on the reference signal without downsampling. Similar to the block matching predictor the adjustments based on the left boundary samples are still valid on the downsampled signal, given that the number of samples on the left boundary is not greater than the downsampling factor k.
[0102] Regardless of the type of prediction, the downsampling may already apply a low-pass filter to the signal and may favor offsets which benefit from such a filter. Hence, it may be more suitable to always choose a low-pass filter in addition to the offset.
[0103] Applying the concept to CC prediction is similar to the BM predictor, and can be also performed in the invention. A set of offset values is searched in the downsampled resolution and the final offset is searched among this set in the original sample resolution. In contrast to BM prediction, the offset does not need to be rescaled for the CC prediction. The only formula that would change is pred(
[0104]
[0105] L)[i] = + *].
[0106] Aspects
[0107] A 1staspect provides an encoder for encoding a digital-time-varying signal (92), configured to encode the digital-time-varying signal (92) into a data stream (16) in temporal blocks (30) by encoding a current temporal block (30) by:FH250308PEP-2026096718. DQCX 19
[0108] predicting (62, 120) the current temporal block (30) [e.g. from previously reconstructed samples], to derive a prediction signal (64),
[0109] deriving (68, 122) a prediction residual signal (66) by comparing the digital-time-varying sig¬ nal (92) with the prediction signal (64); and
[0110] encoding (70, 124), into the data stream (16), prediction residual coefficients (76) which rep¬ resent the prediction residual signal (66),
[0111] configured, in predicting the current temporal block, to:
[0112] downsample (112) an already encoded portion of the digital-time-varying signal [and, in some examples, also downsample the current temporal block] from an original sample reso¬ lution to a downsampled sampling resolution, to derive a downsampled version of the already encoded portion;
[0113] determine (114) a similarity between the current temporal block and the already encoded portion for different offset values at the downsampled sample resolution (downsample sampling resolution), to obtain (116, 117) a final offset value for the current temporal block [e.g. in the original sample resolution]; and
[0114] encode (124) the final offset value for the current temporal block into the data stream.
[0115] [The digital-time-varying signal may have been obtained by at least one of Electrocardiog¬ raphy, Electroencephalography, Electromyography or seismic measurement, or another means. Differently speaking, the digital-time-varying signal might be a bio-physiological waveform data such as an electroencephalography (EEG) signal, an electrocardiogram (ECG), or an electromyography (EMG) signal, or seismic waveform data, or another signal] According to a 2ndaspect when referring back to the 1staspect, the encoder may be configured to [e.g. example 110 in Fig. 1] determine (114) the similarity between the current tem¬ poral block and the already encoded portion for different offset values at the downsampled sample resolution, to obtain the offset value for the current temporal block at the downsam¬ pled sample resolution, and
[0116] to convert (118) the offset value for the current temporal block from the downsampled sample resolution to the original sample resolution, thereby obtaining the final offset value. According to a 3rdaspect when referring back to the 1stor 2ndaspect, the encoder may be configured to [e.g. example 111 in Fig. 2] first, determine (114) the similarity between the cur¬ rent temporal block and the already encoded portion for different offset values at the downsampled sample resolution, to obtain (117) an offset values set [e.g. a test area, or more in general a set of offset values which is less than a global set of offset values (e.g. an
[0117] it.
[0118] FH250308PEP-2026096718. DOCX 20
[0119] unlimited set of offset values)] for the current temporal block [e.g. in the original sample reso¬ lution] at the downsampled sample resolution [e.g. preselection], the offset values set includ¬ ing a plurality of different offset values [e.g., but less than the set of possible offset values]; and
[0120] second, restricted to different offset values at the original sample resolution which corre¬ spond to the different offset values in the offset values set at the downsampled sample resolution, determine a similarity between the current temporal block and the already encoded portion, to obtain (119) the final offset value [e.g. final selection].
[0121] According to a 4thaspect when referring back to the 3rdaspect, offset values set may be or at least may include at least one offset values interval, the encoder being configured to convert information on the offset values interval from the downsampled sample resolution to the original sample resolution [e.g. only the extremities of the offset values interval could be con¬ verted, and the a similarity between the current temporal block and the already encoded por¬ tion in the original sample resolution is performed by using the different offset values be¬ tween the converted extremities of the interval]
[0122] According to a 5thaspect when referring back to any one of the 1stto 4thaspects, the encoder may be configured to determine the similarity, at least in one step [e.g. at least in step 114 of Fig. 1 or at least in step 114 and / or step 119 of Fig. 2], by minimizing a similarity cost.
[0123] According to a 6thaspect when referring back to any one of the 1stto 5thaspects, the encoder may be configured to determine the similarity, at least in one step [e.g. at least in step 114 of Fig. 1 or at least in step 114 and / or step 119 of Fig. 2], through a coding efficiency measure¬ ment.
[0124] According to a 7thaspect when referring back to any one of the 1stto 6thaspects, the encoder may be configured to determine the similarity, at least in one step [e.g. at least in step 114 of Fig. 1 or at least in step 114 and / or step 119 of Fig. 2], through a distortion measurement. According to a 8thaspect when referring back to any one of the 1stto 7thaspects, the encoder may be configured to determine the similarity, at least in one step [e.g. at least in step 114 of Fig. 1 or at least in step 114 and / or step 119 of Fig. 2], through correlation measurements. According to a 9thaspect when referring back to any one of the 1stto 8thaspects, the encoder may be configured to apply a downsampling ratio of k, and, to obtain [e.g. in Fig. 1] the offset value by scaling the final offset value by k, or by scaling by k at least one offset value indica¬ tive of the offset values set [e.g. in case of Fig. 1 and in the case of the offset values set being an interval, and the at least one offset value indicative of the offset values set being, for example, the extremities of the interval, or at least one other offset value indicative of the off-FH250308PEP-2026096718. DOCX 21
[0125] set value (for example, there could be one single offset value which is the center of the inter¬ val, and the length of the interval would be pre-determined; or, there could be one single off¬ set value which is the center of the interval, and the length of the interval would be provided as a second value indicative of the interval)].
[0126] According to a 10thaspect when referring back to any one of the 1stto 9thaspects, the encoder may be configured to apply an interpolation filter [e.g. at least in converting 118 of Fig.
[0127] 1 or in 119 of Fig. 2],
[0128] According to an 11thaspect when referring back to any one of the 1stto 10thaspects, the downsampling may be a temporal downsampling.
[0129] According to a 12thaspect when referring back to any one of the 1stto 11thaspects, the offset values may be or at least may include temporal offset values.
[0130] According to a 13thaspect when referring back to any one of the 1stto 12thaspects, the offset values may be or at least may include channel-wise offset values.
[0131] A 14thaspect provides a method for encoding a digital-time-varying signal (92), to encode the digital-time-varying signal (92) into a data stream (16) in temporal blocks (30) by encoding a current temporal block (30), the method comprising:
[0132] predicting (62, 120) the current temporal block (30) [e.g. from previously reconstructed sam¬ ples], to derive a prediction signal (64);
[0133] deriving (68, 122) a prediction residual signal (66) by comparing the digital-time-varying sig-nal (92) with the prediction signal (64); and
[0134] encoding (70, 124), into the data stream (16), prediction residual coefficients (76) which represent the prediction residual signal (66),
[0135] the method including, in predicting the current temporal block:
[0136] downsampling (112) an already encoded portion of the digital-time-varying signal [and, in some examples, also downsample the current temporal block] from an original sample resolution to a downsampled sampling resolution, to derive a downsampled version of the already encoded portion;
[0137] determining (114) a similarity between the current temporal block and the already encoded portion for different offset values at the downsampled sample resolution, to obtain (116, 117) a final offset value for the current temporal block [e.g. in the original sample resolution]; and encoding (124) the final offset value for the current temporal block into the data stream.FII250308PEP-2026096718.DOCX 22
[0138] A 15thaspect provides a non-transitory storage unit storing instruction which, when executed by a processor, cause the processor to perform the method when referring back to the 14thaspect.
[0139] Implementation alternatives
[0140] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a program¬ mable computer or an electronic circuit. In some embodiments, one or more of the most im¬ portant method steps may be executed by such an apparatus.
[0141] Depending on certain implementation requirements, embodiments of the present tech¬ nique can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
[0142] Some embodiments according to the present technique comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0143] Generally, embodiments of the present technique can be implemented as a computer pro¬ gram product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine-readable carrier.
[0144] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine-readable carrier.
[0145] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0146] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitionary. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods describedFH250308PEP-2026096718. DOCX 23
[0147] herein. The data stream or the sequence of signals may for example be configured to be trans¬ ferred via a data communication connection, for example via the Internet.
[0148] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0149] A further embodiment comprises a computer having installed thereon the computer pro¬ gram for performing one of the methods described herein.
[0150] A further embodiment according to the present technique comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0151] In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a micro¬ processor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
[0152] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0153] The apparatus described herein, or any components of the apparatus described herein, may be implemented at least partially in hardware and / or in software.
[0154] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0155] The methods described herein, or any components of the apparatus described herein, may be performed at least partially by hardware and / or by software.
[0156] The above-described embodiments are merely illustrative for the principles of the present tech¬ nique. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details pre¬ sented by way of description and explanation of the embodiments herein.
Claims
FH250308PEP 2026096718. DOCX 24Claims1. Encoder for encoding a digital-time-varying signal (92), configured to encode the digi¬ tal-time-varying signal (92) into a data stream (16) in temporal blocks (30) by encoding a cur¬ rent temporal block (30) by:predicting (62, 120) the current temporal block (30), to derive a prediction signal (64), deriving (68, 122) a prediction residual signal (66) by comparing the digital-time-vary¬ ing signal (92) with the prediction signal (64); andencoding (70, 124), into the data stream (16), prediction residual coefficients (76) which represent the prediction residual signal (66),configured, in predicting (62, 120) the current temporal block, to:downsample (112) an already encoded portion of the digital-time-varying sig¬ nal from an original sample resolution to a downsampled sampling resolution, to derive a downsampled version of the already encoded portion;determine (114) a similarity between the current temporal block and the al- ready encoded portion for different offset values at the downsampled sample resolu¬ tion, to obtain (116, 117) a final offset value for the current temporal block; and encode (124) the final offset value for the current temporal block into the data stream.
2. The encoder of claim 1, configured to determine (114) the similarity between the cur¬ rent temporal block and the already encoded portion for different offset values at the downsampled sample resolution, to obtain the offset value for the current temporal block at the downsampled sample resolution, andto convert (118) the offset value for the current temporal block from the downsampled sample resolution to the original sample resolution, thereby obtaining the final offset value.
3. The encoder of any of the preceding claims, configured to first, determine (114) the similarity between the current temporal block and the already encoded portion for different offset values at the downsampled sample resolution, to obtain (117) an offset values set for the current temporal block at the downsampled sample resolution, the offset values set including a plurality of different offset values; andsecond, restricted to different offset values at the original sample resolution which correspond to the different offset values in the offset values set at the downsampled sampleFH250308PEP-2026096718. DOCX 25resolution, determine a similarity between the current temporal block and the already encoded portion, to obtain (119) the final offset value.
4. The encoder of claim 3, wherein the offset values set is or at least includes at least one offset values interval, the encoder being configured to convert information on the offset values interval from the downsampled sample resolution to the original sample resolution5. The encoder of any of the preceding claims, configured to determine the similarity, at least in one step, by minimizing a similarity cost.
6. The encoder of any of the preceding claims, configured to determine the similarity, at least in one step, through a coding efficiency measurement.
7. The encoder of any of the preceding claims, configured to determine the similarity, at least in one step, through a distortion measurement.
8. The encoder of any of the preceding claims, configured to determine the similarity, at least in one step, through correlation measurements.
9. The encoder of any of the preceding claims, configured to apply a downsampling ratio of k, and, to obtain the offset value by scaling the final offset value by k, or by scaling by k at least one offset value indicative of the offset values set.
10. The encoder of any of the preceding claims, configured to apply an interpolation filter.
11. The encoder of any of the preceding claims, wherein the downsampling is a temporal downsampling.
12. The encoder of any of the preceding claims, wherein the offset values are or at least include temporal offset values.FH250308PEP-2026096718. DOCX 2613. The encoder of any of the preceding claims, wherein the offset values are or at least include channel-wise offset values.
14. A method for encoding a digital-time-varying signal (92), to encode the digital-time-varying signal (92) into a data stream (16) in temporal blocks (30) by encoding a current temporal block (30), the method comprising:predicting (62, 120) the current temporal block (30), to derive a prediction signal (64); deriving (68, 122) a prediction residual signal (66) by comparing the digital-time-vary¬ ing signal (92) with the prediction signal (64); andencoding (70, 124), into the data stream (16), prediction residual coefficients (76) which represent the prediction residual signal (66),the method including, in predicting the current temporal block:downsampling (112) an already encoded portion of the digital-time-varying signal from an original sample resolution to a downsampled sampling resolution, to derive a downsampled version of the already encoded portion;determining (114) a similarity between the current temporal block and the al¬ ready encoded portion for different offset values at the downsampled sample resolu¬ tion, to obtain (116, 117) a final offset value for the current temporal block; and encoding (124) the final offset value for the current temporal block into the data stream.
15. A non-transitory storage unit storing instruction which, when executed by a processor, cause the processor to perform the method of claim 14.