Methods, apparatus, and programs for utilizing coding distortion measure
By embedding distortion metadata into bitstreams, the method addresses the challenge of unknown coding distortion in lossy coding for utility signals, enabling effective automated analysis by providing distortion information for accurate classification and task suitability evaluation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Existing lossy coding techniques for utility signals, such as medical signals, introduce coding distortion that is unknown on the decoder side, making it challenging to perform automated analysis tasks effectively, as the decoder lacks information about the distortion introduced during coding.
Embedding distortion metadata into bitstreams that include coded time series data, providing information on the coding distortion, such as length and variance of the original data, to enable meaningful processing and assessment of decoded signals.
Enables accurate classification and suitability evaluation of decoded time series data for automated tasks by incorporating distortion metadata, allowing for seamless switching between lossy and lossless coding based on signal characteristics and ensuring compliance with downstream processing requirements.
Smart Images

Figure IMGF000025_0001 
Figure IMGF000025_0002 
Figure IMGF000025_0003
Abstract
Description
[0001] METHODS, APPARATUS, AND PROGRAMS FOR UTILIZING CODING DISTORTION MEASURE
[0002] Cross-Reference to Related Applications
[0003] This application claims priority of the following priority applications: U.S. Provisional Application No. 63 / 713,232, filed October 29, 2024, European Patent Application No.: 24209647.7, filed October 29, 2024, U.S. Provisional Application 63 / 825,025, filed June 17, 2025, and U.S. Provisional Application No. 63 / 881,744, filed September 15, 2025, each of which is hereby incorporated by reference in its entirety.
[0004] Technical Field
[0005] The present disclosure relates to techniques for generating or modifying bitstreams (e.g., framebased, packet-based, or element-based bitstreams). In particular, the present disclosure relates to providing distortion information for coded (e.g., core coded) waveform data in the bitstream.
[0006] Background
[0007] Presently, utility signal waveform data / time series data (for example medical / biomedical signals, haptic signals, seismic signals, accelerometer signals, etc.) is coded in lossless manner.
[0008] On the other hand, lossy coding of such utility signals, while conventionally not explored, may be of interest for many frameworks and use cases. In such cases, it may be assumed that the signals are encoded by an instance of a source codec, which operates in a constrained bitrate setting (for example, a constrained average bitrate per sample), or operates in a constrained distortion mode (in this case, the maximum coding distortion is fixed, and the operating bitrate is variable). In the most typical scenario, either of the constraints may be used to reduce the amount of storage required to store bitstreams associated with the capture of the utility signals. In some cases, it is possible that the bitrate constraint is used to address a throughput constraint.
[0009] However, contrary to multimedia signals, utility signals are seldomly captured to be played out to human observers or listeners. While in some cases these signals may be directly inspected (e.g., by a clinician), the majority of use cases comprise automated analysis, or using the captured signals for training of automated analysis tools. For instance, these signals may be often used to derive performance parameters or to estimate other signals of interest. As a specific example, a photoplethysmogram (PPG) may be captured, encoded, decoded, and reconstructed to perform an estimation of heart rate. According to another specific example, the PPG signal may be used to derive a heart rate variability (HRV) signal, which can be then used to estimate instantaneous HRV.
[0010] The captured utility signals may be encoded, and the bitstreams may be stored in medical records. It may happen that the analysis task to be performed on the decoder side remains unknown on the encoder side. This typically would not be problematic if lossless coding was used, since in this case the source codec would not introduce any coding distortion. However, if lossy coding is used, the distortion that was introduced during coding may become problematic, especially in the absence of any knowledge about the distortion on the decoder / analysis side. For example, some medical signals encoded in a lossy manner may become undesired for including them in analysis performed on the decoder side due to their high coding distortion. As another example, some medical signals coded in a lossy manner may need to be excluded from a signal dataset used for training of Al-based analysis methods.
[0011] There is thus need for improved coding techniques for utility signals, such as waveform signals or time series data in general. There is particular need for such coding techniques that enable meaningful processing of decoded utility signals.
[0012] Summary
[0013] In view of this need, the present disclosure provides methods of generating or modifying bitstreams (e.g., frame-based, packet-based, or element-based bitstreams) and of decoding bitstreams, as well as corresponding apparatus, computer programs, and computer-readable storage media, having the features of respective independent claims.
[0014] One aspect of the present disclosure relates to a method of generating or modifying a bitstream. The method may include obtaining (e.g., receiving or generating) the bitstream. The bitstream may contain coded time series data. The coded time-series data may be core-coded time series data, for example. The method may further include determining a measure of distortion in relation to a portion of the coded time series data. The distortion may be a coding distortion. The method may yet further include embedding distortion metadata into the bitstream. The distortion metadata may include information indicative of the determined measure of distortion. The distortion metadata may further include one or more of an indication of a length of the portion of coded time series data and an indication of a variance of original time series data corresponding to the coded time series data.
[0015] Configured as defined above, the proposed method allows to provide the decoding stage with information on a coding distortion of coded time series data that is transported in a bitstream. Since typically the decoder stage is not aware of the coding distortion, this aids meaningful treatment of any decoded time series data. For example, it may be determined that the decoded time series data is sufficiently close to the original time series data for allowing certain automated tasks to be performed in relation to the decoded time series data, or conversely, that the decoded time series data is not suited to a certain automated task. This may be particularly relevant for utility signals including medical signals that are typically not replayed to a human observer or listener. Thus, the proposed method offers the possibility of using lossy coding for all or part of time series data that had conventionally have to be coded in lossless manner, to ensure compliance with requirements of downstream processing tasks, in particular, automated processing tasks.
[0016] In some embodiments, the coded time series data may include lossy coded time series data. For example, the bitstream may include portions (e.g., segments) of lossy coded time series data and portions (e.g., segments) of lossless coded (or near lossless coded) time series data. For example, the core encoder may seamlessly switch between lossy and lossless coding (e.g., given its bitrate constraint), depending on characteristics of the input signal. In such case, the bitstream may include distortion metadata both for the lossy segments and the lossless segments. For lossless segments, the distortion metadata then may indicate zero distortion, or alternatively, that lossless coding has been used, which may not be apparent to the decoder side from inspecting the bitstream alone.
[0017] In some embodiments, the coded time series data may relate to original time series data, in the sense that the coded time series data may be obtainable from the original time series data via a coding (e.g., core (en-)coding) operation. Then, the method may further include local decoding of the coded time series data to obtain locally decoded time series data. Therein, determining the measure of distortion may be based on (one or more samples of) the original time series data and (one or more samples of) the locally decoded time series data. In this manner, the coding distortion can be accurately determined when the original time series data is available. This is always the case for a first encoding process that is applied to the time series data.
[0018] In some embodiments, the method may further include generating the coded time series data from original time series data via a coding operation (e.g., core (en-)coding operation). The method may yet further include estimating the measure of distortion in relation to the portion of the coded time series data based on a rate allocation process of the coding operation. By doing so, local decoding of the coding time series data can be avoided, to thereby reduce computational complexity and memory use.
[0019] In some embodiments, the coded time series data may be generated from the original time series data via the coding operation subject to a bitrate constraint, and the coding operation may be adapted to minimize a signal distortion given the bitrate constraint. Then, the measure of distortion may be estimated in the course of minimizing the signal distortion given the bitrate constraint.
[0020] In some embodiments, the coded time series data may be generated from the original time series data via the coding operation subject to an average bitrate constraint, and the coding operation may be adapted to optimize a rate-distortion trade-off given the average bitrate constraint, by selection of appropriate operation points of the coding operation. Then, the measure of distortion may be estimated in the course of selecting the appropriate operation points.
[0021] In some embodiments, the coding operation for obtaining the coded time series data may relate to distortion-constrained (en-)coding using a distortion constraint. Then, the measure of distortion may be determined based on the distortion constraint. The distortion constraint may be given in relation to a given distortion metric. The measure of distortion may be indicative of the distortion constraint and optionally the distortion metric.
[0022] In some embodiments, the coded time series data may include one or more channels of coded time series data. Then, the measure of distortion may relate to at least one of a maximum absolute error per channel and a per-channel squared error. In some embodiments, the method may further include quantizing the maximum absolute error per channel and / or the per-channel squared error to a number format of a predetermined number of bits, using a predefined quantization resolution. Therein, the measure of distortion may relate to the quantized maximum absolute error per channel and / or the quantized per-channel squared error. The number format may be an 8-bit number format, for example. The quantization resolution may be IdB, for example.
[0023] In some embodiments, the coded time series data may include a plurality of channel groups. Each channel group may include a plurality of channels of coded time series data. The channels may be grouped in channel groups for purposes of joint coding by a core coder. The measure of distortion may be embedded into the bitstream on a per channel group basis, or in general may be given on a per channel group basis. Each such measure of distortion may relate to a cumulated (e.g., aggregated, or averaged) measure of distortion for the respective channel group.
[0024] In some embodiments, the coded time series data may include a plurality of channel groups. Each channel group may include a plurality of channels of coded time series data. The measure of distortion may be embedded into the bitstream on a per channel group basis. Each such measure of distortion may relate to individual measures of distortion for the channels in the respective channel group.
[0025] In some embodiments, the method may further include quantizing the measure of distortion according to a predefined set of distortion levels (or according to a predefined quantization scheme, to the predefined set of distortion levels). The quantizing may involve comparing the measure of distortion to a plurality (e.g., sequence) of different thresholds. The distortion levels may be represented by respective indices, for example. The distortion metadata may include information indicative of the distortion level to which the measure of distortion has been quantized. Optionally, the distortion metadata may include an indication that the measure of distortion has been quantized according to the predefined set of distortion levels.
[0026] In some embodiments, the distortion metadata may further include information indicative of a length of the portion of the coded time series data to which the measure of distortion applies. This allows to flexibly provide the distortion metadata in the bitstream and to clearly indicate the portions of coded time series data which respective items of distortion metadata relate to. In some embodiments, the distortion metadata may further include information indicative of a variance of the original time series data relating to the portion of the coded time series data. The variance may be a variance of sample values of samples of the original time series data. The variance may be given on a per-channel basis, for example. Providing the decoder stage with information of the variance of the original time series data allows for an accurate assessment of the coding distortion and signal quality based on the measure of distortion.
[0027] In some embodiments, the portion of the coded time series data may relate to a segment including an integer multiple of random access intervals of coded time series data.
[0028] In some embodiments, the portion of coded time series data may relate to the entire bitstream.
[0029] In some embodiments, the embedding may include embedding a data packet into the bitstream that includes an indication of the measure of distortion and an indication of the corresponding portion of the coded time series data. The indication of the corresponding portion of the coded time series data may be a pointer to said portion (e.g., segment, group of frames) or information on a length of said portion, for example.
[0030] In some embodiments where the portion of the coded time series data relates to the segment including the integer multiple of random access intervals, the embedding may include embedding a data packet into the bitstream that includes an indication of the measure of distortion. Therein, the data packet may be embedded at the end of the segment. Embedding the data packet at the end of the segment may avoid a processing delay when encoding as well as decoding the bitstream and referring to the measure of distortion for further processing of the time series data.
[0031] In some embodiments where the portion of the coded time series data relates to the entire bitstream, the embedding may include embedding a data packet into the bitstream that includes an indication of the measure of distortion. Therein, the data packet may be embedded at the end of the bitstream.
[0032] In some embodiments, the time series data may relate to one or more utility signals. The time series data may relate to one or more channels of utility signals. In some embodiments, the time series data may relate to one or more medical or biomedical signals. The time series data may relate to one or more channels of medical or biomedical signals.
[0033] In some embodiments, the time series data may relate to waveform data. In particular, the time series data may relate to one or more channels of waveform data.
[0034] In some embodiments, the method may further include receiving an output of a sensor and applying a coding operation thereto to generate the coded time series data. The output of the sensor may relate to the original time series data.
[0035] Another aspect of the disclosure relates to a method of modifying (e.g., transcoding) a bitstream. The method may include obtaining the bitstream as an input. The bitstream may include coded time series data and distortion metadata. The distortion metadata may include information indicative of a measure of distortion in relation to a portion of the coded time series data. The method may further include extracting the distortion metadata from the bitstream. The method may further include applying transcoding (e.g., a transcoding operation) based on the coded time series data to obtain transcoded time series data. The method may further include determining an updated measure of distortion in relation to a portion of the transcoded time series data based on the measure of distortion. The method may yet further include generating a transcoded bitstream including the transcoded time series data and updated distortion metadata. The updated distortion metadata may include the updated measure of distortion. The updated measure of distortion may further depend on characteristics / properties of the transcoding. The updated measure of distortion may relate to a portion of the transcoded time series data that is obtained from the portion of the coded time series data by the transcoding.
[0036] Configured as defined above, the proposed method ensures that distortion metadata is carried through in the process of transcoding and is appropriately updated. Thereby, even time series data that is transcoded, for example to conserve storage space, may be meaningfully assessed with regard to its suitability for certain (automated) processing tasks after eventual decoding.
[0037] In some embodiments, the coded time series data may be lossless coded time series data. Further, the transcoding may include decoding the coded time series data to obtain decoded time series data. Then, determining the updated measure of distortion may be based on the decoded time series data. The method may further include determining that the coded time series data is lossless coded time series data based on the distortion metadata. Determining the updated measure of distortion may use a locally decoded version of the transcoded time series data or may relate to estimating the measure of distortion without local decoding of the transcoded time series data, as described above.
[0038] In some embodiments, the coded time series data may be lossy coded time series data. The distortion metadata may further include information indicative of a variance of original time series data relating to the portion of the coded time series data. The method may further include determining, based on the distortion metadata, that the measure of distortion is below a predetermined threshold. The method may yet further include estimating a quantization error that results from the transcoding. Then, determining the updated measure of distortion comprises estimating the updated measure of distortion based on the variance of the original time series data and the estimated quantization error.
[0039] In some embodiments, the coded time series data may be lossy coded time series data. The method may further include determining, based on the distortion metadata, that the measure of distortion is above a predetermined threshold. Then, determining the updated measure of distortion may include setting the measure of distortion to a value that indicates high distortion or that indicates that estimating the measure of distortion is not possible.
[0040] Another aspect of the disclosure relates to a method of decoding a bitstream. The bitstream may include coded time series data and distortion metadata. The distortion metadata may include information indicative of a measure of distortion in relation to a portion of the coded time series data. The method may include decoding, from the bitstream, decoded time series data. The method may further include decoding, from the bitstream, the information indicative of the measure of distortion. The method may yet further include outputting the decoded time series data and the information indicative of the measure of distortion. The measure of distortion may relate to an unnormalized squared error over the portion of coded time series data, for example. The measure of distortion may be given on a per-channel basis.
[0041] Configured as defined above, the proposed method ensures that decoded time series data can be classified with regard to its coding distortion and analyzed with regard to its suitability for certain (automated) processing tasks, such as serving as training data for training a neural network, etc. In some embodiments, the coded time series data may include lossy coded time series data. For example, the bitstream may include segments of lossy coded time series data and segments of lossless coded (or near lossless coded) time series data.
[0042] In some embodiments, the distortion metadata may further include information indicative of a variance of original time series data corresponding to the portion of the coded time series data. The method may further include determining a percentage root mean square distortion, PRO, or a channel-normalized percentage root mean square distortion, CPRD, based on the measure of distortion and the variance. The variance may be a variance of sample values of samples of the original time series data. The variance may be given on a per-channel basis, for example.
[0043] In some embodiments, the method may further include determining a mean square error, MSE, for the portion of decoded time series data based on the measure of distortion.
[0044] In some embodiments, the method may further include determining a peak signal to noise ratio, PSNR, value for the portion of decoded time series data based on the measure of distortion. The PSNR value may be a per-channel value or an average (e.g., channel average) value, for example.
[0045] In some embodiments, the distortion metadata may include explicit signaling for the measure(s) of distortion. For example, the distortion metadata may include one or more explicit values for the measure(s) of distortion. Here, explicit distortion signaling may mean that on the encoder side, a measurement or estimation of a distortion (e.g., according to a selected distortion metric) is performed, resulting in a value of the distortion, and that this value of the distortion is explicitly inserted into the bitstream. This facilitates evaluation of coding distortion on the decoder side by only inspecting the decoded bitstreams, without need of knowing the (uncoded) reference signal.
[0046] In some embodiments, the distortion metadata may include implicit signaling for the measure of distortion via a parameterization of the measure(s) of distortion, possibly in addition to explicitly signaled values of the measure(s) of distortion. Here, implicit distortion signaling may mean that on the encoder side, a measurement or estimation of parameters (e.g., a, p) is performed, which are inserted into the bitstream. These parameters may facilitate derivation of several distortion metrics on the decoder side. For instance, given the values of these parameters, the values of several distortion metrics can be computed on the decoder side. On the other hand, no single value of distortion may be transmitted explicitly. This facilitates evaluation of coding distortion on the decoder side by only inspecting the decoded bitstreams, without need of knowing the (uncoded) reference signal.
[0047] In some embodiments, the distortion metadata may include an indication of a distortion level from among a predefined set of distortion levels. Optionally, the distortion metadata may include an indication that the measure of distortion has been quantized according to the predefined set of distortion levels.
[0048] In some embodiments, the method may further include quantizing the measure of distortion according to a predefined set of distortion levels.
[0049] In some embodiments, the decoded time series data may include a plurality of portions (e.g., segments, groups of frames) of decoded time series data. Then, the distortion metadata may include a respective measure of distortion for each of the plurality of measures of distortion, or in some implementations, a respective measure of distortion for each lossy coded portion. In the latter case, absence of a measure of distortion for a certain portion may indicate zero or negligible distortion for that portion. The method may further include classifying the plurality of portions of the decoded time series data in accordance with their respective measures of distortion. The method may further comprise outputting a result of classification. This may be done for each portion of decoded time series data for which a classification result (or distortion metadata in general) is available.
[0050] In some embodiments, the method may further include applying a filtering operation to the plurality of portions of decoded time series data based on a result of classification. In some implementations, the filtering may be applied directly based on the associated measure of distortion. The filtering operation may relate to, for each applicable portion of decoded time series data, applying a “keep or toss” logic, or to assigning different portions to different buckets for further processing, depending on the result of classification or measure of distortion.
[0051] In some embodiments, the method may further include deciding whether to use the portion of decoded time series data for training an Al-based model based on the measure of distortion. For example, the portion of decoded time series data may only be used for training the Al-based model (e.g., neural network-based model) if the measure of distortion is below a predetermined threshold. According to another aspect, an apparatus is provided. The apparatus may include one or more processors and a memory coupled thereto and storing instructions for the one or more processors. The one or more processors may be configured to perform the methods or method steps outlined throughout the present disclosure. This apparatus may relate to an encoder, encoding apparatus, or encoding system, to a transcoder, transcoding apparatus, or transcoding system, or to a decoder, decoding apparatus, or decoding system, as the case may be. The encoder, encoding apparatus, or encoding system may relate to or be part of a wearable device, such as a medical tracker, for example.
[0052] According to a further aspect, a computer program is described. The computer program may comprise executable instructions for performing the methods or method steps outlined throughout the present disclosure when executed by a computing device (e.g., one or more processors).
[0053] According to another aspect, a computer-readable storage medium is described. The storage medium may store a computer program adapted for execution on a computing device (e.g., one or more processors) and for performing the methods or method steps outlined throughout the present disclosure when carried out on the computing device.
[0054] It should be noted that the methods and apparatus including their preferred embodiments as outlined in the present disclosure may be used stand-alone or in combination with the other methods and apparatus disclosed in this document. Furthermore, all aspects of the methods and apparatus outlined in the present disclosure may be arbitrarily combined. In particular, the features of the claims may be combined with one another in an arbitrary manner.
[0055] It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus, and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) (and, e.g., their steps) are understood to likewise apply to the corresponding apparatus (and, e.g., their blocks, stages, units), and vice versa. Brief Description of the Drawings
[0056] The invention is explained below in an exemplary manner with reference to the accompanying drawings, wherein
[0057] Fig. 1 is a flowchart schematically illustrating an example of a method of generating or modifying a bitstream according to embodiments of the disclosure;
[0058] Fig. 2 schematically illustrates an example of a bitstream of coded data to which embodiments of the disclosure may be applied;
[0059] Fig. 3 is a block diagram of an example of an encoder framework according to embodiments of the disclosure;
[0060] Fig. 4 is a flowchart schematically illustrating an example of a method of modifying a bitstream according to embodiments of the disclosure;
[0061] Fig. 5 is a block diagram schematically illustrating a process flow including generation and subsequent modification of a bitstream according to embodiments of the disclosure;
[0062] Fig. 6 is a flowchart schematically illustrating an example of a method of decoding a bitstream according to embodiments of the disclosure;
[0063] Fig. 7 is a block diagram of an example of a decoder framework according to embodiments of the disclosure;
[0064] Fig. 8 is a block diagram schematically illustrating a process flow including generation and subsequent decoding of a bitstream according to embodiments of the disclosure;
[0065] Fig. 9 schematically illustrates an example of a high-level bitstream syntax for a bitstream according to embodiments of the disclosure;
[0066] Fig. 10 schematically illustrates an apparatus suitable for implementing techniques according to embodiments of the disclosure; and
[0067] Fig. 11 schematically illustrates an example of a transmission device according to embodiments of the disclosure. Detailed Description
[0068] The International Telecommunication Union (“ITU”) is an assembly of experts from around the world that work together to develop international standards known as ITU-T Recommendations. These standards enable improved interoperability of communication signals in the global infrastructure of information networks and communication devices, allowing such networks and devices to more easily communicate and operate together. Recognizing the need for a standardized codec for biomedical waveform data, in April 2024 the ITU-T issued a call for proposals for a new ITU-T Recommendation on the coding of biomedical waveform data. One or more of the embodiments disclosed herein describe methods, apparatuses, and systems developed to meet one or more requirements specified in the ITU-T call for proposals for improved coding, compression, storage, transmission, reception, and / or decoding of such signals.
[0069] Technical benefits of one or more of the embodiments disclosed herein enable a standardized lossy, lossless, and / or near-lossless coding format including a transmission-syntax specifically developed for biomedical waveform data, and facilitate clinical neurophysiology data exchange. Such data may include time-based neurophysiology signal data and associated video recordings, if present, from electroencephalography (EEG), video-electroencephalography (VEEG), electromyography (EMG), evoked potentials (EP), polysomnograms (PSGs), electrocardiograms (ECGs), and other types of neurophysiology signals. Additional non-exhaustive examples of biomedical waveform data include photoplethysmogram (PPG). One or more of the embodiments described herein therefore provide features for a standardized codec which facilitates interoperable processing of biomedical waveform data by a wide range of devices.
[0070] In the following, example embodiments of the disclosure will be described with reference to the appended figures. Identical elements in the figures may be indicated by identical reference numbers, and repeated description thereof may be omitted.
[0071] It has been found that lossy coding of time series data relating for example to utility signals may be, while conventionally not explored, of interest for many frameworks and use cases. When applying lossy coding, the coding distortion will typically not only depend on the bitrate, but also on a particular realization of the signal that was encoded. This dependence is generally unknown on the decoder side. Also, if lossy coding was used, it may happen that the coding distortion is prohibitive to accomplish a given decoder-side task (e.g., training of a neural network based on the reconstructed signal). However, this typically cannot be inferred on the decoder side by inspecting the decoded bitstreams alone. Therefore, since the performance of coding depends not only on the operating bitrate, but also on the specific realization of the to-be-coded signal, the present disclosure relates to including suitable side information into the bitstream for characterizing the performance of the encoder (e.g., in terms of distortion) on this realization of the signal.
[0072] Furthermore, the encoding process may utilize many different distortion measures to provide the optimal trade-off point between the distortion of the reconstructed point and the bitrate. It could happen that the distortion measure used by the encoder is not optimal from the point of view of inspecting a collection of signals (e.g., in a medical record) for their usability for a specific task. Therefore, the present disclosure seeks to provide side information included in the bitstream that facilitates derivation of several distortion measures.
[0073] Traditionally, medical signals as an example of utility signals would be encoded in a lossless manner. Once lossy coding is applied to all or portions (e.g., segments) of the utility signals, the present disclosure suggests embedding suitable metadata into the coded bitstreams, which would help to characterize the coding distortion. This enables estimation of coding performance just by inspecting the decoded signals. This would be beneficial to a large variety of workflows involving creating collections of coded bitstreams, which are then analyzed without access to the uncoded reference versions of these signals (for example Digital Imaging and Communication in Medicine (DICOM) workflows, or workflows used for signal acquisition from medical devices (e.g., fitness trackers)).
[0074] Broadly speaking, the present disclosure thus relates to embedding distortion metadata into bitstreams that include coded time series data at an encoder stage, for informing any subsequent coding stages of the coding distortion of the time series data. The present disclosure further relates to updating distortion metadata for example in transcoding, and to further processing of the distortion metadata in decoding for assessing suitability of decoded time series data for certain (possibly automated) processing tasks.
[0075] Fig. 1 is a flowchart schematically illustrating an example of a method 100 of generating or modifying a bitstream according to embodiments of the disclosure. Method 100 comprises steps SI 10 to S130 that may be performed, for example, in the context of encoding of a bitstream. At step S110, the bitstream is obtained. It is understood that the bitstream contains coded (e.g., core coded) time series data. This time series data may relate to one or more utility signals. For example, the time series data may relate to one or more medical or biomedical signals. Likewise, the time series data may relate to waveform data, for example to one or more channels of waveform data.
[0076] In particular, the bitstream may comprise a plurality of channels that may be grouped into one or more (e.g., a plurality) of channel groups. Channels that are grouped into a given channel group may be identified by a channel group identifier of the given channel group that may be carried by the payload packets of the channels. Grouping of channels into channel groups may be performed to account for (e.g., responsive to) core coder requirements, for example to enable joint (core) coding of channels in a given channel group. For instance, channels that are grouped into a given channel group may share some degree of similarity and / or may relate to the same type of waveform data (e.g., ECG data, EEG data, EMG data, PPG data, etc.).
[0077] Time series data in the context of the present disclosure may comprise a series of data points in time order. In some cases, the data points may represent measurements performed by the sensor, for example, values of an observed signal (e.g., waveform signal), or measured quantity. In some cases, the data points may be obtained by sampling, and they may correspond to sample values provided by the sensor. The sampling may be uniform (according to the sampling frequency of the sensor) or non-uniform. The sample values may be provided as a sequence of floating point values, or as a sequence of integer values, for example.
[0078] A biomedical signal in the context of the present disclosure may be a signal that relates to physiological information, and the signal may be electrical, physical or biochemical. The biomedical signal may relate to biological systems and conditions, examples of which may include Electrocardiography (ECG) data, Electroencephalography (EEG) data, Electromyography (EMG) data, and Photoplethysmogram (PPG) data, or signals for blood sugar level, heart rate, body temperature, respiratory rate and oxygen saturation. Further, a biomedical signal may comprise one or more channels of time domain biomedical signal samples. In other examples, the biomedical signals may relate to muscle and / or skin measurements. Any other medical signal and / or physical response would also be understood to be comprised by this definition. A utility signal in the context of the present disclosure may be a signal that is captured with purpose different from playing it out to a human observer / listener. For example, a smartwatch capturing a PPG signal may also capture the signal from its accelerometer (i.e., a mechanical signal), and the acceleration signal may be then used to aid filtering of the PPG signal.
[0079] A mechanical signal may be any signal that can be captured by a sensor of mechanical movement, where the signal can be represented as time series data (e.g., acceleration signals, seismic signals, etc.).
[0080] Further, waveform data in the context of the present disclosure may be based on mechanical motion or may indicate a biomedical signal of a human body. Without intended limitation, in the mechanical context, the waveform data may be data from a seismometer or an accelerometer. For the biomedical signal context, the waveform data may correspond to the examples listed above.
[0081] Obtaining the bitstream may relate to generating the bitstream by an encoding operation (e.g., core encoding) or to receiving the bitstream from another source.
[0082] At step SI 20, a measure of distortion in relation to a portion of the coded time series data is determined. The distortion may be a coding distortion.
[0083] In the context of the present disclosure, distortion or coding distortion is understood to mean the following. Lossy coding of a signal trades off bitrate vs fidelity of signal reconstruction. In other words, if the bitrate to represent the signal is reduced, the decoded signal will only be approximating the encoder’s input signal, and, in general, will differ from the encoded signal. Typically, a codec should maximize the fidelity of the approximation, so that the achievable fidelity will typically be a function of the bitrate. In practice, the fidelity of the reconstruction may be described by a mathematical expression evaluating similarity between the unencoded version of the encoded signal and the reconstruction of that signal performed by the decoder. Several different mathematical expression may be used for this purpose. While a number of examples of distortion measures suitable for evaluating the coding performance is explicitly given throughout the disclosure, this is understood to be without intended limitation.
[0084] Within the portion of the coded time series data, the measure of distortion may be determined on a per channel basis or on a per channel group basis. Alternatively, an aggregate measure of distortion may be determined for all channels in the portion of the bitstream. In some implementations, the determined measure of distortion may be quantized according to a predefined set of distortion levels. These distortion levels may be related to or indicated by respective indices (distortion indices), and / or may be related to qualitative measures such as “good,” “average,” or “poor,” for example, with appropriate granularity.
[0085] At step SI 30, distortion metadata is embedded into the bitstream. This distortion metadata includes information indicative of the determined measure of distortion. The distortion metadata may further include one or more of an indication of a length of the portion of coded time series data and an indication of a variance of original time series data corresponding to the coded time series data. Indicating the length of the portion of coded time series data will not be necessary if the portion of coded time series data relates to a fixed-length portion. One non-limiting example of such fixed-length portion is a segment including an integer multiple of random access intervals, which will be described in more detail below.
[0086] For example, the embedding at step S130 may comprise embedding a data packet into the bitstream that includes an indication of the measure of distortion and optionally an indication of the corresponding portion of the coded time series data. The indication of the corresponding portion of the coded time series data may be a pointer to said portion (e.g., segment, group of frames) or an indication of a length of said portion, for example. In some embodiments, the data packet may be embedded at the end of the respective portion (e.g., segment, group of frames). In this case, including the indication of the corresponding portion of the coded time series data may not be necessary.
[0087] In the above, the coded time series data may comprise lossy coded time series data. For example, the bitstream may include portions (e.g., segments) that are lossy coded, as well as portions (e.g., segments) that are coded in a lossless (or near lossless) manner. In such case, the core coder may seamlessly switch between lossy and lossless coding (e.g., given its bitrate constraint), depending on the realization of the time series data that is to be coded. In one implementation example thereof, the core coder may operate in a near-lossless setting (e.g., where permitted by the bitrate constraint), where some portions (e.g., segments) may be encoded in a lossless fashion. Where necessary, the core coder may switch to a lossy setting.
[0088] The distortion metadata may be included for both types of portions (e.g., segments). For the lossless ones, it may indicate that there is zero distortion or that lossless coding has been used. Otherwise, without the distortion metadata, operation in the lossless setting may not be apparent just by inspecting the core coded bitstreams alone. Alternatively, distortion metadata may be included (only) for lossy portions of coded time series data and the absence of distortion metadata for a given portion of coded time series data may be interpreted at decoding as an indication that lossless coding (or near lossless coding) has been used for said portion.
[0089] If the measure of distortion has been quantized according to the predefined set of distortion levels at step S120, the distortion metadata may include information indicative of the distortion level to which the measure of distortion has been quantized. Optionally, the distortion metadata may include an indication (e.g., flag) that the measure of distortion has been quantized according to the predefined set of distortion levels.
[0090] Although not shown in Fig. 1, method 100 may further comprise a step of performing encoding to generate the coded time series data, or step SI 10 may relate to generating the coded time series data. For example, method 100 may further include a step of receiving an output (e.g., original, or uncoded, time series data) of an apparatus, (e.g., sensor) and applying a coding operation thereto to generate the coded time series data.
[0091] Fig. 2 schematically illustrates an example of a bitstream 200 of coded (e.g., core (en-)coded) time series data and of a temporal scope of the distortion metadata generated for the core coding bitstream. The bitstream 200 may include a plurality of frames (or blocks) 210 of coded (e.g., core (en-)coded) time series data. The temporal scope of the distortion metadata may include single or multiple frames of coded time series data, or a segment 220 of one or more frames of coded time series data. This segment 220 of frames may correspond to the portion of coded time series data defined at step S120 above. In other words, the aforementioned portion of the coded time series data may relate to a segment including integer multiples of frames (or blocks) of coded time series data. In some cases, the portion of the coded time series data may relate to the entire bitstream.
[0092] Furthermore, the distortion metadata may be embedded into the bitstream in an asynchronous manner with respect to the frames of the core coded signals. The frames of the core coded signal do not need to be uniform in length. The coding distortion typically may be indicated for relatively large segments of the signal (e.g., in the order of seconds). Notably, the coding distortion may be generally indicated at the level of frames / blocks (core coder frames / blocks), or with a coarser granularity.
[0093] In some implementations, the portion of the coded time series data may relate to a segment that includes an integer multiple of (i.e., one or more) random access intervals.
[0094] In general, a waveform codec (as an example of a core codec or source codec) may perform encoding of a signal (e.g., a waveform signal, or a signal relating to time series data) using a predefined set of frame sizes, where each frame can be either coded independently or can be coded based on the previously encoded frames. The (fixed) expected interval between two independently decodable frames is referred to as the intra-period at the encoder side. This does not prevent an encoder from sending other independently decodable frames within an intra- period. It is, however, expected that an independently decodable frame follows the next intra- period interval, meaning at least one independently decodable frame exists in an intra-period. A random-access interval may comprise a fraction, single, or multiple of such intra-periods. The random-access interval provides an independently decodable segment of the encoded signal.
[0095] Furthermore, the waveform codec may subdivide every frame into blocks (bitstream elements), where the partition of frames into blocks may differ on a per-frame basis. A particular partition into blocks results in a sequence of blocks. If such sequence of blocks is used to encode an independently decodable frame, its first block must be encoded independently, i.e., may be an independent block. A frame where an independent block resides may be called an independent frame (IF). All subsequent blocks may be encoded based on (e.g., with reference to) the previously encoded block(s), i.e., may be dependent blocks. A frame consisting of one or more dependent block(s) may be called a dependent frame (DF).
[0096] Accordingly, the random access interval (e.g., a single intra period) may comprise one independent block (e.g., independent frame block) that can be decoded without reference to other blocks, and zero or more dependent blocks (e.g., dependent frame blocks) that can only be decoded with reference to other blocks in the random access interval. In some cases, the random access interval may include one or more frames, of which the first comprises an independent block followed by zero or more dependent blocks, and of which the subsequent frames comprise zero independent blocks and one or more dependent blocks. At the decoder side, a sequence of one independent frame followed by zero or more dependent frames may be called a frame sequence. In any case, bitstream portions corresponding to the random access interval can be independently decoded, as noted above, while this may not be true for arbitrary ones of its blocks or frames.
[0097] In this configuration, while the individual blocks may have variable length, the random access interval and accordingly also the segment have a fixed length. Also the aforementioned frames within the random access interval may have a fixed length (e.g., a length among the aforementioned set of frame sizes).
[0098] Notably, the above definition of the segment includes the extreme cases of segments each corresponding to a single independent block, as well as the case of a single segment corresponding to the entire bitstream. That is, the portion of the coded time series data may relate, inter alia, to an individual (independent) block, or to the entire bitstream.
[0099] In line with the above segment definition, without intended limitation, the following implementations may be of particular interest in the context of the present disclosure.
[0100] (1) In one implementation, the distortion metadata may be updated over extents of the randomaccess interval, which could comprise multiple frames. Each frame could in turn consist of multiple blocks (e.g., core coder blocks). In other words, the random access interval (and thus the portion of coded time series data) may comprise a plurality of blocks (bitstream elements).
[0101] (2) In another implementation, which relates to a special case of above implementation (1), the random-access interval comprises a single frame, which consists of a single independent (IF) block (or single block (e.g., core coder block) in general). In other words, the random access interval (and thus the portion of coded time series data) may consist of a single block (bitstream element).
[0102] (3) in yet another implementation, the random access interval may correspond to the entire bitstream.
[0103] A possible use case for implementations (1) and (2) may be as follows. A workflow may involve pulling encoded signals from a remote database, where the unit of signal that is pulled corresponds to one or more random access intervals comprised in a continuous segment of the signal. The distortion metadata associated with such a segment may be used, for example, to decide whether the workflow should use or discard a specific segment. While the above segment definition assumes a fixed-length segment, the present disclosure also is understood to cover potentially non-uniform length of the segments (or portions of coded time series data in general) in relation to which the distortion metadata is provided. In this case, the distortion metadata may further include information indicative of a length of the portion of the coded time series data to which the measure of distortion applies, as indicated earlier.
[0104] Fig. 3 is a block diagram of an example of an encoder framework 300 according to embodiments of the disclosure. An input signal 310 (e.g., original time series data) is provided to a core encoder 320. The core encoder 320 generates a core encoded bitstream 330. At the same time, the core encoder 320 generates side information 360 that may be used for deriving the distortion metadata 380. The core encoder 320 may generate the side information 360 on a per frame basis (e.g., core coder frame) or a per block basis. This side information 360 may be used to compute the distortion metadata 380 for larger portions of the coded signal. The distortion metadata 380 finally is embedded into the bitstream. For example, the core encoded bitstream 330 and the distortion metadata 380 may be provided to a multiplexer 340, which outputs a multiplexed bitstream 350 that includes the coded time series data and the distortion metadata 380.
[0105] Different implementations are feasible for determining (e.g., at step S120) the measure of distortion that is transported in the distortion metadata.
[0106] Determination of Coding Distortion by Local Decoding
[0107] In general, the coded time series data relates to original time series data, and the coded time series data may be obtainable from the original time series data via a coding (e.g., core (en-) coding) operation. The original time series data is typically available for analysis in all encoding scenarios where the original time series data is first encoded.
[0108] Availability of the original time series data may be leveraged in a first implementation of determining the measure of distortion (e.g., at step S120). According to this implementation, method 100 may further include local decoding of the coded time series data to obtain locally decoded time series data. Then, determining the measure of distortion (e.g., at step S120) may be based on (one or more samples of) the original time series data and (one or more samples of) the locally decoded time series data. For example, the measure of distortion may correspond to or be based on a (per-channel) squared error as defined in Eq. (3) below. Estimation of Coding Distortion During Encoding
[0109] On the other hand, in some cases it may be beneficial in terms of encoder complexity and memory consumption to avoid local decoding of the coded time series data. A second implementation of determining the measure of distortion (e.g., at step S120) that avoids local decoding will be described next.
[0110] Also here, the coded time series data may be generated from original time series data via a coding operation. However, instead of performing local decoding, the measure of distortion in relation to the portion of the coded time series data may be estimated (e.g., at step S120) based on a rate allocation process of the coding operation. In other words, the encoder may estimate the coding distortion during its rate allocation process, based on side information that is available to the encoder.
[0111] According to a first example of the second implementation, the coded time series data may be generated from the original time series data via the coding operation (e.g., core (en-)coding operation), subject to a bitrate constraint. Therein, the coding operation may be adapted to minimize a signal distortion given the bitrate constraint. In this case, the distortion per frame of the core codec will typically vary, depending on the actual realization of the input signal (original time series data) to the encoder. The measure of distortion may then be estimated in the course of minimizing the signal distortion given the bitrate constraint. This may, for example, relate to estimating the (per-channel) squared error p? .
[0112] For instance, the encoding algorithm (core encoding algorithm) used for coding the time series data may be characterized by a set of coding tools (e.g., signal quantizers, prediction tools, parametric signal description, etc.). The process of selection of such tools will involve an attempt to minimize signal distortion given the bitrate constraint. As a by-product of encoder operation, an estimate of coding distortion may be provided (e.g., in the form of an expected squared error), for example based on parameters of the selected coding tools.
[0113] According to a second example of the second implementation, the coded time series data may be generated from the original time series data via the coding operation (e.g., core (en-)coding operation), subject to an average bitrate constraint. Therein, the coding operation may be adapted to optimize a rate-distortion trade-off given the average bitrate constraint, by selection of appropriate operation points of the coding operation. Then, the measure of distortion may be estimated in the course of selecting the appropriate operation points.
[0114] For instance, the encoding algorithm (e.g., core encoding algorithm) may be designed in a way that the possible operating points of the encoder are located on a convex function (e.g., a monotonically decreasing convex function, i.e., distortion decreases as bitrate increases). The number of operating points may be very large and may depend on the number of degrees of freedom provided by the encoding algorithm (e.g., associated with selection of quantization step sizes, parameters used to represent the signal, etc.). During the encoding, the encoding algorithm may attempt to efficiently find an appropriate operating point by performing a binary search over the set of points. As a by-product of this process, the coding distortion may be estimated for a segment of the signal, without need for an explicit local decoding operation.
[0115] In a third example of the second implementation, the coding operation (e.g., core encoding) for obtaining the coded time series data may relate to distortion-constrained coding using a distortion constraint. In such scenario, the encoder will use a variable number of bits per frame, depending on the realization of the input signal (original time series data). The measure of distortion may be determined based on the distortion constraint. In fact, the measure of distortion (or the distortion metadata) could be directly provided by the encoder according to the distortion constraint it operates with. As an example, the measure of distortion may be indicative of the distortion constraint. The distortion constraint may be given in relation to a given distortion metric, and in some cases, the constraint may be provided using a different metric than the squared error p?, (e.g., the maximum PRD may be used as a constraint). To account for such cases, the measure of distortion (or the distortion metadata) may optionally be indicative of the distortion metric.
[0116] In some embodiments, the core encoding algorithm may attempt to estimate the coding distortion which would be a function of, for example, the quantization step size. Given the step size, the coding distortion may be estimated. For example, the core coding algorithm may employ a transform approximating a decorrelating transform (e.g., DCT), where scalar quantizers may be used to quantize the transformed signal, and the squared error may be the distortion metric of choice. In this case, the bitrate allocation process may involve, for example, a so-called “reverse water-filling process”, which facilitates optimization of a rate distortion trade-off. As a byproduct, the reverse water-filling process may provide an estimate of the coding distortion. In other words, an encoder employing such a reverse-water filling process may supply an estimate of the coding distortion.
[0117] It is to be noted that the estimation of the distortion by the core encoder will typically happen with a granularity of a single core coder frame (which may correspond to a time unit for which the decisions made by the encoding algorithm are made (e.g., in the constrained bitrate example)), or over an extent of multiple core coder frames (e.g., in the constrained average bitrate example). According to embodiments of the disclosure, the distortion information (e.g., measure of distortion) may be recomputed to correspond to a different segment length, for example, one that is relevant to the application of interest.
[0118] Content of Distortion Metadata: Example 1
[0119] Next, possible content of the distortion metadata according to embodiments of the disclosure will be described. According to a first example, a parametric representation of multiple distortion measures is generated, and the distortion measures are transmitted implicitly. This may correspond to the aforementioned implicit distortion signaling.
[0120] As noted above, the coded time series data may comprise one or more channels of coded time series data. In this context, the following definitions may be made.
[0121] Let Abe the number of channels and let be the number of samples per channel of an input sequence of time series data. Furthermore, let j be the j-th sample (with 0 < j < M) of channel i (with 0 < i < N) and let be the corresponding reconstructed sample after decoding a bitstream. The maximum absolute error (MAE) per channel may then be defined for example as
[0122] Moreover, for m; defined as the mean value of the input data for the i-th channel, i.e. a per-channel squared error may be defined for example as and a per-channel variance may be defined for example as
[0123] With, for example, the above definitions, the measure of distortion (e.g., determined at step S120) may relate to at least one of a maximum absolute error per channel and a per-channel squared error. The distortion metadata may further include information indicative of a variance of the original time series data relating to the portion of the coded time series data. The variance may be a variance of sample values of samples of the original time series data. The variance may be given on a per-channel basis, for example.
[0124] As to the length L of the portion of the coded time series data (e.g., segment) in relation to which the distortion metadata is provided, it may be assumed for example that M is the length of the core coded frame and that the above parameters are updated by the encoder on a per frame basis. The length L of the measurement segment may be introduced and may be typically greater (e.g., much greater) than M. Then, a moving average step (or any other suitable averaging step) could be performed to convert the per-M measurements performed by the encoder into per-L measurements to be included into the distortion metadata in the bitstream.
[0125] As noted above, the estimation of the per-channel squared error may be performed by performing a local decoding operation in the encoder. Alternatively, the encoder may attempt to estimate the squared error without decoding the signal.
[0126] In summary, according to the present disclosure the distortion metadata shall be generated for segments (or portions of the coded time series data in general, possibly spanning multiple frames of coded time series data), where at least a single frame is not coded in a lossless manner.
[0127] The distortion metadata may include for example: a length L of the measurement segment in samples (e.g., expressed in seconds, with quantization step-size of a second, given sampling frequency of the signal, etc.), • af : variance of the input signal (e.g., original time series data), possibly on a per-channel basis over the segment of length L (e.g., represented in log domain and quantized), and
[0128] • p : estimated unnormalized squared error, possibly on a per-channel basis over the segment of length L (e.g., represented in log domain), or any other suitable measure of distortion.
[0129] As an alternative to indicating a quantitative measure of distortion, such as the (unnormalized) square error, the distortion metadata may include an indication (e.g., index) of a distortion level among a predefined set of distortion levels to which the measure of distortion has been quantized. This may be in the form of indices, for example corresponding to distortion levels ranging from “good” to “poor,” with appropriate granularity.
[0130] Indicating the length L in the distortion metadata may not be necessary for a fixed-length portion of the coded time series data, such as the segment defined above as an integer multiple of the random access interval.
[0131] Additionally, in some embodiments, the distortion metadata may include for example:
[0132] • information about the maximum absolute error (MAE) (e.g., reconstructed signal vs input signal), possibly computed on a per-channel basis over the segment of length L.
[0133] Furthermore, in some embodiments, the parameters af and / or p? may be quantized to reduce the number of bits needed to carry their values in the bitstream. For example, the parameters af and / or p? may be normalized with respect to L, and converted to a log domain (e.g., 10 log10(-) domain) as, for example 8 bit int values, where a coarse quantization (e.g., 1 dB resolution) may be used.
[0134] Furthermore, in some embodiments, it may be envisaged to indicate the type of distortion measure that was used in the encoder to optimize its trade-off between the reconstruction quality and the bitrate in the distortion metadata, if applicable.
[0135] Content of Distortion Metadata: Example 2
[0136] According to a second example, the distortion measures are transmitted explicitly. This may correspond to the aforementioned explicit distortion signaling. In general, many different distortion measures may be used. For example, in the case of ECG signals, some of the popular distortion metrics are listed in Nemcova, Andrea, et al. "A comparative analysis of methods for evaluation of ECG signal quality after compression." BioMed Research International 2018 (https: / / pmc.ncbi.nlm.nih.gov / articles / PMC6077674 / .
[0137] According to the second example, we let Abe the number of channels, and consider a segment of a multichannel signal of length L. For such a segment, the distortion metadata may include an indication of a distortion metric that is used. For instance, the set of possible distortion metrics may include one or more of:
[0138] • percentage root mean square distortion (PRO)
[0139] • channel-normalized percentage root mean square distortion (CPRD)
[0140] • signal-to-noise-ratio (SNR), or a variant of SNR (e.g., SNR1, SNR2, etc.)
[0141] • peak-signal-to-noise-ratio (PSNR)
[0142] • percentage similarity (PSim SDNN)
[0143] • quality score (QS)
[0144] • mean squared error (MSE)
[0145] • normalized PRO (e.g., PRDN1, PRDN2, PRN3, etc.)
[0146] • Maximum amplitude error (MAX)
[0147] • Standard error (STDERR)
[0148] • Wavelet-Energy based Diagnostic Distortion (WEDD SWT)
[0149] For a given distortion metric, the value of the distortion metric can be derived (e.g., by estimation or by performing an explicit measurement on the encoder side). For a segment of length £ of a multichannel signal, the distortion values may be derived either on a per-channel basis, or a single value may be derived for a set of N channels. A single bit may be used to indicate availability of per-channel or channel-averaged distortion values.
[0150] Tandem coding may be of interest for improving compression efficiency of coded time series data (e.g., data sets including medical signals). This may involve transcoding from lossless to lossy, or from lossy to lossy. It may be desired to still be able to generate the distortion metadata for such transcoding scenarios. Fig. 4 is a flowchart schematically illustrating an example of a method 400 of modifying a bitstream according to embodiments of the disclosure. This modification of the bitstream may relate for example to transcoding of the bitstream of coded time series data. Method 400 comprises steps S410 to S450 that may be performed, for example, in the context of transcoding a bitstream. It is noted that the order of steps is not necessarily limited to the example order given below and that rearrangements may be made as long as these rearrangements do not interfere with the flow of information.
[0151] At step S410, the bitstream is obtained as an input. The bitstream includes coded time series data and distortion metadata. The distortion metadata in turn includes information indicative of a measure of distortion in relation to a portion of the coded time series data.
[0152] At step S420, the distortion metadata is extracted from the bitstream. This may involve demultiplexing the bitstream, for example.
[0153] At step S430, transcoding (e.g., tandem coding) is applied based on the coded time series data to obtain transcoded time series data.
[0154] At step S440, an updated measure of distortion in relation to a portion of the transcoded time series data is determined based on the measure of distortion.
[0155] The updated measure of distortion may further depend on characteristics / properties of the transcoding operation. The updated measure of distortion may relate to a portion of the transcoded time series data that is obtained from the portion of the coded time series data by the transcoding operation.
[0156] In some implementations, the determined updated measure of distortion is quantized according to a predefined set of distortion levels. These distortion levels could be related to or indicated by respective indices (distortion indices), and / or could be related to qualitative measures such as “good,” “average,” or “poor,” for example. At step S450, a transcoded bitstream including the transcoded time series data and updated distortion metadata is generated. Therein, the updated distortion metadata includes the updated measure of distortion.
[0157] It is understood that the updated distortion metadata may include the same information items as the distortion metadata, but now in relation to a portion of the transcoded time series data. Different implementations are feasible for determining the updated measure of distortion (e.g., at step S440) that is transported in the updated distortion metadata.
[0158] Transcoding of Lossless Segments
[0159] The first implementation of determining the updated measure of distortion relates to a case in which the (relevant portion of the) coded time series data is lossless coded time series data, and the transcoded time series data is lossy-coded time series data. Performing lossy coding as a second coding may be of interest, for example, where storage space for storing the coded time series data is running out.
[0160] In this case, the transcoding may comprise decoding the coded time series data to obtain decoded time series data. Since the time series data had been coded in a lossless manner, the decoded time series data may be used as reference for determining the updated measure of distortion. Thus, determining the updated measure of distortion may be based on the decoded time series data. That the coded time series data had indeed been lossless may be inferred for example from an analysis of the distortion metadata included in the bitstream.
[0161] The actual determination of the updated measure of distortion may use a locally decoded version of the transcoded time series data and may involve comparing it to the decoded time series data, or may relate to estimating the measure of distortion without local decoding of the transcoded time series data, as described above in the context of encoding.
[0162] Transcoding of Lossy Segments with Small Distortion
[0163] According to a second implementation of determining the updated measure of distortion, the (relevant portion of the) coded time series data may be lossy coded time series data, but with comparatively low coding distortion. This may relate to a so-called high bitrate case.
[0164] Here, it is assumed that the distortion metadata further includes information indicative of a variance of original time series data relating to the portion of the coded time series data.
[0165] Given the above, the present implementation of method 400 may further comprise determining, based on the distortion metadata, that the measure of distortion is below a predetermined threshold. If so, a quantization error that results from the transcoding may be estimated in a subsequent step. Specifically, the updated measure of distortion may be estimated (e.g., at step S440) based on the variance of the original time series data and the estimated quantization error. For instance, a so called high -rate theory of quantization may be used for the estimation of the updated measure of distortion. This theory states that the quantization error (e.g., squared error) becomes independent from the signal (e.g., time series data) itself if the bitrate is large enough (or the quantization step size is small enough). In this case, the variance of the reconstructed signal will be smaller than the variance of the uncoded signal with a delta corresponding to the variance of the quantization error. In transcoding, which applies a second encoding, the transcoder has access to the decoded time series data, the variance of the input signal (e.g., original time series data) is known from the distortion metadata, and the distortion introduced by the second stage of encoding can be estimated assuming that the quantization errors are independent from the signal.
[0166] Transcoding of Lossy Segments with Large Distortion
[0167] According to a third implementation of determining the updated measure of distortion, the (relevant portion of the) coded time series data may be lossy coded time series data with comparatively high coding distortion.
[0168] In this case, the implementation of method 400 may further comprise determining, based on the distortion metadata, that the measure of distortion is above a predetermined threshold. If so, determining the updated measure of distortion (e.g., at step S440) may comprise setting the measure of distortion to a value that indicates high distortion or that indicates that estimating the measure of distortion is not possible.
[0169] For instance, if subsequent coding is performed as tandem coding, which further increases the coding distortion, it may not be possible to meaningfully estimate the updated measure of distortion at this point. Nevertheless, it may be of interest to foresee a transcoder that detects this situation, and that then uses the updated distortion metadata to indicate that this situation has occurred (e.g., high distortion, no estimate available). Then such a situation may be detected also at the decoder by inspecting the decoded transcoded time series data alone, in particular without access to the original time series data.
[0170] Fig. 5 is a block diagram schematically illustrating a process flow 500 including generation and subsequent modification (e.g., transcoding) of a bitstream according to embodiments of the disclosure. An apparatus 505 (e.g., sensor, medical sensor, biomedical sensor) outputs an un-coded signal 510 (e.g., original time series data) to an encoder 520 (e.g., according to the encoder framework 300 shown in Fig. 3), which in turn generates a bitstream 525 of encoded signals (e.g., coded time series data) and distortion metadata. The bitstream 525 is transmitted via a transmission medium 530 and, for example, stored in a repository 540 for access at a later point in time. Once the decoded signals (e.g., decoded time series data) are retrieved from the repository 540, the associated distortion metadata may be used to compute the coding distortion and for deciding on suitability of the decoded signals for certain (automated) tasks and further processing steps.
[0171] For example, the coding distortion may be used for classifying the decoded signals or portions thereof, for example according to their coding behavior (e.g., lossless, near lossless, or lossy) or according to their suitability for certain (automated) processing tasks. In other words, such a classification process facilitates decision on whether a segment of coded signal retrieved from a database has low enough coding distortion according to a requirement of the processing task (e.g., training of an Al-based algorithm for analyzing such signals). Note that in such a scenario, there is no access to the uncoded version of the signal. Furthermore, the coding distortion may vary in time for portions of signal, for example, for segments of a signal. Further, the coding distortion or the classification derived therefrom may be used for applying a filtering operation to the decoded signals or portions thereof before applying further processing tasks thereto. That is, the decoded signals or portions thereof can be analyzed with regard to their suitability for certain (automated) processing tasks based on the coding distortion, such as serving as training data for training a neural network, etc. Unsuitable decoded signals or portions thereof may be excluded from these processing tasks via the filtering. Hence, a result of such a filtering of signals is a subset of segments of signals that are suitable for the processing task.
[0172] Further to the above, Fig. 5 also shows a possible follow up step, where the retrieved bitstream 545 (e.g., included in a record retrieved from the repository 540) is transcoded, for example to reduce storage requirements for the database by means of lossy coding, and the distortion metadata is updated accordingly. To this end, a decoded signal 555 (e.g., decoded time series data) is generated by a decoder 550. The decoded signal 555 is subjected to encoding by a second encoder 560 (e.g., according to the encoder framework 300 shown in Fig. 3) for generation of a transcoded signal (e.g., transcoded time series data). The encoder 560 also uses the distortion metadata to generate updated distortion metadata. The resulting bitstream including the transcoded signal and the updated distortion metadata may be stored in a second repository 570 (which may or may not be identical to repository 540) for retrieval at a later point in time.
[0173] In the above, section 580 of the process flow relates to retrieving signals or waveforms (e.g., decoded time series data) and associated distortion metadata, and section 590 of the process flow relates to the transcoding and update of the distortion metadata.
[0174] Fig. 6 is a flowchart schematically illustrating an example of a method 600 of decoding a bitstream according to embodiments of the disclosure. The bitstream includes coded time series data and distortion metadata. The coded time series data may comprise lossy coded time series data, as described above in the context of encoding. The distortion metadata in turn includes information indicative of a measure of distortion in relation to a portion of the coded time series data. Method 600 comprises steps S610 to S630 that may be performed, for example, at a decoder.
[0175] At step 610, decoded time series data is generated (e.g., decoded) from the bitstream.
[0176] At step S620, the information indicative of the measure of distortion is extracted (e.g., decoded) from the bitstream.
[0177] In some implementations, if the distortion metadata does not comprise an indication of a distortion level from among a predefined set of distortion levels (i.e., the distortion metadata does not comprise a quantized version of the measure of distortion), the measure of distortion may be quantized according to the predefined set of distortion levels at this point, if so required for further processing. For example, quantization according to the predefined set of distortion levels may involve comparing the measure of distortion that is extracted from the bitstream to a sequence of thresholds.
[0178] Further details of such quantization have been described above for in the context of encoding and / or transcoding.
[0179] At step S630, the decoded time series data and the information indicative of the measure of distortion are output.
[0180] The following framework for decoding follows Example 1 of content of distortion metadata, where the distortion is signaled implicitly. The measure of distortion may be given in terms of the quantities and information items described above in the context of encoding. Also, as described above, the measure of distortion may be given on a per-channel basis. For example, the measure of distortion may relate to an unnormalized squared error p? over the portion of coded time series data. In some implementations, the distortion metadata may also include information indicative of a variance of original time series data corresponding to the portion of the coded time series data, as described above.
[0181] Such distortion metadata, if made available to the decoder side, may facilitate estimation of performance parameters on the decoder side. For example, the percentage root mean square distortion (PRO) may be determined, which may be defined for example as
[0182] Further, the channel-normalized percentage root mean square distortion (CPRD) may be determined, which may be defined for example as
[0183] Thus, method 600 may for example further comprise determining the PRD or the CPRD based on the measure of distortion and the variance. The variance used for this purpose may be a variance of sample values of samples of the original time series data. The variance may be given on a per-channel basis, for example, as noted above.
[0184] Furthermore, MSE and PSNR-values may be computed for example as
[0185] MSEt= p
[0186] (7) where (B + 1) denotes the bit depth of each sample.
[0187] Accordingly, method 600 may for example further comprise determining the MSE for the portion of decoded time series data based on the measure of distortion, and / or determining the PSNR value for the portion of decoded time series data based on the measure of distortion. The PSNR value may be a per-channel value or an average value, for example.
[0188] The following framework for decoding follows Example 2 of content of distortion metadata, where the distortion is signaled explicitly. The measure of distortion is identified based on the metadata indicating the distortion metric used by the encoder. The availability of average or per- channel distortion values may be signaled by a single bit. Depending on this bit, a single distortion value may be read from the bitstream or individual values for each of the N-channels of the signal segment may be read from the bitstream.
[0189] Fig. 7 is a block diagram of an example of a decoder framework 700 according to embodiments of the disclosure.
[0190] A multiplexed bitstream 710 including coded time series data and associated distortion metadata is demultiplexed by a demultiplexer 720 to output a core coded bitstream 730 and distortion metadata 760. The core coded bitstream 730 is core decoded by a core decoder 740 to generate reconstructed signal 750 (e.g., decoded time series data). The distortion metadata 760 is provided to a distortion computation block 770 that computes the measure of distortion. For instance, a segment of core coded bitstream may be decoded, and the distortion metadata may be retrieved for that segment and used to compute the coding distortion for that segment.
[0191] Fig. 8 is a block diagram schematically illustrating a process flow 800 including generation and subsequent decoding of a bitstream according to embodiments of the disclosure. The first part of this process flow may be identical to the first part of process flow 500 illustrated in Fig. 5. An apparatus 805 (e.g., sensor, medical sensor, biomedical sensor) outputs an un-coded signal 810 (e.g., original time series data) to an encoder 820 (e.g., according to the encoder framework 300 shown in Fig. 3), which in turn generates a bitstream 825 of encoded signals (e.g., coded time series data) and distortion metadata. The bitstream 825 is transmitted via transmission medium 830 and, for example, stored in a repository 840 for access at a later point in time. Once the decoded signals (e.g., decoded time series data) are retrieved from the repository 840, the associated distortion metadata may be used to compute the coding distortion and for deciding on suitability of the decoded signals for certain (automated) tasks and further processing steps.
[0192] For example, the coding distortion may be used for classifying the decoded signals or portions thereof, for example according to their coding behavior (e.g., lossless, near lossless, or lossy) or according to their suitability for certain (automated) processing tasks. Further, the coding distortion or the classification derived therefrom may be used for applying a filtering operation to the decoded signals or portions thereof before applying further processing tasks thereto. That is, the decoded signals or portions thereof can be analyzed with regard to their suitability for certain (automated) processing tasks based on the coding distortion, such as serving as training data for training a neural network, etc. Unsuitable decoded signals or portions thereof may be excluded from these processing tasks via the filtering.
[0193] Further to the above, Fig. 8 also shows a possible follow up step, where the retrieved bitstream 845 (e.g., included in a record retrieved from the repository 840) is decoded by a decoder 850 to provide reconstructed signals 860 (e.g., decoded time series data) and the corresponding distortion metadata 855. A data filter 870 is applied to the reconstructed signals 860, where the data filter 870 uses the values of segment distortion (e.g., the measures of distortion for respective segments of decoded time series data) derived from the distortion metadata 855. The filtered signals may be used for a processing task or estimation task 880 (e.g., computation of parameters of interest) or may be used for training of a trainable algorithm. Filtering in this context may relate to either keeping or discarding a given input, or to assigning each input to one of a plurality of predefined processing bins.
[0194] In the above, section 890 of the process flow relates to retrieving signals or waveforms (e.g., decoded time series data) and associated distortion metadata. In line with the above, in some embodiments, the decoded time series data may comprise a plurality of portions of decoded time series data, and the distortion metadata may include a respective measure of distortion for each of the plurality of portions of decoded time series data. Having available the distortion metadata allows classifying the portions of decoded time series data according to their coding distortion, and by extension, according to their suitability for certain (automated) processing tasks. Accordingly, method 600 may further comprise a step of classifying the plurality of portions of the decoded time series data in accordance with their respective measures of distortion. The method may further comprise outputting a result of classification. This may be done for each portion of decoded time series data for which a classification result (or distortion metadata in general) is available.
[0195] The distortion metadata (e.g., measures of distortion) or the classification results may be used for filtering the plurality of portions of decoded time series data. For example, method 600 may comprise a step of applying a filtering operation to the plurality of portions of decoded time series data based on a result of classification. In some implementations, the filtering may be applied directly based on the associated measure of distortion. The filtering operation may relate, for each applicable portion of decoded time series data, to applying a “keep or toss” logic, or to assigning different portions to different buckets or bins for further processing, depending on the result of classification or measure of distortion.
[0196] For instance, method 600 may further comprise deciding on whether to use the portion of decoded time series data (or any portion of decoded time series data) for training of an Al-based model (as an example of an automated task) based on the measure of distortion. In particular, the portion of decoded time series data may only be used for training the Al-based model (e.g., neural network-based model) if the measure of distortion is below a predetermined threshold.
[0197] Fig. 9 provides an example of high-level view of a bitstream syntax of a bitstream 900 according to embodiments of the disclosure.
[0198] Bitstream 900 may comprise a sync header 910, a sync word 920, a configuration (cfg) header 930, cfg 940, metadata 970 relating to a segment 980, and the segment 980 of core coded frames including payload data 990. The metadata 970 may comprise a payload header 950 and a cyclic redundancy check (CRC) field 960. Without intended limitation, two specific example embodiments of the distortion metadata will be described next. Both example embodiments may relate to a segment of length L of an N- channel signal.
[0199] In the first example embodiment, the following metadata may be generated:
[0200] • a binary flag indicating lossy vs lossless coding if lossy {
[0201] • for each of the N channels: (normalized with respect to L) variance in the 101og10(-) domain as 8 bit int (e.g., quantized with 1 dB resolution)
[0202] • for each of N channels: (normalized with respect to L) squared error in the 101og10(-) domain as 8 bit int (e.g., quantized with 1 dB resolution)
[0203] • the length L for which the above values are computed.
[0204] In the above, the quantized values may come with a sign. Alternatively, a uint8 may be used as an index pointing to a table with signed integer values. If a single channel is lossless or near lossless so that the squared error is -Inf or smaller than the minimum value allowed by the quantization scheme, the lowest allowed value may be used instead.
[0205] In the above, 1 dB quantization has proven very precise for the actual PRO range of interest. It is understood that also coarser quantization may be considered in the context of the present disclosure. Additionally, a coarser quantization may be used for the squared error than for the variance.
[0206] In the second example embodiment, the same metadata as in the first example embodiment may be generated. In addition, the following metadata may be generated:
[0207] • a label indicating what type of distortion optimization was performed by the encoder, for example a 4 bit label to enumerate {MSE, PRO, PSNR, MAE, etc.}
[0208] In principle, the encoder could use any distortion metric of choice to optimize its bitrate vs distortion trade-off. It may be beneficial to label the bitstreams to indicate what type of optimization was used. Then, on the decoder side, this information could be derived solely by inspecting the bitstreams. Next, non-limiting examples of possible bitstream syntax for medical signals as an example of time series data will be described with reference to Table 1 through Table 28. This bitstream syntax may be used for implementing and processing methods of the present disclosure, for example when applied to utility bio-medical signals. The following defines a self-contained stream format to transport medical signals data. The transport mechanism uses a packetized approach. Both configuration data as well as coded payload data is embedded into separate packets.
[0209] Table 1 defines an example of the bitstream syntax of escaped Value().
[0210] Table 1
[0211] Table 2 defines an example of the bitstream syntax of msStreamPacketPayload().
[0212] Table 2
[0213] Table 3 defines an example of the bitstream syntax of msStreamPacket().
[0214] Table 3 Table 4 defines an example of the bitstream syntax of msStream(). Table 4
[0215] Semantics may be as follows: msStreamPacketLabel For values of ‘ T and higher, this element provides an indication of which packets in a stream belong together (so called sub-streams).
[0216] In addition, packets with msStreamPacketLabel set to a value of ‘0’ apply to all sub-streams. msStreamPacketLength This element indicates the length of the msStreamPacketPayload() in Bytes. msStreamPacketPayload() The payload for the actual msStreamPacket. It consists of 1 or more payload frames.
[0217] Table 5 defines example values for bitstream parameter msStreamPacketType.
[0218] Table 5
[0219] Next, an example configuration of the signal of type msSignalType will be described. The configuration shall apply to all packets with the same msStreamPacketLabel as the MS CFG payload.
[0220] Table 6 defines an example of bitstream syntax of msConfig().
[0221] Table 6
[0222] Table 7 defines an example of bitstream syntax of msSignalECGConfig().
[0223] Table 7 Table 8 defines an example of bitstream syntax of msSignalEEGConfig(). Table 8
[0224] Table 10 defines an example of bitstream syntax of msSignalPPGConfig().
[0225] Semantics may be as follows: msIsConfigExtensionPresent
[0226] Indicates the presence of an extended configuration setting. If it is equal to TRUE or 1, then the subsequent signal specific function (e.g., msSignalECGConfigExtension() in msSignalECGConfig()) shall be executed. msSamplingFrequency Coded data sampling frequency. msNchannels Number of coded data input channels. msF rameLength The number of input samples per coded payload frame. msCodecMode Coded data codec mode. msTotalSamples Total input samples corresponding to the whole bitstream. Note that his variable is only temporary. It will be removed in the future. msMeanPerChannel Coded data sample mean per channel. msGlobalGain Coded data global gain. msLPCOrder Coded data LPC order.
[0227] GetLossylndicatorQ Calculates an indicator whether the codec operates in lossless or lossy manner.
[0228] Table 11 describes example values for bitstream parameter msSignalType.
[0229] Table 11
[0230] Therein, msSignalECG indicates coded Electrocardiography (ECG) data, msSignalEEG indicates coded Electroencephalography (EEG) data, msSignalEMG indicates coded Electromyography (EMG) data, and msSignalPPG indicates coded Photoplethysmogram (PPG) data.
[0231] Further, Table 12 defines an example of bitstream syntax for a packet of type msFrame(), Table 13 defines an example of bitstream syntax for a packet of type msSignalECG (), Table 14 defines an example of bitstream syntax for a packet of type msSignalEEG(), Table 15 defines an example of bitstream syntax for a packet of type msSignalEMGQ , Table 16 defines an example of bitstream syntax for a packet of type msSignalPPG() , and Table 17 defines an example of bitstream syntax for a packet of type msCodedDataSideInfo() . Table 12
[0232] Table 13
[0233] Table 14
[0234] Tablel5
[0235] Table 16
[0236] Table 17
[0237] Alternatively, the distortion measure syntax may also be specified generally. For example, it may be specified as given in Table 18.
[0238] Table 18
[0239] Additionally, if a dedicated packet is preferred to identify a feature of the medical signals data, a new packet type may be introduced, e.g., MS FEATURE. This packet may be inserted just prior to the msFeatureSegmentStart frame if there is an actual feature segment associated with the feature. Otherwise, the syntax showing msHasFeature and msNumFeatures is sufficient to indicate the presence of features.
[0240] Table 19 defines an example of alternative bitstream syntax of msFrame().
[0241] Table 19 Table 20 defines an example of bitstream syntax of msFeature() .
[0242] Table 20
[0243] Semantics may be as follows: msSignalType See Table 11. msIndependentFrame Shall be set to 1 if the current MS FRAME payload is decodable without any additional information. getLength() Calculates the payload segment length in samples. msN umF ramesPerSegment
[0244] Indicates the number of payload frames carried by an MS FRAME payload (segment). msFrameSize Indicates the number or size of coded data bits in a payload frame. msDistortionMeasure Indicates the type of distortion measure applicable to an
[0245] MS FRAME segment, e.g., variance of the input signal on a perchannel basis over the segment of length L (msSegmentLength in samples) and estimated unnormalized squared error on a per-channel basis over the segment of length L (msSegmentLength in samples).
[0246] Table 21 describes example values for bitstream parameter msDistortionMeasure.
[0247] Table 21 ms Variance Indicates the signal variance measured per channel. msSquaredError Indicates the signal squared error measured per channel. msPerChannelMeasure Shall be set to ‘ 1’ if the distortion is specified per channel. Shall be set to ‘0’ if the distortion is specified for all channels. msDistortionValue The value associated with msDistortionMeasure . msHasFeature Indicates that this frame has data including a certain feature of interest. The flag may span various frames. msFeatureType Indicates a certain type of feature of the signal. msNumFeatures Indicates the number of features available in the bitstream associated with msSignalType and msFeatureType.
[0248] Table 22, Table 23, and Table 24 describe example values for bitstream parameter msFeatureType depending on the value of msSignalType.
[0249] Table 22
[0250] Table 23
[0251] Table 24 msFeatureSegment Indicates the presence of a feature segment associated with the feature msFeatur eType. msFeatureSegmentStart Indicates the frame index of the start of a feature segment. The index is refering to the frame indexing within an MS FRAME segment of msNumFramesPerSegment payload frames. msFeatureSegmentLen Indicates the length of a feature segment in number of payload frames starting from msFeatureSegmentStart.
[0252] The coded data payload may contain the data resulting from an encoding process on the msFrameLength input samples payload frame. In general, the codec operates in 2 modes, i.e., lossy and lossless mode. Table 25 defines an example of bitstream syntax of msCodedDataPayloadLossy().
[0253] Table 25 bUseLPC Indicates the presence of a feature segment associated with the feature msFeatur eType. bUseMCP Indicates the presence of a feature segment associated with the feature msFeatur eType. Table 26 defines an example of bitstream syntax of msCodedDataPayloadLossless().
[0254] Table 26
[0255] EnableDCT Indicates if coding is performed in the DCT domain or the time domain.
[0256] ICPredEnable Indicates the DCT domain frequency and channel prediction is enabled.
[0257] ICPredSplit Indicates that the DCT domain frequency and channel prediction is separately controlled for the first half of the spectrum and the second half of the spectrum. If the spectrum is split, two additional control bits is sent for each half of the spectrum. If ICPredSplit is 0 then the prediction runs across the entire spectrum. ICPredEnableHigh Indicates the DCT domain prediction is enabled for the high frequency half of the spectrum.
[0258] ICPredEnableLow Indicates the DCT domain prediction is enabled for the low frequency half of the spectrum.
[0259] ICTimePredEnable Indicates if inter-channel prediction is enabled in the time domain. If enabled the reference channel and prediction gain, follow in the bitstream.
[0260] ICRefChannel Indicates the reference channel used to predict the current channel
[0261] (n). The reference channel is less than current channel (n) and carried as a ceil(log2(n)) bit number.
[0262] ICPredGain Indicates the predication gain used in the time domain inter-channel prediction.
[0263] LPCRegionlnc Indicates the number of LPC regions (NumLPCRegions) is incremented by 1.
[0264] LPCRegionStart Indicates the start of the LPC region. The frame can be segmented into 64 possible start locations for each LPC region.
[0265] LPCOrder Indicates the order of the LPC prediction for each channel and LPC region.
[0266] KCoeff The LPC coefficients used in the time domain prediction.
[0267] Num regions Indicates the number of sub-regions of the signal that is transmitted.
[0268] The number of sub-regions transmitted ranges from 0 to max num regions, where max num regions is dependent on the frame length as shown Table 27 below. For lossless operation NumRegions is not transmitted and set to max num regions.
[0269] Table 27
[0270] RegionCb Indicates the index of the codebook used to code the sub-region of the signal.
[0271] DeltaCb Indicates the codebook index change relative to the codebook index in previous sub-region of the signal. unsigned cb Indicates if the codebook codes the sign information explicitly or if the sign is tranmitted seperately. The unsigned cb values are provided in Table 28.
[0272] Table 28
[0273] HGRCode The Huffman or Golomb-Rice codebook used to code the sub-region of the signal.
[0274] QuadSignBits The sign bits for 4 signal values. Sign bits are only transmitted if the absolute signal values are greater than 0. PairSignBits The sign bits for 2 signal values. Sign bits are only transmitted if the absolute signal values are greater than 0.
[0275] SignBit The sign bits for 1 signal values. Sign bits are only transmitted if the absolute signal values are greater than 0. Next, another non-limiting example of possible bitstream syntax for implementing aspects and embodiments of the present disclosure will be described with reference to Table 29.
[0276] Table 29 defines an example of bitstream syntax of a raw byte sequence payload (RBSP) for segment metadata that may be used for signaling distortion metadata according to embodiments of the disclosure.
[0277] Table 29
[0278] Segment metadata shall be used to obtain information on the coded data payload frame sizes (IF SPT, DF SPT) and distortion measure per bitstream segment.
[0279] • By default, a segment refers to the whole coded bitstream • A segment is also defined as a single or multiple of random access intervals. A random access interval consists of a single or multiple of frame sequences, i.e., a sequence of one IF SPT followed by zero or more DF SPT frames. Note a bitstream can be configured comprising only IF SPT frames. In this case, a segment may refer to any partition resulting from the whole waveform partitioning, or it may refer to any arbitrary waveform excerpt based on a given user input or a particular waveform feature.
[0280] The segment metadata packet is inserted at the end of each segment.
[0281] The distortion measure information shall be used for, but not limited to, the following cases:
[0282] • Indicating or classifying the segment’s coding behavior, i.e., lossless, near lossless (small distortion) or lossy.
[0283] • Deriving other distortion measure metrics, e.g., percentage root mean square distortion (PRD), channel-normalized percentage root mean square distortion (CPRD).
[0284] • Enabling assessment of distortion at the decoder side without having access to the original signals. This use case is related to retrieval of lossy coded segments from a dataset by classifying and filtering of the segments in the context of further postprocessing such as Al-based training.
[0285] • Transcoding (encoding of decoded bistream) from lossless to lossy or tandem coding. The resulting distortion shall be updated accordingly. sm channel group parameter set id specifies the value of cgps_channel_groupjparameter_set_id for the CGPS in use. sm channel group id identifies the channel group to which the current segment metadata belongs. When sm_channel_group_id is not present, it is inferred to be equal to 0. sm has feature flag specifies the presence of a feature within a segment. sm_num_features_minus_one plus 1, specifies the number features within a segment. sm feature type specifies the feature type. sm feature segment marking flag indicates the presence of feature marking. sm feature segment start indicates the offset start of a feature marking in samples. sm feature segment length indicates the length of a feature marking in samples. sm segment stat flag indicates whether an information on the frame sizes within a segment is present or not. sm num frames per segment indicates the number of payload frames carried in a segment. sm_frame_size[0] indicates the size / length of the first (n=0) coded data frame (IF SPT) in the segment. Subsequent frame sizes (n>0) are encoded using the Golomb / Rice delta encoding method. sm delta GRjparam specifies the parameter that controls the Golomb / Rice delta encoding of subsequent sm frame sizefn]. It uses the fixed-length (FL) binarization process having cMax = 15, see section Fixed-length binarization process. sm abs delta for a given iteration n, specifies the absolute value of the difference between the n-th and (n-l)-th sm frame size values (encoded absolute delta value). It uses the un-truncated Rice (UTR) binarization process having cRiceParam = sm delta GRjparam, see section Untruncated Rice binarization process. The function, decode(sm_abs_delta, sm delta GRjparam), retrieves the symbolVal which is ths absolute delta value, abs delta. sm sign delta specifies the sign of sm abs delta. It uses the fixed-length (FL) binarization process having cMax = 1, see section Fixed-length binarization process. sm payload extension flag shall be set to 1 to indicate the presence of a segment payload extension in the form of concatenated independent_frame_rbsp( ) and zero or more dependent_frame_rbsp( ) payload sequence(s).
[0286] This utilizes the Golomb / Rice delta encoding syntax elements, i.e., sm num framesjper segment, sm frame size, sm delta GRjparam, sm abs delta and sm sign delta. In a specific transcoding case, where existing frame sequences (bitstream) are already available for processing, the Golomb / Rice syntax elements are determined by extracting the frame sizes / lengths from the stream jpacket header of respective payload packets.
[0287] Given the transcoded bitstream, retreiving the original frame sequences is done by assigning a proper streamjpacket header (type, label, length) to individual payload rbsp, utilizing syntax elements within sm_segment_stat_flag and smjpayload extension flag. In all cases, the bitstream decoding to output the waveform shall be supported. sm num frames per sequence specifies the number of frames in a frame sequence. sm_distortion_measures_per_channel flag indicates whether the per-channel distortion measure metadata is present or not in the bitstream. This shall indicate the coded data coding mode, a lossless codec shall set this to ‘O’. sm_num_distortion_measures_per_channel indicates the number of other per-channel distortion measures calculated for the segment. sm variance indicates the signal variance measured per channel in dB unit, 10 * log 10 (variance). sm squared error indicates the signal squared error measured per channel in dB unit, 10 * loglO (squared error). sm distortion measure type indicates other types of per-channel distortion measure calculated in a segment, e.g., maximum absolute error (MAE), maximum amplitude error (MAX), including it’s unit. sm distortion measiire indicates the per-channel distortion measure value. sm_distortion_measures_in_cg flag shall be set to zero if the measure of distortion per-channel group is not desired. This measure of distortion is active by default. sm_num_distortion_measures_in_cg indicates the number of other per-channel group distortion measures calculated for the segment. sm variance cg indicates the signal variance measured per-channel group in dB unit,
[0288] 10 * log 10 (channel group variance). sm squared error cg indicates the signal squared error measured per-channel group in dB unit, 10 * log 10 (channel group squared error). sm distortion measiire type cg indicates other types of per-channel group distortion measure calculated in a segment, e.g., maximum absolute error (MAE), maximum amplitude error (MAX), including its unit. sm distortion measiire cg indicates the per-channel group distortion measure value.
[0289] The above syntax may be used to inform the coded signal quality, which is targeting at different use cases, e.g., transcoding, archiving for different purposes (such as medical evaluation, research). This could be done either through a post-processing classification, or by directly having a compact representation of such post-processing classification in the syntax, e.g., sm_distortion_measure_type[ch][i] = “research - SNR”, sm_distortion_measure[ch][i] = “good” or “acceptable” (preferably represented by an index pointing to a look-up table, e.g., 0..5 indicating bad to excellent). This would correspond to quantizing the measure of distortion according to a predefined set of distortion levels, as described above. Notably, the distortion measure type may include a specific distortion type (e.g., SNR indicating Signal to Noise Ratio or other signal / coded signal quality indicators) and supporting attributes such as the measurement unit (e.g., dB), the look-up table and other general annotations.
[0290] An example of post-processing classification using a “threshold” is by setting the quality to “good” if the “squared error” is below x dB.
[0291] Untruncated Rice binarization process
[0292] Input to this process is a request for a Untruncated Rice (UTR) binarization and cRiceParam > 0. Output of this process is the UTR binarization associating each value symbolVal with a corresponding bin string.
[0293] A UTR bin string is a concatenation of a prefix bin string and, when present, a suffix bin string.
[0294] For the derivation of the prefix bin string, the following applies:
[0295] - The prefix value of symbolVal, prefixVal, is derived as follows: prefixVal = symbolVal » cRiceParam (10)
[0296] - The prefix of the TR bin string is specified as follows:
[0297] - The prefix bin string is a bit string of length prefixVal + 1 indexed by binldx. The bins for binldx less than prefixVal are equal to 1. The bin with binldx equal to prefixVal is equal to 0. Table 30 illustrates the bin strings of this unary binarization for prefixVal.
[0298] The suffix of the TR bin string is present and it is derived as follows:
[0299] - The suffix value suffixVal is derived as follows: suffixVal = symbolVal - ( prefixVal « cRiceParam ) (11)
[0300] - The suffix of the UTR bin string is specified by invoking the fixed-length (FL) binarization process as specified in Fixed-length binarization process for suffixVal with a cMax value equal to ( 1 « cRiceParam ) - 1.
[0301] Fixed-length binarization process
[0302] Inputs to this process are a request for a fixed-length (FL) binarization and cMax.
[0303] Output of this process is the FL binarization associating each value symbolVal with a corresponding bin string. FL binarization is constructed by using the fixedLength-bit unsigned integer bin string of the symbol value symbol Vai, where fixedLength = Ceil( Log2( cMax + 1 ) ). The indexing of bins for the FL binarization is such that the binldx = 0 relates to the most significant bit with increasing values of binldx towards the least significant bit. Table 30 (informative) gives the bin string of the unary quantization.
[0304] Table 30
[0305] While methods have been described above, it is understood that the present disclosure likewise relates to apparatus (e.g., computer apparatus or apparatus having processing capability in general) for implementing these methods and neural networks (or techniques in general). An example of such apparatus 1000 is schematically illustrated in Fig. 10. The apparatus 1000 comprises a processor 1010 (or multiple processors) and a memory 1020 coupled to the processor 1010. The memory 1020 may store instructions for execution by the processor 1010.
[0306] Processor 1010 may be adapted to implement the apparatus described throughout the disclosure and / or to perform methods (e.g., methods of encoding or decoding, methods of generating or modifying bitstreams, methods of transcoding) described throughout the disclosure. The apparatus 1000 may receive inputs 1030 (e.g., utility signals, waveform data, time series data, bitstreams including coded time series data and / or distortion metadata, etc.) and generate outputs 1040 (e.g., bitstreams including coded metadata and distortion metadata, decoded time series data, information on a measure of distortion, classification results, etc.) as described throughout the disclosure. Accordingly, the apparatus 1000 may relate to any of an encoding apparatus, a transcoding apparatus, and a decoding apparatus, as the case may be.
[0307] Fig. 11 illustrates a transmission device 1100. Transmission device 1100 may comprise a variety of units, including a transmitter unit and / or a receiver unit and / or a coding unit. The coding unit may be composed of at least a processor configured to perform encoding and / or decoding processing. The coding unit may encode, transcode, and / or decode data in accordance with the methods (e.g., methods 100, 400, and 600 as illustrated above with reference to Fig. 1, Fig. 4, and Fig. 6, respectively) throughout the present disclosure.
[0308] Transmission device 1100 may transmit the coded data in the form of a bitstream to a device or to a digital storage medium or through, for example, a network in the form of a file or streaming. The digital storage medium may include various storage mediums such as USB-C, USB, SD, CD, DVD, Blu-ray, HDD, SSD, and equivalent technologies. The digital storage medium may also be part of the coding unit of transmitter device 1100.
[0309] Transmission device 1100 may include an element for generating the bitstream and / or a media file and may include an element for transmission, e.g., through a variety of mediums (Bluetooth, broadcast / communication networks, Internet technologies and equivalents). The transmission may be implemented using a variety of technologies such as, for example, RF, light waves, infrared, Bluetooth, WiFi, and / or acoustic transmission devices.
[0310] The present disclosure further relates to programs (e.g., computer programs) comprising instructions that, when executed by a processor (or multiple processors), cause the processor (or multiple processors) to carry out any of the methods described throughout the disclosure, and to computer-readable storage media storing such programs.
[0311] Aspects of the systems described herein may be implemented in an appropriate computer-based processing network environment (e.g., server or cloud environment) for processing digital or digitized files. Portions of these systems may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof
[0312] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.
[0313] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, computer-implemented neural networks described herein can include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.
[0314] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements. Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings.
[0315] Enumerated Example Embodiments
[0316] Various Aspects and implementations of the invention may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.
[0317] EEE1. A method of generating or modifying a bitstream, comprising: obtaining the bitstream, wherein the bitstream contains coded time series data; determining a measure of distortion in relation to a portion of the coded time series data; and embedding distortion metadata into the bitstream, wherein the distortion metadata includes information indicative of the determined measure of distortion.
[0318] EEE2. The method according to EEE1, wherein the coded time series data comprises lossy coded time series data.
[0319] EEE3. The method according to EEE1 or EEE2, wherein the coded time series data relates to original time series data, the coded time series data being obtainable from the original time series data via a coding operation; wherein the method further comprises local decoding of the coded time series data to obtain locally decoded time series data; and wherein determining the measure of distortion is based on one or more samples of the original time series data and one or more samples of the locally decoded time series data.
[0320] EEE4. The method according to EEE1 or EEE2, further comprising: generating the coded time series data from original time series data via a coding operation; and estimating the measure of distortion in relation to the portion of the coded time series data based on a rate allocation process of the coding operation. EEE5. The method according to EEE4, wherein the coded time series data is generated from the original time series data via the coding operation subject to a bitrate constraint, and the coding operation is adapted to minimize a signal distortion given the bitrate constraint; and wherein the measure of distortion is estimated in the course of minimizing the signal distortion given the bitrate constraint.
[0321] EEE6. The method according to EEE4, wherein the coded time series data is generated from the original time series data via the coding operation subject to an average bitrate constraint, and the coding operation is adapted to optimize a rate-distortion trade-off given the average bitrate constraint, by selection of appropriate operation points of the coding operation; and wherein the measure of distortion is estimated in the course of selecting the appropriate operation points.
[0322] EEE7. The method according to any one of the preceding EEEs, wherein the coding operation for obtaining the coded time series data relates to distortion-constrained coding using a distortion constraint; and wherein the measure of distortion is determined based on the distortion constraint.
[0323] EEE8. The method according to any one of the preceding EEEs, wherein the coded time series data comprises one or more channels of coded time series data; and wherein the measure of distortion relates to at least one of a maximum absolute error per channel and a per-channel squared error.
[0324] EEE9. The method according to EEE8, further comprising: quantizing the maximum absolute error per channel and / or the per-channel squared error to a number format of a predetermined number of bits, using a predefined quantization resolution, wherein the measure of distortion relates to the quantized maximum absolute error per channel and / or the quantized per-channel squared error.
[0325] EEE10. The method according to any one of EEE1 to EEE7, wherein the coded time series data includes a plurality of channel groups, each channel group including a plurality of channels of coded time series data; and wherein the measure of distortion is embedded into the bitstream on a per channel group basis, and wherein each such measure of distortion relates to a cumulated measure of distortion for the respective channel group.
[0326] EEE11. The method according to any one of EEE 1 to EEE7, wherein the coded time series data includes a plurality of channel groups, each channel group including a plurality of channels of coded time series data; and wherein the measure of distortion is embedded into the bitstream on a per channel group basis, and wherein each such measure of distortion relates to individual measures of distortion for the channels in the respective channel group.
[0327] EEE 12. The method according to any one of the preceding EEEs, further comprising: quantizing the measure of distortion according to a predefined set of distortion levels, wherein the distortion metadata includes information indicative of the distortion level to which the measure of distortion has been quantized, and optionally, an indication that the measure of distortion has been quantized according to the predefined set of distortion levels.
[0328] EEE13. The method according to any one of the preceding EEEs, wherein the distortion metadata further includes information indicative of a length of the portion of the coded time series data to which the measure of distortion applies.
[0329] EEE14. The method according to any one of the preceding EEEs, wherein the distortion metadata further includes information indicative of a variance of the original time series data relating to the portion of the coded time series data.
[0330] EEE15. The method according to any one of the preceding EEEs, wherein the portion of the coded time series data relates to a segment including multiple frames of coded time series data.
[0331] EEE16. The method according to any one of the preceding EEEs, wherein the portion of the coded time series data relates to the entire bitstream.
[0332] EEE17. The method according to any one of the preceding EEEs, wherein the embedding comprises embedding a data packet into the bitstream that includes an indication of the measure of distortion and an indication of the corresponding portion of the coded time series data.
[0333] EEE18. The method according to EEE15, wherein the embedding comprises embedding a data packet into the bitstream that includes an indication of the measure of distortion; and wherein the data packet is embedded at the end of the segment.
[0334] EEE19. The method according to EEE16, wherein the embedding comprises embedding a data packet into the bitstream that includes an indication of the measure of distortion; and wherein the data packet is embedded at the end of the bitstream.
[0335] EEE20. The method according to any one of the preceding EEEs, wherein the time series data relates to one or more utility signals.
[0336] EEE21. The method according to any one of the preceding EEEs, wherein the time series data relates to one or more medical or biomedical signals.
[0337] EEE22. The method according to any one of the preceding EEEs, wherein the time series data relates to waveform data.
[0338] EEE23. The method according to any one of the preceding EEEs, further comprising: receiving an output of a sensor and applying a coding operation thereto to generate the coded time series data.
[0339] EEE24. A method of modifying a bitstream, comprising: obtaining the bitstream as an input, wherein the bitstream includes coded time series data and distortion metadata, and wherein the distortion metadata includes information indicative of a measure of distortion in relation to a portion of the coded time series data; extracting the distortion metadata from the bitstream; applying transcoding based on the coded time series data to obtain transcoded time series data; determining an updated measure of distortion in relation to a portion of the transcoded time series data based on the measure of distortion; and generating a transcoded bitstream including the transcoded time series data and updated distortion metadata, wherein the updated distortion metadata includes the updated measure of distortion.
[0340] EEE25. The method according to EEE24, wherein the coded time series data is lossless coded time series data; wherein the transcoding comprises decoding the coded time series data to obtain decoded time series data; and wherein determining the updated measure of distortion is based on the decoded time series data. EEE26. The method according to EEE24, wherein the coded time series data is lossy coded time series data; wherein the distortion metadata further includes information indicative of a variance of original time series data relating to the portion of the coded time series data; wherein the method further comprises: determining, based on the distortion metadata, that the measure of distortion is below a predetermined threshold; and estimating a quantization error that results from the transcoding; and wherein determining the updated measure of distortion comprises estimating the updated measure of distortion based on the variance of the original time series data and the estimated quantization error.
[0341] EEE27. The method according to EEE24, wherein the coded time series data is lossy coded time series data; wherein the method further comprises determining, based on the distortion metadata, that the measure of distortion is above a predetermined threshold; and wherein determining the updated measure of distortion comprises setting the measure of distortion to a value that indicates high distortion or that indicates that estimating the measure of distortion is not possible.
[0342] EEE28. A method of decoding a bitstream, wherein the bitstream includes coded time series data and distortion metadata, and wherein the distortion metadata includes information indicative of a measure of distortion in relation to a portion of the coded time series data, the method comprising: decoding, from the bitstream, decoded time series data; decoding, from the bitstream, the information indicative of the measure of distortion; and outputting the decoded time series data and the information indicative of the measure of distortion.
[0343] EEE29. The method according to EEE28, wherein the coded time series data comprises lossy coded time series data. EEE30. The method according to EEE28 or EEE29, wherein the distortion metadata further includes information indicative of a variance of original time series data corresponding to the portion of the coded time series data; and wherein the method further comprises determining a percentage root mean square distortion, PRD, or a channel -normalized percentage root mean square distortion, CPRD, based on the measure of distortion and the variance.
[0344] EEE31. The method according to any one of EEE28 to EEE30, further comprising determining a mean square error, MSE, for the portion of decoded time series data based on the measure of distortion.
[0345] EEE32. The method according to any one of EEE28 to EEE31, further comprising determining a peak signal to noise ratio, PSNR, value for the portion of decoded time series data based on the measure of distortion.
[0346] EEE33. The method according to any one of EEE28 to EEE32, wherein the distortion metadata comprises explicit signaling for the measure of distortion.
[0347] EEE34. The method according to any one of EEE28 to EEE32, wherein the distortion metadata comprises implicit signaling for the measure of distortion via a parameterization of the measure of distortion.
[0348] EEE35. The method according to any one of EEE28 to EEE32, wherein the distortion metadata comprises an indication of a distortion level from among a predefined set of distortion levels and optionally, an indication that the measure of distortion has been quantized according to the predefined set of distortion levels.
[0349] EEE36. The method according to any one of EEE28 to EEE34, further comprising: quantizing the measure of distortion according to a predefined set of distortion levels.
[0350] EEE37. The method according to any one of EEE28 to EEE36, wherein the decoded time series data comprises a plurality of portions of decoded time series data, and wherein the distortion metadata includes a respective measure of distortion for each of the plurality of measures of distortion; and wherein the method further comprises classifying the plurality of portions of the decoded time series data in accordance with their respective measures of distortion. EEE38. The method according to EEE37, further comprising applying a filtering operation to the plurality of portions of decoded time series data based on a result of classification.
[0351] EEE39. The method according to any one of EEE28 to EEE38, further comprising deciding whether to use the portion of decoded time series data for training an Al-based model based on the measure of distortion.
[0352] EEE40. An encoding apparatus comprising one or more processors and a memory coupled thereto, wherein the one or more processors are configured to perform the method according to any one of EEE 1 to EEE23.
[0353] EEE41. A transcoding apparatus comprising one or more processors and a memory coupled thereto, wherein the one or more processors are configured to perform the method according to any one of EEE24 to EEE27.
[0354] EEE42. A decoding apparatus comprising one or more processors and a memory coupled thereto, wherein the one or more processors are configured to perform the method according to any one of EEE28 to EEE39.
[0355] EEE43. A computer program including instructions that when executed by one or more processors, cause the one or more processors to perform the method according to any one of EEE1 to EEE39.
[0356] EEE44. A computer-readable storage medium storing the computer program according to EEE43.
Claims
1. Claims1. A method of generating or modifying a bitstream, comprising: obtaining the bitstream, wherein the bitstream contains coded time series data; determining a measure of distortion in relation to a portion of the coded time series data; and embedding distortion metadata into the bitstream, wherein the distortion metadata includes information indicative of the determined measure of distortion.
2. The method according to claim 1, wherein the coded time series data comprises lossy coded time series data.
3. The method according to claim 1 or 2, wherein the coded time series data relates to original time series data, the coded time series data being obtainable from the original time series data via a coding operation; wherein the method further comprises local decoding of the coded time series data to obtain locally decoded time series data; and wherein determining the measure of distortion is based on one or more samples of the original time series data and one or more samples of the locally decoded time series data.
4. The method according to claim 1 or 2, further comprising: generating the coded time series data from original time series data via a coding operation; and estimating the measure of distortion in relation to the portion of the coded time series data based on a rate allocation process of the coding operation.
5. The method according to claim 4, wherein the coded time series data is generated from the original time series data via the coding operation subject to a bitrate constraint, and the coding operation is adapted to minimize a signal distortion given the bitrate constraint; and wherein the measure of distortion is estimated in the course of minimizing the signal distortion given the bitrate constraint.
6. The method according to claim 4, wherein the coded time series data is generated from the original time series data via the coding operation subject to an average bitrate constraint, and the coding operation is adapted to optimize a rate-distortion trade-off given the average bitrate constraint, by selection of appropriate operation points of the coding operation; and wherein the measure of distortion is estimated in the course of selecting the appropriate operation points.
7. The method according to any one of the preceding claims, wherein the coding operation for obtaining the coded time series data relates to distortion-constrained coding using a distortion constraint; and wherein the measure of distortion is determined based on the distortion constraint.
8. The method according to any one of the preceding claims, wherein the coded time series data comprises one or more channels of coded time series data; and wherein the measure of distortion relates to at least one of a maximum absolute error per channel and a per-channel squared error.
9. The method according to claim 8, further comprising: quantizing the maximum absolute error per channel and / or the per-channel squared error to a number format of a predetermined number of bits, using a predefined quantization resolution, wherein the measure of distortion relates to the quantized maximum absolute error per channel and / or the quantized per-channel squared error.
10. The method according to any one of claims 1 to 7, wherein the coded time series data includes a plurality of channel groups, each channel group including a plurality of channels of coded time series data; and wherein the measure of distortion is embedded into the bitstream on a per channel group basis, and wherein each such measure of distortion relates to a cumulated measure of distortion for the respective channel group.7511. The method according to any one of claims 1 to 7, wherein the coded time series data includes a plurality of channel groups, each channel group including a plurality of channels of coded time series data; and wherein the measure of distortion is embedded into the bitstream on a per channel group basis, and wherein each such measure of distortion relates to individual measures of distortion for the channels in the respective channel group.
12. The method according to any one of the preceding claims, further comprising: quantizing the measure of distortion according to a predefined set of distortion levels, wherein the distortion metadata includes information indicative of the distortion level to which the measure of distortion has been quantized, and optionally, an indication that the measure of distortion has been quantized according to the predefined set of distortion levels.
13. The method according to any one of the preceding claims, wherein the distortion metadata further includes information indicative of a length of the portion of the coded time series data to which the measure of distortion applies.
14. The method according to any one of the preceding claims, wherein the distortion metadata further includes information indicative of a variance of the original time series data relating to the portion of the coded time series data.
15. The method according to any one of the preceding claims, wherein the portion of the coded time series data relates to a segment including an integer multiple of random access intervals of coded time series data.
16. The method according to any one of the preceding claims, wherein the portion of the coded time series data relates to the entire bitstream.7617. The method according to any one of the preceding claims, wherein the embedding comprises embedding a data packet into the bitstream that includes an indication of the measure of distortion and an indication of the corresponding portion of the coded time series data.
18. The method according to claim 15, wherein the embedding comprises embedding a data packet into the bitstream that includes an indication of the measure of distortion; and wherein the data packet is embedded at the end of the segment.
19. The method according to claim 16, wherein the embedding comprises embedding a data packet into the bitstream that includes an indication of the measure of distortion; and wherein the data packet is embedded at the end of the bitstream.
20. The method according to any one of the preceding claims, wherein the time series data relates to one or more utility signals.
21. The method according to any one of the preceding claims, wherein the time series data relates to one or more medical or biomedical signals.
22. The method according to any one of the preceding claims, wherein the time series data relates to waveform data.
23. The method according to any one of the preceding claims, further comprising: receiving an output of a sensor and applying a coding operation thereto to generate the coded time series data.
24. A method of modifying a bitstream, comprising: obtaining the bitstream as an input, wherein the bitstream includes coded time series data and distortion metadata, and wherein the distortion metadata includes information indicative of a measure of distortion in relation to a portion of the coded time series data; extracting the distortion metadata from the bitstream;applying transcoding based on the coded time series data to obtain transcoded time series data; determining an updated measure of distortion in relation to a portion of the transcoded time series data based on the measure of distortion; and generating a transcoded bitstream including the transcoded time series data and updated distortion metadata, wherein the updated distortion metadata includes the updated measure of distortion.
25. The method according to claim 24, wherein the coded time series data is lossless coded time series data; wherein the transcoding comprises decoding the coded time series data to obtain decoded time series data; and wherein determining the updated measure of distortion is based on the decoded time series data.
26. The method according to claim 24, wherein the coded time series data is lossy coded time series data; wherein the distortion metadata further includes information indicative of a variance of original time series data relating to the portion of the coded time series data; wherein the method further comprises: determining, based on the distortion metadata, that the measure of distortion is below a predetermined threshold; and estimating a quantization error that results from the transcoding; and wherein determining the updated measure of distortion comprises estimating the updated measure of distortion based on the variance of the original time series data and the estimated quantization error.
27. The method according to claim 24, wherein the coded time series data is lossy coded time series data; wherein the method further comprises determining, based on the distortion metadata, that the measure of distortion is above a predetermined threshold; andwherein determining the updated measure of distortion comprises setting the measure of distortion to a value that indicates high distortion or that indicates that estimating the measure of distortion is not possible.
28. A method of decoding a bitstream, wherein the bitstream includes coded time series data and distortion metadata, and wherein the distortion metadata includes information indicative of a measure of distortion in relation to a portion of the coded time series data, the method comprising: decoding, from the bitstream, decoded time series data; decoding, from the bitstream, the information indicative of the measure of distortion; and outputting the decoded time series data and the information indicative of the measure of distortion.
29. The method according to claim 28, wherein the coded time series data comprises lossy coded time series data.
30. The method according to claim 28 or 29, wherein the distortion metadata further includes information indicative of a variance of original time series data corresponding to the portion of the coded time series data; and wherein the method further comprises determining a percentage root mean square distortion, PRO, or a channel-normalized percentage root mean square distortion, CPRD, based on the measure of distortion and the variance.
31. The method according to any one of claims 28 to 30, further comprising determining a mean square error, MSE, for the portion of decoded time series data based on the measure of distortion.
32. The method according to any one of claims 28 to 31, further comprising determining a peak signal to noise ratio, PSNR, value for the portion of decoded time series data based on the measure of distortion.
33. The method according to any one of claims 28 to 32, wherein the distortion metadata comprises explicit signaling for the measure of distortion.
34. The method according to any one of claims 28 to 32, wherein the distortion metadata comprises implicit signaling for the measure of distortion via a parameterization of the measure of distortion.
35. The method according to any one of claims 28 to 32, wherein the distortion metadata comprises an indication of a distortion level from among a predefined set of distortion levels and optionally, an indication that the measure of distortion has been quantized according to the predefined set of distortion levels.
36. The method according to any one of claims 28 to 34, further comprising: quantizing the measure of distortion according to a predefined set of distortion levels.
37. The method according to any one of claims 28 to 36, wherein the decoded time series data comprises a plurality of portions of decoded time series data, and wherein the distortion metadata includes a respective measure of distortion for each of the plurality of measures of distortion; and wherein the method further comprises classifying the plurality of portions of the decoded time series data in accordance with their respective measures of distortion.
38. The method according to claim 37, further comprising applying a filtering operation to the plurality of portions of decoded time series data based on a result of classification.
39. The method according to any one of claims 28 to 38, further comprising deciding whether to use the portion of decoded time series data for training an Al-based model based on the measure of distortion.
40. An encoding apparatus comprising one or more processors and a memory coupled thereto, wherein the one or more processors are configured to perform the method according to any one of claims 1 to 23.
41. A transcoding apparatus comprising one or more processors and a memory coupled thereto, wherein the one or more processors are configured to perform the method according to any one of claims 24 to 27.
42. A decoding apparatus comprising one or more processors and a memory coupled thereto, wherein the one or more processors are configured to perform the method according to any one of claims 28 to 39.
43. A computer program including instructions that when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 39.
44. A computer-readable storage medium storing the computer program according to claim 42.
Citation Information
Patent Citations
Producing and encoding rate-distortion information allowing optimal transcoding of compressed digital image
US20030185453A1
Distributed Architecture for Encoding and Delivering Video Content
US20130343450A1