Channel-group based coding of digital waveform data
By grouping channels into channel groups for encoding and decoding, the method addresses inefficiencies in coding techniques for time-varying signals with varying sampling rates and functions, improving compression and decoding efficiency.
Patent Information
- Application Number
- PCT/EP2025/080534
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-23
- Filing Date
- 2025-10-22
- Publication Date
- 2026-04-30
AI Technical Summary
Existing coding techniques for time-varying signals with multiple channels of differing sampling rates and functions result in inefficiencies due to redundancies and incompatible data structures, necessitating a coding scheme that balances efficiency and complexity.
Channels are grouped into channel groups for encoding and decoding, allowing frame-wise and packet-wise processing, with each packet associated with a specific channel group, enabling separate handling of channels with different sampling rates and functions.
This approach improves coding efficiency by enhancing sample correlation and reducing redundancy, facilitating effective compression and decoding of diverse channel data.
Smart Images

Figure EP2025080534_30042026_PF_FP_ABST
Abstract
Description
[0001] Channel-Group based coding of Digital Waveform Data
[0002] Description
[0003] Embodiments according to the invention relate to coding of digital waveform data comprising one or more channels. For example, embodiments comprise decoders, encoders, methods for decoding, methods for encoding, computer programs and data streams using a high-level syntax for coding of digital waveform data.
[0004] Background
[0005] Time-varying signals are commonly used for representation of media and measurement data such as audio signals, biomedical signals, or seismic measurements. With the increase of signal generation, transportation and storage, there is demand for compression of such time-varying signals.
[0006] A digital waveform data may comprise multiple channels with different characteristics such as different sampling rates and / or channel function._Such different characteristics may result in redundancies or inefficiently structured data during or after encoding. For example different sampling rates may impede coding techniques that use temporal similarities and / or alignment. Similarly, different channel functions may comprise data format or samples values, which are incompatible or show little correlation.
[0007] Therefore, there is a need for a coding scheme that improves a compromise between coding efficiency and coding complexity.
[0008] This is achieved by the subject matter of the independent claims of the present application.
[0009] Further embodiments according to the invention are defined by the subject matter of the dependent claims of the present application
[0010] Summary of the invention
[0011] In accordance with an aspect of the invention, an encoder for encoding into a data stream digital waveform data comprising a one or more channels is provided. The encoder is configured to group the one or more channels into one or more channel groups, and encode the one or more channels into payload packets channel-group wise and frame wise so that each packet is associated with a channel group and has exclusively encoded a frame of one or more channels thereinto which are within the channel group with which the respective packet is associated.
[0012] In accordance with an aspect of the invention, a decoder for decoding from a data stream digital waveform data is provided, wherein channels of the digital waveform data are coded into the data stream in channel groups, and the decoder is configured to locate within the data stream payload packets into which the one or more channels are encoded channel-group wise and frame wise so that each packet is associated with a channel group and has exclusively encoded a frame of one or more channels thereinto which are within the channel group with which the respective packet is associated, and decode a collection of selected channel groups from payload packets associated with a channel group which is contained by the one or more selected channel groups of the collection.
[0013] See, the independent claims enabling a channel group wise and frame wise coding, thereby rendering it possible to separate channels of differing sampling rate. As described later, each payload packet, relating to a certain frame of a certain channel group, may be coded / decoded individually, or, in case of the usage of the example with independent and dependent frame / payload packets, within each channel group, each sequence of immediately successive one or more payload packets consisting of a leading independent payload / frame packet followed by zero, one or more dependent payload / frame packet.
[0014] Similarly, channels may be grouped by channel function, channel type, or sensor type. For example, by separating payload channels from annotation channels, a channel group containing the payload channels can apply corresponding coding schemes accordingly to either of the two channel groups. For example, data of text containing channels may not form a suitable or even compatible data structure for coding techniques (e.g., inter-frame coding) used for sample containing channels. Similarly, channels may be grouped according to sensor source. Temperature sensors may exhibit different patterns (e.g., having a steady signal) and / or sampling rates compared to an Electroencephalography (EEG) sensor, which tends to generate data with a wave pattern. Grouping such different channels can therefore improve sample correlation, which is beneficial for coding schemes that reduce redundancy (e.g., predictive coding). Such grouping also allows signaling information for channels in such groups collectively. The encoder may, for example, be able to signal information that can be applied to a group or a collection of groups.
[0015] Brief
[0016]
[0017] of the
[0018]
[0019] The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings, in which:
[0020] Fig. 1 shows a schematic view of an encoder for encoding into a data stream digital waveform data comprising one or more channels;
[0021] Fig. 2 shows a flow diagram of a method for encoding into a data stream digital waveform data comprising a one or more channels;
[0022] Fig. 3 shows a schematic view of a decoder for decoding from a data stream digital waveform data, wherein channels of the digital waveform data are coded into the data stream in channel groups;
[0023] Fig. 4 shows a flow diagram of a method for decoding from a data stream digital waveform data;
[0024] Fig. 5a shows a schematic view of a data stream with a packet;
[0025] Fig. 5b shows a schematic view of a packet within the data stream, wherein the packet comprises data portion for encapsulation or concatenation;
[0026] Fig. 5c shows a schematic view of an exemplary waveform parameter set, WPS, packet with general coding information; and
[0027] Fig. 6 shows an encoder for encoding a multi-channel digital signal into a datastream as well as decoder for decoding the multi-channel digital signal from datastream.
[0028] Detailed Description of the Embodiments Equal or equivalent elements or elements with equal or equivalent functionality are denoted in the following description by equal or equivalent reference numerals even if occurring in different figures.
[0029] In the following description, a plurality of details is set forth to provide a more throughout explanation of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail in order to avoid obscuring embodiments of the present invention. In addition, features of the different embodiments described herein after may be combined with each other, unless specifically noted otherwise.
[0030] remarks:
[0031] In the following, different inventive embodiments and aspects will be described in “High-level syntax for coding of digital waveform data”, in particular in the sections “Introduction to embodiments”, “Summary”, in particular in respective subsections “Auxiliary metadata concept”, “Waveform parameter set concept”, “Independent frame concept” (and respective sub-subsection), in sections “Dependent frame concept”, “Frame data concept”, “Annotation channel concept”, “Trailing bits concept”, “Channel groups” and “Examples”.
[0032] Also, further embodiments will be defined by the enclosed claims.
[0033] It should be noted that any embodiments as defined by the claims can be supplemented by any of the details (features and functionalities) described in the above-mentioned sections.
[0034] Also, the embodiments described in the above-mentioned sections can be used individually, and can also be supplemented by any of the features in another section, or by any feature included in the claims.
[0035] Also, it should be noted that individual aspects described herein can be used individually or in combination. Thus, details can be added to each of said individual aspects without adding details to another one of said aspects. Moreover, features and functionalities disclosed herein relating to a method can also be used in an apparatus (configured to perform such functionality). Furthermore, any features and functionalities disclosed herein with respect to an apparatus can also be used in a corresponding method. In other words, the methods disclosed herein can be supplemented by any of the features and functionalities described with respect to the apparatuses.
[0036] Moreover, features and functionalities disclosed herein relating to a method, in particular an encoding method can also be used in a data stream or bitstream (e.g. defining a respective data stream or bitstream element). Furthermore, any features and functionalities disclosed herein with respect to an apparatus can also be used in a corresponding data stream, e.g. as a resulting data stream as providing by said encoder. In other words, the data streams disclosed herein can be supplemented by any of the features and functionalities described with respect to the apparatuses and methods.
[0037] Also, any of the features and functionalities described herein can be implemented in hardware or in software, or using a combination of hardware and software, as will be described in the section “Implementation alternatives”.
[0038] Fig. 1 shows a schematic view of an encoder 10 for encoding into a data stream 16 digital waveform data 14 comprising one or more channels 21a-g. The encoder 10 is configured to group the one or more channels 21a-g into one or more channel groups 23a-c, encode the one or more channels 21a-g into payload packets 25a-c channel-group 23a-c wise and frame 27a-c wise so that each packet 25a-c is associated with a channel group 23a-b and has exclusively encoded a frame 27a-c of one or more channels 21a-g thereinto which are within the channel group 23a-c with which the respective packet 25a-c is associated.
[0039] The encoder 10 shown in fig. 1 may be any encoder enclosed herein and / or comprise any functionality described herein (e.g., in context of an encoder and / or decoder).
[0040] Fig. 2 shows a flow diagram of a method 100 for encoding into a data stream 16 digital waveform data 14 comprising a one or more channels 21a-g. The method 100 may be performed by any encoder 10 disclosed herein. Furthermore, any feature disclosed herein related to an encoder (or a decoder) may be applicable to the method and vice versa.
[0041] The method 100 comprises, in step 102, grouping the one or more channels 12a-g into one or more channel groups 23a-c. The method further comprises, in step 104, encoding the one or more channels 21a-g into payload packets 25a-c channel-group wise and frame 27 wise so that each packet 25a-c is associated with a channel group 23a-c and has exclusively encoded a frame 27 of one or more channels 21a-g thereinto which are within the channel group 23a-c with which the respective packet 25a-c is associated.
[0042] The method 100 may comprise transmitting the data stream 16 (e.g., to a decoder 12) and / or storing the data stream 16 (e.g., on a digital storage medium).
[0043] The grouping may be performed according to one or more of sampling rate, sensor type, sensed object (e.g., body part measured for generating the signal of a channel), overall amplitude, and channel name. The grouping may be performed according to channel function (or channel type), e.g., into any combination comprising one or more of payload channels, auxiliary channels, annotation channels, and label channels.
[0044] In the following, the encoder 10 will be described in further detail. However, any feature may be correspondingly applied to the method 100.
[0045] The encoder 10 may be part of or may comprise a computer, smartphone, tablet, smartwatch, server, or any other form of cloud computing resource. The encoder 10 may comprise a processor configured to perform the steps or functions disclosed herein. The encoder 10 may be part of or comprise a medical device (e.g., electroencephalograph and / or electrocardiograph), a seismic device (e.g., seismograph or seismometer), or an audio processing device. The digital waveform data 14 may represent (e.g., may be or may comprise) a biometric signal (e.g., of heart, brain, or eye functions, e.g., body temperature), seismic data (e.g., amplitude of ground motion), or a sound signal (e.g., sound amplitude, e.g., of one or more audio channels). The digital waveform data 14 may be a time varying signal (e.g., in units of ms, e.g., in units of samples).
[0046] The digital waveform data 14 may comprise one channel 21 or more (e.g., two, three, four, five, six, or more) channels 21. In the example shown in fig. 1 , the digital waveform data 14 has seven channels 21a-g. However, any other number of channels 21 may be used. A channel 21 may define a single parameter assuming values over time, e.g., wherein the parameter is sampled over temporally successive samples 19. The channels 21 may identifiable or indexable by an identifier, e.g., seven unique identifiers for seven channels 21a- g. One or more of the channels 21a-g may have different sampling rates (e.g., number of samples 19 per second). Similarly, different channels 21 may have the same sampling rate. For example, in fig. 1, channel 21a may have the same sampling rate as channel 21b or 21 d, but a different sampling rate than channel 21 f (indicated by a different width of the samples 19).
[0047] A plurality of samples 19 may be combined to a block 140. In the example shown in fig. 1, each block 140 comprises twelve samples 19. However, any other number of samples 19 may be combined in a block 140, e.g., such as 4, 8, 16, 32, 64 samples 19.
[0048] The digital waveform data 14 may comprise one or more (e.g., two more more) channel groups 23a-c, wherein each channel group 23a-c comprises one or more (e.g., two more more) channels 21a-g. Two or more channel groups 23a-c may have the same number or different number of channels 21a-g. In the example shown in fig. 1 , the digital waveform data 14 comprises three channel groups 21 a-c, wherein a first channel group 23a comprises three channels 21 a-c, a second channel group 23b comprises two channels 21 d, e, and a third channel group 23c comprises two channels 21 f, g. However, any other number of channel groups 23 and / or size of channel groups 23 may be used.
[0049] Blocks 140 within a channel 21 a-g and / or within a channel group 21 a-c may have the same number of samples 19. Furthermore, all channel groups 21 a-c may have the same number of samples 19 per block 140. Alternatively, the number of samples 19 per block 140 may differ, e.g., between channel groups 21 a-c. Blocks 140 within a channel group 21 a-c may be temporally aligned, e.g., wherein a starting point and end point of a block 140 temporally aligns with a starting point and end point of blocks 140 in other channels 21 a-c of the same channel group 23a-c. Blocks 140 may form a basis for optional predictive coding. However, other units for predictive coding may be used. Furthermore, channels 31 a-g may not be separated into block 140.
[0050] A plurality of blocks 140 may form a temporal block 30. A temporal block 30 may cover one or more channels 21 a-g of a channel group 23a-c (e.g., all three channels 21 a-c of channel group 23a) and or one or more blocks 140 in a temporal direction. Temporal blocks 30 may comprise blocks 140 that are exclusively assigned to the same channel group 23a-c (e.g., no blocks 140 of other channel groups 23a-c). A frame 27a-c (e.g., frame of a group of channels) may indicate a collection (e.g., array) of samples 19 from one group of channels 23a-c. A frame 27a-c may have a common temporal starting and end point of the channels 21a-g of the corresponding channel group 23a-c (e.g., form a rectangular array of samples 19). For example, the first channel group 23a may comprise three channels 21 a-c, wherein each frame 27a comprises 72 samples 19 (24 samples 19 of each channel 21 a-c). However, the frame 27a may comprise any other number of samples 19. A frame 27a-c may comprise samples 19 in units of blocks 140 and / or temporal blocks 30. In other words, a frame 27a-c may comprise blocks 140 and / or 30 only in their respective entirety.
[0051] Two or more (e.g., all) frames 27a-c (e.g., also across collections 31 a, b) of different channel groups 23a-c may have a same temporal length (e.g., measured in ms) and / or have the same temporal starting point (e.g., be temporally aligned). In the example shown in fig. 1, the frames 70a-c have the same temporal lengths. Since channel group 23c has a higher sample rate, frame 70c therefore comprises more samples 19 per channel 21 f, g. In a different example, different channel groups 23a-c may have the same number of samples 19 (e.g., regardless of different sampling rates).
[0052] A frame 27a-c may be encoded into one or more packets (e.g., one or more payload packets). For example, frame 27a may be encoded into a single packet 25a (or payload packet 25) or multiple payload packets 25a (e.g., one or more Network Abstraction Layer, NAL, units). In the example shown in fig. 1 , each frame 27a-c is encoded into one payload packet 25a-c. However, any other number of payload packets 25a-c may be used. Furthermore, encoding of a frame 27a-c may result in an encoding into one or more non-payload packet (not shown in fig. 1).
[0053] In the example shown in fig. 1, the samples 19 may be encoded in the time domain or a transform domain (e.g., frequency domain). Therefore, the samples 19 may be provided in form of transform coefficients. Furthermore, the samples 19 may be encoded with or without predictive coding. For example, the samples 19 may be based on a prediction residual.
[0054] Encoding (and decoding) frames 27 may be performed in an order that uses one or more criteria, wherein, for example, multiple criteria have different priorities. The encoding order of frames 27a-c may use any combination of prioritization based on one or more of the criteria including earliest starting sample, earliest ending sample, and smallest (or highest) channel group index. Encoding frames 27a, b, may be performed in a temporal and channel group order, wherein, for example, a temporal order is prioritized. For example, the encoder 10 may be configured to encode a frame 27a-c with an earliest starting time (e.g., wherein a left most sample 19 has the earliest temporal rank compared to other non-encoded frames 27a-c). In case of two or more frames 27a-c having identical earliest starting times (e.g., the left border of two or more frames 27a-c temporally align), the respective two or more frames 27a-c are encoded according to a channel group priority (e.g., a lowest channel group index being encoded first). However, other priority rules may be used, e.g., the next frame 27a-c to be encoded is the frame 27a-c with a earliest end point (e.g., a frame 27a-c with a leftmost right edge). In the example shown in fig. 1 , frames 27a-c are selected so as to temporally align at the start and end of frames 27a-c. The frames 27a-c may subsequently be encoded in a channel group order (e.g., according to a smallest channel group index), e.g., from top to bottom.
[0055] Fig. 3 shows a schematic view of a decoder 12 for decoding from a data stream 16 digital waveform data 14, wherein channels 21a-g of the digital waveform data 14 are coded into the data stream 16 in channel groups 23a-c. The decoder 12 is configured to locate within the data stream 16 payload packets 25a-c into which the one or more channels 21a-g are encoded channel-group 23a-c wise and frame 27a-c wise so that each packet 25a-c is associated with a channel group 23a-c and has exclusively encoded a frame 27a-c of one or more channels 21a-g thereinto which are within the channel group 23a-c with which the respective packet 25a-c is associated, and decode a collection 31a, b of selected channel groups 23a-c from payload packets 25a-c associated with a channel group 23a-c which is contained by the one or more selected channel groups 23a-c of the collection 31a, b.
[0056] The decoder 12 shown in fig. 3 may be any decoder 12 enclosed herein and / or comprise any functionality described herein (e.g., in context of a decoder 12 and / or encoder 10).
[0057] Fig. 4 shows a flow diagram of a method 110 for decoding from a data stream 16 digital waveform data 14, wherein channels 21a-g of the digital waveform data 14 are coded into the data stream 16 in channel groups 23a-c. The method 110 may be performed by any decoder 12 disclosed herein. Furthermore, any feature disclosed herein related to a decoder (or a encoder) may be applicable to the method and vice versa.
[0058] The method 110 comprises, in step 112, locating within the data stream 16 payload packets 25a-c into which the one or more channels 21a-g are encoded channel-group 23a-c wise and frame 27a-c wise so that each packet 25a-c is associated with a channel group 23a-c and has exclusively encoded a frame 27a-c of one or more channels 23a-c thereinto which are within the channel group 23a-c with which the respective packet 25a-c is associated.
[0059] The method 110 further comprises, in step 114, decoding a collection 31a, b of selected channel groups 23a-c from payload packets 25a-c associated with a channel group 23a-c which is contained by the one or more selected channel groups 23a-c of the collection 31a, b.
[0060] The method 110 may comprise decoding all collections 31 a, b (e.g., further collections 31a, b, if further collections are located within the data stream 16). The method 110 may comprise decoding payload packets 25a-c collection 31a, b wise.
[0061] The method 110 may comprise receiving the data stream 16 (e.g., from an encoder 10) and / or reproduce (e.g., play, render, output) the decoded waveform data 14 (e.g., a portion thereof such as the decoded collection).
[0062] Further may be provided a computer program (e.g., computer program product) for performing the method 100, 110, or any method disclosed herein, when the computer program runs on a computer (e.g., on a processor, e.g., microprocessor). The computer program may be stored on a digital storage medium, e.g., a non-transitory storage medium (e.g., a compact disc, a hard drive, a USB stick, or cloud computing resources). The computer program may comprise instructions that cause a processor to perform the steps of any method disclosed herein. Any encoder 10 and / or decoder 12 disclosed herein may comprise a processor (e.g., microprocessor) configured to perform any method disclosed herein. Any encoder 10 and / or decoder 12 disclosed herein may comprise a data storage for storing the computer program (e.g., instructions thereof).
[0063] Further may be provided a data stream 16 (e.g., stored on a digital data storage, e.g., non-transitory storage medium) having encoded therein a digital waveform data using any encoding method disclosed herein.
[0064] Examples for High-level syntax for coding of digital waveform data are presented herein.
[0065] 1. Introduction to embodiments Consider a sequence of digital waveform data that shall be encoded into a bitstream, e.g., for the purpose of transmission or storage. Such a digital waveform signal may, for example, be an electroencephalogram (EEG), or an electrocardiogram (ECG), or an electromyogram (EMG) or a seismic signal or an audio signal. Typically, a digital waveform signal comprises or contains one or more channels that have same (e.g., temporal) characteristics such as the same sampling rate, e.g., the same number of samples per second in each channel. However, it may also happen that a signal comprises or contains channels with different sampling rates. For example, an EEG recording may have 64 EEG channels recorded at 1000 Hertz and an additional channel containing the body temperature recorded at 0.1 Hertz.
[0066] Assume, such a signal shall be encoded by a lossless or lossy compression scheme that is capable of encoding blocks of n channels (e.g., a time interval or set of samples of channel groups 23a, b with channels 21a-e) where each channel has the same number of samples. This requires to partition the input signal into blocks (e.g., blocks 140 and / or temporal blocks 30) that can be handled by the compression scheme. The encoder of such a compression scheme shall be denoted as block encoder (e.g., any encoder 10 disclosed herein) and the decoder of such a compression scheme shall be denoted block decoder (e.g., any decoder 12 disclosed herein).
[0067] In the following, different embodiments of the encoder 10 and decoder 12 that are described in different chapters below. However, one or more features of any chapter may be combined with any feature or embodiment of any other chapter, unless stated otherwise. The embodiments described in different chapters are therefore not exclusive to each other.
[0068] 2. Summary
[0069] See, the independent claims enabling a channel group wise and frame wise coding, thereby rendering it possible to separate channels, e.g., of differing sampling rate. As described later, each payload packet, relating to a certain frame of a certain channel group, may be coded / decoded individually, or, in case of the usage of the example with independent and dependent frame / payload packets, within each channel group, each sequence of immediately successive one or more payload packets consisting of a leading independent payload / frame packet followed by zero, one or more dependent payload / frame packet. The encoder 10 may be configured to provide the data stream 16 with one or more waveform parameter set (WPS) packets 33a, b, each indicating a collection 31 a, b of one or more selected channel groups 23a-c (e.g. so that each selected channel group collection 31a, b forms a subset out of the channel groups 23a-c, e.g., the subset may be a proper subset out of the channel groups 23a-c or comprise all channel groups 23a-c; in case of more than one WPS (and WPS packet 33a, b, respectively), the subset (i.e. the collection) of different WPS packets may be mutually exclusive, i.e. one channel group 23a-c may solely be included in one collection 31a, b) and comprising a WPS index. In the example shown in fig.
[0070] 1, channel groups 23a, b are in channel group collection 31a and channel group 23c is in channel group collection 31b. In such a case, two WPS packets 33a, b may be provided in the data stream 16, one WPS packet 33a for the channel group collection 31a and one WPS packet 33b for the channel group collection 31b. However, any other number and / or size of channel group collection may be used. The data stream 16 may comprise each WPS packet 33a, b once (e.g., at the start of the data stream 16) or multiple times, e.g., in regular time or packet intervals, e.g., at random access points.
[0071] The WPS packets 33a, b may allow the data stream 16 to be coded in units of collections 31a, b of channel groups 23a-c (e.g., wherein the collections 31a, b, themselves may be coded in units of channel groups 23a-c). For example, the data stream 16 may be decodable such that one of frames 27a-c (or any other unit of sample groups) of each channel group 23a-c of the same collection 31 a, b is decoded before decoding another one of frames 27a-c of each channel 23a-c of the next collection 31a, b. For example, in the data stream 16 shown in fig. 1, a frame 27a and a frame 27b (e.g., one sample of each channel) may be decoded first, since the corresponding channels 21a-e are associated with the first collection 31a. Subsequently, frame 27c of channels 21 f, g may be decoded, which are associated with the second collection 31b. For example, collections 31a, b, of channels groups 23a-c may be encoded in different sampling rates. Decoding in units of collections 31a, b may thusly allow efficient coding of channels 21a-g with different characteristics such as different sampling rate.
[0072] The encoder 10 may be configured to provide each payload packet 25a-c with a WPS reference index which references a corresponding WPS packet via the WPS index of the corresponding WPS packet, wherein the respective payload packet 25a-c is associated with a channel group 23a-c which is contained by the one or more selected channel groups 23a-c of the collection 31a, b indicated by the corresponding WPS packet 33a, b. The encoder 10 may be configured to provide each payload packet 25a-c with, if a number of selected channel groups 23a-c of the collection 31a, b indicated by the corresponding WPS packet 33a, b is larger than one, a channel group index indicting the channel group 23a-c with which the respective payload packet 25a-c is associated. The WPS reference index and / or the channel group index may be indicated in a header of the respective payload packets 25a-c.
[0073] The decoder 12 may be configured to decode from the data stream 16 a predetermined waveform parameter set (WPS) packet 33a, b (the WPS packet encoded by encoder 10) indicating the collection 31a, b of one or more selected channel groups 23a-b and comprising a WPS index, and decode from each payload packet 25a-c a WPS reference index which references the predetermined WPS packet 33a, b via the WPS index of the predetermined WPS packet 33a, b, to derive a selected channel group 23a-c with which the respective payload packet 25a-c is associated and which is contained by the one or more selected channel groups 23a-c of the collection 31 a, b indicated by the predetermined WPS packet 33a, b, with decoding from the respective payload packet 25a-c, if a number of selected channel groups 23a-b of the collection 31a, b indicated by the predetermined WPS packet 33a, b is larger than one, a channel group index indicting the channel group 23a-c with which the respective payload packet 25a-b is associated.
[0074] For example, a payload packet 25a associated with the first channel group 23a may comprise a WPS reference index (e.g., zero out of zero and one) that references a corresponding WPS packet 33a associated with the first channel group 23a. In the example shown in fig. 1 , the first channel group 23a is associated with the first collection 31 a of channel groups 23a-c, which has two channel groups 23a, b and therefore more than one channel group 23a-c. Therefore, the payload packet 25a may further comprise a channel group index indicating the first channel group 23a out of the first and second channel groups 23a, b.
[0075] In a different example, a payload packet 25c is associated with the channel group 23c, which is the only channel group 23c of the second collection 31b of channel groups 23a-c. Therefore the payload packet 25c may comprise only a WPS reference index that references the second collection 31 b of channel groups, but no channel group index, as only the channel group 23c is contained in the second collection 31b. Not signaling a channel group index may reduce the amount of signaled data.
[0076] The encoder 10 may be configured to provide the data stream 16 with a waveform parameter set (WPS) packet 33a, b, indicating a collection 31 a, b of one or more selected channel groups 23a-c (e.g. so that the selected channel group collection 31a, b forms a subset out of the channel groups 23a-c, e.g., the subset may be a proper subset out of the channel groups 23a-c or comprise all channel groups 23a-c). The encoder 10 may be configured to provide each payload packet 25a-c with, if a number of selected channel groups 23a-c of the collection 31 a, b indicated by the WPS packet 33a, b is larger than one, a channel group index indicting the channel group with which the respective payload packet is associated.
[0077] The decoder 12 may be configured to decode from the data stream 16 a predetermined waveform parameter set (WPS) packet 33a. b indicating the collection 31 a, b of one or more selected channel groups 23a-c, and decode from each payload packet 25a-c, if a number of selected channel groups 23a-c of the collection 31a, b indicated by the predetermined WPS packet 33a, b is larger than one, a channel group index indicting the channel group 23a-c with which the respective payload packet 25a-c is associated.
[0078] Some or each payload packet 25a-c may not be provided an explicit index referencing the WPS packet 33a, b. For example, the data stream 16 may be structured such that an association between a payload packet 25a-c and a WPS packet 33a can be inferred. In one example, a lack of a WPS reference index may indicate association with a WPS, wherein, for example, two WPS packets 33a, b (or any other number such as one, three, four, or more WPS packets) are provided and a WPS reference index is only provided for payload packets 25a-c associated with one of the two WPS packets 33a, b, but not for the other one of the two WPS packets 33a, b. Similarly, if only one WPS packet is provided, no WPS reference index may need to be signaled. In another example, a sequence of the payload packets 25a-c may indicate an association with the WPS packet 33a, b. For example, the WPS packets 33a, b, may be provided in a sequence and sets of payload packets 25a-c may be provided in a sequence (e.g., similar sequence) representative of the sequence of WPS packets 33a, b.
[0079] The decoder 12 may be configured to decode from the data stream 16 a predetermined waveform parameter set (WPS) packet 33a, b indicating the collection 31 a, b of one or more selected channel groups 23a-c, and decode from each payload packet 25a-c, if a number of selected channel groups 23a-c of the collection 31a, b indicated by the predetermined WPS packet 33a, b is larger than one, a channel group index indicting the channel group 23a-c with which the respective payload packet 25a-c is associated. The encoder 10 may be configured to perform the grouping so that within each channel group 23a-c, if containing more than one channel, all channels of the respective channel group coincide in a sampling rate. The encoder 10 may be configured to group all channels with the same sampling rate into a single channel group 23 or into more than one channel group 23. In the example shown in fig. 1 , for each of the channel groups 23a-c, channels 23a-g that have the same sampling rate amongst other channels 21a-g within the same channel group 23a-c. For example, all channels 21a-c of the first channel group 23a have the same sampling rate. Similarly, all channels 21 f, g of the third channel group 23c have the same sampling rate. While two different channel groups 23a, b are provided with channels 21a-e that all have the same sampling rate, a single channel group 23c is provided with channels 21 f, g that all have the same sampling rate. However, any other combination of channel groups 23a-c may be provided, e.g., wherein all channels 21 a-g having the same sampling rate are grouped into a single channel group 23 (e.g., channel group 23a comprising channels 21a-e and channel group 23b comprising channels 21 f, g), or one or more channel groups 23 having channels 21 a-g with different sampling rates.
[0080] In case of WPS referencing, the channel groups may be decoded selectively, namely in units of collections of these channel groups, one for each WPS.
[0081] WPS parameter transmission enables to lower the overhead of channel group association of the individual payload packets by adapting the overhead for transmitting the WPS reference. For example, no overhead for aligning different sample rates may be required. Furthermore, the probability of grouping similar wave pattern in the same channel group 23a-c may be increased.
[0082] Fig. 5a shows a schematic view of a data stream 16 with a packet 35. The packet 35 may, for example, be a payload packet or a WPS packet 31 a.
[0083] The encoder 10 may be configured to distinguish the one or more waveform parameter set (WPS) packets 31a, b and the payload packets 25a-c by encoding into each packet 35 (which may be, for example, a payload packet 25a-c or a WPS packet 31a, b) of the data stream 16 a packet type indicator 37 which indicates a packet type of the respective packet 35 out of a plurality of packet types including a WPS packet type and one or more payload packet types. The packet type indicator 37 may be in a header of the packet 35. For example, the packet type indicator 37 may indicate, e.g., in form of a bit sequence mappable to a packet type, that the packet 35 is a WPS packet 33, b. The decoder 12 may be configured to distinguish the predetermined waveform parameter set (WPS) packet 33a, b and the payload packets 25a-c by decoding from each packet 25 of the data stream 16 a packet type indicator 37, which indicates a packet type of the respective packet 35 out of a plurality of packet types including a WPS packet type and one or more payload packet types.
[0084] Therefore, the decoder 12 can identify WPS packets 33a, b and payload packet 25a-c. Compared to coding schemes, in which the packets already use a header structure for identifying a packet type, the use of WPS packets 33a, b may be included with no or little additional bits, while benefitting from coding efficiency of assignment of channels 21a-g to channel groups 23a-c.
[0085] The following describes a packet format, for example, for encoding arbitrary digital waveform signals into one bitstream.
[0086] Note: The bytes in a packet, for example, shall be interpreted as a sequence of bits that starts with the most significant bit (MSB) of the first byte, proceeding to the least significant bit (LSB) of the first byte, then proceeding with the MSB of the second byte and so forth. However, other interpretations of bit sequences are possible as well.
[0087] Pseudo-code notation follows, for example, the method of specifying syntax element in the H.266 / VVC-standard. For example, u(6) means, the next 6 bits are read from the bitstream and interpreted as an unsigned 6 bit integer. ue(v) means, decode an exponential Golomb code of order 0 and interpret it as unsigned integer. See H.266 / VVC for details.
[0088] It is assumed that optionally each packet (e.g., packet 35) comprises or contains a syntax element (comprising or consisting of one or more bits, e.g., packet type indicator 37) that indicates the type of packet out of a list of allowed types. The list of allowed types may be pre-determined (e.g., as part of a standard) or may be signaled or indicated in the bitstream 16 (e.g., at the start). The packet 35 may, for example, be a network abstraction layer (NAL) unit. The packet type indicator 37 may be signaled in the header of such a NAL unit.
[0089] Exemplary pseudo-code where the packet type is signalled as a one byte syntax element nal_unit_type: nal_unit_header( ) { Descriptor nal_unit_type u(8)
[0090] }
[0091]
[0092] It is noted that the packet type indicator 37 may be an 8-bit unsigned integer as exemplarily shown in the above pseudo-code, but may have any other bit length (e.g., two, three, four, five, six, or more bits) and / or have a different data type.
[0093] The encoder 10 may be configured to encode (at least, i.e. at least all of this type) each payload packet 25a-c (e.g., each payload packet 25a-c of one or more or all channel groups 23a-c) without an stream pointer to an end of the respective payload packet 25a-c.
[0094] The decoder 12 may be configured to determine an end of (at least, i.e. at least all of this type) each payload packet by parsing the respective payload packet till the end.
[0095] For example, the end of a payload packet 25a-c may be derivable based on a size of the payload packet 25a-c (e.g., due to a use of a fixed length of payload packets) and / or a structure of a the payload packet 25a-c.
[0096] Furthermore, it may be assumed that each packet can be decoded by a corresponding decoder 12, e.g., without knowing the number of bytes present in the current packet. After decoding of the packet, the decoder may know the position of the last byte of the packet and therefore, it also knows the size of the packet.
[0097] The encoder 10 may be configured to encapsulate each packet 35 (e.g., payload packet 25a-c) of the data stream 16 using a predetermined file format to form the data stream 16. Alternatively, the encoder 10 may be configured to concatenate the packets to form the data stream 16, with separating the packets using start codes and start-code emulation prevention bytes (optionally, the latter codes and bytes may also be used for the encapsulation case).
[0098] Each packet 35 of the data stream may encapsulated using a predetermined file format and the decoder 12 may be configured to decapsulate each packet 35 of the data stream 16. Alternatively, the packets 35 may be concatenatedly written into the data stream 16 and the decoder 12 may be configured to derive the packets from the data stream by sequentially parsing the data stream and detecting the beginning of a new packet using a start code, (for example, using the mechanism as in Annex B of H.266 / VVC), and by performing start- code-emulation-prevention-byte-removal (e.g. from actual packet data). E.g., the detection and removal may optionally also be conducted in case of decapsulation.
[0099] Fig. 5b shows a schematic view of a packet 35 within the data stream 16, wherein the packet comprises data portion 39 for encapsulation or concatenation. The data portion 39 may, for example define a length field or a start code.
[0100] Encapsulation may be or comprise using a length field (e.g., in a header of the packet 35) to indicate a size of the packet. Decapsulation may entail parsing the length field and identifying an end of a packet after reaching a packet length obtained from the length field. Using a start code may comprise a fixed bit sequence (e.g., “000001” or “00000001”) at a start of a packet 35. The decoder 12 may subsequently identify an end of a packet due to a beginning of a new start code of a subsequent packet.
[0101] Such packets can optionally either be encapsulated in an existing file format that handles packets like, e.g., the MP4 file format, or the packet format as specified in Annex B of the H.266 / VVC video compression standard (denoted AnnexB), or they can simply be written into the bitstream one after another (e.g., concatenated). The latter case may be suitable for storage of a bitstream in a file, but it may not allow random access because it may not be possible to efficiently seek a particular packet for example except for the first one in the bitstream. Packet formats like MP4 or Annex B support seeking packets and therefore are a prerequisite for Random access or parallel decoding.
[0102] The plurality of packet types may (e.g., in addition to a WPS packet type and one or more payload packet types) further include an annotation channel packet type (e.g., AC_NUT) and an auxiliary metadata (AM) packet type (e.g., AM NUT) and the one or more payload packet type comprise an independent frame packet type and a dependent frame packet type.
[0103] An annotation channel may, for example, store or represent event annotations, e.g., of medical events such as sleep stages, seizures, or stimulus events. The annotation channel may, for example, hold textual event annotations (e.g., which may be decodable into a text, e.g., “stimulus onset”). The annotation channel may also store non-medical events such as a start of a measurement (e.g., seismic measurement). The auxiliary metadata packet type may include packets for supplemental (e.g., non-essential) data. Such packets may be used for storing, e.g., device configuration data, sub-titles, titles, patent name, and patient data.
[0104] The encoder 10 may be configured to encode the packet type indicator 37 in manner so that the packet type indicator 37 indicates none of the packet types using an all zero bitstring.
[0105] In a preferred embodiment, one or more of the following packet types shall be defined (e.g. one or more of the following):
[0106] Waveform parameter set (WPS)
[0107] Independent frame (IF)
[0108] Dependent frame (DF)
[0109] Annotation channel (AC)
[0110] Auxiliary metadata (AM)
[0111] Exemplary values for syntax element nal_unit_type:
[0112] nal_unit_typ Name of Description and associated syntax structure e nal_unit_type
[0113] 0 FORBIDDEN_NUT Forbidden nal unit type for start code emulation prevention
[0114] 1 WPS_NUT Waveform parameter set waveform_parameter_set_rbsp( )
[0115] 2 IF_NUT Independent frame
[0116] independent_frame_rbsp( )
[0117] 3 DF_NUT Dependent frame
[0118] dependent_frame_rbsp( )
[0119] 4 AC_NUT Annotation channel
[0120] annotation_channel_rbsp( )
[0121] 86 AM_NUT Auxiliary metadata
[0122] auxiliary_metadata_rbsp( )
[0123]
[0124] 2.1 Auxiliary metadata concept
[0125] The encoder 10 may be configured to encode the packet type indicator 37 immediately followed by a predetermined bit sequence 41 (e.g., byte sequence, e.g., of two, three, four, five, or more bytes) into a packet 35 of the data stream 16 which is of the AM packet type, wherein a bitstring of the packet type indicator associated with the AM packet type immediately followed by the predetermined bit sequence equals a predetermined ASCII-sequence (e.g., American Standard Code for Information Interchange) which is indicative of a file format of the data stream 16.
[0126] The decoder 12 may be configured to decode the packet type indicator 37 immediately followed by a predetermined bit sequence from a packet of the data stream which is of the AM packet type, determine from a bitstring composed of the packet type indicator associated with the AM packet type immediately followed by the predetermined bit sequence an ASCII-sequence, and identify a file format of the data stream using the ASCII sequence. Fig. 5a exemplarily shows a predetermined bit sequence 41 immediately following a packet type indicator 37.
[0127] In another preferred embodiment, the first four bytes of a AM packet shall equal a predefined ASCII-Sequence, for example "VAWC", which corresponds to the unsigned 8 bit integer values 86, 65, 87, and 67. Furthermore, this four byte sequence shall not occur for any of the other packet types. In the case where packets are stored one after another in a file (without further packeting syntax), the first for bytes of an AM packet may, for example, act as a four character code to identify the current format. This may, for example, require that an AM packet is the first packet written to the file. However, it is still possible to encapsulate such an AM packet into a regular packet of a packet format like, e.g., AnnexB or MP4. For example, in the ASCII-Sequence “VAWC”, “V” may be the packet type indicator 37 for the AM packet type, followed by a three byte sequence “AWC” for identifying the file format. In a different example, “VAWC” forms a predetermined four byte sequence following a (separate) packet type indicator 37 indicating the AM packet type.
[0128] Fig. 5c shows a schematic view of an exemplary WPS packet 33a with general coding information 43.
[0129] The encoder 10 may be configured to encode into a packet of the data stream 16 which is of the AM packet type general coding information revealing (e.g., one or more of)
[0130] - an upper limit of the number of channels present in the data stream 16 (e.g., the upper limit being a power of two, e.g., 2, 4, 8, 16, 32, 64, 128, e.g., or any other number); - an upper limit of a number of samples contained in each of the channels present in the data stream 16 (e.g., an upper limit for total number of samples within a channel); - an upper limit of a sampling rate of each of the channels present in the data stream 16 (e.g., an upper limit between 1kHz and 192kHz); and
[0131] - a waveform type the digital waveform data 14 relates to.
[0132] The decoder 12 may be configured to decode from a packet 35 of the data stream 16 which is of the AM packet type general coding information revealing on one or more of
[0133] - an upper limit of the number of channels present in the data stream;
[0134] - an upper limit of a number of samples contained in each of the channels present in the data stream;
[0135] - an upper limit of a sampling rate of each of the channels present in the data stream; - a waveform type the digital waveform data relates to.
[0136] The waveform type may indicate a shape of wave forms, e.g., one or more of sine, square, triangle, and sawtooth. The waveform type may indicate a set of signal types or a signal type of such a set, e.g. a set of physiological signal types (e.g., EEG, ECG, EMG, pulse, body temperature), or indicate of these signal types. The waveform type may indicate an audio signal type, astrosignal types, or geophysical signal type.
[0137] The general coding information may indicate any of this information explicitly (e.g., with an encoded value that indicates a limit or waveform type) or an index representative of the information (e.g., an index that indexes an upper limit or a waveform type). The general coding information may indicate a value or syntax element that indicates two or more of the two information indicated above, e.g., a syntax element, which is indicative of both, an upper limit of a number of channels as well as a number of samples contained in each channel. The general coding information may comprise or consist of one or more of flags and syntax elements.
[0138] The AM packets may comprise or may contain auxiliary metadata that comes, for example, from the input file. For example, header information of an 'edf' (e.g., European data format) file (see https: / / www.edfplus.info), or metadata from an audio file in the well-known 'wav' format as specified int ITU-R BS.2088-1. The AM packet may enable a decoder (e.g., decoder 12) to output the decoded waveform sequence in the format it was originally stored in, like, e.g. 'edf' or 'wav'.
[0139] An exemplary pseudo-code and semantics for an AM packet is as follows: auxiliary_metadata_rbsp( ) { Descriptor am_fourcc_id_last_three_bytes / * Equal to 0x415743* / u(24) am_header_crc32 u(32) am_reserved_flag u(1) am waveform type u(2) am_length_signal_mode u(1) am allow reconfig flag u(1) am_copyright_flag u(1) am_original_flag u(1) am_private_flag u(1)
[0140] a m st ream max sam p 1 i n g rate m i n us 1 u(24) am_stream_max_num_channels_minus1 u(16) if( am_length_signal_mode )
[0141] am stream num samples per ch u(32) if( am waveform type = = WT BS2088 ) {
[0142] am_metadata_reserved_flag u(1) am metadata n um bytes m i n us1 u(31) for( i = 0; i <= am_metadata_num_bytes_minus1 ; i++ ) am_metadata_payload_bytes[ i ] u(8)
[0143] } else if( am waveform typt = = WT EDF PLUS ) { am_num_channels_edf u(16) for( i = 0; i < 256 * ( am_num_channels_edf + 1 ); i++ ) am_edf_header_payload_bytes[ i ] u(8)
[0144] }
[0145] }
[0146]
[0147] am_fourcc_id_last_three_bytes must equal 0x415743.
[0148] am_header_crc32 is the CRC (e.g., Cyclic redundancy check) calculated over the byte sequence starting with the byte containing the am_reserved_flag until the end of auxil- iary_metadata_rbsp( ).
[0149] am_reserved_flag shall be ignored.
[0150] am_waveform_type specifies the waveform type according to table 1.
[0151] Table 1 - Name association to am waveform type and type description am wave- Name of am wave- Type of waveform
[0152] form type form type
[0153] 0 WT GENERIC None, not signalled
[0154] 1 WT EDF PLUS EDF+, www.edfplus.info
[0155] 2 WT BS2088 BW64, ITU-R BS.2088-1
[0156] 3 WT RESERVED reserved, astro- or geophysical signal
[0157]
[0158] am_length_signal_mode equal to 1 indicates that syntax element am stream num sam- ples_per_ch is present in the bitstream.
[0159] am_stream_max_sampling_rate_minus1 plus 1 specifies the maximum sampling rate present in the bitstream.
[0160] am_stream_max_num_channels_minus1 plus 1 specifies the maximum number of channels in the bitstream.
[0161] am_stream_num_samples_per_ch specifies the number of samples per channel present in the bitstream.
[0162] am_metadata_reserved_flag shall be ignored.
[0163] am_metadata_num_bytes_minus1 plus 1 specifies the number of metadata payload bytes present in the AM RBSP.
[0164] am_metadata_payload_bytes[ i ] specifies the i-th metadata payload byte. The array am_metadata_payload_bytes is a bitstream according to ITU-R BS.2088-1.
[0165] The pseudo code example above includes multiple syntax elements for indicating general coding information. However, the general coding information may include less syntax elements, e.g., only related to one or more limits or only related to a waveform type, or any other combination.
[0166] 2.2 Waveform parameter set concept
[0167] The encoder 10 may be configured to encode into a packet (e.g., WPS packet 33a, b) of the data stream 16 which is of a WPS packet type a sequence of channel group syntax portions which sequentially (e.g., in coding order), along a channel group order defined among the selected channel groups 23a-c, indicate a number of channels 21a-g contained by the one or more selected channel groups 23a-c of the collection 31a, b.
[0168] For example, each channel group syntax portion may comprise a channel number syntax element indicating the number of channels contained in an associated channel group 23a-c,
[0169] - a repetition number syntax element indicating how many channel groups 23a-c following the associated channel group 23a-c along the channel index order coincide with the associated channel group in the number of channels contained, and - a flag indicating whether the respective channel group syntax is the last in the sequence of channel group syntax portions.
[0170] The decoder 12 may be configured to decode from a packet 35 of the data stream 16 which is of a WPS packet type a sequence of channel group syntax portions which sequentially, along a channel index order defined among the plurality of channels, indicate a number of channels contained by the one or more selected channel groups of the collection.
[0171] For example, the WPS packet 33a may comprise a sequence of channel group syntax portions, which allows determining a number of channel groups 23a-b (e.g., two channel groups) and number of channels 23a-e (e.g., five channels) in the first channel group collection 31a. In one example, the sequence of channel group syntax portions may comprise a channel number syntax element indicating three channels in channel group 23a, a repetition number syntax element indicating no repetition (or no repetition number syntax element indiating in order to indicate no repetition), a flag indicating that first channel group 23a is not the last channel group, a channel number syntax element indicating two channels in channel group 23b, a repetition number syntax element indicating a repetition (or no repetition number syntax element indiating in order to indicate no repetition), and a flag indicating that the respective channel group syntax is the last in the sequence of channel group syntax portions.
[0172] The WPS packets may, for example, store information about the channel groups and their sizes present in the bitstream.
[0173] In a preferred embodiment, each WPS packet contains or comprises an identifier (e.g. denoted wps_waveform_parameter_set_id) that is used by other packets to relate to a particular WPS according to the following pseudo-code:
[0174] waveform_parameter_set_rbsp( ) { Descriptor wps_waveform_parameter_set_id u(4)
[0175] / * Further syntax elements of the WPS * /
[0176] }
[0177]
[0178] In a preferred embodiment, the following pseudo-code is employed to encode the existing channel groups and their sizes:
[0179] NumChannelGroups = 0
[0180] TotalNumChannels = 0
[0181] do {
[0182] wps_num_channels_in_next_group_minus1 ue(v) wps_nu m_chan nel_g rou p_repetitions ue(v) for( j = 0; j <= wps_num_channel_group_repetitions; j++ ) {
[0183] NumChannels[ NumChannelGroups++ ] = wps_num_channels_in_next_group_minus1 + 1
[0184] TotalNumChannels += wps_num_channels_in_next_group_minus1 + 1
[0185] }
[0186] wps_more_channel_groups_present_flag u(1) } while( wps_more_channel_groups_present_flag )
[0187]
[0188] wps_num_channels_in_next_group_minus1 + 1 specifies the number of channels in the next channel group. wps_num_channel_group_repetitions specifies how many channel groups with the same number of channel as the previous channel group are appended to the list of channel groups. wps_more_channel_groups_present_flag specifies whether there are more channel groups present in the bitstream or not.
[0189] After decoding the above syntax structure, the decoder knows there exist NumChan- nelGroups channel groups in the bitstream and the total number of Channels over all channel groups is TotalNumChannels. Each channel group may be associated with a channel group index c with 0 <= c < NumChannelGroups.
[0190] In a preferred embodiment, the output order of the channel of all channel groups starts with the first channel of the first channel group, proceeding to the last channel of the first channel group, then proceeding with the first channel of the next channel group and so forth. In a preferred embodiment, all channels in a channel group have the same sampling rate (e.g., all channels 21a-c in channel group 23a may have the same sampling rate). In a preferred embodiment, different channel groups may have different sampling rates (e.g., channels 21d-e in channel group 23b may have a different sampling rate as the channels 21 f, g in channel group 23c). However, more than one channel group 23a-c may have the same sampling rate. For example, multiple channels 21 having a same sampling rate may be associated with different channel groups 23, e.g., in order to reduce coding complexity, adapt to coding rules (e.g., allowing only for a maximum amount of channels in a channel group), or maintaining other associations (e.g., channels for a similar type of signal, e.g., same or similar sensor).
[0191] The encoder 10 may be configured to encode into a packet 35 (e.g., WPS packet 33a, b) of the data stream 16 which is of a WPS packet type a channel reordering information indicating how the channels 21a-g of the selected channel groups 23a-c are to be reordered for channel output. For example, channels 21a-e of channel group collection 31a may be reordered before encoding. WPS packet 33a channel group collection 31a may comprise channel reordering information on how the channels 21a-e are to be reordered (e.g., after sample decoding). The WPS packet 33a may comprise a syntax element (e.g., flag, e.g., wps_channel_reordering_flag) indicating whether the channels 21a-g of a collection are to be reordered.
[0192] The decoder 12 may be configured to decode from a packet 35 of the data stream which is of the WPS packet type a channel reordering information indicating how the channels of the selected channel groups 23a-c are to reordered for channel output.
[0193] The encoder 10 may be configured to encode the channel reordering information in form of a sequence of channel rank swaps, each indicating two channels whose channel indices are to be swapped so that the two channels swap in channel index order (e.g. wherein a channel order index is determined by immediately consecutively assigning ascending or descending channel indices along the channel group order).
[0194] The decoder 12 may be configured to decode the channel reordering information in form of a sequence of channel rank swaps, each indicating two channels whose channel indices are to be swapped so that the two channels swap in channel index order.
[0195] For example, after decoding samples of channels of a collection, the decoder 12 may obtain channels in a sequence that results from decoding the channels 21. The decoder 12 may subsequently assign the decoded channels 21 an index according to such an intermediate sequence of channels (e.g., indexing 0, 1, 2, 4, and so own). The channel reordering information may provide channel rank swaps, e.g., (0,4) indicating that the channels with indices “0” and “4” are to be swapped. The decoder 12 may subsequently reassign indices or keep using the initial indices.
[0196] In a preferred embodiment, a channel reordering syntax after decoding is optionally decoded according to the following pseudo-code:
[0197] wps_channel_reordering_flag u(1) if( wps_channel_reordering_ flag ) {
[0198] wps_num_channel_swaps_minus1 ue(v) for( i = 0; i <= wps_num_channel_swaps_minus1 ; i++ ) { wps_swap_first_index[ i ] ue(v) wps_swap_second_index_minus_first_index_minus1[ i ] ue(v) }
[0199] }
[0200]
[0201] wps_channel_reordering_flag indicates whether channel reordering after decoding is applied. If it is equal to 1 , a sequence of channel swap operations is decoded that need to be applied in decoding order.
[0202] wps_num_channel_swaps_minus1 + 1 indicates the number of channel swaps to be decoded. A channel swap comprises or consists of two integer values wps_swap_first_index and wps_swap_second_index_minus_first_index_minus1. This means that channel with index values wps_swap_first_index is swapped with channel with index values ( wps_swap_first_index + wps_swap_second_index_minus_first_index_minus1 + 1). Note that in this way, each permutation of the channels can be represented since it is a well known mathematical fact that any permutation can be represented as a composition of swaps.
[0203] The encoder 10 may be configured to encode into a packet (e.g., WPS packet 33a) of the data stream 16 which is of the WPS packet type an information on how many annotation channels are comprised in the data stream 16 which accompany the collection 31a, b of selected channel groups. For example, the number of annotation channels of a collection 31a, b of selected channel groups may be zero, one, two, three, or more.
[0204] The decoder 12 may be configured to decode from a packet (e.g., WPS packet 33a) of the data stream 16 which is of the WPS packet type an information on how many annotation channels are comprised in the data stream which accompany the collection of selected channel groups.
[0205] In a preferred embodiment, the bitstream shall also carry annotation channel data. For example, this may be byte sequences that carry annotation channels as specified in the EDF+ file format (see https: / / www.edfplus.info). The number of annotation channels present in the bitstream can be signalled by a syntax element according to the following pseudo-code:
[0206] wps_num_annotation_channels ue(v)
[0207]
[0208] 2.3 Independent frame concept
[0209] The encoder 10 may be configured to, in providing each payload packet 25a-c with, if a number of selected channel groups 23a-c of the collection 31a, b indicated by the corresponding WPS packet 33a, b is larger than one, a channel group index indicating the channel group 23a-c with which the respective payload packet 25a-c is associated, check whether a number of selected channel groups 23a-c of the collection 31a, b indicated by the corresponding WPS packet 33a, b is larger than one, and, if the number of selected channel groups of the collection indicated by the corresponding WPS packet is larger than one, encode the channel group index into the respective payload packet 25a-c, and, if the number of selected channel groups of the collection indicated by the corresponding WPS packet is not larger than one, leave the respective payload packet 25a-c without the channel group index.
[0210] The encoder 10 may thusly perform a check for whether there are more than one selected channel groups 23a-c in a collection 31a, b and only encode channel group index into the respective payload packet 25a-c if there are two or more channel groups 23a-c in the collection 31a, b. If a collection 31a, b only comprises one channel group, an identification of a channel group by a channel group index may not be necessary, as there is only one channel group 23a-c to choose from. For example, the encoder 10 may check whether the first collection 31a has more than one channel groups 23a-c, identifies two channel groups 23a, b and subsequently encodes a channel group index into the respective payload packet 25a, b of the first collection 31a. Similarly, the encoder 10 may check whether the second collection 31b has more than one channel group 23a-c, identifies only one channel group 23c and leaves the respective payload packet 25c without a channel group index. The decoder 12 may be configured to, in decoding from each payload packet 25a-c, if a number of selected channel groups 23a-c of the collection 31a, b indicated by the corresponding WPS packet is larger than one, a channel group index indicating the channel group 23a-c with which the respective payload packet 25a-c is associated, check whether a number of selected channel groups 23a-c of the collection 31a, b indicated by the corresponding WPS packet 33a, b is larger than one, and, if the number of selected channel groups 23a-c of the collection 31a, b indicated by the corresponding WPS packet 33a, b is larger than one, decode the channel group index from the respective payload packet 25a-c, and, if the number of selected channel groups 23a-c of the collection 31 a, b indicated by the corresponding WPS packet 33a, b is not larger than one, infer that the respective payload packet 25a-c has the one selected channel group 23a-c encoded thereinto.
[0211] For example, when decoding a payload packet 25c of the second collection 31b, the decoder 12 may check the corresponding WPS packet 33b and conclude that the collection 31b has only one channel group 23c. The decoder 12 may thusly infer that the respective payload packet 25c has the one selected channel group 23c encoded thereinto. No signaling of this information may be required, which may therefore improve coding efficiency and may reducing parsing complexity.
[0212] The encoder 10 may be configured to, in providing each payload packet 25a-c with, if a number of selected channel groups 23a-c of the collection 31a, b indicated by the corresponding WPS packet 33a, b is larger than one, a channel group index indicting the channel group with which the respective payload packet is associated, if the number of selected channel groups of the collection indicated by the corresponding WPS packet is larger than one, encode the channel group index into the respective payload packet using a fixed length code whose code length monotonically increases with a number of selected channel groups in the collection (e.g., dependent on NumChannelGroups).
[0213] The decoder 12 may be configured to, in decoding from each payload packet 25a-c, if a number of selected channel groups 23a-c of the collection 31a, b indicated by the corresponding WPS packet 33a, b is larger than one, a channel group index indicting the channel group 23a-c with which the respective payload packet 25a-c is associated, if the number of selected channel groups 23a-c of the collection indicated by the corresponding WPS packet 33a, b is larger than one, decode the channel group index into the respective payload packet 25a-c using a fixed length code whose code length monotonically increases with a number of selected channel groups in the collection.
[0214] The scaling of the code lengths with the number of channel groups allows adjusting the channel group index to the number of channel groups. Since the number of selected channel groups can be derived centrally in the corresponding WPS packet, the channel group index for multiple payload packets 25a-c can be scaled accordingly. Therefore, coding efficiency may be improved, especially for lower numbers of channel groups 23a-c in a collection 31a, b.
[0215] It is noted that in the description above, the term “corresponding WPS packet” is used instead of “WPS packet”. Such terminology is also used in section “2. Summary”, in which payload packets are provided with a WPS reference index. However, principle of signaling a channel group index dependent on whether there is more than one channel group 23a-c in a collection 31a, b, is also applicable to embodiments without a WPS reference index. For the sake of completeness, such embodiments are described below.
[0216] The encoder 10 may be configured to, in providing each payload packet 25a-c with, if a number of selected channel groups 23a-c of the collection 31a, b indicated by the WPS packet 33a, b is larger than one, a channel group index indicting the channel group 23a-c with which the respective payload packet 25a-c is associated, check whether a number of selected channel groups 23a-c of the collection 31a, b indicated by the WPS packet 33a, b is larger than one, and, if the number of selected channel groups 23a-c of the collection 31a, b indicated by the WPS packet 33a, b is larger than one, encode the channel group index into the respective payload packet 25a-c, and, if the number of selected channel groups 23a-c of the collection indicated by the WPS packet is not larger than one, leave the respective payload packet 25a-c without the channel group index.
[0217] The decoder 12 may be configured to, in decoding from each payload packet 25a-c, if a number of selected channel groups 23a-c of the collection 31a, b indicated by the WPS packet 33a, b is larger than one, a channel group index indicting the channel group 23a-c with which the respective payload packet 25a-c is associated, check whether a number of selected channel groups 23a-c of the collection 31a, b indicated by the WPS packet 33a, b is larger than one, and, if the number of selected channel groups 25a-c of the collection 31a, b indicated by the WPS packet 33a, b is larger than one, decode the channel group index from the respective payload packet 25a-c, and, if the number of selected channel groups 25a-c of the collection 31a, b indicated by the WPS packet 33a, b is not larger than one, infer that the respective payload packet 25a-c has the one selected channel group 25a-c encoded thereinto.
[0218] The encoder 10 may be configured to, in providing each payload packet 25a-c with, if a number of selected channel groups 23a-c of the collection indicated by the WPS packet 33a, b is larger than one, a channel group index indicting the channel group 23a-c with which the respective payload packet 25a-c is associated, if the number of selected channel groups 23a-c of the collection 31 a, b indicated by the WPS packet 33a, b is larger than one, encode the channel group index into the respective payload packet 25a-c using a fixed length code whose code length monotonically increases with a number of selected channel groups 23a-c in the collection 31a, b.
[0219] The decoder 12 may be configured to, in decoding from each payload packet 25a-c, if a number of selected channel groups 23a-c of the collection 31a, b indicated by the WPS packet 33a, b is larger than one, a channel group index indicating the channel group 23a-c with which the respective payload packet 25a-c is associated, if the number of selected channel groups 23a-c of the collection 31a, b indicated by the WPS packet 33a, b is larger than one, decode the channel group index into the respective payload packet 25a-c using a fixed length code whose code length monotonically increases with a number of selected channel groups 23a-c in the collection 31a, b.
[0220] The payload packets 25a-c may be each of an independent frame packet type or a dependent frame packet type, wherein the encoder 10 is configured to encode the one or more channels 21a-g into the payload packets 25a-c using block-wise coding (e.g., in units of blocks 140 and / or temporal blocks 30).
[0221] The encoder 10 may be configured to encode into a payload packet 25a-c of the data stream 16 which is of the independent frame packet type a set of parameters comprising one or more of
[0222] - a range of block sizes used for an adaptive setting of a block size (e.g., size of block 140) in the block wise coding (e.g., using if max min block size),
[0223] - a range of bit depths used in the block wise coding (e.g., using if_max_min_bit_depth),
[0224] - a perceptual coding mode switch indicating a use or non-use of perceptual coding for the block wise coding (e.g., using if_perceptual_mode), an inter-channel prediction mode switch indicating a use or non-use of inter-channel prediction for the block wise coding (e.g., using if_allow_cross_channel_pred_flag),
[0225] The encoder 10 may be configured to, in the encoding the one or more channels 21 a-g into the payload packets 25a-c, encode a payload packet 25a-c of the data stream 16 which is of the independent frame packet type and is associated with a predetermined channel group 23a-c (e.g., signaled with if_channel_group_id) and, if present, one or more following payload packets 25 (not shown in fig. 1, but may be arranged to the left of packets 25a-c, assuming a coding direction towards the right; e.g., the packets may still be referred to as payload packets 25a-c, considering their association to channel groups 23a-c) of the data stream 16 being of the dependent frame packet type, being associated with the predetermined channel group 23a-c (e.g. note that the equality in channel group may be determined, for example, based on WPS index and channel group index) and having one or more frames (not shown in fig. 1 , but may be arranged to the left, assuming a temporal direction toward the right, e.g., the frames may still be referred to as frames 27a-c, considering their association with channel groups 23a-c) encoded thereinto temporally immediately following a frame 27a-c coded into the payload packet 25a-c of the data stream 16 which is of the independent frame packet type using the set of parameters encoded into the payload packet of the data stream 16 which is of the independent frame packet type.
[0226] The decoder 12 may be configured to decode the one or more channels 21 a-g from the payload packets 25a-c using block-wise decoding, decode from a payload packet 25a-c of the data stream which is of the independent frame packet type a set of parameters comprising one or more of
[0227] - a range of block sizes used for an adaptive setting of a block size in the block wise decoding,
[0228] - a range of bit depths used in the block wise decoding,
[0229] - a perceptual coding mode switch indicating a use or non-use of perceptual coding for the block wise decoding,
[0230] an inter-channel prediction mode switch indicating a use or non-use of inter-channel prediction for the block wise decoding,
[0231] The decoder 12 may be configured to, in the decoding the one or more channels 21 a-g from the payload packets 25a-c, decode a payload packet 25a-c of the data stream 16 which is of the independent frame packet type and is associated with a predetermined channel group 21 a-c and, if present, one or more following payload packets 25 of the data stream being of the dependent frame packet type, being associated with the predetermined channel group (e.g. note that the equality in channel group may be determined based on WPS index and channel group index) and having one or more frames 27 encoded thereinto temporally immediately following a frame 27a-c coded into the payload packet 25a-c of the data stream 16 which is of the independent frame packet type using the set of parameters encoded into the payload packet 25a-c of the data stream 16 which is of the independent frame packet type.
[0232] An independent frame may be a frame that can be coded without dependency (e.g., in form of prediction dependency) from other previously coded frames. The independent frame may be coded using intra-prediction. An independent frame may be located at a random access point. A dependent frame may be a frame coded using information (e.g., sample information) of previously coded frames, e.g., in form inter-prediction (or inter-frame prediction). For example, within a dependent frame, a block 140 may be predicted by referencing a block 140 of a previously coded frame (e.g., by a vector, e.g., a sample offset), and coding a difference between the block 140 to be coded and the referenced block 140. Predictive coding may be performed in blocks 140, but also in other units (e.g., temporal blocks 30).
[0233] Block wise coding may be performed in a zig-zag pattern, wherein blocks 140 of a same temporal rank are coded in channel order and once all blocks 140 in a channel group 21a-c of a same temporal rank are coded, the coding proceeds to a block 140 of a temporal later rank.
[0234] The range of blocks 140 may range, for example, from 4 to 128 samples 19, e.g., 8 to 32 samples 19, or any other range. Perceptual coding may include considering of human perception. For example, perceptual coding may include removal (or attenuation) of frequencies not perceivable by a human and / or masking (e.g., modifying amplitudes of frequencies according to human perception and / or higher amplitudes in other frequencies). The interchannel prediction mode switch may be provided in form of a flag or may inferred from other signaling. Using the set of set of parameters may entail using all parameters of the set of parameters or a subset thereof. Other parameters may be used in addition or alternatively such as a filtering option, and a switch for cross channel prediction, and block matching prediction. Further parameters are described below in section “2.3.1.1 independent frame RBSP semantics”. Using the set of parameters of the independent frame for dependent frames may allow reducing signaling for the dependent frames. The grouping into channel groups 23a-c may improve similarity between channels and may improve temporal alignment for more efficient referencing of earlier frames 27 and / or blocks 140.
[0235] The payload packets 25a-c may be each of an independent frame packet type or a dependent frame packet type, wherein the encoder 10 is configured to in the encoding the one or more channels 21 a-g into the payload packets 25a-c, encode a payload packet 25a-c of the data stream 16 which is of the independent frame packet type and is associated with a predetermined channel group 23a-c independent from other payload packets 25a-c, and, if present, each of one or more following payload packets 25 of the data stream 16 being of the dependent frame packet type, being associated with the predetermined channel group 23a-c (e.g. note that the equality in channel group may be determined based on WPS index and channel group index) and having one or more frames 27 encoded thereinto temporally immediately following a frame 27a-c coded into the payload packet 25a-c of the data stream 16 which is of the independent frame packet type using coding dependencies from any preceding payload packet associated with the predetermined channel group, preceding the respective payload packet.
[0236] The decoder 12 may be configured to, in the decoding the one or more channels 21 a-g from the payload packets 25a-c, encode a payload packet of the data stream 16 which is of the independent frame packet type and is associated with a predetermined channel group 23a-c independent from other payload packets 25a-c, and, if present, each of one or more following payload packets 25 of the data stream 16 being of the dependent frame packet type, being associated with the predetermined channel group 23a-c (e.g. note that the equality in channel group may be determined based on WPS index and channel group index) and having one or more frames encoded thereinto temporally immediately following a frame 27a-c coded into the payload packet 25a-c of the data stream 16 which is of the independent frame packet type using coding dependencies from any preceding payload packet 25 associated with the predetermined channel group 23a-c, preceding the respective payload packet.
[0237] Using coding dependencies may comprise forming a difference between a frame to be coded (e.g., a sample 19 or a block 140 thereof) and a frame of a preceding payload packet. Coding dependencies may include temporal dependency (e.g. referencing a temporally preceding block 140) and / or inter channel dependency within the same channel group 23a-c (e.g., referencing a block 140 of a different channel 21a-g within the same channel group 23a-c).
[0238] For example, frame 27a of the first channel group 23a may be of an independent frame packet type. A frame following frame 27 of the first channel group 23a may be of a dependent packet type and coding of the following frame 27 may be performed dependent on the frame 27a (e.g., using inter-frame prediction). Similarly, a further frame may follow, which is also of a dependent packet type, which may be coded dependent on the independent frame 27a and / or the previously coded dependent frame.
[0239] Coding dependencies may be limited to the same channel group 23a-c. For example, coding channel 21 f may allow coding dependencies in regards to channel 21g of the same channel group 21c, but not in regards to channel 21a of channel group 23a. Optionally, coding dependencies may be allowed or not allowed between channel groups 23a-c of the same collection 31a, b. For example, coding dependencies between channels 21a-c of channel group 23a and channels 21d-e may be allowed (e.g., if they were distributed into different channel groups due to channel group constrains, but would have the same characteristics, e.g., same sampling rate) or may be not allowed (e.g., in order to reduce coding complexity).
[0240] An IF (independent frame) packet (e.g., payload packet 25a-c) may contain or comprises a sequence of encoded blocks (e.g., blocks 140) that belong to one channel group and of one WPS (e.g., one collection 31a, b). As an example, a block decoder may only require the corresponding WPS (e.g., WPS packet 33a, b) in order to decode an IF.
[0241] Exemplary syntax and semantics for an IF packet:
[0242] independent_frame_rbsp( ) { Descriptor if_waveform_parameter_set_id u(4) if( NumChannelGroups > 1 )
[0243] if_channel_group_id / * e.,g. changing to fixed length depending on ue(v) NumChannelGroups as in dependent:frame_rbsp( )7
[0244] if_length_signal_mode_flag u(1) if_frame_length_shift u(2) if max min block size u(6)
[0245]
[0246] if_perceptual_mode u(2) if _max_min_bit_depth u(6) if_allow_cross_channel_pred_flag u(1) if( if_allow_cross_channel_pred_flag ) {
[0247] if_cc_pred_filtering_mode u(2) if_allow_cc_pred_mult_hyp_flag u(1) for( ch = 1 ; ch < NumChannels[ if_channel_group_id ]; ch++ ) { for( n = 0; n <= if_allow_cc_pred_mult_hyp_flag; n++ ) { CrossChannelPredlnputChDistMinusI [ ch ][ n ] = 0
[0248] }
[0249] }
[0250] }
[0251] if_allow_block _matching_pred_flag u(1) if( if_allow_block_matching_pred_flag ) {
[0252] if_bm_pred_filtering mode u(2) if_allow_bm_pred_mult_hyp_flag u(1) if_allow_bm_offset_pred_prev_ch_flag u(1) for( ch = 0; ch < NumChannels[ if_channel_group_id ]; ch++ ) { for( n = 0; n <= if_allow_bm_pred_mult_hyp_flag; n++ ) { BlockMatchingPredOffsetMinusBlocksSize[ ch][ n ] = 0 Log2BlockMatchingPredBlockSize[ ch][ n ] = 0
[0253] }
[0254] }
[0255] }
[0256] if_allow_lpf u(1) if( if_ allowjpf ){
[0257] if_lpf_allow_prev_ch_flag u(1) LPFMaxNumWeightsNoPrevCh = 16
[0258] for( ch = 0; ch < NumChannels[ if_channel_group_id ]; ch++ ) { for( n = 0; n <= LPFMaxNumWeightsNoPrevCh; n++ ) { LPFWeightsNoPrevChPred[ ch ] [ n ] = 0
[0259] }
[0260] }
[0261] }
[0262] if_residual_quant_mode u(2) if_ch_indep_interval_idx u(4) DepChMask = ( 2 « if_ch_indep_interval_idx ) - 1
[0263] if_max_abs_delta_qp_idx u(3)
[0264]
[0265] MaxAbsDeltaQP = ( 1 « if_max_abs_delta_qp_idx ) - 1
[0266] if( if_length_signal_mode_flag )
[0267] if_indep_num_samples_per_channel_minus1 u(32) if_indep_init_block_qp u(8) for( i = 0; i < NumChannels[ if_channel_group_id ]; i++ ) {
[0268] CurrBlockQP[ i ] = if_indep_init_block_qp
[0269] CurrZeroLSB[ i ] = 0
[0270] }
[0271] byte_alignment( )
[0272] frame_data( NumChannels[ if_channel_group_id ] )
[0273] rbsp_trailing_bits( )
[0274] }
[0275]
[0276] 2.3.1.1 Independent frame RBSP semantics
[0277] if_waveform_parameter_set_id specifies the value of wps_waveform_parameter_set_id for the WPS in use.
[0278] if_channel_group_id identifies the channel group to which the current independent frame belongs. When if_channel_group_id is not present, it is inferred to be equal to 0.
[0279] if_length_signal_mode_flag equal to 1 specifies that a syntax element ifjn- dep_num_samples_per_channel_minus1 is present.
[0280] if_frame_length_shift specifies an offset for deriving the variable Log2FrameLength as follows:
[0281] Log2FrameLength = Log2MaxBlockSize + if framejength shift (1) if_max_min_block_size specifies an index for deriving variable Log2MaxBlockSize as follows:
[0282] Log2MaxBlockSize = LutBlockSizeMaxLog2[ if max min block size ] (2) The value of if max min block size shall be in the range of 0 to 62, inclusive.
[0283] The array LutBlockSizeMaxLog2[ ] is specified as follows:
[0284] LutBlockSizeMaxLog2[ ] = (3) {
[0285] 4, 5, 5, 6, 6, 6, 7, 7, 7, 7, 8, 8, 8, 8, 8, 9,
[0286] 9, 9, 9, 9, 9, 10, 10, 10, 10, 10, 10, 10, 11, 11, 11, 11,
[0287] 11, 11, 11, 11, 12, 12, 12, 12, 12, 12, 12, 12, 12, 13, 13, 13, 13, 13, 13, 13, 13, 13, 14, 14, 14, 14, 14, 14, 14, 14, 14
[0288] }
[0289] The array LutBlockSizeMinLog2[ ] is specified as follows: LutBlockSizeMinLog2[ ] = (4) {
[0290] 4, 4, 5, 4, 5, 6, 4, 5, 6, 7, 4, 5, 6, 7, 8, 4,
[0291] 5, 6, 7, 8, 9, 4, 5, 6, 7, 8, 9, 10, 4, 5, 6, 7,
[0292] 8, 9, 10, 11, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 5, 6, 7,
[0293] 8, 9, 10, 11, 12, 13, 6, 7, 8, 9, 10, 11, 12, 13, 14
[0294] }
[0295] The variable MaxSplitDepth is derived as follows:
[0296] MaxSplitDepth = Log2MaxBlockSize - LutBlockSizeMinLog2[ if max min block size ] (5) if_perceptual_mode specifies that a mode that uses coding tools specifically targeting psychvisual or psychoacoustic qualities of the reconstructed signal (rather than targeting objective signal fideltiy in terms of a metric like 11 or I2 distance) is enabled if_max_min_bit_depth specifies an index for deriving the variables BitDepthMax and BitDepthMin as follows:
[0297] BitDepthMax = LutBitDepthMax[ if_max_min_bit_depth ] (6) BitDepthMin = LutBitDepthMin[ if_max_min_bit_depth ] (7) The value of if_max_min_bit_depth shall be in the range of 0 to 62, inclusive.
[0298] The array LutBitDepthMax[ ] is specified as follows:
[0299] LutBitDepthMax[ ] = (8) {
[0300] 3, 4, 4, 8, 8, 8, 8, 8, 8, 12, 12, 12, 12, 12, 12, 12,
[0301] 12, 12, 12, 16, 16, 16, 16, 16, 16, 16, 16, 16, 16, 16, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 24, 24, 24, 24, 24, 24, 24, 24, 24, 24, 24, 28, 28, 28, 28, 28, 28, 28, 28, 28, 28, 28
[0302] }
[0303] The array LutBitDepthMin[ ] is specified as follows:
[0304] LutBitDepthMin[ ] = (9) {
[0305] 2, 2, 3, 2, 3, 4, 5, 6, 7, 2, 3, 4, 5, 6, 7, 8,
[0306] 9, 10, 11, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 9, 10,
[0307] 11 , 12, 13, 14, 15, 16, 17, 18, 19, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27
[0308] }
[0309] if_allow_cross_channel_pred_flag equal to 1 specifies that the cross channel prediction mode is allowed. if_cc_pred_filtering_mode specifies the allowed filtering options that may be applied to the cross channel prediction signal as follows:
[0310] If if_cc_pred_filtering_mode is equal to 0, no filtering may be applied to the cross channel prediction signal.
[0311] If if_cc_pred_filtering_mode is equal to 1 , a half-pel filtering of the cross channel prediction signal is allowed.
[0312] If if_cc_pred_filtering_mode is equal to 2, a half-pel filtering and a full-pel filtering of the cross channel prediction signal are allowed.
[0313] The value of if_cc_pred_filtering_mode shall lie in the range from 0 to 2 inclusively.
[0314] if_allow_cc_pred_mult_hyp_flag equal to 1 specifies that cross channel prediction with two input channels is allowed.
[0315] if_allow_block_matching_pred_flag equal to 1 specifies that the block matching prediction mode is allowed.
[0316] if_bm_pred_filtering_mode specifies the allowed filtering options that may be applied to the block matching prediction signal as follows:
[0317] If if_bm_pred_filtering_mode is equal to 0, no filtering may be applied to the block matching prediction.
[0318] If if_bm_pred_filtering_mode is equal to 1 , a half-pel filtering of the block matching prediction signal is allowed.
[0319] If if_bm_pred_filtering_mode is equal to 2, a half-pel filtering and a full-pel filtering of the block matching prediction signal are allowed.
[0320] The value of if_bm_pred_filtering_mode shall lie in the range from 0 to 2 inclusively.
[0321] if_allow_bm_pred_mult_hyp_flag equal to 1 specifies that block matching prediction with two hypothesis is allowed.
[0322] if_allow_bm_offset_pred_prev_ch_flag equal to 1 specifies that offsets for the block matching prediction can be predicted from offsets for the block matching prediction of the previous channel.
[0323] if_allow_lpf equal to 1 specifies that linear predictive filtering is allowed.
[0324] if_lpf_allow_prev_ch_flag equal to 1 specifies that linear predictive filtering using input samples from up to 3 previous channels is allowed.
[0325] if_residual_quant_mode specifies the quantization mode.
[0326] if_ch_indep_interval_idx specifies the variable DepChMask = ( 2 « if ch indepjnter-valjdx ) - 1 . For the decoding process of Section 8, the channels can be grouped into consecutive groups of channels, each consisting of at most DepChMask +1 many channels, such that each group of channels can be processed independently from each other group of channels.
[0327] if_max_abs_delta_qp_idx specifies the Variable MaxAbsDeltaQP = ( 1 « if_max_abs_delta_qp_idx ) - 1.
[0328] if_indep_num_samples_per_channel_minus1 plus 1 specifies the number of samples per channel present in the current frame sequence.
[0329] if_indep_init_block_qp specifies the quantization parameter.
[0330] 2.4 Dependent frame concept
[0331] Like an IF packet, a DF (dependent frame) packet also comprises or contains a sequence of encoded blocks (e.g., blocks 140) that belong to one channel group (e.g. channel groups 23a-c) of one WPS (e.g., collections 31a, b). Decoding of a DF packet is the continuation of the decoding process after a preceding DF or IF packet.
[0332] In other words, an encoded channel group sequence starts with one IF packet (e.g., a frame 27 of a channel group 23a-c), followed by zero or more DF packets (e.g., frames of the same channel group 23a-c) and the decoder decodes them in sequence. After the last DF packet, an IF packet may occur in the bitstream for the same channel group. This may act as a random access point to start decoding if the packet format supports seeking to this packet (like it is e.g. the case in the packet format specified in Annex B of H.266 / VVC).
[0333] Exemplary syntax and semantics for a DF packet:
[0334] dependent_frame_rbsp( ) { Descriptor df_waveform_parameter_set_id u(4) if( NumChannelGroups < 16 )
[0335] df_channel_group_id_plus1 u(4) else if( NumChannelGroups < 4096 )
[0336] df_channel_group_id_plus1 u(12) else
[0337] df_channel_group_id_plus1 u(20) frame_data( NumChannels[ df_channel_group_id_plus1 - 1 ] ) rbsp_trailing_bits( )
[0338] }
[0339]
[0340] df_waveform_parameter_set_id specifies the value of wps_waveform_parameter_set_id for the WPS in use.
[0341] df_channel_group_id_plus1 minus 1 identifies the channel group to which the current dependent frame belongs.
[0342] 2. 5 Frame data concept
[0343] A syntax structure frame_data() is called in DF and IF packets and it comprises or contains the actual coded samples associated with the frame. The frame data may comprise payload data for decoding the samples 19 of frames 27. For example, in IF frames the frame data may comprise data for absolute values or samples 19 and / or residual data of intra-predicted samples. In DF frames, the frame data may comprise residual data of inter-predicted frames.
[0344] 2.6 Annotation channel concept
[0345] An AC (annotation channel) packet may contain the (ASCII character) bytes of the corresponding annotation channel, for example, according to the EDF+ format. Exemplary syntax and semantics are as follows:
[0346] annotation_channel_rbsp( ) { Descriptor ac_waveform_parameter_set_id u(4) ac_annotation_channel_id ue(v) ac_num_annotation_bytes_div2_minus1 ue(v) byte_alignment()
[0347] annotation_channel_data( )
[0348] }
[0349]
[0350] ac_waveform_parameter_set_id specifies the value of wps_waveform_parameter_set_id for the WPS in use.
[0351] ac_annotation_channel_id specifies the annotation channel index.
[0352] ac_num_annotation_bytes_div2_minus1 is used to determine the number of syntax elements am annotation bytes present in the current AC RBSP as 2 * ( ac_num_annota- tion_bytes_div2_minus1 + 1). annotation_channel_data( ) { Descriptor offset = AnnotationChannelNumSamples[ ac_annotation_channel_id ]
[0353] for( i = 0; i < 2 * ( ac_num_annotation_bytes_div2_minus1 + 1 ); i++ ) { am annotaion byte u(8) AnnotationChannelBytes[ ac_annotation_channel_id ][ offset + i ] = am_annotation_byte
[0354] AnnotationChannelNumSamples[ ac_annotation_channel_id ]++ }
[0355] }
[0356]
[0357] After decoding a AC packet, the array AnnotationChannelBytes may comprise or contain the decoded annotation channel characters. These can be written into an EDF+ file along with the decoded sample data.
[0358] 2.7 Trailing bits concept
[0359] It may be useful to indicate at the end of a NAL unit that it has ended. This can, for example, be achieved by writing a rbsp_trailing_bits() syntax structure like, for example, in the one of H.266 / VVC according to the following pseudo-code:
[0360] rbsp_trailing_bits( ) { Descriptor rbsp_stop_one_bit / * equal to 1 * / f(1) while( !byte_aligned( ) )
[0361] rbsp_alignment_zero_bit / * equal to 0 * / f(1) }
[0362]
[0363] This also ensures that the bitstream is in a byte-aligned position. A byte aligned position may reduce coding complexity and may facilitate parsing and searching for bytes.
[0364] 2.8 Channel groups
[0365] First, the channels (e.g., 21a-g) in the input signal are partitioned into one or more so-called channel groups (e.g., channel groups 23a-c). The grouping may be performed on similar (or identical) channel criteria, such as a same sampling rate and / or according to a channel type (e.g., channels with payload data and channels with annotation data). For example, if there are channels with three different sampling rates present in the input signal, one option would be to create three channel groups, one for each sampling rate, and put each channel in the channel group with the corresponding sampling rate. Furthermore, channel groups with very many channels could further be split into more channel groups in order to distribute the task of encoding / decoding the corresponding channels to more than one block encoder and in order to limit the burden for an encoder / decoder for example with respect to memory capacities since in an envisioned application predictive decoding across channels might be supported within one channel group. For example, in fig. 1, channels 21a-e may have the same sampling rate, but are distributed into two (i.e. more than one) channel groups 23a, b, e.g., in order to reduce coding complexity.
[0366] An order of channels 21a-g within a channel group 23a-c may be dependent (or based on) an order of the channels 21 a-g within the entire digital wave form data 14. For example, the order of the channels 21 a-g in a respective channel group 23a-c may correspond to the order of the channels 21 a-g in the digital wave form data 14 after removing all channels 21 a-g that are not part of the respective channel group 21a-c.
[0367] 2.9 Examples
[0368] 2.9.1 Example 1 : EDF or EDF+
[0369] Assume, for example, an EDF+ input file that comprises or contains data channels shall be encoded. Annotation channels (if present), shall be encoded in AC packets.
[0370] In a preferred embodiment, the ac_annotation_channel_ids (in ascending order) shall be associated with annotation channels in the order as they occur in the EDF+ file. For example, if the EDF+ file contains 5 channels and channels 2 and 4 are annotation channels, channel 2 is associated with ac_annotation_channel_id equal to 0 and channel 4 is associated with ac_annotation_channel_id equal to 1. In this way it is unambiguously possible to restore this channel association at the decoder if the EDF+ header is known.
[0371] In another preferred embodiment, the data channels are grouped into as many channel groups as different sampling rates occur in the EDF+ input file (the sampling rate is the 'nr of samples in each data record' of the corresponding channel divided by 'duration of a data record'). Channels in the IF or DF packets may be or for example must be in the order as they occur in the EDF+ file. The if / df_channel_group_id’s of the resulting channel groups is defined by the first occurrence of a particular sampling rate in channel order. For example, assume, an EDF+ file contains 6 channels with the following sampling rates:
[0372] Channel 0: 100 Hertz
[0373] Channel 1 : 50 Hertz
[0374] Channel 2: 50 Hertz
[0375] Channel 3: 100 Hertz
[0376] Channel 4: 1000 Hertz
[0377] Channel 5: 100 Hertz
[0378] The first occurrences of a particular sampling rates are in the following channels:
[0379] Channel 0: 100 Hertz
[0380] Channel 1 : 50 Hertz
[0381] Channel 4: 1000 Hertz
[0382] Therefore, if / df_channel_group_id equal to 0 encodes all channels with 100 Hertz, if / df_channel_group_id equal to 1 encodes all channels with 50 Hertz, and if / df_chan-nel_group_id equal to 2 encodes all channels with 1000 Hertz.
[0383] In this way, it is possible to unambiguously associate channels in channel groups with output channels of the EDF or (EDF+) file if the EDF (or EDF+) header is known.
[0384] 3. Further embodiments
[0385] The above description is extended in the following by the presentation of further embodiments. Before this, however, the description proceeds with a presentation of a possible framework or codec into which the embodiments described above as well as the embodiments described further below may be built into. Many details described in this framework are, however, optional when being combined with any of the above or subsequently described embodiments. To be more precise, the framework is described with respect to Fig.
[0386] 6 which shows an encoder 10 (e.g., any encoder 10 disclosed herein) for encoding a multichannel digital signal 14 (e.g., digital waveform data 14) into a datastream 16 as well as decoder 12 (e.g., any decoder 12 disclosed herein) for decoding the multi-channel digital signal 14 from datastream 16. This description of Fig. 6 shall be seen as a presentation of new embodiments of the present application which result when combining any of the embodiments described above or any of the embodiments described subsequently is combined with the decoder 12 or encoder 10 of Fig. 6 either by adopting all details / f u nctionalities described with respect to Fig. 6 or with leaving-out some of the details / functionalities described with respect to Fig. 6. Sometimes such “optional” features of Fig. 6 are explicitly identified as being optional with respect to the combination of the previously and subsequently described embodiments, but the just-mentioned possible combinations of the pre-viously / subsequently explained embodiments with the description of Fig. 6 shall not be restricted to the these explicitly identified variations of Fig. 6 in terms of leaving-out certain features.
[0387] In Fig. 6, the multi-channel digital signal 14 is illustrated by way of an array of samples (e.g., samples 19) with the samples being illustrated as small squares 18. Each line / row corresponds to a certain channel (e.g., channels 21 a-g) of the multi-channel digital signal 14. Each channel of signal 14 may have associated therewith a respective channel ID and Fig.
[0388] 6 shows these channels as being ordered according to their channel ID along vertical axis 20 which, thus, corresponds to a “source” channel axis 20. The horizontal axis 22 corresponds to time so that samples 18 forming one column, or being horizontally aligned, are samples belonging to one common time instant. Such set / column of temporally co-located samples 18 is illustrated in Fig. 6 at 24.
[0389] Each channel, thus, forms a digital time-varying signal or time / amplitude or time-to-ampli-tude signal. The multi-channel digital signal m might have been obtained by at least one of Electrocardiography, Electroencephalography, Electromyography or seismic measurement. Differently speaking, the multi-channel digital signal might be a bio-physiological waveform data such as an electroencephalography (EEG) signal, an electrocardiogram (ECG), or an electromyography (EMG) signal, or seismic waveform data. However, each channel / signal might alternatively be another sort of waveform signal data such as scalar media data such as an audio signal and the signal 14 might be a multi-channel audio signal.
[0390] Fig. 6 illustrates the option according to which signal 14 is not coded directly, i.e., in the original domain 26, but in a so-called “coded domain” 28 which might differ from the original domain 26 by one or more of 1) channel transformation, 2) channel permutation and 3) temporal mutual channel alignment. The channel transformation, if applied, transforms, per sample time instant, a set or column 24 of samples from domain 26 to domain 28. Thus, in domain 28, the sample pitch and the time axis is the same as in domain 26, but the meaning of the channels is different, i.e., the “source” channels of domain 26 become transformed channels in domain 28. Accordingly, the vertical axis in Fig. 6 for domain 28 is denoted as 32. Note that the channel transformation might leave the number of channels unchanged so that there is the same number of channels in domain 26 as well as domain 28, but different approaches are also possible. Generally, the channel transformation would aim at reducing redundancy and trying to condense the channels’ energy onto a fewer number of channels in domain 28. As said, the channel transformation is optional. Accordingly, in general terms, the channels in domain 28 are called “coded channels” in order to distinguish them from the “original” or “source” channels of digital signal 14 in domain 26. The permutation is also optional and may be used in combination with, or without, the channel transformation. If used in combination with the channel transformation, the permutation may be performed prior to and / or or subsequent to the channel transformation in order to per-mute / sort the source channels prior to transformation and the coded channels subsequent to the channel transformation. The channel transformation might be a DCT, DST, FFT or any other transformation. The temporal mutual alignment is also optional and might be seen as a constant temporal alignment between the source channels or the coded channels.
[0391] The signal 14 in fig. 6 may depict a digital waveform data 14 to be grouped (in which case channels of the signal 14 may have different sampling rates, unlike indicating by equally spaced samples in fig. 6). The signal 14 shown in fig. 6 may also be regarded as a channel group out of a plurality of channels (wherein other channel groups are not shown in fig. 6), wherein, the channels of the signal 14 may have the same sampling rate. Similarly, the digital waveform data 14 to be grouped may have been modified, e.g., in form of one or more of channel transformation, channel permutation (e.g., for or in preparation of channel grouping), and temporal alignment. Therefore, the signal in the coded domain 28 shown in fig. 6 may be a digital waveform data 14 to be grouped or a channel group that resulted from grouping.
[0392] The module in encoder 10 performing the one or more of channel transformation, channel permutation and temporal mutual alignment is indicated in Fig. 6 as block 34. Side information 36 might be used in order to signal information on one or more of the following: 1) The channel transformation used, 2) information on the permutation(s) among the source channels and / or coded channels and 3) information on the mutual temporal alignment / de-lays between the source channels or coded channels wherein the temporal mutual alignment might be restricted to full sample precision. A corresponding block 38 in decoder 12 performs the reverse step, i.e., performs one or more of: 1) a channel retransformation, 2) a re-permutation of the source channels and / or coded channels and 3) a temporal re-alignment of the source channels or coded channels. Note, that if no channel transformation takes place, the coded channels are, in fact, equal to the source channels except for being temporally mutually aligned or being differently sorted due to permutation. Block 38 might be controlled by the before-mentioned side information 36.
[0393] Thus, the “actual coding” relates to the coded channels in domain 28. In the coded domain 28, the coded channels are depicted in Fig. 6 as lines or rows of samples 40, each extending along time axis 22, the coded channels being depicted one on top of the other along coded channel axis 32 - potentially ordered according to a coded channel ID they have associated therewith - so as to result into an array of samples 40. Again, although Fig. 6 depicts the case that the number of source channels equals the number of coded channels, the number might be different. Further, if channel transformation is used, while there is no longer a clear association between source channels on the one hand and coded channels on the other hand, the temporal association remains: For each temporally co-located samples 24, there is a corresponding temporally co-located set 42 of samples 40 of the coded channels, wherein the set 42 in domain 28 is a column and might be a set of horizontally mutually offset samples in case of, and according to, the mutual temporal alignment, if applied. In case of Fig. 6, it has been assumed that no such temporal alignment took place so that both sets 42 and 24 are pure columns in the time / channel representation.
[0394] The actual coding is done in units of so-called temporal blocks 30. The term “block” or “temporal block” 30 is used so as to denote both a temporal portion of the multi-channel signal in domain 28, i.e., the set of coded channels, as well as a temporal portion of a certain coded channel. That is, for each temporal block 30, each coded channel has a temporal block such as block 140 depicted for some temporal block 30c and same are mutually colocated. The coding is done sequentially along these blocks 140, by following a coding / de-coding order, which traverses the blocks 140 temporal block 30 by temporal block 30 with traversing temporally co-located blocks of the coded channels along a channel order corresponding to the order of the coded channels along axis 32. This coding / decoding order is illustrated in Fig. 6 at 60. That is, in case of temporal block 140 being the block currently to be coded / decoded, the previously decoded / encoded temporal blocks include all preceding temporal blocks of all coded channels as well as the temporally co-located temporal blocks of coded channels preceding the coded channel 92 of temporal block 140 in channel order. These previously coded / decoded temporal blocks and their samples are illustrated in Fig.
[0395] 6 by way of shading. In this regard, note that in Fig. 6, merely one temporal block 140 has been illustrated explicitly in order to reduce the complexity of Fig. 6. Thus, in the specification herein, reference sign 140 is sometimes used to indicate the currently encoded / de-coded temporal block or to stand representatively for all temporal blocks. Further, as depicted in Fig. 6, the partitioning of signal 14 into temporal blocks 30 and 140, respectively, might be done in a manner so that these blocks 30 and 140, respectively, are non-overlapping.
[0396] The actual coding in units of the temporal blocks 140 is performed predictively. That is, the encoder 10 comprises a block predictor 62 which predicts the samples of the currently coded temporal block 140, thereby yielding a prediction signal 64, and the prediction residual 66 formed by a subtraction between the actual sample values of temporal block 140 and the predicted samples of prediction signal 64 formed at a subtractor 68 is coded into the datastream 16 by residual coder 70. The residual coding in residual coder 70 may, or may not, involve a coding error by means of quantization. In any case, block predictor 62 uses the reconstructable version as being available by previously coded temporal blocks in order to obtain the prediction signal 64. This reconstructable version 72 might be derived at encoder 10 by means of a residual decoder 74 which reverses, potentially under coding loss, such as quantization, e.g. by means of dequantization, the residual signal 76 as coded into datastream 16, and an adder 78 which sums-up prediction signal 64 and the reconstructable residual signal 80 as obtained by residual decoder 74. To be more precise, let’s call the channel-individual temporal blocks 140 subblocks with temporally collocated subblocks of all channels forming a temporal block 30. Then, the prediction in module 62 or, to be more precise, the prediction at encoder and decoder, is performed in units of the subblocks 140, i.e. subblock wise. The encoder is free to choose different prediction modes for the subblocks within one block 30. As explained in more detail herein, within one block 30, one subblock 140 may be predicted based on one or more subblocks previously - according to the decoding order 60 - en / decoded within this block 30, while another subblock 140 within that block 30 might be coded / decoded based on the previously en / decoded subblock 140 of the same channel (but within the previous block 30). The transform residual en / decoding is then performed subblock wise by use of a one-dimensional transform signaled in the data stream as described hereinbelow.
[0397] The decoder 12 decodes the coded channels from data stream 16 in a corresponding manner, i.e., in units of the temporal blocks 30 or in temporal blocks 140, respectively, and using predictive decoding. To this end, the decoder 12 comprises a residual decoder 82, an adder 84 and a block predictor 86 which correspond to, and are mutually connected in the same manner as, elements 74, 78 and 62 of encoder 10. That is, the residual decoder 82 derives from the residual signal 76 in data stream 16 the reconstructable residual signal 80 for a currently decoded temporal block 140 which is then subject to addition with prediction signal 64 derived by block predictor 86 for temporal block 140 on the basis of the reconstructed version 72 of previously decoded temporal blocks at adder 84. The output of adder 84, thus, yields the reconstructed version 72 of the currently decoded temporal block 140 and becomes part of the pool of already decoded samples of previously decoded temporal blocks when the temporal blocks of the coded channels are, in this manner, traversed along cod-ing / decoding order 60 so as to reconstruct the coded channels in the coded domain 28.
[0398] Note that the above description concentrated on the so-called sample prediction where samples of a current block 140 are predicted based on reconstructed samples of one or more previously decoded blocks, but coding inter dependencies, namely intra-channel and inter-channel coding dependencies may be exploited not only in terms of sample prediction, but also in terms of other coding tools involving, for instance, parameter prediction and / or context derivation.
[0399] In order to enable a high degree of random access capability, some of the temporal blocks 30 may be coded in a random access manner meaning that the coded channels therein are coded independent from previous temporal blocks 30. Imagine, for instance, that temporal blocks 30b and 30e are random access temporal blocks. Then, none of the temporal channel blocks 140 in temporal block 30b as well as 30e would depend on any preceding temporal block 140 and no coding dependency would cross these temporal blocks 30b and 30e, that is no temporal block 140 within any of temporal block 30b-30d would be coded depending on any block 140 temporally preceding temporal block 30b, and no temporal block 140 within any of temporal block 30e and following would be coded depending on any block 140 temporally preceding temporal block 30e.
[0400] Thus, in other words, coding dependencies are restricted so as to not reach-out beyond the border of a random access temporal block 30b and 30e towards any preceding temporal block 30. Such restriction might also hold for intermediate temporal blocks 30c to 30d between random access temporal blocks 30b and 30e in that same may not depend on any temporal block preceding the leading one among the random access temporal blocks 30b and 30e, here block 30b. Accordingly, leading temporal borders of the random access temporal blocks 30b and 30e are indicated by bold lines in Fig. 6. In a variant, the restriction is not valid for all en / decoding stages. For instance, while the grouping might hold true for prediction, but the residual en / decoding dependencies might cross borders between channel groups. It might be the case, for instance, that for the entropy coding and decoding, all channels are coded jointly, i.e. using a single arithmetic coding engine, but that for the sake of prediction and reconstruction, the channels are grouped as described into independent groups such that, after entropy decoding, each such group can be reconstructed completely independently from each other group. This means that no prediction of sample values or any other information is supported between different channel groups.
[0401] Further, it might be that the coding of the coded channels also interrupts or restricts interchannel dependencies. For example, one or more of the coded channels might be coded as random access coded channels so that same do not use inter-channel dependencies, but merely intra-channel dependencies. The restriction of inter-channel coding dependencies might follow the channel order 32: that is, coding of these random access coded channels and the intermediate coded channels therebetween would be restricted so as to not reach-out beyond such a random access coded channel toward any coded channel preceding that random access coded channel in channel order along axis 32. Two such random access coded channels 88a and 88b and the resulting inter-channel dependency borders are illustrated in Fig. 6. Note that the restriction of inter-channel dependencies might be differently and is illustrated here merely as an example where the definition of, along channel order 32, interspersed random access channels 88a and 88b defines channel groups covering contiguous channels along the channel order 32. Other groups of channels might be defined, which do not necessarily follow the channel order 32, and inter-channel dependencies might be restricted not to render any channel of one group dependent on a channel of any other group, and within each group the inter-channel dependencies may also by restricted or each channel might by coded inter-channel dependent on any previously coded channel within its channel group.
[0402] The block predictor 62 and 86 of encoder 10 and decoder 12, respectively, operate synchronously, i.e., they generate the same prediction signal 64 based on the previously en-coded / decoded samples of previously encoded / decoded temporal blocks 140. On encoder side 10, the prediction for a certain temporal block 140 may be accompanied or determined by one or more prediction parameters. Same might be determined on encoder side based on a rate / distortion optimization. These prediction parameters 90 are coded into data stream 16 and they are decoded from data stream 16 and used by block predictor 86 so as to perform the same prediction. It might be that encoder 10 and decoder 12 support more than one prediction mode. For instance, encoder 10 and decoder 12 may support an intra prediction mode (which mode may also be called block-copy mode) according to which the currently encoded / decoded temporal block 140 is predicted based on the reconstructable sample values of previously encoded / decoded temporal blocks of the same coded channel to which the currently encoded / decoded temporal block 140 belongs, which is coded channel 92 in the example of Fig. 6. Additionally or alternatively, encoder 10 and decoder 12 may support an inter-prediction mode (which mode may also be called cross-channel prediction mode) according to which the currently encoded / decoded temporal block 140 is predicted based on the reconstructable sample values of previously encoded / decoded temporal blocks of one or more coded channels preceding - in coding order 32 - the coded channel 92 to which the currently encoded / decoded temporal block 140 belongs. Additionally or alternatively, there may be a mixed prediction mode according to which the prediction signal 64 is obtained by both, re-constructed / reconstructable sample values of previously encoded / decoded temporal blocks of coded channel 92 itself as well as reconstructed / reconstructable sample values of one or more coded channels preceding coded channel 92 in channel order along axis 32. Beyond this, there may be temporal blocks 140 which are coded without any prediction at encoder 10 and decoded without any prediction at decoder 12 such as the first temporal blocks 140 in the tiles 94 resulting from mutually separating the temporal blocks by means of the random access borders 96 on the one hand and the random access channel borders 98 on the other hand. This corresponds to the prediction signal 64 being set to zero and this may form an additional mode which could be called bypass mode. Additionally, or alternatively, there may be other modes such as ones deriving a DC predictor or linear function predictor for block 64 based on immediately preceding samples which immediately precede block 140. The prediction parameters 90 may, thus, contain for a currently encoded / decoded temporal block 140 a prediction mode flag or prediction mode indicator indicating the prediction mode to be used for this currently encoded / decoded temporal block 140 and, optionally, one or more parameters parameterizing the prediction mode to be used for this currently encoded / decoded temporal block 140. It might also be that the prediction parameters are themselves coded predictively from already reconstructed blocks 140. In this prediction process, the laid out random-access capabilities in channel- and temporal-direction are, as an example, always maintained, i.e. the mentioned prediction of prediction parameters may never be supported across such a random access segment.
[0403] As mentioned, the aforementioned coding dependencies ought not to cross any of the borders 96 and 98 not only result from the just-described sample prediction capabilities of block predictor 62 and 86, respectively, but may optionally also result from other mechanisms such as parameter prediction according to which parameters such as the aforementioned prediction parameters 90 for a certain temporal block 140 are predicted based on coding parameters conveyed in the data stream 16 for any previous temporal block, or context derivation for context-adaptive entropy coding / decoding any coding parameter such as the prediction parameters 90 or any other side information such as side information 76 and 36 for temporal block 140 based on any coding parameter conveyed in the data stream 16 for any preceding temporal block.
[0404] That is, summarizing, the encoder 10 encodes the multi-channel signal 14 by transferring it into the coded domain 28 and then coding the coded channels into data stream 16 in the just-described block-wise and predictive manner, wherein decoder 12 decodes the coded channels of coded domain 28 from data stream 16 and the corresponding block-wise and predictive manner with then gaining the multi-channel signal 14 in its original form 26 based on the coded channels in coded domain 28 by means of segment 38. As said, the channel transformation is optional and if not used, each sample 40 in the coded domain 28 really corresponds to one sample 18 in the original domain 26. If, further, the temporal mutual alignment is not used, each sample 40 exactly corresponds to a sample 18 in the original domain 26 at exactly the same time instant or, differently speaking, all temporally co-located samples 40 in coded domain 28 remain mutually temporally co-located in the original domain 26.
[0405] It should be noted that the temporal blocks 30 might, other than illustrated in Fig.5, vary in block length rather than being of a constant length as depicted in Fig. 6. For instance, encoder 10 may decide on the length of blocks 30 and signal the block length of blocks 30 (and the corresponding temporal blocks 140 of the coded channels) within data stream 16. Such signaling might be done on block level, such as for each temporal block 30 or, differently speaking for each temporally aligned bundle of blocks 140, so that the encoder may decide on the block size on the fly, or the block length might be signaled in the stream 16 on a larger scope such as for a sequence of blocks or even the whole stream 16.
[0406] As to the residual coder and residual decoder 70 and 82, they may use transform coding / decoding in order to convey the residual signal 76 in data stream 16. That is, the residual signal 80 may be conveyed in data stream 16 in transform or spectral domain by way of transform coefficients in residual signal 76. The transform domain might be a DCT, DST or an FFT. The transform may be non-overlapping, i.e. it may only transform residual signal 80 and its re-transform may only cover residual signal 76 within block 140, and / or may be non-windowed, i.e. the residual signal might be transformed without any transform window used to temporally shape the residual signal 80 before the transform. The transform domain, i.e. the transformation leading from time domain to transform domain which is used by the encoder to transform the prediction residual signal 80 to be coded und the corresponding re-transformation leading from transform domain to time domain which is used by the decoder to derive the prediction residual signal 80, or the transformation, might be selected from a set of available transforms including, for instance, one or more of 1) one or more DCTs, 2) one or more DSTs and 3) an identity transform according to which the prediction residual signal 80 is coded into the data stream 14 in time domain directly. The transform may be critically sampled in that the number of transform coefficients resulting from the samples of one block 140 may equal the number of samples of block 140. Again, the samples might be the residual samples or may be, in case of the bypass mode, the channel samples directly.
[0407] The transform coefficients might be encoded by quantization, i.e. they may be quantized with the quantized coefficients then being coded in the datastream 16. Dequantization may occur at decoding. For quantization, either a scalar uniform reconstruction quantizer or a low complexity vector quantizer might be used. In order to determine the quantization indices, the encoder may perform some optimization algorithm such as a rate-distortion optimized scalar quantization, or a trellis quantization with the goal to approximately minimize an approximated Lagrangian rate-distortion cost. At the decoder, the reconstruction process that yiels the transform coefficients may be conducted by multiplying the coded quantization indices with a certain step-size and, in case of the use of a low-complexity vector quantizer, by additionally invoking a state-machine based on the parity of previously decoded quantization indices in order to reconstruct the current quantization index.
[0408] In order to control the quantization noise, the transform coefficients might be subject to noise shaping. Spectral noise shaping may be used to shape the quantization noise spectrally. This may be done by signaling in the data stream spectral-band scale factors, i.e. a scale factor per spectral band, which represent a transfer function of a spectral filter which approximates the spectral envelope of the signal within the current block 140 (or its prediction residual, respectively), or signaling filter coefficients defining a temporal filter having a filter transfer function which approximates the spectral envelope of the signal within the current block 140 (or its prediction residual, respectively). On encoder side, spectral noise shaping may be applied in spectral domain by multiplying an inverse of scale factors, either directly signaled in the data stream or derivable from the filter coefficients by filter-to-factor conversion, with the transform coefficients before quantization. That is, at encoder, the coefficients are shaped by the inverse of the spectral envelope. At decoder side, spectral shaping may be applied in spectral domain by multiplying scale factors, either directly signaled in the data stream or derived from the filter coefficients by filter-to-factor conversion, with the transform coefficients, with then . That is, at decoder, the coefficients are shaped by the spectral envelope before applying retransformation. Additionally or alternatively, temporal noise shaping might be applied. To this end, TNS filter coefficients might be determined and signaled by the encoder. The TNS filter coefficients may represent a transfer function which approximates the temporal envelope of the current block 140 (or its residual signal). The encoder may apply TNS filtering using the filter coefficients by spectrally filtering the possibly spectrally shaped transform coefficients so as to filter them with a transfer function corresponding to an inverse of the temporal envelope. The TNS filter coefficients might be derived by linear prediction analysis of the possibly spectrally shaped transform coefficients so as to derive a linear prediction filter, then used as TNS filter, which minimizes a prediction residual when spectrally applied on the possibly spectrally shaped transform coefficients. At the encoder, the TNS filtered coefficients are then quantized and entropy coded. At decoder side, the inverse takes place: the possibly spectrally shaped transform coefficients are inversely TNS filtered before applying retransformation. Additionally or alternatively, noise filling might be used. The filling may be applied to zero-quantized portions of the spectrum and controlled by the encoder via corresponding noise filling parameters.
[0409] As to the encoding / decoding the block or sequence of quantized transform coefficients of a current block into / from the data stream 16, arithmetic coding, such as context-adaptive binary arithmetic coding, CABAC, may be used. The CABAC encoding / decoding may by performed frame wise. That is, in each channel, the sequence of blocks 140 may be partitioned into immediately consecutive blocks 140, which form frames. This partitioning may be equal among the channels so that, again, a frame denotes both a temporal portion within each channel individually, as well as a temporal portion of the multi-channel signal, i.e. a collection of temporally aligned frames. Within each frame, the sequence of blocks 140 are CABAC en / decoded with once initializing the contexts and resetting the internal CABAC state at the beginning and then updating the contexts’ probabilities during en / decoding the respective frame. That is, blocks 140 are CABAC decodable merely in units of frames. The context initialization might be done independent from previous frames, or depending on the contexts as manifesting itself at the end of, of during, the en / decoding a previous frame. Some deblocking processing might be used to avoid blocking artifacts. If, alternatively, an overlapped transform is used, an overlap-add processing with re-transforms of immediately preceding / succeeding temporal blocks of the same coded channel might be used in order to completely reconstruct the current temporal block’s 140 residual signal 76.
[0410] Besides such transform-(residual)-coded blocks there might be temporal blocks 140 which, additionally or alternatively, are coded using, besides the block prediction by block predictor 62 / 86 - which could be called a primary prediction - a secondary sample-wise prediction of the residual samples in residual block 66 such as by predicting a current sample’s residual sample by means of already decoded values of preceding - in sample coding order - residual samples in block 66 or 80, with then correcting same by means of a secondary-prediction-residual sample decoded from the data stream 16. The secondary-prediction-residual samples for such a block may coded into the data stream en block in a transform domain or sample-wise in time domain.
[0411] Note that the afore-mentioned spectral shaping of the residual signal of a block 140 might be seen as a sample wise residual prediction, i.e. the case where filter coefficients are signaled for a block which define a temporal filter having a filter transfer function which approximates the spectral envelope of the residual signal within a current block 140. In sample wise residual prediction, the residual predictor on a current block 140 might either be chosen out of a fixed set of prediction modes, where an index to such a residual prediction mode is signaled in the bit-stream, or the residual prediction mode might be ‘signal adaptive’. In the latter case, prediction filter coefficients for the residual predictor are determined at the encoder by solving for example a linear equation, and are then quantized and transmitted to the decoder. At the decoder, the coefficients are inverse quantized and then the samplewise prediction is conducted with these coefficients. The number of used coefficients may vary per block and might also be signaled in the bit-stream. Additionally, it might optionally (i.e. indicated by some information in the bit-stream) be supported to invoke collocated samples from a previous block for the sample wise residual prediction. Finally, the coefficients of the sample wise residual prediction might be coded predictively, i.e. be predicted from used coefficients of a previous block, where only the differences to the current coefficients are transmitted.
[0412] A final note shall be made with respect to the juxtaposition of frames, blocks 140, channels and channel groups and regarding decoding order. The description above already described the fact that the channels (e.g., channels 21 a-g) might be grouped into one or more channel groups (e.g., channel groups 23a-c) with each channel group being coded independently from each other, meaning that the blocks 140 in a certain channel group are coded without dependencies from channels outside their channel group. The decoding order 60, thus, would traverse the channels channel-group individually, channel group by channel group. Within each channel group, the blocks 140 are traversed as described: all temporally aligned blocks 140 of all channels fist, then proceeding with the next blocks 140 and so forth. A frame may have a sequence of blocks of a channel group encoded thereinto along the mentioned decoding order, such as n temporally consecutive blocks 140 for all channels of a channel group (e.g., n = 1, 2, 3, 4, or more). If the channel group had m channels, m*n blocks 104 would, thus, be coded into the frame (e.g., frame 27a-c). As mentioned, there might be dependent frames, for which the CABAC contexts are adopted from the preceding frame of the same channel group, i.e. the one having encoded the immediately preceding block 140. For such dependent frames, not only CABAC contexts may be adopted from the preceding frame, but it may also be allowed to allow for prediction from the preceding frame to the dependent frame. Prediction, and possibly also any coding dependencies, towards channels outside the channel group and, within the channel group, towards frames temporally preceding the mostly recently previously en / decoded independent frame would, for example, be disallowed. Thus, each tile shown in Fig. 6 by bold lines may represent a sequence of an independent frame flowed by zero, one or more dependent frames.
[0413] As mentioned before, Fig. 6 only represents a possible “framework” into which the previously described embodiments and the embodiments described subsequently may be built into. Many modifications may be performed with respect to Fig. 6, and some of these modifications might be mentioned in the subsequent description with respect to certain ones of the subsequently described embodiments, but these modifications shall then be treated as being also applicable with respect to other ones of the subsequently described embodiments.
[0414] The description is now resumed with respect to the announced subsequently described embodiments.
[0415] Further Remarks: Also, it should be noted that individual aspects described herein can be used individually or in combination. Thus, details can be added to each of said individual aspects without adding details to another one of said aspects.
[0416] Moreover, features and functionalities disclosed herein relating to a method can also be used in an apparatus (configured to perform such functionality). Furthermore, any features and functionalities disclosed herein with respect to an apparatus can also be used in a corresponding method. In other words, the methods disclosed herein can be supplemented by any of the features and functionalities described with respect to the apparatuses.
[0417] Also, any of the features and functionalities described herein can be implemented in hardware or in software, or using a combination of hardware and software, as will be described in the section “implementation alternatives”.
[0418] Implementation alternatives
[0419] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus.
[0420] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable. Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0421] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine-readable carrier.
[0422] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine-readable carrier.
[0423] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0424] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-tran-sitionary.
[0425] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.
[0426] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0427] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein. A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0428] In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
[0429] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0430] The apparatus described herein, or any components of the apparatus described herein, may be implemented at least partially in hardware and / or in software.
[0431] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0432] The methods described herein, or any components of the apparatus described herein, may be performed at least partially by hardware and / or by software.
[0433] The above-described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.
Claims
Claims1. Encoder (10) for encoding into a data stream (16) digital waveform data (14) comprising a one or more channels (21a-g), configured togroup the one or more channels (21a-g) into one or more channel groups (23a-c),encode the one or more channels (21a-g) into payload packets (25a-c) channel-group wise and frame (27a-c) wise so that each packet (25a-c) is associated with a channel group (23a-c) and has exclusively encoded a frame (27a-c) of one or more channels (21a-g) thereinto which are within the channel group (23a-c) with which the respective packet is associated.
2. Encoder (10) of claim 1 , configured toprovide the data stream (16) with one or more waveform parameter set, WPS, packets (33a, b), each indicating a collection (31a, b) of one or more selected channel groups (23a-c) and comprising a WPS index,provide each payload packet (25a-c) witha WPS reference index which references a corresponding WPS packet (33a, b) via the WPS index of the corresponding WPS packet (33a, b), wherein the respective payload packet (25a-c) is associated with a channel group (23a-c) which is contained by the one or more selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b), andif a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, a channel group (23a-c) index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associated.
3. Encoder (10) of claim 1 , configured toprovide the data stream (16) with a waveform parameter set, WPS, packet, indicating a collection (31a, b) of one or more selected channel groups (23a-c),provide each payload packet (25a-c) withif a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the WPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associated.
4. Encoder (10) of any of the preceding claims, configured toperform the grouping so that within each channel group (23a-c), if containing more than one channel (21a-g), all channels (21a-g) of the respective channel group (23a-c) coincide in a sampling rate.
5. Encoder (10) of claim 2 or any claim depending thereon, configured to, in providing each payload packet (25a-c) with, if a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associatedcheck whether a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, and,if the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, encode the channel group index into the respective payload packet (25a-c), and,if the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is not larger than one, leave the respective payload packet (25a-c) without the channel group index.
6. Encoder (10) of claim 2 or any claim depending thereon, configured to, in providing each payload packet (25a-c) with, if a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associatedif the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, encode the channel group index into the respective payload packet (25a-c) using a fixed length code whose codelength monotonically increases with a number of selected channel groups (23a-c) in the collection (31a, b).
7. Encoder (10) of claim 3 or any claim depending thereon, configured to, in providing each payload packet (25a-c) with, if a number of selected channel groups (23a-c) of the collection (31 a, b) indicated by the WPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associatedcheck whether a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the WPS packet (33a, b) is larger than one, and,if the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the WPS packet (33a, b) is larger than one, encode the channel group index into the respective payload packet (25a-c), and,if the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the WPS packet (33a, b) is not larger than one, leave the respective payload packet (25a-c) without the channel group index.
8. Encoder (10) of claim 3 or any claim depending thereon, configured to, in providing each payload packet (25a-c) with, if a number of selected channel groups (23a-c) of the collection (31 a, b) indicated by the WPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associatedif the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the WPS packet (33a, b) is larger than one, encode the channel group index into the respective payload packet (25a-c) using a fixed length code whose code length monotonically increases with a number of selected channel groups (23a-c) in the collection (31a, b).
9. Encoder (10) of claim 2 or 3 or any claim depending on any of same, configured to distinguish the one or more waveform parameter set, WPS, packets and the payload packets (25a-c) byencoding into each packet of the data stream (16) a packet type indicator (37) which indicates a packet type of the respective packet out of a plurality of packet types including a WPS packet type and one or more payload packet types.
10. Encoder (10) of claim 9, wherein the plurality of packet types further include an annotation channel packet type and an auxiliary metadata, AM, packet type and the one or more payload packet type comprise an independent frame packet type and a dependent frame packet type.11 . Encoder (10) of claim 10, configured toencode the packet type indicator (37) immediately followed by a predetermined bit sequence (41 ) into a packet of the data stream (16) which is of the AM packet type, wherein a bitstring of the packet type indicator (37) associated with the AM packet type immediately followed by the predetermined bit sequence (41) equals a predetermined ASCII-sequence which is indicative of a file format of the data stream (16).
12. Encoder (10) of claim 10 or 11 , configured toencode into a packet of the data stream (16) which is of the AM packet type general coding information (43) revealingan upper limit of the number of channels (21a-g) present in the data stream (16); an upper limit of a number of samples contained in each of the channels (21a-g) present in the data stream (16);an upper limit of a sampling rate of each of the channels (21a-g) present in the data stream (16);a waveform type the digital waveform data (14) relates to.
13. Encoder (10) of claim 9 or 10, configured toencode the packet type indicator (37) in manner so that the packet type indicator (37) indicates none of the packet types using an all zero bitstring.
14. Encoder (10) of any previous claim, configured to encode each payload packet (25a-c) without an stream pointer to an end of the respective payload packet (25a-c).
15. Encoder (10) of any previous claim, configured toencapsulate each packet of the data stream (16) using a predetermined file format to form the data stream (16), orconcatenate the packets to form the data stream (16), with separating the packets using start codes and start-code emulation prevention bytes.into encode each payload packet (25a-c) without an stream pointer to an end of the respective payload packet (25a-c).
16. Encoder (10) of claim 2 or 3 or any claim depending on any of same, configured to encode into a packet of the data stream (16) which is of a WPS packet type a sequence of channel group syntax portions which sequentially, along a channel group order defined among the selected channel groups (23a-c), indicate a number of channels (21a-g) contained by the one or more selected channel groups (23a-c) of the collection (31a, b).
17. Encoder (10) of claim 16, wherein each channel group syntax portion comprises a channel number syntax element indicating the number of channels (21a-g) contained in an associated channel group (23a-c),a repetition number syntax element indicating how many channel groups (23a-c) following the associated channel group (23a-c) along the channel index order coincide with the associated channel group (23a-c) in the number of channels (21a-g) contained, and a flag indicating whether the respective channel group syntax is the last in the sequence of channel group syntax portions.
18. Encoder (10) of claim 2 or 3 or any claim depending on any of same, configured to encode into a packet of the data stream (16) which is of a WPS packet type a channel reordering information indicating how the channels (21a-g) of the selected channel groups (23a-c) are to reordered for channel output.
19. Encoder (10) of claim 18, configured toencode the channel reordering information in form of a sequence of channel rank swaps, each indicating two channels (21a-g) whose channel indices are to be swapped so that the two channels (21a-g) swap in channel index order.
20. Encoder (10) of claim 2 or 3 or any claim depending on any of same, configured to encode into a packet of the data stream (16) which is of the WPS packet type information on how many annotation channels are comprised in the data stream (16) which accompany the collection (31a, b) of selected channel groups (23a-c).21 . Encoder (10) of any previous claim, wherein the payload packets (25a-c) are each of an independent frame packet type or a dependent frame packet type,wherein the encoder is configured toencode the one or more channels (21a-g) into the payload packets (25a-c) using block-wise coding,encode into a payload packet (25a-c) of the data stream (16) which is of the independent frame packet type a set of parameters comprising one or more ofa range of block sizes used for an adaptive setting of a block size in the block wise coding,a range of bit depths used in the block wise coding,a perceptual coding mode switch indicating a use or non-use of perceptual coding for the block wise coding,an inter-channel prediction mode switch indicating a use or non-use of inter-channel prediction for the block wise coding,in the encoding the one or more channels (21a-g) into the payload packets (25a-c), encode a payload packet (25a-c) of the data stream (16) which is of the independent frame packet type and is associated with a predetermined channel group (23a-c) and, if present, one or more following payload packets (25a-c) of the data stream (16) being of the dependent frame packet type, being associated with the predetermined channel group (23a-c) and having one or more frames (27a-c) encoded thereinto temporally immediately following a frame (27a-c) coded into the payload packet (25a-c) of the data stream (16) which is of the independent frame packet type using the set of parameters encoded into the payload packet (25a-c) of the data stream (16) which is of the independent frame packet type.
22. Encoder (10) of any previous claim, wherein the payload packets (25a-c) are each of an independent frame packet type or a dependent frame packet type,wherein the encoder is configured toin the encoding the one or more channels (21a-g) into the payload packets (25a-c), encode a payload packet (25a-c) of the data stream (16) which is of the independent frame packet type and is associated with a predetermined channel group (23a-c) independent from other payload packets (25a-c), and,if present, each of one or more following payload packets (25a-c) of the data stream (16) being of the dependent frame packet type, being associated with the predetermined channel group (23a-c) and having one or more frames (27a-c) encoded thereinto temporally immediately following a frame (27a-c) coded into thepayload packet (25a-c) of the data stream (16) which is of the independent frame packet type using coding dependencies from any preceding payload packet (25a-c) associated with the predetermined channel group (23a-c), preceding the respective payload packet (25a-c).
23. Decoder (12) for decoding from a data stream (16) digital waveform data (14), wherein channels (21a-g) of the digital waveform data (14) are coded into the data stream (16) in channel groups (23a-c), and the decoder (12) is configured tolocate within the data stream (16) payload packets (25a-c) into which the one or more channels (21a-g) are encoded channel-group wise and frame (27a-c) wise so that each packet is associated with a channel group (23a-c) and has exclusively encoded a frame (27a-c) of one or more channels (21 a-g) thereinto which are within the channel group (23a-c) with which the respective packet is associated, anddecode a collection (31a, b) of selected channel groups (23a-c) from payload packets (25a-c) associated with a channel group (23a-c) which is contained by the one or more selected channel groups (23a-c) of the collection (31a, b).
24. Decoder (12) of claim 23, wherein the decoder (12) is configured todecode from the data stream (16) a predetermined waveform parameter set, WPS, packet (33a, b) indicating the collection (31a, b) of one or more selected channel groups (23a-c) and comprising a WPS index,decode from each payload packet (25a-c)a WPS reference index which references the predetermined WPS packet (33a, b) via the WPS index of the predetermined WPS packet (33a, b), to derive a selected channel group (23a-c) with which the respective payload packet (25a-c) is associated and which is contained by the one or more selected channel groups (23a- c) of the collection (31a, b) indicated by the predetermined WPS packet (33a, b), with decoding from the respective payload packet (25a-c), if a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the predeterminedWPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associated.
25. Decoder (12) of claim 23, wherein the decoder (12) is configured todecode from the data stream (16) a predetermined waveform parameter set, WPS, packet indicating the collection (31a, b) of one or more selected channel groups (23a-c),decode from each payload packet (25a-c)if a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the predetermined WPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associated.
26. Decoder (12) of claim 23, 24, or 25, wherein, within each of the one or more selected channel groups (23a-c), if containing more than one channel (21a-g), all channels of the respective selected channel group (23a-c) coincide in the sampling rate27. Decoder (12) of claim 24 or any claim depending thereon, configured to, in decoding from each payload packet (25a-c), if a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associatedcheck whether a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, and,if the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, decode the channel group index from the respective payload packet (25a-c), and,if the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is not larger than one, infer that the respective payload packet (25a-c) has the one selected channel group (23a-c) encoded thereinto.
28. Decoder (12) of claim 24 or any claim depending thereon, configured to, in decoding from each payload packet (25a-c), if a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associatedif the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the corresponding WPS packet (33a, b) is larger than one, decode the channel group index into the respective payload packet (25a-c) using a fixed length code whose code length monotonically increases with a number of selected channel groups (23a-c) in the collection (31a, b).
29. Decoder (12) of claim 25 or any claim depending thereon, configured to, in decoding from each payload packet (25a-c), if a number of selected channel groups (23a-c) of the collection (31 a, b) indicated by the WPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associatedcheck whether a number of selected channel groups (23a-c) of the collection (31a, b) indicated by the WPS packet (33a, b) is larger than one, and,if the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the WPS packet (33a, b) is larger than one, decode the channel group index from the respective payload packet (25a-c), and,if the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the WPS packet (33a, b) is not larger than one, infer that the respective payload packet (25a-c) has the one selected channel group (23a-c) encoded thereinto.
30. Decoder (12) of claim 25 or any claim depending thereon, configured to, in decoding from each payload packet (25a-c), if a number of selected channel groups (23a-c) of the collection (31 a, b) indicated by the WPS packet (33a, b) is larger than one, a channel group index indicting the channel group (23a-c) with which the respective payload packet (25a-c) is associatedif the number of selected channel groups (23a-c) of the collection (31a, b) indicated by the WPS packet (33a, b) is larger than one, decode the channel group index into therespective payload packet (25a-c) using a fixed length code whose code length monotonically increases with a number of selected channel groups (23a-c) in the collection (31a, b).31 . Decoder (12) of claim 24 or 25 any claim depending on any of same, configured to distinguish the predetermined waveform parameter set, WPS, packet and the payload packets (25a-c) bydecoding from each packet of the data stream (16) a packet type indicator (37) which indicates a packet type of the respective packet out of a plurality of packet types including a WPS packet type and one or more payload packet types.
32. Decoder (12) of claim 31 , wherein the plurality of packet types further include an annotation channel packet type and an auxiliary metadata, AM, packet type and the one or more payload packet type comprise an independent frame packet type and a dependent frame packet type.
33. Decoder (12) of claim 32, configured todecode the packet type indicator (37) immediately followed by a predetermined bit sequence (41 ) from a packet of the data stream (16) which is of the AM packet type, determine from a bitstring composed of the packet type indicator (37) associated with the AM packet type immediately followed by the predetermined bit sequence (41) an ASCII-sequence, and identify a file format of the data stream (16) using the ASCII sequence.
34. Decoder (12) of claim 32 or 33, configured todecode from a packet of the data stream (16) which is of the AM packet type general coding information (43) revealing on one or more ofan upper limit of the number of channels (21a-g) present in the data stream (16); an upper limit of a number of samples contained in each of the channels (21a-g) present in the data stream (16);an upper limit of a sampling rate of each of the channels (21a-g) present in the data stream (16);a waveform type the digital waveform data (14) relates to.
35. Decoder (12) of claim 31 or 32, wherein the packet type indicator (37) is encoded into the data stream (16) in manner so that the packet type indicator (37) indicates none of the packet types using an all zero bitstring.
36. Decoder (12) of any of the previous claims 23 to 35, configuredto determine an end of each payload packet (25a-c) by parsing the respective payload packet (25a-c) till the end.
37. Decoder (12) of any of the previous claims 23 to 36, whereineach packet of the data stream (16) is encapsulated using a predetermined file format and the decoder (12) is configured to decapsulate each packet of the data stream (16), orthe packets are concatenatedly written into the data stream (16) and the decoder (12) is configured to derive the packets from the data stream (16) by sequentially parsing the data stream (16) and detecting the beginning of a new packet using a start code, , and by performing start-code-emulation-prevention-byte-removal.
38. Decoder (12) of claim 24 or 25 or any claim depending on any of same, configured todecode from a packet of the data stream (16) which is of a WPS packet type a sequence of channel group syntax portions which sequentially, along a channel index order defined among the plurality of channels (21a-g), indicate a number of channels (21a-g) contained by the one or more selected channel groups (23a-c) of the collection (31a, b).
39. Decoder (12) of claim 38, wherein each channel group syntax portion comprises a channel number syntax element indicating the number of channels (21a-g) contained in an associated channel group (23a-c),a repetition number syntax element indicating how many channel groups (23a-c) following the associated channel group (23a-c) along the channel index order coincide with the associated channel group (23a-c) in the number of channels (21a-g) contained, and a flag indicating whether the respective channel group syntax is the last in the sequence of channel group syntax portions.
40. Decoder (12) of claim 24 or 25 or any claim depending on any of same, configured todecode from a packet of the data stream (16) which is of the WPS packet type a channel reordering information indicating how the channels (21 a-g) of the selected channel groups (23a-c) are to reordered for channel output.41 . Decoder (12) of claim 40, configured todecode the channel reordering information in form of a sequence of channel rank swaps, each indicating two channels (21a-g) whose channel indices are to be swapped so that the two channels (21a-g) swap in channel index order.
42. decoder (12) of claim 24 or 25 or any claim depending on any of same, configured to decode from a packet of the data stream (16) which is of the WPS packet type an information on how many annotation channels are comprised in the data stream (16) which accompany the collection (31a, b) of selected channel groups (23a-c).
43. Decoder (12) of any of the previous claims 23 to 42, wherein the payload packet types are each of an independent frame packet type or a dependent frame packet type, wherein the decoder (12) is configured todecode the one or more channels (21a-g) from the payload packets (25a-c) using block-wise decoding,decode from a payload packet (25a-c) of the data stream (16) which is of the independent frame packet type a set of parameters comprising one or more ofa range of block sizes used for an adaptive setting of a block size in the block wise decoding,a range of bit depths used in the block wise decoding,a perceptual coding mode switch indicating a use or non-use of perceptual coding for the block wise decoding,an inter-channel prediction mode switch indicating a use or non-use of interchannel prediction for the block wise decoding,in the decoding the one or more channels (21 a-g) from the payload packets (25a-c), decode a payload packet (25a-c) of the data stream (16) which is of the independent frame packet type and is associated with a predetermined channel group (23a-c) and, if present, one or more following payload packets (25a-c) of the data stream (16) being of the dependent frame packet type, being associated with the predetermined channel group (23a-c) and having one or more frames (27a-c) encoded thereinto temporally immediately following a frame (27a-c) coded into the payload packet (25a-c) of the data stream (16) which is of the independent frame packet type using the set of parameters encoded into the payload packet (25a-c) of the data stream (16) which is of the independent frame packet type.
44. Decoder (12) of any of the previous claims 23 to 43, wherein the payload packets (25a-c) are each of an independent frame packet type or a dependent frame packet type,wherein the decoder (12) is configured toin the decoding the one or more channels (21a-g) from the payload packets (25a-c), encode a payload packet (25a-c) of the data stream (16) which is of the independent frame packet type and is associated with a predetermined channel group (23a-c) independent from other payload packets (25a-c), and,if present, each of one or more following payload packets (25a-c) of the data stream (16) being of the dependent frame packet type, being associated with the predetermined channel group (23a-c) and having one or more frames (27a-c) encoded thereinto temporally immediately following a frame (27a-c) coded into the payload packet (25a-c) of the data stream (16) which is of the independent frame packet type using coding dependencies from any preceding payload packet (25a-c) associated with the predetermined channel group (23a-c), preceding the respective payload packet (25a-c).
45. Method (100) for encoding into a data stream (16) digital waveform data (14) comprising a one or more channels (21a-g), the method comprising:grouping (102) the one or more channels (21 a-g) into one or more channel groups (23a-c),Encoding (104) the one or more channels (21 a-g) into payload packets (25a-c) channel-group wise and frame (27a-c) wise so that each packet is associated with a channel group (23a-c) and has exclusively encoded a frame (27a-c) of one or more channels (21 a-g) thereinto which are within the channel group (23a-c) with which the respective packet is associated.
46. Method (110) for decoding from a data stream (16) digital waveform data (14), wherein channels (21 a-g) of the digital waveform data (14) are coded into the data stream (16) in channel groups (23a-c), and wherein the method comprises:locating (112) within the data stream (16) payload packets (25a-c) into which the one or more channels (21 a-g) are encoded channel-group wise and frame (27a-c) wise so that each packet is associated with a channel group (23a-c) and has exclusively encoded a frame (27a-c) of one or more channels (21 a-g) thereinto which are within the channel group (23a-c) with which the respective packet is associated, anddecoding (114) a collection (31a, b) of selected channel groups (23a-c) from payload packets (25a-c) associated with a channel group (23a-c) which is contained by the one or more selected channel groups (23a-c) of the collection (31a, b).
47. Computer program for performing the method according to claim 45 or 46, when the computer program runs on a computer.
48. Data stream (16) having encoded therein a digital waveform data (14) using the method of 45.