Signalling features in a bitstream

The bitstream with a tone mapping parameter in the decoder configuration addresses bandwidth and decoding capability challenges, enabling flexible and high-quality video delivery across diverse devices by controlling tone mapping operations for colour space conversion.

GB2638653APending Publication Date: 2025-09-03V NOVA INT LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
GB2023018015
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-09-03

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

There is provided a bitstream, 600, for transmitting one or more enhancement residuals planes suitable to be combined with a set of preliminary pictures obtained from a decoder reconstructed video. The bitstream comprises a decoder configuration, 630, for controlling a decoding process of the bitstream, the decoder configuration comprising a tone mapping parameter, 632-635. The bitstream may be in accordance with the MPEG5 Part 2 LCEVC (Low Complexity Enhancement Video Coding) standard and / or ISO / IEC 23094-2. Also disclosed are associated methods of generating and decoding the claimed bitstream and a signal carrying the claimed bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field The invention relates to signalling features in a bitstream comprising a video signal. In particular, but not exclusively, the bitstream is in accordance with the MPEG5 Part 2 LCEVC standard. The invention is implementable in hardware or software. Background For video broadcasters and streamers, the delivery of high-quality video experiences to end users is very important. However, broadcasters and streamers face significant challenges in managing limitations such as bandwidth constraints and varied decoding capabilities across different distribution channels and end user devices. This problem becomes pronounced with technical innovations in the field of video coding. Hierarchical coding schemes for example, including that introduced by the LCEVC standard, have mitigated the need to simulcast multiple streams with diverse characteristics to suit different bandwidths and decoding capabilities. Yet, there is a necessity to further improve coding schemes. A critical aspect that needs addressing is the colour conversion operation to allow video to be delivered in wide colour gamut formats and / or high dynamic ranges. Colour conversion operations typically involve tone mapping procedures, which transition the signal from one colour space to another. In current coding schemes, there is a need for improved tone mapping procedures to aid in the precise rendition of video data in a plurality of colour spaces across various end user devices potentially with different decoding capabilities. Therefore, this disclosure aims at providing an improved tone mapping procedure for coding schemes. Summary According to a first aspect of the invention, there is provided a bitstream for transmitting one or more enhancement residuals planes suitable to be combined with a set of preliminary pictures obtained from a decoder reconstructed video. The bitstream comprising a decoder configuration for controlling a decoding process of the bitstream. The decoder configuration comprises a tone mapping parameter. In this way, there is provided an improved approach to video broadcasting by enabling control over the tone mapping operations that are used during the colour conversion process in coding schemes. By including configuration settings comprising a tone mapping parameter in a bitstream, the invention provides a bitstream signalling control of tone mapping operations to a decoder and thus enables higher quality videos to be produced with more precise rendition of video data in a plurality of colour spaces across various end user devices potentially with different decoding capabilities. By signalling the tone mapping parameter in the bitstream, it is possible to send a single-cast stream that can be decoded by legacy decoders into a base signal, such as for example an SDR signal in a BT.709 colour space at a bit depth of 8 bits, and also decoded by a subset of users with an enhancement decoder that can use the tone mapping parameter to obtain an improved signal, such as for example an HDR signal in a BT.2020 colour space at a bit depth of 10 bits. Other colour spaces and other bit depths could be achieved using the principle disclosed. Therefore, simulcasting two separate streams each at different quality levels (e.g. colour space, bit depth) can be avoided, while at the same time retaining the retro-compatibility of the bitstream for use with legacy decoders that cannot understand the tone mapping parameter. Preferably, the tone mapping parameter indicates whether a decoder should apply the one or more enhancement residuals planes in a first colour space or a second colour space. In this way, added flexibility is achieved in how to convert to each colour space in an encoding pipeline and a corresponding decoding pipeline. In one scenario the enhancement residuals planes will be generated at the encoder in for example a BT.709 colour space and be used at a decoder on a signal in the same colour space to correct that signal. In another scenario the enhancement residuals planes will be generated at the encoder in for example a BT.2020 colour space and be used at a decoder on a signal in the same colour space to correct that signal. The tone mapping parameter enables a decoder to decode the bitstream appropriately. Preferably, the first colour space is a colour space of the decoder reconstructed video. Preferably, the second colour space is a colour space of a preliminary picture of the set of preliminary pictures. Preferably, a first value of the tone mapping parameter indicates to a decoder to apply a tone mapping operation on a combined plane obtained by combining an enhancement residuals plane of the one or more enhancement residuals planes with a preliminary picture of the set of preliminary pictures. Preferably, a second value of the tone mapping parameter indicates to a decoder to apply a tone mapping operation on a preliminary picture of the set of preliminary pictures, before combining with an enhancement residuals plane of the one or more enhancement residuals planes. Preferably, the tone mapping parameter indicates a tone mapping algorithm that should be used during a decoding / processing of the bitstream. Preferably, a first value of the tone mapping parameter indicates a decoder should use a tone mapping algorithm associated with a first entry of a table. Preferably, a second or other value the tone mapping parameter indicates a decoder should use a tone mapping algorithm associated with a second or other entry of a table. Preferably, a further value of the tone mapping parameter indicates a decoder should use a further tone mapping parameter to determine a tone mapping algorithm to be used during a decoding / processing of the bitstream. Preferably, the bitstream comprises the further tone mapping parameter. Preferably, the further tone mapping parameter specifies a tone mapper algorithm to be used for tone mapper operation. Preferably, a first or other value of the further tone mapping parameter indicates a decoder should use a tone mapping algorithm associated with a first or other entry of a table. Preferably, the tone mapping parameter indicates whether additional data relating to a tone mapping operation is present in the bitstream. Preferably, a first value of the tone mapping parameter indicates that the additional data relating to a tone mapping operation is present in the bitstream and a second value of the tone mapping parameter indicates that the additional data relating to a tone mapping operation is not present in the bitstream. Preferably, the bitstream comprises a further tone mapping parameter specifying a size of the additional data. Preferably, the bitstream comprises a further tone mapping parameter comprising the additional data. Preferably, the tone mapping parameter specifies a size of an additional data relating to a tone mapping operation. Preferably, the bitstream comprises additional data relating to a tone mapping operation. Preferably, the tone mapping parameter indicates a change in bit depth in a signal associated with the tone mapping operation. Preferably, the tone mapping parameter indicates a value representing the difference in bit depth from a first bit depth to a second bit depth. Preferably, the tone mapping parameter indicates a bit depth of decoded base layer pictures. Preferably, the tone mapping parameter indicates a bit depth of decoded enhancement layer pictures. Preferably, the bitstream indicates that the tone mapping parameter or parameters are present in the bitstream using a bitstream version. Preferably, the tone mapping parameter is located in a payload of the bitstream. Preferably, the bitstream version is in compliance with standard Itu_t_35. Preferably, the tone mapping parameter is located in an additional_info_type field of the bitstream. Preferably, the additionaljnfojype field is in compliance with standard ISO / IEC 23094-2. Preferably, the bitstream is in accordance with MPEG5 Part 2 LCEVC standard and / or ISO / IEC 23094-2. Preferably, the tone mapping parameter comprises configuration information usable by a decoder to deploy a tone mapping operation. Brief Description of the Drawings The invention shall now be described, by way of example only, with reference to the accompanying drawings in which: FIG. 1 shows a high-level schematic of a hierarchical encoding and decoding process; FIG. 2 shows a high-level schematic of an encoding process of a hierarchical coding technology; FIG. 3 shows a high-level schematic of a decoding process suitable for decoding the output of FIG. 2 FIG. 4 shows a high-level schematic of a modification showing a first variation of an encoding process of a hierarchical coding technology shown in FIG. 2; FIG. 5 shows a high-level schematic of a second variation of the encoding process of FIG. 4; FIG. 6 illustrates an example bitstream for signalling features of the first and second variations. FIG. 7 shows a high-level schematic of a modification showing a first variation of a decoding process; and FIG. 8 shows a high-level schematic of a second variation of the decoding process of FIG. 7. Detailed Description FIG. 1 shows a high-level schematic of a hierarchical encoding and decoding process. Data 101 to be encoded is retrieved by a hierarchical encoder 102 which outputs encoded data 103. Subsequently, the encoded data 103 is received by a hierarchical decoder 104 which decodes the data and outputs decoded data 105. Typically, the hierarchical coding schemes used in examples herein create a base or core level, which is a representation of the original data at a lower level of quality and one or more levels of residuals which can be used to recreate the original data at a higher level of quality using a decoded version of the base level data. In general, the term "residuals" as used herein refers to a difference between a value of a reference array or reference frame and an actual array or frame of data. The array may be a one or two-dimensional array that represents a coding unit. For example, a coding unit may be a 2x2 or 4x4 set of residual values that correspond to similar sized areas of an input video frame. It should be noted that the generalised examples are agnostic as to the nature of the input signal. Reference to "residual data" as used herein refers to data derived from a set of residuals, e.g. a set of residuals themselves or an output of a set of data processing operations that are performed on the set of residuals. Throughout the present description, generally a set of residuals includes a plurality of residuals or residual elements, each residual or residual element corresponding to a signal element, that is, an element of the signal or original data. In specific examples, the data may be an image or video. In these examples, the set of residuals corresponds to an image or frame of the video, with each residual being associated with a pixel of the signal, the pixel being the signal element. The methods described herein may be applied to so-called planes of data that reflect different colour components of a video signal. For example, the methods may be applied to different planes of YUV or RGB data reflecting different colour channels. Different colour channels may be processed in parallel. The components of each stream may be collated in any logical order. A further hierarchical coding technology with which the principles of the present invention may be utilised is illustrated in FIGS. 2 and 3. This technology is a flexible, adaptable, highly efficient and computationally inexpensive coding format which combines a different video coding format, a base codec, (e.g., AVC, HEVC, or any other present or future codec) with at least two enhancement levels of coded data. The general structure of the encoding scheme uses a down-sampled source signal encoded with a base codec, adds a first level of correction data to the decoded output of the base codec to generate a corrected picture, and then adds a further level of enhancement data to an up-sampled version of the corrected picture. Thus, the streams are considered to be a base stream and an enhancement stream, which may be further multiplexed or otherwise combined to generate an encoded data stream. References to an encoded data as described herein may refer to the enhancement stream or a combination of the base stream and the enhancement stream. The base stream may be decoded by a hardware decoder while the enhancement stream may be suitable for software processing implementation with suitable power consumption. This general encoding structure creates a plurality of degrees of freedom that allow great flexibility and adaptability to many situations, thus making the coding format suitable for many use cases including OTT transmission, live streaming, live ultra-high-definition UHD broadcast, and so on. Although the decoded output of the base codec is not intended for viewing, it is a fully decoded video at a lower resolution, making the output compatible with existing decoders and, where considered suitable, also usable as a lower resolution output. Returning to the initial process described above, where a base stream is provided along with two levels (or sub-levels) of enhancement within an enhancement stream, an example of a generalised encoding process is depicted in the block diagram of FIG. 2. An input video 200 at an initial resolution is processed to generate various encoded streams 201, 202, 203 to form bitstream 230. A first encoded stream 201 (encoded base stream) is produced by feeding a base codec (e.g., AVC, HEVC, or any other codec) with a down-sampled version of the input video 200. The encoded base stream may be referred to as the base layer or base level. A second encoded stream 303(encoded level 1 stream) is produced by processing the residuals obtained by taking the difference between a reconstructed base codec video and the down-sampled version of the input video. A third encoded stream 203 (encoded level 2 stream) is produced by processing the residuals obtained by taking the difference between an up-sampled version of a corrected version of the reconstructed base coded video and the input video. In certain cases, the components of FIG. 2 may provide a general low complexity encoder. In certain cases, the enhancement streams may be generated by encoding processes that form part of the low complexity encoder and the low complexity encoder may be configured to control an independent base encoder and decoder (e.g., as packaged as a base codec). In other cases, the base encoder and decoder may be supplied as part of the low complexity encoder. In one case, the low complexity encoder of FIG. 2 may be seen as a form of wrapper for the base codec, where the functionality of the base codec may be hidden from an entity implementing the low complexity encoder. A down-sampling operation illustrated by down-sampling component 205 may be applied to the input video to produce a down-sampled video to be encoded by a base encoder 213 of a base codec. The down-sampling can be done either in both vertical and horizontal directions, or alternatively only in the horizontal direction. The base encoder 213 and a base decoder 214 may be implemented by a base codec (e.g., as different functions of a common codec). The base codec, and / or one or more of the base encoder 213 and the base decoder 214 may comprise suitably configured electronic circuitry (e.g., a hardware encoder / decoder) and / or computer program code that is executed by a processor. Each enhancement stream encoding process may not necessarily include an upsampling step. In FIG. 2 for example, the first enhancement stream is conceptually a correction stream while the second enhancement stream is upsampled to provide a level of enhancement. Looking at the process of generating the enhancement streams in more detail, to generate the encoded level 1 stream, the encoded base stream is decoded by the base decoder 214 (i.e. a decoding operation is applied to the encoded base stream to generate a decoded base stream). Decoding may be performed by a decoding function or mode of a base codec. The difference between the decoded base stream and the down-sampled input video is then created at a level 1 comparator 210 (i.e. a subtraction operation is applied to the down-sampled input video and the decoded base stream to generate a first set of residuals). The output of the comparator 210 may be referred to as a first set of residuals, e.g. a surface or frame of residual data, where a residual value is determined for each picture element at the resolution of the base encoder 213, the base decoder 214 and the output of the down-sampling block 205. The difference is then encoded by a first encoder 215 (i.e. a level 1 encoder) to generate the encoded level 1 stream 202 (i.e. an encoding operation is applied to the first set of residuals to generate a first enhancement stream). As noted above, the enhancement stream may comprise a first level of enhancement 202 and a second level of enhancement 203. The first level of enhancement 202 may be considered to be a corrected stream, e.g. a stream that provides a level of correction to the base encoded / decoded video signal at a lower resolution than the input video 200. The second level of enhancement 203 may be considered to be a further level of enhancement that converts the corrected stream to the original input video 200, e.g. that applies a level of enhancement or correction to a signal that is reconstructed from the corrected stream. In the example of FIG. 2, the second level of enhancement 203 is created by encoding a further set of residuals. The further set of residuals are generated by a level 2 comparator 219. The level 2 comparator 219 determines a difference between an upsampled version of a decoded level 1 stream, e.g. the output of an upsampling component 217, and the input video 200. The input to the up-sampling component 217 is generated by applying a first decoder (i.e. a level 1 decoder 218) to the output of the first encoder 215. This generates a decoded set of level 1 residuals. These are then combined with the output of the base decoder 214 at summation component 220. This effectively applies the level 1 residuals to the output of the base decoder 214. It allows for losses in the level 1 encoding and decoding process to be corrected by the level 2 residuals. The output of summation component 220 may be seen as a simulated signal that represents an output of applying level 1 processing to the encoded base stream 201 and the encoded level 1 stream 202 at a decoder. As noted, an upsampled stream is compared to the input video which creates a further set of residuals (i.e. a difference operation is applied to the upsampled re-created stream to generate a further set of residuals). The further set of residuals are then encoded by a second encoder 221 (i.e. a level 2 encoder) as the encoded level 2 enhancement stream (i.e. an encoding operation is then applied to the further set of residuals to generate an encoded further enhancement stream). Thus, as illustrated in FIG. 2 and described above, the output of the encoding process is a base stream 201 and one or more enhancement streams 202, 203 which preferably comprise a first level of enhancement and a further level of enhancement. It should be noted that the components shown in FIG. 2 may operate on blocks or coding units of data, e.g. corresponding to 2x2 or 4x4 portions of a frame at a particular level of resolution. The components operate without any inter-block dependencies, hence they may be applied in parallel to multiple blocks or coding units within a frame. This differs from comparative video encoding schemes wherein there are dependencies between blocks (e.g., either spatial dependencies or temporal dependencies). The dependencies of comparative video encoding schemes limit the level of parallelism and require a much higher complexity. A corresponding generalised decoding process is depicted in the block diagram of FIG. 3. FIG. 3 may be said to show a low complexity decoder that corresponds to the low complexity encoder of FIG. 6. The low complexity decoder receives the three streams 201, 202, 203 generated by the low complexity encoder together with headers 304 containing further decoding information as part of a bitstream 330. The encoded base stream 301 is decoded by a base decoder 310 corresponding to the base codec used in the low complexity encoder. The encoded level 1 stream 302 is received by a first decoder 311 (i.e. a level 1 decoder), which decodes a first set of residuals as encoded by the first encoder 215 of Figure 1. At a first summation component 312, the output of the base decoder 310 is combined with the decoded residuals obtained from the first decoder 311. The combined video, which may be said to be a level 1 reconstructed video signal, is upsampled by upsampling component 313. The encoded level 2 stream 303 is received by a second decoder 314 (i.e. a level 2 decoder). The second decoder 314 decodes a second set of residuals as encoded by the second encoder 221 of FIG. 2. Although the headers 304 are shown in FIG. 3 as being used by the second decoder 314, they may also be used by the first decoder 311 as well as the base decoder 310. The output of the second decoder 314 is a second set of decoded residuals. These may be at a higher resolution to the first set of residuals and the input to the upsampling component 313. At a second summation component 315, the second set of residuals from the second decoder 314 are combined with the output of the up-sampling component 313, i.e. an up-sampled reconstructed level 1 signal, to reconstruct decoded video 350. As per the low complexity encoder, the low complexity decoder of FIG. 3 may operate in parallel on different blocks or coding units of a given frame of the video signal. Additionally, decoding by two or more of the base decoder 310, the first decoder 311 and the second decoder 314 may be performed in parallel. This is possible as there are no inter-block dependencies. In the decoding process, the decoder may parse the headers 304 (which may contain global configuration information, picture or frame configuration information, and data block configuration information) and configure the low complexity decoder based on those headers. In order to re-create the input video, the low complexity decoder may decode each of the base stream, the first enhancement stream and the further or second enhancement stream. The frames of the stream may be synchronised and then combined to derive the decoded video 350. The decoded video 350 may be a lossy or lossless reconstruction of the original input video 100 depending on the configuration of the low complexity encoder and decoder. In many cases, the decoded video 350 may be a lossy reconstruction of the original input video 200 where the losses have a reduced or minimal effect on the perception of the decoded video 350. In each of FIGS. 2 and 3, the level 2 and level 1 encoding operations may include the steps of transformation, quantization and entropy encoding (e.g., in that order). The encoding operations may also include residual ranking, weighting and filtering. Similarly, at the decoding stage, the residuals may be passed through an entropy decoder, a dequantizer and an inverse transform module (e.g., in that order). Any suitable encoding and corresponding decoding operation may be used. Preferably however, the level 2 and level 1 encoding steps may be performed in software (e.g., as executed by one or more central or graphical processing units in an encoding device). The transform as mentioned herein may use a directional decomposition transform such as a Hadamard-based transform. Both may comprise a small kernel or matrix that is applied to flattened coding units of residuals (i.e. 2x2 or 4x4 blocks of residuals). More details on the transform can be found for example in patent WO 2013 / 171173 Alor WO 2018 / 046941 Al, which are incorporated herein by reference. The encoder may select between different transforms to be used, for example between a size of kernel to be applied. The transform may transform the residual information to four surfaces. For example, the transform may produce the following components or transformed coefficients: average, vertical, horizontal and diagonal. A particular surface may comprise all the values for a particular component, e.g. a first surface may comprise all the average values, a second all the vertical values and so on. As alluded to earlier in this disclosure, these components that are output by the transform may be taken in such embodiments as the coefficients to be quantized in accordance with the described methods. A quantization scheme may be useful to create the residual signals into quanta, so that certain variables can assume only certain discrete magnitudes. Entropy encoding in this example may comprise run length encoding (RLE), then processing the encoded output is processed using a Huffman encoder. In certain cases, only one of these schemes may be used when entropy encoding is desirable. In summary, the methods and apparatuses herein are based on an overall approach which is built over an existing encoding and / or decoding algorithm (such as MPEG standards such as AVC / H.264, HEVC / H.265, etc. as well as non-standard algorithm such as VP9, AVI, and others) which works as a baseline for an enhancement layer which works accordingly to a different encoding and / or decoding approach. The idea behind the overall approach of the examples is to hierarchically encode / decode the video frame as opposed to the use block-based approaches as used in the MPEG family of algorithms. Hierarchically encoding a frame includes generating residuals for the full frame, and then a decimated frame and so on. As indicated above, the processes may be applied in parallel to coding units or blocks of a colour component of a frame as there are no inter-block dependencies. The encoding of each colour component within a set of colour components may also be performed in parallel (e.g., such that the operations are duplicated according to (number of frames) * (number of colour components) * (number of coding units per frame)). It should also be noted that different colour components may have a different number of coding units per frame, e.g. a luma (e.g., Y) component may be processed at a higher resolution than a set of chroma (e.g., U or V) components as human vision may detect lightness changes more than colour changes. Thus, as illustrated and described above, the output of the decoding process is an (optional) base reconstruction, and an original signal reconstruction at a higher level. This example is particularly well-suited to creating encoded and decoded video at different frame resolutions. For example, the input signal 30 may be an HD video signal comprising frames at 1920 x 1080 resolution. In certain cases, the base reconstruction and the level 2 reconstruction may both be used by a display device. For example, in cases of network traffic, the level 2 stream may be disrupted more than the level 1 and base streams (as it may contain up to 4x the amount of data where down-sampling reduces the dimensionality in each direction by 2). In this case, when traffic occurs the display device may revert to displaying the base reconstruction while the level 2 stream is disrupted (e.g., while a level 2 reconstruction is unavailable), and then return to displaying the level 2 reconstruction when network conditions improve. A similar approach may be applied when a decoding device suffers from resource constraints, e.g. a set-top box performing a systems update may have an operation base decoder 220 to output the base reconstruction but may not have processing capacity to compute the level 2 reconstruction. The encoding arrangement also enables video distributors to distribute video to a set of heterogeneous devices; those with just a base decoder 320 view the base reconstruction, whereas those with the enhancement level may view a higher-quality level 2 reconstruction. In comparative cases, two full video streams at separate resolutions were required to service both sets of devices. As the level 2 and level 1 enhancement streams encode residual data, the level 2 and level 1 enhancement streams may be more efficiently encoded, e.g. distributions of residual data typically have much of their mass around 0 (i.e. where there is no difference) and typically take on a small range of values about 0. This may be particularly the case following quantization. In contrast, full video streams at different resolutions will have different distributions with a non-zero mean or median that require a higher bit rate for transmission to the decoder. In the examples described herein residuals are encoded by an encoding pipeline. This may include transformation, quantization and entropy encoding operations. It may also include residual ranking, weighting and filtering. Residuals are then transmitted to a decoder, e.g. as L-l and L-2 enhancement streams, which may be combined with a base stream as a hybrid stream (or transmitted separately). In one case, a bit rate is set for a hybrid data stream that comprises the base stream and both enhancements streams, and then different adaptive bit rates are applied to the individual streams based on the data being processed to meet the set bit rate (e.g., high-quality video that is perceived with low levels of artefacts may be constructed by adaptively assigning a bit rate to different individual streams, even at a frame by frame level, such that constrained data may be used by the most perceptually influential individual streams, which may change as the image data changes). The sets of residuals as described herein may be seen as sparse data, e.g. in many cases there is no difference for a given pixel or area and the resultant residual value is zero. When looking at the distribution of residuals much of the probability mass is allocated to small residual values located near zero - e.g. for certain videos values of -2, -1, 0, 1, 2 etc. occur the most frequently. In certain cases, the distribution of residual values is symmetric or near symmetric about 0. In certain test video cases, the distribution of residual values was found to take a shape similar to logarithmic or exponential distributions (e.g., symmetrically or near symmetrically) about 0. The exact distribution of residual values may depend on the content of the input video stream. Residuals may be treated as a two-dimensional image in themselves, e.g. a delta image of differences. Seen in this manner the sparsity of the data may be seen to relate features like "dots", small "lines", "edges", "corners", etc. that are visible in the residual images. It has been found that these features are typically not fully correlated (e.g., in space and / or in time). They have characteristics that differ from the characteristics of the image data they are derived from (e.g., pixel characteristics of the original video signal). As the characteristics of residuals differ from the characteristics of the image data they are derived from it is generally not possible to apply standard encoding approaches, e.g. such as those found in traditional Moving Picture Experts Group (MPEG) encoding and decoding standards. For example, many comparative schemes use large transforms (e.g., transforms of large areas of pixels in a normal video frame). Due to the characteristics of residuals, e.g. as described above, it would be very inefficient to use these comparative large transforms on residual images. For example, it would be very hard to encode a small dot in a residual image using a large block designed for an area of a normal image. Certain examples described herein address these issues by instead using small and simple transform kernels (e.g., 2x2 or 4x4 kernels - the Directional Decomposition and the Directional Decomposition Squared - as presented herein). The transform described herein may be applied using a Hadamard matrix (e.g., a 4x4 matrix for a flattened 2x2 coding block or a 16x16 matrix for a flattened 4x4 coding block). This moves in a different direction from comparative video encoding approaches. Applying these new approaches to blocks of residuals generates compression efficiency. For example, certain transforms generate uncorrelated transformed coefficients (e.g., in space) that may be efficiently compressed. While correlations between transformed coefficients may be exploited, e.g. for lines in residual images, these can lead to encoding complexity, which is difficult to implement on legacy and low-resource devices, and often generates other complex artefacts that need to be corrected. Pre-processing residuals by setting certain residual values to 0 (i.e. not forwarding these for processing) may provide a controllable and flexible way to manage bitrates and stream bandwidths, as well as resource use. In FIGS. 4 and 5, we present high-level schematics illustrating modifications to the encoding processes implemented by the exemplary hierarchical coding technology initially introduced in FIG. 2. The encoding processes are in accordance with MPEG-5 Part 2 Low Complexity Enhancement Video Coding (LCEVC) standard delineated in ISO / IEC 23094-2:2021(en). Notwithstanding, the disclosed techniques retain applicability across various other hierarchical or non-hierarchical coding technologies. To facilitate comprehension, like features between FIGS. 2, 4 and 5 are denoted using consistent reference signs. Herein, we emphasise the distinctions pertaining to tone mapping in the encoding processes depicted in FIGS. 4 and 5. FIG. 4 shows a high-level schematic of a modified encoding process of the hierarchical coding technology shown in FIG. 2 according to a first variation. In FIG. 4, the modified encoding process includes a tone mapping operation 404 which is performed on the input video 200 before the input video 200 branches to go to down sampling operation 205 and the Level 2 comparator 219. The tone mapping operation 404 changes the input video 200 from a first colour space (for example BT.2020) to a second colour space (for example BT.709). The modified encoding process of FIG. 4 uses the tone mapping operation 404 to change the colour space before the encoding scheme of FIG.2 operates. FIG. 5 shows a high-level schematic of a second variation of the encoding process of FIG. 4. In FIG. 5, the modified encoding process includes a tone mapping operation 508 which is performed on the output of the down sampling operation 205 after the video signal branches to comparator 210. In this way, the tone mapping operation 506 changes the down sampled input video from a first colour space (e.g. BT.2020) to a second colour space (e.g. BT.709). The base encoder will now operate on the video signal in the second colour space (and also on the down sampled video signal). However, the encoded level 1 stream 202 and the encoded level 2 stream 203 are generated in the first colour space (e.g. BT.2020) and so the enhancements available to a decoder are in the second colour space. The enhancements may be known as enhancement residuals planes and are applied to preliminary pictures in the decoder. The modified encoding process of FIG. 5 also includes an inverse tone map operation 507 or pseudo inverse tone mapping operation located between the output of based decoder 214 and comparator 210 before the signal branches also to comparator 220. The inverse tone map changes the video signal output from the base decoder 214 from the second colour space (e.g. BT.709) to the first colour space (e.g. BT.2020) so that the encoded level 1 and 2 streams 202, 203 can be generated. The modified encoding processes of FIG. 4 and FIG. 5 each generate and send information related to the tone mapping operation or operations used, such as tone mapping parameter(s) as part of a modified bitstream 430. The modified bitstream 430 comprises in addition to the information of bitstream 230 the tone mapping parameter(s), or information related to the tone mapping operation. The tone mapping parameter(s) can then be used by a decoder to perform an inverse tone mapping operation or pseudoinverse tone mapping operation or other tone mapping operation that is useful to picture reconstruction at the decoder to the decoded video signal 200 to change the colour space of the decoded video signal 200 at the decoder from for example the second colour space (e.g. BT.709) to the first colour space (e.g. BT.2020), at the appropriate place in the decoding process (e.g. before or after application of one or more residuals enhancement planes such as encoded level 1 stream and encoded level 2 stream). An encoder performing the encoding processes of FIGS. 4 and 5 can be deployed which is reconfigurable to switch between the first variation and the second variation as needed, sending suitable tone mapping parameter(s) as needed. As will now be apparent to the skilled reader, the tone mapping parameter(s) are used to indicate to a decoder whether a tone mapping operation is to be performed or not during a decoding of the video signal to change a signal being processed from a first colour space to a second colour space. Also, in some instances, the tone mapping parameter(s) indicate a location of the tone mapping operation in the decoding process i.e., whether the tone mapping operation is in accordance with the first variation of FIG. 4 or the second variation of FIG. 5. This location of the tone mapping operation may be signalled in the bitstream by a value that indicates whether the enhancement levels or residuals planes are provided in the first or second colour spaces. In the examples of FIGS. 4 and 5, the first colour space is a wide colour gamut (WCG) such as BT.2020 or 2100, DCI-P3, or Adobe (RTM) RGB colour space, and may also be part of a high dynamic range (HDR) signal. The second colour space is typically a standard colour gamut such as BT.709 and may also be part of a standard dynamic range (SDR) signal. FIG. 6 illustrates an example bitstream for signalling features of the first and second variations. The bitstream 600 comprises a version 610, additional information payload 620, HDR payload 630 and video signal 640. Bitstreams 230, 330 or 430 are examples of bitstream 600. Version 610 represents the bitstream version in accordance with a bitstream standardisation. The version 600 indicates the version of the bitstream to a decoder. The version 600 facilitates retro-compatibility of the bitstream with legacy decoders that are not up-to-date with processing newer bitstreams, and also signals to newer decoders additional information and functionality. For example, the version 600 may indicate a tone mapping operation and may be a tone mapping parameter itself simply by virtue of the version number indicating that a particular tone mapping operation is to be applied by a decoder, and / or the version 600 may indicate whether or not a decoder is to look for other tone mapping parameters in the bitstream which are used to signal to a decoder how to apply tone mapping to arrive at an improved rendition of the encoded video signal received at the decoder. The version 610 may in some deployments indicate to at least some decoders that a tone mapping operation is to be applied when decoding the video signal 640. For those decoders that understand the version 610, they may deploy a tone mapping operation during decoding that corresponds to the version 600. In other words, one version number may indicate to decoders to perform a particular tone mapping operation, or perform the tone mapping operation at a particular location in the decoding process, or some other detail of the tone mapping operation to be performed, depending on how the video signal 640 was encoded. For legacy decoders that do not understand the version 610 or are incompatible with the version 610, the tone mapping operations would be ignored and legacy video decoding measures would take place at the legacy decoders. The version 610 may in some deployments instead indicated to at least some decoders that further additional tone mapping parameters are included in the bitstream 600. The additional information payload 620 comprises additional information payload flags to indicate that other additional information may be included in the bitstream. In this example bitstream, the additional information payload 620 comprises a HDR INFO PRESENT flag 621 to indicate whether a HDR payload 630 is present in the bitstream 600. In other words, an example of tone mapping parameters includes SEI (Supplemental Enhancement Information) + PAYLOAD. The HDR payload 630 comprises one or more tone mapping parameters for indicating to a decoder how to apply a tone mapping operation, and in this illustrated embodiment comprises an enhancement_dynamic_range_flag 631, a tone_mapper_type parameter 632, a tone_mapper_data_present_flag 633, tone_mapper_data 634 and a tone^mapperjype^extended parameter 635. The following table shows an example configuration of a HDR PAYLOAD: Table 1 — Configuration of HDR PAYLOAD Syntax Descriptor hdr payload global config(payload size) { enhancement dynamic range flag u(l) tone mapper type u(5) tone mapper data present flag u(l) if (tone mapper data present flag == 1) | tone mapper data mb J if (tone mapper type == 31) { tone mapper type extended u(8) } While the term HDR payload is used, and enhancement dynamic range flag, the term may be broadly applied to tone mapping without dynamic range adjustment, or tone mapping plus some dynamic range adjustment, or dynamic range adjustment without tone mapping. The enhancement_dynamic_range_flag 631 indicates whether the enhancement layer is to be applied in a first colour space or in a second colour space at a decoder. The location of the tone map is indicated using binary notation. For example, if the enhancementjynamic^rangejag 631 equals 1 this indicates that enhancement data shall be applied by the decoder in the enhancement colour space (e.g. BT.2020), or if the enhancement_dynamic_range_flag 631 equals 0 this indicates that enhancement data shall be applied by the decoder in the base colour space (e.g. BT.709). If enhancement_dynamic_range_flag 631 is not present, it may be inferred to be equal to 0. The tone^mapperjype parameter 632 specifies the type of tone mapper algorithm to be used for tone mapper operations at the decoder. In this example, the tone_mapper_type 632 parameter comprises a predefined number which indicates to a decoder which tone mapping algorithm to apply in the decoding process. See Table 2 below for example tone_mapper_type_extended 635 parameters and algorithms. If tone_mapper_type parameter 632 is equal to 0, the tone mapping operation is disabled. If tone_mapper_type parameter 632 is equal to 1, then the algorithm may be SDR BT.709 to PQ BT.2020 vOl, for example. If tone^mapperjype 632 is equal to 31, then the algorithm is identified from tone^mapperjype^extended parameter 635. The tone^mapperjype 632 would typically be the inverse of the corresponding tone mapper operation performed at the encoder, but not necessarily so. Table 2 — tone mapper type 0 Disabled 1 SDR BT.709 to PQ BT.2020 vOl 2 SDR BT.709 to HLG BT.2020 vOl (Hable) 3 30 Reserved 31 Extended range The tone_mapper_type_extended parameter 635 provides an extended range of possible types of tone mapping algorithm to be used for the tone mapping operation. The extended range of fields may be 0 to 255 providing a possible 256 further tone mapping algorithms that may be signalled by the bitstream 600. The tone mapper d 633 indicates that additional data for the tone mapping operation is present in the bitstream 600. Tone_mapper_data_present_flag 633 is represented in binary. For example, in one implementation if the tone_mapper_data_present_flag 633 is equal to 1 then additional data for the tone mapping operation is present in the bitstream 600 and if the tone^m 633 is equal to 0 then no additional data for the tone mapping operation is present in the bitstream 600. If tone^mapperjata_present_flag 633 is not present, it may be inferred to be equal to 0. The tone_mapper_data 634 indicates data that the tone mapping algorithm may need during the tone mapping operation at the decoder, for example in order to apply the tone mapping operation more accurately. For example, the encoder may perform a tone mapping operation using an algorithm having particular parameters, and so the decoder is provided with those parameters in order to perform a corresponding tone mapping operation (for example an approximation of an inverse tone map to that used in the encoder). The tone mapping parameter or parameters may also indicate a change in bit depth in a signal associated with the tone mapping operation. For example, the tone mapping parameter or parameters may indicate a value representing the difference in bit depth from a first bit depth to a second bit depth. In the example of FIG. 6, the tone mapping parameter or parameters indicates a bit depth of decoded base layer pictures and the bit depth of decoded enhancement layer pictures. As the above defined signalling and their respective algorithms are proprietary, a few clarifications are added to ensure any bitstream is still compliant with ISO / IEC 23094-2, even if a decoder is not aware of this proprietary signalling. • base_depth_type in ISO / IEC 23094-2 defines the bit depth of the decoded base picture. The value of base_depth_type shall be the same as the value of the bit depth used for the decoded base picture. Note: The value of base^depthjype shall always be the actual bit depth of the decoded base pictures. Following a tone mapping operation, the input bit depth to a decoder may be different. This change shall be identified in the description of the specific tone mapping algorithm. • enhancement_depth_type defines the bit depth of the enhanced decoded picture. The value of enhancemenLdepthJype shall be the actual bit depth of the decoded enhancement pictures. A subsequent tone mapping operation may change the bit depth of the final output pictures. This change shall be identified in the description of the specific tone mapping algorithm. If the base and enhancement layers are signalled with different bit depths, the following two scenarios can apply: 1. If enhancemenLdynamic^rangeJag is true (or 1), the "in-loop" tone mapping operation converts the bit depth between the base and the enhancement layer. 2. If enhancement_dynamic_range_flag is false (or 0), the tone mapping operation is performed "out-of-Ioop". The bit depth is converted separately between the base and enhancement layer. Video signal 640 comprises an encoded video signal, such as the video signal shown in Figures 2, 3, 4 and 5 comprising encoded base stream 201, encoded level 1 stream 202 and encoded level 2 stream 203. A preliminary picture may be a picture waiting to be enhanced. In other words, a preliminary picture may be a to be enhanced picture. In embodiments, the preliminary picture may be one or more of: a picture received from a base decoder, an upsampled uncorrected rendition of a picture received from a base layer, an upsampled corrected version of a base layer, a double-upsampled uncorrected version, a double upsampled corrected rendition (e.g. upsampled and corrected then upsampled again). A picture received from a base decoder may be referred to as a picture received from the base layer. Upsampling may be one or more of spatial, temporal, quality, bit depth upsampling. A preliminary picture may be referred to as a reconstructed picture derived from a base decoder. An enhancement residuals plane may be a plane of correction data. An enhancement residuals plane may be for combining with the preliminary picture, to obtain an enhanced picture. An enhancement residuals plane may be referred to as a residuals surface. Encoded data within an encoded bitstream may be separated into chunks. More particularly, a bitstream generated by an enhancement encoder may have a particular data structure (e.g. level 1 and level 2 encoded data). We describe a plurality of planes (of number nPlanes). Each plane relates to a particular colour component. We describe an example with YUV colour planes (e.g. where a frame of input video has three colour channels, i.e. three values for every pixel). In the examples, the planes may be encoded separately. The data for each plane is further organised into a number of levels (nLevels). We describe two levels, relating to each of enhancement levels 1 and 2. The data for each level is then further organised as a number of layers (nLayers). These layers are separate from the base and enhancement layers; in this case, they refer to data for each of the coefficient groups that result from the transform. For example, an 2x2 transform results in four different coefficients that are then quantized and entropy encoded and an 4x4 transform results in sixteen different coefficients that are then likewise quantized and entropy encoded. In these cases, there are thus respectively 4 and 16 layers, where each layer represents the data associated with each different coefficient. In cases, where the coefficients are referred to as A, H, V and D coefficients then the layers may be seen as A, H, V and D layers. In certain examples, these "layers" are also referred to as "surfaces", as they may be viewed as a "frame" of coefficients in a similar manner to a set of two-dimensional arrays for a set of colour components. The data for the set of layers may be considered as "chunks". As such each payload may be seen as ordered hierarchically into chunks. That is, each payload is grouped into planes, then within each plane each level is grouped into layers and each layer comprises a set of chunks for that layer. A level represents each level of enhancement (first or further) and layer represents a set of transform coefficients. In any decoding process, the method may comprise retrieving chunks for two levels of enhancement for each plane. The method may comprise retrieving 4 or 16 layers for each level, depending on size of transform that is used. Thus, each payload is ordered into a set of chunks for all layers in each level and then the set of chunks for all layers in the next level of the plane. Then the payload comprises the set of chunks for the layers of the first level of the next plane and so on. As such, in the methods described herein, the pictures of a video may be partitioned, e.g. into a hierarchical structure with a specified organisation. Each picture may be composed of three different planes, organized in a hierarchical structure. A decoding process may seek to obtain a set of decoded base picture planes and a set of residuals planes. A decoded base picture corresponds to the decoded output of a base decoder. The base decoder may be a known or legacy decoder, and as such the bitstream syntax and decoding process for the base decoder may be determined based on the base decoder that is used. In contrast, the residuals planes are new to the enhancement layer and may be partitioned as described herein. A "residuals plane" may comprise a set of residuals associated with a particular colour component. For example, although the planes are described as relating to YUV planes of an input video, it should be noted the data does not comprise YUV values, e.g. as for a comparative coding technology. Rather, the data comprises encoded residuals that were derived from data from each of the YUV planes. In certain examples, a residuals plane may be divided into coding units whose size depends on the size of the transform used. For example, a coding unit may have a dimension of 2x2 if a 2x2 directional decomposition transform is used or a dimension of 4x4 if a 4x4 directional decomposition transform is used. The decoding process may comprise outputting one or more set of residuals surfaces, that is one or more sets of collections of residuals. For example, these may be output by the level 1 decoding component and the level 2 decoding component. A first set of residual surfaces may provide a first level of enhancement. A second set of residual surfaces may be a further level of enhancement. Each set of residual surfaces may combine, individually or collectively, with a reconstructed picture derived from a base decoder. The input to the described decoding processes is an enhancement bitstream (also called a low complexity enhancement video coding bitstream) that contains an enhancement layer consisting of up to two sub-layers. The outputs of the decoding process may therefore be: 1) an enhancement residuals planes (sub-layer 1 residual planes) to be added to a set of preliminary pictures that are obtained from the base decoder reconstructed pictures; and 2) an enhancement residuals planes (sub-layer 2 residual planes) to be added to the preliminary output pictures resulting from upscaling, and modifying via predicted residuals, the combination of the preliminary pictures 1 and the sub-layer 1 residual planes. Data may be arranged in chunks or surfaces. Each chunk or surface may be decoded according to an example process substantially similar to that described elsewhere. As such the decoding process operates on data blocks as described in the sections above. An overview of a decoding method will now be set out. The decoding may omit one or more of these steps. One or more of the steps of the decoding method may be combined with the other described decoding methods. In other words one or more of the steps of the decoding method may be added to, or substituted for, a step of other described decoding methods: • A set of payload data block units are decoded. This allows portions of the bitstream following the NAL unit headers to be identified and extracted (i.e. the payload data block units). • A decoding process for the picture receives the payload data block units and starts decoding of a picture using the syntax elements set out above. Pictures may be decoded sequentially to output a video sequence following decoding. We describe extracting a set of (data) surfaces and a set of temporal surfaces as described above. In certain cases, entropy decoding may be applied at this block. • A decoding process for base encoding data extraction is applied to obtain a set of reconstructed decoded base samples (recDecodedBaseSamples). This may comprise applying the base decoder of previous examples. If the base codec or decoder is implemented separately, then the enhancement codec may instruct the base decoding of a particular frame (including sub-portions of a frame and / or particular planes for a frame). The set of reconstructed decoded base samples are then passed to a further block where an optional first set of upscaling may be applied to generate a preliminary intermediate picture. The output of this block is a set of reconstructed level 1 base samples (where level 0 may comprise to the base level resolution). • A decoding process for the enhancement sub-layer 1 (i.e. level 1) encoded data is performed. This may receive variables that indicate a transform size (nTbs), a user data enabled flag (userDataEnabled) and a step-width (i.e. for dequantization), as well as blocks of level 1 entropy-decoded quantized transform coefficients (TransformCoeffQ) and the reconstructed level 1 base samples (recLIBaseSamples). A plane index (IdxPlanes) may also be passed to indicate which plane is being decoded (in monochrome decoding there may be no index). The variables and data may be extracted from the payload data units of the bitstream using the above syntax. This block may comprise a number of sub-blocks that correspond to the inverse quantization, inverse transform and level 1 filtering (e.g. deblocking) components of previous examples. At a first sub-block, a decoding process for the dequantization is performed. This may receive a number of control variables from the above syntax that are described in more detail below. A set of dequantized coefficient coding units or blocks may be output. At a second sub-block, a decoding process for the transform is performed. A set of reconstructed residuals (e.g. a first set of level 1 residuals) may be output. At a third sub-block, a decoding process for a level 1 filter may be applied. The output of this process may be a first set of reconstructed and filtered (i.e. decoded) residuals. In certain cases, the residual data may be arranged in NxM blocks so as to apply an NxM filter. • The reconstructed level 1 base samples and the (e.g. filtered) residuals are combined. This may be is referred to as residual reconstruction for a level 1 block. At output of this block is a set of reconstructed level 1 samples. These may be viewed as a video stream (if multiple planes are combined for colour signals). • A second up-scaling process is applied. This up-scaling process takes a combined intermediate picture and generates a preliminary output picture. It may comprise an application of the up-scaler or any of the previously described up-sampling components. • Switching is implemented depending on a signalled up-sampler type. Respective implementations of a nearest sample up-sampling process, a bilinear up-sampling process, a cubic up-sampling process and a modified cubic up-sampling process. Sub-blocks may be extended to accommodate new up-sampling approaches as required (e.g. such as the neural network up-sampling described herein). The output from sub-blocks is provided in a common format, e.g. a set of reconstructed up-sampled samples, and is passed, together with a set of lower resolution reconstructed samples to a predicted residuals process. This may implement the modified up-sampling described herein to apply predicted average portions. The output of block is a set of reconstructed level 2 modified up-sampled samples (recL2ModifiedUpsampledSamples). • A decoding process for the enhancement sub-layer 2 (i.e. level 2) encoded data. It receives variables that indicate a step-width (i.e. for dequantization), as well as blocks of level 2 entropy-decoded quantized transform coefficients (TransformCoeffQ) and the set of reconstructed level 2 modified up-sampled samples (recL2ModifiedUpsampledSamples). A plane index (IdxPlanes) is also passed to indicate which plane is being decoded (in monochrome decoding there may be no index). The variables and data may again be extracted from the payload data units of the bitstream using the above syntax. • Temporal prediction is applied for enhancement sub-layer 2 (i.e. level 2). Further variables may thus be received, such as variable that relate to temporal processing including the variables temporal_enabled, temporal_refresh_bit, temporal_signalling_present, and temporal_step_width_modifier as well as the data structures TransformTempSig and TileTempSig that provide the temporal signalling data. Two temporal processing sub-blocks are described : a first subblock where a decoding process for temporal prediction is applied using the TransformTempSig and TileTempSig data structures and a second sub-block that applies a tiled temporal refresh. A sub block is configured to set the contents of a temporal buffer to zero depending on the refresh signalling. • Decoding processes for the dequantization and transform are applied to the level 2 data in a similar manner to the level 1 data decoding. A second set of reconstructed residuals that are output from the inverse transform processing are then added to a set of temporally predicted level 2 residuals; this implements partof the temporal prediction. The output is a set of reconstructed level 2 residuals (resL2Residuals). • The reconstructed level 2 residuals (resL2Residuals) and the reconstructed level 2 modified up-sampled samples (recL2ModifiedUpsampledSamples) are combined in a residual reconstruction process for the enhancement sub-layer 2. The output of this block is a set of reconstructed picture samples at level 2 (recL2PictureSamples). These reconstructed picture samples at level 2 may be subject to a dithering process that applies a dither filter. The output to this process is a set of reconstructed dithered picture samples at level 2 (recL2DitheredPictureSamples). These may be viewed as an output video sequence (e.g. for multiple consecutive pictures making up the frames of a video, where planes may be combined into a multi-dimensional array for viewing on display devices). FIG. 7 shows a high-level schematic of a modification showing a first variation of a decoding process. FIG. 7 has many similarities to FIG. 3 and only the differences are described, with like reference signs denoting like features. In the decoding process of FIG. 7, bitstream 430 is received which comprises the tone mapping parameter(s). In accordance with the tone mapping parameter(s), FIG. 7 shows an inverse tone mapping operation 716 which is performed on the output of the base decoder 310 before the first summation component 312. FIG. 7 shows a complementary decoding process to the encoding process of FIG. 4, thereby fostering a reciprocal data handling mechanism that ensures integrity of data at the decoder. FIG. 8 shows a high-level schematic of a second variation of the decoding process of FIG. 7. FIG. 8 has many similarities to FIG. 3 and only the differences are described, with like reference signs denoting like features. In the decoding process of FIG. 8, bitstream 530 is received which comprises the tone mapping parameter(s). In accordance with the tone mapping parameter(s), FIG. 8 shows an inverse tone mapping operation 817 which is performed on the output of the second summation component 315. FIG. 8 shows a complementary decoding process to the encoding process of FIG. 5. A decoder performing the decoding processes of FIGS. 7 and 8 is reconfigurable to switch between the first variation and the second variation as needed based on the signalled tone mapping parameter(s), and to perform the appropriate tone mapping based on the signalled tone mapping parameter(s). Although the above examples are focussed on LCEVC, the above disclosure can also be applied to other coding schemes, such as but not limited to other coding schemes where enhancement residuals are applied to preliminary pictures. The skilled person will understand from this disclosure that the encoding of video data in the way disclosed is not graphics rendering, nor is the disclosure related to transcoding. Instead, the video encoding disclosed relates to the creation of an encoded video stream from an input video source. Definitions and Terms In certain examples described herein the following terms are used: "base layer" - this is a layer pertaining to a coded base picture, where the "base" refers to a codec that receives processed input video data. It may pertain to a portion of a bitstream that relates to the base. "bitstream" - this is sequence of bits, which may be supplied in the form of a NAL unit stream or a byte stream. It may form a representation of coded pictures and associated data forming one or more coded video sequences (CVSs). "block" - an MxN (M-column by N-row) array of samples, or an MxN array of transform coefficients. The term "coding unit" or "coding block" is also used to refer to an MxN array of samples. These terms may be used to refer to sets of picture elements (e.g. values for pixels of a particular colour channel), sets of residual elements, sets of values that represent processed residual elements and / or sets of encoded values. The term "coding unit" is sometimes used to refer to a coding block of luma samples or a coding block of chroma samples of a picture that has three sample arrays, or a coding block of samples of a monochrome picture or a picture that is coded using three separate colour planes and syntax structures used to code the samples. A coding unit may comprise an M by N array R of elements with elements R[x][y]. For a 2x2 coding unit, there may be 4 elements. For a 4x4 coding unit, there may be 16 elements. "chroma" - this is used as an adjective to specify that a sample array or single sample is representing a colour signal. This may be one of the two colour difference signals related to the primary colours, e.g. as represented by the symbols Cb and Cr. It may also be used to refer to channels within a set of colour channels that provide information on the colouring of a picture. The term chroma is used rather than the term chrominance in order to avoid the implication of the use of linear light transfer characteristics that is often associated with the term chrominance. "coded picture" - this is used to refer to a set of coding units that represent a coded representation of a picture. "coded base picture" - this may refer to a coded representation of a picture encoded using a base encoding process that is separate (and often differs from) an enhancement encoding process. "coded representation" - a data element as represented in its coded form "decoded base picture" - this is used to refer to a decoded picture derived by decoding a coded base picture. "decoded picture" - a decoded picture may be derived by decoding a coded picture. A decoded picture may be either a decoded frame, or a decoded field. A decoded field may be either a decoded top field or a decoded bottom field. "decoder" - equipment or a device that embodies a decoding process. "decoding process" - this is used to refer to a process that reads a bitstream and derives decoded pictures from it. "encoder" - equipment or a device that embodies a encoding process. "encoding process" - this is used to refer to a process that produces a bitstream (i.e. an encoded bitstream). "enhancement layer" - this is a layer pertaining to a coded enhancement data, where the enhancement data is used to enhance the "base". It may pertain to a portion of a bitstream that comprises planes of residual data. The singular term is used to refer to encoding and / or decoding processes that are distinguished from the "base" encoding and / or decoding processes. "video frame or frame" - in certain examples a video frame may comprise a frame composed of an array of luma samples in monochrome format or an array of luma samples and two corresponding arrays of chroma samples. The luma and chroma samples may be supplied in 4:2:0, 4:2:2, and 4:4:4 colour formats (amongst others). A frame may consist of two fields, a top field and a bottom field (e.g. these terms may be used in the context of interlaced video). References to a "frame" in these examples may also refer to a frame for a particular plane, e.g. where separate frames of residuals are generated for each of YUV planes. As such the terms "plane" and "frame" may be used interchangeably. "layer" - this term is used in certain examples to refer to one of a set of syntactical structures in a non-branching hierarchical relationship, e.g. as used when referring to the "base" and "enhancement" layers, or the two (sub-) "layers" of the enhancement layer. "luma" - this term is used as an adjective to specify a sample array or single sample that represents a lightness or monochrome signal, e.g. as related to the primary colours. Luma samples may be represented by the symbol or subscript Y or L. The term "luma" is used rather than the term luminance in order to avoid the implication of the use of linear light transfer characteristics that is often associated with the term luminance. The symbol L is sometimes used instead of the symbol Y to avoid confusion with the symbol y as used for vertical location. "network abstraction layer (NAL) unit (NALU)" - this is a syntax structure containing an indication of the type of data to follow and bytes containing that data in the form of a raw byte sequence payload (RBSP). The RBSP is a syntax structure containing an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits containing syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0. The RBSP may be interspersed as necessary with emulation prevention bytes. "network abstraction layer (NAL) unit stream" - a sequence of NAL units. "picture" - this is used as a collective term for a field or a frame. In certain cases, the terms frame and picture are used interchangeably. "residual" - this term is defined in further examples below. It generally refers to a difference between a reconstructed version of a sample or data element and a reference of that same sample or data element. "source" - this term is used in certain examples to describe the video material or some of its attributes before encoding. "transform coefficient" (or just "coefficient") - this term is used to refer to a value that is produced when a transformation is applied to a residual or data derived from a residual (e.g. a processed residual). It may be a scalar quantity, that is considered to be in a transformed domain. In one case, an M by N coding unit may be flattened into an M*N one-dimensional array. In this case, a transformation may comprise a multiplication of the one-dimensional array with an M by N transformation matrix. In this case, an output may comprise another (flattened) M*N one-dimensional array. In this output, each element may relate to a different "coefficient", e.g. for a 2x2 coding unit there may be 4 different types of coefficient. As such, the term "coefficient" may also be associated with a particular index in an inverse transform part of the decoding process, e.g. a particular index in the aforementioned one-dimensional array that represented transformed residuals. According to another aspect of the invention, there is provided a method of creating a bitstream comprising a hierarchical video data signal. The method comprising: processing a source video data signal to create an enhancement layer useable to enhance a base layer in a hierarchical video data signal. The processing comprising: in dependence on a configuration setting or settings, performing or not performing a tone mapping operation at a configured location to change a signal being processed in the processing from a first colour space to a second colour space; setting a tone mapping parameter or parameters to indicate whether a tone mapping operation is to be performed or not during a decoding of the hierarchical video data signal, and if the tone mapping operation is to be performed a location of the tone mapping operation in the decoding; outputting a bitstream comprising the encoded hierarchical video data signal and the tone mapping parameter or parameters. Preferably, setting the tone mapping parameter or parameters to comprise a tone mapping data present flag and an indication of the tone mapping algorithm to be used by a decoder. Preferably, setting the tone mapping parameter or parameters to indicate a tone mapping algorithm to be used during a decoding of the hierarchical video data signal. Preferably, the tone mapping parameter or parameters indicate indicates additional data for usage by the tone mapping operation. Wherein the additional data is independent of the tone mapping algorithm. Preferably, generating the tone mapping parameter or parameters to indicate a tone mapping algorithm which is the inverse or pseudo-inverse of a tone mapping algorithm used during the processing. Preferably, the tone mapping parameter or parameters indicate a tone mapping algorithm to be used by using one of a set of pre-defined numbers, each number corresponding to one of a set of tone mapping algorithms. Preferably, the set of tone mapping algorithms are selected from a list of tone mapping algorithms which are pre-configured in corresponding decoders. Preferably, the tone map parameter or parameters indicates a change in bit depth of the video signal associated with the tone mapping operation. Preferably, the tone mapping parameter or parameters indicates a value representing the difference in bit depth from a first bit depth to a second bit depth. Preferably, the value is signalled in a description of the tone mapping algorithm.Preferably, performing the tone mapping operation in one of the following locations: before creating the enhancement layer; during creating the enhancement layer; and after creating the enhancement layer. Preferably, the location of the tone mapping operation in the decoding indicates whether the enhancement layer is to be applied in a first colour space or in a second colour space at a decoder. Preferably, the tone mapping parameter or parameters comprises a binary indication of the location of the tone mapping. Preferably, signalling in the bitstream that the tone mapping parameter or parameters are present in the bitstream using a bitstream version. Preferably, locating the tone mapping parameter or parameters in a payload of the bitstream. Preferably, locating the tone mapping parameter or parameters in an additionaljnfojype field of the bitstream. Preferably, the additionaljnfojype field is in compliance with standard ISO / IEC 23094-2. Preferably, the bitstream version is in compliance with standard Itu_t_35. Preferably, the first colour space is an HDR colour space. Preferably, the first colour space is in compliance with BT.2020 or 2100. Preferably, the second colour space is an SDR. colour space. Preferably, the second colour space is in compliance with BT.709. Preferably, the enhancement layer is useable to provide enhancements when the base layer spatial resolution is increased. Preferably, the enhancement layer is useable to provide enhancements when the base layer bit depth is increased. Preferably, the enhancement layer is useable when increasing the base layer bit depth from 8 bits to 10 bits. Preferably, the method comprises generating the enhancement layer using an encoding process that is different from the encoding process used to generate the base layer. According to another aspect of the invention, there is provided a method of decoding a bitstream. The method comprising processing a hierarchical video data signal contained within the bitstream. The processing comprises: in dependence on a tone mapping parameter or parameters contained in the bitstream, performing or not performing a tone mapping operation at a configured location to change a signal being processed in the processing from a first colour space to a second colour space. According to another aspect of the invention, there is provided a computer-readable medium comprising instructions which when executed cause a processor to perform the method of any preceding statement. According to another aspect of the invention, there is provided a bitstream comprising a hierarchical video data signal, the hierarchical video data signal comprising a base layer and an enhancement layer. The bitstream comprising a tone mapping parameter or parameters, wherein the tone mapping parameter or parameters indicate whether a tone mapping operation is to be performed or not during a decoding of the hierarchical video data signal to change a signal being processed from a first colour space to a second colour space. The tone mapping parameter or parameters may also indicate that if the tone mapping operation is to be performed a location of the tone mapping operation in the decoding. Preferably, the tone mapping parameter or parameters comprises a tone mapping data present flag and an indication of the tone mapping algorithm to be used by a decoder. Preferably, the tone mapping parameter or parameters indicate a tone mapping algorithm to be used during a decoding of the hierarchical video data signal. Preferably, the tone mapping parameter or parameters indicate additional data for usage by the tone mapping operation. Wherein the additional data is independent of the tone mapping algorithm. Preferably, wherein the tone mapping parameter or parameters indicate a tone mapping algorithm which is the inverse or pseudo-inverse of a tone mapping algorithm used when creating the enhancement layer. Preferably, the tone mapping parameter or parameters indicate a tone mapping algorithm to be used by using one of a set of pre-defined numbers, each number corresponding to one of a set of tone mapping algorithms. Preferably, the set of tone mapping algorithms are selected from a list of tone mapping algorithms which are pre-configured in corresponding decoders. Preferably the tone mapping parameter or parameters indicates a change in bit depth in a signal associated with the tone mapping operation. Preferably the tone mapping parameter or parameters indicates a value representing the difference in bit depth from a first bit depth to a second bit depth. Preferably, the value is signalled in a description of the tone mapping algorithm. Preferably, the tone mapping parameter or parameters indicates that the location of the tone mapping operation was in one of the following locations: before creating the enhancement layer; during creating the enhancement layer; and after creating the enhancement layer. Preferably, the tone mapping parameter or parameters indicates whether the enhancement layer is to be applied in a first colour space or in a second colour space at a decoder. Preferably the tone mapping parameter or parameters comprises a binary indication of the location of the tone mapping. Preferably, the bitstream indicates that the tone mapping parameter or parameters are present in the bitstream using a bitstream version. Preferably, the tone mapping parameter or parameters are located in a payload of the bitstream. Preferably, the bitstream version is in compliance with standard Itu_t_35. Preferably, the tone mapping parameter or parameters are located in an additional_info_type field of the bitstream. Preferably, the additional_info_type field is in compliance with standard ISO / IEC 23094-2. Preferably the bitstream is in accordance with MPEG5 Part 2 LCEVC standard and / or SO / IEC 23094-2. According to another aspect of the invention, there is provided a computer program product and / or a computer-readable medium comprising instructions which when executed cause a processor to perform any of the methods disclosed herein. There is also provided a data signal comprising the computer program product. According to another aspect of the invention, there is provided a method of decoding a bitstream. The method comprises processing a hierarchical video data signal contained within the bitstream. The processing comprises: in dependence on a tone mapping parameter or parameters contained in the bitstream, performing or not performing a tone mapping operation at a configured location to change a signal being processed in the processing from a first colour space to a second colour space. 5 Preferably, the method comprises processing the above-mentioned bitstream or bitstreams and parsing the tone mapping parameter or parameters described and preforming a tone mapping operation in accordance therewith.

Claims

1. A bitstream for transmitting one or more enhancement residuals planes suitable to be combined with a set of preliminary pictures obtained from a decoder reconstructed video, the bitstream comprising:a decoder configuration for controlling a decoding process of the bitstream; and, wherein the decoder configuration comprises a tone mapping parameter.

2. The bitstream of claim 1 wherein the tone mapping parameter indicates whether a decoder should apply the one or more enhancement residual planes in a first colour space or a second colour space.

3. The bitstream of claim 2 wherein the first colour space is a colour space of the decoder reconstructed video.

4. The bitstream of any one of claims 2 to 3 wherein the second colour space is a colour space of a preliminary picture of the set of preliminary pictures.

5. The bitstream of any preceding claim, wherein a first value of the tone mapping parameter indicates to a decoder to apply a tone mapping operation on a combined plane obtained by combining an enhancement residuals plane of the one or more enhancement residuals planes with a preliminary picture of the set of preliminary pictures.

6. The bitstream of any preceding claim wherein a second value of the tone mapping parameter indicates to a decoder to apply a tone mapping operation on a preliminary picture of the set of preliminary pictures, before combining with an enhancement residuals plane of the one or more enhancement residuals planes.

7. The bitstream of any preceding claim wherein the tone mapping parameter indicates a tone mapping algorithm that should be used during a decoding / processing of the bitstream.

8. The bitstream of any preceding claim wherein a first value of the tone mapping parameter indicates a decoder should use a tone mapping algorithm associated with a first entry of a table.

9. The bitstream of claim 8 wherein a second or other value the tone mapping parameter indicates a decoder should use a tone mapping algorithm associated with a second or other entry of a table.

10. The bitstream of any preceding claim wherein a further value of the tone mapping parameter indicates a decoder should use a further tone mapping parameter to determine a tone mapping algorithm to be used during a decoding / processing of the bitstream.

11. The bitstream of any claim 10 wherein the bitstream comprises the further tone mapping parameter.

12. The bitstream of any claim 10 or 11 wherein the further tone mapping parameter specifies a tone mapper algorithm to be used for tone mapper operation.

13. The bitstream of any one of claims 10 to 12 wherein a first or other value of the further tone mapping parameter indicates a decoder should use a tone mapping algorithm associated with a first or other entry of a table.

14. The bitstream of any preceding claim wherein the tone mapping parameter indicates whether additional data relating to a tone mapping operation is present in the bitstream.

15. The bitstream of claim 14 wherein a first value of the tone mapping parameter indicates that the additional data relating to a tone mapping operation is present in thebitstream and a second value of the tone mapping parameter indicates that the additional data relating to a tone mapping operation is not present in the bitstream.

16. The bitstream of claim 14 or 15 wherein the bitstream comprises a further tone mapping parameter specifying a size of the additional data.

17. The bitstream of any one of claims 14 to 16 wherein the bitstream comprises a further tone mapping parameter comprising the additional data.

18. The bitstream of any preceding claim wherein the tone mapping parameter specifies a size of an additional data relating to a tone mapping operation.

19. The bitstream of any preceding claim wherein the bitstream comprises additional data relating to a tone mapping operation.

20. The bitstream of any preceding claim wherein the tone mapping parameter indicates a change in bit depth in a signal associated with the tone mapping operation.

21. The bitstream of any preceding claim wherein the tone mapping indicates a value representing the difference in bit depth from a first bit depth to a second bit depth.

22. The bitstream of any preceding claim wherein the tone mapping parameter indicates a bit depth of decoded base layer pictures.

23. The bitstream of any preceding claim wherein the tone mapping parameter indicates a bit depth of decoded enhancement layer pictures.

24. The bitstream of any preceding claim wherein the bitstream indicates that the tone mapping parameter is present in the bitstream using a bitstream version.

25. The bitstream of any preceding claim the tone mapping parameter is located in a payload of the bitstream.

26. The bitstream of claim 25 wherein the bitstream version is in compliance with standard Itu_t_35.

27. The bitstream of any preceding claim wherein the tone mapping parameter is located in an additionaljnfojype field of the bitstream.5 28. The bitstream of any preceding claim wherein the additional_info_type field is incompliance with standard ISO / IEC 23094-2.

29. The bitstream of any preceding claim wherein the bitstream is in accordance with MPEG5 Part 2 LCEVC standard and / or ISO / IEC 23094-2.

30. A method of decoding the bitstream of any preceding claim using the decoder 10 configuration.

31. A method of generating the bitstream of any of claims 1 to 29.

32. An apparatus configured to perform the method of either of claims 30 or 31.

33. A computer product comprising instruction which when executed by a processor, cause the processor to perform the method of either of claims 30 or 31.15 34. A computer readable media comprising instructions when executed by a processor,cause the processor to perform the method of either of claims 30 or 31.

35. A signal carrying the bitstream of any of claims 1 to 29.

Citation Information

Patent Citations

  • Apparatus and method for encoding and decoding multilayer videos

    US20100226427A1

  • Apparatus and method for multilayer picture encoding / decoding

    US20130300923A1