Encoding and decoding of pre-processing renditions of input videos
Patent Information
- Application Number
- GB2025004196
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-22
- Filing Date
- 2023-08-22
- Publication Date
- 2025-09-03
AI Technical Summary
The increasing number of High Dynamic Range (HDR) video standards leads to incompatibility issues between content providers and device manufacturers, resulting in costly bandwidth and equipment requirements for simulcasting multiple HDR streams, and increased storage overheads for on-demand video providers, while also necessitating inefficient output of multiple video renditions.
A method that uses a single base stream combined with alternative enhancement streams to represent multiple renditions of an input video, where the base stream and enhancement streams are generated by comparing different renditions of the video, allowing for efficient reconstruction of various HDR standards without needing separate streams for each rendition.
This approach reduces resource and data intensity of enhancement streams, simplifies decoders, and decreases signaling overheads by enabling the transmission or storage of a single base stream with multiple enhancement streams, facilitating efficient support of multiple HDR standards and reducing the need for multiple pre-processing formats.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] ENCODING AND DECODING OF PRE-PROCESSING RENDITIONS OF INPUT VIDEOS
[0002] BACKGROUND
[0003] The encoding techniques in the following specification are particularly suited to be used with existing Low Complexity Enhancement Video Coding (LCEVC) techniques.
[0004] A standard specification for LCEVC is provided in the Text of ISO / IEC 23094-2 Ed 1 Low Complexity Enhancement Video Coding published in November 2021 , and many possible implementation details of LCEVC are described in patent publications WO 2020 / 188273 and WO 2020 / 188229. Each of these earlier documents is incorporated herein by reference.
[0005] Broadly speaking, LCEVC enhances the reproduction fidelity of a decoded video after encoding and decoding using an existing codec. This is achieved by combining a base layer with an enhancement layer, where the base layer contains the video encoded using the existing codec, and the enhancement layer indicates a residual difference between the original video and a predicted decoded video produced by decoding the base layer using the existing codec. The enhancement layer can be combined with the decoded base layer to more accurately reproduce the original video.
[0006] Colour conversion within a hierarchical video coding scheme has previously been described in W02020 / 074896, the contents of which are incorporated herein by reference.
[0007] Examples of implementation of LCEVC may be described in WO2022 / 023747 and WO2022 / 023739, which are incorporated herein by reference.
[0008] It remains an objective to effectively and efficiently integrate enhancement coding into existing ecosystems, especially in the case of high dynamic range (HDR) video. In recent years a number of different standards have been developed for use in providing HDR video. These standards each allow for a video to be rendered in different ways which affect the display of that video, and as the demands of users for high quality video content increase so too new standards aim to improve in various ways on earlier standards. However, a downside associated with this rise in the number of HDR standards available is that it can lead to incompatibility between the standards used by content providers and those supported by device manufacturers.
[0009] In the case of broadcast providers, such as television channels, what this has meant in practice is that it is often necessary to simultaneously transmit (or “simulcast”) multiple video streams corresponding to different HDR renditions of a video. In this way, the likelihood of a user’s device not being able to display HDR content due to incompatibility with the HDR standards used by the broadcast provider is reduced.
[0010] In the case of on demand video providers, such as streaming services, although the HDR rendition required by a user’s device can be selected and individually streamed, the on demand video providers nevertheless need to support multiple HDR standards. This requires multiple HDR renditions of content to be stored, which increases overheads for the on demand video provider. In a broadcast scenario, simulcast is very costly in terms of bandwidth and equipment.
[0011] It would therefore be advantageous to provide a more efficient means of supporting multiple HDR standards.
[0012] Similarly, in some cases an input video can undergo multiple different preprocessing processes to produce multiple renditions of the same video. It would be advantageous to provide a means of outputting representations of these renditions that is more efficient than simply outputting multiple streams each corresponding to one of the renditions of the video.
[0013] One such example is when providing videos that include watermarked logos. The watermarks used can be region-specific, so it would be advantageous to provide a more efficient means for supporting alternate watermarks. Another such example is when there are multiple renditions of the same input video, each of which is formatted differently from the other renditions of the input video.
[0014] SUMMARY OF INVENTION
[0015] According to a first aspect of the invention a method for generating a representation of an input video is provided, the method comprising: receiving a base stream and a first enhancement stream, the base stream and the first enhancement stream together providing a representation of a first rendition of the input video; comparing a second rendition of the input video and the base stream to generate a second enhancement stream, the second enhancement stream being an alternative enhancement stream to the first enhancement stream such that the base stream and the second enhancement stream together provide a representation of the second rendition of the input; and outputting the base stream and second enhancement stream. The method may further comprise outputting the first enhancement stream.
[0016] Whereas conventionally it is necessary to treat each rendition of an input video separately by generating an independent stream for each rendition and then outputting one or more of these independent streams, in the present invention a single base stream can be used with multiple alternative enhancement streams. Because the renditions are all based on the same input video, there are many common features of the renditions which can be represented by the base stream. The features specific to each rendition can then be represented by the corresponding enhancement stream. It has been found that because the differences between the renditions are typically small compared with the features the renditions have in common, the enhancement streams are less resource and data intensive than if each rendition were to be represented by separate, independent streams. The present invention therefore allows for multiple renditions of an input video to be represented more efficiently.
[0017] Moreover, advantageously, the decoder may be simplified as it only needs to support one pre-processing format and enhancement coding in order to effectively process the transmitted video. The decoder does not need to support the format which was used to create the common base stream, but only one or more of the pre-processing formats used to create the one or more enhancement streams.
[0018] It should be noted that the step of outputting the base stream, first enhancement stream, and second enhancement stream could comprise transmitting the base stream, first enhancement stream, and second enhancement stream to a receiver or could comprise storing the base stream, first enhancement stream, and second enhancement stream. As such, improvements can be made both to implementations in which these streams are transmitted by reducing the signalling resources required to transmit the streams and to implementations in which the streams are stored for subsequent transmission.
[0019] It should also be noted that the first and second renditions are different. One advantageous example is where each rendition corresponds to a different HDR standard, but there are other implementations of the invention, for example a layering pre-processing. A counterexample to this would be the case of LCEVC in which a base stream and one or more residual streams are used to represent the same rendition of a video. Although the base stream and the first residual stream could be used to provide a representation of this rendition of the video without the second residual stream, this would be a representation of the same rendition of the video as when the base stream is used to provide a representation of the rendition of the video in combination with both residual streams. In contrast, in the present invention the base stream and the first enhancement stream together provide a representation of a different rendition of the input video to that provided by the base stream and the second enhancement stream.
[0020] Another difference is that, in the first aspect of the invention, the first and second enhancement streams are alternatives and the base stream is not used in combination with both of these to provide a representation of a rendition of the input video. To put it another way, the base stream can be combined with either the first enhancement stream or the second enhancement stream. In some instances, the base stream and the first enhancement stream may have already been generated in order to output a representation of the first rendition of the input video. In such instances, the second enhancement stream may be compared with the second rendition of the input video to generate the second enhancements stream, which is then output along with the base stream and the first enhancement stream. However, in other embodiments receiving the first enhancement stream comprises comparing the first rendition of the input video and the base stream to generate the first enhancement stream. In this way, a first enhancement stream may be generated in a similar manner to the second enhancements stream.
[0021] Likewise, the method may comprise further steps for generating a representation of further renditions of the input video. In such embodiments, the method preferably further comprises a step of comparing a third rendition of the input video and the base stream to generate a third enhancement stream, the base stream and the third enhancement stream together providing a representation of the third rendition of the input video. Similarly, the method may comprise still further steps of comparing one or more further renditions of the input video and the base stream to generate one or more further enhancement streams, the base stream and said one or more further enhancement streams together providing representations of the one or more further renditions of the input video. As will be understood by the skilled person, in each of these embodiments the method will typically further comprise outputting any further enhancement streams along with the base stream, first enhancement stream, and second enhancement stream.
[0022] As has been discussed, an advantage of the invention is that a single common base stream may be used which can be combined with each of the enhancement streams to reconstruct the renditions of the input video. This base stream will typically be generated based on one of the renditions of the input video, and it has been found that, in some implementations, one or more renditions of the input video may be more suitable for generating the base stream than others. However, any of the renditions of the input video could be used to generate the base stream. For example, in some embodiments there may be a 0threndition of the input video which is particularly suited to generating the base stream, which is often because the 0threndition of the input video has a high degree of similarity with the other renditions of the input video. However, it is not necessary to generate or to output an enhancement stream corresponding to the 0threndition of the input video. As such, the base stream may be derived from the 0threndition of the input video, in which case the step of receiving the base stream may comprise instructing an encoding of the 0threndition of the input video using a base codec to generate the base steam. This base stream can then be outputted for use in providing representations of the other renditions of the input video in combination with the corresponding enhancement streams.
[0023] It may be advantageous to instruct the encoding of the 0threndition of the input video at full resolution and without changing colour space or otherwise reducing the quality of the 0threndition of the input video. In many instances this will be because the 0threndition is at a lower resolution or is otherwise at a lower quality than the other renditions of the input video. For example, the 0threndition of the input video could itself be generated based on one or more of the other renditions of the input video for the purpose of generating a base stream. However, often it will be more advantageous to reduce the size of the base stream in which case instructing an encoding of a 0threndition of the input video using a base codec to generate the base steam comprises: down-sampling the first rendition of the input video to generate a down-sampled version of a 0threndition of the input video; and instructing an encoding of a down-sampled version of the 0threndition of the input video using a base codec to generate the base stream.
[0024] While the base stream could be generated based on a rendition of the input video for which the enhancement stream is not to be generated or outputted, in many embodiments the base stream will be derived from, which is to say generated based on, the first rendition of the input video. In these embodiments the step of receiving the base stream may comprise instructing an encoding of the first rendition of the input video using a base codec to generate the base steam without down-sampling the first rendition of the input video, and preferably the step of instructing an encoding of the first rendition of the input video using a base codec to generate the base steam comprises: down-sampling the first rendition of the input video to generate a down-sampled version of the first rendition of the input video; and instructing an encoding of the down-sampled version of the first rendition of the input video using a base codec to generate the base stream.
[0025] However the base stream is generated, its purpose is to provide a common base from which representations of one or more renditions of the input video may be constructed in combination with corresponding enhancement streams. As explained above, for each rendition of the input video this involves comparing the base stream with said rendition of the input video. Often the base stream will be encoded, although this may not be the case in some embodiments, such as those in which outputting the base stream comprises storing the base stream in which case the base stream may or may not be encoded. In those embodiments in which the base stream is encoded, comparing a rendition of the input video with the base stream preferably involves comparing the rendition of the input video with a decoded version of the base stream, the base stream having been decoded using a base codec. For example, comparing the second rendition of the input video with the base stream may involve comparing the second rendition of the input video with a decoded version of the base stream, the base stream having been decoded using a base codec.
[0026] The comparison of the base stream, decoded or otherwise, with a rendition of the input video can be carried out in various manners, but comparing said rendition of the input video and the base stream to generate the corresponding enhancement stream advantageously comprises, for each rendition of the input video: downsampling said rendition of the input video to generate a down-sampled version of said rendition of the input video; and comparing the base stream with the down- sampled version of said rendition of the input video to generate a corresponding first residual stream. The first residual stream will be comprised in the corresponding enhancement stream and can then be used in combination with the base stream when reconstructing the rendition of the input video. The skilled person will of course understand that, if the base stream has been encoded, the step of comparing the base stream with the down-sampled version of said rendition of the input video will preferably be carried out using a decoded version of the base stream, in the same way as the comparison with a rendition of the input video with the base stream preferably involves comparing the rendition of the input video with a decoded version of the base stream, the base stream having been decoded using a base codec.
[0027] While the corresponding enhancement stream may solely comprise the first residual stream, because this is generated using a down-sampled version of the corresponding rendition of the input video then the corrected reconstructed video generated when applying the first residual stream to the base stream will be at a lower level of quality. This is typically addressed by up-sampling the corrected reconstructed video, but as will be appreciated there will be differences between the up-sampled reconstructed video and the corresponding rendition of the input video. In order to address this, the step of comparing said rendition of the input video and the base stream to generate the corresponding enhancement stream preferably further comprises, for each rendition of the input video: applying the first residual stream to the reconstructed video to generate a corrected reconstructed video; up-sampling the corrected reconstructed video to generate an up-sampled reconstructed video; comparing the up-sampled reconstructed video with said rendition of the input video to generate a corresponding second residual stream. This second residual stream may then be comprised in the corresponding enhancement stream and can be used in combination with the first residual stream when reconstructing the rendition of the input video to improve the quality of the reconstructed video.
[0028] In some embodiments, for each rendition of the input video, comparing said rendition of the input video and the base stream to generate the corresponding enhancement stream further comprises: instructing an encoding of the corresponding first and second residual streams using an enhancement encoder to generate the corresponding enhancement stream. In other words, the first and second residual streams are generated as described above and then encoded to generate the corresponding enhancement stream. The first and second residual streams will typically be encoded separately, but could be encoded together.
[0029] In other embodiments, the first and second residual streams are separately generated by encoding corresponding sets of residuals, in which case further steps are involved in generating these residual streams. Specifically, comparing the base stream with the down sampled version of said rendition of the input video comprises: instructing the decoding of the base stream to generate a reconstructed video, comparing the decoded version of the base stream with the down-sampled version of said rendition of the input video to generate a corresponding first set of residuals, and instructing an encoding of the first set of residuals using an enhancement encoder to generate the first residual stream; applying the first residual stream to the reconstructed video comprises: instructing a decoding of the corresponding first residual stream using an enhancement decoder to generate a corresponding first decoded residual stream, and applying the first decoded residual stream to the reconstructed video; and comparing the up-sampled reconstructed video with said rendition of the input video comprises: comparing the up-sampled reconstructed video with said rendition of the input video to generate a corresponding second set of residuals, and instructing an encoding of the second set of residuals using an enhancement encoder to generate the second residual stream; wherein the enhancement stream comprises the first and second residual streams.
[0030] Down-sampling a rendition of the input video to generate a down-sampled version of said rendition of the input video serves to reduce the size of the base stream and, in embodiments where the corresponding enhancement stream comprises at least a first residual stream, the first residual stream. It may comprise reducing a resolution of said rendition of the input video and / or implementing other techniques for reducing the size of the rendition of the input video. For example, it may comprise converting said rendition of the input video from a second colour space to a first colour space. The first colour space differs from the second colour space in one or more aspects which allow for a reduction in the size of the rendition of the input video, and converting from a second colour space to a first colour space may include one or more of: mapping between colour spaces (e.g. YCbCr / LMS to RGB); changing a sampling pattern (e.g. 4:4:4 to 4:2:2); converting from a second bit depth to a first bit depth, wherein the second bit depth is greater than the first bit depth; reducing a dynamic range of the luminance component; reducing the colour gamut; changing the electro-optic transfer function model; changing the lowest luminance level (e.g. the black level); changing the highest luminance level (e.g. the white level); and tone mapping.
[0031] The up-sampling performed will then mirror the down-sampling in order to improve the quality of a reconstructed video. For example, up-sampling the corrected reconstructed video to generate an up-sampled reconstructed video may comprise increasing a resolution of the reconstructed video and / or implementing other techniques for improving the quality of the reconstructed video. For example, it may comprise converting the reconstructed video from a first colour space to a second colour space. The second colour space differs from the first colour space in one or more aspects which allow for an improvement in the quality of the reconstructed video, and converting from a first colour space to a second colour space may include one or more of: mapping between colour spaces (e.g. RGB to YCbCr / LMS); changing a sampling pattern (e.g. 4:2:2 to 4:4:4); converting from a first bit depth to a second bit depth, wherein the second bit depth is greater than the first bit depth; increasing a dynamic range of the luminance component; reducing the colour gamut; changing the electro-optic transfer function model; changing the lowest luminance level (e.g. the black level); changing the highest luminance level (e.g. the white level); and tone mapping.
[0032] Preferably, each rendition of the input video corresponds to a respective different format. One or more of these formats may be a standardised format, such as HDR in which case each rendition of the input video corresponds to a respective different HDR standard.
[0033] In some embodiments, each rendition of the input video may correspond to a respective different pre-processing of the input video. For example, two or more renditions could correspond to a same format, such as a same HDR standard, but may have been pre-processed in different ways such that the display of these renditions is altered, by changing the process of displaying the renditions and / or by changing the appearance of the displayed rendition. This could, for example, involve each rendition being associated with different metadata used when displaying the rendition. One specific example is providing two renditions corresponding to a particular HDR standard where one rendition is associated with static metadata and the other rendition is associated with dynamic metadata.
[0034] The renditions of the input video may be generated from the input video in an earlier method, or the method may further comprise instructing the pre-processing of the input video to generate one or more renditions of the input video. Preferably, a different pre-processing module will be instructed for each rendition, which in the case of the first and second renditions of the input video means instructing a first pre-processing module to generate the first rendition of the input video and instructing a second, different pre-processing module to generate the second rendition of the input video.
[0035] The input video itself typically comprises a sequence of frames of data, which may additionally be supplemented by metadata. In such case, comparing renditions of the input video, which includes reconstructions of the input video generated using the base stream and / or an enhancement stream as well as corrected and / or decoded versions of said reconstructions, preferably comprises comparing said renditions on frame by frame basis, which is to say that each frame of one of the renditions being compared is compared with a corresponding frame of the other rendition.
[0036] As has been discussed, the display of the renditions of the input data may be influenced by metadata associated with each of the renditions. In such embodiments, corresponding metadata, which may be different from or the same as the metadata associated with each of the renditions, may be output for use in reconstructing a video from the one or more streams. To this end, the base stream may comprise metadata relating to the display of one or more renditions of the input video. Likewise, one or more of the enhancement streams may comprise metadata relating to the display of the corresponding rendition of the input video. The method may further comprise steps of transmitting data to a receiver allowing for a video to be reconstructed. How this data is transmitted will depend on the use case.
[0037] In a broadcast implementation, such as for display of a video for television, the method may further comprise transmitting the base stream, first enhancement stream, and second enhancement stream to a receiver. The receiver can then select which of the enhancement streams to apply to the base stream to reconstruct the corresponding rendition of the input video.
[0038] In implementations where videos are provided on demand, such as during streaming, it is not necessary to transmit multiple enhancement streams. As such, in these implementations the method may further comprise: receiving a signal from a receiver indicating that one of the first and second enhancement streams should be transmitted; and transmitting the base stream and said one of the first and second enhancement streams to the receiver. In this way, only the enhancement stream corresponding to the rendition of the input video to be reconstructed needs to be transmitted.
[0039] In both cases, the transmitting could be performed by the same entity as the other steps of the method, or the streams could be output to a different entity for transmitting to a receiver, with the transmitting steps being performed by said different entity.
[0040] The enhancement stream may be an enhancement stream encoded according to the LCEVC standard. The base stream may be encoded according to any number of video standard formats, example base codecs include VC-6 or SMPTE ST- 2117, include MPEG standards such as AVC / H.264, HEVC / H.265, etc. as well as nonstandard algorithm such as VP9, AV1 , and others. Each enhancement stream may be instructed to be encoded by a corresponding enhancement encoder. Example instructions are described in WO2022 / 023747, which is incorporated herein by reference. Each enhancement stream may be a separate enhancement stream that does not rely on features from any other enhancement stream. The enhancement stream may comprise further enhancement streams, such as the levels of quality of the LCEVC coding approach.
[0041] In further implementations, the method may comprise generating additional enhancement streams from further renditions of the input signal. For example, the method may comprise comparing a third rendition of the input video and the base stream to generate a third enhancement stream, the third enhancement stream being an alternative enhancement stream to the first and second enhancement streams such that the base stream and the third enhancement stream together provide a representation of the third rendition of the input video; and, additionally and optionally, comparing a fourth rendition of the input video and the base stream to generate a fourth enhancement stream, the fourth enhancement stream being an alternative enhancement stream to the first, second and third enhancement streams such that the base stream and the third enhancement stream together provide a representation of the third rendition of the input video.
[0042] In all embodiments, input video may be downscaled prior to the base coding operation (either before or after pre-processing) and upscaled prior to the difference operation with the pre-processed renditions of the input video. In general signals when comparing two signals, one of the signals may be up or downscaled to match a scaling of the second of the two signals, or vice versa.
[0043] According to a second aspect of the invention, a method for reconstructing a representation of a video is provided, the method comprising: receiving a base stream; selecting a first enhancement stream from two or more enhancement streams, each enhancement stream providing a representation of a different rendition of the video in combination with the base stream; processing the first enhancement stream and the base stream to reconstruct a representation of a first rendition of the video. In a broadcast implementation, such as where the receiver is a television, the method may further comprise, before the step of selecting the first enhancement stream, receiving the two or more enhancement streams. In such an implementation, the receiver will typically receive all of the two or more enhancement streams and will then select which to apply to the base stream to recover a representation of the corresponding rendition of the video.
[0044] In implementations where videos are provided on demand, such as during streaming, the transmitter will reduce signalling overhead by transmitting only those streams required to reconstruct a selected rendition of the video. As such, in these implementations the method may further comprise, after the step of selecting the first enhancement stream, receiving the first enhancement stream. The selection is typically made by the receiver based on signalling received from the transmitter, such as metadata relating to the display of the renditions of the input video corresponding to the two or more enhancement streams. Often, this is transmitted as part of the base stream and the base stream is therefore received prior to the step of selecting a first enhancement stream. However, this signalling may be transmitted separately in which case the base stream could be received after the step of selecting the first enhancement stream.
[0045] In order to indicate to the transmitter which streams are required, in the implementations where videos are provided on demand, the method preferably further comprises, after the step of selecting the first enhancement stream and before the step of receiving the first enhancement stream, sending a signal to a transmitter indicating that the first enhancement stream should be transmitted.
[0046] As with the first aspect of the invention, the base stream may comprise metadata and the first enhancement stream may be selected based on said metadata. Likewise, the two or more enhancement streams may each comprise metadata and the first enhancement stream may be selected based on said metadata.
[0047] As discussed above, it can be advantageous for each enhancement stream to comprise a corresponding first residual stream. In such embodiments, processing the first enhancement stream and the base stream to recover a representation of a first rendition of the video advantageously comprises applying the first residual stream to the base stream to generate a corrected reconstructed video and up- sampling the corrected reconstructed video to generate and up-sampled reconstructed video.
[0048] The first residual stream can be applied to the base stream in a number of ways, but the step of applying the first residual stream to the base stream advantageously comprises: instructing a decoding of the base stream using base codec to generate a reconstructed video; instructing a decoding of the first residual stream using an enhancement decoder to generate a decoded first residual stream; and applying the decoded first residual stream to the reconstructed video.
[0049] Each enhancement stream preferably further comprises a corresponding second residual stream, and in these embodiments processing the first enhancement stream and the base stream to recover a representation of a first rendition of the video may further comprise applying the second residual stream to the up- sampled reconstructed video. Using the second residual stream in this way allows the quality of reconstructed video to be improved, as has been discussed above in relation to the first aspect of the invention.
[0050] Similarly, to the use of the first residual stream, the second residual stream can be applied to the up-sampled reconstructed video in a variety of ways. However, this step typically comprises instructing a decoding of the second residual stream using an enhancement decoder to generate a decoded second residual stream; applying the decoded second residual stream to the up-sampled reconstructed video.
[0051] The up-sampling process used when reconstructing the representation of the video serves to improve the quality of the reconstructed video and corresponds to the up-sampling and down-sampling processes used to generate the enhancement streams. As such, up-sampling the corrected reconstructed video to generate an up-sampled reconstructed video may comprise increasing a resolution of the reconstructed video and / or implementing other techniques for improving the quality of the reconstructed video. For example, it may comprise converting the reconstructed video from a first colour space to a second colour space. The second colour space differs from the first colour space in one or more aspects which allow for an improvement in the quality of the reconstructed video, and converting from a first colour space to a second colour space may include one or more of: mapping between colour spaces (e.g. RGB to YCbCr / LMS); changing a sampling pattern (e.g. 4:4:4 to 4:2:2); mapping between colour spaces (e.g. RGB to YCbCr / LMS); changing a sampling pattern (e.g. 4:2:2 to 4:4:4); converting from a first bit depth to a second bit depth, wherein the second bit depth is greater than the first bit depth; increasing a dynamic range of the luminance component; reducing the colour gamut; changing the electro-optic transfer function model; changing the lowest luminance level (e.g. the black level); changing the highest luminance level (e.g. the white level); and tone mapping.
[0052] Preferably, each rendition of the input video corresponds to a respective different format. One or more of these formats may be a standardised format, such as HDR in which case each rendition of the input video corresponds to a respective different HDR standard.
[0053] In some embodiments, each rendition of the video may correspond to a respective different pre-processing of the video. For example, two or more renditions could correspond to a same format, such as a same HDR standard, but may have been pre-processed in different ways such that the display of these renditions is altered, by changing the process of displaying the renditions and / or by changing the appearance of the displayed rendition. This could, for example, involve each rendition being associated with different metadata used when displaying the rendition. One specific example is providing two renditions corresponding to a particular HDR standard where one rendition is associated with static metadata and the other rendition is associated with dynamic metadata.
[0054] The video itself typically comprises a sequence of frames of data, which may additionally be supplemented by metadata.
[0055] As has been discussed, the display of the renditions of the input data may be influenced by metadata associated with each of the renditions. In such embodiments, corresponding metadata, which may be different from or the same as the metadata associated with each of the renditions, may be output for use in reconstructing a video from the one or more streams. To this end, the base stream may comprise metadata relating to the display of one or more renditions of the video. Likewise, one or more of the enhancement streams may comprise metadata relating to the display of the corresponding rendition of the video.
[0056] The metadata relating to the display of the renditions of the video may be different to the metadata used to select an enhancement stream or may be the same as the metadata used to select an enhancement stream, while in some embodiments the metadata used to select an enhancement stream comprises the metadata relating to the display of the renditions. In a particularly preferred example, the base stream comprises the metadata used to select an enhancement stream while the enhancement streams each comprise relating to the display of a corresponding rendition of the video.
[0057] According to a third aspect of the invention, an apparatus is provided which is configured to perform a method according to any of the embodiments of the first aspect of the invention.
[0058] According to a fourth aspect of the invention, an apparatus is provided which is configured to perform a method according to any of the embodiments of the second aspect of the invention.
[0059] According to a fifth aspect of the invention, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform a method according to any of the embodiments of the first aspect of the invention.
[0060] According to a sixth aspect of the invention, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform a method according to any of the embodiments of the second aspect of the invention.
[0061] According to a seventh aspect of the invention, a bitstream is provided, the bitstream representing two or more renditions of a video, the bitstream comprising: a base stream; a first enhancement stream, the base stream and the first enhancement stream together providing a representation of a first rendition of the input video; and a second enhancement stream, the base stream and the second enhancement stream together providing a representation of a first rendition of the input video.
[0062] This bitstream advantageously allows representations of two or more renditions of a video to be reconstructed while reducing the signalling overheads associated with transmitting two or more video streams. It is especially advantageous in a broadcast implementation, such as for displaying a video on a television, in which representations of multiple renditions of a video may need to be transmitted simultaneously.
[0063] According to an eight aspect of the invention, a bitstream is provided, the bitstream representing a rendition of a video, the bitstream comprising: a base stream; and a first enhancement stream or a second enhancement stream, the base stream and the first enhancement stream together providing a representation of a first rendition of the input video, and the base stream and the second enhancement stream together providing a representation of a first rendition of the input video
[0064] This bitstream allows for a selected rendition of a video to be reconstructed from a base stream that may also be used to reconstruct a different rendition of the video. It is especially advantageous in implementations in which videos are provided on demand, such as during streaming, and may be transmitted in response to an indication of which rendition of the video has been selected.
[0065] Each bitstream can be produced using any of the embodiments of the first aspect of the invention, and as such each bitstream may comprise features corresponding to the steps performed in those embodiments. The bitstreams may also comprise elements generated from any the respective methods described herein. For brevity we do not repeat the elements that may be comprised within the bitstream. For example, the bitstream may comprise an enhancement stream generated by comparison of a rendition of the input video and the base stream, preferably a base decoded version of the base encoded input video.
[0066] According to a ninth aspect of the invention, a method for generating a representation of an input video is provided, the method comprising: receiving a base stream and a first enhancement stream, the base stream and the first enhancement stream together providing a representation of a first rendition of the input video; processing the base stream and the first enhancement stream to generate a reconstructed representation of the first rendition of the input video; comparing a second rendition of the input video with the reconstructed representation of the first rendition of the input video to generate a second enhancement stream, such that the base stream, the first enhancement stream, and the second enhancement stream together provide a representation of the second rendition of the input video; and outputting the base stream, first enhancement stream, and second enhancement stream; wherein each rendition of the input video corresponds to a respective different pre-processing of the input video.
[0067] In this aspect of the invention, the base stream and the first enhancement stream are generated based on the first rendition of the input video such that the base stream and the first enhancement stream together provide a representation of a first rendition of the input video. It has been found that it can be advantageous to generate each additional enhancement stream based on a comparison between a corresponding rendition of the input video and both the base stream and the first enhancement stream, such that each further rendition of the input video is represented by a corresponding further enhancement stream in combination with both the base stream and the first enhancement stream. In this way, fewer features of the further renditions of the input video need to be represented by the further enhancement streams. According to a tenth aspect of the invention, a method for reconstructing a rendition of a video is provided, the method comprising: receiving a base stream and a first enhancement stream; selecting a rendition of a video to reconstruct from two or more renditions of the video; wherein if a first rendition is selected the method comprises processing the first enhancement stream and the base stream to reconstruct a representation of the first rendition of the video; wherein if a second rendition is selected the method comprises receiving a second enhancement stream and processing the second enhancement stream, the first enhancement stream, and the base stream to reconstruct a representation of the second rendition of the video; wherein each rendition of the video corresponds to a respective different pre-processing of the video.
[0068] Although in the ninth and tenth aspects of the invention the first and second enhancement streams are not alternatives but are rather used in combination to provide a representation of the second rendition of the input video, in all other respects the ninth and tenth aspects of the invention correspond to the first and second aspects of the invention, respectively. As such, the preferable features of those aspects discussed above may likewise be applied to the ninth and tenth aspects of the invention.
[0069] According to an eleventh aspect of the invention, a method for embedding representations of two or more renditions of an input video in a bitstream is provided, the method comprising: embedding a representation of the input video in the bitstream; and embedding multiple metadata, each metadata associated with a different pre-processing of the input video; wherein a first metadata and the embedded representation of the input video together provide a representation of a first rendition of the input video; wherein a second metadata and the embedded representation of the input video together provide a representation of a second rendition of the input video.
[0070] This aspect of the invention is particularly advantageous in implementations in which multiple renditions of a video differ only in the metadata associated with the video file. For example, static metadata is used in the HDR10 standard, whereas dynamic metadata may be used in developments of this standard, such as the HDR10+ standard. The eleventh aspect of the invention allows these standards to all be embedded in the same bitstream such that a single bitstream can be used to transmit multiple renditions of a video corresponding to different HDR standards. In these embodiments the multiple metadata preferably each correspond to different HDR standards.
[0071] Preferably, the multiple metadata are each embedded in one or more sei messages and embedding multiple metadata comprises adding said one or more sei messages to the bitstream.
[0072] According to a twelfth aspect of the invention, a method for reconstructing a representation of a video from a bitstream, the bitstream comprising a representation of the video and multiple metadata, each metadata associated with a different pre-processing of the video, wherein a first metadata and the embedded representation of the input video together provide a representation of a first rendition of the input video and a second metadata and the embedded representation of the input video together provide a representation of a second rendition of the input video, the method comprising: selecting the first metadata or the second metadata; extracting the representation of the video and the selected first metadata or second metadata; and reconstructing a representation of the video using the extracted representation of the video and the extracted first metadata or second metadata.
[0073] In this way, a bitstream such as that provided by the eleventh aspect of the invention may be used to reconstruct a representation of a video.
[0074] As with the eleventh aspect of the invention, the multiple metadata preferably each correspond to different HDR standards.
[0075] Preferably, the multiple metadata are each embedded in one or more sei messages and embedding multiple metadata comprises adding said one or more sei messages to the bitstream.
[0076] According to a further aspect there may be provided a bitstream comprising: a base stream representing an embedded representation of an input video; and, multiple metatdata, each metadata associated with a different pre-processing of the input video, wherein a first metadata and the embedded representation of the input video together provide a representation of a first rendition of the input video; wherein a second metadata and the embedded representation of the input video together provide a representation of a second rendition of the input video. The multiple metadata may each embedded in one or more SEI messages. The multiple metadata may each correspond to different HDR standards.
[0077] According to a further aspect there may be provided a method for generating a representation of an input video, the method comprising: obtaining a first rendition of an input video having been pre-processed according to a first pre-processing format; instructing a base encoding of first rendition of an input video; obtaining a second rendition of the input video having been pre-processed according to a second pre-processing format different from the first pre-processing format; comparing the second rendition of the input video with a version of the base encoded version first rendition to generate a set of residuals; and, encoding the set of residuals according to an enhancement coding. The version may be a base decoded version, a base encoded version or a reconstructed version of the first rendition reconstructed from an enhancement layer combined with the base decoded version.
[0078] According to a further aspect there may be provided a method comprising: generating residuals by comparing: a first rendition of an HDR video signal, the first rendition having a first HDR format; and a second rendition of the HDR video signal, the second rendition having a second HDR format wherein the residuals are useable for combining with a third rendition of the HDR video signal to produce a display rendition of the HDR video signal, wherein the first HDR format is different to the second HDR format.
[0079] There may be provided a method comprising generating residuals by comparing a reference rendition of a input video signal an alternative rendition of the input video signal. The reference rendition may be obtained via a processing of the input video signal. The alternative rendition may be obtained via an alternative processing of the input video signal. Said residuals may be usable for combination with the reference rendition to reconstruct the alternative rendition of the input video signal.
[0080] The processing of the input video signal may comprise pre-processing. The alternative processing of the input video signal may comprise alternative preprocessing.
[0081] The pre-processing may comprise pre-processing the video signal in conformance with a first HDR format.
[0082] The alternative pre-processing may comprising pre-processing the video signal in conformance with a second HDR format. The second HDR format may be a different HDR format to the first HDR format.
[0083] The reference rendition may comprise decoded data. The alternative rendition may comprise decoded data.
[0084] The processing (and / or alternate processing) may comprise, before or after said pre-processing, encoding and (e.g. corresponding) decoding the input video signal.
[0085] The comparing may comprise comparing a frame of the reference rendition of the video signal with a corresponding (e.g. this could be one or more of corresponding in time, corresponding in resolution, and so forth) frame of the alternative rendition of the video signal. The result of this comparison may result in generating residuals associated with said frame.
[0086] The method may comprise generating residuals by comparing the reference rendition of the input video signal a further alternative rendition of the input video signal. The further alternative rendition may be obtained via a further alternative processing of the input video signal. The further alternative processing of the input video signal may comprise further alternative pre-processing. The further alternative pre-processing may comprising pre-processing the video signal in conformance with a third HDR format. The third HDR format may be a different HDR format to the first and / or second HDR format. Although we have described this method in relation to ‘alternative’ and ‘further alternative’ renditions, the method may be performed in relation to a plurality of yet further alternative renditions.
[0087] Although methods are described in relation to video signals, the method may be performed in relation with other signals , for example signals suitable for a virtual reality display. In other words, the signal(s) may comprise point cloud data and / or mesh data.
[0088] An apparatus, such as an encoder or decoder, may be provided which is configured to perform a method according to any of the above aspects. Further, a computer-readable medium may be provided, computer-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform a method according to any of the above aspects.
[0089] BRIEF DESCRIPTION OF DRAWINGS
[0090] Embodiments of the invention will now be described with reference to the figures, in which:
[0091] Figure 1 illustrates a schematic flow diagram in which two renditions of a video provided according to different HDR standards are encoded;
[0092] Figure 2 illustrates a further schematic flow diagram in which two renditions of a video provided according to different HDR standards are encoded with enhancement being relative to a base coded from a different HDR standard;
[0093] Figure 3 illustrates a version of Figure 2 in a schematic flow diagram in which two renditions of a video provided according to different HDR standards are encoded and the base corresponds to the first rendition without enhancement;
[0094] Figure 4 illustrates an alternative of Figure 2 in which two renditions of a video provided according to different HDR standards are encoded and in which the common base corresponds to the original PQ input; Figure 5 illustrates an alternative schematic flow diagram in which two renditions of a video provided according to different HDR standards are encoded, with the enhancement of the further rendition being relative to the reconstructed final resolution reference of the first rendition; and,
[0095] Figure 6 illustrates an alternative schematic flow diagram in which two renditions of a video provided according to different HDR standards are encoded relative to a final resolution reference coded using a single-layer codec.
[0096] DETAILED DESCRIPTION
[0097] Figure 1 illustrates an embodiment in which an HDR input video is pre-processed to generate a first rendition in the SL-HDR2 format, while the second rendition comprises the PQ format original input video itself.
[0098] Although one of the two renditions in this embodiment is the input video itself, the base stream is derived from the other rendition. In the case of renditions corresponding to the PQ and SL-HDR2 formats it may be preferable to use the rendition in the SL-HDR2 format to derive the base stream, but this is not necessarily the case. Indeed PQ10 may also provide the most appropriate reference rendition or any other suitable format. We use SL-HDR2 as a reference here, merely for explanation only.
[0099] The next step after the pre-processing of the input video is to instruct the encoding of the first rendition in the SL-HDR2 format using a base encoder, also referred to herein as using a base codec, to generate the base stream. Nevertheless, it would be possible to instead use the rendition in the PQ format to derive the base stream.
[0100] A first enhancement stream LCEVC (A) is then generated by instructing the decoding of the base stream using a base decoder, also referred to herein as using a base codec, to generate a decoded version of the base stream and comparing the decoded version of the base stream with the first rendition in the SL-HDR2 format. A second enhancement stream LCEVC (B) is generated in a similar way by comparing the decoded version of the base stream with the second rendition of the video in the PQ format (i.e. the input video itself).
[0101] A representation of the input video in the SL-HDR2 format can then be generated for display by applying the LCEVC (A) enhancement stream to the base stream, while a representation of the input video in the PQ format can be generated for display by applying the LCEVC (B) enhancement stream to the base stream.
[0102] The base stream, LCEVC (A) enhancement stream, and LCEVC (B) enhancement stream are preferably generated using standard LCEVC techniques such that each of the LCEVC (A) enhancement stream and LCEVC (B) enhancement stream respectively comprise first and second residual streams. The base stream can then be combined with the residual stream of the LCEVC (A) enhancement stream to generate a reconstructed version of the SL-HDR2 video using standard LCEVC reconstruction techniques, while the same base stream can be combined with the residual streams of the LCEVC (B) enhancement stream to generate a reconstructed version of the PQ original input video using standard LCEVC reconstruction techniques.
[0103] The same approach can be used when there are more renditions of the input video, such as is illustrated in Figure 2 in which there are three renditions of the input video in different HDR formats.
[0104] The three renditions are generated by pre-processing an HDR input video, which in this embodiment is in the PQ format, to generate a first rendition in the Technicolor format, a second rendition in the HDR10+ format, and a third rendition in the Dolby Vision format.
[0105] In the embodiment of Figure 2, the base stream is derived from the Technicolor rendition and LCEVC (A), LCEVC (B), and LCEVC (C) enhancement streams are respectively generated for each of the Technicolor, HDR10+, and Dolby Vision renditions of the input video. This process is performed in the same manner as for the renditions discussed in Figure 1 , and likewise could involve using standard LCEVC techniques. In each figure we present examples using different HDR formats and terminologies. This is intentional to demonstrate the utility of the invention across different formats and examples. What is important is that each corresponds to a respective pre-processing. The exact formats used for the common base reference and the residuals may change in use, and may be dependent on many different factors such as complexity and bandwidth.
[0106] In this case a representation of the PQ format is not output by the method, so a different format is used to generate the base stream. Although the Technicolor format has been chosen, one of the other formats could be chosen instead. It would also be possible to generate the base stream using the PQ format input video itself, even though this video is not to be displayed. Likewise, the input video could be pre-processed to provide a 0threndition of the input video in a format that is not to be output by the method and this rendition could then be used to generate the base stream. An example is illustrated in Figure 3 in which the Technicolor rendition is not to be output but where this is used to generate the base stream.
[0107] Figure 4 illustrates an example of the concept in which the input HDR video, e.g. in PQ format is encoded and transmitted. This may then be combined with one of the enhancement streams to recreate one of the HDR standards.
[0108] There exists many different formats of High Dynamic Range (HDR) video. Dynamic range describes the ratio between the smallest and largest possible values of a changeable quantity. In practice, this is the range of tonal difference between the lightest light and the darkest dark of an image. In HDR, bright tones are made brighter without overexposing. Dark tones are made dark without underexposing.
[0109] The differences between Standard Dynamic Range (SDR) and HDR content are mainly related to the colour gamut and the allowed peak brightness, being both much higher in HDR. Specifically, SDR is based on the colorimetric parameters described in Rec. ITU-R BT.709, that covers only 35.9% of the spectrum visible. On the other hand, HDR uses colour parameters described in Rec. ITU-R BT.2020, covering 75.8% of the spectrum. The majority of SDR content uses a colour depth of 8 bits, which allows for representing almost 17 millions of colours. In HDR the colour depth should be at least 10 bits, which allows for representing more than 1 billion colours. In terms of luminance, SDR is limited to 100 cd / m2, while HDR standards theoretically could achieve 10000 cd / m2, although in practice typical consumer displays the peak value is only 1000 cd / m2while the professional ones could achieve up to 4000 cd / m2.
[0110] There are three key components of an HDR video. One is an increased bit-depth to represent more luminance levels. Another is the wide colour gamut, i.e. the range of colours that can be represented. Typically this is defined in a standard such as BT.709 for SDR and BT.2020 for HDR. The third is the electro-optical transfer function which is the transformation from luma pixel values to light values (nits).
[0111] Regular SDR displays use the BT.1886 Gamma transfer function (up to 100 nits). Whereas there are typically two types of transfer function used with HDR, indeed the ITU-R BT.2100 specifies the use of PQ or HLG as transfer functions for HDR- TV. The perceptual quantizer (PQ), published by SMPTE as SMPTE ST 2084, is a transfer function that allows for HDR display by replacing the gamma curve used in SDR. PQ is a display-referred signal that needs a tone mapping operation to adapt the light levels using content metadata. The hybrid log-gamma (HLG) transfer function is a transfer function designed to be backwards compatible with the transfer function of SDR (the gamma curve). HLG is a scene-referred signal that enables automatically adaptation of light levels based on the content and the display capabilities.
[0112] With a PQ transfer function, an HDR system needs metadata to adapt to the display. Examples include static metadata, where there is one set of metadata for the entire video, e.g. HDR10, or dynamic metadata presented on a frame-by- frame orscene-by-scene basis, e.g. HDR10+, Dolby Vision, Technicolor, SL-HDR.
[0113] Few devices support all HDR standards. In fact many devices support a subset but there is no common subset across devices. There is an overlap in the formats supported by different devices. HDR10 was the first standard and it has wide acceptance in the market of manufacturers and video streaming services. HLG was also widely adopted, especially in European countries. HDR10+ presents several improvements compared to HDR10, but it is still adopted by just a few companies. Dolby Vision is gaining popularity, although it is considered the most complex. Other formats include Technicolor Advanced HDR, SL-HDR1 , SL-HDR2 and SL-HDR3 (the SL- HDR standards typically being referred to under the umbrella of Technicolor).
[0114] In a typical implementation scenario, in order to create a video in each of the required standards, an HDR format video having a PQ transfer function is pre- processed to generate a video with the accompanying metadata according to the particular standard. In other words, the pre-processing step transforms the signal to be inline with specific HDR format, calculates specific metadata and the data needed for playback according to specific standard, i.e. specific metadata and / or data to facilitate a device supporting the standard to playback the video. Different coding schemes handle and process the HDR metadata in different ways.
[0115] HDR10 includes static metadata that is applied to the entire video sequence with mastering display colour volume and content light level information. The main difference compared to HDR 10 is that HDR 10+ uses dynamic metadata per frame or scene. However, it offers backward compatibility with HDR10. In case the display does not support dynamic metadata, if static metadata is present, the display could use it. The Dolby Vision standard can use two layers in one video file: base layer (BL), and enhanced layer (EL). The composer that receives these layers should reconstruct the HDR signal from the BL image, the associated EL image, and the related metadata information.
[0116] The metadata can be utilised with existing codecs in different ways. For example, at the MP4 or MKV / WebM media container level, at the elementary video stream level in the corresponding SEI headers, signalled using the OBU, included in SEI messages and / or in NAL unit at the elementary stream level in the corresponding NAL and SEI. Returning to Figure 1 , there is illustrated a process for encoding and transmitting HDR data according to the different available standards. As noted above, this obviates the existing need to simultaneously transmit an encoded video in all HDR standards to ensure display on all devices or otherwise allow the maximum possible range of devices to display HDR content within given bandwidth limitations.
[0117] In general, we refer to pre-processing in the context of video and HDR content pre-processing as a way of explaining the concepts. Other pre-processing may be performed in conjunction with the invention, such as layering pre-processing or point cloud pre-processing. The video signals may be images, video or audio signals, point cloud video or other video suitable for rendering in XR or VR.
[0118] Accordingly, the disclosure may also relate to systems in which images are generated and displayed on a display device and may for example virtually represent objects in a 3D space. The display device may for example be a user extended Reality (“XR”, comprising Virtual Reality and / or Augmented Reality) headset, a pair of XR smart-glasses, an auto-stereoscopic display, a TV display, a mobile device, a PC, etc. According to so-called “split computing” or “remote rendering”, the images are commonly generated remotely from the display device, and there are typically limitations on the communication speed and capacity between the image source and the display device.
[0119] In this latter example, the pre-processing may be any form of point cloud data processing. Additionally or alternatively, where the frames comprise point cloud data, the encoders may apply a point cloud data encoding technique such as described in European patent application EP21386059.6, which is incorporated herein by reference.
[0120] Figure 1 illustrates an input HDR video signal 101. In view of the above, input 101 may be any type of input image, video or point cloud data (or other signal that is to be pre-processed in different manners). Input signal 101 is not limited to HDR video signals. In an example this input video signal may be processed using the PQ transfer function. In further examples, this may be PQ10, sometimes referred to as the PQ format, which is an HDR format that can be used for video and still images. It is equivalent to the HDR10 format without any metadata. PQ10 uses the perceptual quantizer (PQ) transfer function, Rec. 2020 colour primaries and a bit depth of 10-bits. It will of course be understood that PQ10 is merely an example and any suitable HDR input format may be used, which can be used as the starting point for pre-processing the HDR video into a video signal in respective HDR video formats.
[0121] The HDR input video is pre-processed into a first HDR video format 102. In this example, we consider the SL-HDR2 video format but any pre-processing may be performed to generate a pre-processed video in a first format. By first, we don’t mean any temporal significance but instead a label to different the format from the other formats.
[0122] In each of the figures, where a particular HDR format is mentioned, it will be understood that this may be any suitable HDR format or any suitable preprocessing step performed on an input video, such as layering or point cloud preprocessing.
[0123] A base encoding is performed 103 to generate a base encoded version of the video in the HDR format. In examples, the video may be passed to a base encoder which is instructed to encode the video and return the encoded video. As noted above, the HDR metadata may be inserted in different ways according to the base video coding format. For example the HDR metadata may be stored in SEI messages.
[0124] In order for a decoder to regenerate the original PQ input video signal, a base decoding of the base encoded video signal is performed 104 (or instructed by a base decoder) and then compared 105 to the PQ original input. Typically this comparison is a difference operation, but it may be any suitable comparison necessary to generate a set of residuals, the set of residuals representing the difference between the base decoded version of the base encoded HDR format video and the original PQ input video. The set of residuals are then encoded according to an enhancement coding scheme 106, such as LCEVC. The enhancement coding may be performed or instructed by an enhancement encoder. The enhancement coding facilitates reconstruction of the PQ original input by combination of a decoded enhancement stream to generate the set of residuals and a decoded base stream to generate the base video stream. In this way, the encoded residuals represents the difference between the original PQ input signal and a base encoded version of a video in a first HDR format.
[0125] Optionally the base decoded version of the base encoded video in the first HDR format may then be compared to the HDR video signal in the first HDR format 107, here SL-HDR2, so as to correct errors in the base coding. As above, this comparison is preferably a difference operation but may be any suitable comparison operation. The difference between the base decoded version (i.e. the base decoded reconstruction) of the base encoded video signal in the HDR format and the video in the HDR format may be thought of a set of residuals. This set of residuals may be encoded by an enhancement encoder so as to generate an enhancement stream 108. The encoded residuals, once decoded, can be combined with the decoded base stream to correct for errors in the base decoding process. Accordingly, this is an optional step.
[0126] What is important here though, is that the enhancement stream encodes residuals representing the difference between a video of one pre-processing step (here the original PQ input video) and a video of another, different or alternative, preprocessing step (here the SL-HDR2 video format). In this way, a low-cost, low- complexity signal may be used to transmit the signals by transmitting only the encoded difference between the two HDR formats, rather than the full video in each format. In typical enhancement coding schemes, the enhancement layer corresponds to the same format as the base layer.
[0127] In short, the SL-HDR2 video may be reconstructed as the sum of the reconstructed residuals of the LCEVC(B) enhancement stream and the reconstructed base video. It may be desirable to transmit the original PQ input format (or in examples the PQ10 format) to enable certain displays to display HDR video. Similarly, further data or metadata may be sent alongside the base and LCEVC (B) enhancement layer which enables the display to recreate the HDR format from the PQ input format and certain other data for display. For example, this implementation could be combined with the implementation discussed below in which multiple metadata corresponding to different pre-processing steps are embedded in SEI messages of one bitstream. That is, the SEI messages could be embedded with the PQ or PQ10 video.
[0128] Figure 2 illustrates the primary example of the proposed concepts described herein.
[0129] An HDR video 201 , for example a PQ input or other HDR input video, may be pre- processed 202 to generate a video signal in a first HDR video format. Here the video format is given as Technicolor as an example, but any suitable initial HDR video format may be used for the initial pre-processing step. The most suitable initial pre-processing step could be selected so as to ensure the least difference between the other HDR format signals, or alternatively as a balance of objectives depending on the overall use case depending on the size and complexity of the overall transmissions.
[0130] The HDR video in the first format is then passed to a base encoder (or base encoding module) to generate a base stream.
[0131] A further pre-processing step 209 is performed on the original PQ input video. In this example, the pre-processing step is described as the pre-processing of the input video to generate a video signal in an HDR10+ or Dolby Vision video format.
[0132] The base encoded version of the first HDR video signal is then decoded by a base decoder 204. The base decoder generates a base decoded version of the base encoded video signal in the first HDR format. The base decoded version is then compared 205 to the video signal in the further HDR format. The comparison is preferably a difference operation but any suitable comparison may be used, as above. The comparison generates a set of residuals representing the comparison between the video signals in each coding format. That is, the residuals represent the difference between the video signal in the further HDR video format (i.e. after the further pre-processing step) and the base decoded version of the base encoded video signal in the first HDR video format (i.e. after the first preprocessing step).
[0133] The set of residuals are coded 206 according to an enhancement coding scheme such as LCEVC. That is, the residuals are passed to an enhancement encoder or an enhancement coding module.
[0134] In all embodiments, the input video may be downscaled prior to the base coding operation (either before or after pre-processing) and upscaled prior to the difference operation with the pre-processed renditions of the input video. In general signals when comparing two signals, one of the signals may be up or downscaled to match a scaling of the second of the two signals, or vice versa.
[0135] In the figures, we show the pre-processing of different further pre-processing steps, i.e. HDR10+ and Dolby Vision, separately to generate individual and separate alternative enhancement streams LCEVC (B) and LCEVC (C). This illustrates the repeatability of the concepts described herein, but we do not repeat each step in the description above as they are largely similar for each repeated pre-processing pipeline. The different enhancement streams are alternative and separate for each pre-processing step.
[0136] There may be any number of additional pre-processing pipelines relative to the common base reference. For example, there may be an LCEVC (D) or an LCEVC (E) which are based on a different pre-processing step.
[0137] The alternative enhancement streams, once decoded, can be combined with a decoded base to generate an input video in the HDR format. That is, each combination of base and enhancement stream together form an encoded version of the input video signal in one HDR format, although the base stream is an encoded version of a stream from a different pre-processing format. The concepts may be thought of as separate LCEVC encoders and separate LCEVC encoded signals, using a common base. In examples, each enhancement stream may be separate layers of one hierarchy (e.g. one LCEVC instance) or separate (i.e. multiple LCEVC instances). For example, the enhancements could be sub layer 1 and 2 of the same enhancement instance (combined independently or sequentially), or a sub layer of separate enhancement instances (such as for example sub layer 1 or sub layer 2, indeed the metadata could be included in sub layer 1 in one instance and sub layer 2 in another instance, each instance corresponding to a respective pre-processing step).
[0138] By decoding the base stream, it may be possible to recreate the video signal in the first HDR format. However, to correct for errors in the base coding, an LCEVC enhancement stream may be generated based on the first coding format. This involves a comparison operation 207 between a base decoded version of the base encoded signal in the first HDR format and the video signal in the first HDR format. The comparison operation may be a difference operation, as above, that generates a set of residuals which can then be used in an enhancement coding operation 208 to generate an enhancement stream.
[0139] It is anticipated that each alternative enhancement stream may be able to operate at different resolutions, using suitable upscaling and downscaling blocks in the pipeline. Each enhancement stream may support multiple levels of quality, for example as set out in the LCEVC coding scheme.
[0140] In Figure 3, the step of outputting the first enhancement is omitted, as is the difference operation 207 used to generate the first enhancement.
[0141] In Figure 4, the base encoder operation is performed directly on the HDR input (or PQ input) rather than on a pre-processed rendition of the HDR input. As such, pre-processing step 202 is omitted with base coding operation 203 being performed on the HDR input and difference operation 207 being performed between the HDR input and the output of the base decoding operation 204. As discussed above, each pre-processing step may generate metadata. This metadata may be inserted in the enhancement stream. In a typical LCEVC implementation of an HDR coded video signal, the HDR metadata may be inserted in the enhancement layer or the base layer, depending on the intended implementation use case. If the base is being re-used, it is anticipated that the HDR metadata is inserted in the enhancement layer, but it remains possible to insert certain metadata in the base stream.
[0142] In the above explanation, there is a reference video and a series of enhancements each encoding the differences between the reference video and the different preprocessing videos. Each layer may be used to reconstruct differently pre- processed content.
[0143] It is further contemplated that an enhancement may encode different metadata, even if the underlying content is the same. That is, instead of each pre-processing corresponding to an HDR pre-processing of the input video, a pre-processing may produce different metadata, with a bitstream comprising that different metadata and a common reference video.
[0144] Thus, in an implementation, the image may remain the same but the reference and pre-processed signals may have different metadata, the metadata could be encoded in an enhancement stream relative to the reference signal encoded using the base codec. One implementation of this concept would be for different SEI to be included in a stream. Accordingly, this implementation may be thought of as embedding within the same bitstream, multiple metadata, each metadata associated with a specific pre-processing technology. An example includes the SEI messages for Dolby Vision and the SEI messages for HDR 10+ being included. That metadata may be to inform a display how to interpret HDR values, for example, as typical with the HDR format.
[0145] According an implementation, there may be provided one bitstream with multiple sets of SEI messages being embedded associated with the same content, each set of SEI messages associated with a different pre-processing step, e.g. different HDR formats. On the decoding side, this may comprise selecting and decoding one set of metadata transmitted in different enhancement streams or layers.
[0146] Returning to Figure 2, it can be seen that HDR10+ residuals may be generated by comparing an HDR10+ input with a decoded Technicolor / SL-HDR2 base. If the base was obtained at a lower resolution, the base would need to be upsampled to be brought to the same resolution as the HDR10+ input. The relative advantage of this implementation example is that only one enhancement decoding is needed providing a low latency and reducing complexity. Since there is a comparison before or after upsampling, this may lead to relatively high residuals.
[0147] Note, any decoding can be used as a reference for the HDR10+ comparison. Similarly, HDR10+ and Technicolor / SL-HDR are just examples, which can be swapped or used with other HDR schemes (e.g., Dolby Vision, etc.) instead.
[0148] In an alternative example shown in Figure 5, the HDR10+ residuals may be generated by comparing an HDR10+ video with a final reconstructed Technicolor / SL-HDR2 sequence, reconstructed using the base and Technicolor / SL-HDR2 enhancement.
[0149] That is, compared to Figure 2, instead of the difference operation being performed between the base decoded version of the base encoded first pre-processed signal and the further pre-processed signal, instead the difference operation is between the further pre-processed signal and a reconstructed version of the first pre- processed signal, reconstructed using the residuals from the enhancement of the first pre-processed pipeline. That is, the residuals representing the difference between the base decoded version of the base encoded first pre-processed signal and the first pre-processed signal.
[0150] In Figure 5, the output of the first enhancement stream is shown being directed to the further enhancement stream. For brevity, separate base coding operations and the difference operations are not shown to more simply demonstrate the key concept of the alternative implementation. The relative advantage of this implementation examples is there are low residual values as comparison is performed at a high resolution. However, two enhancement decodings may be needed, providing a high latency and complexity.
[0151] In a further alternative example shown in Figure 6, the HDR10+ residuals may be generated by comparing the HDR10+ signal with a regular HDR stream, such as for example an x264 encoding 603 at full resolution.
[0152] That is, each alternative enhancement stream encodes residuals between the further pre-processing signal and a base coded signal of the first pre-processing signal.
[0153] As with Figure 5, the difference operations are not shown for simplicity of illustration. The relative advantage of this implementation example is that only one LCEVC decoding may be needed, with a low disruption to current workflows. However, a higher bandwidth may be needed as there is no enhancement layer used for the first HDR format stream, so inevitably lower compression and / or higher encoding complexity.
[0154] LCEVC as a standard, typically starts from one input video, and based on that and a reconstructed base layer it generates one bitstream formed of maximum two enhancement sub-layers.
[0155] Now, in the case of having two different input videos, it is possible to make the LCEVC scheme fit this scenario by feeding a first video (e.g., SL-HDR2) to the base encoder and a second video (HDR 10+) to the enhancement encoder. In that case, the enhancement layer (namely, what is called enhancement sub-layer 2 in the standard) could be encoding the residuals — these being the differences between the decoded first video and the second input video. This is potentially a different implementation to that of the standard, as a standard encoder would expect one input video, not two (even though the standard doesn’t really specify the decoder, but the test model does).
[0156] Now let’s look at the case of three video inputs, say there is an SL-HDR2, HDR10 and Dolby Vision (DV) video. Now, taking each couple of input videos, it is possible to arrive at the above situation. However, what one would like to do is to encode one of them as reference (say SL-HDR2 to be consistent with the above example), and then at the same time generate the two other bitstreams, one corresponding to the residuals difference between the decoded reference video and a second input video (say HDR10), and another corresponding to the residuals difference between the decoded reference video and a third input video (DV).
[0157] At the encoder side, the LCEVC scheme could be made to fit this scenario by feeding the reference video (e.g., SL-HDR2) to the base encoder, the second video (HDR10) to the enhancement encoder sub-layer 1 and the third video (DV) to the enhancement encoder sub-layer 2.
[0158] Note that sub-layer 1 doesn’t have any temporal prediction, so it is likely that sublayer 1 would be less efficient in encoding as it would need to use intra-frame prediction only.
[0159] When you move to four input videos, say SL-HDR2, HDR10, DV and HDR10+, there is no current configuration of LCEVC which would enable sending three enhancement sub-layers using the same base encoder. In this case, a base encoder could be used with a second instance of LCEVC (for example, the base encoder would have as base layer SL-HDR2 and as enhancement layer HDR10), and then DV and HDR10+ would be encoded as in the above step.
[0160] In all examples described herein, the base video may be output together with, or separately from the LCEVC encoded signal. That is, the base may be encoded with the LCEVC signal or the LCEVC may be encoding the residuals as a separate signal for independent transmission and the LCEVC coder does not include the base signal in the output enhancement sequence.
[0161] In preferred examples, the encoders or decoders are part of a tier-based hierarchical coding scheme or format. Examples of a tier-based hierarchical coding scheme include LCEVC: MPEG-5 Part 2 LCEVC (“Low Complexity Enhancement Video Coding”) and VC-6: SMPTE VC-6 ST-2117, the former being described in PCT / GB2020 / 050695, published as WO 2020 / 188273, (and the associated standard document) and the latter being described in PCT / GB2018 / 053552, published as WO 2019 / 111010, (and the associated standard document), all of which are incorporated by reference herein. However, the concepts illustrated herein need not be limited to these specific hierarchical coding schemes. In certain cases, a base codec may be used. The base codec may comprise an independent codec that is controlled in a modular or “black box” manner. The methods described herein may be implemented by way of computer program code that is executed by a processor and makes function calls upon hardware and / or software implemented base codecs.
[0162] In general, the term “residuals” as used herein refers to a difference between a value of a reference array or reference frame and an actual array or frame of data. The array may be a one or two-dimensional array that represents a coding unit. For example, a coding unit may be a 2x2 or 4x4 set of residual values that correspond to similar sized areas of an input video frame. It should be noted that this generalised example is agnostic as to the encoding operations performed and the nature of the input signal. Reference to “residual data” as used herein refers to data derived from a set of residuals, e.g. a set of residuals themselves or an output of a set of data processing operations that are performed on the set of residuals. Throughout the present description, generally a set of residuals includes a plurality of residuals or residual elements, each residual or residual element corresponding to a signal element, that is, an element of the signal or original data. The signal may be an image or video. In these examples, the set of residuals corresponds to an image or frame of the video, with each residual being associated with a pixel of the signal, the pixel being the signal element. Examples disclosed herein describe how these residuals may be modified (i.e. processed) to impact the encoding pipeline or the eventually decoded image while reducing overall data size. Residuals or sets may be processed on a per residual element (or residual) basis, or processed on a group basis such as per tile or per coding unit where a tile or coding unit is a neighbouring subset of the set of residuals. In one case, a tile may comprise a group of smaller coding units. A tile may comprise a 16x16 set of picture elements or residuals (e.g. an 8 by 8 set of 2x2 coding units or a 4 by 4 set of 4x4 coding units). Note that the processing may be performed on each frame of a video or on only a set number of frames in a sequence. In general, enhancement streams may be encapsulated into one or more enhancement bitstreams using a set of Network Abstraction Layer Units (NALUs). The NALUs are meant to encapsulate the enhancement bitstream in order to apply the enhancement to the correct base reconstructed frame. The NALU may for example contain a reference index to the NALU containing the base decoder reconstructed frame bitstream to which the enhancement has to be applied. In this way, the enhancement can be synchronised to the base stream and the frames of each bitstream combined to produce the decoded output video (i.e. the residuals of each frame of enhancement level are combined with the frame of the base decoded stream). A group of pictures may represent multiple NALUs.
Claims
CLAIMS1. A method for generating a representation of an input video, the method comprising: receiving a base stream and a first enhancement stream, the base stream and the first enhancement stream together providing a representation of a first rendition of the input video; comparing a second rendition of the input video and the base stream to generate a second enhancement stream, the second enhancement stream being an alternative enhancement stream to the first enhancement stream such that the base stream and the second enhancement stream together provide a representation of the second rendition of the input video; and outputting the base stream and second enhancement stream.
2. A method according to claim 1 , wherein each rendition of the input video corresponds to a respective different format.
3. A method according to claim 2, wherein each rendition of the input video corresponds to a respective different HDR standard.
4. A method according to any of claims 1 to 3, wherein each rendition of the input video corresponds to a respective different pre-processing of the input video.
5. A method according to any of claims 1 to 4, wherein the method further comprises instructing the pre-processing of the input video to generate the first and second renditions of the input video.
6. A method according to any of claims 1 to 5, wherein receiving the first enhancement stream comprises comparing the first rendition of the input video and the base stream to generate the first enhancement stream.
7. A method according to any of claims 1 to 6, the method further comprising: comparing a third rendition of the input video and the base stream to generate a third enhancement stream, the base stream and the thirdenhancement stream together providing a representation of the third rendition of the input video.
8. A method according to any of claims 1 to 7, wherein the base stream is derived from a 0threndition of the input video.
9. A method according to claim 8, wherein receiving the base stream comprises instructing an encoding of the 0threndition of the input video using a base codec to generate the base steam.
10. A method according to claim 9, wherein instructing an encoding of a 0threndition of the input video using a base codec to generate the base steam comprises: down-sampling the first rendition of the input video to generate a down- sampled version of a 0threndition of the input video; and instructing an encoding of a down-sampled version of the 0threndition of the input video using a base codec to generate the base stream.
11. A method according to any of claims 1 to 7, wherein the base stream is derived from the first rendition of the input video.
12. A method according to claim 11 , wherein receiving the base stream comprises instructing an encoding of the first rendition of the input video using a base codec to generate the base steam.
13. A method according to claim 12, wherein instructing an encoding of the first rendition of the input video using a base codec to generate the base steam comprises: down-sampling the first rendition of the input video to generate a down- sampled version of the first rendition of the input video; and instructing an encoding of the down-sampled version of the first rendition of the input video using a base codec to generate the base stream.
14. A method according to any of claims 1 to 13, wherein for each rendition of the input video, comparing said rendition of the input video and the base stream to generate the corresponding enhancement stream comprises:down-sampling said rendition of the input video to generate a down- sampled version of said rendition of the input video; and comparing the base stream with the down-sampled version of said rendition of the input video to generate a corresponding first residual stream.
15. A method according to claim 14, wherein for each rendition of the input video, comparing said rendition of the input video and the base stream to generate the corresponding enhancement stream further comprises: applying the first residual stream to the reconstructed video to generate a corrected reconstructed video; up-sampling the corrected reconstructed video to generate an up- sampled reconstructed video; comparing the up-sampled reconstructed video with said rendition of the input video to generate a corresponding second residual stream.
16. A method according to claim 15, wherein for each rendition of the input video, comparing said rendition of the input video and the base stream to generate the corresponding enhancement stream further comprises: instructing an encoding of the corresponding first and second residual streams using an enhancement encoder to generate the corresponding enhancement stream.
17. A method according to claim 15, wherein: comparing the base stream with the down sampled version of said rendition of the input video comprises: instructing a decoding of the base stream using a base codec to generate a reconstructed video; comparing the decoded version of the base stream with the down-sampled version of said rendition of the input video to generate a corresponding first set of residuals; and instructing an encoding of the first set of residuals using an enhancement encoder to generate the first residual stream; applying the first residual stream to the reconstructed video comprises: instructing a decoding of the corresponding first residual stream using an enhancement decoder to generate a corresponding first decoded residual stream; and applying the first decoded residual stream to the reconstructed video; andcomparing the up-sampled reconstructed video with said rendition of the input video comprises: comparing the up-sampled reconstructed video with said rendition of the input video to generate a corresponding second set of residuals; and instructing an encoding of the second set of residuals using an enhancement encoder to generate the second residual stream; wherein the enhancement stream comprises the first and second residual streams.
18. A method according to any of claims 15 to 17, wherein up-sampling the corrected reconstructed video to generate an up-sampled reconstructed video comprises increasing a resolution of the reconstructed video.
19. A method according to any of claims 15 to 18, wherein up-sampling the corrected reconstructed video to generate an up-sampled reconstructed video comprises converting the reconstructed video from a first colour space to a second colour space.
20. A method according to any of claims 10 or 13 to 19, wherein downsampling a rendition of the input video to generate a down-sampled version of said rendition of the input video comprises reducing a resolution of said rendition of the input video.
21. A method according to any of claims 10 or 13 to 19, wherein downsampling a rendition of the input video to generate a down-sampled version of said rendition of the input video comprises converting said rendition of the input video from a second colour space to a first colour space.
22. A method according to any of claims 1 to 21 , wherein the input video comprises a sequence of frames of data.
23. A method according to any of claims 1 to 22, wherein the base stream comprises metadata relating to the display of one or more renditions of the input video.
24. A method according to any of claims 1 to 23, wherein each enhancement stream comprises metadata relating to the display of the corresponding rendition of the input video.
25. A method according to any of claims 1 to 24, wherein the method further comprises transmitting the base stream, first enhancement stream, and second enhancement stream to a receiver.
26. A method according to any of claims 1 to 25, wherein the method further comprises: receiving a signal from a receiver indicating that one of the first and second enhancement streams should be transmitted; and transmitting the base stream and said one of the first and second enhancement streams to the receiver.
27. A method for reconstructing a representation of a video, the method comprising: receiving a base stream; selecting a first enhancement stream from two or more enhancement streams, each enhancement stream providing a representation of a different rendition of the video in combination with the base stream; processing the first enhancement stream and the base stream to reconstruct a representation of a first rendition of the video.
28. A method according to claim 27, wherein the method further comprises, before the step of selecting the first enhancement stream, receiving the two or more enhancement streams.
29. A method according to claim 27, wherein the method further comprises, after the step of selecting the first enhancement stream, receiving the first enhancement stream.
30. A method according to claim 29, wherein the method further comprises, after the step of selecting the first enhancement stream and before the step ofreceiving the first enhancement stream, sending a signal to a transmitter indicating that the first enhancement stream should be transmitted.31 . A method according to any of claims 27 to 30, wherein the base stream comprises metadata and the first enhancement stream is selected based on said metadata.
32. A method according to any of claims 27 to 31 , wherein the two or more enhancement streams each comprise metadata and the first enhancement stream is selected based on said metadata.
33. A method according to any of claims 27 to 32, wherein each enhancement stream comprises a corresponding first residual stream and processing the first enhancement stream and the base stream to recover a representation of a first rendition of the video comprises: applying the first residual stream to the base stream to generate a corrected reconstructed video; and up-sampling the corrected reconstructed video to generate and up- sampled reconstructed video.
34. A method according to claim 33, wherein applying the first residual stream to the base stream comprises: instructing a decoding of the base stream using a base codec to generate a reconstructed video; instructing a decoding of the first residual stream using an enhancement decoder to generate a decoded first residual stream; and applying the decoded first residual stream to the reconstructed video.
35. A method according to claim 33 or claim 34, wherein each enhancement stream further comprises a corresponding second residual stream and processing the first enhancement stream and the base stream to recover a representation of a first rendition of the video further comprises: applying the second residual stream to the up-sampled reconstructed video.
36. A method according to claim 35, wherein applying the second residual stream to the up-sampled reconstructed video comprises: instructing a decoding of the second residual stream using an enhancement decoder to generate a decoded second residual stream; applying the decoded second residual stream to the up-sampled reconstructed video.
37. A method according to any of claims 33 to 36, wherein up-sampling the corrected reconstructed video to generate an up-sampled reconstructed video comprises increasing a resolution of the reconstructed video.
38. A method according to any of claims 33 to 37, wherein up-sampling the corrected reconstructed video to generate an up-sampled reconstructed video comprises converting the reconstructed video from a first colour space to a second colour space.
39. A method according to any of claims 27 to 38, wherein each rendition of the video corresponds to a respective different format.
40. A method according to claim 39, wherein each rendition of the video corresponds to a respective different HDR standard.41 . A method according to any of claims 27 to 40, wherein each rendition of the video corresponds to a respective different pre-processing of the video.
42. A method according to any of claims 27 to 41 , wherein the video comprises a sequence of frames of data43. A method according to any of claims 27 to 42, wherein the base stream comprises metadata relating to the display of one or more renditions of the video.
44. A method according to any of claims 27 to 43, wherein each enhancement stream comprises metadata relating to the display of the corresponding rendition of the video.
45. An apparatus configured to perform a method according to any of claims 1 to 26 or 27 to 44.
46. A non-transitory computer-readable storage medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform a method according to any of claims 1 to 26 or 27 to 44.
47. A bitstream representing two or more renditions of a video, the bitstream comprising: a base stream; a first enhancement stream, the base stream and the first enhancement stream together providing a representation of a first rendition of the input video; and a second enhancement stream, the base stream and the second enhancement stream together providing a representation of a first rendition of the input video.
48. A bitstream representing a rendition of a video, the bitstream comprising: a base stream; and a first enhancement stream or a second enhancement stream, the base stream and the first enhancement stream together providing a representation of a first rendition of the input video, and the base stream and the second enhancement stream together providing a representation of a first rendition of the input video.
49. A method for generating a representation of an input video, the method comprising: receiving a base stream and a first enhancement stream, the base stream and the first enhancement stream together providing a representation of a first rendition of the input video; processing the base stream and the first enhancement stream to generate a reconstructed representation of the first rendition of the input video; comparing a second rendition of the input video with the reconstructed representation of the first rendition of the input video to generate a secondenhancement stream, such that the base stream, the first enhancement stream, and the second enhancement stream together provide a representation of the second rendition of the input video; and outputting the base stream, first enhancement stream, and second enhancement stream; wherein each rendition of the input video corresponds to a respective different pre-processing of the input video.
50. A method for reconstructing a rendition of a video, the method comprising: receiving a base stream and a first enhancement stream; selecting a rendition of a video to reconstruct from two or more renditions of the video; wherein if a first rendition is selected the method comprises processing the first enhancement stream and the base stream to reconstruct a representation of the first rendition of the video; wherein if a second rendition is selected the method comprises receiving a second enhancement stream and processing the second enhancement stream, the first enhancement stream, and the base stream to reconstruct a representation of the second rendition of the video; wherein each rendition of the video corresponds to a respective different pre-processing of the video.51 . A method for embedding representations of two or more renditions of an input video in a bitstream, the method comprising: embedding a representation of the input video in the bitstream; and embedding multiple metadata, each metadata associated with a different pre-processing of the input video; wherein a first metadata and the embedded representation of the input video together provide a representation of a first rendition of the input video; wherein a second metadata and the embedded representation of the input video together provide a representation of a second rendition of the input video.
52. A method according to claim 51 , wherein the multiple metadata are each embedded in one or more SEI messages and embedding multiple metadata comprises adding said one or more SEI messages to the bitstream.
53. A method according to claim 51 or claim 52, wherein the multiple metadata each correspond to different HDR standards.
54. A method for reconstructing a representation of a video from a bitstream, the bitstream comprising a representation of the video and multiple metadata, each metadata associated with a different pre-processing of the video, wherein a first metadata and the embedded representation of the input video together provide a representation of a first rendition of the input video and a second metadata and the embedded representation of the input video together provide a representation of a second rendition of the input video, the method comprising: selecting the first metadata or the second metadata; extracting the representation of the video and the selected first metadata or second metadata; and reconstructing a representation of the video using the extracted representation of the video and the extracted first metadata or second metadata.
55. A method according to claim 54 wherein the multiple metadata are each embedded in one or more SEI messages in the bitstream.
56. A method according to claim 54 or 55 wherein the multiple metadata each correspond to different HDR standards.
57. A bitstream comprising: a base stream representing an embedded representation of an input video; and, multiple metatdata, each metadata associated with a different preprocessing of the input video, wherein a first metadata and the embedded representation of the input video together provide a representation of a first rendition of the input video;wherein a second metadata and the embedded representation of the input video together provide a representation of a second rendition of the input video.
58. A bitstream according to claim 57 wherein the multiple metadata are each embedded in one or more SEI messages.
59. A bitstream according to claim 57 or 58 wherein the multiple metadata each correspond to different HDR standards.
Citation Information
Patent Citations
Generation of high dynamic range images from low dynamic range images in multiview video coding
US20130108183A1
Backward-compatible HDR codecs with temporal scalability
US20170295382A1
Video Coding and Delivery with Both Spatial and Dynamic Range Scalability
US20180220144A1
Colour conversion within a hierarchical coding scheme
US20210360253A1