Rendering of layered video signals

By overlaying enhancement streams onto a first rendered video signal without accessing its data, the method enhances video quality while addressing the challenges of multi-layer video coding schemes and security protections.

WO2025109332A1PCT designated stage expired Publication Date: 2025-05-30V NOVA INT LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2024/052953
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-23
Filing Date
2024-11-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing multi-layer video coding schemes face challenges in widespread adoption due to difficulties in adapting existing and new decoders to process multi-layer encoded streams, particularly when security protections restrict access to the base video data.

Method used

A method of rendering a representation of a reference video signal by obtaining a first rendered video signal at a first level of quality and overlaying enhancement overlay streams onto it to generate a second rendered video signal at a higher quality, without requiring access to the data in the first rendered video signal.

Benefits of technology

This approach allows for the enhancement of video quality without compromising security protections, enabling more benefits of multi-layer coding schemes to be realized, particularly in scenarios where access to the base video data is restricted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2024052953_30052025_PF_FP_ABST
    Figure GB2024052953_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A method of rendering a representation of a reference video signal is disclosed, the method comprising: obtaining a first rendered video signal representing the reference video signal at a first level of quality; and overlaying one or more enhancement overlay streams on the first rendered video signal to generate a second rendered video signal representing the reference video signal at a second level of quality, the second level of quality being a higher level of quality than the first level of quality.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] RENDERING OF LAYERED VIDEO SIGNALS

[0002] FIELD OF THE INVENTION

[0003] The present invention relates to a method of rendering a representation of a reference video signal, a system for the same, and a computer readable medium.

[0004] BACKGROUND

[0005] Encoding and decoding of video content is a consideration in many known systems. Video content may be encoded for transmission, for example over a data communications network. When such video content is decoded, it may be desired to increase a level of quality of the video and / or recover as much of the information contained in the original video as possible. Many video coding formats, and their associated codecs, have been developed that attempt to achieve these desired characteristics, but often require significant software updates at the level of an operating system and / or hardware upgrades. Furthermore, to increase the quality of decoded video content, it is typically required to increase the complexity of the encoding and decoding procedures, which can increase power usage and increase the latency with which video content can be delivered.

[0006] One approach to increasing the quality of decoded video content is to employ multi-layer video coding schemes. Although these have existed for a number of years, they have experienced problems with widespread adoption. Much of the video content on the Internet is still encoded using H.264 (also known as MPEG- 4 Part 10, Advanced Video Coding - MPEG-4 AVC), with this format being used for between 80-90% of online video content. This content is typically supplied to decoding devices as a single video stream that has a one-to-one relationship with available hardware and / or software video decoders, e.g. a single stream is received, parsed, and decoded by a single video decoder to output a reconstructed video signal. Many video decoder implementations are thus developed according to this framework. To support different encodings, decoders are generally configured with a simple switching mechanism that is driven based on metadata identifying a stream format. Existing multi-layer coding schemes include the Scalable Video Coding (SVC) extension to H.264, Scalable extensions to H.265 (MPEG-H Part 2 High Efficiency Video Coding - SHVC), and newer standards such as MPEG-5 Part 2 Low Complexity Enhancement Video Coding (LCEVC). While H.265 is a development of the coding framework used by H.264, LCEVC takes a different approach to scalable video. SVC and SHVC operate by creating different encoding layers and feeding each of these with a different spatial resolution. Each layer encodes the input according to a normal AVC or HEVC encoder with the possibility of leveraging information generated by lower encoding layers. LCEVC, on the other hand, generates one or more layers of enhancement residuals as compared to a base encoding, where the base encoding may be of a lower spatial resolution.

[0007] One reason for the slow adoption of multi-layer coding schemes has been the difficulty adapting existing and new decoders to process multi-layer encoded streams. As discussed above, video streams are typically single streams of data that have a one-to-one pairing with a suitable decoder, whether implemented in hardware or software or a combination of the two. Client devices and media players, including Internet browsers, are thus built to receive a stream of data, determine what video encoding the stream uses, and then pass the stream to an appropriate video decoder. Within this framework, multi-layer schemes such as SVC and SHVC have typically been packaged as larger single video streams containing multiple layers, where these streams may be detected as “SVC” or “SHVC” and the multiple layers extracted from the single stream and passed to an SVC or SHVC decoder for reconstruction. However, this approach fails to offer some of the benefits of multi-layer encodings. Hence, many developers and engineers have concluded that multi-layer coding schemes are too cumbersome and return instead to a multicast of single H.264 video streams.

[0008] It may also be desirable to embed the video content within a web page for playback by an end user using the World Wide Web. To display video content within a webpage, a media element can be included in a Hypertext Markup Language (HTML) document that embeds a media player into the webpage and through which video content can be played. For example, the latest version of HTML, HTML5, includes a video element to embed video content. However, a browser may be unable to render video content of a particular video coding format. In particular, a browser may be unable to render video content encoded using a multi-layer coding scheme.

[0009] Further problems arise when existing multi-layer coding schemes are used to encode protected content.

[0010] In general, there are two common approaches taken to secure video delivery, Conditional Access (CA) and Digital Rights Management (DRM). Conditional Access is used in the more traditional broadcast world with a physical authentication system (typically a smartcard). For online content distribution operators generally use Digital Rights Management (DRM). The aim of both of these approaches is to prevent consumers gaining illegal access to the content, as this would allow them to freely distribute that content to other people.

[0011] With regard to protection, there are generally three areas that need to be considered: the compressed video stream should be encrypted; the output of the then decrypted video stream to the display should be through a protected pipe, e.g. using High-bandwidth Digital Content Protection (HDCP); and it should not be possible for software to capture the content from the decoded and decrypted video. The latter can be achieved by having a secure platform that prevents the execution of unapproved software or by utilising a secure memory controller that prevents general access to the secure memory.

[0012] Implementing multi-level coding scheme, such as a LCEVC, to protected content therefore has to contend with the security protections used to prevent the unauthorised capturing of content.

[0013] In the standard approach taken to encoding protected content using an LCEVC implementation, it is only the base layer which presents a security risk, as the residual data included in the enhancement layers is sparse and is not of interest without the decoded base video. However, the encryption of the base layer and the security protections applied to the display of the rendered base video make it difficult to apply enhancements to that base video. For example, standard approaches to implementing LCEVC involve combining frames of a base video with frames of one or more enhancement layers, something that is not possible when security protections do not allow access to the individual frames of the base video.

[0014] It is thus desired to obtain an improved method and system for decoding multilayer video data that overcomes some of the disadvantages discussed above and that allows more of the benefits of multi-layer coding schemes to be realised.

[0015] SUMMARY OF INVENTION

[0016] According to a first aspect of the invention there is provided a method of rendering a representation of a reference video signal, the method comprising: obtaining a first rendered video signal representing the reference video signal at a first level of quality; and overlaying one or more enhancement overlay streams on the first rendered video signal to generate a second rendered video signal representing the reference video signal at a second level of quality, the second level of quality being a higher level of quality than the first level of quality.

[0017] In applying enhancements to the first rendered video signal by overlaying one or more enhancement overlay streams, the present invention obviates the need for access to the data in the first rendered video signal itself. As compared to conventional hierarchical coding schemes, in which access to a base layer is required in order to render an enhanced video signal, in the present invention the first rendered video signal and the one or more enhancement overlay streams can be processed entirely separately. This means that the present invention finds particular benefit when rendering a representation of a reference video signal which is subject to security protections, such as Conditional Access (CA) or Digital Rights Management (DRM).

[0018] This invention finds particular utility when rendering a representation of a reference video signal which has been encoded according to LCEVC techniques, in which case the first rendered video signal corresponds to the rendered base video signal. In conventional approaches to decoding a video encoded according to LCEVC techniques, the enhanced video signal would be generated frame by frame by combining each frame of the base video signal with corresponding frames of one or more enhancement streams. However, as this requires access to the frames of the base video signal, this approach may not be possible if the base video signal is subject to security protections. In contrast, the present invention allows for the enhancements to be applied without access to the base video signal.

[0019] The step of obtaining the first rendered video signal preferably comprises: obtaining a first decoded video stream; and rendering the first rendered video signal in a first markup video display region based on the first decoded video stream.

[0020] More preferably, obtaining the first decoded video stream comprises: receiving a first encoded video stream; and, in a first decoding process, decoding the first encoded video stream to obtain the first decoded video stream. The first decoding process may be performed in accordance with a standard coding scheme, such as H.264. However, the security protections associated with protected content may limit the decoding processes which may be used to obtain the first rendered video signal.

[0021] Advantageously, the present method is used to render a first encoded video signal in the form of an HTMLVideoElement. As such, the first markup video display region may be an HTML <video> element.

[0022] The step of overlaying the one or more enhancement overlay streams preferably comprises: obtaining one or more decoded enhancement streams; and rendering the one or more enhancement overlay streams in a second markup video display region based on the one or more decoded enhancement streams, the second markup video display region being coincident with the first markup video display region, which is to say that the second markup video display region is overlayed on the first markup video display region and the first and second markup video display regions are the same size. As will be appreciated, the first rendered video signal fills the first markup video display region and the one or more enhancement overlay streams fill the second markup video display region, such that one or more enhancement overlay streams are aligned with the first rendered video signal.

[0023] More preferably, obtaining the one or more decoded enhancement streams comprises: receiving one or more encoded enhancement streams; and, in a second decoding process: decoding the one or more encoded enhancement streams to obtain the one or more decoded enhancement streams.

[0024] The one or more encoded enhancement streams may be received as part of the same data stream as the first encoded video stream or may be received in one or more separate data streams.

[0025] The first decoding process and the second decoding process are, in most embodiments, performed according to different video decoding schemes. Further, the first and second decoding processes are preferably performed by different decoding elements. For example, if the method is implemented within a browser, the first and second decoding processes may be performed by separate elements of the browser implementation. In one such implementation, the first decoding process may be performed by an HTML media element while the second decoding process may be performed by an enhancement stream decoder.

[0026] The second markup video display region may be an HTML <canvas> element. This is preferable when the first markup video display region is an HTML <video> element.

[0027] As noted above, the present invention finds particularly utility when security protections used to prevent the unauthorised capturing of content restrict access to data in the first rendered video signal. As such, the first rendered video signal may comprise data that is inaccessible to the second decoding method. That said, the second decoding method may be based on other data obtained from the first rendered video signal, such as metadata including a playback time of the rendered video signal. The one or more encoded enhancement streams may comprise one or more sets of encoded residuals and decoding the one or more encoded enhancement streams may comprise decoding the one or more sets of encoded residuals to obtain one or more sets of decoded residuals. The one or more sets of decoded residuals may then be used to render the one or more enhancement overlay streams.

[0028] In particular, the one or more encoded enhancement streams may comprise at least a second set of encoded residuals in a second enhancement sublayer of a hierarchical coding scheme. As noted above, the present invention may be used to render a video which has been encoded using LCEVC techniques. As will be understood, the LCEVC coding scheme is based on using a base stream and first and second enhancement streams to represent a reference video signal. However, a representation of the reference video signal may be rendered based on only the base stream and the second enhancement stream. The present invention is compatible with this approach.

[0029] In preferred embodiments, at least one of the one or more sets of decoded residuals comprises positive and negative residuals and decoding the one or more encoded enhancement streams comprises: obtaining one or more decoded positive residual streams from the positive residuals; and obtaining one or more decoded negative residual streams from the negative residuals.

[0030] In still more preferred embodiments, rendering the one or more enhancement overlay streams comprises: rendering one or more positive residual overlay streams in the second markup video display region based on the decoded positive residual streams; and rendering one or more negative residual overlay streams in the second markup video display region based on the decoded negative residual streams.

[0031] For example, the decoded residuals may represent modifications to the luma components of a video signal, with positive residuals indicating that the brightness of a pixel should be increased and negative residuals indicating that the brightness of a pixel should be decreased. In order for the one or more enhancement overlay streams to implement these positive and negative residuals most effectively, it is preferable to separate the residuals into separate decoded positive and negative residual streams which may then be used separately to render positive and negative residual overlay streams in the second markup video display region.

[0032] While there may, in some embodiments, only be one positive residual overlay stream and one negative residual overlay stream, there could be any number of positive and negative residual overlay streams. For example, there could be two pairs of positive and negative residual overlay streams, each pair corresponding to an enhancement sublayer of a hierarchical coding scheme, such as LCEVC. There could also be more or fewer positive residual overlay streams than negative residual overlay streams.

[0033] In some embodiments, at least one positive residual overlay stream of the one or more positive residual overlay streams may encode modifications to the luma components of the rendered base video, said at least one positive residual overlay stream comprising an array of values of between 0 and 1 , each of said values corresponding to a pixel of the second rendered video signal, and rendering said at least one positive residual overlay stream in the second markup video display region comprises, for each pixel of the second rendered video signal, rendering a white pixel having an opacity equal to the corresponding value of the array of values.

[0034] In further embodiments, at least one negative residual overlay stream of the one or more negative residual overlay streams may encode modifications to the luma components of the rendered base video, said at least one negative residual overlay stream comprising an array of values of between 0 and 1 , each of said values corresponding to a pixel of the second rendered video signal, and rendering said at least one negative residual overlay stream in the second markup video display region comprises, for each pixel of the second rendered video signal: averaging the corresponding value from the array of values with the luma value of the underlying pixel in the first rendered video signal, and rendering a pixel having said averaged luma value. Advantageously, overlaying the one or more enhancement overlay streams may comprise synchronizing the one or more enhancement overlay streams with the first rendered video signal based on a playback time of the first rendered video signal. Preferably, each of the one or more enhancement overlay streams comprises a series of frames, the series of frames being indexed using a corresponding series of timestamps; and, for each of the one or more enhancement overlay streams, synchronizing said enhancement overlay stream with the first rendered video signal comprises: comparing a current playback time of the first rendered video signal with the timestamps indexing the frames of said enhancement overlay stream; identifying the frame of said enhancement overlay stream corresponding to the current playback time based on said comparison; and overlaying the frame of said enhancement overlay stream corresponding to the current playback time on the first rendered video signal. Said timestamps may be obtained from the first rendered video signal. In embodiments in which the enhancement overlay streams are derived from one or more encoded enhancement streams, the timestamps may be derived from the one or more encoded enhancement streams, in addition or as an alternative to obtaining the timestamps from the first rendered video signal.

[0035] Comparing the current playback time of the first rendered video signal with the timestamps indexing the frames of said enhancement overlay stream may comprise: searching for a frame of said enhancement overlay stream that has a timestamp that falls within a range defined with reference to the current playback time. This range may be set based on a configurable drift offset, which may in turn be based on a framerate of the first rendered video signal. In this case, searching for a frame of said enhancement overlay stream that has a timestamp that falls within a range defined with reference to the current playback time may comprise: searching for a frame of said enhancement overlay stream where the timestamp of the frame plus the configurable drift offset equals the current playback time.

[0036] The second rendered video signal represents the reference video signal at a second level of quality, which is a higher level of quality than the first level of quality. For example, the second rendered video signal may at a higher resolution than the first rendered video signal, and / or the second rendered video signal may have a higher dynamic range than the first rendered video signal. The second rendered video signal may, for example, be an HDR video while the first rendered video signal may be an SDR video.

[0037] The one or more of the one or more enhancement overlay streams may encode modifications to the luma and / or chroma components of the rendered base video.

[0038] In some embodiments, the method may advantageously be implemented within a browser.

[0039] According to a second aspect of the invention, a system for rendering a representation of a reference video signal may be provided, the system comprising: a video Tenderer configured to obtain a first rendered video signal representing the reference video signal at a first level of quality; and an enhancement decoder configured to overlay one or more enhancement overlay streams on the first rendered video signal to generate a second rendered video signal representing the reference video signal at a second level of quality, the second level of quality being a higher level of quality than the first level of quality.

[0040] The system provided herein may be configured to implement a method according the first aspect of the invention.

[0041] According to a third aspect of the invention, a computer-readable medium may be provided, the computer-readable medium comprising instructions which, when executed, cause a processor to perform a method according to the first aspect of the invention.

[0042] BRIEF DESCRIPTION OF DRAWINGS

[0043] The invention will now be described with reference to the figures, in which:

[0044] Figure 1 are a schematic diagram illustrating how encoded video data may be transported within various data streams; Figures 2A and 2B are schematic diagrams illustrating conventional systems for receiving and decoding multiple streams of encoded data;

[0045] Figures 3A and 3B are schematic diagrams illustrating example browser implementations of a system for receiving and decoding multiple streams of encoded data;

[0046] Figures 4A and 4B are schematic diagrams illustrating example browser implementations of a system for receiving and decoding multiple streams of encoded data provided according to embodiments of the present invention;

[0047] Figure 5 is a flow diagram showing an example method of decoding a multi-layer video stream according to embodiments of the present invention; and

[0048] Figure 6 is a schematic diagrams showing an example multilayer encoder configuration.

[0049] DETAILED DESCRIPTION

[0050] Certain examples described herein allow decoding devices to be easily adapted to handle multi-layer video coding schemes. Certain examples are described with reference to an LCEVC multi-layer video stream, but the general concepts may be applied to other multi-layer video schemes including SVC and SHVC, as well as multi-layer watermarking and content delivery schemes. Certain examples described herein are particularly useful in cases where a first layer of a multi-layer video stream is decoded by a first-layer decoder, which implements a first decoding method, and a second layer of the multi-layer video stream is decoded by a second-layer decoder, which implements a second decoding method. To allow flexibility in the multi-layer configuration, the first layer may be encoded using a variety of video coding methods, such as H.264 and H.265 as described above, as well as new and / or yet to be implemented video coding methods such Versatile Video Coding (WC or H.266). Hence, the first-layer decoder may vary for different encoded video streams. In example described herein, support is provided for fully encapsulated first-layer decoders that restrict access to internal data. For example, a first-layer decoder may comprise a hardware component such as a secure hardware decoder chipset where other processes within a client device performing the decoding cannot access data supplied in packets for the first layer. Further, protections applied to the data in the first layer may apply to the decoded output of the first layer decoder, for example by preventing access to the individual frames of a rendered video. Examples described herein thus allow data for a second layer of the multi-layer video stream to be decoded separately using a different decoding method but then combined with the appropriate output of the first-layer decoder, e.g. to provide an enhanced video output.

[0051] In particular, examples described herein enable an output of a second-layer decoder to be combined with an output of a first-layer decoder, where the first- layer decoder uses data that is inaccessible to the second-layer decoder. For example, whereas conventional approaches to combining the output of a second layer decoder with the output of a first layer decoder may involve combining a decoded frame of video data from the first-layer decoder with one or more decoded frames from the second-layer decoder, the second-layer decoder may not have access to the data of individual frames of the video data from the first- layer decoder. The present disclosure provides techniques for addressing these limitations. The present examples may be beneficial in cases where the second- layer comprises residual data, watermarking data, and / or localised embedded metadata.

[0052] In certain variations, the second-layer decoded data may have two or more sublayers at two or more resolutions, e.g. spatial resolutions. In these cases, the matched first-layer decoded data, with or without correction at the decoded resolution, may be upsampled to a higher resolution to provide enhancement via the second layer. In some implementations, the first-layer decoded data may also be generally available to output processes as well as the combined reconstruction, thus providing different options for viewing. For example, a video may be rendered by the first-layer decoder and presented without enhancements. In certain examples described herein, different layers of a multi-layer video coding may be transmitted as separate packets that are multiplexed within a transport stream. This allows different layers to be effectively supplied separately and for enhancement layers to be easily added to pre-existing or pre-configured base layers. At a decoding device, different packet sub-streams may be received and parsed, e.g. based on packet identifiers (PI Ds) within packet headers.

[0053] In the description below, a first example of an encoded video stream is described with reference to Figure 1 . Examples of conventional systems and methods of decoding a multi-layer video stream are then described with reference to Figures 2A, 2B, 3A, and 3B. Examples of systems and methods of decoding a multi-layer video stream according to embodiments of the present invention are then described with reference to Figures 4 and 5. An example of a specific multi-layer coding scheme is then described with reference to Figure 6.

[0054] Figure 1 shows an example 100 of a T ransport Stream (TS) 102 that may be used to transmit encoded video data to one or more decoding devices. The Transport Stream 102 comprises a sequence of fixed-length 188-byte TS packets 110. Each TS packet 110 has a header 112, which may have a variable length, and a payload 114. The header 112 includes one or more data fields. One of these data fields provides a Packet Identifier (PID) 116. The PID is used to distinguish different substreams within the T ransport Stream. The PID may be a number of bits of a fixed length (e.g., 13 bits) that stores a numeric or alphanumeric identifier (typically represented as hexadecimal value). For example, the PID 116 may be used to identify different video streams that are multiplexed together into a single stream that forms the Transport Stream 102. An example Transport Stream specification is set out in MPEG-2 Part 1 and defined as part of ISO / IEC standard 13818-1 or ITU-T Rec. H.222.0.

[0055] Figure 1 also shows one of the so-called PID streams 104 that may be extracted from the Transport Stream 102. The PID stream 104 comprises a stream of consecutive packets 110 that have a common (i.e., shared) PID value. The PID stream 104 may be created by demultiplexing the Transport Stream 102 based on the PID value. The PID stream 104 thus represents a sub-stream of the Transport Stream 102.

[0056] In certain cases, there may be special PID values that are reserved for indexing tables. In one case, one PID value may be reserved for a program association table (PAT) that contains a directory listing of a set of program map tables, a program map table (PMT) comprising a mapping between one or more PID values and a particular “program”. Originally a “program” related to a particular broadcast program but with Internet streaming, the term is used broadly to relate to the content of a particular video stream.

[0057] Figure 1 also shows a Packetised Elementary Stream (PES) 106 that is constructed based on the payload data of a plurality of TS packets 110. A PES comprises data from payloads of a PID stream that carries media sample data. Media sample data may comprise video data as well as other modalities, such as audio data, subtitle data, or volumetric data. In Figure 1 , a PES is generated by combining the payloads 114 of multiple media TS packets 110 that are associated with a common (i.e., shared) PID value. The PES comprises a packet stream where each PES packet consists of a header (i.e., a PES Header) 122 and a payload 124, the payload 124 carrying the combined data. The start of a new PES packet is indicated by a one-bit field from the TS header 112, called a Payload Unit Start Indicator (PUSI) 118. When the PUSI is set, the first byte of the TS packet payload 114 indicates where a new PES payload unit starts. This allows a decoding device that starts receiving data mid-transmission to determine when to start extracting data. The PES header 122 contains a Presentation Time Stamp (PTS) 128. This indicates a time of presentation for the corresponding piece of media encapsulated within the payload 124.

[0058] Figure 1 lastly shows the contents of the PES payload 124 for a video stream. In this case, the PES payload 124 comprises a sequence 108 of NAL units 130 (i.e., a NALU stream). These may form part of an Access Unit for the video stream, i.e. a set of NAL units that are associated with a particular output time, are consecutive in decoding order, and contain a coded picture or frame. Figure 1 shows a NALU stream 108 that may be provided to a suitable video decoder for decoding. The example of Figure 1 shows how different layers of a multi-layer video stream, e.g. an encoded video stream, may be communicated to decoding devices. The Transport Stream 102, for example, may be transmitted over one or more communication channels, including over-the-air transmissions as well as network communications. In the example of Figure 1 , different layers of a multi-layer video stream may be delivered as different PID sub-streams, e.g. packets with a PI D of “B” may represent encoded data for a first layer (“base”) stream, whereas packets with a PID of “L” may represent encoded data for a second layer (“LCEVC” or enhancement) stream. The different sub-streams for each layer may be demultiplexed and provided as a PES 106 to a video decoder. Although Figure 1 shows an example Transport Stream, other digital media containers, such as “tracks” on computer-readable media such as discs or solid-state storage, may also be used to provide encoded layer data. In certain cases, encoded data for different layers may be read using a client device file system (e.g., as different tracks from a provided medium).

[0059] Before describing examples in which access to data in a first layer stream is restricted, it will be instructive to first describe the conventional techniques used in cases where there are fewer such restrictions.

[0060] Figure 2A shows an example of a conventional system 200 for decoding a multilayer video stream. According to this example, the access to data in a first layer stream is not restricted, and first-layer decoded data may be enhanced by pairing first-layer decoded data for a particular frame with second-layer decoded data for that frame.

[0061] The system 200 is typically implemented as part of a client computing device, such as a smartphone, laptop, smart television, or other media receiver and / or player. The multi-layer video stream encodes a video signal and comprises at least a first layer and a second layer. The first layer may comprise an encoding according to a first coding method, scheme, or standard, such as one of H.264, H.265 or H.266, amongst others. The first layer may be referred to as a “base” layer. The first layer may be encoded according to a first level of quality, such as a first spatial resolution, a first level of quantisation, a first specified or desired bit rate, and / or a first temporal resolution. The first layer may be a complete video encoding, i.e. encoded data may be received, decoded, and rendered regardless of the presence of further layers. The second layer may comprise an encoding according to a second coding method, scheme, or standard. For example, the second layer may comprise an “enhancement” layer for enhancing or otherwise augmenting the “base” layer. The second layer may be encoded using an enhancement coding method such as LCEVC. The second layer may comprise an encoding of a residual data stream, e.g. for combination with the first layer. The second layer may be encoded according to a second level of quality, such as a second spatial resolution, a second level of quantisation, a second specified or desired bit rate, and / or a second temporal resolution. The second level of quality may be higher than the first level of quality, e.g. to provide an enhancement. Each layer may be received as a series of NAL units, such as 108 in Figure 1 , where the NAL units comprise encoded data for the layer for a particular frame. Each layer may comprise a sequence of encoded data for consecutive frames. In certain cases, the sequence of encoded data may relate to a group of pictures.

[0062] In Figure 2A, the system 200 receives a first layer stream 202 and a second layer stream 204. The term “stream” is used herein to refer to consecutive portions of data that are received or accessed. The first- and second-layer streams 202, 204 may result from a demultiplexing operation, such as an operation performed on Transport Stream 102 to extract one of PID stream 104 or PES 106, or may comprise data read from one or more files.

[0063] In the system 200 of Figure 2A, the first layer stream 202 is provided to a first layer video decoder 212 and the second layer stream 204 is provided to a second-layer decoder 214. The first layer video decoder 212 is configured to decode the first layer of the multi-layer video stream and the second layer video decoder 214 is configured to decode the second layer of the multi-layer video stream. The first layer video decoder 212 may be implemented using a video codec. The video codec may be hardware and / or software based. In certain cases, the first layer video decoder 212 may use hardware acceleration, wherein one or more actions performed as part of the decoding are implemented using a specifically configured hardware device (such as a particular video decoding chipset). In certain cases, the first layer video decoder 212 may be implemented using one or more operating system services, e.g. functions provided as part of an operating system such as iOS®, Windows®, or Linux®. The second layer video decoder 214 may be an LCEVC decoder, e.g. a software decoder implemented according to the LCEVC standard.

[0064] In Figure 2A, the second layer video decoder 214 is communicatively coupled to a memory 216. The memory 216 is configured to store an output of the second- layer decoder 214. The memory 216 may comprise a dedicated hardware buffer and / or a portion of shared system memory reserved for video decoding. The second layer video decoder 214 is configured to decode frames of data 224 for the second layer and store these in the memory 216. If the second layer video decoder 214 is configured to receive and decode one or more sub-layers of residual data (as in LCEVC), the frames of data 224 may comprise one or more frames at one or more respective resolutions.

[0065] Figure 2A also shows a decoding controller 230. The decoding controller 230 may form part of a decoder integration layer that is configured to control the decoding of a multi-layer video, e.g. in association with the first layer video decoder 212 and the second layer video decoder 214. In one case, the first layer video decoder 212 may be implemented as an independent video decoder (e.g., an independent codec) for the first layer (e.g., that is able to decode first layer data in the absence of second layer data) and the decoding controller 230 and the second layer video decoder 214 may be implemented as part of a multi-layer decoder that is configured to enhance the first layer with one or more additional layers of video encoding. The decoding controller 230 may be implemented in software (e.g., as executed by a processor of a client device) and / or using dedicated hardware. In one case, the decoding controller 230 and the second layer video decoder 214 may be implemented as a software (including firmware) enhancement to an existing or legacy video decoding system (including those with hardware acceleration for the first layer). One example decoding system that may be adapted to provide the functionality of the second layer video decoder 214 and the decoding controller 230 is described in PCT / GB2021 / 051940, which is incorporated by reference herein.

[0066] In Figure 2A, the decoding controller 230 is communicatively coupled to the first- layer decoder212. The decoding controller 230 is configured to receive a call back 222 indicating an availability of first-layer decoded data from the first-layer decoder 212 for a frame of the first layer. For example, the decoding controller 230 may request a call back from the first-layer decoder 212 whenever decoded data for the first layer is ready for rendering. The term “call back” is used herein to refer to a communication or signal that is sent between hardware and / or software components to indicate an event has occurred. In hardware, a call back may comprise a physical signal sent over a communication bus or channel and / or a change in a register value (e.g., representing a binary flag). In software, a call back may comprise an asynchronous function return and / or a change in a monitored value in memory. The call back 222 may comprise a call back that is used, in comparative cases, for rendering an output of the first-layer decoder 212. For example, the call back 222 may indicate that a frame of video encoded using the first layer has been decoded from the first layer stream 202 and is ready for display as part of a rendered video. In certain cases, the call back 222 may comprise the data for the decoded frame (e.g., Frame Layer 1 - FL1); in other cases, the call back 222 may comprise a reference that indicates where the decoding controller 230 may access the decoded data, such as a memory address.

[0067] On receipt of the call back 222, the decoding controller 230 is configured to further obtain timing metadata 232 for the first-layer decoded data, e.g. the decoded first layer frame. The timing metadata 232 is associated with a rendering of the frame for the first layer. The timing metadata 232 may be generated by the first-layer decoder 212. For example, the timing metadata 232 may be generated to help a downstream process render or otherwise display the first-layer decoded data. In one case, the timing metadata 232 may comprise one of a media time for the frame derived from the first-layer decoded data or a current playback time for the frame derived from the first-layer decoded data. The timing metadata 232 may be provided as part of the call back 222, and / or may be accessible with the first-layer decoded data, e.g. from a memory address associated with a memory address for the first-layer decoded data.

[0068] Following receipt of the call back 222, and having obtained the timing metadata 232, the decoding controller 230 is configured to compare the timing metadata with one or more timestamps for the output of the second-layer decoder to pair the first-layer decoded data for the frame with second-layer decoded data for the frame. This is possible in conventional techniques because access is given to data in the first layer stream 202.

[0069] In the example of Figure 2A, the decoding controller 230 performs a query 236 on the second-layer decoded data 224 that is stored in the memory 216. In particular, the decoding controller 230 looks for a match between a timestamp stored with each portion of the second-layer decoded data, such as each decoded frame of second layer data, and the timing metadata 232, e.g. a media or playback time, associated with the ready from of first-layer decoder data. For example, the decoding controller 230 may be configured to search for second-layer decoded data that has a timestamp that falls within a range defined with reference to a time indicated by the timing metadata. This range may be set based on a configurable drift offset. In this case, the decoding controller 230 may search for data within the memory where the timestamp of the data plus a small drift offset equals a time indicated within the timing metadata. The configurable drift offset may be a small number of milliseconds (e.g., 10ms) and may be set based on a framerate of the rendering (e.g., may be reduced for higher framerates).

[0070] In Figure 2A, the decoding controller 230 retrieves second-layer decoded data 234 based on the comparison and combines this with the first-layer decoded data 222 to output a reconstruction 238 of the frame of the video signal. The reconstruction 238 may be an enhanced frame whereby second-layer decoded data 234 in the form of residual data for the frame is combined with a decoded frame for the first layer. The decoding controller 230 may apply one or more sublayers of enhancement based on the second-layer decoded data 234, e.g. the second-layer decoded data 234 may comprise two sub-layers of residual data as described later with reference to Figure 6. In certain configurations, if no second- layer decoded data 234 is located, e.g. because there has been an issue with receipt of data for the second layer stream or network congestion, then the decoding controller 230 may simply output the first-layer decoded data 222 (e.g., without enhancement), e.g. act in a pass-through mode. The output of the decoding controller 230 may be rendered on a display forming part of, or communicatively coupled to, a client device implementing system 200. In other cases, the output of the decoding controller 230 may be made available to other processes, e.g. in memory 216 or another buffer.

[0071] Figure 2B shows another configuration 240 of the conventional system 200. In the configuration of the system of Figure 2A, each of the first-layer video decoder 212, the second-layer video decoder 214, and the decoding controller 230 access data stored within the memory 216. As with the configuration of the system of Figure 2A, the access to data in a first layer stream is not restricted, and first-layer decoded data may be enhanced by pairing first-layer decoded data for a particular frame with second-layer decoded data for that frame.

[0072] In this case, the decoding controller 230 is configured to receive call backs from both the first-layer video decoder 212 and the second-layer video decoder 214. The first-layer video decoder 212 sends a first call back 232 to the decoding controller 230 to indicate a new frame of first layer data 242 is ready. The second- layer video decoder 214 sends a call back 252 to the decoding controller 230 to indicate a new frame of second layer data 224 is ready. Both the first-layer video decoder 212 and the second-layer video decoder 214 may be configured to buffer decoded frames within the memory 216. Although the memory 216 is shown as a shared memory in Figure 2B, it may comprise separate memories that are accessible by the appropriate components (e.g., dedicate hardware or software frame buffers for each decoder).

[0073] As with Figure 2A, the frames of first layer data 242 are indexed using timing metadata values, which in Figure 2B are shown as a media playback time tMPB, and the frames of the second layer data 224 are indexed using the timestamps from the second layer stream 204, e.g. the PTS values. In Figure 2B, the decoding controller 230 is configured to coordinate receipt of the call backs 232 and 252 and compare timing metadata and timestamp values to locate corresponding frames of the first layer data 242 and the second layer data 224. In certain cases, the decoding controller 230 may act conditionally on receipt of both call backs 232 and 252. Following the comparison, the corresponding frames are then combined, e.g. either by the decoding controller 230 or via arithmetic performed in memory 216, to output the multi-layer reconstruction 238. In Figure 2B, the multi-layer reconstruction 238 is then available for rendering, e.g. may be copied to a dedicated frame buffer of a display device. As described above, the frames of the first layer data 242 and the second layer data 224 may be matched by looking for timing metadata values that equal the timestamp values plus a configurable drift offset, e.g. tMPB = PTS + offset.

[0074] Figure 3A is a schematic diagram of another conventional system 300 for decoding an encoded multi-layer video stream, such as an encoded multi-layer video stream encoded using LCEVC. The system 300 may implement one of the systems 200 and 240 in Figures 2A and 2B within a browser. The browser may be any browser capable of accessing information on the World Wide Web, examples of which include, but are not limited to, Google Chrome®, Microsoft Edge®, Safari®, Firefox® and Opera®. The browser may be implemented in a client device. Example client devices include, but are not limited to, mobile devices, computing devices, tablet devices, smart televisions, and so on. The client device typically comprises an operating system (OS) and the OS comprises the browser.

[0075] One function of the browser is to transform documents written in a markup scripting language (sometimes referred to as a markup language) into a visual representation of a webpage. The markup scripting language is used to control a display of data in a rendered webpage. The markup language may include a markup video element which in turn becomes a video display region when processed by the browser. For example, a user of the browser may navigate to a web page that includes an embedded video. When the browser renders the webpage, it receives data corresponding to the video. The browser may include resources necessary to decode and playback the video, so as to display the video to the user within a video display region rendered by the browser on a display of a client device, for example. Examples of a markup scripting language include any versions of Hypertext Markup Language (HTML), such as HTML5, and Extensible HyperText Markup Language (XHTML).

[0076] The markup video element, for example, indicates properties associated with display of the video in the webpage, such as the size of the video within the webpage and whether the video will autoplay upon loading of the webpage. The markup video element may also include an indication of the video coding format used to encode the video. This indicates to the browser which decoder(s) to use to decode the encoded video. The browser may then perform a call to at least one of a decoding function within the resources of the browser itself (which may be considered browser-native resources, which are native to the browser), or to a decoding function implemented in the OS, as discussed further below.

[0077] The system 300 of Figure 3A comprises a source buffer 302 to receive an encoded multi-layer video stream. In this example, the source buffer 302 may receive, and be the source of, the first layer stream 202 and the second layer stream 204 in Figures 2A and 2B. In other examples, each stream may have a separate corresponding source buffer rather than a joint buffer for both streams. The source buffer is a section of memory, which is for example accessible to the browser. The source buffer may be a Media Source Extensions (MSE) application programming interface (API) SourceBuffer, for example. The encoded multi-layer video stream in this example comprises an encoded base stream and an encoded enhancement stream. The encoded base stream comprises video content encoded by any base encoder, also known as a compressor, such as an Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), VP9, MPEG-5 Essential Video Coding (EVC), orAOMedia Video 1 (AV1) encoder.

[0078] The system 300 further comprises an HTML media element 304, which implements a first layer (e.g., base stream) video decoder, also known as a base stream decompressor. The HTML media element 304 may instruct or otherwise implement a first layer video decoder 212 as described with reference to Figures 2Aand 2B. In this example, the encoded base stream is extracted from the source buffer 302 and decoded using the HTML media element 304. The HTML media element 304 is a markup video element that provides an interface between the markup language and decoding resources. HTML5, for example, includes a markup video element that can be used to embed video content in a webpage. Another example is a JavaScript library that builds a custom set of controls over top of the HTML5 video element, which may be considered to function as a JavaScript player. It is to be appreciated that a markup video element such as the HTML5 video element can be modified by adding additional resources, such as a multi-layer video stream decoding library, a WebAssembly decoding library and / or a web worker function that can be accessed by the markup video element.

[0079] In a case, where LCEVC is used, the enhancement stream may be carried within a set of Supplemental Enhancement Information (SEI) messages that accompany and are associated with the base stream or within a separate network abstraction layer (NAL) unit stream, e.g. as carried within a PID stream as shown in Figure 1 . Base stream decoders may be configured to ignore SEI messages or a NAL unit stream if they contain information they cannot interpret, such as header information indicating a particular packet type. Hence, in this case, the HTML media element 304 may retrieve data relating to the base stream from the source buffer 302 in a default manner, wherein both enhanced and non-enhanced base streams are processed in a common manner. In this case, the HTML media element 304 may ignore SEI messages or NAL units that carry the enhancement stream that reside within the source buffer 302.

[0080] In certain implementations of the example system 300 of Figure 3A, the markup video element includes an indication to the video coding format associated with the encoded base stream. The markup video element, when processed by the browser, locates the appropriate base stream decoder associated with the video coding format, and decodes the encoded base stream. The HTML media element 304 may be implemented within the browser, using functionality of the OS of a client device comprising the browser, or utilising resources of both the browser and the OS. For example, the OS may utilise hardware acceleration to decode the encoded base stream which can reduce power consumption and the number of computations performed by a CPU compared to software-only decoding.

[0081] The decoded base stream is rendered in a first markup video display region 306. The first markup video display region 306, for example, corresponds to a region of the webpage at which it is desired to display a video. The first markup video display region 306 may comprise a markup <video> element, such as an HTML <video> element. In this conventional approach, the rendering of the decoded base stream allows access to the base stream video data, e.g. decoded frames of the base encoded video. Even if the decoding of the base stream is performed by an inaccessible or protected method, rendering the base stream video data makes, in this conventional approach, the base stream video data accessible to other decoding processes within the browser. For example, the HTML media element 304 may provide an option of registering a call-back when a frame of first- layer decoded data is ready, e.g. as indicated by 232 in Figures 2A and 2B. However, as will be discussed further, this approach relies on access to the base stream video data being made available by the rendering of the decoded base stream. In particular, enhancement of the base stream video data is based on access to individual frames of the base stream video.

[0082] The rendered decoded base stream is subsequently combined with a decoded enhancement stream to generate a reconstructed video stream. In certain cases, as the rendered base stream does not include enhancement data from the enhancement stream at this point, the first markup video display region is hidden. This ensures that the rendered video content corresponding to the base stream is not displayed in the webpage and so is not visible to a viewer of the webpage. However, in certain cases, there may be a user option to view this content. Rendering the decoded base stream ensures that the system 300 can still decode and render video streams that are not encoded using a multi-layer video coding format, e.g. if this is the case, the first markup video display region may be set as visible and the decoded base stream may be displayed as per comparative nonenhancement video rendering. For example, if the webpage included a singlelayer video stream that lacked an enhancement stream, the system 300 of Figure 3A could be used to display the decoded single-layer video stream, e.g. by unhiding the first markup video display region 306.

[0083] The system 300 further comprises an enhancement stream decoder 308. The enhancement stream decoder 308 may implement functionality of one or more of the second layer video decoder 214 and the decoding controller 230 described with reference to Figures 2A and 2B. In the present case, the source buffer 302, which may comprise a MSE component, issues a call-back to indicate that data is ready for decoding. This call-back may be received by the enhancement stream decoder 308 (e.g., the enhancement stream decoder 308 may register for this callback as part of an initial configuration). On receipt of the call-back, the encoded enhancement stream is extracted from the source buffer 302 and decoded by the enhancement stream decoder 308. For example, the enhancement stream decoder 308 may retrieve the encoded enhancement stream from data for a set of SEI messages or a portion of a PI D stream that is stored within the source buffer 302. In certain cases, the enhancement stream decoder 308 may obtain both the first layer and the second layer streams (i.e., base and enhancement streams), demultiplex the two streams and then discard the first layer stream (e.g., as this is being obtained and decoded by way of the HTML media element 304). At this stage, a timestamp may be obtained by way of sourcing the combined stream from the source buffer 302. For example, a PTS time stamp may be obtained from the base stream or from the source buffer call-back.

[0084] In the example of Figure 3A, the enhancement stream decoder 308 also obtains the decoded base stream from the first markup video display region 306 and combines the decoded base stream with the decoded enhancement stream to generate a reconstructed video stream. The reconstructed video stream may then be rendered in a second markup video display region 310 within the webpage that is visible to a viewer of the webpage. The second markup video display region 310 may comprise a markup <canvas> element, such as an HTML <canvas> element. In the present example, the enhancement stream decoder 308 receives a call-back from the first markup video display region 306 when a frame of the base (i.e., first layer) stream is ready. This may comprise a requestVideoFrame call-back. When the enhancement stream decoder 308 obtains the call-back from the first markup video display region 306, it obtains timing metadata for the base frame as discussed above. For example, the timing metadata may comprise a “media time” variable or a current playback time that is provided with the frame that is rendered in the first markup video display region 306. As described above, the enhancement stream decoder 308 then compares this timing metadata with the originally sourced timestamp to pair first-layer decoded data for the frame with second-layer decoded data for the frame. The combined paired data is then rendered in the second markup video display region 310 as an enhanced frame.

[0085] The enhancement stream decoder 308 may be a multi-layer video stream decoder plugin (DPI) such as an LCEVC decoder plugin, configured to decode an LCEVC- encoded video stream. The enhancement stream decoder 308 may provide the decoding capabilities of the second layer decoder 214 and the control capabilities of the decoding controller 230 as described with reference to Figures 2A and 2B. The enhancement stream decoder 308 may comprise a single component or two separate components depending on the implementation, with the same functional effect. One or more components of the system 300 may be implemented in a browser. In one example, a browser is provided comprising the enhancement stream decoder 308.

[0086] Figure 3B is a schematic diagram of another conventional system 300 for decoding an encoded multi-layer video stream, such as an encoded multi-layer video stream encoded using LCEVC. Components and features common to those of Figure 3A function in the same way as previously described and so will not be described again in detail.

[0087] As before, the system 300 of Figure 3B comprises a source buffer 302 to receive an encoded multi-layer video stream, for example the first layer stream 202 and the second layer stream 204 in Figures 2A and 2B. The encoded base stream is extracted from the source buffer 302 and decoded using the HTML media element 304. The decoded base stream is rendered in a first markup video display region 306. The source buffer 302 issues a call-back to indicate that data is ready for decoding and the call-back is received by the enhancement stream decoder 308. On receipt of the call-back, the encoded enhancement stream is extracted from the source buffer 302 and decoded by the enhancement stream decoder 308. A timestamp, e.g. a PTS time stamp, is obtained from the base stream or from the source buffer call-back.

[0088] The enhancement stream decoder 308 also obtains the decoded base stream from the first markup video display region 306 and combines the decoded base stream with the decoded enhancement stream to generate a reconstructed video stream. The reconstructed video stream is then be rendered in a second markup video display region 310.

[0089] As previously discussed, when LCEVC is used, the enhancement stream extracted from the source buffer 302 is carried within a set of SEI messages that accompany and are associated with the base stream or within a separate NAL unit stream.

[0090] Base stream decoders are often configured to ignore SEI messages or a NAL unit stream if these contain information the base stream decoder cannot interpret. This means that the HTML media element 304 retrieves data relating to the base stream from the source buffer 302 in a default manner, wherein both enhanced and non-enhanced base streams are processed in a common manner. The HTML media element 304 ignores SEI messages or NAL units that carry the enhancement stream that reside within the source buffer 302.

[0091] As seen in Figure 3B, the enhancement stream extracted from the source buffer 302, and carrying the set of SEI messages are first passed to an integration layer 312 before being passed to the enhancement stream decoder 308 for decoding. In particular, the integration layer 312 extracts LCEVC data from the enhancement stream based on the type of source buffer. The extracted data can be NAL units with LCEVC data and a realtime transport protocol (RTP) timestamp or data appended to the source buffer from media source extensions. The LCEVC data is indexed using the PTS and stored for use when a corresponding base stream is to be matched and combined with the LCEVC data using the PTS. As before this is done by comparing timing metadata with the PTS and then the matched LCEVC data can be combined with the base stream to generate a reconstructed video stream. As shown in Figure 3B, the relevant LCEVC data is passed to the enhancement stream decoder 308 along with an offset which is generated based on the environment of playback.

[0092] The offset is calculated by an offset calculation block 314 and can be done using two methods. The first method involves calculating the offset during the source buffer append and the second method involves calculating the offset during the enhancement stream decoding.

[0093] Using the first method, during the source buffer append, there are instances where data having the same timestamp gets appended to the source buffer more than once. In this case, the frame time is calculated based on the FPS of the enhancement stream and then this frame time is added to the provided timestamp the number of times the append was repeated. For example, if data was appended three times, the frame time is added to the provided timestamp three times. This is then used to store the LCEVC data used by the integration layer .

[0094] Using the second method, during the enhancement layer decoding, the offset is calculated based on one or more of the operating system, browser, player, and container format. The offset is then added to the timestamp provided by the video stream in order to fetch the relevant LCEVC data from the stored LCEVC data.

[0095] It will be understood from the foregoing description of the conventional systems 200 and 240 of Figures 2A and 2B that the rendering of a video based on the first layer stream 202 and the second layer stream 204 is based on access by the system 200 to the individual frames of the first layer stream 202 and second layer stream 204. Similarly, in the system 300 of Figures 3A and 3B the rendering of the decoded base stream allows access to the base stream video data, and specifically to the decoded frames of the base encoded video. However, the first layer stream may be subject to protections that restrict access to individual frames of the first layer stream as well as to individual frames of the video rendered from the first layer stream, such as to the decoded frames of the base encoded video in Figures 3A and 3B. These protections mean that it is difficult to use the approach taken in these conventional systems to apply enhancements to the first layer stream.

[0096] To address these issues, an alternative approach is used in embodiments of the present invention.

[0097] Figure 4A is a schematic diagram of an example system 400 for decoding an encoded multi-layer video stream, such as an encoded multi-layer video stream encoded using LCEVC, according to embodiments of the present invention. The system 400 is similar in many respects to the system 300 shown in Figures 3A and 3B and comprises corresponding elements to those shown in Figure 3A. System 400 may also be implemented within a browser. Indeed, in some embodiments the system 400 is also capable of implementing the conventional approaches discussed above.

[0098] As with the conventional approaches discussed above, one function of the browser is to transform documents written in a markup scripting language (sometimes referred to as a markup language) into a visual representation of a webpage. The markup scripting language is used to control a display of data in a rendered webpage. The markup language may include a markup video element which in turn becomes a video display region when processed by the browser. For example, a user of the browser may navigate to a web page that includes an embedded video. When the browser renders the webpage, it receives data corresponding to the video. The browser may include resources necessary to decode and playback the video, so as to display the video to the user within a video display region rendered by the browser on a display of a client device, for example. Examples of a markup scripting language include any versions of Hypertext Markup Language (HTML), such as HTML5, and Extensible HyperText Markup Language (XHTML). The markup video element, for example, indicates properties associated with display of the video in the webpage, for example the size of the video within the webpage and whether the video will autoplay upon loading of the webpage. The markup video element, for example, also includes an indication of the video coding format used to encode the video. This indicates to the browser which decoder(s) to use to decode the encoded video. The browser may then perform a call to at least one of a decoding function within the resources of the browser itself (which may be considered browser-native resources, which are native to the browser), or to a decoding function implemented in the OS, as discussed further below.

[0099] Similar to the system 300 of Figures 3A and 3B, the system 400 of Figure 4A comprises a source buffer 402 to receive an encoded multi-layer video stream. In this example, the source buffer 402 may receive, and be the source of, a first layer stream and a second layer stream. In other examples, each stream may have a separate corresponding source buffer rather than a joint buffer for both streams. The source buffer may be a section of memory, which may for example be accessible to the browser. The source buffer may be a Media Source Extensions (MSE) application programming interface (API) SourceBuffer, for example. The encoded multi-layer video stream may comprise an encoded base stream and an encoded enhancement stream, with the encoded base stream comprising video content encoded by any base encoder, also known as a compressor, such as an Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), VP9, MPEG-5 Essential Video Coding (EVC), or AOMedia Video 1 (AV1) encoder.

[0100] The system 400 further comprises an HTML media element 404, which implements a first layer (e.g., base stream) video decoder, also known as a base stream decompressor. The HTML media element 404 may instruct or otherwise implement a first layer video decoder, which may be similar to that described with reference to Figures 2A and 2B. In this example, the encoded base stream is extracted from the source buffer 402 and decoded using the HTML media element 404. The HTML media element 404 may be a markup video element that provides an interface between the markup language and decoding resources. HTML5, for example, includes a markup video element that can be used to embed video content in a webpage. Another example is a JavaScript library that builds a custom set of controls over top of the HTML5 video element, which may be considered to function as a JavaScript player. It is to be appreciated that a markup video element such as the HTML5 video element can be modified by adding additional resources, such as a multi-layer video stream decoding library, a WebAssembly decoding library and / or a web worker function that can be accessed by the markup video element.

[0101] Where LCEVC is used, the enhancement stream may be carried within a set of Supplemental Enhancement Information (SEI) messages that accompany and are associated with the base stream or within a separate network abstraction layer (NAL) unit stream, e.g. as carried within a PID stream as shown in Figure 1. Base stream decoders may be configured to ignore SEI messages or a NAL unit stream if they contain information they cannot interpret, such as header information indicating a particular packet type. Hence, in this case, the HTML media element 404 may retrieve data relating to the base stream from the source buffer 402 in a default manner, wherein both enhanced and non-enhanced base streams are processed in a common manner. In this case, the HTML media element 404 may ignore SEI messages or NAL units that carry the enhancement stream that reside within the source buffer 402.

[0102] In certain implementations of the example system 400 of Figure 4A, the markup video element includes an indication to the video coding format associated with the encoded base stream. The markup video element, when processed by the browser, locates the appropriate base stream decoder associated with the video coding format, and decodes the encoded base stream. The HTML media element 404 may be implemented within the browser, using functionality of the OS of a client device comprising the browser, or utilising resources of both the browser and the OS. For example, the OS may utilise hardware acceleration to decode the encoded base stream which can reduce power consumption and the number of computations performed by a CPU compared to software-only decoding.

[0103] The decoded base stream is rendered in a first markup video display region 406.

[0104] The first markup video display region 406, for example, corresponds to a region of the webpage at which it is desired to display a video. The first markup video display region 406 may comprise a markup <video> element, such as an HTML <video> element. However, whereas in the example described above with reference to Figure 3A the rendering of the decoded base stream allows access to the base stream video data, such as to decoded frames of the base encoded video, in the system 400 of Figure 4A access to the base stream video data may be restricted. This in turn means that, rather than combining the rendered decoded base stream with a decoded enhancement stream to generate a reconstructed video stream, a different approach is taken.

[0105] As the rendered base stream is not directly combined with enhancement data, the first markup video display region is not hidden, in contrast to the system 300 of Figures 3A and 3B. This in turn means that the rendered video content corresponding to the base stream may be visible to a viewer of the webpage prior to the application of any enhancements. However, similar to the system 300 described above with reference to Figures 3A and 3B, the system 400 of Figure 4A may be used to decode and render video streams that are not encoded using a multi-layer video coding format. For example, if a webpage included a singlelayer video stream that lacked an enhancement stream, the system 400 of Figure 4A could be used to display the decoded single-layer video stream.

[0106] The system 400 further comprises an enhancement stream decoder 408. The enhancement stream decoder 408 may implement functionality of one or more of a second layer video decoder and a decoding controller, each of which may be similar in function to those described with reference to Figures 2A and 2B. In the present case, the source buffer 402, which may comprise a MSE component, issues a call-back to indicate that data is ready for decoding. This call-back may be received by the enhancement stream decoder 408 (e.g., the enhancement stream decoder 408 may register for this call-back as part of an initial configuration). On receipt of the call-back, the encoded enhancement stream is extracted from the source buffer 402 and decoded by the enhancement stream decoder 408. For example, the enhancement stream decoder 408 may retrieve the encoded enhancement stream from data for a set of SEI messages or a portion of a PI D stream that is stored within the source buffer 402. In certain cases, the enhancement stream decoder 408 may obtain both the first layer and the second layer streams (i.e. , base and enhancement streams), demultiplex the two streams and then discard the first layer stream (e.g., as this is being obtained and decoded by way of the HTML media element 404). At this stage, a timestamp may be obtained by way of sourcing the combined stream from the source buffer 402. For example, a PTS time stamp may be obtained from the base stream or from the source buffer call-back. This time stamp allows a playback time of the rendered base video to be identified, which can then be used to synchronize enhancements with the base video.

[0107] In the case that enhancements are to be applied to the rendered base video, the enhancement stream decoder 408 receives the encoded enhancement stream, decodes this to obtain a decoded enhancement stream, and renders one or more enhancement overlay streams based on the decoded enhancement stream. Whereas in the conventional approach described above with reference to Figure 3A the enhancements are first combined with the base stream before being displayed, in the system 400 of Figure 4Athe one or more enhancement overlay streams are rendered in a second markup video display region 410, the second markup video display 410 region being coincident with the first markup video display region 406, to generate an enhanced rendered video. For example, the first markup video display region 406 may comprise a markup <video> element, such as an HTML <video> element, and the second markup video display region 410 may be a markup <canvas> element, such as an HTML <canvas> element, used to overlay the one or more enhancement overlay streams on the base video displayed in the first markup video display region 406. As will be appreciated, the enhancement stream decoder 408 may also receive more than one encoded enhancement stream and decode these to obtain more than one decoded enhancement stream.

[0108] Thus, whereas in the example of Figure 3A the enhancement stream decoder 308 obtains decoded base video data from the first markup video display region 306 and combines this with enhancement data to generate an enhanced rendered video, in the embodiment of Figure 4B the enhancement stream decoder 408 renders one or more enhancement overlay streams in a second markup video display region 410 and the enhanced rendered video is obtained as a result of this second markup video display region 410 being coincident with the first markup video display region 406. However, the enhancement stream decoder 408 may obtain metadata from the first markup video display region 406, such as a playback time of the rendered base video which may be used to syncrhonise the one or more enhancement overlay streams with the rendered base stream.

[0109] The one or more enhancement overlay streams may encode modifications to the luma and / or chroma components and may be at a higher resolution than the base video. For example, the one or more enhancement overlay streams may encode positive or negative modifications to the luma components of the base video.

[0110] Positive modifications to the luma components of the base video serve to selectively brighten pixels of the base video. This is achieved by using an enhancement overlay stream in which each frame comprises an array of values between 0 and 1 corresponding to pixels of the enhanced rendered video. These values represent the degree of opacity. When rendering the one or more enhancement overlay streams in the second markup video display region 410, the enhancement stream decoder 408 will treat a value of 0 as representing no modification to the underlying base video and a value of 1 as a fully opaque white pixel overlayed on the underlying base video. For values between 0 and 1 , the enhancement stream decoder 408 applies a white pixel overlay having an opacity equal to the respective value of the frame enhancement overlay stream. The enhancement stream decoder 408 may comprise a shader configured to perform this rendering of the positive luma modifications.

[0111] Negative modifications to the luma components of the base video serve to selectively darken pixels of the base video. This is achieved by using an enhancement overlay stream in which each frame comprises an array of values between 0 and 1 corresponding to pixels of the enhanced rendered video. However, whereas positive luma modifications are treated as transparency values, for negative modifications the enhancement stream decoder 408 will render the corresponding enhancement overlay stream by, for each pixel of each frame of said enhancement overlay stream, averaging between the value of the pixel of the enhancement overlay stream and the underlying luma value of the base video. The enhancement stream decoder 408 may comprise a cascading style sheets (CSS) filter to perform this rendering of the negative luma modifications.

[0112] Modifications to the luma components of a base video are typically performed to improve the resolution of the image, in which case the resolution of the one or more enhancement overlay streams is higher than the resolution of the base video. However, one more enhancement overlay streams may also be used to correct for errors introduced by the encoding and decoding of a reference video signal into the rendered base video, in which case the one or more enhancement overlay streams may be at the same resolution as the rendered base video. In another example, the one or more enhancement overlay streams encode modifications to the chroma components of the rendered base video so as to increase the colour space, in which case the one or more enhancement overlay streams may be at the same resolution as the rendered base video.

[0113] The enhancement stream decoder 408 may be a multi-layer video stream decoder plugin (DPI) such as an LCEVC decoder plugin, configured to decode an LCEVC- encoded video stream. For example, the encoded enhancement stream may comprise one or sets of encoded residuals which are then decoded by the enhancement stream decoder 408 to obtain one or more sets of decoded residuals. The enhancement stream decoder 408 then uses the one or more sets of decoded residuals to render one or more enhancement overlay streams. For example, when at least one of the sets of decoded residuals comprises positive and negative residuals, the enhancement stream decoder 408 may render at least a first positive residual overlay stream and at least a first negative residual overlay stream.

[0114] The enhancement stream decoder 408 may provide the decoding capabilities of a second layer decoder and the control capabilities of a decoding controller, such as those described with reference to Figures 2Aand 2B. The enhancement stream decoder 408 may comprise a single component or two separate components depending on the implementation, with the same functional effect. One or more components of the system 400 may be implemented in a browser. In one example, a browser is provided comprising the enhancement stream decoder 408.

[0115] Figure 4B is a schematic diagram of an alternative exemplary system 400 for decoding an encoded multi-layer video stream, such as an encoded multi-layer video stream encoded using LCEVC. Components and features common to those of Figure 4A function in the same way as previously described and so will not be described again in detail.

[0116] As before, the system 400 of Figure 4B comprises a source buffer 302 to receive an encoded multi-layer video stream, for example a first layer stream and a second layer stream. The encoded base stream is extracted from the source buffer 402 and decoded using the HTML media element 404. The decoded base stream is rendered in a first markup video display region 406.

[0117] The source buffer 402 issues a call-back to indicate that data is ready for decoding and the call-back is received by the enhancement stream decoder 408. On receipt of the call-back, the encoded enhancement stream is extracted from the source buffer 402 and decoded by the enhancement stream decoder 408. A timestamp, e.g. a PTS time stamp, is obtained from the base stream or from the source buffer call-back.

[0118] The HTML media element 404 renders a base video signal in first markup video display region 406 and the enhancement stream decoder 408 overlays one or more enhancement overlay streams in the second markup video display region 410 which is coincident with the first markup video display region 406 to generate a reconstructed video stream, as discussed above with reference to Figure 4A.

[0119] As previously discussed, when LCEVC is used, the encoded enhancement stream extracted from the source buffer 402 is carried within a set of SEI messages that accompany and are associated with the base stream or within a separate NAL unit stream. Base stream decoders are often configured to ignore SEI messages or a NAL unit stream if these contain information the base stream decoder cannot interpret. This means that the HTML media element 404 retrieves data relating to the base stream from the source buffer 402 in a default manner, wherein both enhanced and non-enhanced base streams are processed in a common manner. The HTML media element 404 ignores SEI messages or NAL units that carry the enhancement stream that reside within the source buffer 402.

[0120] As seen in Figure 4B, the encoded enhancement stream extracted from the source buffer 402, and carrying the set of SEI messages are first passed to an integration layer 412 before being passed to the enhancement stream decoder 408 for decoding. In particular, the integration layer 412 extracts LCEVC data from the enhancement stream based on the type of source buffer. The extracted data can be NAL units with LCEVC data and a realtime transport protocol (RTP) timestamp or data appended to the source buffer from media source extensions. The LCEVC data is indexed using the PTS and stored for use when overlaying the one or more enhancement overlay streams on the rendered base video. This is done by comparing timing metadata, such as a playback time of the rendered base video, with the PTS synchronizing the one or more enhancement overlay streams with the rendered base video signal accordingly. As shown in Figure 4B, the relevant LCEVC data is passed to the enhancement stream decoder 408 along with an offset which is generated based on the environment of playback.

[0121] The offset is calculated by an offset calculation block 414, advantageously by using one of two methods. The first method involves calculating the offset during the source buffer append and the second method involves calculating the offset during the enhancement stream decoding, as described above with reference to offset calculation block 314 of Figure 3B.

[0122] Notably, whereas in conventional approaches to rendering video the individual frames of the base video may be synced with corresponding frames of the enhancement streams, this may not be possible due to security protections applied to the base video. As such, the system 400 of Figure 4B is configured to synchronize the one or more enhancement overlay layers with the rendered base video using timing metadata of the rendered base video signal, preferably a playback time of the rendered base video signal, and timestamps of the frames of the one or more enhancement overlay layers.

[0123] An example method 500 of rendering a representation of a reference video signal will now be described with reference to Figure 5. The method 500 may be performed with system 400 as described with reference to Figures 4A or 4B or, for example, may be implemented as a set of instructions that are executed by a processor. As in the examples above, the multi-layer video stream encodes a video signal and comprises at least a first layer and a second layer. The first layer is decoded using a first decoding method and the second layer is decoded using a second decoding method. The first decoding method uses data that is inaccessible to the second decoding method. For example, a PTS for decoded frames within a first layer of the multi-layer video stream may be available to the first decoding method but not available to the second decoding method, e.g. due to the use of secure or restricted internal data.

[0124] Preferably, the method 500 is implemented within a browser. As discussed above, the browser may be any browser that processes documents in a markup scripting language to generate a visual representation of a webpage. The visual representation of the webpage may be made visible through a user interface associated with a client device associated with the browser.

[0125] At block 502, an encoded multi-layer video stream is received in a source buffer, such as in the source buffer 402 of Figures 4A or 4B. The encoded multi-layer video stream comprises an encoded base stream and an encoded enhancement stream.

[0126] The encoded base stream may be a down-sampled source signal encoded using a base encoder or codec, and decodable by a decoder, such as a hardware-based decoder. The base encoder or codec can be any base encoder or codec, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), VP9, MPEG-5 Essential Video Coding (EVC), and AOMedia Video 1 (AV1) encoders and codecs. Using existing base encoders and codecs as part of the encoding (and decoding) procedure ensures that systems that are not capable of rendering multi-layer video content may still decode the base stream using the existing base codec. This means that no updates to hardware are required to decode the encoded multi-layer video stream using the method 500, and future base codecs may also be used without further hardware upgrades, should the hardware of a system be upgraded to become compatible with the future base codec.

[0127] In some examples, the encoded enhancement stream comprises an encoded set of residuals which correct or enhance the base stream. There may be multiple levels of enhancement data in a hierarchical structure. The encoded enhancement stream may be encoded using a dedicated encoder configured to generate an encoded enhancement stream from uncompressed full resolution video.

[0128] An LCEVC-enhanced stream is an example of such a video stream encoded using the multi-layer coding scheme. In this case, the video stream may be encoded by an LCEVC encoder. Other examples are also envisaged, though.

[0129] At block 504 of Figure 5, the encoded base stream is extracted from the source buffer and decoded using a markup video element. The markup video element may be an element of a markup scripting language, such as HTML and XHTML. For example, the markup video element may be the HTML media element 404 shown in Figure 4B. Then, encoded base stream may be decoded using a suitable decoder. The markup video element receives or otherwise obtains an indication of the video coding format of the encoded base stream. When processed by the browser, the markup video element locates a decoder, or codec, associated with the video coding format and causes the encoded base stream to be decoded by the decoder. HTML5, for example, includes a <video> element within the markup language to embed video content in a webpage, which is an example of a markup video element. The base decoder is any decoder capable of decoding the encoded base stream.

[0130] The decoded base stream comprises a plurality of individual frames. A frame, for example, corresponds to a still image or picture. In some examples, a video is composed of a series of frames. A frame may include a plurality of pixels. Each frame comprises data representing properties of the video content. For example, a frame may comprise data defining the colour of each pixel in the frame. This data can be used by the markup video element to form a visual representation of the video stream when rendered in the final webpage.

[0131] However, while the decoded base stream comprises a plurality of individual frames, the decoded base stream may be subject to security protections that restrict access to the individual frames of the decoded base stream.

[0132] At block 506 of Figure 5, the decoded base stream is rendered in a first markup video display region. The browser for example processes the markup video element corresponding to a markup video display region in the generated webpage by embedding a media player in the webpage and rendering the video content in the media player. The first video display region is not hidden meaning that the unenhanced video content may be not displayed.

[0133] The first markup video display region may be defined in the markup scripting language.

[0134] At block 508, the encoded enhancement stream is extracted from the source buffer and decoded. The encoded enhancement stream comprises the enhancement data associated with one or more enhancement layers of the multilayer video stream. The encoded enhancement stream may be decoded by a multi-layer video stream DPI, such as the enhancement stream decoder 408 shown in Figures 4A and 4B, and used to obtain one or more enhancement overlay streams.

[0135] Finally, at block 510, the one or more enhancement overlay streams are rendered in a second markup video display region which is overlayed on the first markup video display region. A <canvas> tag within the markup language is preferably used to render the one or more enhancement overlay streams. In this way, the enhancement data comprised in the one or more enhancement overlay streams may be used to overlay enhancements on the decoded base stream so as to generate an enhanced video signal without requiring access to the data of the decoded base stream.

[0136] The one or more enhancement overlay streams may encode modifications to the luma and / or chroma components and may be at a higher resolution than the base video. For example, the one or more enhancement overlay streams may encode positive or negative modifications to the luma components of the base video.

[0137] Positive modifications to the luma components of the base video serve to selectively brighten pixels of the base video. This is achieved by using an enhancement overlay stream in which each frame comprises an array of values between 0 and 1 corresponding to pixels of the enhanced rendered video. These values represent the degree of opacity. When rendering the one or more enhancement overlay streams in the second markup video display region, a value of 0 is treated as representing no modification to the underlying base video and a value of 1 as a fully opaque white pixel overlayed on the underlying base video. For values between 0 and 1 , the method comprises applying a white pixel overlay having an opacity equal to the respective value of the frame enhancement overlay stream. This may be performed using a shader.

[0138] Negative modifications to the luma components of the base video serve to selectively darken pixels of the base video. This is achieved by using an enhancement overlay stream in which each frame comprises an array of values between 0 and 1 corresponding to pixels of the enhanced rendered video. However, whereas positive luma modifications are treated as transparency values, for negative modifications the method comprises rendering the corresponding enhancement overlay stream by, for each pixel of each frame of said enhancement overlay stream, averaging between the value of the pixel of the enhancement overlay stream and the underlying luma value of the base video. This may be performed using a cascading style sheets (CSS) filter.

[0139] Modifications to the luma components of a base video are typically performed to improve the resolution of the image, in which case the resolution of the one or more enhancement overlay streams is higher than the resolution of the base video. However, one more enhancement overlay streams may also be used to correct for errors introduced by the encoding and decoding of a reference video signal into the rendered base video, in which case the one or more enhancement overlay streams may be at the same resolution as the rendered base video. In another example, the one or more enhancement overlay streams encode modifications to the chroma components of the rendered base video so as to increase the colour space, in which case the one or more enhancement overlay streams may be at the same resolution as the rendered base video.

[0140] Figure 6 shows a spatially scalable coding scheme that uses a down-sampled source signal encoded with a base codec, adds a first level of correction or enhancement data to the decoded output of the base codec to generate a corrected picture, and then adds a further level of correction or enhancement data to an up-sampled version of the corrected picture. Thus, the spatially scalable coding scheme may generate an enhancement stream with two spatial resolutions (higher and lower), which may be combined with a base stream at the lower spatial resolution.

[0141] In the spatially scalable coding scheme, the methods and apparatuses may be based on an overall algorithm which is built over an existing encoding and / or decoding algorithm (e.g., MPEG standards such as AVC / H.264, HEVC / H.265, etc. as well as non-standard algorithms such as VP9, AV1 , and others) which works as a baseline for an enhancement layer. The enhancement layer works accordingly to a different encoding and / or decoding algorithm. The idea behind the overall algorithm is to encode / decode hierarchically the video frame as opposed to using block-based approaches as done in the MPEG family of algorithms. Hierarchically encoding a frame includes generating residuals for the full frame, and then a reduced or decimated frame and so on.

[0142] Figure 6 shows a system configuration for an example spatially scalable encoding system 600. The encoding process is split into two halves as shown by the dashed line. Each half may be implemented separately. Below the dashed line is a base level and above the dashed line is the enhancement level, which may usefully be implemented in software. The encoding system 600 may comprise only the enhancement level processes, or a combination of the base level processes and enhancement level processes as needed. The encoding system 600 topology at a general level is as follows. The encoding system 600 comprises an input I for receiving an input signal 601. The input I is connected to a down-sampler 605D. The down-sampler 605D outputs to a base encoder 620E at the base level of the encoding system 600. The down-sampler 605D also outputs to a residual generator 610-S. An encoded base stream is created directly by the base encoder 620E, and may be quantised and entropy encoded as necessary according to the base encoding scheme. The encoded base stream may be the base layer as described above, e.g. a lowest layer in a multi-layer coding scheme.

[0143] Above the dashed line is a series of enhancement level processes to generate an enhancement layer of a multi-layer coding scheme. In the present example, the enhancement layer comprises two sub-layers. In other example, one or more sublayers may be provided. In Figure 6, to generate an encoded sub-layer 1 enhancement stream, the encoded base stream is decoded via a decoding operation that is applied at a base decoder 620D. In preferred examples, the base decoder 620D may be a decoding component that complements an encoding component in the form of the base encoder 620E within a base codec. In other examples, the base decoding block 620D may instead be part of the enhancement level. Via the residual generator 610-S, a difference between the decoded base stream output from the base decoder 620D and the down-sampled input video is created (i.e., a subtraction operation 610-S is applied to a frame of the down- sampled input video and a frame of the decoded base stream to generate a first set of residuals). Here, residuals represent the error or differences between a reference signal or frame and a desired signal or frame. The residuals used in the first enhancement level can be considered as a correction signal as they are able to ‘correct’ a frame of a future decoded base stream. This is useful as this can correct for quirks or other peculiarities of the base codec. These include, amongst others, motion compensation algorithms applied by the base codec, quantisation and entropy encoding applied by the base codec, and block adjustments applied by the base codec. In Figure 6, the first set of residuals are transformed, quantised and entropy encoded to produce the encoded enhancement layer, sub-layer 1 stream. In Figure 6, a transform operation 610-1 is applied to the first set of residuals; a quantisation operation 620-1 is applied to the transformed set of residuals to generate a set of quantised residuals; and, an entropy encoding operation 630-1 is applied to the quantised set of residuals to generate the encoded enhancement layer, sub-layer 1 stream (e.g., at a first level of enhancement). However, it should be noted that in other examples only the quantisation step 620-1 may be performed, or only the transform step 610-1. Entropy encoding may not be used, or may optionally be used in addition to one or both of the transform step 610-1 and quantisation step 620-1. The entropy encoding operation can be any suitable type of entropy encoding, such as a Huffmann encoding operation or a run-length encoding (RLE) operation, or a combination of both a Huffmann encoding operation and a RLE operation (e.g., RLE then Huffmann or prefix encoding).

[0144] To generate the encoded enhancement layer, sub-layer 2 stream, a further level of enhancement information is created by producing and encoding a further set of residuals via residual generator 600-S. The further set of residuals are the difference between an up-sampled version (via up-sampler 605U) of a corrected version of the decoded base stream (the reference signal or frame), and the input signal 601 (the desired signal or frame).

[0145] To achieve a reconstruction of the corrected version of the decoded base stream as may be generated at a decoder, at least some of the sub-layer 1 encoding operations are reversed to mimic the processes of the decoder, and to account for at least some losses and quirks of the transform and quantisation processes. To this end, the first set of residuals are processed by a decoding pipeline comprising an inverse quantisation block 620-1 i and an inverse transform block 610-1 i. The quantised first set of residuals are inversely quantised at inverse quantisation block 620-1 i and are inversely transformed at inverse transform block 610-1 i in the encoding system 600 to regenerate a decoder-side version of the first set of residuals. The decoded base stream from decoder 620D is then combined with the decoder-side version of the first set of residuals (i.e. , a summing operation 610-C is performed on the decoded base stream and the decoder-side version of the first set of residuals). Summing operation 610-C generates a reconstruction of the down-sampled version of the input video as would be generated in all likelihood at the decoder - i.e. a reconstructed base codec video). The reconstructed base codec video is then up-sampled by up-sampler 605U.

[0146] Processing in this example is typically performed on a frame-by-frame basis. Each colour component of a frame may be processed as shown in parallel or in series. Notably, even if the content is protected such that access to individual frames of the base stream is restricted at a decoder, the encoder may still generate the base layer and enhancement layer on a frame by frame basis. The structure of the base and enhancement streams is the same whether or not security protections are applied during decoding.

[0147] The up-sampled signal (i.e., reference signal or frame) is then compared to the input signal 601 (i.e., desired signal or frame) to create the further set of residuals (i.e., a difference operation is applied by the residual generator 600-S to the up- sampled re-created frame to generate a further set of residuals). The further set of residuals are then processed via an encoding pipeline that mirrors that used for the first set of residuals to become an encoded enhancement layer, sub-layer 2 stream (i.e., an encoding operation is then applied to the further set of residuals to generate the encoded further enhancement stream). In particular, the further set of residuals are transformed (i.e., a transform operation 610-0 is performed on the further set of residuals to generate a further transformed set of residuals). The transformed residuals are then quantised, and entropy encoded in the manner described above in relation to the first set of residuals (i.e., a quantisation operation 620-0 is applied to the transformed set of residuals to generate a further set of quantised residuals; and, an entropy encoding operation 630-0 is applied to the quantised further set of residuals to generate the encoded enhancement layer, sub-layer 2 stream containing the further level of enhancement information). In certain cases, the operations may be controlled, e.g. such that, only the quantisation step 620-1 may be performed, or only the transform and quantisation step. Entropy encoding may optionally be used in addition. Preferably, the entropy encoding operation may be a Huffmann encoding operation or a run-length encoding (RLE) operation, or both (e.g., RLE then Huffmann encoding). The transformation applied at both blocks 610-1 and 610-0 may be a Hadamard transformation that is applied to 2x2 or 4x4 blocks of residuals.

[0148] The encoding operation in Figure 6 does not result in dependencies between local blocks of the input signal (e.g., in comparison with many known coding schemes that apply inter or intra prediction to macroblocks and thus introduce macroblock dependencies). Hence, the operations shown in Figure 6 may be performed in parallel on 4x4 or 2x2 blocks, which greatly increases encoding efficiency on multicore central processing units (CPUs) or graphical processing units (GPUs).

[0149] As illustrated in Figure 6, the output of the spatially scalable encoding process is one or more enhancement streams for an enhancement layer which preferably comprises a first level of enhancement and a further level of enhancement. This is then combinable (e.g., via multiplexing or otherwise) with a base stream at a base level. The first level of enhancement (sub-layer 1) may be considered to enable a corrected video at a base level, that is, for example to correct for encoder quirks. The second level of enhancement (sub layer 2) may be considered to be a further level of enhancement that is usable to convert the corrected video to the original input video or a close approximation thereto. For example, the second level of enhancement may add fine detail that is lost during the downsampling and / or help correct from errors that are introduced by one or more of the transform operation 610-1 and the quantisation operation 620-1.

Claims

CLAIMS1. A method of rendering a representation of a reference video signal, the method comprising: obtaining a first rendered video signal representing the reference video signal at a first level of quality; and overlaying one or more enhancement overlay streams on the first rendered video signal to generate a second rendered video signal representing the reference video signal at a second level of quality, the second level of quality being a higher level of quality than the first level of quality.

2. A method according to claim 1 , wherein obtaining the first rendered video signal comprises: obtaining a first decoded video stream; and rendering the first rendered video signal in a first markup video display region based on the first decoded video stream.

3. A method according to claim 2, wherein obtaining the first decoded video stream comprises: receiving a first encoded video stream; and, in a first decoding process, decoding the first encoded video stream to obtain the first decoded video stream.

4. A method according to claim 2 or claim 3, wherein the first markup video display region is an HTML <video> element.

5. A method according to any of claims 2 to 4, wherein overlaying the one or more enhancement overlay streams comprises: obtaining one or more decoded enhancement streams; and rendering the one or more enhancement overlay streams in a second markup video display region based on the one or more decoded enhancement streams, the second markup video display region being coincident with the first markup video display region.

6. A method according to claim 5, wherein obtaining the one or more decoded enhancement streams comprises: receiving one or more encoded enhancement streams; and, in a second decoding process, decoding the one or more encoded enhancement streams to obtain the one or more decoded enhancement streams.

7. A method according to claim 5 or claim 6, wherein the first rendered video signal comprises data that is inaccessible to the second decoding method.

8. A method according to any of claims 5 to 7, wherein the one or more encoded enhancement streams comprise one or more sets of encoded residuals and decoding the one or more encoded enhancement streams comprises decoding the one or more sets of encoded residuals to obtain one or more sets of decoded residuals.

9. A method according to claim 8, wherein the one or more encoded enhancement streams comprise at least a second set of encoded residuals in a second enhancement sublayer of a hierarchical coding scheme.

10. A method according to claim 8 or clam 9, wherein at least one of the one or more sets of decoded residuals comprises positive and negative residuals and decoding the one or more encoded enhancement streams comprises: obtaining one or more decoded positive residual streams from the positive residuals; and obtaining one or more decoded negative residual streams from the negative residuals.

11. A method according to claim 10, wherein rendering the one or more enhancement overlay streams comprises: rendering one or more positive residual overlay streams in the second markup video display region based on the decoded positive residual streams; and rendering one or more negative residual overlay streams in the second markup video display region based on the decoded negative residual streams.

12. A method according to claim 11 , wherein at least one positive residual overlay stream of the one or more positive residual overlay streams encodes modifications to the luma components of the rendered base video, said at least one positive residual overlay stream comprising an array of values of between 0 and 1 , each of said values corresponding to a pixel of the second rendered video signal, and rendering said at least one positive residual overlay stream in the second markup video display region comprises, for each pixel of the second rendered video signal, rendering a white pixel having an opacity equal to the corresponding value of the array of values.

13. A method according to claim 11 or claim 12, wherein at least one negative residual overlay stream of the one or more negative residual overlay streams encodes modifications to the luma components of the rendered base video, said at least one negative residual overlay stream comprising an array of values of between 0 and 1 , each of said values corresponding to a pixel of the second rendered video signal, and rendering said at least one negative residual overlay stream in the second markup video display region comprises, for each pixel of the second rendered video signal: averaging the corresponding value from the array of values with the luma value of the underlying pixel in the first rendered video signal, and rendering a pixel having said averaged luma value.

14. A method according to any of claims 5 to 13, wherein the second markup video display region is an HTML <canvas> element.

15. A method according to any of claims 1 to 14, wherein overlaying the one or more enhancement overlay streams comprises synchronizing the one or more enhancement overlay streams with the first rendered video signal based on a playback time of the first rendered video signal.

16. A method according to claim 15, wherein each of the one or more enhancement overlay streams comprises a series of frames, the series of frames being indexed using a corresponding series of timestamps; and for each of the one or more enhancement overlay streams, synchronizing said enhancement overlay stream with the first rendered video signal comprises:comparing a current playback time of the first rendered video signal with the timestamps indexing the frames of said enhancement overlay stream; identifying the frame of said enhancement overlay stream corresponding to the current playback time based on said comparison; and overlaying the frame of said enhancement overlay stream corresponding to the current playback time on the first rendered video signal.

17. A method according to claim 16 and any of claims 5 to 14, wherein the timestamps are derived from the one or more encoded enhancement streams.

18. A method according to claim 16 or claim 17, wherein comparing the current playback time of the first rendered video signal with the timestamps indexing the frames of said enhancement overlay stream comprises: searching for a frame of said enhancement overlay stream that has a timestamp that falls within a range defined with reference to the current playback time.

19. A method according to claim 18, wherein the range is set based on a configurable drift offset.

20. A method according to claim 19, wherein the configurable drift offset is based on a framerate of the first rendered video signal.

21. A method according to claim 18 or claim 19, wherein searching for a frame of said enhancement overlay stream that has a timestamp that falls within a range defined with reference to the current playback time comprises: searching for a frame of said enhancement overlay stream where the timestamp of the frame plus the configurable drift offset equals the current playback time.

22. A method according to any of the preceding claims, wherein the second rendered video signal is at a higher resolution than the first rendered video signal.

23. A method according to any of the preceding claims, wherein one or more of the one or more enhancement overlay streams encode modifications to the luma and / or chroma components of the rendered base video.

24. A method according to any of the preceding claims, wherein the method is implemented within a browser.

25. A system for rendering a representation of a reference video signal, the system comprising: a video Tenderer configured to obtain a first rendered video signal representing the reference video signal at a first level of quality; and an enhancement decoder configured to overlay one or more enhancement overlay streams on the first rendered video signal to generate a second rendered video signal representing the reference video signal at a second level of quality, the second level of quality being a higher level of quality than the first level of quality.

26. A computer-readable medium comprising instructions which, when executed, cause a processor to perform a method according to any of claims 1 to

Citation Information

Patent Citations

  • Integrating a decoder for hierachical video coding

    WO2022023739A1

  • Decoding a video stream within a browser

    GB2601364A

  • Synchronising frame decoding in a multi-layer video stream

    GB2613886A

  • Method and apparatus for video coding and decoding

    US20150304665A1

  • Secure decoder and secure decoding methods

    WO2022243672A1