Low latency quality switching for adaptive video streaming
The method employs a layered coding scheme with metadata-driven switching points to enable low latency quality transitions in adaptive video streaming, addressing latency issues in existing technologies by allowing immediate quality changes without excessive bitrates.
Patent Information
- Application Number
- PCT/EP2024/088648
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-03
AI Technical Summary
Existing adaptive video streaming technologies suffer from high latency in quality switching, particularly in immersive multimedia systems, leading to disruptions in user experience due to the need to wait for the end of CMAF fragments for quality adaptations, which can be exacerbated by the requirement for all-intra encoding with smaller fragments or frames, resulting in excessive bitrates.
A method and system utilizing a layered coding scheme with a base layer and enhancement layers, where video frames in the enhancement layer have dependencies only on current or past frames of the base layer, allowing low latency switching by retrieving encoded frames at specific switching points without waiting for the end of a fragment, using metadata to identify these points and enable seamless quality transitions.
Enables low latency quality switching in adaptive video streaming without increasing bitrate, allowing immediate transitions to higher quality without the need to wait for fragment ends, thereby maintaining a smooth user experience and reducing encoding inefficiencies.
Smart Images

Figure EP2024088648_03072025_PF_FP_ABST
Abstract
Description
[0001] Low latency quality switching for adaptive video streaming
[0002] Technical field
[0003] The embodiments relate to low latency quality switching for adaptive video streaming, and, in particular, though not exclusively, to methods and systems for low latency quality switching for adaptive video streaming, a video data structure for low latency adaptive video streaming and a computer program product for executing such methods.
[0004] Background
[0005] Quality adaptations based on network conditions and client resources are essential to maintain an undisrupted video timeline. Especially for immersive multimedia systems, such quality adaptations should be performed with very low latency to maintain a high Quality of Experience (QoE) for the user. For example, in 360 tiled streaming or, more generally, XR content consumption where only a part of the content is inspected in high quality, a user action may trigger the consumption of another part of the content that is delivered only in low quality. In this case, being able to fetch a high-quality version of that part of the content as fast as possible is crucial to avoid disruptions in the quality perceived by the user.
[0006] Quality switching is addressed in media standards such as Dynamic Adaptive Streaming of HTTP (MPEG-DASH), HTTP Live Streaming (HLS) and Common Media Applications Format (CMAF). For example, CMAF defines a data format for encoding and packaging media data for adaptive media streaming applications. Quality switching in CMAF may be performed through adaptive switching between CMAF tracks associated with different rendering quality. Each CMAF track comprises encoded video data in CMAF fragments that can be requested by a client device. Different CMAF tracks associated with encoded video data of different qualities may define a CMAF switching. By default, a client device needs to wait until the end of the CMAF fragment for quality adaptations, e.g. switching to a higher video quality. This latency hinders optimal user experience.
[0007] Hence, in CMAF quality switching is performed at a CMAF fragment level so that if faster quality switching is desired, CMAF fragment lengths should be made smaller. Considering that each CMAF fragment should start with an l-frame (or a frame similar to an I- frame), decreasing the length of the CMAF fragment would introduce more l-frames in the bitstream, leading to a substantial increase of the bitrate. In an extreme case, wherein an application demands quality switching on the granularity of a single frame or a few frames, would require to use a single frame as a CMAF fragment. In other terms, an all-intra encoded bitstream would be required with each frame being placed in a CMAF fragment. Such a solution however would request excessive bitrates. Hence, from the above it follows there is a need in the art for improved methods and systems that provide low latency quality switching for adaptive streaming applications.
[0008] Summary of the invention
[0009] As will be appreciated by one skilled in the art, aspects of the present invention may be embodied as a system, method or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a "circuit," "module" or "system." Functions described in this disclosure may be implemented as an algorithm executed by a microprocessor of a computer. Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code embodied, e.g., stored, thereon.
[0010] Any combination of one or more computer readable medium(s) may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0011] A computer readable signal medium may include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium may be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber, cable, RF, etc., or any suitable combination of the foregoing. Computer program code for carrying out operations for aspects of the present invention may be written in any combination of one or more programming languages, including an object-oriented programming language such as Java(TM), Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0012] Aspects of the present invention are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor, in particular a microprocessor or central processing unit (CPU), of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer, other programmable data processing apparatus, or other devices create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0013] These computer program instructions may also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks.
[0014] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. Additionally, the Instructions may be executed by any type of processors, including but not limited to one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FP- GAs), or other equivalent integrated or discrete logic circuitry.
[0015] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0016] In an aspect, the embodiments may relate to a method of processing video data which may include one or more of the following steps: receiving by a client apparatus one or more video base parts, each video base part comprising encoded video frames representing video data of a base layer and each video base part being associated with one or more corresponding video enhancement parts, the one or more corresponding video enhancement parts comprising encoded video frames representing video data of one or more enhancement layers respectively for enhancing video data of the base layer, the encoded video frames of the one or more video base parts and the one or more video enhancement parts having a common timeline; receiving switching information, the switching information identifying a corresponding video enhancement part for a video base part and one or more locations of one or more video frames in the corresponding video enhancement part that are associated with one or more switching points on the common timeline, wherein video frames of the corresponding enhancement layer are encoded such that a video frame associated with a switching point in the video enhancement part can have one or more coding dependencies on video frames in the associated video base part on or preceding the switching point and such that the video frame associated with a switching point in the video enhancement part has no coding dependencies on video frames in the video enhancement part preceding the switching point; and, switching to decoding video data of a higher quality, the switching including: based on the switching information determining for a current video base part at least one corresponding enhancement part and a position of a video frame in the corresponding enhancement part that is associated with a switching point; retrieving encoded video frames of the corresponding enhancement part, the encoded video frames comprising encoded video frames having a position on the common timeline on or succeeding the switching point; and, decoding encoded video frames of the video base part and the encoded video frames of the corresponding enhancement part for determining video frames of a quality that is higher than the quality of video frames associated with the base layer.
[0017] In an embodiment, video frames of the corresponding enhancement layer may be encoded such that video frames in the enhancement part succeeding the switching point have no coding dependency on video frames in the video fragment of the enhancement layer preceding the switching point.
[0018] In an embodiment, video frames of the corresponding enhancement layer may be encoded such that a video frame associated with a switching point in the video enhancement part only have coding dependencies on one or more video frames in the associated video base part at or preceding the switching point.
[0019] The method allows low latency switching from a first resolution to a second, higher resolution at one of the one or more switching points signalled in metadata associated with the one or more video base parts, e.g. one or more CMAF fragments carrying video data of the base layer. This way, a video processing apparatus, e.g. a client device, does not have to wait until the end of a video base part, such as a CMAF fragment, for quality adaptations of the video data that is played out by a client apparatus. To that end, a layered coding scheme is used, which includes a base layer and one or more enhancement layers, wherein a base layer carries metadata to fetch part of an enhancement layer that depends on already-delivered base layer video data. The method allows bitrate reduction by allowing the use of longer data containers, e.g. CMAF fragments, and larger GOP sizes, without increasing the delay for switching to a higher quality. In other words, compared to using smaller CMAF fragments in order to maintain the same delay in quality switching, the method offers better encoding efficiency.
[0020] In an embodiment, the one or more video base parts and / or the one or more video enhancement parts may be formatted as Common Media Application Format (CMAF) fragments.
[0021] In an embodiment, each CMAF fragment may be subdivided into a plurality of non-overlapping CMAF chunks.
[0022] In an embodiment, the CMAF chunks may be received using an HTTP chunked transfer mode. In an embodiment, each of the one or more switching points may coincide with the start of an CMAF chunk, e.g. the first video frame of a CMAF chunk, in one of the one or more video enhancement parts.
[0023] In an embodiment, the one or more video base parts and / or the one or more video enhancement parts may be formatted as DASH segments.
[0024] In an embodiment, the CMAF chunks may be received using an HTTP chunked transfer mode.
[0025] In an embodiment, the one or more locations of one or more video frames in the corresponding video enhancement part that is associated with the one or more switching points may define the start of one or more chunks, e.g. the first video frame of the one or more CMAF chunks, in the corresponding enhancement part.
[0026] In an embodiment, the switching information may include an identifier or information for determining an identifier for identifying the at least one corresponding enhancement part.
[0027] In an embodiment, the switching information may include a parameter value indicative of the position of the video frame in the corresponding enhancement part that is associated with the switching point.
[0028] In an embodiment, the parameter value may define a byte range in the corresponding enhancement part. The byte range may define the position of the video frame relative to the start of the corresponding enhancement part.
[0029] In an embodiment, encoded video frames of the corresponding enhancement part may be retrieved by the client apparatus based on a manifest file, such as a media presentation description (MPD), comprising one or more video part identifiers for identifying the one or more video base parts and / or one or more corresponding video enhancement parts associated with each of the one or more video base parts.
[0030] In an embodiment, retrieving encoded video frames of the corresponding enhancement part may include: determining a resource locator based on the one or more video part identifiers for locating the corresponding enhancement part; and, determining a request message based on the resource locator and information associated with the position of a video frame in the corresponding enhancement part for requesting the encoded video frames of the corresponding enhancement part. In an embodiment, the request message may be a HTTP message, such as a HTTP byte range request message,.
[0031] In an embodiment, at least part of the switching information may be provided to the client apparatus using one or more in-band messages, e.g. one or more DASH in band event messages. In an embodiment, at least part of the one or more in-band messages may be inserted in the one or more video base parts, and, optionally, in at least part of the one or more corresponding enhancement parts. In an embodiment, at least part of the one or more in-band messages may be inserted in one or more chunks of the one or more video base parts.
[0032] In another embodiment, at least part of the one or more in-band messages may be inserted in at least part of the one or more chunks of the one or more corresponding enhancement parts.
[0033] In an embodiment, the client apparatus may requests the video enhancement part associated with the video base part if the client apparatus receives a request for switching to a higher quality and / or if the client apparatus determines to switch to a higher quality.
[0034] In an embodiment, each of the one or more video base parts and / or one or more video enhancement parts may comprise at least a group of pictures (GOP), wherein the video frames of a GOP have no coding dependency on video frames outside of the GOP.
[0035] In an embodiment, the video data may be encoded using a layered coding scheme, e.g. scalable HEVC or multilayer VVC.
[0036] In an aspect, the embodiments in this disclosure may relate to an apparatus for processing video data comprising: a computer readable storage medium having at least part of a program embodied therewith; and, a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, wherein the processor may be configured to perform one or more of the following executable operations: receiving by a client apparatus one or more video base parts, each video base part comprising encoded video frames representing video data of a base layer and each video base part being associated with one or more corresponding video enhancement parts, the one or more corresponding video enhancement parts comprising encoded video frames representing video data of one or more enhancement layers respectively for enhancing video data of the base layer, the encoded video frames of the one or more video base parts and the one or more video enhancement parts having a common timeline; receiving switching information, the switching information identifying a corresponding video enhancement part for a video base part and one or more locations of one or more video frames in the corresponding video enhancement part that are associated with one or more switching points on the common timeline, wherein video frames of the corresponding enhancement layer are encoded such that a video frame associated with a switching point in the video enhancement part can have one or more coding dependencies on video frames in the associated video base part at or preceding the switching point and has no coding dependencies on video frames in the video enhancement part preceding the switching point; and, switching to decoding video data of a higher quality, the switching including: based on the switching information determining for a current video base part at least one corresponding enhancement part and a position of a video frame in the corresponding enhancement part that is associated with a switching point; retrieving encoded video frames of the corresponding enhancement part, the encoded video frames comprising encoded video frames having a position on the common timeline at or succeeding the switching point; decoding encoded video frames of the video base part and the encoded video frames of the corresponding enhancement part for determining video frames of a quality that is higher than the quality of video frames associated with the base layer.
[0037] In an embodiment, the processor may be further configured to perform any of the method steps as described with reference to the above embodiments
[0038] In a further aspect, the embodiments may relate to a video preparation module comprising: a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, the processor is configured to perform executable operations comprising: encoding video data into encoded video frames representing video data of a base layer and encoded video frames representing video data of one or more enhancement layers for enhancing video data of the base layer, packaging encoded video frames representing video data of the base layer into one or more video base parts and packaging encoded video frames representing video data of one or more enhancement layers into one or more video enhancement parts, each video base part being associated with one or more corresponding video enhancement parts, wherein encoded video frames of the one or more video base parts and the one or more video enhancement parts have a common timeline; generating switching information, the switching information identifying a corresponding video enhancement part for a video base part and one or more locations of one or more video frames in the corresponding video enhancement part that are associated with one or more switching points on the common timeline, wherein video frames of the corresponding enhancement layer are encoded such that a video frame associated with a switching point in the video enhancement part can have one or more coding dependencies on video frames in the associated video base part at or preceding the switching point and has no coding dependencies on video frames in the video enhancement part preceding the switching point; and, storing the one or more video base parts of the base layer, the one or more video enhancement parts of the one or more enhancement layers and the switching information on a storage medium.
[0039] In yet a further aspect, the embodiments may relate to a computer readable storage medium comprising stored data thereon, wherein the data may define for controlling video processing by a client device comprising, wherein the data may include: one or more video base parts, each video base part comprising encoded video frames representing video data of a base layer; one or more video enhancement parts, each video enhancement part comprising encoded video frames representing video data of one or more enhancement layers, each video base part being associated with one or more corresponding video enhancement parts, wherein encoded video frames of the one or more video base parts and the one or more video enhancement parts have a common timeline; and, switching information for the client device, the switching information identifying a corresponding video enhancement part for a video base part and one or more locations of one or more video frames in the corresponding video enhancement part that are associated with one or more switching points on the common timeline, wherein video frames of the corresponding enhancement layer are encoded such that a video frame associated with a switching point in the video enhancement part can have one or more coding dependencies on video frames in the associated video base part at or preceding the switching point and has no coding dependencies on video frames in the video enhancement part preceding the switching point.
[0040] The invention may also relate to a computer program product comprising software code portions configured for, when run in the memory of a computer, executing the method steps according to any of process steps described above.
[0041] The invention will be further illustrated with reference to the attached drawings, which schematically will show embodiments according to the invention. It will be understood that the invention is not in any way restricted to these specific embodiments.
[0042] Brief description of the drawings
[0043] Fig. 1 illustrates a known quality switching scheme for adaptive video streaming.
[0044] Fig. 2 illustrates a quality switching scheme for adaptive video streaming according to an embodiment;
[0045] Fig. 3 depicts an example of an encoding dependency between video data in a base layer and an enhancement layer;
[0046] Fig. 4 depicts a flow chart of a method of processing video data according to an embodiment.;
[0047] Fig. 5 depicts a schematic of a system for streaming media data based on a low latency quality switching scheme according to an embodiment;
[0048] Fig. 6 illustrates a quality switching scheme for adaptive video streaming according to another embodiment;
[0049] Fig. 7 illustrates a quality switching scheme for adaptive video streaming according to further embodiment; Fig. 8 illustrates a quality switching scheme for adaptive video streaming according to yet further embodiment;
[0050] Fig. 9 depicts a block diagram illustrating an exemplary data processing system that may be used with embodiments described in this disclosure.
[0051] Description of the embodiments
[0052] The embodiments in this disclosure relate to fast, low-latency quality switching for adaptive video streaming. An example of a media packaging standard supporting quality switching is the Common Media Application Format (CMAF). CMAF is built on top of the ISO Base Media File Format (ISOBMFF) and provides a common way to package the content in CMAF fragments such that it is suitable for streaming protocols such as Dynamic Adaptive Streaming HTTP (DASH) and HTTP Live Streaming (HLS). Further, it supports chunked encoding (i.e. , splitting a segment into smaller units referred to as chunks, which can be delivered upon encoding) and chunked transfer encoding, which is a data transfer mechanism in HTTP 1.1, referred to as the HTTP 1.1 chunked transfer encoding mode, wherein the data stream is split into non-overlapping chunks that are sent and received independently.
[0053] Benefitting from the chunked encoding and chunked transfer encoding capabilities of CMAF, and making use of CMAF chunks instead of CMAF fragments, CMAF enables to achieve low latency streaming in DASH. Specifically, the encoder and the packager do not need to wait for encoding of an entire CMAF fragment to be completed before transmitting it to the origin server of a e.g. a content delivery network. Instead, CMAF chunks are encoded, packaged and transmitted to the origin server (using the HTTP 1.1 chunked transfer encoding mode) and a client device can request a CMAF fragment using an HTTP request based a manifest file, which includes information for accessing video data, (e.g. information regarding resource locators which are needed to request CMAF chunks). This way, a client device can receive corresponding CMAF chunks, as soon as their encoding is completed. This decreases the end-to-end latency, which is proportional to the length of a CMAF chunk (which is much shorter than the length of a CMAF fragment).
[0054] Quality switching in CMAF is performed through adaptive switching between so-called CMAF tracks comprising CMAF switching sets. Switching sets are formed by CMAF tracks of CMAF fragments wherein the video frames in the CMAF fragments or the chunks in the CMAF fragments are synchronized and time-aligned (hereafter referred to time-aligned CMAF fragments and CMAF chunks).
[0055] Fig. 1 illustrates an example of a CMAF switching set 102i,2 of time-aligned CMAF fragments, including a first CMAF track 201 (a base layer) comprising a sequence of first CMAF fragments 204i^, and a second CMAF track 200 (an enhancement layer) comprising a sequence of second CMAF fragments 202i^. Each track may comprise video data associated with different characteristics, e.g. qualities. The video data of the base layer and enhancement layers may generated using a layered video coding scheme such as scalable video coding (SVC), Scalable HEVC or multilayer VVC. Further, a CMAF fragment may be divided in CMAF chunks. Other layered video coding schemes may include video coding schemes that are based on super-resolution techniques as e.g. described in WO2019 / 197674 with title Block-level super-resolution based video coding, which is hereby incorporated by reference in this disclosure.
[0056] Each CMAF fragment in a layer may be time aligned with a corresponding CMAF fragment in another layer. Each CMAF fragment in a layer may contain multiple nonoverlapping CMAF chunks. In some embodiments, these CMAF chunk in the layer may be time aligned CMAF chunks of a corresponding CMAF fragment in another layer. The first CMAF chunk in each CMAF fragment is configured as a stream access point (SAP) of type 1 or 2 (e.g., an IDR in an AVC-generated CMAF fragment), which allows independent decoding of each CMAF fragment. Here, the stream access point may relate to an SAP frame of type 1 or 2, e.g. an l-frame or an IDR frame or the like. Additionally, timestamps may be used to simply switching between tracks by sequencing CMAF fragments from different CMAF tracks during playback.
[0057] The CMAF format depicted in the figure allows a client device to consume the video data at low quality and to subsequently continue based on the high-quality video data based on an event trigger 106, e.g. a user action, network conditions, and / or resources of the client. As shown in the figure, by default, after the event trigger, the client device needs to wait until the end 108 of CMAF fragment 104i for a quality adaptation by switching to a higher resolution, i.e. retrieving both encoded video data of the base layer and the enhancement layer. Thus, upon switching, the client will start requesting CMAF chunks from both the base layer 110i and the enhancement layer 1102, wherein the video data of the enhancement layer are used to enhance the quality of the base layer.
[0058] In the scheme of Fig. 1 , smaller CMAF segment / fragment lengths should be used in order to achieve faster quality switches. Considering that each CMAF fragment should start with a SAP frame of type 1 or 2, by decreasing the length of the CMAF fragment, more SAP frames of type 1 or 2 are introduced to the bitstream, which leads to bitrate increase. In an extreme case, wherein an application demands granularity of a single frame for quality switching, based on the current CMAF specification would require to use a single frame in a CMAF fragment, or in other terms, an all-intra encoding would be required with each frame being placed in a CMAF fragment. However, it is clear that such a solution would request excessive bitrate. Fig. 2 illustrates an example of time-aligned fragments for low latency switching according to an embodiment of the invention. In particular, the figure illustrates CMAF tracks comprising switching sets of time-aligned fragments comprising media data which may be encoded using a layered coding scheme (e.g., Scalable HEVC or Multilayer VVC). In such encoding scheme, encoded media data may be formatted in a base layer 201 and one or more enhancement layers 200. The dependency of the layers may be specified in a manifest file which identifies layers, each comprising a sequence of CMAF fragments, wherein the dependency may be signalled using for example a specific parameter, e.g. the parameter flag @dependencylD in DASH MPD.
[0059] Media data associated with the base layer may be packaged in CMAF fragments 204i^ wherein each fragment is further divided in chunks, e.g. chunks 205i.sof fragment 2042. Media data of the enhancement layer may be packaged CMAF fragments 202I-4. In some embodiments, CMAF fragments of the enhancement layer may be further divided in CMAF chunks, e.g. chunks 205i-s of fragment 2022, which is associated with fragment 2042 of the base layer.
[0060] In this scheme, when the client device receives a trigger 206 for switching, the client device is able to switch to a higher quality based on metadata that is associated with the base layer 201. The metadata may include switching information for signalling one or more switching points 214i^ for each CMAF fragment, wherein the switching points are time- aligned with the timeline of a CMAF fragment 2042 of the base layer 201 and an associated CMAF fragment 2022 of the enhancement layer 200. Each switching point may identify at least one video frame in the CMAF fragment of the enhancement layer. This video frame identified by a switching point may be referred to as a switching frame. In an embodiment, a switching point identifying a switching frame may include an offset parameter relative to the start of the CMAF fragment (e.g. an I DR). This offset parameter may be implemented as a byte-range in the CMAF fragment of the enhancement layer. In another embodiment, the switching point may include a time stamp associated with the content timeline of the enhancement layer that identifies a switching frame. In that case, the timestamp needs to be translated to a byte range in the CMAF fragment of the enhancement layer in order to address the switching frame. In an embodiment, switching points may be defined such that they coincide with the start of a chunk in the CMAF fragment of the base layer as shown in Fig. 2. Other configurations of switching points are also possible. For example, in an embodiment, only one or a limited number of switching points per CMAF fragment may be defined.
[0061] Hence, if the client device receives an action trigger during playout of video frames of a CMAF fragment of the base layer, it will continue playout video frames of this CMAF fragment until playout of the common timeline reaches a switching point (in this particular example switching point 208) which coincides with the start of the third chunk 205s of CMAF fragment 2042, which has an associated CMAF fragment 2022 in the enhancement layer. At that moment, the client device may use the metadata associated with the base layer to identify a switching frame in an associated CMAF fragment of the enhancement layer. The switching frame defines the first video frame of a sequence of video frames that are needed for switching to rendering the media data at a higher quality.
[0062] To enable seamless switching to a higher quality at switching points signalled in the metadata associated with the base layer as illustrated in Fig. 2, the media data of the enhancement layer may be encoded such that during switching coding dependencies associated with video frames of the enhancement layer can be handled by the decoder. This means that a switching video frame in a CMAF fragment of the enhancement layer has one or more coding dependencies 2161.3 on one or more earlier video frames in an associated CMAF fragment (and, optionally, in one or more CMAF fragments preceding the associated CMAF fragment) of the base layer, wherein the earlier video frames are located on the common timeline preceding the switching point or on the switching point. Further, a switching video frame in a CMAF fragment of the enhancement layer has no coding dependencies on later video frames in the CMAF fragment of the enhancement layer wherein the later video frames are located on the common timeline succeeding the switching point. Hence, media data of the enhancement layer that are packaged in either CMAF fragments or CMAF chunks, depend only on media data of the base layer data packaged in CMAF chunks of current or past time. Hence, in this embodiment, for an enhancement layer only inter-layer encoding dependencies on media data of current and past video frames of the base layer are allowed. These encoding constraints allow many different implementations of layered encoding schemes that could be used.
[0063] Fig. 3 depicts an example of an encoding dependency between video data in a base layer and an enhancement layer that is suitable for quality switching schemes according to the embodiments in this disclosure. The figure illustrates a base layer 302 comprising chunks 302I,2, of a CMAF fragment and an enhancement layer 304 comprising chunks 304I,2 of a CMAF fragment corresponding to the CMAF fragment of the base layer. Hence, the chunks may be packaged into time-aligned fragments as described in detail with reference to Fig. 2. Each of the chunks may include a plurality of video frames, wherein metadata associated with the base layer may indicate that some video frames 302i^ in the enhancement layer are associated with a switching point (i.e. they are identified as switching frames). As shown in the figure, a switching frame in the enhancement layer only has coding dependencies on a current frame in the base layer or a past video frame in the base layer. Further, a switching frame in the enhancement layer has no coding dependencies on current or past video frames in the enhancement layer. For example, switching frame 306s (in this example a P frame) in the enhancement layer has a dependency on an associated P frame in the base layer 308s, which has a dependency on an earlier video frame 3082 (in this example a P frame) in the base layer. In turn, this earlier video frame 3082 has a dependency on yet an earlier video frame 308i (in this example an l-frame) in the base layer. Hence, the switching frame, i.e. the video frame associated with a switching point in a fragment of the enhancement layer, can have one or more coding dependencies on video frames in the associated fragment of the base layer, which have are located on the timeline on or preceding the switching point. Further, switching frame 306s in the enhancement layer has no coding dependency on earlier video frames in the enhancement layer, i.e. no coding dependencies on video frames in the video fragment of the enhancement layer preceding the switching point. Additionally, video frames in the enhancement layer succeeding the switching point also have no coding dependency on video frames in the video fragment of the enhancement layer preceding the switching point. This encoding scheme may be used for low latency quality switching schemes as described with reference to the embodiments in this disclosure.
[0064] Hence, from the above it follows that based on the switching information contained in the metadata associated with a base layer, a client device that is processing encoded video data of a chunk of the base layer is able to identify a switching frame in a corresponding chunk or a corresponding part of a fragment from the enhancement layer. One or more switching point in a fragment of the enhancement layer may be defined so that, the decoder can switch to generating video frames of a higher quality when the video data of a chunk in the enhancement layer is combined with video data of the present or past chunks of the base layer as explained with reference to Fig. 2 and 3.
[0065] The switching information including the information about the switching points may be signalled in or associated with chunks of the base layer chunks using for example event message (e.g.,an ‘emsg’ entry of a chunk). In an embodiment, each chunk of a fragment of the base layer may comprise an event message that carries metadata about the identification of corresponding enhancement video data, e.g. a corresponding chunk, of the enhancement layer. To that end, the event message may include a byte offset (relative to the start of a chunk) or a time stamp that corresponds to a switching frame in a fragment of the enhancement layer that is associated with the fragment in the base layer. The event message may also include an identifier of the CMAF fragment to which the byte offset belongs to.
[0066] Hence, in case of an action, triggered by e.g. a user action or prompted from an adaptive bitrate streaming algorithm, a client may decide to switch to rendering content at a higher quality. The client device may then parse the metadata associated with the base layer chunk and fetch the corresponding chunk or part of the fragment of the enhancement layer using a suitable request protocol, such as an HTTP request, in particular an HTTP byte range request.
[0067] The identifier of the enhancement layer and the identifier of the fragment (included in the metadata associated with the base layer) may be matched with information published in the manifest file, allowing the client to construct a resource locator, e.g. an URL, for the corresponding fragment. The metadata may further include information for identifying a switching frame in the CMAF fragment, such as a byte offset. This information may be used by the client to perform an HTTP byte range request and fetch the enhancement layer chunk or part of the fragment that is required for the quality switch.
[0068] After the switch, the client device receives video data of the base and the enhancement layer, which are fed to the decoder for decoding. Typically, the base layer is decoded independent from the enhancement layer. Further, the decoding of the enhancement layer may depend on the decoding of the base layer to improve coding efficiency. Hence, there are inter-layer dependencies the decoder takes care off when decoding the enhancement layer.
[0069] The decoder needs to be configured using a profile and level that is able to support both the base and the enhancement layer. Further, the media samples from the base and the enhancement layer may be formatted based on the same NAL Access Unit structure. The profile and level of the decoder may be configured for the high-quality representation and, therefore, no re-configuration is needed to support the decoding of the high quality version.
[0070] Fig. 4 depicts a flow chart of a method of processing video data according to an embodiment. In particular, encoded video data of a base layer and an enhancement layer are processed to allow low latency quality switching.
[0071] In a first step 402, the method may include receiving one or more video base parts, each video base part comprising encoded video frames representing video data of a base layer and each video base part being associated with one or more corresponding video enhancement parts, the one or more corresponding video enhancement parts comprising encoded video frames representing video data of one or more enhancement layers respectively for enhancing video data of the base layer, the encoded video frames of the one or more video base parts and the one or more video enhancement parts having a common timeline. Here each base part may define encoded video data, e.g. encoded video frames, that are packaged according to a certain format, e.g. the CMAF format.
[0072] In some embodiments, the base layer parts and enhancement layer parts may form a switching set, comprising time aligned and synchronized fragments, e.g. CMAF fragments. Fragments of the base layer may be divided in chunks, e.g. CMAF chunks. In some embodiments also the fragments of the enhancement layer may be divided in chunks. In other embodiment, the base layer parts may be CMAF fragments, and the enhancement layer parts may be DASH segments.
[0073] Further, switching information may be received (step 404) wherein the switching information may identify a corresponding video enhancement part for a video base part and one or more locations of one or more video frames in the corresponding video enhancement part that are associated with one or more switching points on the common timeline. Video frames of the corresponding enhancement layer may be encoded such that a video frame associated with a switching point in the video enhancement part can have one or more coding dependencies on video frames in the associated video base part on or preceding the switching point. Further, video frames of the corresponding enhancement layer may be encoded such that a video frame associated with a switching point in the video enhancement part has no coding dependencies on video frames in the video enhancement part preceding the switching point.
[0074] Then, the decoding may be switched to decoding video data of a higher quality (step 406). The switching may include: based on the switching information determining for a current video base part at least one corresponding enhancement part and a position of a video frame in the corresponding enhancement part that is associated with a switching point. Encoded video frames of the corresponding enhancement part may be retrieved, wherein the encoded video frames may comprise encoded video frames having a position on the common timeline on or succeeding the switching point. Then, encoded video frames of the video base part and the encoded video frames of the corresponding enhancement part may be decoded for determining video frames of a quality that is higher than the quality of video frames associated with the base layer.
[0075] Hence, the embodiments provide low latency switching to rendering of video data from a first resolution to a second resolution without the need to wait for the end of a fragment (or an equivalent data container such as a DASH segment) and without the need to substantially increase the bitrate as is known from conventional quality switching schemes.
[0076] Fig. 5 depicts a schematic of an exemplary streaming system for streaming media data based on the low latency quality switching schemes as described with reference to the embodiments in this disclosure. The system may include a content preparation system 502, a server system 504, and media playout device 506 comprising a client device 542, which may communicate via a network 536 (including the Internet) to the server system. In some embodiments the content preparation system may be part of the server system. In other embodiments, the content preparation system may be connected to the server system via e.g. the network 536 or another network, or may be directly communicatively coupled. The server system 504 may include a server processor 532 and one or more network interfaces 534 which are configured to send and receive data via network 536. In some embodiment, functionalities of the server system and the content preparation system may be implemented in the form of a distributed system including a plurality of communicatively connected network devices, including but not limited to routers, bridges, proxy devices, switches, etc. In an embodiment, the server system may be part of a CDN.
[0077] The content preparation system 502 may include a media source 508, e.g. a media storage and / or one or more devices to generate media data, e.g. video and audio sources such as camera’s, microphones and computers for generating, e.g. synthesizing, computer-generated graphics and / or audio. Media data may include video data, e.g. in the form of sequences of video frames, which may be associated other data, e.g. audio data in the form of audio frames and / or metadata such as a recording time and / or information about temporal order of video data, in particular encoded video data. The media server system may include an encoder system 514 comprising one or more encoder instances for encoding the media data.
[0078] The encoder system may produce one or more encoded media data streams, in short media streams. In some embodiments, an individual stream may be referred to as an elementary stream representing a single, digitally coded component (e.g. video or audio) of a media representation. The content preparation system may further include a packetizer 516 for converting elementary streams comprising encoded media data into a packetized stream, e.g. a packetized elementary stream (PES). The PES streams may be formatted, e.g. encapsulated, by an encapsulator 517 for transport so that that encoded media data can be transmitted to the server system using a suitable media streaming standard such as MPEG- DASH, HLS or CMAF and stored as one or more media files at a storage medium of the server system 504. The encapsulator may be configured to generate media files, which are formatted according to a predetermined data format for example CMAF fragment or DASH segments.
[0079] Media data which are encoded in one or more elementary streams according to a certain bitrate or quality may form a media representation or in short a representation. The encoder system 514 may be configured to encode media data of a media title in different ways using a video coding standard to produce different representations of a media title at various bitrates and various characteristics, such as pixel resolutions, frame rates, conformance to various coding standards, etc. These different representations may be used for adaptive bitrate streaming as known from streaming protocols such as DASH. In particular, the encoder system may encode the media data according to any suitable standardized coding scheme such as H.264 / AVC, HEVC or VVC. In an embodiment, the media data may be encoded using a layered coding scheme such as Scalable HEVC or Multilayer VVC. Layered coding is a type of data encoding for digital video or digital audio where the result of encoding source video data is not just one compressed data stream, but multiple streams, referred to as layers (typically a base layer and one or more enhancement layers). During decoding all layers may be combined to recreate an original high-quality video stream. The stream can be decoded even if some layers are missing (although usually a layer hierarchy has to be respected, e.g. a base layer associated with a certain quality must be available). If one or more enhancement layers are missing, the resulting stream may have reduced visual quality, but may still be usable.
[0080] Layered coding is particularly suitable when the same media stream needs to be available in different qualities, for example for adaptive bitrate streaming. Without layered coding, the source video stream needs to be encoded multiple times to obtain compressed streams with different qualities and bitrates. In contrast, layered coding provides the advantage of only encoding a single time, because streams with different qualities can be obtained based on the base layer and the one or more enhancement layers.
[0081] The encapsulator 516 may be configured may be configured to format packets of elementary systems into network abstraction layer (NAL) units. NAL units, which are defined as part of the H.264 / AVC and HEVC video coding standards, include Video Coding Layer (VCL) NAL units comprising video data payload and non-VCL NAL units, which may comprise metadata such as parameter sets (important header data that can apply to a large number of VCL NAL units) and supplemental enhancement information (timing information and other supplemental data that may enhance usability of the decoded video signal). Non- VCL NAL units may include sequence parameter sets (SPS), which apply to a series of consecutive coded video pictures called a coded video sequence and picture parameter sets (PPS), which apply to the decoding of one or more individual pictures within a coded video sequence. Non-VCL NAL units may further include Supplemental Enhancement Information (SEI) messages which may contain information for assisting the decoding process.
[0082] A set of NAL units may define a so-called access unit which together may form a coded picture (a video frame), This way, the decoding of an access unit generally results in one decoded picture (a decoded video frame). A coded video sequence consists of a series of access units that are sequential in the NAL unit stream and use only one sequence parameter set. Each coded video sequence can be decoded independently of any other coded video sequence, given the necessary parameter set information, which may be conveyed "in-band" or "out-of-band". At the beginning of a coded video sequence is an instantaneous decoding refresh (I DR) access unit. An I DR access unit comprises an intra picture (l-frame) which is a coded picture that can be decoded without decoding any previous pictures in the NAL unit stream. The presence of an I DR access unit indicates that no subsequent picture in the stream will require reference to pictures prior to the intra picture it contains in order to be decoded. The encapsulator may use coded video sequences to produce short non-overlapping short video files, such as DASH segments and CMAF fragments, that are used by adaptive streaming protocols to provide adaptive streaming functionality.
[0083] The encapsulator 516 may be further configured to determine one or more manifest files 524 associated with media representations of a media title. An example of a manifest file is a media presentation descriptor (MPD) in case MPEG-DASH is used for streaming the media data. The manifest file identifies the media representations 526-530 that are associated with a media title using a suitable language such as the extensible markup language (XML). In some embodiments, media representations may be divided into so-called adaptation sets. An adaptation set may define media data associated with a common set of characteristics, including but not limited to e.g. codec, profile and level, resolution, number of views, file format for segments, etc. The manifest file may include data identifying such adaptation sets and further information associated with characteristics, such as bitrates, of specific representations of adaptation sets.
[0084] In some embodiments, the manifest file may comprise information about the availability of CMAF fragments or chunks of a representations or DASH segment or subsegments. For example, in an embodiment, the manifest file may include information indicating the wall-clock time at which a first fragment of one of the representations becomes available, as well as information indicating the durations of fragments within representations. This way, the client processor may determine when each fragment is available, based on the starting time as well as the durations of the fragments. The packetized and encapsulated media files, e.g. CMAF fragments and / or DASH segments, associated with a media title prepared by the content preparation system may be stored at the server system as one or more representations 526-530 and an associated manifest file 524.
[0085] For example, the content preparation system may prepare and store CMAF tracks comprising switching sets of time-aligned fragments comprising media data which may be encoded using a layered coding scheme as described with reference to Fig. 2-4. Thus, when encoding the media data, the content preparation system may generate a track of CMAF fragments comprising video data associated with a base layer and one or more tracks comprising CMAF fragments associated with one or more enhancement layers respectively.
[0086] The content preparation system may further generate metadata associated with the base layer wherein the metadata may include switching information for signalling one or more switching points identifying at least one video frame - a switching frame - in one or more CMAF fragments of the enhancement layer. A switching point may be signalled a time stamp associated with the common timeline of the enhancement layer and / or as an offset relative to the start of the CMAF fragment (e.g. an IDR). As explained with reference to Fig. 2 and Fig. 3, when introducing switching points in the media data of the enhancement layer, the content preparation system makes such that the encoder encodes the media data of the enhancement layer such that, when switching to a higher quality at a switching point, coding dependencies associated with video frames of the enhancement layer can be handled by the decoder.
[0087] The server processor 532 may be configured to receive network requests from client devices, such as client device 542. The client device may comprise a client processor 548 configured to request media data and / or metadata, e.g. a manifest filed, that is stored at the server system. Based on a manifest file 546 stored at the client device, the client device may request (via the client processor) media data from the server system and store the media data in a client buffer 544. The buffered (encoded) media data may be provided to the input of a decoder system 550 which may be configured to execute one or more decoder instances for decoding the encoded media data in to decoded media data, which may be rendered by a display device 552. The display device may be implemented as any type of display devices, including display devices such as a head-mounted device for rendering XR- type of media data (e.g. tiled 360 video data).
[0088] The client device may receive information about the decoding capabilities of video decoder and rendering capabilities of the media playout device. The functionality of the client device or portions of the functionality of the client device may be implemented in hardware, or a combination of hardware, software, and / or firmware, where requisite hardware may be provided to execute instructions for software or firmware. Further, in some embodiments, the client device, the decoder system and the display device may be implemented in different devices, e.g. the client device may be implemented in a media box and the decoder and the display maybe implemented in a display system.
[0089] The server processor 532 and client processor 548 may be implemented to process requests based on the hypertext transfer protocol (HTTP), for example HTTP 1.1, HTTP 2, HTTP 3 and future versions. For example HTTP version 1.1, allows transmission of encoded media data based on the chunked transfer encoding mode. This way, the server request processor may be configured to receive HTTP messages, such as HTTP GET or partial GET requests and send media data in response to the requests back to the client device. The requests may specify a video file or a part of a specific part, e.g. a fragment or a chunk of a fragment, of one of the representations 526-530, e.g., using a resource locator, such as an URL. In some examples, the requests may also specify a byte range for identifying a chunk in a fragment. Instead of HTTP other client-server communication protocols may be used to handle request and response messages. For example, in an embodiment, a bi-directional communication channel between the client and the sever system may be realized based on a WebSocket protocol. In that case, a handshake request may be used to set up a WebSocket connection between the client and server. Request and response messages may be exchanged between the client and the server over the WebSocket connection. Other protocols that may be used to communicate between the server and the client include Long polling, WebRTC, SignaIR, FTP, MQPTT, etc.
[0090] Additionally and / or alternatively, the request processors of the server system and the client device may be configured to deliver media data via a broadcast or multicast protocol. In that case, the server device may deliver encapsulated media data (e.g. segments or fragments or parts thereof) using a broadcast or multicast network transport protocol. For example, the server request processor may be configured to receive a multicast group join request from the client device. For example, the server system may advertise an Internet protocol (IP) address associated with a multicast group to client device 542, associated with particular a media title. The client device may subsequently submit a request to join the multicast group. This request may be propagated throughout network 536, e.g., routers defining a network, such that the routers are configured to media data destined for the IP address associated with the multicast group to the client device that has joined the multicast group.
[0091] The client device may request encoded video data from the server system based on the information in the manifest file and based on information about the available network bandwidth. Higher bitrate representations may yield higher quality video playback, while lower bitrate representations may provide sufficient quality video playback when available network bandwidth decreases. Accordingly, when available network bandwidth is relatively high, the client device may retrieve media data associated with (relatively) high bitrate representations, whereas when available network bandwidth is low, the client device may request media data associated with (relatively) low bitrate representations. This way, media data may be streamed over the network to the client device while adapting to changing network bandwidth availability of network.
[0092] The network interface 540 of the client device may receive media and buffer media, e.g. encapsulated and packetized media data such as CMAF fragments or chunks of a selected representation. The client processor 548 may be configured to decapsulate the media files in to PES streams and depacketize the PES streams into encoded media data which may be to the decoder system. The decoded media data may be subsequently sent to display device 552.
[0093] The devices and modules depicted in Fig. 5 such as encoder, packetizer, encapsulator, server processor, client processor, etc. may be implemented as any of a variety of suitable processing circuitry, as applicable, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware or any combinations thereof. Alternatively and / or additionally these devices and modules may comprise an integrated circuit, a microprocessor, and / or a wireless communication device.
[0094] In further embodiments, more than one enhancement layer may be used to provide switching between multiple qualities. In case of a plurality of enhancement layers, to enable quality switching between more than one layer representations, the dependencies of enhancement layers with other enhancement layers and with base layer need to be constrained so that upon switching the decoder can handle dependencies associated with video frames of the new enhancement layer. Depending on these constrains, different implementations of quality switching quality can be realized. Below some non-limiting examples of quality switching based on a plurality of enhancement layers are provided.
[0095] Fig. 6 depicts an example of time-aligned fragments for low latency switching according to a further embodiment. In particular, the figure illustrates CMAF tracks defining a switching set of time-aligned fragments including a base layer 601 and N enhancement layers. The example shows two (N=2) enhancement layers, a first enhancement layer 6OO1 associated with a first quality comprising CMAF fragments 604i^ and a second enhancement layer 6OO1 associated with a second quality comprising CMAF fragments 603i^.
[0096] Combining video data of the base layer with video data of the first enhancement layer will result - after decoding - in video frames of a first quality which is higher than a base quality, i.e. a quality of video frames which are generated based on decoding video data of the base layer only. In a similar way, combining video data of the base layer with video data of the first enhancement layer and second enhancement layer will result - after decoding - in video frames of a second quality which is higher than video frames of the first quality and video frames of the base quality.
[0097] The CMAF fragments may be divided into chunks in a similar way as described with reference to Fig. 2. The switching information for all enhancement layers may be provided as metadata to the base layer. This allows instant switching from a representation only associated with the base layer to a representation of an enhancement layer N by fetching chunks associated with the enhancement layer N and all the chunks associated with enhancement layers between the base layer and the enhancement layer N. This allows to immediate switch to any representation but the gives more signaling overhead in the base layer.
[0098] In this embodiment, the constrains of the first and second dependencies may be similar to the dependencies explained with reference to Fig. 2 wherein at a switching point video data of an enhancement layer depend only on video data of current and past CMAF chunks of the base layer. This way, first encoding dependencies 610 exists between the base layer and the first enhancement layer and second encoding dependencies 612 between the base layer and the second enhancement layer. Fig. 7 depicts an example of time-aligned fragments for low latency switching according another embodiment. The figure illustrates CMAF tracks defining a switching set of time-aligned fragments including a base layer 701, comprising CMAF fragments 704i^, and N enhancement layers. The example shows two (N=2) enhancement layers, a first enhancement layer 700i associated with a first quality comprising CMAF fragments 704i^ and a second enhancement layer 700i associated with a second quality comprising CMAF fragments 703i^ similar to the example in Fig. 6.
[0099] In this embodiment, the switching information may be configured to provide incremental switching of representation, wherein switching information of a chunk for enhancement layer N is provided in metadata associated with chunks of enhancement layer N-1. Hence, switching information for chunks of an enhancement layer N can be signaled using metadata associated with chunks of another lower enhancement layer, for example enhancement layer N-1. Hence, based on the metadata in the base layer the starting point of a chunk in the first enhancement layer 700i is determined and fetched. Then, based on the metadata in the received chunk of the first enhancement layer the starting point of a chunk in the second enhancement 700s is determined and fetched.
[0100] This way, video data of enhancement layers may be incrementally fetched by extracting switching information from an enhancement layer N-1 to fetch the corresponding chunks in the enhancement layer N. In this embodiment, the switching information for switching to higher representation may be distributed over the different layers but provide less flexibility to switch immediately to any higher representation.
[0101] In this embodiment, the constrains of the first and second dependencies may be similar to the dependencies explained with reference to Fig. 2 wherein at a switching point video data of an enhancement layer depend only on video data of current and past CMAF chunks of the base layer. Hence, first encoding dependencies 710 exists between the base layer and the first enhancement layer and second encoding dependencies 712 between the base layer and the second enhancement layer. This way, it is possible to directly switch to any enhancement layer in the switching set.
[0102] Fig. 8 depicts an example of time-aligned fragments for low latency switching according another embodiment. The figure illustrates CMAF tracks defining a switching set of time-aligned fragments including a base layer 801, comprising CMAF fragments 804i^, and N enhancement layers. The example shows two (N=2) enhancement layers, a first enhancement layer 8OO1 associated with a first quality comprising CMAF fragments 804i^ and a second enhancement layer 8OO2 associated with a second quality comprising CMAF fragments 803i^ similar to the example in Fig. 6.
[0103] In this embodiment, a chunk in enhancement layer N depends on present and past CMAF chunks of the enhancement layer N-1 and so forth. Hence, in this case, it is only possible to only switch quality in the middle of a fragment to a representation that is one level higher. For example, if the client starts with a representation of the base layer 801 , it can only switch quality to the representation of the firs enhancement layer 8OO1 by fetching corresponding chunks of enhancement layer 1. If the client starts with the base layer and the first enhancement layer it can only switch quality to the representation of the next, second enhancement layer by fetching corresponding chunks of enhancement layer 2. In this case, metadata associated with the CMAF chunks of the enhancement layer N-1 may comprise switching information for the CMAF chunks of the enhancement layer N. This information may for example be signalled in an event message (i.e. ‘emsg’).
[0104] It is submitted that different aspects of the switching sets described with reference to Fig. 6-8 may be combined. Other implementations than those illustrated in the figures are also possible.
[0105] Hence, the CMAF low latency switching set described with reference to the embodiments in this disclosure comprise one base layer and at least one enhancement layer, where CMAF chunks of the enhancement layer depend only on CMAF chunks of the base layer. Further, preferably CMAF fragment header parameters are the same between CMAF fragments of the base layer and CMAF fragments of the enhancement layer. If this is not the case, metadata may be signaled so that during switching a client device can still fetch the CMAF fragment headers. Further, the CMAF chunks of the base layer may contain switching information associated with CMAF chunks of the enhancement layer. When using multiple enhancement layers, the metadata may also be associated with (part of) these enhancement layers.
[0106] As described above, the switching information may be signalled in an event message (i.e. ‘emsg’). In an embodiment, switching information may be signalled in-band in a DASH in-band event message cmafllqswitch. An exemplary syntax of such DASH in-band event message for CMAF low latency switching set may look as follows: aligned(8) class DASHEventMessageBox extends FullBox('emsg', version, flags=0) if (version==0) { string scheme_id_uri; string value; unsigned int(32) timescale; unsigned int(32) presentation_time_delta; unsigned int(32) event_duration; unsigned int(32) id;
[0107] } else if (version==1) { unsigned int(32) timescale; unsigned int(64) presentation_time; unsigned int(32) event_duration; unsigned int(32) id; string scheme_id_uri; string value;
[0108] } unsigned int(8) message_data[ ];
[0109] }
[0110] Some of the parameters in the event message are described hereunder in more detail. Since CMAF chunks of the base layer comprise address information of CMAF chunks or parts of CMAF fragments in the enhancement layer, it is not required to have the same in-band message inserted in both representations. So the AdaptationSet@segmentAlignment should be cleared to false.
[0111] The event message parameter scheme_id_uri may define an identifier, an URI, for example, "urn:mpeg:dash:event:cmafllqswitch:202x" for signalling the client device that low latency switching is enabled and that the metadata associated with the base layer comprises switching information, e.g. one or more event messages, for switching to a higher quality based on media data of the base layer and an enhancement layer.
[0112] The event message parameter presentati on_time defines a time instance on the common timeline that is linked to a video frame in the enhancement layer. Typically, the time instance will point to the start of a chunk in a CMAF fragment of the enhancement layer.
[0113] The event message parameter value relates to a byte range for fetching a portion of a fragment of an enhancement layer that is needed to make a desired quality switch. Different values may define different ways of providing the byte range. Other addressing schemes are possible and the value should be extended to cover the schemes.
[0114] The event message parameter message_data[] defines the ID of the corresponding enhancement fragment, and byte range of the desired part of the enhancement fragment when desired to switch to higher quality at the next chunk.
[0115] Table 1 provides examples of schemes for signalling the CMAF fragment and a part in the CMAF fragment of the enhancement layer that is needed for a quality switch based on the value parameter in the event message.
[0116] Table 1 - interpretation of values field of the cmafllqswitch (CMAF low latency quality switch) event
[0117] Hence, as shown in this table, if the event message signals value=1 then the message_data[ 7 field contains the ID of the corresponding enhancement fragment, and the absolute address of the desired part of the enhancement fragment to use for the byte-range request to switch quality at the next chunk. Similarly, when value=3 the message_data[ ] field contains the ID of the corresponding enhancement fragment, and the offset address (instead of an absolute address) of the desired part of the enhancement fragment
[0118] When the client device receives an event message, it only uses this event when it desires to switch to higher quality in the next chunk, otherwise this event is skipped. When switching to higher quality is desired, the client device may process the event and uses the byte ranges in the event to make an HTTP range request to get the corresponding part of enhancement layer fragment that corresponds to the enhancement layer and perform switch to higher quality in the next chunk.
[0119] While the embodiments in this disclosure are illustrated based on implementation in CMAF, the embodiments are not limited thereto. For example, the enhancement layers may also be implemented as a sequence of DASH segments. Selfinitialized DASH segments, i.e. self-decodable DASH segments that start with an l-frame) can be further divided to subsegments. Segmentations and subsegmentations may be performed in a ways that make switching simpler. According to clause 4.3 of ISO / IEC 23009- 1 “Segmentation and Subsegmentation may be performed in ways that make switching simpler. For example, in the very simplest cases, each Segment or Subsegment begins with a SAP and the boundaries of Segments or Subsegments are aligned across the Representations of one Adaptation Set
[0120] In this case, switching Representation involves playing to the end of a (Sub)Segment of one Representation and then playing from the beginning of the next (Sub)Segment of the new Representation.”. The way DASH provides a way to switch between subsegments is like segments. When the subsegments are time aligned and start with a SAP frame of type 1 or 2, then switching between different representations can be achieved seamlessly. Hence, similar to the situation in Fig.1 , low latency switching would require division of the video data in small subsegments, wherein each subsegments includes an l-frame. Decreasing the length of the DASH (sub) segments, would require more SAP frames of type 1 or 2 in the bitstream, leading to excessive bitrates.
[0121] In order to address this problem, in an embodiment, an enhancement layer may be formatted as DASH segments (which may not be compliant with the CMAF standard) and switching information in the metadata of the base layer may be used to identify switching frames in the enhancement layer. Similarly, an enhancement layer may be formatted as DASH segments comprising sub-segments which do not start with an I frame. That is, even if the subsegments do not conform to SAP 1 or 2, switching in sub-segments can still be enabled by using layered coding and using switching information (using offsets or absolute address of the subsegments of the enhancement layer) in the metadata of the base layer subsegment. The dependency between layers can still be specified in the manifest (e.g., flag @dependencylD in DASH MPD).
[0122] Hence, the embodiments in this disclosure generally allow different non-CMAF based DASH implementations which uses sub-segmentations to still benefit from this solution by allowing them to have subsegments to start with non-SAP frames and achieve low latency switching to high quality without the need to wait for SAP boundaries and gain similar advantage that can be provided by the CMAF embodiments described with reference to the figures in this disclosure.
[0123] The devices and systems described with reference to embodiments in this disclosures, such as the client device, the server system and the content preparation system are typically implemented as one or more communicatively connected data processing systems. Fig. 9 is a block diagram illustrating an exemplary data processing system that may be used in as described in this disclosure. Data processing system 900 may include at least one processor 902 coupled to memory elements 904 through a system bus 906. As such, the data processing system may store program code within memory elements 904. Further, processor 902 may execute the program code accessed from memory elements 904 via system bus 906. In one aspect, data processing system may be implemented as a computer that is suitable for storing and / or executing program code. It should be appreciated, however, that data processing system may be implemented in the form of any system including a processor and memory that is capable of performing the functions described within this specification.
[0124] Memory elements 904 may include one or more physical memory devices such as, for example, local memory 908 and one or more bulk storage devices 910. Local memory may refer to random access memory or other non-persistent memory device(s) generally used during actual execution of the program code. A bulk storage device may be implemented as a hard drive or other persistent data storage device. The data processing system 900 may also include one or more cache memories (not shown) that provide temporary storage of at least some program code in order to reduce the number of times program code must be retrieved from bulk storage device 910 during execution.
[0125] Input / output (I / O) devices depicted as input device 912 and output device 914 optionally can be coupled to the data processing system. Examples of input device may include, but are not limited to, for example, a keyboard, a pointing device such as a mouse, or the like. Examples of output device may include, but are not limited to, for example, a monitor or display, speakers, or the like. Input device and / or output device may be coupled to data processing system either directly or through intervening I / O controllers. A network adapter 916 may also be coupled to data processing system to enable it to become coupled to other systems, computer systems, remote network devices, and / or remote storage devices through intervening private or public networks. The network adapter may comprise a data receiver for receiving data that is transmitted by said systems, devices and / or networks to said data and a data transmitter for transmitting data to said systems, devices and / or networks. Modems, cable modems, and Ethernet cards are examples of different types of network adapter that may be used with data processing system.
[0126] As pictured in Fig. 9, memory elements 904 may store an application 918. It should be appreciated that data processing system may further execute an operating system (not shown) that can facilitate execution of the application. Application, being implemented in the form of executable program code, can be executed by data processing system, e.g., by processor 902. Responsive to executing application, data processing system may be configured to perform one or more operations to be described herein in further detail.
[0127] In one aspect, for example, data processing system may represent a client data processing system. In that case, application 918 may represent a client application that, when executed, configures data processing system to perform the various functions described herein with reference to a "client". Examples of a client can include, but are not limited to, a personal computer, a portable computer, a mobile phone, or the like. In other aspects, data processing system may represent a server data processing system. In that case, application 918 may represent a server application that, when executed, configures data processing system to perform the various functions described herein with reference to a "server".
[0128] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a codec hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.
[0129] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0130] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiment was chosen and described in order to best explain the principles of the invention and the practical application, and to enable others of ordinary skill in the art to understand the invention for various embodiments with various modifications as are suited to the particular use contemplated.
Claims
CLAIMS1. Method of processing video data including: receiving by a client apparatus one or more video base parts, each video base part comprising encoded video frames representing video data of a base layer and each video base part being associated with one or more corresponding video enhancement parts, the one or more corresponding video enhancement parts comprising encoded video frames representing video data of one or more enhancement layers respectively for enhancing video data of the base layer, the encoded video frames of the one or more video base parts and the one or more video enhancement parts having a common timeline; receiving switching information, the switching information identifying a corresponding video enhancement part for a video base part and one or more locations of one or more video frames in the corresponding video enhancement part that are associated with one or more switching points on the common timeline, wherein video frames of the corresponding enhancement layer are encoded such that a video frame associated with a switching point in the video enhancement part can have one or more coding dependencies on video frames in the associated video base part at or preceding the switching point and has no coding dependencies on video frames in the video enhancement part preceding the switching point; and, switching to decoding video data of a higher quality, the switching including:- based on the switching information determining for a current video base part at least one corresponding enhancement part and a position of a video frame in the corresponding enhancement part that is associated with a switching point;- retrieving encoded video frames of the corresponding enhancement part, the encoded video frames comprising encoded video frames having a position on the common timeline at or succeeding the switching point;- decoding encoded video frames of the video base part and the encoded video frames of the corresponding enhancement part for determining video frames of a quality that is higher than the quality of video frames associated with the base layer.
2. Method according to claim 1 wherein the one or more video base parts and / or the one or more video enhancement parts are formatted as Common Media Application Format (CMAF) fragments.
3. Method according to claims 1 or 2 wherein each CMAF fragment is subdivided into a plurality of non-overlapping CMAF chunks, preferably the CMAF chunks being received using an HTTP chunked transfer mode.
4. Method according to claim 3 wherein the one or more locations of one or more video frames in the corresponding video enhancement part that is associated with the one or more switching points defining the start of one or more HTTP chunks in the corresponding enhancement part.
5. Method according to any of claims 1-4 wherein the switching information includes an identifier for identifying the at least one corresponding enhancement part and / or a parameter value, preferably a byte range, indicative of the position of the video frame in the corresponding enhancement part that is associated with the switching point.
6. Method according to any of claims 1-5, wherein the encoded video frames of the corresponding enhancement part are retrieved by the client apparatus based on a manifest file, preferably a media presentation description (MPD), comprising one or more video part identifiers for identifying the one or more video base parts and / or one or more corresponding video enhancement parts associated with each of the one or more video base parts.
7. Method according to claim 6 wherein retrieving encoded video frames of the corresponding enhancement part includes: determining a resource locator based on the one or more video part identifiers for locating the corresponding enhancement part; and, determining a request message, preferably a HTTP byte range request message, based on the resource locator and information associated with the position of a video frame in the corresponding enhancement part for requesting the encoded video frames of the corresponding enhancement part.
8. Method according to any of claims 1-7 wherein at least part of the switching information is provided to the client apparatus using one or more in-band messages, e.g. one or more DASH in band event messages, in at least part of the one or more video base parts and / or in one or more chunks of the one or more video base parts, and, optionally, in at least part of the one or more corresponding enhancement parts and / or in one or more chunks of the one or more corresponding enhancement parts.
9. Method according to any of claim 1-8 wherein the client apparatus requests the video enhancement part associated with the video base part if the client apparatus receives a request for switching to a higher quality and / or if the client apparatus determines to switch to a higher quality.
10. Method according to any of claims 1-9 wherein each of the one or more video base parts and / or one or more video enhancement parts comprise at least a group of pictures (GOP), wherein the video frames of a GOP have no coding dependency on video frames outside of the GOP; and / or wherein the video data are encoded using a layered coding scheme, e.g. scalable HEVC or multilayer VVC.
11. An apparatus for processing video data comprising: a computer readable storage medium having at least part of a program embodied therewith; and, a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, the processor is configured to perform executable operations comprising: receiving by a client apparatus one or more video base parts, each video base part comprising encoded video frames representing video data of a base layer and each video base part being associated with one or more corresponding video enhancement parts, the one or more corresponding video enhancement parts comprising encoded video frames representing video data of one or more enhancement layers respectively for enhancing video data of the base layer, the encoded video frames of the one or more video base parts and the one or more video enhancement parts having a common timeline; receiving switching information, the switching information identifying a corresponding video enhancement part for a video base part and one or more locations of one or more video frames in the corresponding video enhancement part that are associated with one or more switching points on the common timeline, wherein video frames of the corresponding enhancement layer are encoded such that a video frame associated with a switching point in the video enhancement part can have one or more coding dependencies on video frames in the associated video base part at or preceding the switching point and has no coding dependencies on video frames in the video enhancement part preceding the switching point; and, switching to decoding video data of a higher quality, the switching including:- based on the switching information determining for a current video base part at least one corresponding enhancement part and a position of a video framein the corresponding enhancement part that is associated with a switching point;- retrieving encoded video frames of the corresponding enhancement part, the encoded video frames comprising encoded video frames having a position on the common timeline at or succeeding the switching point;- decoding encoded video frames of the video base part and the encoded video frames of the corresponding enhancement part for determining video frames of a quality that is higher than the quality of video frames associated with the base layer.
12. Apparatus according to claim 11 wherein the processor is further configured to perform any of the method steps according to claims 2-10.
13. A video preparation module comprising: a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, the processor is configured to perform executable operations comprising: encoding video data into encoded video frames representing video data of a base layer and encoded video frames representing video data of one or more enhancement layers for enhancing video data of the base layer, packaging encoded video frames representing video data of the base layer into one or more video base parts and packaging encoded video frames representing video data of one or more enhancement layers into one or more video enhancement parts, each video base part being associated with one or more corresponding video enhancement parts, wherein encoded video frames of the one or more video base parts and the one or more video enhancement parts have a common timeline; generating switching information, the switching information identifying a corresponding video enhancement part for a video base part and one or more locations of one or more video frames in the corresponding video enhancement part that are associated with one or more switching points on the common timeline, wherein video frames of the corresponding enhancement layer are encoded such that a video frame associated with a switching point in the video enhancement part can have one or more coding dependencies on video frames in the associated video base part at or preceding the switching point and has no coding dependencies on video frames in the video enhancement part preceding the switching point; and,storing the one or more video base parts of the base layer, the one or more video enhancement parts of the one or more enhancement layers and the switching information on a storage medium.
14. A computer readable storage medium comprising stored data thereon, the data defining a data format for steaming video data to a client device comprising, the data including: one or more video base parts, each video base part comprising encoded video frames representing video data of a base layer; one or more video enhancement parts, each video enhancement part comprising encoded video frames representing video data of one or more enhancement layers, each video base part being associated with one or more corresponding video enhancement parts, wherein encoded video frames of the one or more video base parts and the one or more video enhancement parts have a common timeline; and, switching information for the client device, the switching information identifying a corresponding video enhancement part for a video base part and one or more locations of one or more video frames in the corresponding video enhancement part that are associated with one or more switching points on the common timeline, wherein video frames of the corresponding enhancement layer are encoded such that a video frame associated with a switching point in the video enhancement part can have one or more coding dependencies on video frames in the associated video base part at or preceding the switching point and has no coding dependencies on video frames in the video enhancement part preceding the switching point.
15. Computer program product comprising software code portions configured for, when run in the memory of a computer, executing the method steps according to any of 1-10.
Citation Information
Patent Citations
Block-level super-resolution based video coding
WO2019197674A1
Efficient hypertext transfer protocol (HTTP) adaptive bitrate (ABR) streaming based on scalable video coding (SVC)
US20230283789A1
An apparatus, a method and a computer program for video coding and decoding
WO2018146376A1