Low-latency quality switching for adaptive video streaming

CN122720131APending Publication Date: 2026-09-08KONINK KPN NV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480087478.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-30
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

然而,这种解决方案将要求过高的比特率

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122720131A_ABST
    Figure CN122720131A_ABST
Patent Text Reader

Abstract

A method and apparatus for processing video data are described, wherein the method includes: receiving one or more video base portions by a client device, each video base portion including encoded video frames representing base layer video data, and each video base portion being associated with one or more corresponding video enhancement portions, the one or more corresponding video enhancement portions including encoded video frames representing one or more enhancement layer video data respectively used to enhance the base layer video data, the encoded video frames of the one or more video base portions and the one or more video enhancement portions having a common timeline; receiving switching information, the switching information identifying the corresponding video enhancement portion used for the video base portion, and the corresponding video enhancement portion associated with one or more switching points on the common timeline. One or more positions of one or more video frames in a strong portion, wherein the video frames corresponding to the enhancement layer are encoded such that: the video frames associated with a switching point in the video enhancement portion have one or more encoding dependencies on video frames at or before the switching point in the associated video base portion, and have no encoding dependencies on video frames before the switching point in the video enhancement portion; and, switching to decoding higher quality video data, the switching including, based on switching information, determining at least one corresponding enhancement portion and the position of the video frame associated with the switching point in the corresponding enhancement portion for the current video base portion; retrieving encoded video frames of the corresponding enhancement portion, the encoded video frames including encoded video frames on a common timeline at or after the switching point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments relate to low-latency quality switching for adaptive video streaming, and particularly, but not exclusively, to methods and systems for low-latency quality switching for adaptive video streaming, video data structures for low-latency adaptive video streaming, and computer program products for performing such methods. Background Technology

[0002] Quality adaptation based on network conditions and client resources is crucial for maintaining an uninterrupted video timeline. This is especially true for immersive multimedia systems, where such quality adaptation should be performed with very low latency to maintain a high quality of experience (QoE) for the user. For example, in 360-degree tiled streaming or more generally, in XR content consumption where only a portion of the content is inspected in high quality, user actions can trigger the consumption of another portion of the content, delivered only in low quality. In such cases, being able to obtain a high-quality version of that portion of the content as quickly as possible is essential to avoid perceived quality interruptions for the user.

[0003] Quality switching is addressed in media standards such as HTTP-based Dynamic Adaptive Streaming (MPEG-DASH), HTTP Live Streaming (HLS), and Common Media Application Format (CMAF). For example, CMAF defines the data format used to encode and package media data for adaptive media streaming applications. Quality switching in CMAF can be performed by adaptively switching between CMAF tracks associated with different rendering qualities. Each CMAF track includes encoded video data in a CMAF segment that can be requested by the client device. Different CMAF tracks associated with encoded video data of different qualities define CMAF switching. By default, the client device needs to wait until the CMAF segment ends for quality adaptation, such as switching to a higher video quality. This latency hinders the optimal user experience.

[0004] Therefore, in CMAF, quality switching is performed at the CMAF segment level, meaning that if faster quality switching is desired, the CMAF segment length should be smaller. Considering that each CMAF segment should begin with an I-frame (or an I-frame-like frame), reducing the CMAF segment length introduces more I-frames into the bitstream, resulting in a substantial increase in bitrate. In extreme cases where applications require quality switching at the granularity of a single frame or a few frames, it will be necessary to use single frames as CMAF segments. In other words, a fully intra-coded bitstream would be required, where each frame is placed within a CMAF segment. However, this solution would demand excessively high bitrates.

[0005] Therefore, as can be seen from the above, there is a need in the art for improved methods and systems that provide low-latency quality switching for adaptive streaming applications. Summary of the Invention

[0006] As those skilled in the art will appreciate, aspects of the present invention can be implemented as systems, methods, or computer program products. Therefore, aspects of the present invention can take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects, all of which are generally referred to herein as “circuit,” “module,” or “system.” The functionality described in this disclosure can be implemented as algorithms executed by a computer’s microprocessor. Furthermore, aspects of the present invention can take the form of computer program products implemented on one or more computer-readable media having computer-readable program code implemented thereon (e.g., stored thereon).

[0007] Any combination of one or more computer-readable media may be used. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media will include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the context of this document, a computer-readable storage medium may be any tangible medium that may contain or store programs for use by or in connection with an instruction execution system, device, or apparatus.

[0008] A computer-readable signal medium may contain a propagated data signal in which computer-readable program code is implemented, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be a computer-readable medium of any non-computer-readable storage medium, and it may communicate, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or apparatus.

[0009] Program code implemented on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic, cable, RF, etc., or any suitable combination of the foregoing. Computer program code for performing the operations of various aspects of the invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Java™, Smalltalk, C++, or similar languages) and conventional procedural programming languages ​​(such as the "C" programming language or similar programming languages). The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0010] The aspects of the invention are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor (particularly a microprocessor or central processing unit (CPU)) of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of said computer, other programmable data processing apparatus, or other means, create components for implementing the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams.

[0011] These computer program instructions may also be stored in a computer-readable medium that directs a computer, other programmable data processing device or other means to function in a particular manner such that the instructions stored in the computer-readable medium produce an article of writing comprising instructions that implement the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0012] The computer program instructions may also be loaded onto a computer, other programmable data processing device, or other apparatus to cause a series of operational steps to be performed on the computer, other programmable device, or other apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable device, provide a process for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. Furthermore, the instructions may be executed by any type of processor, including, but not limited to, one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits.

[0013] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, comprising one or more executable instructions for implementing one or more specified logical functions. It should also be noted that in some alternative implementations, the functions indicated in the blocks may not occur in the order indicated in the drawings. For example, two blocks shown consecutively may, in fact, execute substantially concurrently, or these blocks may sometimes execute in reverse order, depending on the functionality involved. It will also be noted that each block illustrated in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0014] In one aspect, embodiments may relate to a method for processing video data, which may include one or more of the following steps: receiving one or more video base portions by a client device, each video base portion including encoded video frames representing base layer video data, and each video base portion being associated with one or more corresponding video enhancement portions, the one or more corresponding video enhancement portions including encoded video frames representing one or more enhancement layer video data respectively used to enhance the base layer video data, the one or more video base portions and the one or more video enhancement portions having a common timeline; receiving switching information, the switching information identifying one or more positions of one or more video frames in the corresponding video enhancement portion used for the video base portion, and associated with one or more switching points on the common timeline, wherein the video frames of the corresponding enhancement layers are encoded such that: The video frames associated with a switch point in the video enhancement portion can have one or more coding dependencies on video frames at or before the switch point in the associated video base portion, and such that the video frames associated with the switch point in the video enhancement portion do not have coding dependencies on video frames before the switch point in the video enhancement portion; and switch to decoding higher quality video data, the switch comprising: determining, based on the switch information, at least one corresponding enhancement portion and the position of the video frame associated with the switch point in the corresponding enhancement portion for the current video base portion; retrieving the encoded video frames of the corresponding enhancement portion, the encoded video frames including encoded video frames located at or after the switch point on the common timeline; and decoding the encoded video frames of the video base portion and the encoded video frames of the corresponding enhancement portion to determine a higher quality video frame compared to the quality of the video frames associated with the base layer.

[0015] In an embodiment, the video frames of the corresponding enhancement layer can be encoded such that the video frames after the switching point in the enhancement portion have no encoding dependency on the video frames before the switching point in the video segment of the enhancement layer.

[0016] In an embodiment, the video frames of the corresponding enhancement layer may be encoded such that the video frames associated with the switching point in the video enhancement portion have an encoding dependency only on one or more video frames at or before the switching point in the associated video base portion.

[0017] The method allows a switch from a first low-latency resolution to a second higher resolution at one of the one or more switching points signaled in the metadata associated with the one or more video base components (e.g., one or more CMAF segments carrying video data of the base layer). This eliminates the need for the video processing device (e.g., a client device) to wait until the end of the video base component (such as a CMAF segment) for quality adaptation of the video data played by the client device. For this purpose, a layered coding scheme comprising a base layer and one or more enhancement layers is used, where the base layer carries metadata to obtain a portion of the enhancement layer that depends on the video data already delivered to the base layer. The method allows for bitrate reduction by allowing the use of longer data containers (e.g., CMAF segments) and larger GOP sizes without increasing the latency for switching to higher quality. In other words, the method provides better coding efficiency compared to using smaller CMAF segments to maintain the same latency during quality switching.

[0018] In an embodiment, the one or more video base portions and / or the one or more video enhancement portions may be formatted as Common Media Application Format (CMAF) fragments.

[0019] In an embodiment, each CMAF segment can be subdivided into multiple non-overlapping CMAF blocks.

[0020] In this embodiment, the CMAF block can be received using HTTP chunked transfer mode.

[0021] In an embodiment, each of the one or more switching points may coincide with the start of a CMAF block in one of the one or more video enhancement sections (e.g., the first video frame of the CMAF block).

[0022] In an embodiment, the one or more video base portions and / or the one or more video enhancement portions may be formatted as DASH segments.

[0023] In this embodiment, the CMAF block can be received using HTTP chunked transfer mode.

[0024] In an embodiment, the position of one or more video frames associated with one or more switching points in the corresponding video enhancement portion may define the start of one or more blocks in the corresponding enhancement portion (e.g., the first video frame of the one or more CMAF blocks).

[0025] In an embodiment, the switching information may include an identifier or information for determining the identifier, the identifier being used to identify the at least one corresponding enhancement.

[0026] In an embodiment, the switching information may include parameter values ​​indicating the position of the video frame associated with the switching point in the corresponding enhancement portion.

[0027] In this embodiment, the parameter value may define a byte range within the corresponding enhancement portion. The byte range may define the position of the video frame relative to the beginning of the corresponding enhancement portion.

[0028] In an embodiment, the encoded video frames of the corresponding enhancement portion can be retrieved by a client device based on a manifest file (such as a Media Presentation Description (MPD)) that includes one or more video portion identifiers for identifying the one or more video base portions and / or one or more corresponding video enhancement portions associated with each of the one or more video base portions.

[0029] In an embodiment, retrieving the encoded video frame of the corresponding enhancement portion may include: determining a resource locator based on the one or more video portion identifiers for locating the corresponding enhancement portion; and determining a request message based on the resource locator and information associated with the position of the video frame in the corresponding enhancement portion for requesting the encoded video frame of the corresponding enhancement portion. In an embodiment, the request message may be an HTTP message, such as an HTTP byte range request message.

[0030] In an embodiment, at least a portion of the switching information may be provided to the client device using one or more in-band messages (e.g., one or more DASH in-band event messages). In an embodiment, at least a portion of the one or more in-band messages may be inserted into the one or more video base portions, and optionally, into at least a portion of the one or more corresponding enhancement portions.

[0031] In an embodiment, at least a portion of the one or more in-band messages may be inserted into one or more blocks of the one or more video base portions.

[0032] In another embodiment, at least a portion of the one or more in-band messages may be inserted into at least a portion of one or more blocks of the one or more corresponding enhancement portions.

[0033] In an embodiment, if the client device receives a request to switch to higher quality, and / or if the client device determines to switch to higher quality, the client device may request a video enhancement portion associated with the video base portion.

[0034] In an embodiment, each of the one or more video base portions and / or one or more video enhancement portions may include at least one group of pictures (GOP), wherein video frames of the GOP have no encoding dependency on video frames outside the GOP.

[0035] In this embodiment, the video data may be encoded using a layered coding scheme (e.g., scalable HEVC or multi-layer VVC).

[0036] In one aspect, embodiments of this disclosure may relate to an apparatus for processing video data, comprising: a computer-readable storage medium having at least a portion of a program implemented thereon; and a computer-readable storage medium having computer-readable program code implemented thereon; and a processor, preferably a microprocessor, coupled to the computer-readable storage medium, wherein, in response to executing the computer-readable program code, the processor may be configured to perform one or more of the following executable operations: receiving by a client device one or more video base portions, each video base portion including encoded video frames representing base layer video data, and each video base portion being associated with one or more corresponding video enhancement portions, the one or more corresponding video enhancement portions including encoded video frames representing one or more enhancement layer video data respectively used to enhance the base layer video data, the encoded video frames of the one or more video base portions and the one or more video enhancement portions having a common timeline; receiving switching information, the switching information identifying a corresponding video for the video base portion. The enhancement portion, and one or more positions of one or more video frames in the corresponding video enhancement portion associated with one or more switching points on the common timeline, wherein the video frames of the corresponding enhancement layer are encoded such that: the video frames associated with the switching points in the video enhancement portion have one or more encoding dependencies on video frames at or before the switching points in the associated video base portion, and have no encoding dependencies on video frames before the switching points in the video enhancement portion; and switch to decoding higher quality video data, the switching comprising: determining at least one corresponding enhancement portion and the position of the video frame associated with the switching point in the corresponding enhancement portion for the current video base portion based on the switching information; retrieving encoded video frames of the corresponding enhancement portion, the encoded video frames including encoded video frames with positions at or after the switching points on the common timeline; and decoding the encoded video frames of the video base portion and the encoded video frames of the corresponding enhancement portion to determine higher quality video frames compared to the quality of video frames associated with the base layer.

[0037] In an embodiment, the processor may be further configured to perform any of the method steps described with reference to the above embodiments.

[0038] In another aspect, embodiments may relate to a video preparation module comprising: a computer-readable storage medium having computer-readable program code implemented thereon, and a processor, preferably a microprocessor, coupled to the computer-readable storage medium, wherein, in response to executing the computer-readable program code, the processor is configured to perform executable operations including: encoding video data into encoded video frames representing base layer video data and encoded video frames representing one or more enhancement layer video data for enhancing the base layer video data; packaging the encoded video frames representing the base layer video data into one or more video base portions; and packaging the encoded video frames representing the one or more enhancement layer video data into one or more video enhancement portions, each video base portion associated with one or more corresponding video enhancement portions, wherein the one or more... The encoded video frames of a video base portion and the one or more video enhancement portions have a common timeline; switching information is generated, the switching information identifying one or more positions of one or more video frames in the corresponding video enhancement portion associated with one or more switching points on the common timeline, wherein the video frames of the corresponding enhancement layer are encoded such that: the video frames associated with the switching points in the video enhancement portions have one or more encoding dependencies on video frames at or before the switching points in the associated video base portion, and have no encoding dependencies on video frames before the switching points in the video enhancement portions; and the one or more video base portions of the base layer, the one or more video enhancement portions of the one or more enhancement layers, and the switching information are stored on a storage medium.

[0039] In another aspect, embodiments may relate to a computer-readable storage medium having stored data thereon, wherein the data may be defined for controlling video processing performed by a client device, including: one or more video base portions, each video base portion including encoded video frames representing base layer video data; one or more video enhancement portions, each video enhancement portion including encoded video frames representing one or more enhancement layer video data, each video base portion being associated with one or more corresponding video enhancement portions, wherein the encoded video frames of the one or more video base portions and the one or more video enhancement portions have a common timeline; and switching information for the client device, the switching information identifying one or more positions of one or more video frames in the corresponding video enhancement portion for the video base portion and associated with one or more switching points on the common timeline, wherein the video frames of the corresponding enhancement layers are encoded such that: the video frames associated with the switching points in the video enhancement portions may have one or more encoding dependencies on video frames at or before the switching points in the associated video base portion, and have no encoding dependencies on video frames in the video enhancement portions before the switching points.

[0040] The present invention may also relate to a computer program product including a software code portion configured to, when run in the memory of a computer, perform method steps according to any of the process steps described above.

[0041] The invention will be further illustrated with reference to the accompanying drawings, which will schematically illustrate embodiments according to the invention. It will be understood that the invention is not limited in any way to these specific embodiments. Attached Figure Description

[0042] Figure 1 The diagram illustrates a known quality switching scheme for adaptive video streaming.

[0043] Figure 2 The illustration shows a quality switching scheme for adaptive video streaming according to an embodiment; Figure 3 An example of the coding dependencies between video data in the base layer and the enhancement layer is depicted; Figure 4 A flowchart is depicted illustrating a method for processing video data according to an embodiment; Figure 5 A schematic diagram of a system for streaming media data based on a low-latency quality switching scheme according to an embodiment is depicted. Figure 6 The illustration shows a quality switching scheme for adaptive video streaming according to another embodiment; Figure 7The illustration shows a quality switching scheme for adaptive video streaming according to a further embodiment; Figure 8 The illustration shows a quality switching scheme for adaptive video streaming according to a further embodiment; Figure 9 A block diagram illustrating an exemplary data processing system that can be used with the embodiments described in this disclosure is depicted. Detailed Implementation

[0044] The embodiments in this disclosure relate to fast, low-latency quality switching for adaptive video streaming. An example of a media packaging standard that supports quality switching is the Common Media Application Format (CMAF). CMAF is built on top of the ISO Basic Media File Format (ISOBMFF) and provides a universal way to package content within CMAF segments, making it suitable for streaming protocols such as Dynamic Adaptive Streaming HTTP (DASH) and HTTP Live Streaming (HLS). Furthermore, it supports chunked encoding (i.e., breaking segments into smaller units called chunks, which can be delivered at encoding time) and chunked transport encoding, a data transfer mechanism in HTTP 1.1 known as the HTTP 1.1 Chunked Transport Encoding Mode, where the data stream is broken into non-overlapping chunks that are sent and received independently.

[0045] Benefiting from CMAF's chunked encoding and chunked transfer encoding capabilities, and by utilizing CMAF blocks instead of CMAF fragments, CMAF enables low-latency streaming in DASH. Specifically, the encoder and packer do not need to wait for the encoding of an entire CMAF fragment to be completed before transmitting it to an origin server, such as a content delivery network. Instead, CMAF blocks are encoded, packed, and transmitted to the origin server (using HTTP 1.1 chunked transfer encoding mode), and client devices can request CMAF fragments using HTTP requests based on a manifest file containing information for accessing video data (e.g., information about resource locators required to request CMAF blocks). Thus, client devices can receive them immediately upon completion of the encoding of the corresponding CMAF block. This reduces end-to-end latency, which is proportional to the length of the CMAF block (which is much shorter than the length of a CMAF fragment).

[0046] Quality switching in CMAF is performed through adaptive switching between so-called CMAF tracks, which include a set of CMAF switching segments. The switching set is formed by CMAF tracks of CMAF segments, where video frames or blocks within a CMAF segment are synchronized and time-aligned (hereinafter referred to as time-aligned CMAF segments and CMAF blocks).

[0047] Figure 1The illustration shows a CMAF switching set 102 for time-aligned CMAF segments. 1,2 The example includes the first CMAF fragment 204. 1-4 The sequence includes the first CMAF orbital 201 (basal layer) and the second CMAF segment 202. 1-4 The sequence includes a second CMAF track 200 (enhancement layer). Each track may include video data associated with different characteristics, such as quality. The video data for the base layer and enhancement layer may be generated using a layered video coding scheme, such as Scalable Video Coding (SVC), Scalable HEVC, or Multi-Layer VVC. Furthermore, CMAF segments may be divided into CMAF blocks. Other layered video coding schemes may include video coding schemes based on super-resolution techniques, such as those described, for example, in WO2019 / 197674 entitled "Block-level super-resolution based video coding," which is hereby incorporated by reference.

[0048] Each CMAF segment in one layer can be time-aligned with a corresponding CMAF segment in another layer. Each CMAF segment in one layer can contain multiple non-overlapping CMAF blocks. In some embodiments, these CMAF blocks in that layer can be time-aligned CMAF blocks of corresponding CMAF segments in another layer. The first CMAF block in each CMAF segment is configured as a Type 1 or 2 Stream Access Point (SAP) (e.g., an IDR in an AVC-generated CMAF segment), which allows for independent decoding of each CMAF segment. The Stream Access Point here can involve a Type 1 or 2 SAP frame, such as an I-frame or an IDR frame. Additionally, timestamps can be used to easily switch between tracks by sorting CMAF segments from different CMAF tracks during playback.

[0049] The CMAF format depicted in the figure allows a client device to consume video data at low quality and then continue with high-quality video data based on event triggering 106 (e.g., user action, network conditions, and / or client resources). As shown, by default, after an event trigger, the client device needs to wait until the end 108 of the CMAF segment 1041 for quality adaptation by switching to a higher resolution (i.e., retrieving encoded video data from both the base layer and the enhancement layer). Therefore, upon switching, the client will begin requesting CMAF blocks from both the base layer 1101 and the enhancement layer 1102, where the video data from the enhancement layer is used to enhance the quality of the base layer.

[0050] exist Figure 1In this proposed solution, smaller CMAF segment / fragment lengths should be used to achieve faster quality switching. Considering that each CMAF fragment should begin with a Type 1 or 2 SAP frame, reducing the length of the CMAF fragment introduces more Type 1 or 2 SAP frames into the bitstream, leading to an increase in bit rate. In extreme cases, where the application requires single-frame granularity for quality switching, current CMAF specifications would require single-frame encoding within CMAF fragments, or in other words, full intra-frame encoding, where each frame is placed within a CMAF fragment. However, it is clear that such a solution would request excessively high bit rates.

[0051] Figure 2 An example of a timing alignment segment for low-latency handover according to an embodiment of the present invention is illustrated. Specifically, the figure illustrates a CMAF track comprising a set of handover segments including timing alignment segments, which include media data encoded using a layered coding scheme (e.g., Scalable HEVC or Multi-Layer VVC). In such a coding scheme, the encoded media data may be formatted in a base layer 201 and one or more enhancement layers 200. Layer dependencies may be specified in a manifest file that identifies the layers, each layer including a sequence of CMAF segments, wherein dependencies may be specified using, for example, specific parameters (e.g., parameter flags in DASH MPD). @dependencyID (Use signals to notify.)

[0052] Media data associated with the base layer can be packaged in CMAF segment 204 1-4 In this context, each segment is further divided into blocks, for example, block 205 of segment 2042. 1-5 The media data of the enhancement layer can be packaged into CMAF fragments 202. 1-4 In some embodiments, the CMAF fragments of the enhancement layer may be further divided into CMAF blocks, such as block 205 of fragment 2022. 1-5 It is associated with fragment 2042 of the base layer.

[0053] In this scheme, when the client device receives a handover trigger 206, the client device can switch to higher quality based on metadata associated with the base layer 201. This metadata may contain handover information used to signal one or more handover points 214 for each CMAF segment. 1-4The switching points are time-aligned with the timelines of the CMAF segment 2042 of the base layer 201 and the associated CMAF segment 2022 of the enhancement layer 200. Each switching point may identify at least one video frame in the CMAF segment of the enhancement layer. The video frame identified by the switching point may be referred to as the switching frame. In an embodiment, the switching point identifying the switching frame may include an offset parameter relative to the start of the CMAF segment (e.g., IDR). This offset parameter may be implemented as a byte range in the CMAF segment of the enhancement layer. In another embodiment, the switching point may include a timestamp associated with the content timeline of the enhancement layer, which identifies the switching frame. In this case, the timestamp needs to be converted to a byte range in the CMAF segment of the enhancement layer in order to address the switching frame. In an embodiment, the switching points may be defined such that they are time-aligned with, for example, the timeline of the base layer 201 and the associated CMAF segment 2022 of the enhancement layer 200. Figure 2 The blocks in the CMAF segment of the base layer are shown to begin overlapping. Other configurations of the switching points are also possible. For example, in an embodiment, only one or a limited number of switching points may be defined for each CMAF segment.

[0054] Therefore, if the client device receives an action trigger during the playback of a video frame of a CMAF segment in the base layer, it will continue playing the video frames of that CMAF segment until the playback of the common timeline reaches a switching point (switching point 208 in this specific example) that coincides with the start of the third block 2053 of CMAF segment 2042, which has an associated CMAF segment 2022 in the enhancement layer. At that moment, the client device can use the metadata associated with the base layer to identify the switching frame in the associated CMAF segment of the enhancement layer. This switching frame defines the first video frame in the video frame sequence required to switch to rendering the media data at a higher quality.

[0055] In order to be able to Figure 2 The diagram illustrates a seamless switch to higher quality at the switching point signaled by the metadata associated with the base layer. The media data of the enhancement layer can be encoded so that the coding dependencies of the video frames associated with the enhancement layer during the switch can be handled by the decoder. This means that the switched video frames in the enhancement layer's CMAF segment have one or more coding dependencies on one or more earlier video frames in the associated CMAF segment of the base layer (and optionally, one or more CMAF segments preceding the associated CMAF segment).216 1-3The earlier video frame is located before or at the switching point on the common timeline. Furthermore, the switched video frames in the enhancement layer's CMAF segments do not have a coding dependency on later video frames in the enhancement layer's CMAF segments, where the later video frames are located after the switching point on the common timeline. Therefore, the enhancement layer's media data packaged in CMAF segments or CMAF blocks depends only on the base layer's media data packaged in the current or past time CMAF blocks. Thus, in this embodiment, for the enhancement layer, inter-layer coding dependencies are only allowed for the media data of the base layer's current and past video frames. These coding constraints allow for many different implementations of the layered coding schemes that can be used.

[0056] Figure 3 An example of the coding dependency between video data in the base layer and enhancement layer is depicted, suitable for a quality switching scheme according to embodiments of this disclosure. The figure illustrates block 302 including a CMAF segment. 1,2 The base layer 302 and the block 304 including CMAF segments corresponding to the base layer CMAF segments. 1,2 The enhancement layer 304. Therefore, as referenced Figure 2 In detail, these blocks can be packaged into time-aligned segments. Each block can contain multiple video frames, where metadata associated with the base layer indicates some video frames 302 in the enhancement layer. 1-4 These are associated with switching points (i.e., they are identified as switching frames). As shown in the figure, switching frames in the enhancement layer only have encoding dependencies on the current frame or past video frames in the base layer. Furthermore, switching frames in the enhancement layer do not have encoding dependencies on the current or past video frames in the enhancement layer.

[0057] For example, a switching frame 3063 in the enhancement layer (a P-frame in this example) depends on an associated P-frame 3083 in the base layer, which in turn depends on an earlier video frame 3082 (a P-frame in this example) in the base layer. This earlier video frame 3082, in turn, depends on an even earlier video frame 3081 (an I-frame in this example) in the base layer. Therefore, a switching frame, i.e., a video frame associated with a switching point in an enhancement layer segment, may have one or more coding dependencies on video frames in the associated segments of the base layer that are located at or before the switching point on the timeline. Furthermore, the switching frame 3063 in the enhancement layer does not have a coding dependency on earlier video frames in the enhancement layer, i.e., it does not have a coding dependency on video frames preceding the switching point in the video segment of the enhancement layer. Additionally, video frames in the enhancement layer after the switching point also do not have a coding dependency on video frames preceding the switching point in the video segment of the enhancement layer. This coding scheme can be used for low-latency quality handover schemes as described with reference to embodiments of this disclosure.

[0058] Therefore, as can be seen from the above, based on the switching information contained in the metadata associated with the base layer, the client device processing the encoded video data of the base layer block can identify the switching frames in the corresponding portion of the corresponding block or segment from the enhancement layer. One or more switching points in the enhancement layer segment can be defined such that, as referenced... Figure 2 and 3 When the video data of the interpreted enhancement layer block is combined with the video data of the current or past blocks of the base layer, the decoder can switch to generating higher quality video frames.

[0059] Switching information, including information about the switching point, can be signaled or associated with blocks in the base layer block using, for example, event messages (e.g., 'emsg' entries for blocks). In an embodiment, each block of a base layer segment may include an event message carrying metadata identifying the corresponding enhanced video data (e.g., the corresponding block of the enhancement layer). For this purpose, the event message may include a byte offset (relative to the start of the block) or a timestamp corresponding to the switching frame in the enhancement layer segment associated with the base layer segment. The event message may also include an identifier of the CMAF segment to which the byte offset belongs.

[0060] Therefore, in cases such as those triggered by user actions or prompted by adaptive bitrate streaming algorithms, the client can decide to switch to rendering content at a higher quality. The client device can then parse the metadata associated with the base layer chunk and retrieve portions of the corresponding chunk or enhancement layer fragment using an appropriate request protocol, such as an HTTP request, particularly an HTTP byte-range request.

[0061] The identifiers of enhancement layers and fragments (contained in the metadata associated with the base layer) can be matched with information published in the manifest file, allowing clients to construct resource locators (e.g., URLs) for the corresponding fragments. The metadata may further contain information used to identify the switching frames within a CMAF fragment, such as byte offsets. This information can be used by the client to perform an HTTP byte range request and retrieve portions of the enhancement layer block or fragment required for quality switching.

[0062] After the switch, the client device receives video data from the base layer and enhancement layer, which is then fed to the decoder for decoding. Typically, the base layer is decoded independently of the enhancement layer. However, the decoding of the enhancement layer may depend on the decoding of the base layer to improve coding efficiency. Therefore, there are inter-layer dependencies that the decoder handles when decoding the enhancement layer.

[0063] The decoder needs to be configured using profiles and hierarchies that support both the base and enhancement layers. Furthermore, media samples from both the base and enhancement layers can be formatted based on the same NAL access unit structure. The decoder's profile and hierarchies can be configured for high-quality rendering, and therefore, reconfiguration is not required to support high-quality decoding.

[0064] Figure 4 A flowchart illustrating a method for processing video data according to an embodiment is provided. Specifically, encoded video data from the base layer and enhancement layer are processed to allow for low-latency quality switching.

[0065] In the first step 402, the method may include receiving one or more video base portions, each video base portion including encoded video frames representing base layer video data, and each video base portion being associated with one or more corresponding video enhancement portions, the one or more corresponding video enhancement portions including encoded video frames representing one or more enhancement layer video data respectively used to enhance the base layer video data, the encoded video frames of the one or more video base portions and the one or more video enhancement portions having a common timeline. Each base portion may define encoded video data, such as encoded video frames, packaged according to a certain format (e.g., CMAF format).

[0066] In some embodiments, the base layer portion and the enhancement layer portion may form a switching set, including time-aligned and synchronized segments, such as CMAF segments. Base layer segments may be divided into blocks, such as CMAF blocks. In some embodiments, enhancement layer segments may also be divided into blocks. In other embodiments, the base layer portion may be a CMAF segment, and the enhancement layer portion may be a DASH segment.

[0067] Furthermore, switching information (step 404) may be received, wherein the switching information may identify one or more positions of one or more video frames in the corresponding video enhancement portion used for the video base portion, and one or more video frames in the corresponding video enhancement portion associated with one or more switching points on a common timeline. Video frames in the corresponding enhancement layer may be encoded such that the video frames associated with the switching points in the video enhancement portion may have one or more coding dependencies on video frames at or preceding the switching points in the associated video base portion. Furthermore, video frames in the corresponding enhancement layer may be encoded such that the video frames associated with the switching points in the video enhancement portion do not have coding dependencies on video frames preceding the switching points in the video enhancement portion.

[0068] Subsequently, decoding can be switched to decoding higher-quality video data (step 406). The switching may include: determining at least one corresponding enhancement portion for the current video base portion and the position of the video frame associated with the switching point within the corresponding enhancement portion, based on the switching information. Encoded video frames of the corresponding enhancement portion can be retrieved, wherein the encoded video frames may include encoded video frames on a common timeline that are located at or after the switching point. Subsequently, the encoded video frames of the video base portion and the encoded video frames of the corresponding enhancement portion can be decoded to determine higher-quality video frames compared to the quality of the video frames associated with the base layer.

[0069] Therefore, the embodiments provide low-latency switching in video data rendering from a first resolution to a second resolution without waiting for the segment (or equivalent data container, such as a DASH segment) to end, and without substantially increasing the bit rate (as is known from conventional quality switching schemes).

[0070] Figure 5 A schematic diagram of an exemplary streaming system is depicted for streaming media data based on a low-latency quality switching scheme as described with reference to embodiments of this disclosure. The system may include a content preparation system 502, a server system 504, and a media playback device 506 including a client device 542, which may communicate with the server system via a network 536 (including the Internet). In some embodiments, the content preparation system may be part of the server system. In other embodiments, the content preparation system may be connected to the server system via, for example, network 536 or another network, or may be directly communicatively coupled. The server system 504 may include a server processor 532 and one or more network interfaces 534 configured to send and receive data via network 536.

[0071] In some embodiments, the functionality of the server system and content preparation system can be implemented as a distributed system, which includes multiple network devices with communication connections, including but not limited to routers, bridges, proxy devices, switches, etc. In some embodiments, the server system may be part of a CDN.

[0072] Content preparation system 502 may include media source 508, such as media storage devices and / or one or more means for generating media data, such as video and audio sources, such as cameras, microphones, and computers for generating (e.g., synthesizing) computer-generated graphics and / or audio. The media data may include video data (e.g., in the form of a sequence of video frames), which may be associated with other data, such as audio data in the form of audio frames and / or metadata such as recording time and / or information about the temporal order of the video data (particularly encoded video data). The media server system may include encoder system 514, which includes one or more encoder instances for encoding the media data.

[0073] The encoder system can generate one or more encoded media data streams, simply referred to as media streams. In some embodiments, a single stream may be referred to as a basic stream representing a single digitally encoded component (e.g., video or audio) of media presentation. The content preparation system may further include a packetizer 516 for converting the basic stream comprising encoded media data into a packetized stream, such as a packetized basic stream (PES). A containerizer 517 can format (e.g., encapsulate) the PES stream for transmission, such that the encoded media data can be transmitted to a server system using an appropriate media streaming standard (such as MPEG-DASH, HLS, or CMAF) and stored as one or more media files on the storage medium of the server system 504. The containerizer can be configured to generate media files formatted according to a predetermined data format (e.g., CMAF segments or DASH segments).

[0074] Media data encoded in one or more primary streams, depending on a certain bitrate or quality, can form a media presentation, or simply a presentation. Encoder system 514 can be configured to encode media data of media headers in different ways using video coding standards, thereby producing different presentations of media headers with different bitrates and different characteristics (such as pixel resolution, frame rate, compliance with various coding standards, etc.). These different presentations can be used for adaptive bitrate streaming, as known from streaming protocols such as DASH. In particular, the encoder system can encode media data according to any suitable standardized coding scheme (such as H.264 / AVC, HEVC, or VVC). In embodiments, media data can be encoded using a layered coding scheme such as Scalable HEVC or Multi-Layer VVC. Layered coding is a type of data encoding for digital video or digital audio where the result of encoding source video data is not just a compressed data stream, but multiple streams called layers (typically a base layer and one or more enhancement layers). During decoding, all layers can be combined to recreate the original high-quality video stream. Even if some layers are missing, the stream can still be decoded (although layer-by-layer delimitation is generally required; for example, a base layer associated with a certain quality must be available). If one or more enhancement layers are missing, the resulting stream may have reduced visual quality, but may still be usable.

[0075] Layered coding is particularly suitable when the same media stream needs to be available at different qualities (e.g., for adaptive bitrate streaming). Without layered coding, the source video stream would need to be encoded multiple times to obtain compressed streams with different qualities and bitrates. In contrast, layered coding offers the advantage of encoding only once, as streams of different qualities can be obtained based on a base layer and one or more enhancement layers.

[0076] Encapsulator 516 can be configured to format packets of the basic system into Network Abstraction Layer (NAL) units. NAL units, defined as part of the H.264 / AVC and HEVC video coding standards, contain Video Coding Layer (VCL) NAL units including video data payloads, and non-VCL NAL units that may include metadata such as parameter sets (important header data applicable to a large number of VCL NAL units) and supplemental enhancement information (timing information and other supplementary data that enhance the usability of the decoded video signal). Non-VCL NAL units may contain Sequence Parameter Sets (SPS) (applied to a sequence of consecutive coded video images known as a coded video sequence) and Picture Parameter Sets (PPS) (applied to the decoding of one or more individual images within the coded video sequence). Non-VCL NAL units may further contain Supplemental Enhancement Information (SEI) messages that may contain information to assist the decoding process.

[0077] A set of NAL units defines so-called access units, which together form an encoded picture (video frame). Decoding an access unit typically results in a decoded picture (decoded video frame). An encoded video sequence consists of a series of access units sequentially within the NAL unit stream, using only one set of sequence parameters. Given the necessary set of parameters (which can be transmitted "in-band" or "out-of-band"), each encoded video sequence can be decoded independently of any other encoded video sequence. At the beginning of an encoded video sequence is an Instantaneous Decoding Refresh (IDR) access unit. An IDR access unit includes an intra-frame picture (I-frame), which is an encoded picture that can be decoded without decoding any previous pictures in the NAL unit stream. The presence of an IDR access unit indicates that subsequent pictures in the stream will not need to reference pictures preceding their contained intra-frame pictures for decoding. Wrappers can use encoded video sequences to produce short, non-overlapping short video files, such as DASH segments and CMAF fragments used by adaptive streaming protocols to provide adaptive streaming functionality.

[0078] The encapsulator 516 may be further configured to determine one or more manifest files 524 for media presentations associated with a media title. An example of a manifest file is a Media Presentation Descriptor (MPD) in the case of using MPEG-DASH for streaming media data. The manifest files identify media presentations 526-530 associated with a media title using an appropriate language such as Extensible Markup Language (XML). In some embodiments, media presentations may be divided into so-called adaptive sets. Adaptive sets may define media data associated with a common set of characteristics, including but not limited to, codecs, profiles and hierarchies, resolution, number of views, segmented file formats, etc. The manifest files may contain data identifying such adaptive sets and further information about the characteristics (such as bitrate) of a particular presentation associated with the adaptive set.

[0079] In some embodiments, the manifest file may include information about the availability of CMAF segments or blocks or DASH segments or sub-segments in a presentation. For example, in an embodiment, the manifest file may include clockwork information indicating the time at which a first segment of one of the presentations becomes available, and information indicating the duration of segments within a presentation. Thus, the client processor can determine when each segment is available based on the start time and the duration of the segment. Grouped and encapsulated media files, such as CMAF segments and / or DASH segments, associated with media titles prepared by the content preparation system, may be stored at the server system as one or more presentations 526-530 and an associated manifest file 524.

[0080] For example, as referenced Figure 2-4The described content preparation system can prepare and store CMAF tracks comprising a set of time-aligned segments (including media data that can be encoded using a layered coding scheme). Thus, when encoding media data, the content preparation system can generate tracks comprising CMAF segments associated with video data of a base layer, and one or more tracks comprising CMAF segments respectively associated with one or more enhancement layers.

[0081] The content preparation system can further generate metadata associated with the base layer, which may contain switching information to signal one or more switching points of at least one video frame (switching frame) in one or more CMAF segments of the enhancement layer. The switching point may be signaled as a timestamp of a common timeline associated with the enhancement layer, and / or as an offset relative to the start of a CMAF segment (e.g., IDR).

[0082] For reference Figure 2 and Figure 3 As explained, when a switching point is introduced into the media data of the enhancement layer, the content preparation system causes the encoder to encode the media data of the enhancement layer so that when switching to higher quality at the switching point, the encoding dependency of the video frames associated with the enhancement layer can be handled by the decoder.

[0083] Server processor 532 can be configured to receive network requests from client devices, such as client device 542. Client devices may include client processor 548, configured to request media data and / or metadata, such as manifest files, stored at the server system. Based on manifest file 546 stored at the client device, the client device can request media data from the server system (via the client processor) and store the media data in a client buffer 544. The buffered (encoded) media data can be provided as input to decoder system 550, which can be configured to execute one or more decoder instances to decode the encoded media data into decoded media data, which can be rendered by display device 552. The display device can be implemented as any type of display device, including displays such as head-mounted displays for rendering XR type media data (e.g., tiled 360-degree video data).

[0084] The client device can receive information about the decoding capabilities of the video decoder and the rendering capabilities of the media playback device. The functionality of the client device, or a portion thereof, can be implemented in hardware, or a combination of hardware, software, and / or firmware, wherein the necessary hardware can be provided to execute instructions for the software or firmware. Furthermore, in some embodiments, the client device, decoder system, and display device can be implemented in different devices; for example, the client device can be implemented in a media box, and the decoder and display can be implemented in a display system.

[0085] Server processor 532 and client processor 548 can be implemented to handle requests based on the Hypertext Transfer Protocol (HTTP) (e.g., HTTP 1.1, HTTP 2, HTTP 3, and future versions). For example, HTTP version 1.1 allows encoded media data to be transmitted based on a chunked transfer encoding mode. Thus, the server request processor can be configured to receive HTTP messages (such as HTTP GET or partial GET requests) and, in response to these requests, send media data back to the client device. These requests can, for example, use resource locators (such as URLs) to specify a portion of a video file or a specific part of one of 526-530, such as a clip or a block of a clip. In some examples, these requests may also specify a range of bytes used to identify a block within a clip. Other client-server communication protocols can be used instead of HTTP to handle request and response messages. For example, in an embodiment, a bidirectional communication channel between the client and server systems can be implemented based on the WebSocket protocol. In that case, a handshake request can be used to establish a WebSocket connection between the client and server. Request and response messages can be exchanged between the client and server over the WebSocket connection. Other protocols that can be used for communication between servers and clients include Long polling, WebRTC, SignalR, FTP, MQPTT, etc.

[0086] Additionally and / or alternatively, the request processors of the server system and client devices can be configured to deliver media data via broadcast or multicast protocols. In that case, the server device may use broadcast or multicast network transport protocols to deliver encapsulated media data (e.g., segments or fragments or portions thereof). For example, the server request processor may be configured to receive multicast group join requests from client devices. For example, the server system may advertise the Internet Protocol (IP) address associated with the multicast group to client device 542 associated with a specific media header. The client device may then submit a request to join the multicast group. This request may be propagated across network 536, such as routers defining the network, such that these routers are configured to deliver media data destined for the IP address associated with the multicast group to client devices that have already joined the multicast group.

[0087] Client devices can request encoded video data from the server system based on information in the manifest file and information about available network bandwidth. Higher bitrate presentations can produce higher quality video playback, while lower bitrate presentations can provide sufficient quality video playback when available network bandwidth decreases. Accordingly, when available network bandwidth is relatively high, client devices can retrieve media data associated with (relatively) high bitrate presentations, and when available network bandwidth is low, client devices can request media data associated with (relatively) low bitrate presentations. In this way, media data can be streamed to client devices over the network while adapting to changes in network bandwidth availability.

[0088] The network interface 540 of the client device can receive and buffer media, such as encapsulated and packetized media data, like selected CMAF segments or blocks for presentation. The client processor 548 can be configured to decapsulate the media file into a PES stream and depacket the PES stream into encoded media data, which can then be sent to a decoder system. The decoded media data can then be sent to a display device 552.

[0089] Figure 5 The devices and modules described herein (such as encoders, packetizers, packagers, server processors, client processors, etc.) can be implemented as any various suitable processing circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof, if applicable. Alternatively and / or additionally, these devices and modules may include integrated circuits, microprocessors, and / or wireless communication devices.

[0090] In further embodiments, more than one enhancement layer can be used to provide switching between multiple qualities. In the case of multiple enhancement layers, in order to enable quality switching between more than one layer presentation, the dependencies of each enhancement layer on other enhancement layers and on the base layer need to be constrained so that the decoder can handle the dependencies of video frames associated with the new enhancement layer during switching. Based on these constraints, different implementations of quality switching can be achieved. Some non-limiting examples of quality switching based on multiple enhancement layers are provided below.

[0091] Figure 6 An example of a timing alignment segment for low-latency handover according to a further embodiment is depicted. Specifically, the figure illustrates a CMAF track that defines a set of handover segments comprising a base layer 601 and N enhancement layers. This example shows two (N=2) enhancement layers, with a first enhancement layer 6001 associated with a first quality comprising a CMAF segment 604. 1-4 And the second enhancement layer 6001 associated with the second mass includes CMAF fragment 603 1-4 .

[0092] Combining the base layer video data with the first enhancement layer video data will produce a first-quality video frame after decoding. This first quality is higher than the base quality; that is, the quality of the video frame generated solely based on decoding the base layer video data. Similarly, combining the base layer video data with the first and second enhancement layers video data will produce a second-quality video frame after decoding. This second-quality video frame is higher than both the first and base quality video frames.

[0093] CMAF fragments can be referenced as follows Figure 2 The described approach is divided into blocks. Switching information for all enhancement layers can be provided to the base layer as metadata. This allows for access to information associated with the enhancement layers. N The blocks and their associations in the base layer and enhancement layer N All blocks in the enhancement layer immediately switch from rendering only associated with the base layer to rendering in the enhancement layer. N This allows for an immediate switch to any rendering, but introduces more signaling overhead in the base layer.

[0094] In this embodiment, the constraints of the first and second dependencies can be similar to those of the reference. Figure 2 The dependencies are explained, where at the switching point, the video data of the enhancement layer depends only on the video data of the current and past CMAF blocks of the base layer. Thus, there is a first coding dependency 610 between the base layer and the first enhancement layer, and a second coding dependency 612 between the base layer and the second enhancement layer.

[0095] Figure 7 An example of a timing alignment segment for low-latency handover according to another embodiment is depicted. The figure illustrates a CMAF track defining a set of handover segments for the timing alignment segments, which includes a base layer 701 (which includes CMAF segment 704). 1-4 And N enhancement layers. This example shows two (N=2) enhancement layers, with the first enhancement layer 7001 associated with the first mass including CMAF fragment 704. 1-4 And the second enhancement layer 7001 associated with the second mass includes CMAF fragment 703 1-4 This is similar to Figure 6 Examples are shown in the text.

[0096] In this embodiment, the switching information can be configured to provide a progressive switching effect, wherein the switching information for blocks in enhancement layer N is provided in the metadata of blocks associated with enhancement layer N-1. Therefore, the switching information for blocks in enhancement layer N can be signaled using metadata associated with blocks in another lower enhancement layer (e.g., enhancement layer N-1). Thus, based on the metadata in the base layer, the starting point of blocks in the first enhancement layer 7001 is determined and obtained. Then, based on the metadata in the blocks received from the first enhancement layer, the starting point of blocks in the second enhancement layer 7002 is determined and obtained.

[0097] In this way, the video data of the enhancement layer can be progressively obtained by extracting switching information from enhancement layer N-1 to obtain the corresponding block in enhancement layer N. In this embodiment, the switching information for switching to a higher rendering can be distributed across different layers, but this provides less flexibility in immediately switching to any higher rendering.

[0098] In this embodiment, the constraints of the first and second dependencies can be similar to those of the reference. Figure 2 The dependencies are explained, where at the switching point, the video data of the enhancement layer depends only on the video data of the current and past CMAF blocks of the base layer. Therefore, there is a first coding dependency 710 between the base layer and the first enhancement layer, and a second coding dependency 712 between the base layer and the second enhancement layer. This allows for direct switching to any enhancement layer in the switching set.

[0099] Figure 8 An example of a timing alignment segment for low-latency handover according to another embodiment is depicted. The figure illustrates a CMAF track defining a set of handover segments for the timing alignment segment, which includes a base layer 801 (which includes CMAF segment 804). 1-4 And N enhancement layers. This example shows two (N=2) enhancement layers, with the first enhancement layer 8001 associated with the first mass including CMAF fragment 804. 1-4And the second enhancement layer 8002 associated with the second mass includes CMAF fragment 803 1-4 This is similar to Figure 6 Examples are shown in the text.

[0100] In this embodiment, blocks in enhancement layer N depend on the current and past CMAF blocks of enhancement layer N-1, and so on. Therefore, in this case, it is only possible to switch the quality to a higher-level rendering layer in the middle of a segment. For example, if the client starts with the rendering of base layer 801, it can only switch the quality to the rendering of the first enhancement layer 8001 by fetching the corresponding block of enhancement layer 1. If the client starts with the base layer and the first enhancement layer, it can only switch the quality to the rendering of the next second enhancement layer by fetching the corresponding block of enhancement layer 2. In this case, the metadata associated with the CMAF blocks of enhancement layer N-1 can include switching information for the CMAF blocks of enhancement layer N. This information can be signaled, for example, in an event message (i.e., "emsg").

[0101] The argument is that it can be referenced Figure 6-8 The description combines different aspects of the switching set. Other implementations, different from those illustrated in the attached diagram, are also possible.

[0102] Therefore, the CMAF low-latency handover set described with reference to the embodiments of this disclosure includes a base layer and at least one enhancement layer, wherein the CMAF blocks of the enhancement layers depend only on the CMAF blocks of the base layer. Furthermore, preferably, the CMAF fragment header parameters are identical between the CMAF fragments of the base layer and the CMAF fragments of the enhancement layers. If this is not the case, metadata can be signaled so that the client device can still obtain the CMAF fragment header during handover. Additionally, the CMAF blocks of the base layer may contain handover information associated with the CMAF blocks of the enhancement layers. When multiple enhancement layers are used, metadata may also be associated with (a portion of) these enhancement layers.

[0103] As described above, the switching information can be signaled in an event message (i.e., "emsg"). In this embodiment, the switching information can be sent via an in-band DASH event message. cmafllqswitch In-band signal notification. An exemplary syntax for such DASH in-band event messages used for CMAF low-latency switching sets can be shown below: The following section describes some parameters in the event message in more detail. Since the base layer's CMAF block includes a portion of the CMAF fragment in the enhancement layer or the address information of the CMAF block, it is not necessary to insert the same in-band message in both renderings. Therefore, AdaptationSet@segmentAlignment It should be cleared as false.

[0104] Event message parameters scheme_id_uri You can define identifiers, i.e., URIs, such as " urn:mpeg:dash: event:cmafllqswitch:202x "" is used to signal to the client device that low-latency handover has been enabled, and the metadata associated with the base layer includes handover information, such as one or more event messages, for switching to higher quality based on media data from the base and enhancement layers.

[0105] Event message parameters presentation_time A time instance is defined on a common timeline, which is linked to a video frame in the enhancement layer. Typically, this time instance will point to the beginning of a block in the CMAF segment of the enhancement layer.

[0106] Event message parameters value This pertains to the byte range of the segment used to obtain the enhancement layer required for the desired quality switch. Different values ​​can define different ways of providing this byte range. Other addressing schemes are possible, and this value should be extended to cover these schemes.

[0107] Event message parameters message_data[] The ID of the corresponding enhancement fragment is defined, as well as the byte range of the expected portion of the enhancement fragment when a higher quality is expected to be switched at the next block.

[0108] Table 1 provides information for use in event-based messages. value Examples of schemes for using parameters to signal the required enhancement layer CMAF segment and parts within the CMAF segment for quality switching.

[0109] Table 1 – cmafllqswitch Interpretation of the value fields of the (CMAF low latency quality handover) event Therefore, as shown in the table, if the event message is notified by a signal... value=1 ,but message_data[ ] The field contains the ID of the corresponding enhancement fragment, and the absolute address of the expected portion of the enhancement fragment used for a byte-range request to switch quality at the next block. Similarly, when value=3, the message_data[] field contains the ID of the corresponding enhancement fragment, and the offset address (instead of the absolute address) of the expected portion of the enhancement fragment.

[0110] When a client device receives an event message, it only uses the event if it expects to switch to higher quality in the next block; otherwise, it skips the event. When a switch to higher quality is expected, the client device can process the event and use the byte range in the event to make an HTTP range request to obtain the corresponding portion of the enhancement layer fragment corresponding to the enhancement layer, and then perform the switch to higher quality in the next block.

[0111] While the embodiments in this disclosure are illustrated based on an implementation in CMAF, the embodiments are not limited thereto. For example, the enhancement layer can also be implemented as a sequence of DASH segments. A self-initialized DASH segment (i.e., a self-decoding DASH segment starting with an I-frame) can be further divided into sub-segments. Segments and sub-segments can be implemented in a way that simplifies switching. According to Clause 4.3 of ISO / IEC 23009-1, “Segments and sub-segments can be implemented in a way that simplifies switching. For example, in the simplest case, each segment or sub-segment begins with an SAP, and the boundaries of the segment or sub-segment span a presentation alignment of an adaptive set. In this case, switching presentations involves playing to the end of a (sub)segment of a presentation, and then playing from the beginning of the next (sub)segment of the new presentation.” DASH provides a similar way of switching between sub-segments as segmentation. When a sub-segment is time-aligned and begins with a Type 1 or 2 SAP frame, switching between different presentations can be seamlessly implemented. Therefore, similar to… Figure 1 In this scenario, low-latency handover would require dividing the video data into small sub-segments, each containing an I-frame. Reducing the length of the DASH (sub)segments would require more Type 1 or 2 SAP frames in the bitstream, resulting in excessively high bitrates.

[0112] To address this issue, in one embodiment, the enhancement layer can be formatted as DASH segments (which may not conform to the CMAF standard), and switching information in the base layer's metadata can be used to identify switching frames in the enhancement layer. Similarly, the enhancement layer can be formatted as DASH segments comprising sub-segments that do not begin with an I-frame. That is, even if a sub-segment does not conform to SAP 1 or 2, switching within a sub-segment can still be enabled by using layered encoding and switching information in the base layer sub-segment's metadata (using the sub-segment's offset or absolute address). Dependencies between layers can still be specified in the manifest (e.g., the flag @dependencyID in the DASH MPD).

[0113] Therefore, the embodiments in this disclosure generally allow different non-CMAF-based DASH implementations using sub-segments to still benefit from the solution (by allowing them to have sub-segments that start with non-SAP frames and achieve high-quality, low-latency handover without waiting for SAP boundaries), and to obtain similar advantages that can be provided by the CMAF embodiments described with reference to the accompanying drawings in this disclosure.

[0114] The apparatuses and systems (such as client devices, server systems, and content preparation systems) described with reference to embodiments in this disclosure are typically implemented as one or more data processing systems that are communicatively connected. Figure 9 This is a block diagram illustrating an exemplary data processing system that can be used as described in this disclosure. The data processing system 900 may include at least one processor 902 coupled to a memory element 904 via a system bus 906. Therefore, the data processing system can store program code within the memory element 904. Furthermore, the processor 902 can execute program code accessed from the memory element 904 via the system bus 906. In one aspect, the data processing system may be implemented as a computer suitable for storing and / or executing program code. However, it should be understood that the data processing system may be implemented in the form of any system including a processor and memory capable of performing the functions described herein.

[0115] Memory element 904 may include one or more physical memory devices, such as, for example, local memory 908 and one or more mass storage devices 910. Local memory may refer to random access memory or other non-persistent memory(s) typically used during the actual execution of the program code. Mass storage devices may be implemented as hard disk drives or other persistent data storage devices. Data processing system 900 may also include one or more cache memories (not shown) that provide temporary storage for at least some program code in order to reduce the number of times program code must be retrieved from mass storage device 910 during execution.

[0116] Input / output (I / O) devices, depicted as input device 912 and output device 914, may optionally be coupled to the data processing system. Examples of input devices may include, but are not limited to, keyboards, pointing devices (such as mice), etc. Examples of output devices may include, but are not limited to, monitors or displays, speakers, etc. The input and / or output devices may be coupled to the data processing system directly or via an intermediary I / O controller. Network adapter 916 may also be coupled to the data processing system to enable it to be coupled to other systems, computer systems, remote network devices, and / or remote storage devices via an intermediary private or public network. The network adapter may include a data receiver for receiving data transmitted to the system, device, and / or network, and a data transmitter for transmitting data to the system, device, and / or network. Modems, cable modems, and Ethernet cards are examples of different types of network adapters that can be used with the data processing system.

[0117] like Figure 9 As illustrated, memory element 904 can store application 918. It should be understood that the data processing system can further execute an operating system (not shown) that facilitates the execution of the application. The application, implemented as executable program code, can be executed by the data processing system (e.g., by processor 902). In response to executing the application, the data processing system can be configured to perform one or more operations, which will be described in further detail herein.

[0118] In one respect, for example, the data processing system may represent a client-side data processing system. In this case, application 918 may represent a client application that, when executed, configures the data processing system to perform the various functions described herein with reference to "client". Examples of clients may include, but are not limited to, personal computers, laptops, mobile phones, etc. In another respect, the data processing system may represent a server-side data processing system. In this case, application 918 may represent a server application that, when executed, configures the data processing system to perform the various functions described herein with reference to "server".

[0119] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or IC sets (e.g., chip sets). The various components, modules, or units described in this disclosure are intended to emphasize functional aspects of a device configured to perform the disclosed techniques, but do not necessarily require implementation by different hardware units. Rather, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with suitable software and / or firmware.

[0120] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that the terms “comprising” and / or “including,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0121] In the following claims, the corresponding structures, materials, actions, and equivalents of all components or steps plus functional elements are intended to include any structure, material, or action that performs the function in conjunction with other claimed elements as specifically claimed. The description of the invention has been presented for illustrative and descriptive purposes but is not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The embodiments have been chosen and described in order to best explain the principles and practical application of the invention and to enable others skilled in the art to understand the invention with various modifications suitable for the particular uses contemplated.

Claims

1. A method for processing video data, comprising: The client device receives one or more video base portions, each video base portion including encoded video frames representing base layer video data, and each video base portion is associated with one or more corresponding video enhancement portions, the one or more corresponding video enhancement portions including encoded video frames representing one or more enhancement layer video data respectively used to enhance the base layer video data, the one or more video base portions and the one or more video enhancement portions having a common timeline; Receive switching information, the switching information identifying a corresponding video enhancement portion of a video base portion, and one or more positions of one or more video frames in the corresponding video enhancement portion associated with one or more switching points on the common timeline, wherein the video frames of the corresponding enhancement layer are encoded such that: the video frames associated with the switching points in the video enhancement portion have one or more encoding dependencies on video frames at or before the switching point in the associated video base portion, and have no encoding dependencies on video frames before the switching point in the video enhancement portion; and Switching to decoding higher quality video data, the switching includes: - Based on the switching information, determine at least one corresponding enhancement portion and the position of the video frame associated with the switching point in the current video base portion; - Retrieve the encoded video frames of the corresponding enhancement portion, the encoded video frames including encoded video frames that have a position on the common timeline at or after the switching point; - Decode the encoded video frames of the video base portion and the encoded video frames of the corresponding enhancement portion to determine a higher quality video frame compared to the quality of the video frames associated with the base layer.

2. The method according to claim 1, wherein, The one or more video base portions and / or the one or more video enhancement portions are formatted as Common Media Application Format (CMAF) segments.

3. The method according to claim 1 or 2, wherein, Each CMAF fragment is subdivided into multiple non-overlapping CMAF blocks, preferably received using HTTP chunked transfer mode.

4. The method according to claim 3, wherein, The one or more positions of one or more video frames in the corresponding video enhancement portion associated with the one or more switching points define the start of one or more HTTP blocks in the corresponding enhancement portion.

5. The method according to any one of claims 1-4, wherein, The switching information includes an identifier and / or preferably a parameter value in the range of bytes, the identifier being used to identify the at least one corresponding enhancement portion, and the parameter value indicating the position of the video frame in the corresponding enhancement portion associated with the switching point.

6. The method according to any one of claims 1-5, wherein, The encoded video frames of the corresponding enhancement portions are retrieved by the client device based on a manifest file, preferably a Media Presentation Description (MPD), which includes one or more video portion identifiers for identifying the one or more video base portions and / or one or more corresponding video enhancement portions associated with each of the one or more video base portions.

7. The method according to claim 6, wherein, Retrieving the encoded video frame corresponding to the enhancement portion includes: A resource locator is determined based on the one or more video portion identifiers, used to locate the corresponding enhancement portion; as well as Based on the resource locator and the information associated with the position of the video frame in the corresponding enhancement section, a request message, preferably an HTTP byte range request message, is determined to request the encoded video frame of the corresponding enhancement section.

8. The method according to any one of claims 1-7, wherein, At least a portion of the switching information is provided to the client device using one or more in-band messages, such as one or more DASH in-band event messages, in at least a portion of the one or more video base portions and / or in one or more blocks of the one or more video base portions, and optionally in at least a portion of the one or more corresponding enhancement portions and / or in one or more blocks of the one or more corresponding enhancement portions.

9. The method according to any one of claims 1-8, wherein, If the client device receives a request to switch to higher quality, and / or if the client device determines to switch to higher quality, the client device requests the video enhancement portion associated with the video base portion.

10. The method according to any one of claims 1-9, wherein, Each of the one or more video base components and / or one or more video enhancement components includes at least one group of pictures (GOP), wherein video frames of the GOP have no coding dependency on video frames outside the GOP; and / or wherein... The video data is encoded using a layered coding scheme, such as Scalable HEVC or Multi-Layer VVC.

11. An apparatus for processing video data, comprising: A computer-readable storage medium having at least a portion of a program implemented thereon; And, a computer-readable storage medium having computer-readable program code implemented thereon, and a processor, preferably a microprocessor, coupled to said computer-readable storage medium, wherein, in response to executing said computer-readable program code, said processor is configured to perform executable operations including: The client device receives one or more video base portions, each video base portion including encoded video frames representing base layer video data, and each video base portion is associated with one or more corresponding video enhancement portions, the one or more corresponding video enhancement portions including encoded video frames representing one or more enhancement layer video data respectively used to enhance the base layer video data, the one or more video base portions and the one or more video enhancement portions having a common timeline; Receive switching information, the switching information identifying a corresponding video enhancement portion for a video base portion, and one or more positions of one or more video frames in the corresponding video enhancement portion associated with one or more switching points on the common timeline, wherein the video frames of the corresponding enhancement layer are encoded such that: the video frames associated with the switching points in the video enhancement portion have one or more encoding dependencies on video frames at or before the switching point in the associated video base portion, and have no encoding dependencies on video frames before the switching point in the video enhancement portion; and Switching to decoding higher quality video data, the switching includes: - Based on the switching information, determine at least one corresponding enhancement portion and the position of the video frame associated with the switching point in the corresponding enhancement portion for the current video base portion; - Retrieve the encoded video frames of the corresponding enhancement portion, the encoded video frames including encoded video frames that have a position on the common timeline at or after the switching point; - Decode the encoded video frames of the video base portion and the encoded video frames of the corresponding enhancement portion to determine a higher quality video frame compared to the quality of the video frames associated with the base layer.

12. The device according to claim 11, wherein, The processor is further configured to perform any one of the steps of the method according to claims 2-10.

13. A video preparation module, comprising: A computer-readable storage medium having computer-readable program code implemented thereon, and a processor, preferably a microprocessor, coupled to said computer-readable storage medium, wherein, in response to executing said computer-readable program code, the processor is configured to perform executable operations including: The video data is encoded into encoded video frames representing base layer video data and encoded video frames representing one or more enhancement layer video data used to enhance the base layer video data. Encoded video frames representing the base layer video data are packaged into one or more video base portions, and encoded video frames representing one or more enhancement layer video data are packaged into one or more video enhancement portions. Each video base portion is associated with one or more corresponding video enhancement portions, wherein the encoded video frames of the one or more video base portions and the one or more video enhancement portions have a common timeline. Generate switching information, which identifies a corresponding video enhancement portion for a video base portion, and one or more positions of one or more video frames in the corresponding video enhancement portion associated with one or more switching points on the common timeline, wherein the video frames of the corresponding enhancement layer are encoded such that: the video frames associated with the switching points in the video enhancement portion have one or more encoding dependencies on video frames at or before the switching point in the associated video base portion, and have no encoding dependencies on video frames before the switching point in the video enhancement portion; and The one or more video base components of the base layer, the one or more video enhancement components of the one or more enhancement layers, and the switching information are stored on the storage medium.

14. A computer-readable storage medium thereon including data storing data, the data defining a data format for streaming video data to a client device, comprising: One or more video base portions, each video base portion including encoded video frames representing base layer video data; One or more video enhancement portions, each video enhancement portion including encoded video frames representing one or more enhancement layer video data, each video base portion being associated with one or more corresponding video enhancement portions, wherein the encoded video frames of the one or more video base portions and the one or more video enhancement portions have a common timeline; as well as Switching information for the client device, the switching information identifying a corresponding video enhancement portion for a video base portion, and one or more positions of one or more video frames in the corresponding video enhancement portion associated with one or more switching points on the common timeline, wherein the video frames of the corresponding enhancement layer are encoded such that: the video frames associated with the switching points in the video enhancement portion have one or more encoding dependencies on video frames at or before the switching points in the associated video base portion, and have no encoding dependencies on video frames before the switching points in the video enhancement portion.

15. A computer program product comprising a software code portion configured to, when run in a computer's memory, perform the method steps according to any one of claims 1-10.

Citation Information

Patent Citations

  • Block-level super-resolution based video coding

    WO2019197674A1