Method, apparatus and computer program product for video streaming

By distinguishing between the viewport and other areas in 360-degree video streaming and prioritizing the transmission of tile streams from the viewport area, the problem of insufficient bandwidth utilization in existing technologies is solved, thereby improving user experience and image quality.

CN115023955BActive Publication Date: 2026-05-01NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2021-01-21
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing 360-degree video streaming technologies struggle to efficiently handle the differences between the viewport area and non-viewport areas, resulting in insufficient bandwidth utilization and a poor user experience.

Method used

By identifying the viewport region and other regions of the 360-degree video, the set of tile streams covering these regions is inferred, and a high-priority transmission is requested for the set of tile streams in the viewport region, while the set of tile streams in the edge and background regions is transmitted with low priority. The transmission strategy is optimized using metadata.

Benefits of technology

It improves bandwidth utilization and enhances the user experience, especially in terms of image quality and response speed in the viewport area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115023955B_ABST
    Figure CN115023955B_ABST
Patent Text Reader

Abstract

Embodiments relate to a method comprising: determining a foreground region covering a viewport of a 360-degree video and one or more other regions of the 360-degree video that do not contain the foreground region as a whole; inferring a first set of tile streams among available tile streams of the 360-degree video to cover the foreground region; inferring a second set of tile streams among the available tile streams of the 360-degree video to cover the one or more other regions; and requesting transmission of a first set of portions of the first set of tile streams and a second set of portions of the second set of tile streams, wherein portions in the first set of portions have a shorter duration than portions in the second set of portions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technical solution generally involves the streaming of media content (especially video). Background Technology

[0002] This section is intended to provide background or context for the invention as set forth in the claims. The description herein may include concepts that can be obtained but are not necessarily concepts that were previously conceived or obtained. Therefore, unless otherwise indicated herein, the content described in this section is not prior art to the specification and claims of this application, and is not acknowledged as prior art by virtue of its inclusion in this section.

[0003] Devices capable of capturing images and videos have evolved from those capturing limited angular fields of view to those capturing 360-degree content. These devices are capable of capturing all visual and audio content around them; that is, they can capture the entire angular field of view, which can be referred to as a 360-degree field of view. More precisely, devices are capable of capturing a spherical field of view (i.e., 360 degrees in all spatial directions). In addition to new types of image / video capture devices, new types of output technologies have also been invented and developed, such as head-mounted displays. These devices allow an individual to see all the visual content around him / her, thus giving the feeling of being “immersed” in a scene captured by a 360-degree camera. With a spherical field of view, this new capture and display paradigm is often referred to as virtual reality (VR) and is considered a common way for people to experience media content in the future. Summary of the Invention

[0004] The scope of protection sought by the various embodiments of the present invention is set forth in the independent claims. Embodiments and features (if any) described in this specification that do not fall within the scope of the independent claims should be interpreted as examples that aid in understanding the various embodiments of the invention.

[0005] Various aspects include a method, an apparatus, and a computer-readable medium comprising a computer program product stored therein, characterized by the content set forth in the independent claim. Various embodiments are disclosed in the dependent claims.

[0006] According to the first aspect, a method is provided, comprising:

[0007] - Identify the foreground region of the viewport covering the 360-degree video and one or more other regions of the 360-degree video that do not include the entire foreground region;

[0008] - Infer the first set of available tile streams in the 360-degree video, where the first set of tile streams covers the foreground region;

[0009] - Infer a second set of tile streams from the available tile streams of the 360-degree video, wherein the second set of tile streams covers one or more other regions; and

[0010] - Request the transmission of a first portion of a first tile stream set and a second portion of a second tile stream set, wherein portions of the first portion have a shorter duration than portions of the second portion.

[0011] According to an embodiment, one or more other regions include an edge region and a background region, and the method further includes:

[0012] - Identify the edge regions adjacent to the foreground region and the background regions that cover the 360-degree video but are not included in the foreground or edge regions;

[0013] - Infer the first subset of the second tile flow set, where the first tile flow set covers the foreground region;

[0014] - Infer the second subset of the tile flow to include tile flows that are in the second tile flow set but not in the first subset; and

[0015] - Request the transfer of a first subset of a portion of a first subset of a tile stream and a second subset of a portion of a second subset of a tile stream, wherein the portion of the first subset of the portion has a shorter duration than the portion of the second subset of the portion.

[0016] According to an embodiment, the method further includes requesting the transfer of an extractor track or the like in a portion having a duration longer than the first set of portions.

[0017] According to an embodiment, the method further includes: obtaining metadata on the area covered by the tile stream from media presentation description and / or stream initialization data and / or stream index data; and inferring a first tile stream set and a second tile stream set based on the metadata.

[0018] According to a second aspect, an apparatus is provided, comprising at least one processor and a memory including computer program code, the memory and the computer program code being configured to use the at least one processor to cause the apparatus to perform at least the following operations:

[0019] - Identify the foreground region of the viewport covering the 360-degree video and one or more other regions of the 360-degree video that do not include the entire foreground region;

[0020] - Infer the first set of available tile streams in the 360-degree video, where the first set of tile streams covers the foreground region;

[0021] - Infer a second set of tile streams from the available tile streams of the 360-degree video, wherein the second set of tile streams covers one or more other regions; and

[0022] - Request the transmission of a first portion of a first tile stream set and a second portion of a second tile stream set, wherein portions of the first portion have a shorter duration than portions of the second portion.

[0023] According to an embodiment, one or more other regions include an edge region and a background region, and the device further includes computer program code for causing the device to perform the following operations:

[0024] - Identify the edge regions adjacent to the foreground region and the background regions that cover the 360-degree video but are not included in the foreground or edge regions;

[0025] - Infer the first subset of the second tile flow set, where the first tile flow set covers the foreground region;

[0026] - Infer the second subset of the tile flow to include tile flows that are in the second set of tile flows but not in the first subset; and

[0027] - Request the transfer of a first subset of a portion of a first subset of a tile stream and a second subset of a portion of a second subset of a tile stream, wherein the portion of the first subset of the portion has a shorter duration than the portion of the second subset of the portion.

[0028] According to an embodiment, the device also includes computer program code configured to request the transfer of an extractor track or the like in a portion having a duration longer than that of the first set of portions.

[0029] According to an embodiment, the device further includes computer program code configured to cause the device to perform the following operations:

[0030] - Obtain metadata on the area covered by the tile stream from media presentation description and / or stream initialization data and / or stream index data; and

[0031] - Infer the first tile flow set and the second tile flow set based on metadata.

[0032] According to a third aspect, a computer program product is provided, including computer program code configured to cause a device or system to perform the following operations when executed on at least one processor:

[0033] - Identify the foreground region of the viewport covering the 360-degree video and one or more other regions of the 360-degree video that do not include the entire foreground region;

[0034] - Infer the first set of available tile streams in the 360-degree video, where the first set of tile streams covers the foreground region;

[0035] - Infer the second set of tile streams from the available tile streams in the 360-degree video, wherein the second set of tile streams covers one or more other regions; and

[0036] - Request the transmission of a first portion of a first tile stream set and a second portion of a second tile stream set, wherein portions of the first portion have a shorter duration than portions of the second portion. Attached Figure Description

[0037] In the following description, various embodiments will be described in more detail with reference to the accompanying drawings, in which:

[0038] Figure 1 An example of the OMAF system architecture is shown;

[0039] Figure 2 This shows an example of packing 360-degree content into the same frame;

[0040] Figure 3 An example of a 360-degree sphere is shown from the top;

[0041] Figure 4 An example of a video tile track placed at a bitrate is shown;

[0042] Figure 5 A simplified illustration shows the switching between low-quality and high-quality streams in a DASH on-demand profile;

[0043] Figure 6 A simplified illustration shows the switching between low-quality and high-quality streams in the DASH live profile;

[0044] Figure 7 This is a flowchart illustrating a method according to an embodiment; and

[0045] Figure 8 An example of a device according to an embodiment is shown. Detailed Implementation

[0046] Since the invention of photography and cinematography, the most common types of image and video content have been captured by cameras with relatively narrow fields of view and displayed as rectangular scenes on flat-panel displays. In this application, such content is referred to as "flat content," "flat image," or "flat video." Cameras are primarily directional, thus capturing only a limited angular field of view (the field of view they are pointing towards). Such flat video is output by display devices capable of displaying two-dimensional content.

[0047] Recently, new image and video capture devices have become available. These devices are capable of capturing all visual and audio content around them; that is, they can capture the entire angular field of view, which can be referred to as a 360-degree field of view. More precisely, the devices are capable of capturing a spherical field of view (i.e., 360 degrees in all spatial directions). In addition, new types of output such as head-mounted displays and other devices allow individuals to see 360-degree visual content.

[0048] 360-degree video, or virtual reality (VR) video, generally refers to visual content with a large field of view (FOV) where only a portion of the video is displayed at a single point in time in a typical display setup. For example, VR video can be viewed on a head-mounted display (HMD) capable of displaying, for example, a 100-degree field of view. The spatial subset of VR video content to be displayed can be selected based on the orientation of the HMD. In another example, consider a typical flat-panel viewing environment where, for example, a field of view of up to 40 degrees can be displayed. When displaying wide FOV content (e.g., fisheye) on such a display, a spatial subset, rather than the entire image, can be displayed.

[0049] Available media file format standards include the International Organization for Standardization (ISO) Basic Media File Format (ISO / IEC 14496-12, which can be abbreviated as ISOBMFF), the Moving Picture Experts Group (MPEG)-4 file format (ISO / IEC 14496-14, also known as MP4 format), the file format for video constructed using NAL (Network Abstraction Layer) units (ISO / IEC 14496-15), and the High Efficiency Video Coding Standard (HEVC or H.265 / HEVC).

[0050] Some concepts, structures, and specifications of ISOBMFF are described below as examples of container file formats, based on embodiments that can be implemented thereon. The invention is not limited to ISOBMFF, but rather a description is given of a possible foundation upon which the invention can be implemented, in part or in whole.

[0051] High Efficiency Image File Format (HEIF) is a standard developed by the Moving Picture Experts Group (MPEG) for storing images and image sequences. Furthermore, this standard promotes the encapsulation of data encoded according to the High Efficiency Video Coding (HEVC) standard. HEIF includes features built upon the ISO Basic Media File Format (ISOBMFF).

[0052] The ISOBMFF structure and features are largely used in the design of HEIF. The basic design of HEIF involves storing still images as items and image sequences as tracks.

[0053] In the following text, the term "omnidirectional" can refer to media content that can have a spatial range greater than the field of view of the device rendering the content. Omnidirectional content can, for example, cover essentially 360 degrees in the horizontal dimension and essentially 180 degrees in the vertical dimension, but omnidirectional can also refer to covering less than 360 degrees of view in the horizontal direction and / or covering 180 degrees of view in the vertical direction.

[0054] Panoramic images covering a 360-degree horizontal field of view and a 180-degree vertical field of view can be represented by a sphere that has been mapped onto a two-dimensional image plane using an equal rectangular projection (ERP). In this case, the horizontal coordinates can be considered equivalent to longitude, and the vertical coordinates can be considered equivalent to latitude, without any transformation or scaling applied. In some cases, panoramic content with a 360-degree horizontal field of view but less than 180 degrees of vertical field of view can be considered a special case of ERP, where the polar coordinate region of the sphere has not yet been mapped onto the two-dimensional image plane. In some cases, panoramic content can have less than 360 degrees of horizontal field of view and up to 180 degrees of vertical field of view, while in others it exhibits characteristics of an equal rectangular projection format.

[0055] Compared to consuming 2D content, immersive multimedia, such as omnidirectional content consumption, is more complex for end users. This is due to the greater degrees of freedom available to the end user. This freedom also leads to greater uncertainty. MPEG Omnidirectional Media Format (OMAF) v1 standardizes the omnidirectional streaming of a single 3DoF (3 degrees of freedom) content (where the viewer is located at the center of a unit sphere and has three degrees of freedom (yaw-pitch-roll)). At the time of writing, OMAF version 2 is nearing completion. Furthermore, OMAF v2 includes units for optimizing viewport-dependent streaming (VDS) operation and bandwidth management.

[0056] A viewport can be defined as an area of ​​omnidirectional image or video suitable for display and viewing by a user. The current viewport (which may sometimes be simply referred to as the viewport) can be defined as the portion of spherical video currently displayed and therefore viewable by the user. At any given point in time, a portion of the 360-degree video rendered by an application on a head-mounted display (HMD) is called the viewport. Similarly, when viewing a spatial portion of 360-degree content on a regular display, the currently displayed spatial portion is the viewport. The viewport is a window represented in the 360-degree world of omnidirectional video displayed via a rendering display. A viewport can be characterized by a horizontal field of view (VHFoV) and a vertical field of view (VVFoV).

[0057] In the context of viewport-dependent 360-degree video streaming, the term "tile" can refer to an isolated region, defined as an isolated region juxtaposed in a reference frame that depends only on other regions in the current or reference frames. Sometimes, a term tile or a sequence of term tiles can refer to a sequence of juxtaposed isolated regions. It is worth noting that video codecs can use term tiles as spatial image segmentation units, which may not have the same constraints or properties as isolated regions.

[0058] In the context of viewport-dependent 360-degree video streaming, an example of how to obtain tiles using the High Efficiency Video Coding (HEVC) standard is called a Motion Constrained Tile Set (MCTS). MCTS constrains the inter-frame prediction process in the coding such that no sample values ​​outside the MCTS are used, and no sample values ​​at fractional sample locations derived from one or more sample values ​​outside the Motion Constrained Tile Set are used for inter-frame prediction of any samples within the Motion Constrained Tile Set. Additionally, the coding of the MCTS is constrained in such a way that motion vector candidates are not derived from blocks outside the MCTS. This can be enforced by disabling temporal motion vector prediction in HEVC or by disallowing the encoder from using temporal motion vector prediction (TMVP) candidates or any motion vector prediction candidates after the TMVP candidates in the motion vector candidate list for prediction units located directly to the left of the right tile boundary of the MCTS (except for the last one at the bottom right of the MCTS). In general, an MCTS can be defined as a tile set independent of any sample values ​​and coded data (such as motion vectors) outside the MCTS. An MCTS sequence can be defined as a sequence of corresponding MCTS in one or more coded video sequences or the like. In some cases, it may be required that the Motion Constraint Tile Set (MCTS) form a rectangular region. It should be understood that, depending on the context, MCTS can refer to a set of tiles within an image or a corresponding set of tiles in a sequence of images. The corresponding tile set may, but generally does not, need to be juxtaposed within the sequence of images. A motion-constrained tile set can be considered an independently coded tile set because it can be decoded without other tile sets.

[0059] The following describes another example of how tiles can be obtained using the Universal Video Coding (VVC) standard in the context of viewport-dependent 360-degree video streaming. The draft VVC standard supports subpictures (also called sub-pictures). A subpicture can be defined as a rectangular region of one or more slices within a picture, where one or more slices are complete. Therefore, a subpicture consists of one or more slices that collectively cover a rectangular region of the picture. It is possible to require that the slices of a subpicture be rectangular slices. The segmentation of the picture to subpicture (also called subpicture layout or subpicture arrangement) can be indicated in and / or decoded from the SPS. One or more of the following attributes can be indicated collectively or individually (e.g., by the encoder) or decoded (e.g., by the decoder) or inferred (e.g., by the encoder and / or decoder) for each subpicture: i) whether the subpicture is treated as a picture during decoding; in some cases, this attribute does not include in-loop filtering operations, which can be indicated / decoded / inferred individually; ii) whether in-loop filtering operations are performed across subpicture boundaries. Treating sub-pictures as pictures during decoding can be included in inter-frame prediction so that sample locations outside the sub-picture would otherwise be saturated onto the sub-picture boundaries. When sub-picture boundaries are treated as picture boundaries, and sometimes when loop filtering across sub-picture boundaries is disabled, sub-pictures can be considered as tiles within the context of viewport-dependent 360° video streaming.

[0060] When streaming VR video, a subset of the 360-degree video content covering the viewport (i.e., the current view orientation) can be sent at optimal quality / resolution, while the remainder of the 360-degree video can be sent at lower quality / resolution. This is a characteristic of VDS systems, as opposed to viewport-independent streaming systems, where omnidirectional video is streamed at the same quality in all directions.

[0061] Figure 1 The diagram illustrates the OMAF system architecture. For example, the system can reside within a video camera or a network server. Figure 1 As shown, omnidirectional media (A) is acquired. If the OMAF system is part of the video source, omnidirectional media (A) is acquired from the camera device. If the OMAF system is located on a network server, omnidirectional media (A) is acquired from the video source via the network.

[0062] Omnidirectional media includes image data (B) i ) and audio data (B a The source media's images / videos are processed separately. In image stitching, rotation, projection, and region-by-region packing, the source media's images / videos are provided as input (B). iThe spheres are then stitched together to generate a sphere image on a unit sphere for each global coordinate axis. The unit sphere is then rotated relative to the global coordinate axis. The amount of rotation used to transform the local coordinate axis to the global coordinate axis can be specified by the rotation angle indicated in the RotationBox. The local coordinate axis of the unit sphere is the axis of the coordinate system that has been rotated. The absence of the RotationBox indicates that the local coordinate axis is the same as the global coordinate axis. The sphere image on the rotated unit sphere is then converted into a 2D projection image, for example, using an isometric projection. When spatial packing of the stereoscopic content is applied, the two sphere images for the two views are converted into two composition images, after which frame packing is applied to pack the two composition images into a single projection image. Rectangular region packing can then be applied to obtain the packed image from the projection image. The packed image (D) is then provided for video and image encoding to obtain the encoded image (E). i ) and / or encoded video stream (E v The audio from the source media is provided as input to the audio encoding (B). a Audio encoding provides encoded audio (E) a ) as output. Encoded data (E) i E v E a It is then encapsulated for playback (F) and delivery (i.e., streaming) (F) s (The document is missing from the original text.)

[0063] In OMAF Player 200, such as in HMD, the file unpacker processes files (F', F'). s And extract the encoded bitstream (E' i E' v E' a And then parse the metadata. Audio, video, and / or images are then decoded into decoded data (D', B'). a Based on the viewport and orientation sensed by the head / eye-tracking device, the decoded image (D') is projected onto the display. Similarly, the decoded audio (B')... a Rendered via speakers / headphones.

[0064] The basic building blocks in the ISO Basic Media File Format are called boxes. Each box has a header and a payload. The box header indicates the type of box and its size in bytes. The box type can be identified by an unsigned 32-bit integer interpreted as a four-character code (4CC). Boxes can surround other boxes, and the ISO file format specifies which box types are allowed within a given type of box. Additionally, the presence of some boxes may be mandatory in every file, while the presence of others may be optional. Furthermore, for some box types, having more than one box in a file may be permitted. Therefore, the ISO Basic Media File Format can be considered as specifying a hierarchy of boxes.

[0065] According to the ISO Basic Media File Format, the file includes media data and metadata encapsulated in a box.

[0066] In files conforming to the ISO Basic Media File Format, media data can be provided in one or more instances of MediaDataBox('mdat') and MovieBox('moov') and can be used to encapsulate metadata for timing media. In some cases, both the 'mdat' and 'moov' boxes can be required for an operable file. A 'moov' box can contain one or more tracks, and each track can reside in a corresponding TrackBox('trak'). Each track is associated with a processor identified by a four-character code specifying the track type. Video, audio, and image sequence tracks can be collectively referred to as media tracks, and they contain the basic media stream. Other track types include cue tracks and timing metadata tracks.

[0067] Movie fragments can be used, for example, when recording content to an ISO file to avoid data loss in case the recording application crashes, runs out of storage space, or other accidents. Without movie fragments, data loss can occur because the file format may require all metadata (e.g., movie boxes) to be written to a contiguous area of ​​the file. Additionally, when recording a file, there may not be enough storage space (e.g., RAM) to buffer movie boxes for the size of available storage, and recalculating the contents of movie boxes when the movie is closed may be too slow. Furthermore, movie fragments enable simultaneous recording and playback of files using a regular ISO file parser. Additionally, for progressive downloads, a shorter initial buffer duration is required, such as when simultaneously receiving and playing back a file with movie fragments, and the initial movie box is smaller compared to a structured file with the same media content but without movie fragments.

[0068] The movie fragment feature allows metadata that would otherwise reside in a movie box to be split into multiple fragments. Each fragment can correspond to a specific time period of a track. In other words, the movie fragment feature allows file metadata to be interleaved with media data. Therefore, the size of the movie box can be limited, and the use cases mentioned in paragraph

[0031] can be implemented.

[0069] In some examples, the media sample of a movie fragment may reside in a 'mdat' box. However, for the metadata of a movie fragment, a 'moof' box can be provided. The 'moof' box can include information about a specific playback duration that was previously in a 'moov' box. The 'moov' box itself may still represent a valid movie, but additionally, it can include a 'mvex' box indicating that the movie fragment will follow in the same file. Movie fragments can be extended to temporally associated presentations with 'moov' boxes.

[0070] Within a movie fragment, there can be sets of track fragments, including anywhere from zero to multiple tracks per track. Track fragments can in turn include anywhere from zero to multiple track extensions, each track extension being a continuous extension of samples of that track (and thus similar to chunks). Within these structures, many fields are optional and can be defaulted. Metadata that can be included in a 'moof' box can be limited to a subset of metadata that can be included in a 'moov' box and can be encoded differently in some cases. Details about the boxes that can be included in a 'moof' box can be found in the ISOBMFF specification. A self-contained movie fragment can be defined as consisting of consecutive 'moof' boxes and 'mdat' boxes in file order, where the mdat box contains samples of the movie fragment (for which the 'moof' box provides metadata) and does not contain samples of any other movie fragments (i.e., any other 'moof' boxes).

[0071] Tracks consist of samples, such as audio or video frames. For video tracks, media samples may correspond to encoded images or access units.

[0072] A media track refers to a sample formatted according to a media compression format (and its encapsulation of the ISO Basic Media File Format); it may also be called a media sample. A prompt track refers to a prompt sample containing instructions for constructing packets to be transmitted via the indicated communication protocol. A timing metadata track may refer to a sample describing the referenced media and / or prompt sample.

[0073] The 'trak' box includes a SampleDescriptionBox in its box hierarchy, which provides details about the encoding type used, as well as any initialization information required for that encoding. The SampleDescriptionBox contains the number of entries and as many sample entries as indicated by the entry count. The format of the sample entries is track-type specific but derived from a generic class (e.g., VisualSampleEntry, AudioSampleEntry). The type of the sample entry form used to derive the track-type specific sample entry format is determined by the track's media handler.

[0074] In the ISOBMFF container format of HEVC and AVC (Advanced Video Coding) bitstreams (ISO / IEC 14496-15), an extractor track is specified. Samples in the extractor track contain instructions to reconstruct a valid HEVC or AVC bitstream by including rewritten parameter sets and slice header information and referencing byte ranges of encoded video data from other tracks. Therefore, the player only needs to follow the instructions of the extractor track to obtain a decodeable bitstream from the track containing the MCTS.

[0075] A Uniform Resource Identifier (URI) can be defined as a string used to identify the name of a resource. Such identification enables interaction with the resource's representation over a network using a specific protocol. A URI is defined by specifying the specific syntax used for the URI and the scheme of the associated protocol. Uniform Resource Locators (URLs) and Uniform Resource Names (URNs) are forms of URIs. A URL can be defined as a URI that identifies a network resource and specifies the way to act on or obtain the resource's representation, specifying both its primary access mechanism and network location. A URN can be defined as a URI that identifies a resource by its name within a specific namespace. A URN can be used to identify a resource without implying its location or how to access it.

[0076] Hypertext Transfer Protocol (HTTP) has been widely used for delivering real-time multimedia content over the Internet, such as in video streaming applications. Several commercial technologies for adaptive streaming over HTTP exist, such as... Smooth streaming transmission Adaptive HTTP live streaming and Dynamic streaming has been initiated and standardization projects have been implemented. Adaptive HTTP Streaming (AHS) was first standardized in Release 9 of the 3GPP Packet Switched Streaming Service (PSS) (3GPP TS 26.234 Release 9: "Transparent end-to-end packet-switched streaming service (PSS); protocols and codecs"). MPEG uses 3GPP AHS Release 9 as the starting point for the MPEG DASH standard (ISO / IEC 23009-1: "Dynamic adaptive streamingover HTTP (DASH) - Part 1: Media presentation description and segment formats," International Standard, 2nd Edition, 2014). MPEG DASH and 3GP-DASH are technically similar and can therefore be collectively referred to as DASH.

[0077] Some concepts, structures, and specifications of DASH are described below as examples of container file formats, upon which embodiments can be implemented. The invention is not limited to DASH, but rather a description is given of a possible foundation upon which the invention can be implemented, in part or in whole.

[0078] In DASH, multimedia content can be stored on an HTTP server and delivered using HTTP. Content can be stored on the server in two parts: a Media Presentation Description (MPD), which describes the available content, its various alternatives, their URLs, and other characteristics; and segments, which are actual multimedia bitstreams contained in chunks within single or multiple files. The MPD provides the client with the necessary information to establish dynamic adaptive streaming over HTTP. The MPD contains information describing the media presentation, such as the HTTP Uniform Resource Locator (URL) for each segment for which a GET segment request is made. To play content, a DASH client can obtain the MPD, for example, via HTTP, email, thumb drives, broadcast, or other delivery methods. By parsing the MPD, the DASH client can learn about program timing, media content availability, media type, resolution, minimum and maximum bandwidth, the presence of various encoding alternatives for multimedia components, accessibility features and required Digital Rights Management (DRM), the location of media components on the network, and other content characteristics. Using this information, the DASH client can select appropriate encoding alternatives and begin streaming the content by retrieving segments, for example, using an HTTP GET request. After appropriate buffering to allow for changes in network throughput, the client can continue retrieving subsequent segments and also monitor network bandwidth fluctuations. Clients can determine how to adapt to available bandwidth to maintain a sufficient buffer by extracting different alternative segments (with lower or higher bit rates).

[0079] In the context of DASH, the following definitions can be used: A media content component or a media component can be defined as a contiguous component of media content with a media component type that can be individually encoded into a media stream. Media content can be defined as a media content segment or a contiguous sequence of media content segments. A media content component type can be defined as a single type of media content, such as audio, video, or text. A media stream can be defined as an encoded version of a media content component.

[0080] In DASH, the hierarchical data model is used to construct media presentations as follows: A media presentation consists of a sequence of one or more time segments, each time segment containing one or more groups, each group containing one or more adaptation sets, each adaptation set containing one or more representations, and each representation consisting of one or more segments. A group can be defined as a collection of adaptation sets that are not expected to be presented simultaneously. An adaptation set can be defined as a set of interchangeable encoded versions of one or more media content components. A representation is one of the alternative choices of media content or a subset thereof that can be differentiated by encoding choices (e.g., by bitrate, resolution, language, codec, etc.). A segment contains a specific duration of media data, as well as metadata of the media content included in decoding and presentation. Segments are identified by a URI and can be requested by an HTTP GET request. A segment can be defined as a unit of data associated with an HTTP-URL and, optionally, a byte range specified by the MPD.

[0081] DASH MPDs conform to Extensible Markup Language (XML) and are therefore specified using elements and attributes as specified in XML. MPDs can be specified using the following conventions: elements in an XML document can be identified by their capitalized first letter and can be displayed in bold as Element. To indicate that element Element1 is protected within another element Element2, it can be written as Element2.Element1. Camel case can be used if the element name consists of two or more compound words, such as ImportantElement. An element can exist only once, or the minimum or maximum number of occurrences can be specified by [the specified number of occurrences]. <minoccurs> ... <maxoccurs>Definition. Attributes in an XML document can be identified by their lowercase initial letter and can be preceded by the '@' symbol, for example, @attribute. To refer to a specific attribute @attribute contained within an element, it can be written as Element@attribute. If the attribute name consists of two or more words, camelCase can be used after the first word, such as @veryImportantAttribute. Attributes can also be assigned status in XML as mandatory (M), optional (O), optional with default value (OD), and conditionally mandatory (CM).

[0082] In DASH, all descriptor elements can be constructed in the same way: they contain the `@schemeIdUri` attribute, which provides the URI of the identification scheme, along with optional attributes `@value` and `@id`. The semantics of the element are scheme-specific. The URI of the identification scheme can be a URN or a URL. Some descriptors are specified in MPEG-DASH (ISO / IEC 23009-1), while descriptors are additionally or alternatively specified in other specifications. When specified in a specification other than MPEG-DASH, MPD does not provide any specific information about how to use the descriptor element. This depends on the application or specification that uses the DASH format to instantiate the descriptor element using the appropriate scheme information. The application or specification using one of these elements defines the scheme identifier in the form of a URI and the value space for that element when the scheme identifier is used. The scheme identifier appears in the `@schemeIdUri` attribute. In the case of a simple set of enumerated values, a text string can be defined for each value and this string can be included in the `@value` attribute. If structured data is required, any extended elements or attributes can be defined in a separate namespace. The `@id` value can be used to refer to a unique descriptor or a group of descriptors. In the latter case, descriptors with the same value for the `@id` attribute can be required to be synonymous; that is, processing one of the descriptors with the same value for `@id` is sufficient. Two elements of type `DescriptorType` are equivalent if the element name, the value of `@schemeIdUri`, and the value of the `@value` attribute are equivalent. If `@schemeIdUri` is a URN, equivalence can refer to semantic equivalence as defined in section 5 of RFC 2141. If `@schemeIdUri` is a URL, equivalence can refer to character-to-character equality as defined in section 6.2.1 of RFC 3986. If the `@value` attribute is not present, equivalence can be determined solely by equivalence for `@schemeIdUri`. Attributes and elements in extended namespaces may not be used to determine equivalence. The `@id` attribute can be ignored for equivalence determination.

[0083] MPEG-DASH specifies the descriptors EssentialProperty and SupplementalProperty. For the EssentialProperty element, the media presentation author states that successful processing of the descriptor is essential for the proper use of information in the parent element containing the descriptor, unless the element shares the same @id with another EssentialProperty element. If EssentialProperty elements share the same @id, it is sufficient to process one of the EssentialProperty elements with the same value for @id. At least one EssentialProperty element with each distinct @id value is expected to be processed. If a scheme or value for an EssentialProperty descriptor is not recognized, the DASH client expects to ignore the parent element containing that descriptor. Multiple EssentialProperty elements with the same value for @id and different values ​​for @id can exist in MPD.

[0084] For the SupplementalProperty element, the media presentation authors indicate that the descriptor contains supplementary information that can be used by the DASH client for optimized processing. If no scheme and value are identified for the SupplementalProperty descriptor, the DASH client expects to ignore the descriptor. Multiple SupplementalProperty elements can exist in MPD.

[0085] MPEG-DASH specifies viewpoint elements formatted as attribute descriptors. The @schemeIdUri attribute of a viewpoint element is used to identify the viewpoint scheme employed. Adaptation sets containing non-equivalent viewpoint elements contain different media content components. Viewpoint elements can also be applied to media content types that are not video. Adaptation sets with equivalent viewpoint element values ​​are intended to be presented together. This processing should apply equally to both identified and unidentified @schemeIdUri values.

[0086] DASH services can be provided as either on-demand or live services. In the former, the MPD (Multi-Party Description) is static, and all segments of the media presentation are available when the content provider publishes the MPD. However, in the latter, the MPD can be static or dynamic, depending on the segment URL construction method used by the MPD, and segments are continuously created as content is generated and published to the DASH client by the content provider. The segment URL construction method can be a template-based method or a segment list generation method. In the former, the DASH client can construct segment URLs without updating the MPD before requesting segments. In the latter, the DASH client must periodically download updated MPDs to obtain segment URLs. Therefore, for live services, the template-based segment URL construction method is preferable to the segment list generation method.

[0087] An initialization fragment can be defined as a fragment containing metadata necessary for rendering a media stream encapsulated within a media fragment. In the ISOBMFF-based fragment format, an initialization fragment may include a movie box ('moov'), which may not include metadata for any sample; that is, any metadata for a sample is provided in the 'moov' box.

[0088] A media segment contains media data for a specific duration used for playback at normal speed; this duration is called the media segment duration or segment duration. Content creators or service providers can choose the segment duration based on the desired characteristics of their service. For example, a relatively short segment duration can be used in live services to achieve short end-to-end latency. This is because the segment duration can be a lower bound on the end-to-end latency perceived by the DASH client, since a segment is a discrete unit for generating media data for DASH. Content generation can be done in a way that makes the entire segment of media data available to the server. Furthermore, many client implementations use segments as units for GET requests. Therefore, in a live service setup, a segment can be requested by the DASH client only when the entire duration of the media segment is available and when it is encoded and encapsulated into a segment. For on-demand services, different strategies for selecting segment durations can be used.

[0089] Fragments can be further subdivided into sub-fragments, for example, to enable downloading fragments in multiple parts. Sub-fragments can be required to contain complete access units. Sub-fragments can be indexed by fragment index boxes, which contain information mapping the rendering time range and byte range for each sub-fragment. Fragment index boxes can also describe sub-fragments and stream access points within a fragment by signaling their duration and byte offset. DASH clients can use the information obtained from the fragment index boxes to make HTTP GET requests for specific sub-fragments using byte-range HTTP requests. If a relatively long fragment duration is used, sub-fragments can be used to keep the size of the HTTP response reasonably and flexibly adapted to the bitrate. The fragment indexing information can be placed in a single box at the beginning of the fragment or scattered among many index boxes within the fragment. Different scattering methods are possible, such as hierarchical, daisy-chain, and mixed. This technique avoids adding large boxes at the beginning of the fragment and thus prevents potential initial download delays.

[0090] Sub-representations are embedded within regular representations and described by the `SubRepresentation` element. The `SubRepresentation` element is contained within the representation element. The `SubRepresentation` element describes the properties of one or more media content components embedded in the representation. It may, for example, describe the exact properties of an embedded audio component (such as codec, sample rate, etc.), an embedded subtitle (such as codec), or it may describe some embedded lower-quality video layer (such as some lower frame rate, or, for example, other cases). Sub-representations and representations share common properties and elements. If the `@level` attribute exists in the sub-representation element, the following applies:

[0091] Sub-representations provide the ability to access lower-quality versions of the representations they contain. In this case, sub-representations, for example, allow the extraction of audio tracks from multiplexed representations or enable efficient fast-forward or rewind operations when provided at a lower frame rate.

[0092] The initialization fragment and / or media fragment and / or index fragment should provide sufficient information to allow the data to be easily accessed via an HTTP partial GET request. Details regarding the provision of such information are defined by the media format in use.

[0093] When ISOBMFF fragments are used and sub-representations are described in MPD, the following applies:

[0094] a) Initialize the fragment containing the level assignment box.

[0095] b) A sub-fragment index box ('ssix') exists for each sub-fragment.

[0096] c) The @level attribute specifies the level to which the described sub-representation is associated in the sub-fragment index. The information in the representation, sub-representation, and level assignment ('leva') boxes contains information about the allocation of media data to levels.

[0097] d) Media data should be ordered so that each level provides enhancement compared to lower levels.

[0098] If the @level attribute is not present, the SubRepresentation element is only used to provide a more detailed description of the media stream in the embedded representation.

[0099] Typically, DASH can be used in two operating modes: live and on-demand. For both modes, the DASH standard provides profiles specifying constraints regarding the MPD and fragment formats. In the live profile, the MPD contains sufficient information for requesting media fragments, and the client can adapt to the streaming bitrate by selecting the representation from which it receives media fragments. In the on-demand profile, in addition to the information in the MPD, the client typically obtains a fragment index from each media fragment of each representation. In most cases, only a single media fragment exists for the entire representation. The fragment index includes SegmentIndexBoxes ('sidx'), which describes the sub-fragments included in the media fragment. The client selects a representation, extracts the sub-fragments from that representation, and requests them using byte-range requests.

[0100] A spherical region can be defined as a region on a sphere. The type can be indicated or inferred for the spherical region. A first spherical region type can be specified by four large circles, where each large circle can be defined as the intersection of the sphere and a plane passing through the center point of the sphere. A second spherical region type is defined by two azimuth circles and two altitude circles, where each azimuth circle can be defined as a circle on the sphere connecting all points with the same azimuth value, and each altitude circle can be defined as a circle on the sphere connecting all points with the same altitude value. The first and / or second spherical region types can additionally override the defined type of a region on a sphere rotated after applying specific amounts of yaw, pitch, and roll rotation.

[0101] Region-by-region quality ranking metadata can exist in or accompany the video or image bitstream. The quality ranking value of a quality ranking region can be relative to other quality ranking regions within the same bitstream or track, or to quality ranking regions on other tracks. Region-by-region quality ranking metadata can be indicated, for example, using a SphereRegionQualityRankingBox or a 2DRegionQualityRankingBox, which is specified as part of the MPEG omnidirectional media format. The SphereRegionQualityRankingBox provides the quality ranking value for a sphere region (i.e., a region defined on a sphere), while the 2DRegionQualityRankingBox provides the quality ranking value for a rectangular region on the decoded picture (and the remaining region that may cover all regions not covered by any of the rectangular regions). The quality ranking value indicates the relative quality order of the quality ranking regions. When quality ranking region A has a non-zero quality ranking value that is less than that of quality ranking region B, quality ranking region A has a higher quality than quality ranking region B. When the quality ranking value is non-zero, the picture quality within the entire indicated quality ranking region can be defined as approximately constant. Generally, the boundaries of the quality-ordered sphere or 2D region may or may not match the boundaries of the packed region or projected region specified in the per-region packed metadata. DASH MPD or other streaming manifests may include per-region quality-ordering signaling. For example, OMAF specifies per-spherical region quality ordering (SRQR) and per-2D region quality ordering (2DQR) descriptors for carrying quality-ordering metadata for spherical regions and 2D regions on the image for decoding, respectively.

[0102] Content coverage can be defined as one or more spherical regions represented by tracks or image items. Content coverage metadata can exist in or accompany the video or image bitstream, such as in the CoverageInformationBox specified in OMAF. Content coverage can be indicated to apply to either single-image content, stereoscopic content (as indicated), or both views of stereoscopic content. When indicated for two views, the content coverage of the left view may or may not match the content coverage of the right view. DASH MPD or other streaming manifests can include content coverage signaling. For example, OMAF specifies a Content Coverage (CC) descriptor carrying content coverage metadata.

[0103] Recent trends in VR video streaming to reduce bitrates include sending a subset of 360-degree video content covering the primary viewport (i.e., the current view orientation) at optimal quality / resolution, while sending the remainder of the 360-degree video at lower quality / resolution. Generally, two methods exist for viewport-adaptive streaming: viewport-specific encoding and streaming; and VR viewport video.

[0104] Viewport-specific coding and streaming is also known as viewport-dependent coding and streaming, or asymmetric projection or packaged VR video. In such methods, 360-degree image content is packaged into the same frame, with emphasis (e.g., larger spatial areas) placed on the main viewport. The packaged VR frames are encoded into a single bitstream. For example, in a cubemap projection format (also called cube mapping), spherical video is projected onto the six faces (also called sides) of a cube. The cubemap can be generated, for example, by first rendering the spherical scene six times from the viewpoint, where the view is defined by a 90-degree view frustum representing each cube face. Cube sides can be packaged into the same frame, or each cube side can be (i.e., in encoding) processed individually. There are many possible orders for positioning cube sides onto frames and / or for cube sides to be rotated or mirrored. In cubemap formats, the front face of the cubemap can be sampled at a higher resolution compared to other cube faces, and the cube face can be mapped to the same packaged VR frames, such as... Figure 2 As shown.

[0105] VR viewport video is also known as tile-based encoding and streaming. In this approach, 360-degree video is encoded and available in a way that allows selective streaming from different encodings to the viewport. The projected image is encoded as a series of tiles. Several versions of the content are encoded at different bitrates and / or resolutions. The encoded tile sequence can be streamed along with metadata describing the location of the tiles on the omnidirectional video. The client selects which tiles to receive so that the viewport is covered at a higher quality and / or resolution than other tiles. For example, each cube face can be encoded and encapsulated in its own track (and representation). More than one encoded bitstream can be provided for each cube face, for example, each with a different spatial resolution. The player is able to select tracks (or representations) to decode and play based on the current viewing orientation. High-resolution tracks (or representations) can be selected for use to render the cube faces of the current viewing orientation, while the remaining cube faces can be obtained from their lower-resolution tracks (or representations). In another example, isorectangular panoramic content is encoded using tiles. More than one encoded bitstream can be provided, for example, with different spatial resolutions and / or image qualities. Each tile sequence is available in its own track (and representation). The player can select a track (or representation) to decode and play based on the current viewing orientation. A high-resolution or high-quality track (or representation) can be selected to cover tiles in the current primary viewport, while the remaining area of ​​360-degree content can be obtained from a low-resolution or low-quality track (or representation).

[0106] While the VR viewport video discussed above is more flexible because it implements mixed tiles and introduces higher-quality tiles when needed, and may use already cached tiles as lower-quality tiles, instead of switching the entire 360-degree video when the viewport changes, it also introduces overhead because the tiles are being transmitted as separate streams through the transport layer.

[0107] Furthermore, to optimize motion for high-quality latency, video tile segments need to be relatively short, preferably less than 1 second. This means short transport layer packets, and therefore a high packet rate. Together with a large number of parallel download streams, this results in a high network packet rate; leading to high MPEG-DASH rates for HTTP requests.

[0108] This embodiment focuses on the transmission of tiles in a solution, such as in VR viewport video (i.e., tile-based encoding and streaming).

[0109] Viewport relevance can be achieved by having at least two quality regions: foreground (content within the current viewport) and background (content outside the current viewport in a 360-degree video). Instead of two types, quality regions can include a third type: edges surrounding the viewport. It should be understood that the embodiments are not limited to two or three types of quality regions but are generally applicable to any number of types of quality regions. Figure 3 The top view shows a 360-degree sphere. Region 310 represents the current viewport (where region 320 represents edges due to mismatch between the viewport and tile boundaries). Region 330 represents edges that can have lower quality than the viewport. Region 340 represents the background region, i.e., content outside the current viewport, which has even lower video quality and therefore occupies the least bandwidth.

[0110] From a transport system perspective, tile-based delivery of 360-degree video content can be challenging due to the combination of a large number of tiles and short tile fragments that together generate high-rate HTTP requests. For example, if a 360-degree video is divided into 16 tiles, and the tile fragment duration is 500 milliseconds or less, the number of tile fragments to download is 16 / 0.5 seconds = 32 HTTP requests / second just for the video. Furthermore, the rate is even higher when the user is moving their head, as this triggers high-quality tile requests for the new viewport. This can cause throughput issues in the client's HTTP / TCP / WiFi stack and / or network. Moreover, in DASH live profile streaming, the Content Delivery Network (CDN) needs to handle even millions of ultra-short segments (both in terms of time and byte size). Because the problem involves the number of requests, it cannot be solved using an Adaptive Bitrate (ABR) scheme. Additionally, other tracks, such as extractor tracks and audio tracks, may exist, further increasing HTTP traffic.

[0111] HTTP traffic is not a problem for traditional non-tile video streaming because only a few streams are downloaded in parallel. Additionally, non-360-degree live video broadcasts can use 2–3 second long segments. Such segment durations represent a trade-off between the increased latency of live broadcasting and the load on the CDN (traffic from the client to the CDN and from the CDN back to the content origin, as well as the storage of small segments).

[0112] On the other hand, for on-demand non-360-degree streaming, the client selects the size of the downloaded sub-segments, without necessarily adhering to possible sub-segment boundaries at the file format level. However, clients that support scenarios where streaming representations may need to be switched (such as adaptive bitrate) need to utilize sub-segments to achieve switching between representations.

[0113] HTTP / 2 can reuse HTTP requests and responses, thus reducing the number of TCP connections required. However, in practice, HTTP / 2 seems to be widely supported only with the help of HTTPS, which may not be optimal for all media streaming use cases because HTTPS encryption is specifically for each stream sent by each client, and therefore the stream cannot be cached in intermediate network elements to serve multiple clients.

[0114] This embodiment focuses on using different (sub)segment sizes (in terms of duration) for different streams, instead of using a fixed size for all streams.

[0115] According to an embodiment, a media player device:

[0116] - Identify the foreground region of the viewport covering the 360-degree video and one or more other regions of the 360-degree video that do not include the entire foreground region;

[0117] - Infer the first set of available tile streams in the 360-degree video, where the first set of tile streams covers the foreground region;

[0118] - Infer a second set of tile streams from the available tile streams of the 360-degree video, wherein the second set of tile streams covers one or more other regions; and

[0119] - Request the transmission of a first portion of a first tile stream set and a second portion of a second tile stream set, wherein portions of the first portion have a shorter duration than portions of the second portion.

[0120] In the embodiments, a portion is a sub-fragment as defined in DASH.

[0121] In the embodiments, a portion is a segment as defined in DASH.

[0122] In this embodiment, a stream is represented as specified in DASH.

[0123] Short (sub)segment sizes (in terms of duration) allow for rapid viewport switching as the user turns their head. This can also be achieved using long transport-level segments with several random access points within the segment. However, this can lead to wasted bandwidth if only a portion of the downloaded segment is used for the active viewport.

[0124] The viewport consumes most of the bandwidth, and therefore short (sub)segment sizes should be used during content creation to avoid waste in the download and / or playback of (sub)s that contain the viewport before the viewport changes. Furthermore, the player device should only extract such (sub)s that are truly needed.

[0125] In 360-degree video, the background area (content outside the current viewport) can occupy significantly less bandwidth than the foreground area (viewport). Therefore, there is less bandwidth wastage when some segments are replaced by higher-quality segments due to viewport changes.

[0126] If an edge region exists, the edge region can have a (sub)fragment size between the background (sub)fragment size and the foreground (viewport) (sub)fragment size.

[0127] In this embodiment, one or more other regions include an edge region and a background region, and the player device:

[0128] - Identify the edge regions adjacent to the foreground region and the background regions that cover the 360-degree video but are not included in the foreground or edge regions;

[0129] - Infer the first subset of the second tile flow set, where the first tile flow set covers the foreground region;

[0130] - Infer that the second subset of the tile flow includes tile flows that are in the second set of tile flows but not in the first subset; and

[0131] - Request the transfer of a first subset of a portion of a first subset of a tile stream and a second subset of a portion of a second subset of a tile stream, wherein the portion of the first subset of the portion has a shorter duration than the portion of the second subset of the portion.

[0132] Extractor tracks (if present) can also use very little bandwidth, so using long clip sizes is not a problem. The same applies to audio, as well as other tracks that do not affect the user experience when viewport switching occurs.

[0133] In one embodiment, the player device requests the transmission of an extractor track or the like in a portion having a duration longer than the first set of portions.

[0134] The increase in latency when receiving longer background (sub) segments should not be significant, because background (sub) segments have a lower bit rate and are downloaded faster than foreground (sub) segments.

[0135] A more detailed description in the context of OMAF is given below.

[0136] OMAF includes two main types of viewport-related operations based on tiles:

[0137] - Region-by-Region Mixed Quality (RWMQ): Several versions of the content are encoded using the same tile format, each with a different bitrate and / or image quality. All versions have the same resolution, and the tile sizes are generally identical, regardless of their position within the image. The player selects high-quality tiles to cover the viewport, and low-quality tiles cover the remainder of the sphere.

[0138] - Mixed-Resolution-Per-Region (RWMR): Tiles are encoded at multiple resolutions. The player device selects a combination of high-resolution tiles covering the viewport and low-resolution tiles for the remaining areas.

[0139] RWMQ can include, for example, a 6x4 tiled grid, thus resulting in 24 tiled tracks being downloaded in parallel along with an extractor, audio, and possibly other tracks. Utilizing short clip sizes (e.g., 300 milliseconds), this can result in close to 100 HTTP requests per second or even more.

[0140] Several options exist for implementing the RWMR scheme, but the number of tile tracks is comparable to that of the proposed RWMQ case.

[0141] If each tile is decoded using a dedicated decoder in the player device, then tile fragments do not need to be aligned. However, this may not be the actual setting, and therefore OMAF specifies the mechanism for merging fragment bitstreams and decoding the entire image using a single decoder. However, this does require tile fragments to be aligned, i.e., the frame type of each frame must be aligned.

[0142] DASH On-Demand Configuration Files:

[0143] In the ISOBMFF-based DASH on-demand profile, an entire video tile track at a bitrate can be placed in a single file as a single segment and accessed via a single URL. Figure 4 An example of file 400 is shown, comprising a movie head box ('moov'), a fragment index box ('sidx'), and a sequence of movie fragments ('moof' + 'mdat'). 'sidx' defines the offset of each 'moof' box from the origin of the file.

[0144] First, the player device downloads a fragment index table ('sidx') indicating the location of sub-fragments in the file. After that, the player device uses HTTP byte-range requests to extract sub-fragments from the file. The player device can decide how much data it wants to download for each request. The requested URL remains the same, only the byte range differs. Therefore, the server doesn't even need to know about possible fragment extraction optimizations, as long as it creates short fragments suitable for fast viewport switching.

[0145] There is no requirement to follow any standard for sub-fragment boundaries during download, but the improvement in this embodiment involves selecting the range of bytes to download based on the role of the video tile: background contrast, foreground contrast, and edges.

[0146] According to an embodiment, for a background tile, the player device can extract several sub-fragments per HTTP request; however, for a foreground tile, it extracts a smaller number of (sub)fragments per HTTP request, such as only a single sub-fragment per HTTP request.

[0147] Figure 5 The illustration shows the sub-segments switching between low-quality and high-quality streams. Initially, the player device has downloaded four low-quality sub-segments as a combination. A viewport switch occurs during sub-segment 3, so the player device quickly fetches high-quality sub-segment 4 and continues fetching high-quality sub-segments 5 and 6. The already downloaded low-quality sub-segment 4 is no longer needed. When another viewport switch occurs at sub-segment 5, the player device continues playing the already buffered high-quality sub-segments 5 and 6, and then switches to low quality at the next sub-segment boundary.

[0148] DASH Live - Configuration File:

[0149] In a DASH live-profile configuration file, each segment (including one or more ISOBMFF movie fragments) can have a unique URL. Therefore, a player device cannot extract many segments in a single HTTP request. Instead, for a DASH live-profile configuration file, some modifications are needed to the process of how the DASH stream is generated. This can be achieved through any of the following embodiments:

[0150] Migrating non-360-degree live video deployments for 360-degree viewport-related streaming enables

[0151] - Use short DASH fragments to keep viewport switching latency low but with high HTTP request traffic and CDN (Content Delivery Network) load.

[0152] - Use long DASH segments to keep CDN load low but with significant viewport switching latency.

[0153] According to an embodiment, (without any changes to existing standards) the optimized live deployment for viewport-related stream transmission can instead be configured to generate long DASH segments consisting of sub-segments that match viewport switching delay targets. For example, a 3-second long segment with 5–10 sub-segments.

[0154] According to an embodiment, the player device is configured as follows:

[0155] - Obtain the available portions of the tile stream from the media presentation description and / or stream initialization data and / or stream index data;

[0156] - Request the transfer of the first part set and the second part set based on the available parts.

[0157] According to an embodiment, the player device is configured as follows:

[0158] - Obtain metadata on the area covered by the tile stream from media presentation description and / or stream initialization data and / or stream index data; and

[0159] - Infer the first tile flow set and the second tile flow set based on metadata.

[0160] In this embodiment, the media presentation description conforms to the DASH MPD.

[0161] In this embodiment, the stream initialization data conforms to the initialization fragments specified in DASH and / or OMAF v2.

[0162] In this embodiment, the stream index data conforms to the index fragments specified in DASH and / or OMAF v2.

[0163] According to an embodiment, the player device is configured to obtain metadata as a spherical region from one or more of the following:

[0164] - Sphere Region Quality Ranking Box

[0165] Content Coverage Box

[0166] - Spherical Region Quality Sort Descriptor (SRQR Descriptor)

[0167] -Content Override Descriptor (CC Descriptor)

[0168] According to an embodiment, the player device is configured to obtain metadata over the area covered by the tile stream as a two-dimensional region, such as a rectangle, for example, an image relative to the projection, from one or more of the following:

[0169] - 2D Region Quality Ranking Box

[0170] - 2D Region Quality Ranking (2DQR) Descriptor

[0171] - Such as the Spatial Relationship Descriptor (SRD Descriptor) specified in DASH.

[0172] The following describes several embodiments for implementing the proposed operations on a player device:

[0173] In the conventional approach, presented only for reference, the player device is able to download complete segments for all tiles, as well as complete segments for the foreground, and use them sub-segments one by one. However, this requires additional bandwidth because the player needs to download high-quality additional complete segments to provide low-latency response to viewport switching.

[0174] According to the first embodiment, the player device is configured to first extract a segment index for each foreground tile fragment, and then extract foreground sub-fragments one by one as needed. Extracting the segment index individually (i.e., the 'sidx' box) increases the rate of HTTP requests (because it causes an additional HTTP GET request per segment for each DASH representation). However, if the segment duration is, for example, 3 seconds, and the sub-fragments are ~300 milliseconds long, there are 10 sub-fragments, resulting in a maximum of 11 HTTP requests per segment, unless there is a large amount of header movement, in which case the segment index request portion increases. When a switch occurs within a segment, the number of requests required is equal to the number of sub-fragments plus one.

[0175] MPDs can include a byte range containing the fragment index in the @indexRange attribute. The player device can obtain the byte range for the fragment index from the @indexRange attribute.

[0176] Additionally, the player device can estimate the fragment index and the size of the first sub-fragment using applicable values ​​such as @indexRange (if available) and known properties of the tile bitstream (such as bitrate and fragment duration) or fragment sizes from other similar tilestreams belonging to the same process, and combine them into a single request, thus saving an HTTP request.

[0177] In the proposed embodiment, the HTTP request rate is not very high compared to the case where all representations are extracted as short fragments. This is because the lower request rate for non-critical flows compensates for latency-critical and bandwidth-starved foreground flows.

[0178] The second embodiment is based on a new standardized technology. OMAF version 2 introduces the use of index fragments, which were previously introduced into the ISOBMFF and DASH specifications. Index fragments provide access to a fragment index, as if through an index fragment URL separate from the media fragment URL.

[0179] OMAF version 2 includes new fragment formats: initialization fragments for base tracks, tile index fragments, and tile data fragments, which are described below. A base track can be, for example, an extractor track. The initialization fragment for a base track contains the track header for the base track and all referenced tile tracks. Therefore, clients can download only the initialization fragment for the base track without needing to download the initialization fragments for the tile tracks.

[0180] In this embodiment, the tile index fragment is logically an index fragment as defined in the DASH standard. It contains a SegmentIndexBox for the base track and all referenced tile tracks, and a MovieFragmentBox for the base track and all referenced tile tracks. The MovieFragmentBox indicates a byte range on a sample basis. Therefore, the client can selectively request content on units smaller than the (sub)fragment.

[0181] Tile data fragments contain only media data encapsulated within a specific media data box called an IdentifiedMediaDataBox. The byte offset within a MovieFragmentBox is relative to the beginning of the IdentifiedMediaDataBox. Therefore, unlike in regular DASH fragment formats, MovieFragmentBoxes and media data can reside in separate resources.

[0182] As described above, the tile index fragment contains (sub)fragments for all tile stream indices, thus eliminating the need to extract the index separately for each tile stream. It is expected to add a new HTTP connection, but only one for the entire DASH stream. Similar to the first embodiment, in the second embodiment, the player device can download the foreground tile as a sub-fragment with a byte-range HTTP request as needed, utilizing the 'sidx' information and / or MovieFragmentBox in the tile index fragment, and extract the foreground and other non-critical streams as complete fragments.

[0183] The third embodiment is based on segments with multiple durations. The server can create tile segments with multiple durations, eliminating the need to extract sub-segment indexes. Each duration can be its own DASH representation, and therefore each segment of duration has a unique URL. Representations for tiles at background quality can have longer durations than representations for tiles at foreground quality. All representations can have the same video characteristics, such as random access periods (video GOPs) to enable partial decoding, but they are simply DASH segments packed with different segment durations (1 GOP per DASH segment versus, for example, 4 GOPs == 4 sub-segments per DASH segment). This may be trivial for RWMR schemes because background tiles are in a different fitness set than foreground tiles. Therefore, segment durations differ between fitness sets but are constant between representations within a fitness set. For multi-quality VD schemes, the same fitness set can include representations with different segment durations. Additionally, background-to-foreground contrast can be indicated within the MPD metadata for the representation. Therefore, the lowest quality representation (or possibly several lowest quality representations) can have longer DASH segments than others.

[0184] The DASH@segmentAlignment requirement in OMAF (additional B / HEVC VD profile definition) that requires each segment to have the same duration can be relaxed.

[0185] Figure 6 The diagram illustrates the interaction with on-demand streaming. Figure 5 The same situation, in which Figure 6 A live stream with four sub-segments in one DASH segment is shown, with the number corresponding to the number of sub-segments. The difference is that when switching back to low quality, the player needs to fetch the low-quality segment with sub-segments 5, 6, 7, and 8 to be able to switch to low quality at sub-segment 7. This can be a small overhead, but it can be avoided using the second embodiment, where the segment index for the low-quality stream can also be obtained using a single indexed segment stream for the high-quality stream downloaded in any way. Additionally, because sub-segments 4 and 5 come from different DASH segments, the player device needs to have separate segment indexes for them. In the first embodiment, this would require two additional HTTP requests for this situation. In the reference embodiment, the player device will fetch two complete segments, a total of eight sub-segments, and will skip sub-segments 1, 2, and 3, and will show the high-quality sub-segments 4, 5, 6, 7, and 8. Furthermore, since the download of the first high-quality segment starts late, it is possible that the first high-quality segment cannot be downloaded in time before sub-segment 4 needs to be rendered. This is because downloading large segments takes time, so only high-quality rendering is possible at segment 5, which increases the switching latency.

[0186] exist Figure 7 The diagram illustrates a method according to an embodiment. The method generally includes: determining a foreground region covering a viewport of a 360-degree video and one or more other regions of the 360-degree video that do not entirely contain the foreground region; inferring a first set of tile streams among available tile streams of the 360-degree video, wherein the first set of tile streams covers the foreground region; inferring a second set of tile streams among available tile streams of the 360-degree video, wherein the second set of tile streams covers one or more other regions; and requesting transmission of a first portion of the first set of tile streams and a second portion of the second set of tile streams, wherein portions of the first set have a shorter duration than portions of the second set. Each step can be implemented by a corresponding module of a computer system.

[0187] The apparatus according to an embodiment includes the following components:

[0188] - Components used to determine the foreground region of the viewport covering the 360-degree video and one or more other regions of the 360-degree video that do not include the entire foreground region;

[0189] - A component for inferring a first set of tile streams among the available tile streams in a 360-degree video, wherein the first set of tile streams covers the foreground region;

[0190] - A component for inferring a second set of tile streams from the available tile streams of the 360-degree video, wherein the second set of tile streams covers one or more additional regions; and

[0191] - A component for requesting the transmission of a first portion of a first tile stream set and a second portion of a second tile stream set, wherein a portion of the first portion has a shorter duration than a portion of the second portion.

[0192] These components include at least one processor and a memory comprising computer program code, wherein the processor may further include processor circuitry. According to various embodiments, the memory and computer program code are configured to enable the device to execute using at least one processor. Figure 7 The method.

[0193] exist Figure 8 An example of the device is shown. If needed, several functionalities can be executed using a single physical device, such as a single processor. Device 90 includes a main processing unit 91, a memory 92, a user interface 94, and a communication interface 93. Figure 8 The apparatus according to the embodiment shown also includes a camera module 95. A memory 92 stores data including computer program code from the apparatus 90. The computer program code is configured to implement the... Figure 7 The flowchart method is as follows. Camera module 95 receives input data in the form of a video stream for processing by processor 91.

[0194] Communication interface 93 forwards processed data, for example, to the display of another device, such as an HMD. When device 90 is a video source including camera module 95, user input can be received from the user interface. If device 90 is an intermediate box in a network, the user interface is optional, such as the camera module.

[0195] Various embodiments can be implemented using computer program code residing in memory and causing the associated apparatus to perform methods. For example, a device may include circuitry and electronics for processing, receiving, and transmitting data, computer program code in memory, and a processor that, when executing the computer program code, causes the device to perform the features of the embodiments. Additionally, a network device, such as a server, may include circuitry and electronics for processing, receiving, and transmitting data, computer program code in memory, and a processor that, when executing the computer program code, causes the network device to perform the features of various embodiments.

[0196] The computer program product according to the embodiments can be embodied on a non-transitory computer-readable medium. According to another embodiment, the computer program product can be downloaded in a data packet via a network.

[0197] If necessary, the different functions discussed herein can be executed in different orders and / or concurrently with other functions. Additionally, if necessary, one or more of the functions and embodiments described above can be optional or can be combined.

[0198] Although various aspects of the embodiments are set forth in the independent claims, other aspects include other combinations of features from the described embodiments and / or dependent claims with features of the independent claims, not just those explicitly set forth in the claims.

[0199] It should also be noted that although exemplary embodiments have been described above, these descriptions should not be considered limiting. Rather, several variations and modifications are possible without departing from the scope of this disclosure as defined by the appended claims.< / maxoccurs> < / minoccurs>

Claims

1. A method comprising: - Identify the foreground region of the viewport covering the 360-degree video and one or more other regions of the 360-degree video that do not include the entire foreground region; - Infer a first set of tile streams among the available tile streams of the 360-degree video, wherein the first set of tile streams covers the foreground region; - Infer a second set of tile streams from the available tile streams of the 360-degree video, wherein the second set of tile streams covers the one or more other regions; as well as - A request for transmission of a first portion of the first tile stream set and a second portion of the second tile stream set, wherein a portion of the first portion has a shorter duration than a portion of the second portion, wherein the transmission of the first portion of the first tile stream set and the second portion of the second tile stream set is requested by the player before a viewport switch.

2. The method of claim 1, wherein the one or more other regions include edge regions and background regions, and the method further comprises: - Identify the edge region adjacent to the foreground region and the background region that covers the 360-degree video but is not included in the foreground region or the edge region; - Infer a first subset of the second tile flow set, wherein the first subset of the second tile flow set covers the edge region; - Infer a second subset of the second tile flow to include tile flows that are in the second tile flow set but not in the first subset of the second tile flow set; as well as - Request the transmission of a first subset of a portion of a first subset of a tile stream and a second subset of a portion of a second subset of a tile stream, wherein a portion of the first subset of a portion of a first subset of a tile stream has a shorter duration than a portion of the second subset of a portion of a second subset of a tile stream.

3. The method according to claim 1, further comprising: A request is made for the transfer of an extractor track or the like in a portion of the set that has a longer duration than the first set.

4. The method according to claim 1, further comprising: - Obtain metadata on the area covered by the tile stream from media presentation description and / or stream initialization data and / or stream index data; as well as - The first tile flow set and the second tile flow set are inferred based on the metadata.

5. A computer program product comprising computer program code configured to, when executed on at least one processor, cause a device or system to: - Identify the foreground region of the viewport covering the 360-degree video and one or more other regions of the 360-degree video that do not include the entire foreground region; - Infer a first set of tile streams among the available tile streams of the 360-degree video, wherein the first set of tile streams covers the foreground region; - Infer a second set of tile streams from the available tile streams of the 360-degree video, wherein the second set of tile streams covers the one or more other regions; as well as - A request for transmission of a first portion of the first tile stream set and a second portion of the second tile stream set, wherein a portion of the first portion has a shorter duration than a portion of the second portion, wherein the transmission of the first portion of the first tile stream set and the second portion of the second tile stream set is requested by the player before a viewport switch.

6. The computer program product of claim 5, wherein the one or more other regions comprise edge regions and background regions, and the computer program product comprises computer program code configured to perform the following: - Identify the edge region adjacent to the foreground region and the background region that covers the 360-degree video but is not included in the foreground region or the edge region; - Infer a first subset of the second tile flow set, wherein the first subset of the second tile flow set covers the edge region; - Infer a second subset of the second tile flow to include tile flows that are in the second tile flow set but not in the first subset of the second tile flow set; as well as - Request the transmission of a first subset of a portion of a first subset of a tile stream and a second subset of a portion of a second subset of a tile stream, wherein a portion of the first subset of a portion of a first subset of a tile stream has a shorter duration than a portion of the second subset of a portion of a second subset of a tile stream.

7. The computer program product according to claim 5, further comprising: Computer program code configured to request the transfer of extractor tracks or the like in a portion having a duration longer than the first set of portions.

8. The computer program product of claim 5, further comprising computer program code configured to perform the following: - Obtain metadata on the area covered by the tile stream from media presentation description and / or stream initialization data and / or stream index data; and - The first tile flow set and the second tile flow set are inferred based on the metadata.

9. An apparatus comprising at least: - Components for determining the foreground region of a viewport covering a 360-degree video and one or more other regions of the 360-degree video that do not include the entire foreground region; - A component for inferring a first set of tile streams among the available tile streams of the 360-degree video, wherein the first set of tile streams covers the foreground region; - A component for inferring a second set of tile streams among the available tile streams of the 360-degree video, wherein the second set of tile streams covers the one or more other regions; as well as - A component for requesting the transmission of a first portion of the first tile stream set and a second portion of the second tile stream set, wherein a portion of the first portion has a shorter duration than a portion of the second portion, wherein the transmission of the first portion of the first tile stream set and the second portion of the second tile stream set is requested by the player before a viewport switch.

10. The apparatus of claim 9, wherein the one or more other regions comprise an edge region and a background region, and the apparatus further comprises: - Components for determining the edge region adjacent to the foreground region and the background region that is not included in the foreground region or edge region and covers the 360-degree video; - A component for inferring a first subset of the second tile stream set, wherein the first subset of the second tile stream set covers the edge region; - A component for inferring a second subset of the second tile stream to include tile streams that are in the second tile stream set and not in the first subset of the second tile stream set; as well as - A component for requesting the transmission of a first subset of a portion of a first subset of a tile stream and a second subset of a portion of a second subset of a tile stream, wherein a portion of the first subset of a portion of a first subset of a tile stream set has a shorter duration than a portion of the second subset of a portion of a second subset of a tile stream set.

11. The apparatus of claim 9, further comprising: A component used to request the transfer of an extractor track or the like in a portion having a longer duration than the first set of portions.

12. The apparatus of claim 9, further comprising: - A component for obtaining metadata on the area covered by the tile stream from media presentation description and / or stream initialization data and / or stream index data; as well as - A component for inferring the first tile stream set and the second tile stream set based on the metadata.

Citation Information

Patent Citations

  • Video data processing method and apparatus

    CN107888993A

  • Systems and methods for providing content

    US20180190327A1