Viewport-based tiled streaming
By employing semantic segmentation maps and image synthesis models to generate high-quality video tiles, the method addresses latency and quality degradation issues in viewport-based streaming, ensuring smooth transitions and improved user experience.
Patent Information
- Application Number
- PCT/EP2024/088659
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-31
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-03
AI Technical Summary
Existing viewport-based video streaming systems experience latency issues and quality degradation when switching between high and low-quality video tiles, leading to a loss of quality of experience and motion sickness, particularly in head-mounted displays.
Implement a method that uses semantic segmentation maps and image synthesis models to generate high-quality video tiles by decoding and synthesizing peripheral tiles based on metadata, allowing seamless transitions without waiting for high-quality tiles to be fetched.
Minimizes latency in quality switching and maintains consistent video quality, reducing motion sickness and optimizing bandwidth usage in viewport-based video streaming.
Smart Images

Figure EP2024088659_03072025_PF_FP_ABST
Abstract
Description
[0001] Viewport-based tiled streaming
[0002] Technical field
[0003] The embodiments relate to viewport-based tiled streaming, and, in particular, though not exclusively, to methods and systems for viewport-based tiled streaming, a video data structure and a computer program product for executing such methods.
[0004] Background
[0005] Omnidirectional Media Format (OMAF) is a video streaming standard developed by MPEG for enabling video streaming for media applications with omnidirectional content including video, images, audio and timed text. OMAF uses the ISO Base Media File Format (ISOBMFF) for encapsulation, Dynamic Adaptive Streaming over HTTP (DASH) for media delivery and MPEG Media Transport (MMT) for signaling. Additionally, OMAF supports standardized codecs, such as AVC, HEVC and VVC, which use Supplemental Enhancement Information (SEI) messages to deliver arbitrary user defined data such as timing information, video properties and postprocessing enhancement methods that can be used along with Video Usage Information (VUI) for high level properties like color space.
[0006] To efficiently deliver omnidirectional content adapted to the user's viewport OMAF and DASH make use of so-called Motion-Constrained Tile Sets (MCTS) to spatially partition the video frames of a source video file into tiles. A video encoder supporting MCTS restricts inter-temporal prediction at tile boundaries and eliminates correlation between tiles. This way each tile can be extracted from video data of an original video source file and transmitted to the client. Encoded video data associated with a tile may be formatted in video frames, which can be streamed (as a tile stream or an MCTS stream) to a client device. Hereunder, a tile refers to a sequence of video frames comprising a content of a predetermined region, typically a rectangular region, in video frames of a video source. Based on video data of the tile streams video frames or a part of the original source file can be constructed. Metadata associated with a tile may determine the spatial position of the tile in video frames of the video source. Several resolution versions of tiles may be encoded at different bitrates and made available for streaming as a so-called representation.
[0007] This scheme is used by the client device to request one or more high-quality tiles (i.e. tiles associated with high-quality video data which is e.g. required for immersive rending of omnidirectional video) that lie within the viewport of the viewport-based video playout device, e.g. a head-mounted device (HMD) or the like. When the viewport moves, the client device associated with the playout device may determine which new tiles within the viewport are required display and send one or more requests to retrieve these new tiles in high-quality which typically can be displayed up to 5 seconds thereafter. To address this latency problem, conventionally low-quality video data of (part of) the source video is sent to the client device so that these low-quality video data can be used to temporally construct low-quality versions of the required new tiles. This latency in switching quality results in rendering low quality tiles along with high quality tiles, which can lead to a loss of quality of experience (QoE) and induce motion sickness when the content is viewed using head mounted displays (HMDs).
[0008] Hence, from the above it follows there is a need in the art for improved methods and systems for viewport switching.
[0009] Summary of the invention
[0010] As will be appreciated by one skilled in the art, aspects of the present invention may be embodied as a system, method or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a "circuit," "module" or "system." Functions described in this disclosure may be implemented as an algorithm executed by a microprocessor of a computer. Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code embodied, e.g., stored, thereon.
[0011] Any combination of one or more computer readable medium(s) may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0012] A computer readable signal medium may include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium may be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0013] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber, cable, RF, etc., or any suitable combination of the foregoing. Computer program code for carrying out operations for aspects of the present invention may be written in any combination of one or more programming languages, including an object-oriented programming language such as Java(TM), Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0014] Aspects of the present invention are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor, in particular a microprocessor or central processing unit (CPU), of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer, other programmable data processing apparatus, or other devices create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0015] These computer program instructions may also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks. The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. Additionally, the Instructions may be executed by any type of processors, including but not limited to one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FP- GAs), or other equivalent integrated or discrete logic circuitry.
[0016] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0017] The embodiments in this disclosure describe scheme to reduce the latency of quality switching after viewport movements of a viewport based display device and to optimize how available bandwidth is spent while streaming so-called tiled omnidirectional video wherein pictures (video frames) of omnidirectional video data are spatially divided in independently decodable non-overlapping tiles - video tiles - that are made available at multiple qualities.
[0018] In an aspect, the embodiments relate to a method of processing tiled video data by a client device comprising a step of: requesting by a client device transmission of video tiles based on the position of a viewport of a display device, the video tiles including one or more viewport tiles having a position within the viewport and one or more peripheral tiles, each peripheral tile having a position outside the viewport, the one or more viewport tiles comprising encoded video data of a first resolution and the one or more peripheral tiles comprising encoded data of one or more semantic segmentation maps, preferably grayscale semantic segmentation maps. The method may also comprise the step of receiving requested video tiles and metadata associated with the requested video tiles by the client device, the metadata including semantic labels for labelling pixel regions in the one or more semantic segmentation maps, wherein the client device is associated with an image synthesis model, preferably an image synthesis model based on one or more artificial neutral networks, which is trained to determine a synthesized peripheral tile of the first resolution based on decoded data of a semantic segmentation map and semantic labels associated with the semantic segmentation map. The method may further comprise the steps of decoding the encoded video data of the one or more viewport tiles into decoded video data of the one or more viewport tiles; and, storing the decoded video data in a decoded playback buffer of a display device associated with the client device for displaying at least part of the one or more viewport tiles in the viewport.
[0019] In an embodiment, the semantic segmentation map is a grayscale semantic segmentation map comprising encoded pixel values in greyscale.
[0020] In an embodiment, a pixel region may comprise pixels with the same greyscale pixel value.
[0021] Hence, a client fetches tiles associated with viewport of a high quality (viewport tiles) and tiles outside the viewport (peripheral tiles) wherein the peripheral tiles may be represented by a semantic segmentation map and semantic labels associated with the semantic segmentation map. During playout, if the viewport moves, peripheral tiles enter into the viewport along with old tiles that remain in the viewport. In that case, the peripheral tiles may be used to generate synthesized tiles of a quality that approximate the quality of the viewport tiles. Here, the synthesized tiles are generated by an image synthesis model which is trained to receive a semantic segmentation map and associated semantic labels at its input and to output a synthesized tile which has a quality that approximates the quality of the viewport tiles. This way, viewport tiles in high quality can be displayed next to synthesized tiles so that tiles visible in the viewport are presented at the same final quality.
[0022] In a further embodiment, client device may request the peripheral tiles that were used for generating the synthesized tiles in the high quality (the first resolution). Once these are decoded and available for playout, the synthesized tiles may be replaced with high quality tiles.
[0023] In an embodiment, the method may comprise decoding the encoded video data of the one or more peripheral tiles into decoded data of one or more semantic segmentation maps, preferably the encoded video data being decoded while decoding the encoded video data of the one or more viewport tiles and storing the decoded video data of the one or more viewport tiles in the decoded playback buffer.
[0024] In an embodiment, the encoded video data of the one or more peripheral tiles may be stored until one or more decoder instances of a decoder system, preferably a video decoding engine, associated with the client device are available for the decoding of the encoded video data.
[0025] In an embodiment, the method may comprise: generating one or more synthesized peripheral tiles of the first resolution by providing the decoded data of the one of the one or more semantic segmentation maps and associated semantic labels to the input of the image synthesis model; and, storing video data associated with the one or more synthesized peripheral tiles in the decoded playback buffer of the display device.
[0026] In an embodiment, the decoded data of the one or more peripheral tiles may be stored until computational resources are available for processing the decoded data of the one of the one or more semantic segmentation maps and associated semantic labels by the image synthesis model.
[0027] In an embodiment, the method may further comprise: displaying at least part of the one or more viewport tiles of the first resolution in the viewport of the display device if at least part of the one or more viewport tiles are positioned in the viewport of the display device.
[0028] In an embodiment, the method may further comprise: displaying at least part of the one or more synthesized peripheral tiles of the first resolution in the viewport of the display device if at least part of the one or more synthesized peripheral tiles have entered the viewport of the display device.
[0029] In an embodiment, the one or more peripheral tiles may further comprise encoded data of one or more peripheral video tiles of a second resolution which is lower than the first resolution.
[0030] In an embodiment, the method may further comprise decoding the encoded video data of the one or more peripheral video tiles of the second resolution; generating one or more upscaled peripheral video tiles of the first resolution by upscaling decoded video data of the of the one or more peripheral video tiles of the second resolution; storing the decoded video data in the decoded playback buffer of the display device; and, displaying at least part of the one or more upscaled peripheral tiles of the first resolution if the one or more upscaled peripheral tiles have entered the viewport of the display device.
[0031] In an embodiment, the requested video tiles and metadata associated with the requested video tiles may be transmitted to the client device based on a video streaming standard.
[0032] In an embodiment, the video streaming may be an adaptive video streaming standards such as MPEG Dynamic Adaptive Streaming over HTTP (DASH) or MPEG Omnidirectional Media Format (OMAF) and / or wherein the video tiles are formatted as motion-constrained tile sets (MOTS). In an embodiment, the semantic labels or information to determine the semantic labels may be transmitted using a one or more in-band messages, e.g. one or more DASH in band event messages, associated with the video tiles.
[0033] In an embodiment, the video tiles may be requested by the client device based on information in a manifest file, preferably a media presentation description (MPD), the manifest file including information for identifying video tiles, preferably one or more resource locators, e.g. URLs, or information for constructing such resource locators.
[0034] In an further aspect, the embodiments in this disclosure may relate to an apparatus for processing video data comprising: a computer readable storage medium having at least part of a program embodied therewith; and, a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, wherein the processor may be configured to perform executable operations comprising, wherein the executable operations may include requesting by a client device transmission of video tiles based on the position of a viewport of a display device, the video tiles including one or more viewport tiles having a position within the viewport and one or more peripheral tiles, each peripheral tile having a position outside the viewport, the one or more viewport tiles comprising encoded video data of a first resolution and the one or more peripheral tiles comprising encoded data of one or more semantic segmentation maps, preferably grayscale semantic segmentation maps. The executable operations may also include receiving requested video tiles and metadata associated with the requested video tiles by the client device, the metadata including semantic labels for labelling pixel regions in the one or more semantic segmentation maps, wherein the client device is associated with an image synthesis model, preferably an image synthesis model based on one or more artificial neutral networks, which is trained to determine a synthesized peripheral tile of the first resolution based on decoded data of a semantic segmentation map and semantic labels associated with the semantic segmentation map. The executable operations may further include decoding the encoded video data of the one or more viewport tiles into decoded video data of the one or more viewport tiles; and, storing the decoded video data in a decoded playback buffer of a display device associated with the client device for displaying at least part of the one or more viewport tiles in the viewport.
[0035] In various other embodiments, the processor may be further configured to perform any of the method steps as described with reference to the above-described embodiments.
[0036] In yet a further aspect, the embodiments may relate to a video preparation module comprising: a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, wherein the processor may be configured to perform executable operations comprising: generating one or more sets of tiles, e.g. Motion-Constrained Tile Sets (MCTSs), based on video source file comprising video data of a first resolution; analyzing each of the one or more tile sets to determine if a tile is suitable for image synthesis using an image synthesis model, the image synthesis model being trained to determine a synthesized tile of the first resolution based on a semantic segmentation map of the tile and semantic labels associated with the semantic segmentation map; determining at least a high resolution tile representation comprising encoded and packetized video data of the first resolution and a baseline tile representation comprising encoded and packetized data representing a semantic segmentation map of a tile that is suitable for image synthesis and encoded and packetized data representing semantic labels associated with the semantic segmentation map; and, storing the high resolution tile representation and the a baseline tile representation and an manifest file identifying the high resolution tile representation and the baseline representation representations at a server
[0037] The invention may also relate to a computer program product comprising software code portions configured for, when run in the memory of a computer, executing the method steps according to any of process steps described above.
[0038] The invention will be further illustrated with reference to the attached drawings, which schematically will show embodiments according to the invention. It will be understood that the invention is not in any way restricted to these specific embodiments.
[0039] Brief description of the drawings
[0040] Fig. 1 illustrates a tiled video streaming data format;
[0041] Fig. 2A-2C illustrate the use of a video streaming data format according to an embodiment;
[0042] Fig. 3 depicts a post-processing scheme for a video streaming data format according to an embodiment;
[0043] Fig. 4 depicts a flow diagram of a method of processing video based on a video streaming data format according to an embodiment;
[0044] Fig. 5 depicts a system for video streaming based on a video steaming data format according to an embodiment;
[0045] Fig. 6 depicts a schematic of a decoder system for decoding video data associated with a video steaming data format according to an embodiment; Fig. 7 depicts a schematic of a post-processor block according to an embodiment;
[0046] Fig. 8 depicts a schematic of processing of video tiles by a video decoding engine according to an embodiment;
[0047] Fig. 9 depicts a block diagram illustrating an exemplary data processing system that may be used with embodiments described in this disclosure.
[0048] Detailed description
[0049] Omnidirectional Media Format (OMAF) is a video streaming standard developed by MPEG for enabling video streaming for media applications with omnidirectional content including video, images, audio and timed text. OMAF uses the ISO Base Media File Format (ISOBMFF) for encapsulation, Dynamic Adaptive Streaming over HTTP (DASH) for media delivery and MPEG Media Transport (MMT) for signaling. Additionally, OMAF supports standardized codecs, such as AVC, HEVC and VVC, which use Supplemental Enhancement Information (SEI) messages to deliver arbitrary user defined data such as timing information, video properties and postprocessing enhancement methods that can be used along with Video Usage Information (VUI) for high level properties like color space.
[0050] To efficiently deliver omnidirectional content adapted to the user's viewport OMAF and DASH make use of so-called Motion-Constrained Tile Sets (MCTS) to spatially partition the video frames of a source video file into tiles. A video encoder supporting MCTS restricts inter-temporal prediction at tile boundaries and eliminates correlation between tiles. This way each tile can be extracted from video data of an original video source file and transmitted to the client. Encoded video data associated with a tile may be formatted in video frames, which can be streamed (as a tile stream or an MCTS stream) to a client device. Hereunder, a tile refers to a sequence of video frames comprising a content of a predetermined region, typically a rectangular region, in video frames of a video source. Based on video data of the tile streams video frames or a part of the original source file can be constructed. Metadata associated with a tile may determine the spatial position of the tile in video frames of the video source. For supporting adaptive streaming of tile streams, a tile stream may be temporally divided in segments, i.e. short media filed comprising approx. 1-20 seconds of video data. In DASH such segments are referred to as DASH segments and in CMAF such segments are referred to as CMAF fragment.
[0051] Several resolution versions of tiles may be encoded at different bitrates and made available for streaming as a so-called representation. Each MCTS sequence can be made available for example as an HEVC-compliant sub-picture track. HEVC allows partitioning along a grid of rows and columns. The encoding is then restricted in such a way that each MCTS can reference pixels only within the same MCTS in the current and reference video frame. AVC does not support tiles, however, MCTS behavior can be achieved by arranging regions vertically as slices. This approach allows the use of one decoder instance per sub-picture track. OMAF specifies the use of so-called extractor tracks to process the MCTS into a single bitstream. The client device is configured to extract this extractor track from the bitstream to obtain the decodable bitstream.
[0052] This scheme is used by the client device to request one or more high-quality tiles (i.e. tiles associated with high-quality video data which is e.g. required for immersive rending of omnidirectional video) that lie within the viewport of the viewport-based video playout device, e.g. a head-mounted device (HMD) or the like. When the viewport moves, the client device associated with the playout device may determine which new tiles within the viewport are required display and send one or more requests to retrieve these new tiles in high-quality which typically can be displayed up to 5 seconds thereafter. To address this latency problem, conventionally low-quality video data of (part of) the source video is sent to the client device so that these low-quality video data can be used to temporally construct low-quality versions of the required new tiles. This latency in switching quality results in rendering low quality tiles along with high quality tiles, which can lead to a loss of quality of experience (QoE) and induce motion sickness when the content is viewed using head mounted displays (HMDs).
[0053] Hence, from the above it follows that omnidirectional video may be delivered to users on bandwidth limited networks using independently decodable video tiles of varying quality. The quality of each video tile is decided based on the user's viewport that defines a region of interest (ROI) to optimize how the available bandwidth is spent. To that end, tiles are created in a data format so that adaptive tiled video streaming can be realized. Fig. 1 illustrates a data format 102 for tiled video streaming wherein pictures of omnidirectional video data are spatially divided in sub-pictures, referred to as tiles. Each tile may be available in different qualities, e.g. high resolution and low resolution. A client device may determine tiles that fall within the viewport 106 of a viewport-based display device and request these tiles in high-resolution. These tiles may also be referred to as viewport tiles. Further, the client device may request tiles outside the viewport (referred to as periphery tiles) as low- resolution tiles. When the viewport moves some of the periphery tiles may enter the viewport at least low-resolution versions of the periphery tiles that enter the viewport are available for rendering without the need to wait for requesting and receiving the new high-resolution tiles. Once the high-quality tiles are received, the low-quality versions are replaced. Hence, in this case at least temporarily, both high quality and low-quality tiles are displayed in the viewport thus degrading the quality of experience and inducing motion sickness when the content is viewed using head mounted displays (HMDs). Fig. 2A-2C illustrate the use of a data format for adaptive tiled streaming according to an embodiment. In particular, the figures depict a data format for tiled video streaming that provides fast high-quality viewport switching which eliminates or at least reduces part of the problems related to known adaptive tiled streaming schemes as for example illustrated in Fig. 1. As shown in Fig. 2A, each tile is available in a high-resolution tile and an associated baseline tile. These baseline tiles comprise encoded video data which after decoding can be post-processed in a high-resolution version of the tile that is substantially similar to the associated high-resolution tile.
[0054] In an embodiment, a baseline tile 208i may represent a semantic segmented map with associated semantic labels which can be synthesized using a trained image synthesis model into a synthesized video tile of a resolution that is similar or identical to the high-resolution version of the tile. In another embodiment, other types of baseline tiles may also be used. For example, in an embodiment, a second type of baseline tile 2082 may be a low-resolution version of the associated high-resolution tile which can be upscaled into an upscaled video tile. An image upscaling model, e.g. a convolutional neural network, may be trained to receive a low-resolution image and to generate an upscaled video tile that has a resolution that is similar or identical to the associated high-resolution version of the tile. Baseline tiles may be associated with metadata for signaling a baseline tile type and information that is needed for post-processing a baseline tile into a video tile that is similar or identical to the associated high-resolution tile. For example, in an embodiment, metadata associated with baseline tile for identifying groups of pixels that are associated with a semantic class. For example, the baseline tiles in the new viewport include semantic segmented maps identifying pixels related to “sandy ground”. A semantic segmented map may be associated with many different semantic labels so that a synthesized tile can be generated with sufficient detail. Thus, when using the data format depicted in Fig. 2 for tiled video streaming, the current viewport 206i is used to identify high-quality viewport tiles, which are requested and displayed by a client device as high-quality tiles. The rest of the tiles or at least the tiles around the periphery of the viewport may be streamed as peripheral tiles.
[0055] Fig. 2B shows that the current viewport 2062 has shifted and encompasses new tiles. The high-resolution video data associated with these tiles are not yet available in the buffer and need to be requested. As however described with reference to Fig. 2A, the new tiles however are available in the buffer as peripheral tiles, in this example peripheral tiles comprising encoded video data of semantic segmentation maps. Hence, the client device may decode encoded data associated with the peripheral tiles and use the decoded data and associated semantic labels as input to a trained image synthesis model for generating synthesized tiles that is similar or identical to the associated high-resolution tiles. The synthesized tiles may be stored in the decoded playback buffer and displayed in the viewport of the display device, while the client device may request the current viewport tiles in high quality from the server. In Fig. 2C, the newly requested high-quality tiles in the current viewport 206s have been successfully retrieved from the server by the client, decoded and displayed in the viewport instead of the synthesized tiles.
[0056] The concept of baseline tiles is further illustrated in Fig. 3, which schematically illustrates a first baseline tile representing a semantic segmentation map 304i and a second baseline tile representing a low-resolution video tile 3042 that is suitable for upscaling. The semantic segmentation map may be a grayscale segmentation map including different pixel areas 314I-3 associated with different grayscale values 313I-3. Each grayscale level is associated with a semantic label 312I-3 which defines a semantic class indicating the type of content associated with pixels of a grayscale level. The semantic labels and the grayscale semantic segmentation map may be provided to the input of an image synthesis model 308 that is trained to generate a synthesized video tile 3101 that is similar or substantially identical to its associated high quality video tile. Similarly, the low-resolution video tile 3042 may be provided to an image upscaling model 306 that is trained for upscaling the low-resolution video tile into an upscaled video tile 31 O2 that is similar or substantially identical to its associated high quality video tile. This way, the synthesized video tile and / or the upscaled video tile, which are available in the display buffer of a display device, may be displayed in the viewport together with high quality viewport tiles without any delay and with minimal distortion of the user experience.
[0057] The client device may take the advantage of synthesis and upscaling to postprocess baseline quality tiles that enter the viewport before or at the moment these tiles are presented to the viewer into high-resolution tiles. When the viewport moves and new tiles become visible, the user is presented with synthesized and upscaled tiles at the same or similar resolution as existing tiles, while high-quality tiles are being fetched. This way the latency associated with which the quality switching is minimized or at least reduced thereby including the quality of experience, reducing VR motion sickness, and saving or freeing additional bandwidth.
[0058] Tile streams, including baseline tile streams that need postprocessing, may be stored as different tracks of an ISOBMFF compliant media file. The high-resolution tiles and the baseline tiles may be made available to a client device using a manifest file. Based on the viewport and the information in the manifest file, the client device can request streaming of video data representing one or more tiles and associated metadata, including high resolution viewport tiles positioned within the viewport and peripheral tiles of a baseline quality (a grayscale semantic segmented map or a low-quality tile that requires upscaling) that are positioned outside the viewport. Fig. 4 depicts a flow chart of a method of processing tiled video data according to an embodiment. The process may start with requesting by a client device transmission of video tiles based on the position of a viewport of a display device, the video tiles including one or more viewport tiles having a position within the viewport and one or more peripheral tiles, each peripheral tile having a position outside the viewport (step 402). Here, the one or more viewport tiles may comprise encoded video data of a first resolution, typically a high-resolution. Further, the one or more peripheral tiles may comprise encoded data one or more semantic segmentation maps. In an embodiment, the one or more semantic segmentation maps may be grayscale semantic segmentation maps, which can be efficiently encoded. This may be used to optimize the bandwidth usage when streaming tiled omnidirectional video data to client devices.
[0059] Thereafter, the requested video tiles and metadata associated with the requested video tiles may be received by the client device wherein the metadata may include semantic labels for labelling pixel regions in the one or more semantic segmentation maps (step 404). The client device may be associated with an image synthesis model which is configured to post-process the decoded data representing the semantic segmentation map. In an embodiment, the image synthesis model may be based on one or more artificial neutral networks, which is trained to determine a synthesized peripheral tile of the first resolution based on decoded data of a semantic segmentation map and semantic labels associated with the semantic segmentation map.
[0060] In step 406, the encoded video data of the one or more viewport tiles may be decoded into decoded video data of the one or more viewport tiles which may be stored in a decoded playback buffer of the display device for displaying at least part of the one or more viewport tiles in the viewport.
[0061] Hence, the method allows a client device to request tiles associated with viewport of a high quality (viewport tiles) and tiles outside the viewport (peripheral tiles) wherein the peripheral tiles may be represented by a semantic segmentation map and semantic labels associated with the semantic segmentation map. During playout, if the viewport of the display device moves, peripheral tiles may enter into the viewport along with some of the viewport tiles that remain in the viewport. In that case, video data associated with peripheral tiles may be used to generate synthesized tiles of a quality that approximate the quality of the viewport tiles.
[0062] Here, the synthesized tiles may be generated by an image synthesis model which is trained to receive a semantic segmentation map and associated semantic labels at its input and to output a synthesized tile which has a quality that approximates the quality of the viewport tiles. This way, viewport tiles in high quality can be displayed next to synthesized tiles so that tiles visible in the viewport are presented at the same final quality. If sufficient decoder instances are available, the encoded video data of the one or more peripheral tiles received by the client device may be decoded into decoded data of one or more semantic segmentation maps, while the encoded video data of the one or more viewport tiles are decoded into decoded video data of the one or more viewport tiles and stored in the decoded playback buffer of the display device.
[0063] Further, one or more synthesized peripheral tiles of the first resolution may be generated by providing the decoded data of the one of the one or more semantic segmentation maps and associated semantic labels to the input of the image synthesis model and receiving video data associated with the one or more synthesized peripheral tiles at the output the image synthesis mode. The thus synthesized video data may be stored in the decoded playback buffer of the display device. This way, at least part of video data of the one or more synthesized peripheral tiles of the first resolution may be displayed in the viewport of the display device without any latency if at least part of the one or more synthesized peripheral tiles enter the viewport of the display device.
[0064] To allow a client device processing of tiled video data as shown in Fig. 2A-2C, a content preparation server may be configured to prepare different representation of a tile, e.g. a high-resolution tile representation and associated baseline tile representations and to provide the prepared tile representations to a server for distribution. The content preparation process may include the steps of:
[0065] - generating sets of tiles, e.g. Motion-Constrained Tile Sets (MCTSs) based on a high-resolution video source file;
[0066] - analyzing each tile of a tile set for suitability for synthesis based on a semantic segmentation map and a trained image synthesis model;
[0067] - preparing different representations for each tile (tile representations) and associated metadata that is needed for post-processing baseline tiles, wherein the different tile representations include at least a high-resolution tile representation and a baseline tile representation for image synthesis or a baseline tile representation for image upscaling.
[0068] - storing the tile representations, the associated metadata and an manifest file identifying the different tile representations at a server, e.g. a server which is part of a content delivery network (CDN).
[0069] Further, an image synthesis model may be prepared for a client device that is trained to receive a baseline tile representing a semantic segmentation map.
[0070] The content preparation server may be configured to analyze high-resolution video tiles based on semantic content and spatio-temporal texture complexity and prepare baseline tiles based on the analysis of the video tiles. For baseline tiles that are suitable for upscaling (baseline tiles of the first type) and for synthesis (baseline tiles of the second type) the server will generate metadata that is need for the post-processing at the client device.
[0071] Preparation of the baseline tiles based on the analysis of the video tiles may include partitioning omnidirectional video into independently decodable tiles (MCTSs). The server may categorize each tile as suitable for synthesis or suitable for upscaling based on video complexity metrics, which are used in video coding to determine encoding parameters. Alternatively, and / or in addition semantic segmentation maps of tiles may be analyzed using an algorithm may determine if a semantic segmentation map can be determined that has sufficient semantic labels to identify all relevant features in a tile that is needed for synthesis.
[0072] A semantic segmentation map for a tile may be determined using any suitable semantic segmentation algorithm, including for example well-known auto-encoder models which are based on encoder-decoder type deep neural networks. Non-limiting examples may include an ll-Net type of convolutional network as described in the article by Ronneberger et al, U-Net: Convolutional Networks for Biomedical Image Segmentation arXiv: 1505.04597 15 May 2015; one of the You-Only-Look-Once models, e.g. YOLOv5 as developed by Ultralytics or the Segment Anything model as described by Kirillov et al, Segment Anything arXiv:2304.026432023. In some embodiments, the semantic segmentation map may be a grayscale semantic segmentation map, wherein pixels with different grayscale values are associated with different semantic labels. In other embodiment, the semantic segmentation map may be a color semantic segmentation map, wherein pixels with different color values are associated with different semantic labels.
[0073] Hence, for tiles that are marked as suitable for synthesis a semantic segmentation map and associated semantic labels (or information for determining associated semantic labels) may be generated on the basis of a high-resolution tile and a known semantic segmentation model as referred above. Further, for tiles that are marked as suitable for upscaling a low-resolution version of a high resolution tile may be generated.
[0074] Additionally, the server may prepare metadata for post-processing baseline tiles. The metadata may be generated as supplementary enhancement information in the form of one or more SEI messages for each baseline tile. Baseline tiles that have been identified as suitable for synthesis may include metadata informing a client device that an image needs to be synthesized based on the semantic segmentation map and the semantic labels for color values or grayscale values in the color or grayscale semantic segmentation map. The SEI message may comprise the semantic labels and, optionally, other information about the synthesis, e.g. the type of image synthesis model that the client device can use for the synthesis process. In an embodiment, semantic labels may be ordered from lowest to highest grayscale value. This way, based on the color or grayscale values of the semantic segmentation map an image synthesis model can determine which part or parts of the semantic segmentation map belongs or belong to a certain semantic label. For example, if a part of the semantic segmentation map is represented by a grayscale value that is 2nd darkest when compared to the other values in the tile (sorted from darkest to lightest) then the corresponding label will also be the 2nd label). Baseline tiles that have been identified as suitable for upscaling may include metadata informing a client device that the tile requires upscaling as a post-processing step and, optionally, other information about the upscaling, e.g. the type of image upscaling model that the client device can use for the upscaling process. The tile representations, high-resolution tiles and the baseline tiles may be encoded using an encoding scheme and the encoded data may encapsulated and packaged based on a suitable media format and stored on a server for further distribution.
[0075] As described above, to accurately synthesize tiles of a high quality based on semantic segmented map and a trained image synthesis model, first it needs to be determined which tiles are suitable for synthesizing. Then, one or more image synthesis models may be established which can execute image synthesis at the client-side, e.g. at a view-port display device such as a HMD or the like.
[0076] In an embodiment, determining the suitability of a tile for image synthesis may be based on known video complexity metrics, which are used in video coding to determine encoding parameters. For example, in an embodiment VQEG Spatial Information (SI) and Temporal Information (Tl) metrics may be used as described in ITU Recommendation P.910 : Subjective video quality assessment methods for multimedia applications. In another embodiment, the Video Complexity Analysis (VGA) metric may be used as described in the article by Menon et al, VGA: video complexity analyser. In Proceedings of the 13th ACM Multimedia Systems Conference (MMSys '22). Association for Computing Machinery, New York, NY, USA, 259-264. https: / / doi.org / 10.1145 / 3524273.3532896.
[0077] The SI metric may be computed as the maximum standard deviation in pixel values after applying a filter, e.g. a Sobel filter, to frames in a video sequence. Similarly, the Tl metric may be computed as the maximum standard deviation of the difference between pixel values in consecutive video frames. The VGA metric may be derived using discrete cosine transform (DCT) based energy function that determines the block-wise texture of each video frame. The VGA metric has been shown to be very efficient at determining the optimal encoding parameters for a video. Hence, the SI and / or VGA metrics may be used for determining a spatial complexity threshold parameter. Tiles below the threshold parameter may be determined as candidates for synthesis. Further, the temporal complexity of the video tile sequence using the Tl metric and set a temporal complexity threshold parameter. Video tiles candidates (that were selected based on the spatial complexity threshold parameter) that are below the temporal complexity threshold parameter may be selected for synthesis. The thus selected video tiles are likely to produce limited artifacts when switching from synthesized video to high-quality video. This way, it may be determined which tiles are suitable for postprocessing using synthesis. Remaining video tiles may be post processed using an image upscaling model. An example of such upscaling model is the Efficient Sub-Pixel Convolutional Neural Network (ESPCN) as described in the article by Shi et al, Real-Time Single Image and Video SuperResolution Using an Efficient Sub-Pixel Convolutional Neural Network, arXiv:1609.05158 16 September 2016. It is noted that above example is just one example of many ways for determining if synthesis is possible. For example, dedicated models may be used to classify labels of a tile and determine whether their coverage is sufficient for high quality image synthesis.
[0078] As already discussed above, the content preparation server may generate semantic segmentation maps and associated semantic labels of tiles using on a semantic segmentation model. Such model may include generic labels which can be used for semantics-guided image synthesis, which are typically based on an encoder-decoder model. Examples of such models are described above. These models are based on trained neural networks which take an image (at a certain resolution) at the input and output a semantic segmentation map (at the same resolution) in which pixel of the segmentation map corresponds to a pixel at the same location in the original image and has a value which labels the pixel as a certain semantic category. The semantic labels are sufficiently finegrained so that the semantic segmentation map captures sufficient details in a tile (for example, not only labelling ‘elephant’ but also details ‘elephant ear’).
[0079] One or more image synthesis models for the client device may be trained to receive a decoded semantic segmentation map and semantic labels that are provided to the client device via messages, e.g. SEI messages, in the tiles so that it can accurately and precisely reconstruct an image through the inclusion of detailed context information. The image synthesis models can be specialized towards specific tasks or subject matters (for example one model might be specialized to work with faces, another may be fine-tuned to cityscapes, sky, grass, etc.). This process can be done on a segment-by-segment basis. This selection may be done in several ways: using a ‘master’ image synthesis model which is suitable for a wide range of topics, contexts and styles; - using a brute-force method of testing a plurality of different image synthesis models that are available on the server and determining which performs best at the time of encoding, for example using the metrics described above.
[0080] - using a pre-trained classifier to select a suitable image synthesis model (for example, a quality evaluation model which is trained to predict an evaluation score that evaluate the quality of the video was reconstructed when using an image synthesis model). At the time of encoding, quality evaluation model will be input with the video (or a segment of it) and may output a predicted performance score for each model. Based on the performance score the system can choose which image synthesis model is most likely to generate the best results.
[0081] To setup a model for image synthesis at the client device, both the content preparation server and the client device may be provided with a plurality of image synthesis models. The content preparation server may test the plurality of image synthesis models for a particular tile and selects an image synthesis model that is producing a synthesized tile that matches the original high-resolution tile the best. SEI messages in tiles sent to the client device may be used to select best performing image synthesis model that needs to be used by client during post-processing.
[0082] The image synthesis models can be finetuned using model finetuning, wherein the model at the client device and the model at the server are seeded with the same fixed set of first frame(s) of a video tile. The process of fine tuning is described in the article by Wang et al, Few-shot video-to-video synthesis. arXiv:1910.12713. Oct 28, 2019 which is hereby incorporated by reference into this disclosure. Few-shot Video-to-Video Synthesis is a technique which may be used to fine-tune a more generic video synthesis model using only a few video frames, that are given at inference time. This way, the image synthesis can be tuned to subjects that were not originally in the dataset that was used to train the model. This way, context information such as style, time of day, landscape, the appearance of a specific face or body and other visual qualities in a video tile may be extracted at run-time and used in subsequent to-be-generated frames.
[0083] Using few-shot learning, the first few frames of a target video may be used to seed the few-shot image synthesis model of the client device. In this way, more (detailed) context information may be provided to the labels and the masks upon which the image synthesis model can more accurately infer which image and style to produce. This process may be used both by the content preparation server and the client device. On the serverside, it may be used to seed an image synthesis model to synthesize the video frames for testing and determine if the tile is suitable for synthesis or not. On the client-side, it is used to seed the same image synthesis model in reconstructing the video frames. In this manner, the same model is used at the server side and at the client side to synthesize the video data.
[0084] Fig. 5 depicts a schematic of an exemplary system for streaming video data as described with reference to the embodiments in this disclosure. The system may include a content preparation system 502, a server system 504, and media playout device 506 comprising a client device 542, which may communicate via a network 536 (including the Internet) to the server system. In some embodiments the content preparation system may be part of the server system. In other embodiments, the content preparation system may be connected to the server system via e.g. the network 536 or another network, or may be directly communicatively coupled. The server system 504 may include a server processor 532 and one or more network interfaces 534 which are configured to send and receive data via network 536.
[0085] In some embodiment, functionalities of the server system and the content preparation system may be implemented in the form of a distributed system including a plurality of communicatively connected network devices, including but not limited to routers, bridges, proxy devices, switches, etc. In an embodiment, the server system may be part of a CDN.
[0086] The content preparation system 502 may include a media source 508, e.g. a media storage and / or one or more devices to generate media data, e.g. video and audio sources such as camera’s, microphones and computers for generating, e.g. synthesizing, computer-generated graphics and / or audio. Media data may include video data, e.g. in the form of sequences of video frames, which may be associated other data, e.g. audio data in the form of audio frames and / or metadata such as a recording time and / or information about temporal order of video data, in particular encoded video data.
[0087] The content preparation system may include a tile preparation module 510, which is configured to spatially partition video frames of a video source into different tile representations, including but not limited to high-resolution tile representations and tile presentations that require post-processing as described with reference to the embodiments in this disclosure, e.g. low-resolution tile representations for upscaling and tile representations associated with a semantic segmentation map for image synthesis
[0088] The content preparation system may further include an encoder system 514 comprising one or more encoder instances for encoding the media data associated with the tiles. The encoder system may produce one or more encoded media data tile streams, in short tile streams. In some embodiments, an individual stream may be referred to as an elementary stream representing a single, digitally coded component (e.g. video or audio) of a tile representation. The content preparation system may further include a packetizer 516 for converting elementary streams comprising encoded media data into a packetized stream, e.g. a packetized elementary stream (PES). The PES streams may be formatted, e.g. encapsulated, by an encapsulator 517 for transport so that that encoded media associated with different tiles can be transmitted to the server system using a suitable media streaming standard such as MPEG-DASH, HLS or CMAF and stored as one or more media files at a storage medium of the server system 504. The encapsulator may be configured to generate media files, which are formatted according to a predetermined data format for example CMAF fragment or DASH segments.
[0089] Media data associated with video tile which may be encoded in one or more elementary streams according to a certain bitrate or quality may form a tile representation or in short a representation. The encoder system 514 may be configured to encode media data of a media title in different ways using a video coding standard to produce different tile representations of a media title at various bitrates and various characteristics, such as pixel resolutions, frame rates, conformance to various coding standards, etc. These different tile representations include the baseline tile representations which need post-processing by the client device. These different representations may be used for adaptive bitrate streaming as known from streaming protocols such as DASH. In particular, the encoder system may encode the media data according to any suitable standardized coding scheme such as H.264 / AVC, HEVC or VVC. In an embodiment, the media data may be encoded using a layered coding scheme such as Scalable HEVC or Multilayer VVC.
[0090] Layered coding is particularly suitable when the same media stream needs to be available in different qualities, for example for adaptive bitrate streaming. Without layered coding, the source video stream needs to be encoded multiple times to obtain compressed streams with different qualities and bitrates. In contrast, layered coding provides the advantage of only encoding a single time, because streams with different qualities can be obtained based on the base layer and the one or more enhancement layers.
[0091] The encapsulator 516 may be configured may be configured to format packets of elementary systems into network abstraction layer (NAL) units. NAL units, which are defined as part of the H.264 / AVC and HEVC video coding standards, include Video Coding Layer (VCL) NAL units comprising video data payload and non-VCL NAL units, which may comprise metadata such as parameter sets (important header data that can apply to a large number of VCL NAL units) and supplemental enhancement information (timing information and other supplemental data that may enhance usability of the decoded video signal). Non- VCL NAL units may include sequence parameter sets (SPS), which apply to a series of consecutive coded video pictures called a coded video sequence and picture parameter sets (PPS), which apply to the decoding of one or more individual pictures within a coded video sequence. Non-VCL NAL units may further include Supplemental Enhancement Information (SEI) messages which may contain information for assisting the decoding process. A set of NAL units may define a so-called access unit which together may form a coded picture (a video frame), This way, the decoding of an access unit generally results in one decoded picture (a decoded video frame). A coded video sequence consists of a series of access units that are sequential in the NAL unit stream and use only one sequence parameter set. Each coded video sequence can be decoded independently of any other coded video sequence, given the necessary parameter set information, which may be conveyed "in-band" or "out-of-band". At the beginning of a coded video sequence is an instantaneous decoding refresh (I DR) access unit. An I DR access unit comprises an intra picture (l-frame) which is a coded picture that can be decoded without decoding any previous pictures in the NAL unit stream. The presence of an I DR access unit indicates that no subsequent picture in the stream will require reference to pictures prior to the intra picture it contains in order to be decoded. The encapsulator may use coded video sequences to produce short non-overlapping short video files, such as DASH segments and CMAF fragments, that are used by adaptive streaming protocols to provide adaptive streaming functionality.
[0092] The encapsulator 516 may be further configured to determine one or more manifest files 524 associated with the tile representations of a media title. An example of a manifest file is a media presentation descriptor (MPD) in case MPEG-DASH is used for streaming the media data. The manifest file may identify different tile representations 526- 530 that are associated with a media title using a suitable language such as the extensible markup language (XML). In some embodiments, media representations may be divided into so-called adaptation sets. An adaptation set may define media data associated with a common set of characteristics, including but not limited to e.g. codec, profile and level, resolution, number of views, file format for segments, etc. The manifest file may include data identifying such adaptation sets and further information associated with characteristics, such as bitrates, of specific representations of adaptation sets.
[0093] In some embodiments, the manifest file may comprise information about the availability of CMAF fragments or chunks of a representations or DASH segment or subsegments. For example, in an embodiment, the manifest file may include information indicating the wall-clock time at which a first fragment of one of the representations becomes available, as well as information indicating the durations of fragments within representations. This way, the client processor may determine when each fragment is available, based on the starting time as well as the durations of the fragments. The packetized and encapsulated media files, e.g. CMAF fragments and / or DASH segments, associated with a media title prepared by the content preparation system may be stored at the server system as one or more representations 526-530 and an associated manifest file 524. For example, the content preparation system may prepare and store CMAF tracks comprising switching sets of time-aligned fragments comprising media data of tile representations as described with reference to Fig. 2-4. Thus, when encoding the media data, the content preparation system may generate a track of CMAF fragments comprising video data associated a high-resolution tile and one or more tracks comprising CMAF fragments associated with one or more baseline tile representations respectively.
[0094] The content preparation system may further generate metadata associated with the tile representations which may include one or more SEI messages associated with the baseline tiles for signalling the client device that the video data of the baseline tiles need post-processing as e.g. described with reference to Fig. 3.
[0095] The server processor 532 may be configured to receive network requests from client devices, such as client device 542. The client device may comprise a client processor 548 configured to request media data and / or metadata, e.g. a manifest file, that is stored at the server system. Based on a manifest file 546 stored at the client device, the client device may request (via the client processor) tiles form the server system and store the tiles in a client buffer 544. The buffered (encoded) media data may be provided to the input of a decoder system 550 which may be configured to execute one or more decoder instances for decoding the encoded media data in to decoded media data, which may be rendered by a display device 552. As described with reference to Fig. 6 and 7, the decoder system may include or may be associated with a post-processor block 512 which is configured to postprocess decoded data of baseline tiles using a certain post-processing scheme based on an image synthesis model 564 or an image upscaling model 562. Decoded and post-processed tiles may be stored in a viewport buffer that is used by the media playout device to display tiles that are positioned in the viewport. Post-processed tiles which have a quality that is comparable to the high-quality tiles and which are not positioned in the viewport are directly available for display as soon these tiles are moved into the viewport.
[0096] The display device may be implemented as any type of display devices, including display devices such as a head-mounted device for rendering XR-type of media data (e.g. tiled 360 video data).
[0097] The client device may receive information about the decoding capabilities of video decoder and rendering capabilities of the media playout device. The functionality of the client device or portions of the functionality of the client device may be implemented in hardware, or a combination of hardware, software, and / or firmware, where requisite hardware may be provided to execute instructions for software or firmware. Further, in some embodiments, the client device, the decoder system and the display device may be implemented in different devices, e.g. the client device may be implemented in a media box and the decoder and the display maybe implemented in a display system. The server processor 532 and client processor 548 may be implemented to process requests based on the hypertext transfer protocol (HTTP), for example HTTP version 1.1, which allow transmission of encoded media data based on the chunked transfer encoding mode. This way, the server request processor may be configured to receive HTTP messages, such as HTTP GET or partial GET requests and sent media data in response to the requests back to the client device. The requests may specify a video file or a part of a specific part, e.g. a fragment or a chunk of a fragment, of one of the representations 526- 530, e.g., using a resource locator, such as an URL. In some examples, the requests may also specify a byte range for identifying a chunk in a fragment. Instead of HTTP other clientserver communication protocols may be used to handle request and response messages. For example, in an embodiment, a bi-directional communication channel between the client and the sever system may be realized based on a WebSocket protocol. In that case, a handshake request may be used to set up a WebSocket connection between the client and server. Request and response messages may be exchanged between the client and the server over the WebSocket connection. Other protocols that may be used to communicate between the server and the client include Long polling, WebRTC, SignaIR or the like.
[0098] Additionally, and / or alternatively, the request processors of the server system and the client device may be configured to deliver media data via a broadcast or multicast protocol. In that case, the server device may deliver encapsulated media data (e.g. segments or fragments or parts thereof) using a broadcast or multicast network transport protocol. For example, the server request processor may be configured to receive a multicast group join request from the client device. For example, the server system may advertise an Internet protocol (IP) address associated with a multicast group to client device 542, associated with particular a media title. The client device may subsequently submit a request to join the multicast group. This request may be propagated throughout network 536, e.g., routers defining a network, such that the routers are configured to media data destined for the IP address associated with the multicast group to the client device that has joined the multicast group.
[0099] The client device may request encoded video data from the server system based on the information in the manifest file and based on information about the available network bandwidth. Higher bitrate representations may yield higher quality video playback, while lower bitrate representations may provide sufficient quality video playback when available network bandwidth decreases. Accordingly, when available network bandwidth is relatively high, the client device may retrieve media data associated with (relatively) high bitrate representations, whereas when available network bandwidth is low, the client device may request media data associated with (relatively) low bitrate representations. This way, media data may be streamed over the network to the client device while adapting to changing network bandwidth availability of network.
[0100] The network interface 540 of the client device may receive media and buffer media, e.g. encapsulated and packetized media data such as CMAF fragments or chunks of a selected representation. The client processor 548 may be configured to decapsulate the media files into PES streams and depacketize the PES streams into encoded media data which may be to the decoder system. The decoded media data may be subsequently sent to display device 552.
[0101] The devices and modules depicted in Fig. 5 such as encoder, packetizer, encapsulator, server processor, client processor, etc. may be implemented as any of a variety of suitable processing circuitry, as applicable, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware or any combinations thereof. Alternatively, and / or additionally these devices and modules may comprise an integrated circuit, a microprocessor, and / or a wireless communication device.
[0102] Fig. 6 depicts a schematic of a decoder system for decoding video data associated with a video steaming data format according to an embodiment. The decoder system may be implemented based on the Video Decoding Interface for Immersive Media as described in ISO / IEC 23090-13:222 Information technology — Coded representation of immersive media — Part 13: Video decoding interface for immersive media. In particular, the decoding system may be implemented as a Video Decoding Engine (VDE) 602 which is configured to decode, synchronize and format media streams that comprise of one or more aggregated elementary streams. The VDE may be associated with post-processing block as shown in Fig. 7 which includes functionality for post-processing decoded tiles as described with reference to the embodiments in this disclosure. Media streams 622i-nand metadata streams 624i-massociated with the media streams may be retrieved by a client device based on a manifest file as described with reference to Fig. 5. Video data associated with the media streams may enter a so-called rendering pipeline via the input video decoding interface 606 (I VDI) of the VDE and are provided to the subsequent elements of the rendering pipeline.
[0103] Decoded video at the output video decoding interface 608 (OVDI) may be post-processed by the post-processing block and subsequently stored in a decoded playback buffer. As will be described hereunder in more detail, the decoder system may include a post-processing block that is placed at the OVDI so that synthesis and / or upscaling operations can be applied to at least part of the decoded streams before placing the streams in the decoded playback buffer.
[0104] An input formatting function 612 may be configured to extract encoded video data of independently decodable video tiles from the media streams into elementary streams 611i-m. These elementary streams are then fed to the video decoder instances 61 On that run inside the engine. The input formatting functions allows the VDE to execute either one or more merging operations or one or more extraction operations to allow the engine to have a different number of decoder instances running as compared to the number of media streams required by the application. For example, the VDE be configured to receive n media streams 622i-nand m metadata streams 624i.mwhile running / decoder instances 610i.j for processing / elementary streams wherein n^i, for example n > i. Similarly, an output function may be configured to execute either one or more merging operations or one or more extraction operations on metadata streams and / or decoded elementary streams so that the output of the VDE may comprise q decoded sequences 626i.qand p metadata streams 628i-p, wherein q^p.
[0105] Fig. 7 depicts a schematic of a post-processing block 700 associated with the VDE as shown in Fig. 6. The post-processing block may include a decoded playback buffer 702 for a playback device configured to process video data based on a video steaming data format according to an embodiment. The post-processing block 700 positioned after the OVDI of the VDE may include one or more post-processing units 704,706 for postprocessing decoded data originating from the OVDI. For example, a first post processing unit 704 may be configured to apply image synthesis operations based on decoded data and associated metadata provided to the input of the first processing unit. Post-processed data, e.g. a synthesized video tile of a certain resolution, at the output of the first post-processing unit may be stored in the decoded playback buffer. Similarly, a second post processing unit 704 may be configured to apply upscaling operations based on decoded data and associated metadata provided to the input of the second processing unit. Post-processed data, e.g. a upscaled video tile of a certain resolution, at the output of the second post-processing unit may be stored in the decoded playback buffer.
[0106] As shown in Fig. 6 the decoder system may be further configured to receive information 620 regarding applications and capabilities. This information may be formatted according to a VDI interface definition language syntax for controlling the interface of the VDE. This information can be used to inform a client device so that the appropriate buffers can be allocated after the OVDI and / or that post-processing units can be initialized. The VDI supplementary enhancement information (SEI) envelope specified in Annex C of the VDI standard is registered as a SEI payload in video codecs such as ISO / IEC 23090-3: Versatile video coding. The independent layer info SEI messages provide the spatial alignment of the different independent streams by expressing the relative position using matching boundary identifiers. This SEI message may be extracted by the VDE and used for the output formatting function to correctly place the decoded pictures from each layer in the final output picture. For example, such SEI message may include a boundary identifier flag specifying that the SEI message contains a boundary identifier for tiles that need to be placed in the decoded picture buffer. For example, it may include boundary identifiers to identify the north, east, south and west boundaries of a 2x2 mosaic of tiles: boundary_identifier_north, boundary_identifier_east, boundary_identifier_south and boundary_identifier_west. This enables parallel processing of video blocks as illustrated in the example of Fig. 8. This figure shows a video decoder engine 802 comprising decoder instances 808M for decoding encoded video data of media streams 806M that are received by the input interface 804 of the VDE. The encoded video data are decoded into decoded video data of video objects 810M, in this example four video tiles. Based on metadata (not shown), e.g. above-mentioned SEI messages, the tiles may be spatially formatted for an output formatting function 812 into a mosaic for rendering into the viewport of a display device.
[0107] Hence, the embodiments in this disclosure include postprocessing operations to be applied to decoded video data of tiles that lie outside the current viewport. The client device may decide on employing these operations based on certain information, e.g. the number of decoder instances for decoding encoded data that are available at a certain point in time and / or the computational resources that are available for postprocessing operations such as imagine synthesis of a tile that is outside the current viewport based on a trained image synthesis model and / or upscaling of a tile that is outside the current viewport based on trained image upscaling model as described with reference to the embodiments in this application.
[0108] If the client device does not have additional decoder instances available, the client device may choose to store encoded video frames of tiles that are positioned outside the viewport boundary. These encoded pictures can be stored as elementary streams generated by the input formatting function of the Video Decoding Interface that are later fed to video decoder instances if these tiles enter the viewport after a viewport movement.
[0109] If the client device has decoder instances available but does not have additional compute resources available to post-process the decoded video data associated with the tiles, the client device may choose to store decoded the decoded video data that are output by the Video Decoding Engine. Decoded video data and associated metadata may be stored in one or more output buffers. Part of the metadata may include instructions for the postprocessing block of the decoder system to post processing the decoded video data. If these tiles enter the viewport, the buffered decoded video data may be post-processed before displayed in the viewport.
[0110] If the client device has both decoder instances and computer resources to decoded the tiles and post-process the decoded tiles, the client device can choose to decode and post-process tiles that lie outside the viewport boundary and stored the post-processed tiles in a decoded playback buffer. If these tiles enter the viewport after viewport movements then these tiles can be directly displayed, while high quality tiles for the new viewport are fetched.
[0111] Information in the manifest file may be used to signal if the video data identified in the manifest filed can be post-processed using the schemes as described with reference to the embodiments in this disclosure. For example, in an embodiment, a binary flag postprocessing_fastswitching in the manifest file may be used to signal the client device whether the media supports post-processing of baseline quality tiles in the buffer to allow fast quality switching after viewport movements.
[0112] Additionally, metadata may be sent in-band along with the media stream to allow post-processing at the client side. A possible syntax of an in-band message may look gas follows:
[0113] This message may include the following parameters: lnpaint_flag 1 may indicate that the encoded video data of the tile is a semantic segmentation mask and that this tile should be synthesized. A zero may indicate that the tile should not be synthesized. semantic_mask_persistence_flag 1 may indicate that the semantic mask from the previous segment for this tile can be reused, 0 indicates the semantic mask for this tile needs to be fetched inpainting_approach may specify the image synthesis scheme to be used. This can be determined server-side using image quality metrics comparing the synthesized tile to the original tile greyscalejeve . provides a list of each greyscale level present in the semantic segmentation mask semantic_category: provides a semantic category corresponding to the greyscale level semantic_category_descriptor: provides additional information to describe the semantic category upscale_flag-. 1 indicates the baseline quality of this tile can be upscaled, 0 indicates it cannot. Note that for every tile in the segment either the upscale_flag is 1 or the inpaint_flag is 1. upscale_approach'. specifies the upscaling scheme to be used. This can be determined server-side using image quality metrics comparing the upscaled tile to the original tile.
[0114] The devices and systems described with reference to embodiments in this disclosure, such as the client device, the server system and the content preparation system are typically implemented as one or more communicatively connected data processing systems. Fig. 9 is a block diagram illustrating an exemplary data processing system that may be used in as described in this disclosure. Data processing system 900 may include at least one processor 902 coupled to memory elements 904 through a system bus 906. As such, the data processing system may store program code within memory elements 904. Further, processor 902 may execute the program code accessed from memory elements 904 via system bus 906. In one aspect, data processing system may be implemented as a computer that is suitable for storing and / or executing program code. It should be appreciated, however, that data processing system may be implemented in the form of any system including a processor and memory that is capable of performing the functions described within this specification.
[0115] Memory elements 904 may include one or more physical memory devices such as, for example, local memory 908 and one or more bulk storage devices 910. Local memory may refer to random access memory or other non-persistent memory device(s) generally used during actual execution of the program code. A bulk storage device may be implemented as a hard drive or other persistent data storage device. The data processing system 900 may also include one or more cache memories (not shown) that provide temporary storage of at least some program code in order to reduce the number of times program code must be retrieved from bulk storage device 910 during execution.
[0116] Input / output (I / O) devices depicted as input device 912 and output device 914 optionally can be coupled to the data processing system. Examples of input device may include, but are not limited to, for example, a keyboard, a pointing device such as a mouse, or the like. Examples of output device may include, but are not limited to, for example, a monitor or display, speakers, or the like. Input device and / or output device may be coupled to data processing system either directly or through intervening I / O controllers. A network adapter 916 may also be coupled to data processing system to enable it to become coupled to other systems, computer systems, remote network devices, and / or remote storage devices through intervening private or public networks. The network adapter may comprise a data receiver for receiving data that is transmitted by said systems, devices and / or networks to said data and a data transmitter for transmitting data to said systems, devices and / or networks. Modems, cable modems, and Ethernet cards are examples of different types of network adapter that may be used with data processing system.
[0117] As pictured in Fig. 9, memory elements 904 may store an application 918. It should be appreciated that data processing system may further execute an operating system (not shown) that can facilitate execution of the application. Application, being implemented in the form of executable program code, can be executed by data processing system, e.g., by processor 902. Responsive to executing application, data processing system may be configured to perform one or more operations to be described herein in further detail.
[0118] In one aspect, for example, data processing system may represent a client data processing system. In that case, application 918 may represent a client application that, when executed, configures data processing system to perform the various functions described herein with reference to a "client". Examples of a client can include, but are not limited to, a personal computer, a portable computer, a mobile phone, or the like. In other aspects, data processing system may represent a server data processing system. In that case, application 918 may represent a server application that, when executed, configures data processing system to perform the various functions described herein with reference to a "server".
[0119] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a codec hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.
[0120] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0121] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiment was chosen and described in order to best explain the principles of the invention and the practical application, and to enable others of ordinary skill in the art to understand the invention for various embodiments with various modifications as are suited to the particular use contemplated.
Claims
CLAIMS1. Method of processing video tiles comprising:Requesting, preferably by a client device, transmission of video tiles based on the position of a viewport of a display, the video tiles including one or more viewport tiles having a position within the viewport and one or more peripheral tiles, each peripheral tile having a position outside the viewport, the one or more viewport tiles comprising encoded video data of a first resolution and the one or more peripheral tiles comprising encoded data of one or more semantic segmentation maps, preferably grayscale semantic segmentation maps comprising encoded pixel values in greyscale; receiving, preferably by the client device, the video tiles and metadata associated with the video tiles, the metadata including semantic labels for labelling pixel regions in the one or more semantic segmentation maps, a pixel region preferably comprising pixels with the same greyscale pixel value, wherein the client device is associated with an image synthesis model, preferably an image synthesis model based on one or more artificial neutral networks, which is trained to determine a synthesized peripheral tile of the first resolution based on decoded data of a semantic segmentation map and semantic labels associated with the semantic segmentation map; decoding the encoded video data of the one or more viewport tiles into decoded video data of the one or more viewport tiles; and, storing the decoded video data in a decoded playback buffer of the display device for displaying at least part of the one or more viewport tiles in the viewport.
2. Method according to claims 1 further comprising: decoding the encoded video data of the one or more peripheral tiles into decoded data of one or more semantic segmentation maps, preferably the encoded video data being decoded while decoding the encoded video data of the one or more viewport tiles and storing the decoded video data of the one or more viewport tiles in the decoded playback buffer.
3. Method according to claim 2 wherein the encoded video data of the one or more peripheral tiles are stored until one or more decoder instances of a decoder system, preferably a video decoding engine, associated with the client device are available for the decoding of the encoded video data.
4. Method according to claims 2 or 3 further comprising:generating one or more synthesized peripheral tiles of the first resolution by providing the decoded data of the one of the one or more semantic segmentation maps and associated semantic labels to the input of the image synthesis model; and, storing video data associated with the one or more synthesized peripheral tiles in the decoded playback buffer of the display device.
5. Method according to claim 4 wherein the decoded data of the one or more peripheral tiles are stored until computational resources are available for processing the decoded data of the one of the one or more semantic segmentation maps and associated semantic labels by the image synthesis model.
6. Method according to any of claim 1-5 further comprising: displaying at least part of the one or more viewport tiles of the first resolution in the viewport of the display device if at least part of the one or more viewport tiles are positioned in the viewport of the display device.
7. Method according to any of claims 2-5 further comprising: displaying at least part of the one or more synthesized peripheral tiles of the first resolution in the viewport of the display device if at least part of the one or more synthesized peripheral tiles have entered the viewport of the display device.
8. Method according to any of claims 1-7 wherein the one or more peripheral tiles further comprise encoded data of one or more peripheral video tiles of a second resolution which is lower than the first resolution, the method further comprising: decoding the encoded video data of the one or more peripheral video tiles of the second resolution; generating one or more upscaled peripheral video tiles of the first resolution by upscaling decoded video data of the of the one or more peripheral video tiles of the second resolution; storing the decoded video data in the decoded playback buffer of the display device; and, displaying at least part of the one or more upscaled peripheral tiles of the first resolution if the one or more upscaled peripheral tiles have entered the viewport of the display device.
9. Method according to any of claims 1-8 wherein the tiled requested video tiles and metadata associated with the requested video tiles are transmitted to the clientdevice based on a video streaming standard, preferably an adaptive video streaming standards such as MPEG Dynamic Adaptive Streaming over HTTP (DASH) or MPEG Omnidirectional Media Format (OMAF) and / or wherein the video tiles are formatted as motion-constrained tile sets (MOTS).
10. Method according to any of claims 1-9 wherein the semantic labels or information to determine the semantic labels are transmitted using a one or more in-band messages, e.g. one or more DASH in band event messages, associated with the video tiles.
11. Method according to any of claims 1-10 wherein the video tiles are requested by the client device based on information in a manifest file, preferably a media presentation description (MPD), the manifest file including information for identifying video tiles, preferably one or more resource locators, e.g. URLs, or information for constructing such resource locators.
12. An apparatus for processing video tiles comprising: a computer readable storage medium having at least part of a program embodied therewith; and, a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, the processor is configured to perform executable operations comprising: requesting transmission of video tiles based on the position of a viewport of a display device, the video tiles including one or more viewport tiles having a position within the viewport and one or more peripheral tiles, each peripheral tile having a position outside the viewport, the one or more viewport tiles comprising encoded video data of a first resolution and the one or more peripheral tiles comprising encoded data of one or more semantic segmentation maps, preferably grayscale semantic segmentation maps comprising encoded pixel values in greyscale; receiving requested video tiles and metadata associated with the requested video tiles by the client device, the metadata including semantic labels for labelling pixel regions in the one or more semantic segmentation maps, a pixel region preferably comprising pixels with the same greyscale pixel value, wherein the client device is associated with an image synthesis model, preferably an image synthesis model based on one or more artificial neutral networks, which is trained to determine a synthesized peripheral tile of the first resolution based on decoded data of a semantic segmentation map and semantic labels associated with the semantic segmentation map;decoding the encoded video data of the one or more viewport tiles into decoded video data of the one or more viewport tiles; and, storing the decoded video data in a decoded playback buffer of a display device associated with the client device for displaying at least part of the one or more viewport tiles in the viewport.
13. Apparatus according to claim 12 wherein the processor is further configured to perform any of the method steps according to claims 2-11.
14. A video preparation module comprising: a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, the processor is configured to perform executable operations comprising: generating one or more sets of tiles, e.g. Motion-Constrained Tile Sets (MCTSs), based on video source file comprising video data of a first resolution; analyzing each of the one or more tile sets to determine if a tile is suitable for image synthesis using an image synthesis model, the image synthesis model being trained to determine a synthesized tile of the first resolution based on a semantic segmentation map, of the tile and semantic labels associated with the semantic segmentation map; determining at least a high-resolution tile representation comprising encoded and packetized video data of the first resolution and a baseline tile representation comprising encoded and packetized data representing a semantic segmentation map of a tile that is suitable for image synthesis and encoded and packetized data representing semantic labels associated with the semantic segmentation map; and, storing the high-resolution tile representation and the a baseline tile representation and an manifest file identifying the high resolution tile representation and the baseline representation representations at a server.
15. Computer program product comprising software code portions configured for, when run in the memory of a computer, executing the method steps according to any of 1-11.
Citation Information
Patent Citations
Methods and apparatus for sub-picture adaptive resolution change
US20220191543A1