Signaling of the size and position of the picture-in-picture region

By signaling the size and position of the target picture-in-picture area in the MPD, the method addresses the lack of effective integration of supplementary videos within main videos in DASH, enhancing the video coding process and picture-in-picture experience.

JP7712032B2Active Publication Date: 2025-07-23LEMON CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024083739
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-01
Filing Date
2024-05-23
Publication Date
2025-07-23
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

Existing video coding technologies lack effective methods to indicate the size and position of a target picture-in-picture area, which is necessary for seamless integration of supplementary videos within main videos, especially in Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) protocols.

Method used

The method involves signaling the size and position of the target picture-in-picture area through attributes such as @x, @y, @width, and @height in the MPD, and using elements like PicnPic to specify the position and size, enabling accurate overlay of supplementary videos on main videos.

Benefits of technology

This approach enhances the video coding process by allowing precise placement and integration of supplementary videos within main videos, improving the picture-in-picture experience in DASH streaming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007712032000014
    Figure 0007712032000014
  • Figure 0007712032000015
    Figure 0007712032000015
  • Figure 0007712032000016
    Figure 0007712032000016
Patent Text Reader

Abstract

To provide a method for processing media data and the like.SOLUTION: A method includes a step of determining whether the size and location of a target picture-in-picture region in a main video exist in a media data file for conversion between media data and the media data file. When the main video is displayed, a supplemental video appears so as to be overlaid on the target picture-in-picture region. The method further includes a step of performing conversion between the media data and the media data file on the basis of the determined size and location. Corresponding video coding devices and non-transitory computer-readable recording media are also disclosed.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application is a divisional application of Japanese Patent Application No. 2022 - 105487, filed on Jun. 30, 2022, for appropriately claiming the priority and benefits of U.S. Provisional Patent Application No. 63 / 216,975, filed on Jun. 30, 2021, and U.S. Provisional Patent Application No. 63 / 217,665, filed on Jul. 1, 2021. All of the above - mentioned patent applications are incorporated herein by reference in their entireties.

[0002] [Technical Field] The present disclosure generally relates to video streaming, and in particular, to the support of picture - in - picture services in the Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) protocol.

Background Art

[0003] Digital video occupies the largest bandwidth usage on the Internet and other digital communication networks. As the number of user devices capable of receiving and displaying video increases, the bandwidth demand for digital video utilization is expected to continue to grow.

Summary of the Invention

[0004] The disclosed aspects / embodiments provide techniques for indicating the size and position of a target picture - in - picture area. Thus, the video coding process is improved over conventional video coding techniques.

[0005] The first aspect relates to a method for processing media data. The method includes determining whether the size and position of a target picture-in-picture area in a main video exist in the media data file for conversion between the media data and the media data file, and when the main video is displayed, a supplementary video appears to be overlaid on the target picture-in-picture area, and executing the conversion between the media data and the media data file based on the determined size and position.

[0006] Optionally, in any of the above aspects, another implementation of the aspect provides that the size of the supplementary video is smaller than the size of the main video.

[0007] Optionally, in any of the above aspects, another implementation of the aspect provides that when the size and position of the target picture-in-picture area in the main video exist in the media data file, they are specified by x, y, width, and height, where x and y specify the upper left corner of the target picture-in-picture area in luma samples, and the width and height specify the width and height of the target picture-in-picture area in the luma samples.

[0008] Optionally, in any of the above aspects, another implementation of the aspect provides that when the size and position of the target picture-in-picture area in the main video exist in the media data file, one or more of them are indicated by elements or attributes of a preselected element.

[0009] Optionally, in any of the above aspects, another implementation of the aspect provides that when the size and position of the target picture-in-picture area in the main video exist in the media data file, one or more of them are indicated by elements or attributes of a main adaptation set of a preselected element.

[0010] Optionally, in any of the above aspects, other implementations of the aspect provide that when the size and the position exist in the media data file and the @dataUnitsReplaceable attribute is true, the size and the position specify the exact position of the target picture-in-picture area within the main video.

[0011] Optionally, in any of the above aspects, other implementations of the aspect provide that when the size and the position exist in the media data file and the @dataUnitsReplaceable attribute is false, the size and the position specify the proposed position of the target picture-in-picture area within the main video.

[0012] Optionally, in any of the above aspects, other implementations of the aspect provide that when the size and the position do not exist in the media data file and the @dataUnitsReplaceable attribute is false, the media data file does not supply a proposal for the position of the target picture-in-picture area within the main video.

[0013] Optionally, in any of the above aspects, other implementations of the aspect provide that the media data file includes a preselected element that includes a PicnPic element, and the PicnPic element includes all or none of the @x attribute, @y attribute, @width attribute, and @height attributes.

[0014] Optionally, in any of the above aspects, other implementations of that aspect provide that the media data file includes a preselected element that includes a PicnPic element, the PicnPic element includes @x and @y attributes, the @x attribute specifies the horizontal position of the encoded video sample at the upper left of the target picture-in-picture region, and the @y attribute specifies the vertical position of the encoded video sample at the upper left of the target picture-in-picture region.

[0015] Optionally, in any of the above aspects, other implementations of that aspect provide that the media data file includes a preselected element that includes a PicnPic element, the PicnPic element includes @width and @height attributes, the @width attribute specifies the width of the target picture-in-picture region, and the @height attribute specifies the height of the target picture-in-picture region.

[0016] Optionally, in any of the above aspects, other implementations of that aspect provide that the media data file includes a preselected element that includes a PicnPic element, and the PicnPic element includes a region element.

[0017] Optionally, in any of the above aspects, other implementations of that aspect provide that the media data file includes a preselected element that includes a PicnPic element, the PicnPic element includes a region element, the region element includes @x and @y attributes, the @x attribute specifies the horizontal position of the encoded video sample at the upper left of the target picture-in-picture region, and the @y attribute specifies the vertical position of the encoded video sample at the upper left of the target picture-in-picture region.

[0018] Optionally, in any of the above aspects, other implementations of that aspect provide that the media data file includes a preselected element that includes a PicnPic element, the PicnPic element includes a region element, the region element includes an @width attribute and an @height attribute, the @width attribute specifies the width of the target picture-in-picture region, and the @height attribute specifies the height of the target picture-in-picture region.

[0019] Optionally, in any of the above aspects, other implementations of that aspect provide that the size and the position are arranged in a media presentation description (MPD) file.

[0020] Optionally, in any of the above aspects, other implementations of that aspect provide that the size and the position are arranged in a dynamic adaptive streaming over hypertext transfer protocol (DASH) preselected element.

[0021] Optionally, in any of the above aspects, other implementations of that aspect provide that the conversion includes encoding the media data into a bitstream.

[0022] Optionally, in any of the above aspects, other implementations of that aspect provide that the conversion includes decoding the media data from a bitstream.

[0023] The second aspect relates to an apparatus for processing media data. The apparatus has a processor and a non-transitory memory having instructions. When executed by the processor, the instructions cause the processor to determine whether the size and position of a target picture-in-picture area in a main video exist in a media data file for conversion between the media data and the media data file, and when the main video is displayed, a supplementary video appears to be overlaid on the target picture-in-picture area; and to execute the conversion between the media data and the media data file based on the determined size and position.

[0024] The third aspect relates to a non-transitory computer-readable recording medium storing an MPD of a video generated by a method executed by a video processing apparatus. The method includes determining whether the size and position of a target picture-in-picture area in a main video exist in a media data file, and when the main video is displayed, a supplementary video appears to be overlaid on the target picture-in-picture area; and generating the MPD based on the determined size and position.

[0025] For clarity, any one of the above embodiments may be combined with one or more of the other above embodiments to create new embodiments within the scope of the present disclosure.

[0026] These and other features will be more clearly understood from the following detailed description in conjunction with the accompanying drawings and the claims.

[0027] For a more complete understanding of the present disclosure, reference is now made to the following brief description in connection with the accompanying drawings and detailed description. Here, like reference numerals represent like parts.

Brief Description of the Drawings

[0028]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Embodiments for Carrying Out the Invention

[0029] Exemplary implementations of one or more embodiments are given below, but it should first be understood that the disclosed system and / or method may be implemented using any number of techniques, whether currently known or existing. The present disclosure should in no way be limited to the exemplary implementations, drawings, and techniques described below, including the designs and implementation examples illustrated or described herein, but may be modified within the full scope of the appended claims and their equivalents.

[0030] The terms of H.266 are used in some descriptions not to limit the scope of the disclosed technology, but to facilitate its understanding. As such, the technology described herein is also applicable to other video codec protocols and designs.

[0031] Video coding standard specifications have mainly evolved through the development of well-known ITU-T (International Telecommunication Union - Telecommunication) and ISO (International Organization for Standardization) / IEC (International Electrotechnical Commission) standard specifications. ITU-T created H.261 and H.263, ISO / IEC created MPEG (Moving Picture Experts Group)-1 and MPEG-4 Visual, and the two organizations jointly created the H.262 / MPEG-2 Video, H.264 / MPEG-4 AVC (Advanced Video Coding), and H.265 / HEVC (High Efficiency Video Coding) standard specifications.

[0032] Since H.262, video coding standard specifications have been based on a hybrid video coding structure that combines transform coding with temporal prediction. To explore future video coding technologies beyond HEVC, the JVET (Joint Video Exploration Team) was jointly established in 2015 by VCEG (Video Coding Experts Group) and MPEG. Since then, many new methods have been adopted by the JVET and placed in a reference software called JEM (Joint Exploration Model).

[0033] In April 2018, the Joint Video Expert Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was formed to research the Versatile Video Coding (VVC) standard, which aims for a 50% bitrate reduction compared to HEVC.

[0034] The VVC standard (ITU-T H.266|ISO / IEC 23090-3) and the related Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IEC 23002-7) are designed for use in the broadest possible range of applications, including both traditional uses such as television broadcasting, video conferencing, or playback from storage media, and newer and more advanced uses such as adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple coded video bitstreams, multi-view video, scalable layer coding, and viewport-adaptive 360° immersive media.

[0035] The Essential Video Coding (EVC) standard (ISO / IEC 23094-1) is another recently developed video coding standard by MPEG.

[0036] Media streaming applications typically rely on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP) transport methods and usually depend on file formats such as the ISO base media file format (ISOBMFF). One such streaming system is DASH. To use video formats together with ISOBMFF and DASH, file format specifications specific to video formats such as the AVC file format and the HEVC file format in ISO / IEC 14496-15: “Information technology - Coding of audio-visual objects - Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format” are required for the encapsulation of video content within ISOBMFF tracks and within DASH representations and segments. Important information regarding the video bitstream, such as profile, tier, and level, as well as much other information, needs to be exposed as file format level metadata and / or video content media presentation description (MPD) within DASH representations and segments for the selection of appropriate media segments for content selection, for example, for both initialization at the start of a streaming session and stream adaptation during the streaming session.

[0037] Similarly, in order to use an image format together with ISOBMFF, file format specifications specific to image formats such as the AVC image file format and the HEVC image file format in ISO / IEC 23008-12: “Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 12: Image File Format” are required.

[0038] The VVC video file format, which is a file format for storing VVC video content based on ISOBMFF, is currently being developed by MPEG. The latest draft specification of the VVC video file format is included in ISO / IEC JTC 1 / SC 29 / WG 03 output document N0035, “Potential improvements on Carriage of VVC and EVC in ISOBMFF”, November 2020.

[0039] The VVC image file format, which is a file format for storing image content coded using VVC based on ISOBMFF, is currently being developed by MPEG. The latest draft specification of the VVC image file format is included in ISO / IEC JTC 1 / SC 29 / WG 03 output document N0038, “Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 12: Image File Format - Amendment 3: Support for VVC, EVC, slideshows and other improvements (CD stage)”, November 2020.

[0040] FIG. 1 is a schematic diagram illustrating the description of video or media data by an MPD 100 used in DASH. The MPD 100 describes a media stream with respect to a period 110, an adaptation set 120, a representation 130, and a segment 140.

[0041] In DASH, there may be multiple representations for the video and / or audio data of multimedia content, and different representations may correspond to different coding characteristics (e.g., different profiles or levels of video coding standards, different bitrates, different spatial resolutions, etc.). The manifest of such representations may be defined in the MPD data structure. A media presentation may correspond to a structured set of data that can be accessed by a DASH streaming client device (e.g., a smartphone, an electronic tablet, a laptop, etc.). The DASH streaming client device may request and download media data information to present the streaming service to the user of the client device. The media presentation may be described in the MPD data structure, which may include updates to the MPD.

[0042] A media presentation may include a sequence of one or more periods 110. Each period 110 may extend until the start of the next period 110, or, in the case of the last period 110, until the end of the media presentation. Each period 110 may include one or more representations 130 for the same media content. A representation 130 may be one of a number of alternative encoded versions of audio, video, timed text, or other such data. A representation 130 may differ by the coding type, e.g., by the bitrate, resolution, and / or codec of video data, and by the bitrate, language, and / or codec of audio data. The term representation 130 may be used to refer to a section of encoded audio or video data that corresponds to a particular period 110 of multimedia content and is encoded in a particular way.

[0043] The representations 130 of a particular period 110 may be assigned to a group indicated by an attribute in the MPD 100 that indicates the adaptation set 120 to which those representations 130 belong. Representations 130 within the same adaptation set 120 are generally considered to be alternatives to each other in that a client device can switch between them dynamically and seamlessly, for example, to perform bandwidth adaptation. For example, each representation 130 of the video data for a particular period may be assigned to the same adaptation set 120, whereby any one of the representations 130 can be selected for decoding to present the media data of the multimedia content for the corresponding period 110, such as video data or audio data. The media content within one period 110 may, in some examples, be represented by either one representation 130 from group 0, if present, or a combination of at most one representation from each non-zero group. The timing data of each representation 130 of a period 110 may be represented relative to the start time of that period 110.

[0044] A representation 130 may include one or more segments 140. Each representation 130 may include an initialization segment 140, or each segment 140 of a representation may be self-initializing. If present, the initialization segment 140 may include initialization information for accessing the representation 130. Generally, the initialization segment 140 does not include media data. A segment 140 may be uniquely referenced by an identifier such as a URL (Uniform Resource Locator), URN (Uniform Resource Name), or URI (Uniform Resource Identifier). The MPD 100 may supply an identifier for each segment 140. In some examples, the MPD 100 may also supply, in the form of a range attribute, the byte range corresponding to the data of a segment 140 within a file accessible by a URL, URN, or URI.

[0045] The different representations 130 may be selected for substantially simultaneous reading of different types of media data. For example, the client device may select an audio representation 130, a video representation 130, and a timed text representation 130 that reads the segment 140. In some examples, the client device may select a specific adaptation set 120 to perform bandwidth adaptation. That is, the client device may select an adaptation set 120 that includes a video representation 130, an adaptation set 120 that includes an audio representation 130, and / or an adaptation set 120 that includes a timed text representation 130. Alternatively, the client device may select an adaptation set 120 for a specific type of media (e.g., video), and then directly select a representation 130 for other types of media (e.g., audio and / or timed text).

[0046] A typical DASH streaming procedure is shown by the following steps.

[0047] First, the client obtains the MPD. In an embodiment, the client device receives the MPD from the content source after the client device requests the MPD.

[0048] Second, the client estimates the downlink bandwidth and selects the video representation 130 and the audio representation 130 according to the estimated downlink bandwidth, codec, decoding ability, display size, audio language setting, etc.

[0049] Third, unless the end of the media presentation is reached, the client requests the media segment 140 of the selected representation 130 and presents the streaming content to the user.

[0050] Fourth, the client continues to estimate the downlink bandwidth. When the bandwidth changes significantly (e.g., decreases), the client selects a different video representation 130 to adapt to the newly estimated bandwidth, and the procedure returns to the third step.

[0051] In VVC, a picture is partitioned (e.g., partitioned) into one or more tile rows and one or more tile columns. A tile is a sequence of Coding Tree Units (CTUs) that cover a rectangular region of the picture. The CTUs within a tile are scanned in raster scan order within that tile.

[0052] A slice consists of a number of complete tiles, or an integer number of consecutive complete CTUs within a tile of the picture.

[0053] Two modes of slices are supported, namely, the raster scan slice mode and the rectangular slice mode. In the raster scan slice mode, a slice contains a sequence of complete slices in the tile raster scan of the picture. In the rectangular slice mode, a slice contains either a number of complete tiles that collectively form a rectangular region of the picture, or a number of consecutive complete CTU rows of one tile that collectively form a rectangular region of the picture. The tiles within a rectangular slice are scanned in raster scan order within the rectangular region corresponding to that slice.

[0054] A sub-picture contains one or more slices that collectively cover a rectangular region of the picture.

[0055] Figure 2 is an example of a picture 200 partitioned into tiles 202, sub - pictures / slices 204, and CTUs 206. As shown, picture 200 is partitioned into 18 tiles 202 and 24 sub - pictures / slices 204. In VVC, each sub - picture consists of one or more complete rectangular slices that collectively cover the rectangular area of the picture, as shown in Figure 2. A sub - picture may be specified as either extractable (i.e., coded independently of other sub - pictures of the same picture and previous pictures in the decoding order) or non - extractable. Regardless of whether a sub - picture is extractable or not, the encoder can control whether in - loop filtering (de - blocking, Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF)) is applied individually for each sub - picture across sub - picture boundaries.

[0056] Functionally, sub - pictures are similar to Motion - Constrained Tile Sets (MCTS) in HEVC. Both enable independent coding and extraction of consecutive rectangular subsets of the coded picture for use cases such as view - port - dependent 360° video streaming optimization and Region of Interest (ROI) applications.

[0057] In the streaming of 360° video (also known as omnidirectional video), at any given point in time, only a portion of the entire omnidirectional video sphere (e.g., the current viewport) is rendered to the user. Meanwhile, the user can rotate their head at any time to change the viewing direction and, as a result, the current viewport. While it is desirable to render at least somewhat lower quality the representation of areas that are not covered by the current viewport available to the client but that are ready to be rendered to the user in case the user suddenly changes their viewing direction to somewhere on the sphere, high-quality representation of the omnidirectional video is only required for the current viewport being rendered to the user at any given point in time. Dividing the high-quality representation of the entire omnidirectional video into sub-pictures at an appropriate granularity enables an optimization as shown in FIG. 2, which includes 12 high-resolution sub-pictures on the left side and the remaining 12 sub-pictures of the low-resolution omnidirectional video on the right side.

[0058] FIG. 3 is an example of a viewport-dependent 360° video delivery method 300 based on sub-pictures. As shown in FIG. 3, only the high-resolution representation of the entire video uses sub-pictures. The low-resolution representation of the entire video does not use sub-pictures and can be coded at random access points (RAPs) that are less frequent than the high-resolution representation. The client receives the entire video at low resolution, while for the high-resolution video, the client only receives and decodes the sub-pictures that cover the current viewport.

[0059] There are several design differences between sub-pictures and MCTS. First, the sub-picture features in VVC allow the motion vectors of coding blocks to point outside the sub-picture even when the sub-picture is extractable. To do so, the sub-picture features apply sample padding at the sub-picture boundary which in this case is similar to the picture boundary. Second, additional changes have been introduced for the selection and derivation of motion vectors in the merge mode and in the decoder side motion vector refinement process of VVC. This enables higher coding efficiency compared to the non-standard motion constraints applied on the encoder side for MCTS. Third, rewriting the Slice Headers (SH) (and Picture Header (PH) Network Abstraction Layer (NAL) units if present) is not necessary when extracting one or more extractable sub-pictures from the picture succession to generate a sub-bitstream which is a conforming bitstream. SH rewriting is required for sub-bitstream extraction based on HEVC MCTS. Note that rewriting of the Sequence Parameter Sets (SPS) and Picture Parameter Sets (PPS) is required for both HEVC MCTS extraction and VVC sub-picture extraction. However, usually there are only a few parameter sets in the bitstream while each picture has at least one slice. Thus, SH rewriting can be a significant burden for the application system. Fourth, slices of different sub-pictures within a picture are allowed to have different NAL unit types. This is a feature often referred to as mixed NAL unit types or mixed sub-picture types within a picture which is discussed in more detail below.Fifthly, since VVC specifies the Hypothetical Reference Decoder (HRD) and level definitions for sub-picture sequences, the conformance of the sub-bitstream of each extractable sub-picture sequence can be ensured by the encoder.

[0060] In AVC and HEVC, all VCL NAL units within a picture need to have the same NAL unit type. Since VVC introduces the option of mixing specific different VCL NAL unit types and sub-pictures within a picture, it provides support for random access not only at the picture level but also at the sub-picture level. In VVC, the VCL NAL units within a sub-picture still need to have the same NAL unit type.

[0061] The ability of random access from Intra Random Access Pictures (IRAP) sub-pictures is beneficial for 360° video applications. In a viewport-dependent 360° video delivery method similar to that shown in Figure 3, the content of spatially adjacent viewports overlaps significantly, that is, only a part of the sub-picture within the viewport is replaced by a new sub-picture during the change of the viewport orientation, while most sub-pictures remain within the viewport. The newly introduced sub-picture sequence into the viewport must start from an IRAP slice, but when the remaining sub-pictures can perform inter prediction during the viewport change, a significant reduction in the overall transmission bitrate can be achieved.

[0062] The indication of whether a picture contains only one type of NAL unit or more than one type is provided in the PPS referred to by the picture (i.e., using the flag called pps_mixed_nalu_types_in_pic_flag). A picture may simultaneously have a sub-picture containing an IRAP slice and a sub-picture containing a trailing slice. Some other combinations of different NAL unit types within a picture are also permitted. This includes RASL (Random Access Skipped Leading) and RADL (Random Access Decodable Leading) which are leading picture slices of the NAL unit type, thereby enabling the sub-picture sequence to be merged with open group of pictures (GOP) and closed GOP structures extracted from different bitstreams into one bitstream.

[0063] Since the layout of sub-pictures within VVC is signaled in the SPS, it is constant within a Coded Layer Video Sequence (CLVS). Each sub-picture is signaled by the position of its top-left CTU and its width and height expressed in terms of the number of CTUs, ensuring that the sub-picture covers a rectangular region of the picture at the CTU granularity. The order in which sub-pictures are signaled in the SPS determines the index of each sub-picture within the picture.

[0064] To enable the extraction and merging of sub - picture sequences without rewriting SH or PH, the slice addressing scheme in VVC is based on the sub - picture ID and the sub - picture - specific slice index that associates a slice with a sub - picture. In the slice header (SH), the sub - picture identifier (ID) of the sub - picture containing the slice and the slice index at the sub - picture level are signaled. Note that the value of the sub - picture ID for a particular sub - picture may be different from the value of its sub - picture index. The mapping between the two is either signaled in the SPS or PPS (but never both) or is implicitly inferred. When present, the sub - picture ID mapping needs to be rewritten or added when rewriting the SPS and PPS during the sub - picture sub - bitstream extraction process. The sub - picture ID and the sub - picture - level slice index together indicate to the decoder the exact position of the first decoded CTU of the slice within the decoded picture buffer (DPB) slot of the decoded picture. After sub - bitstream extraction, the sub - picture ID of the sub - picture remains unchanged, but the sub - picture index may change. Even if the raster - scan CTU address of the first CTU of a slice within a sub - picture has changed compared to its value in the original bitstream, the invariant values of the sub - picture ID and the sub - picture - level slice index in each SH still accurately determine the position of each CTU within the decoded picture of the extracted sub - bitstream.

[0065] Figure 4 is an example of sub - picture extraction 400. In particular, Figure 4 illustrates the use of the sub - picture ID, sub - picture index, and sub - picture - level slice index to enable an example of sub - picture extraction 400. In the illustrated embodiment, sub - picture extraction 400 includes two sub - pictures and four slices.

[0066] Similar to sub-picture extraction 400, under the condition that different bitstreams are generated cooperatively (e.g., using mostly aligned SPS, PPS, and PH parameters such as individual sub-picture IDs, but if not, CTU size, chroma format, coding tools, etc.), the signaling of sub-pictures enables the merging of several sub-pictures from different bitstreams into one bitstream by only rewriting the SPS and PPS.

[0067] While sub-pictures and slices are signaled independently in the SPS and PPS respectively, there are inherent mutual constraints between the sub-picture layout and the slice layout to form a compliant bitstream. First, the existence of a sub-picture obligates rectangular slices and prohibits raster-scan slices. Second, the slices of a given sub-picture are consecutive NAL units in decoding order, that is, the sub-picture layout constrains the order of the coded slice NAL units within the bitstream.

[0068] The picture-in-picture service provides the ability to include a picture with a lower resolution within a picture with a higher resolution. Such a service can be useful for showing two videos to the user simultaneously. Thereby, the video with a higher resolution is regarded as the main video, and the video with a lower resolution is regarded as the supplementary video. Such a picture-in-picture service can be used to provide an accessibility service where the main video is supplemented by a signage video.

[0069] VVC sub - pictures can be used for picture - in - picture services by using the characteristics of both VVC sub - picture extraction and merge. For such services, the main video is encoded using a number of sub - pictures. One of the sub - pictures is the same size as the supplementary video, and it is located at the exact position where the supplementary video is intended to be overlaid on the main video. The supplementary video is independently coded to enable extraction. When the user chooses to view a version of the service that includes the supplementary video, the sub - picture corresponding to the picture - in - picture area of the main video is extracted from the main video bitstream, and the supplementary video bitstream is merged with the main video bitstream at that location.

[0070] Figure 5 is an example of picture - in - picture support 500 based on VVC sub - pictures. The part of the main video labeled Subpic ID 0 can be referred to herein as the target picture - in - picture area because that part of the main video will be replaced with the supplementary video having Subpic ID 0. That is, the supplementary video is embedded in or overlaid on the main video shown in Figure 5. In the example, the pictures of the main video and the supplementary video share the same video characteristics. In particular, the bit depth, sample aspect ratio, size, frame rate, color space and transfer characteristics, and chroma sample position of the main video and the supplementary video must be the same. The main video bitstream and the supplementary video bitstream do not need to use the same NAL unit type within each picture. However, the coding order of the pictures in the main video bitstream and the supplementary video bitstream needs to be the same.

[0071] Since subpicture merging is used here, the subpicture IDs used in the main video and the supplementary video cannot overlap. Even if the supplementary video bitstream consists of only one subpicture without any further tile or slice partitioning, the subpicture information (e.g., in particular, the subpicture ID and the subpicture ID length) needs to be signaled in order to enable the merging of the supplementary video bitstream with the main video bitstream. The subpicture ID length used to signal the length of the subpicture ID syntax element within a slice NAL unit of the supplementary video bitstream must be the same as the subpicture ID length used to signal the subpicture ID within a slice NAL unit of the main video bitstream. Furthermore, in order to simplify the merging of the supplementary video bitstream with the main video bitstream without the need to rewrite the PPS partitioning information, it may be beneficial to use only one slice and one tile for encoding the supplementary video within the corresponding area of the main video. The main video bitstream and the supplementary video bitstream need to signal the same coding tools in the SPS, PPS, and picture headers. This includes using the same maximum and minimum allowable sizes for block partitioning and the same value of the initial quantization parameter (the same value of the pps_init_qp_minus26 syntax element) as signaled in the PPS. The utilization of coding tools can be changed at the slice header level.

[0072] When both the main video bitstream and the supplementary video bitstream are available via a DASH-based delivery system, DASH Preselection may be used to signal the main video bitstream and the supplementary video bitstream that are intended to be merged and rendered together. DASH Preselection defines a subset of the media components of the MPD that are expected to be consumed together by a DASH player. Here, consumption may have decoding and rendering. The adaptation set containing the main media component for DASH Preselection is called the main Adaptation Set. Further, each DASH Preselection may include one or more partial adaptation sets. The partial adaptation sets may need to be processed in combination with the main adaptation set. The main adaptation set and the partial adaptation sets may be indicated by a preselection descriptor or by a preselection element.

[0073] Unfortunately, the following problems were identified when attempting to support picture-in-picture services with DASH. First, while it is possible to use DASH preselection for the picture-in-picture experience, there is a lack of such purpose indication.

[0074] Second, while it is possible to use VVC sub-pictures for the picture-in-picture experience (as described above, for example), there is also a possibility of using other codecs and methods where the coded video data unit representing the target picture-in-picture region within the main video cannot be replaced with the corresponding video data unit of the supplementary video. Therefore, it is necessary to indicate whether such replacement is possible.

[0075] Thirdly, when the above replacement is possible, the client needs to know which coded video data unit within each picture of the main video represents the target picture-in-picture area so that the replacement can be performed. Therefore, this information needs to be signaled.

[0076] Fourthly, for content selection (and optionally also for other purposes), it is useful to signal the position and size of the target picture-in-picture area within the main video.

[0077] In this specification, techniques for solving one or more of the above problems are disclosed. For example, the present disclosure provides techniques for indicating the size and position of the target picture-in-picture area. Thus, the video coding process is improved compared to conventional video coding techniques.

[0078] The following detailed embodiments should be regarded as examples for explaining the overview. These embodiments should not be construed in a narrow sense. Furthermore, these embodiments can be combined in any way.

[0079] In the following description, a video unit (also known as a video data unit) may be a sequence of pictures, a picture, a slice, a tile, a block, a sub-picture, a CTU / Coding Tree Block (CTB), a CTU / CTB row, one or more Virtual Pipeline Data Units (VPDUs), a sub-region within a picture / slice / tile / block. A parent video unit (also known as a parent video unit) represents a unit larger than a video unit. Usually, a parent video unit includes several video units. For example, when the video unit is a CTU, the parent video unit may be a slice, a CTU row, multiple CTUs, etc. In some embodiments, the video unit may be a sample / pixel. In some embodiments, the video unit may sometimes be called a video data unit.

[0080] [Example 1] 1) To solve the first problem, the indication is signaled in the MPD. The indication indicates that the preselection (also known as Preselection or DASH Preselection) is for providing a picture-in-picture experience. That is, the indicator within the preselection element indicates that the purpose of the preselection is to provide a picture-in-picture experience where the supplementary video appears overlaid on the target picture-in-picture area within the main video.

[0081] a. In one example, the indication is signaled by a specific value of the @tag attribute of the preselection (e.g., through the CommonAttributesElements). In an embodiment, the value of the @tag is "PicInPic".

[0082] b. In one example, the indication is signaled by a specific value of the @value attribute of the Role element of the preselection. In an embodiment, the specific value of the @value is "PicInPic".

[0083] [Example 2] 2) To solve the second problem, the indication is signaled in the MPD. The indication indicates whether the coded video data unit representing the target picture-in-picture area within the main video can be replaced by the corresponding video data unit of the supplementary video.

[0084] a. In one example, the indication is signaled by an attribute. In an embodiment, the attribute is specified or named by @dataUnitsReplaceable of the preselection.

[0085] b. In one example, the indication is specified as being optional.

[0086] i. In one example, the indication can exist only if the @tag attribute of the preselection indicates that the purpose of the preselection is to provide a picture-in-picture experience.

[0087] c. In the absence of an indication, it is unknown whether such a replacement is possible.

[0088] d. In one example, when @dataUnitsReplaceable is true, it is specified that the client may choose to replace the coded video data unit representing the target picture - in - picture area in the main video with the corresponding video data unit of the supplementary video before sending it to the video decoder. In this way, separate decoding of the main video and the supplementary video can be avoided.

[0089] e. In one example, for a specific picture in the main video, it is specified that the corresponding video data unit of the supplementary video is all the coded video data units within the decoding - time - synchronized samples in the supplementary video representation.

[0090] [Example 3] 3) To solve the third problem, a list of region IDs is signaled in the MPD. The list of region IDs indicates which coded video data units within each picture of the main video represent the target picture - in - picture area.

[0091] a. In one example, the list of region IDs is signaled as a pre - selected attribute. In an embodiment, the attribute is named @regionsIds.

[0092] b. In one example, the indication is specified to be optional.

[0093] i. In one example, the indication can exist only if a pre - selected @tag attribute indicates that the pre - selected purpose is to provide a picture - in - picture experience.

[0094] c. It is specified that the specific semantics of the region IDs need to be explicitly specified for a particular video codec.

[0095] i. In one example, for VVC, it is specified that the region ID is the sub-picture ID and the coded video data unit is the VCL NAL unit. The VCL NAL units representing the target picture-in-picture regions in the main video have the same sub-picture IDs as those in the corresponding VCL NAL units in the supplementary video. Usually, all the VCL NAL units of one picture in the supplementary video share the same sub-picture ID that is explicitly signaled. In this case, there is only one region ID in the list of region IDs.

[0096] A layer is a set of VCL NAL units all having a specific value of nuh_layer_id and the associated non-VCL NAL units. A Network Abstraction Layer (NAL) unit is a syntax structure that includes an indication of the type of the following data and the bytes containing that data in the form of a Raw Byte Sequence Payload (RBSP) in which emulation prevention bytes are interleaved as necessary. A Video Coding Layer (VCL) NAL unit is a general term for coded slice NAL units and a subset of NAL units having reserved values of nal_unit_type classified as VCL NAL units in the VVC standard.

[0097] ii. In one example, for VVC, when a client chooses to replace the coded video data unit (which is a VCL NAL unit) representing the target picture-in-picture region in the main video with the corresponding VCL NAL unit in the supplementary video before sending it to the video decoder, for each sub-picture ID, the VCL NAL unit in the main video is replaced with the corresponding VCL NAL unit having that sub-picture ID in the supplementary video without changing the order of the corresponding VCL NAL units.

[0098] [Example 4] 4) To solve the fourth problem, information regarding the position and size of the main video is signaled in the MPD. In an embodiment, the information on the position and size of the main video can be used when embedding / overlaying a supplementary video that is smaller in size than the main video.

[0099] a. In one example, the position and size information is signaled by signaling four values, namely x, y, width, and height. In this embodiment, x and y specify the position of the upper left corner of the region, and width and height specify the width and height of the region. The unit can be in lumasamples / pixels.

[0100] b. In one example, the position and size information is signaled by pre - selected attributes or elements.

[0101] c. In one example, the position and size information is signaled by attributes and elements of a pre - selected main adaptation set.

[0102] d. In one example, when @dataUnitsReplaceable is true and position and size information exists, it is specified that the position and size should accurately represent the target picture - in - picture region within the main video.

[0103] e. In one example, when @dataUnitsReplaceable is false and position and size information exists, it is specified that the position and size information indicates the desired region for embedding / overlaying the supplementary video (i.e., the client can choose to overlay the supplementary video on a different region of the main video).

[0104] f. In one example, when @dataUnitsReplaceable is false and position and size information does not exist, it is specified that no information or recommendation regarding where to overlay the supplementary video is suggested, and it is entirely up to the client to choose where to overlay the supplementary video.

[0105] [Example 5] 5) Instead, an element, for example, called the PicnPic element, is added to the preselected elements. This PicnPic element includes at least one or more of the following: a. Similar to the above, the @dataUnitsReplaceable attribute, b. Similar to the above, the @regionsIds attribute, c. The @x attribute that specifies the horizontal position of the top-left coded video pixel (sample) of the target picture-in-picture region within the main video. The unit is video pixel (sample). Either all four attributes @x, @y, @width, and @height exist or none of them exist, d. The @y attribute that specifies the vertical position of the top-left coded video pixel (sample) of the target picture-in-picture region within the main video. The unit is video pixel (sample), e. The @widhth attribute that specifies the width of the target picture-in-picture region within the main video. The unit is video pixel (sample), f. The @height attribute that specifies the height of the target picture-in-picture region within the main video. The unit is video pixel (sample).

[0106] [Example 6] 6) Instead, an element, for example, called the PicnPic element, is added to the preselected elements. This PicnPic element includes at least one or more of the following: a. Similar to the above, the @dataUnitsReplaceable attribute, b. Similar to the above, the @regionsIds attribute, c. An element, for example, called the Region element, that includes at least the following: i. The @x attribute that specifies the horizontal position of the top-left coded video pixel (sample) of the target picture-in-picture region within the main video. The unit is video pixel (sample), ii. The @y attribute that specifies the vertical position of the top - left encoded video pixel (sample) of the target picture - in - picture area within the main video. The unit is video pixel (sample). iii. The @widhth attribute that specifies the width of the target picture - in - picture area within the main video. The unit is video pixel (sample). iv. The @height attribute that specifies the height of the target picture - in - picture area within the main video. The unit is video pixel (sample).

[0107] The following are some exemplary embodiments corresponding to the examples described above. The embodiments are applicable to DASH. The most relevant parts that are added or changed in the following syntax and / or semantics are shown in italics. There may be some additional other changes that are essentially editorial and thus not highlighted.

[0108] The semantics of the pre - selection element are provided. Instead of the pre - selection descriptor, the pre - selection may also be defined by the pre - selection elements provided in Table 25. The selection of the pre - selection is based on the attributes and elements included in the pre - selection element.

Table 1

[0109] The Extensible Markup Language (XML) syntax is provided.

Table 2

[0110] Picture - in - picture based on pre - selection.

Table 3

[0111] The semantics of the pre-selection elements are provided in other embodiments. As an alternative to the pre-selection descriptor, the pre-selection may also be defined by the pre-selection elements provided in Table 25. The selection of the pre-selection is based on the attributes and elements included in the pre-selection elements.

Table 4

[0112] An XML syntax is provided.

Table 5

[0113] Picture-in-picture based on pre-selection.

Table 6

Table 7

Table 8

[0114] FIG. 6 is a block diagram showing an exemplary video processing system 600 in which various techniques disclosed in this specification can be implemented. Various implementations may include some or all of the components of video processing system 600. Video processing system 600 may include an input section 602 that receives video content. The video content may be received in a raw or uncompressed format, such as 8- or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input section 602 may correspond to a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wireless interfaces such as Ethernet®, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular networks.

[0115] Video processing system 600 may include a coding component 604 that can implement various coding or encoding methods described herein. The coding component 604 can reduce the average bit rate of the video from the input unit 602 to the output unit of the coding component 604 so as to generate a coded representation of the video. Coding techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the coding component 604 may be stored or transmitted via a connected communication as represented by the component 606. The stored or communicated bitstream (or coded) representation of the video received at the input unit 602 may be used by a component 608 that generates a displayable video that is sent to the pixel values or the display interface 610. The process of generating a video that a user can view from the bitstream representation is sometimes referred to as video decompression. Further, while certain video processing operations are called “coding” operations or tools, it will be understood that such coding tools or operations are used in an encoder and the corresponding decoding tools or operations that reverse the results of the coding are to be performed by a decoder.

[0116] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI (registered trademark)) or Displayport (registered trademark), etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI (Peripheral Component Interconnect), IDE (Integrated Drive Electronics) interface, etc. The techniques described herein may be embodied in various electronic devices such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0117] FIG. 7 is a block diagram of a video processing apparatus 700. The video processing apparatus 700 can be used to implement one or more of the methods described herein. The video processing apparatus 700 may be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The video processing apparatus 700 may include one or more processors 702, one or more memories 704, and video processing hardware 706 (also known as a video processing circuit). The processor 702 may be configured to implement one or more of the methods described herein. The memory (memories) 704 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 706 may be used to implement some of the techniques described herein in a hardware circuit. In some embodiments, the video processing hardware 706 may be located partially or entirely within the processor 702, e.g., a graphics processor.

[0118] FIG. 8 is a block diagram illustrating an example of a video coding system 800 that can utilize the techniques of the present disclosure. As shown in FIG. 8, the video coding system 800 may include a source device 810 and a destination device 820. The source device 810 generates encoded video data, which may be referred to as video encoding. The destination device 820 can decode the encoded video data generated by the source device 810 and may be referred to as a video decoding device.

[0119] The source device 810 may include a video source 812, a video encoder 814, and an input / output (I / O) interface 816.

[0120] The video source 812 may include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system that generates video data, or a combination of such sources. The video data may have one or more pictures. The video encoder 814 encodes the video data from the video source 812 to generate a bitstream. The bitstream may include a sequence of bits that forms a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded picture is a coded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 816 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly through the network 830 to the destination device 820 via the I / O interface 816. The encoded video data may also be stored in the storage medium / server 840 for access by the destination device 820.

[0121] The destination device 820 may include an I / O interface 826, a video decoder 824, and a display device 822.

[0122] The I / O interface 826 may include a receiver and / or a modem. The I / O interface 826 may obtain the encoded video data from the source device 810 or the storage medium / server 840. The video decoder 824 may decode the encoded video data. The display device 822 may display the decoded video data to the user. The display device 822 may be integrated with the destination device 820, or may be configured to interface with an external display device and be outside the destination device 820.

[0123] The video encoder 814 and the video decoder 824 may operate according to video compression standards such as the HEVC (High Efficiency Video Coding) standard, the VVC (Versatile Video Coding) standard, and other current and / or further standard specifications.

[0124] FIG. 9 is a block representing an example of a video encoder 900, which may be the video encoder 814 of the video coding system 800 shown in FIG. 8.

[0125] The video encoder 900 may be configured to execute any or all of the techniques of the present disclosure. In the example of FIG. 9, the video encoder 900 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video encoder 900. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.

[0126] The functional components of the video encoder 900 may include a partitioning unit 901, a prediction unit 902 that may include a mode selection unit 903, a motion estimation unit 904, a motion compensation unit 905, and an intra prediction unit 906, a residual generation unit 907, a transformation unit 908, a quantization unit 909, an inverse quantization unit 910, an inverse transformation unit 911, a reconstruction unit 912, a buffer 913, and an entropy encoding unit 914.

[0127] In other examples, the video encoder 900 may include more, fewer, or different functional components. In an example, the prediction unit 902 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode where at least one reference picture is the picture in which the current video block is located.

[0128] Furthermore, some components such as motion estimation unit 904 and motion compensation unit 905 may be highly integrated, but are shown separately in the example of FIG. 9 for the sake of explanation.

[0129] Partition unit 901 may partition a picture into one or more video blocks. Video encoder 814 and video decoder 824 of FIG. 8 may support various video block sizes.

[0130] Mode selection unit 903 may select, for example, one of an intra or inter coding mode based on an error result, and supply the resulting intra or inter-coded block to residual generation unit 907 that generates residual block data, and to reconstruction unit 912 that reconstructs the coded block for use as a reference picture. In some examples, mode selection unit 903 may select an Intra-Inter combined prediction (CIIP) mode where the prediction is based on an inter prediction signal and an intra prediction signal. Mode selection unit 903 may also select a resolution (e.g., sub-pixel or integer pixel accuracy) for the motion vector of the block in the case of inter prediction.

[0131] To perform inter prediction for the current video block, motion estimation unit 904 may generate motion information of the current video block by comparing one or more reference frames from buffer 913 with the current video block. Motion compensation unit 905 may determine a predicted video block of the current video block based on the motion information and the decoded samples of the picture from buffer 913 other than the picture associated with the current video block.

[0132] The motion estimation unit 904 and the motion compensation unit 905 may perform different operations for the current video block, for example, according to whether the current video block is an I slice, a P slice, or a B slice. An I slice (or an I frame) has the lowest compression ratio and does not require other video frames for decoding. A P slice (or a P frame) can use data from the previous frame for decompression and has a higher compression ratio than an I frame. A B slice (or a B frame) can use both the previous frame and the next frame for data reference to obtain the highest amount of data compression.

[0133] In some examples, the motion estimation unit 904 may perform one-directional prediction for the current video block, and the motion estimation unit 904 may search for a reference video block for the current video block from the reference pictures in list 0 or list 1. The motion estimation unit 904 may then generate a reference index indicating the reference picture in list 0 or list 1 that includes the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 904 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 905 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.

[0134] In another example, the motion estimation unit 904 may perform bidirectional prediction for the current video block. The motion estimation unit 904 may search for a reference video block for the current video block from the reference pictures in list 0, and may also search for another reference video block for the current video block from the reference pictures in list 1. The motion estimation unit 904 may then generate a reference index indicating the reference pictures in lists 0 and 1 that include the reference video blocks, and a motion vector indicating the spatial displacement between those reference video blocks and the current video block. The motion estimation unit 904 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 905 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.

[0135] In some examples, the motion estimation unit 904 may output a full set of motion information for decoder decoding processing.

[0136] In some examples, the motion estimation unit 904 may not output a full set of motion information of the current video. Rather, the motion estimation unit 904 may signal the motion information of the current video block by referring to the motion information of other video blocks. For example, the motion estimation unit 904 may determine that the motion information of the current video block is sufficiently similar to the motion information of adjacent video blocks.

[0137] In one example, the motion estimation unit 904 may indicate a value to the video decoder 824 indicating that the current video block has the same motion information as other video blocks in the syntax structure related to the current video block.

[0138] In another example, the motion estimation unit 904 may identify other video blocks and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 824 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0139] As described above, the video encoder 814 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 814 are advanced motion vector prediction (AMVP) and merge mode signaling.

[0140] The intra prediction unit 906 may perform intra prediction on the current video block. When the intra prediction unit 906 performs intra prediction on the current video block, the intra prediction unit 906 may generate predicted data for the current video block based on the decoded samples of other video blocks within the same picture. The predicted data for the current video block may include the predicted video block and various syntax elements.

[0141] The residual generation unit 907 may generate residual data for the current video block by subtracting the predicted video block of the current video block from the current video block (e.g., indicated by a negative sign). The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples within the current video block.

[0142] In another example, for example, in skip mode, there may be no residual data for the current video block, and the residual generation unit 907 may not perform a subtraction operation.

[0143] The transform processing unit 908 may generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0144] After the transform processing unit 908 generates the transform coefficient video block associated with the current video block, the quantization unit 909 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0145] The inverse quantization unit 910 and the inverse transform unit 911 may each apply inverse quantization and inverse transform to the transform coefficient video block to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 912 may generate a reconstructed video block associated with the current block for storage in the buffer 913 by adding the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 902.

[0146] After the reconstruction unit 912 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0147] The entropy encoding unit 914 may receive data from other functional components of the video encoder 900. When the entropy encoding unit 914 receives the data, the entropy encoding unit 914 may perform one or more entropy encoding operations to generate entropy encoded data and may output a bitstream including the entropy encoded data.

[0148] FIG. 10 is a block diagram representing an example of a video decoder 1000, which may be the video decoder 824 of the video coding system 800 represented in FIG. 8.

[0149] Video decoder 1000 may be configured to execute any or all of the techniques of the present disclosure. In the example of FIG. 10, video decoder 1000 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of video decoder 1000. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.

[0150] In the example of FIG. 10, video decoder 1000 includes an entropy decoding unit 1001, a motion compensation unit 1002, an intra prediction unit 1003, an inverse quantization unit 1004, an inverse transform unit 1005, a reconstruction unit 1006, and a buffer 1007. Video decoder 1000 may, in some examples, execute a decoding path generally inverse to the encoding path described with respect to video encoder 814 (FIG. 8).

[0151] Entropy decoding unit 1001 may extract the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., an encoded block of video data). Entropy decoding unit 1001 may decode the entropy-coded video data, and from the entropy-decoded video data, motion compensation unit 1002 may determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 1002 may determine such information, for example, by executing AMVP and merge mode.

[0152] Motion compensation unit 1002 may optionally perform interpolation based on an interpolation filter to generate a motion-compensated block. An identifier for the interpolation filter used at sub-pixel precision may be included in the syntax element.

[0153] The motion compensation unit 1002 may use the interpolation filter used by the video encoder 814 during the encoding of a video block to calculate interpolation values for sub-integer pixels of a reference block. The motion compensation unit 1002 may determine the interpolation filter used by the video encoder 814 according to the received syntax information, and use the interpolation filter to generate a prediction block.

[0154] The motion compensation unit 1002 may use some of the syntax information to determine the size of the blocks used to encode the frames and / or slices of the encoded video sequence, the partition information describing how each macroblock of the pictures of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information for decoding the encoded video sequence.

[0155] The intra prediction unit 1003 may use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 1004 inverse quantizes, i.e., dequantizes, the quantized video block coefficients supplied in the bitstream and decoded by the entropy decoding unit 1001. The inverse transform unit 1005 applies an inverse transform.

[0156] The reconstruction unit 1006 may add the corresponding prediction block generated by the motion compensation unit 1002 or the intra prediction unit 1003 to the residual block to form the decoded block. If desired, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 1007, which supplies reference blocks for subsequent motion compensation / intra prediction and further generates the decoded video for presentation on a display device.

[0157] FIG. 11 is a media data processing method 1100 according to an embodiment of the present disclosure. The method 1100 may be executed by a coding device (e.g., an encoder) having a processor and a memory. The method 1100 may be implemented when a picture-in-picture service needs to be instructed.

[0158] In block 1102, the coding device determines whether the size and position of the target picture-in-picture region in the main video exist in the media data file for conversion between the media data and the media data file, where the supplementary video appears to be overlaid on the target picture-in-picture region when the main video is displayed.

[0159] In block 1104, the coding device performs the conversion between the media data and the media data file based on the determined size and position. When implemented in an encoder, the conversion includes receiving a media file (e.g., a video unit) and encoding the media file into a bitstream. When implemented in a decoder, the conversion includes receiving a bitstream including the media file and decoding the bitstream to obtain the media file.

[0160] In an embodiment, the size of the supplementary video is smaller than the size of the main video.

[0161] In an embodiment, when the size and position of the target picture-in-picture area in the main video exist in the media data file, they are specified by x, y, width, and height. x and y specify the upper left corner of the target picture-in-picture area in luma samples, and the width and height specify the width and height of the target picture-in-picture area in luma samples.

[0162] In an embodiment, one or more of the size and position of the target picture-in-picture area of the main video are indicated by an element or attribute of a preselected element when the size and position exist in the media data file.

[0163] In an embodiment, one or more of the size and position of the target picture-in-picture area of the main video are indicated by an element or attribute of the main adaptation set of the preselected element when the size and position exist in the media data file.

[0164] In an embodiment, when the size and position exist in the media data file and the @dataUnitsReplaceable attribute is true, the size and position specify the exact position of the target picture-in-picture area in the main video.

[0165] In an embodiment, when the size and position exist in the media data file and the @dataUnitsReplaceable attribute is false, the size and position specify the proposed position of the target picture-in-picture area in the main video.

[0166] In an embodiment, when the size and position do not exist in the media data file and the @dataUnitsReplaceable attribute is false, the media data file does not supply a proposal for the position of the target picture-in-picture area in the main video.

[0167] In an embodiment, the media data file includes a preselected element that includes a PicnPic element, and the PicnPic element either includes all of the @x attribute, @y attribute, @width attribute, and @height attribute or none of them.

[0168] In an embodiment, the media data file includes a preselected element that includes a PicnPic element, the PicnPic element includes the @x attribute and the @y attribute, the @x attribute specifies the horizontal position of the encoded video sample at the upper left of the target picture-in-picture area, and the @y attribute specifies the vertical position of the encoded video sample at the upper left of the target picture-in-picture area.

[0169] In an embodiment, the media data file includes a preselected element that includes a PicnPic element, the PicnPic element includes the @width attribute and the @height attribute, the @width attribute specifies the width of the target picture-in-picture area, and the @height attribute specifies the height of the target picture-in-picture area.

[0170] In an embodiment, the media data file includes a preselected element that includes a PicnPic element, and the PicnPic element includes a region element.

[0171] In an embodiment, the media data file includes a preselected element that includes a PicnPic element, the PicnPic element includes a region element, the region element includes the @x attribute and the @y attribute, the @x attribute specifies the horizontal position of the encoded video sample at the upper left of the target picture-in-picture area, and the @y attribute specifies the vertical position of the encoded video sample at the upper left of the target picture-in-picture area.

[0172] In an embodiment, the media data file includes a pre-selection element that includes a PicnPic element, the PicnPic element includes a region element, the region element includes an @width attribute and an @height attribute, the @width attribute specifies the width of a target picture-in-picture region, and the @height attribute specifies the height of the target picture-in-picture region.

[0173] In an embodiment, the size and position are arranged in a Media Presentation Description (MPD) file. In an embodiment, the size and position are arranged in a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) pre-selection element.

[0174] In an embodiment, the conversion includes encoding media data into a bitstream. In an embodiment, the conversion includes decoding media data from a bitstream.

[0175] In an embodiment, method 1100 may utilize or incorporate one or more of the features or processes of other methods disclosed herein.

[0176] A list of preferred solutions is given below by some embodiments.

[0177] The following solutions illustrate exemplary embodiments of the techniques described in this disclosure (e.g., Example 1).

[0178] Solution 1. A method for processing video data, comprising: performing a conversion between video data and a descriptor of the video data, wherein the descriptor conforms to format rules, and the format rules specify that the descriptor includes a syntax element indicating the use of picture-in-picture in a pre-selection syntax structure of the descriptor. A method.

[0179] Solution 2. The method according to Solution 1, comprising: The syntax element is a tag attribute of the preselected syntax structure. Method.

[0180] Solution 3. The method according to Solution 1, wherein the syntax element is a role attribute of the preselected syntax structure. Method.

[0181] Solution 4. A method for processing video data, comprising performing a conversion between video data and a descriptor of the video data, wherein the descriptor conforms to a format rule, the format rule selectively includes a syntax element indicating whether a video data unit of a main video in the video data corresponding to a picture-in-picture area can be replaced by a video data unit of a supplementary video in the video data. Method.

[0182] Solution 5. The method according to Solution 4, wherein the syntax element is an attribute field in the descriptor. Method.

[0183] Solution 6. The method according to Solution 4, wherein the syntax element is selectively included based on a value of a tag attribute in the descriptor. Method.

[0184] Solution 7. A method for processing video data, comprising performing a conversion between video data and a descriptor of the video data, wherein the descriptor conforms to a format rule, the format rule specifies that the descriptor includes a list of region identifiers indicating video data units in a picture of a main video in the video data corresponding to a target picture-in-picture area. Method.

[0185] Solution 8. The method according to Solution 7, wherein the list is included as an attribute of a preselected syntax structure within the descriptor, method.

[0186] Solution 9. The method according to Solution 7 or 8, wherein the region identifier corresponds to a syntax field used to indicate the video data unit according to a coding scheme used for coding the main video, method.

[0187] Solution 10. A method for processing video data, comprising performing a conversion between video data and a descriptor of the video data, wherein the descriptor follows format rules, the format rules specifying that the descriptor includes one or more fields indicating the position and / or size information of a region within the main video used for overlaying or embedding supplementary video, method.

[0188] Solution 11. The method according to Solution 10, wherein the position and size information has four values including the position coordinates, height, and width of the region, method.

[0189] Solution 12. The method according to Solution 10 or 11, wherein the one or more fields have attributes or elements of a preselected syntax structure, method.

[0190] Solution 13. The method according to any one of Solutions 10 to 12, wherein whether the region is an exact replaceable region or a preferred replaceable region is determined based on other syntax elements, method.

[0191] Solution 14. A method according to any one of Solutions 1 to 13, wherein the descriptor is a Media Presentation Description (MPD). Method.

[0192] Solution 15. A method according to any one of Solutions 1 to 14, wherein the format rule specifies that a specific syntax element is included in the descriptor, and the specific syntax element includes picture-in-picture information. Method.

[0193] Solution 16. A method according to any one of Solutions 1 to 15, wherein the conversion includes generating a bitstream from the video. Method.

[0194] Solution 17. A method according to any one of Solutions 1 to 15, wherein the conversion includes generating video from the bitstream. Method.

[0195] Solution 18. A video decoding apparatus having a processor configured to implement a method according to one or more of Solutions 1 to 17.

[0196] Solution 19. A video encoding apparatus having a processor configured to execute a method according to one or more of Solutions 1 to 17.

[0197] Solution 20. A computer program product storing computer code that, when executed by a processor, causes the processor to implement a method according to any one of Solutions 1 to 17.

[0198] Solution 21. A method for video processing, generating a bitstream according to any one or more of Solutions 1 to 17; storing the bitstream in a computer-readable medium; and A method comprising the steps of:

[0199] Solution 22. A method, apparatus, or system described herein.

[0200] The following documents are incorporated by reference in their entirety: [1] ITU-T and ISO / IEC, “High efficiency video coding”, Rec. ITU-T H.265 | ISO / IEC 23008-2 (in force edition) [2] J. Chen, E. Alshina, G. J. Sullivan, J.-R. Ohm, J. Boyce, “Algorithm description of Joint Exploration Test Model 7 (JEM7),” JVET-G1001, Aug. 2017 [3] Rec. ITU-T H.266 | ISO / IEC 23090-3, “Versatile Video Coding”, 2020 [4] B. Bross, J. Chen, S. Liu, Y.-K. Wang (editors), “Versatile Video Coding (Draft 10),” JVET-S2001 [5] Rec. ITU-T Rec. H.274 | ISO / IEC 23002-7, “Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams”, 2020 [6] J. Boyce, V. Drugeon, G. Sullivan, Y.-K. Wang (editors), “Versatile supplemental enhancement information messages for coded video bitstreams (Draft 5),” JVET-S2007 [7] ISO / IEC 14496-12: "Information technology - Coding of audio-visual objects - Part 12: ISO base media file format" [8] ISO / IEC 23009-1: "Information technology - Dynamic adaptive streaming over HTTP (DASH) - Part 1: Media presentation description and segment formats"(The 4th edition of the DASH standard specification is available in MPEG input document m52458.) [9] ISO / IEC 14496-15: "Information technology - Coding of audio-visual objects - Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format"

[10] ISO / IEC 23008-12: "Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 12: Image File Format"

[11] ISO / IEC JTC 1 / SC 29 / WG 03 output document N0035, "Potential improvements on Carriage of VVC and EVC in ISOBMFF", Nov. 2020

[12] ISO / IEC JTC 1 / SC 29 / WG 03 output document N0038, "Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 12: Image File Format - Amendment 3: Support for VVC, EVC, slideshows and other improvements (CD stage)", Nov. 2020.

[0201] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware that includes the structures disclosed in this specification and their structural equivalents, or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing apparatus” includes, by way of example, all apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., an electrical, optical, or electromagnetic signal generated by a machine, that is generated to encode information for transmission to an appropriate receiver device.

[0202] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code), in portions of files that hold other programs or data (e.g., one or more scripts stored in a markup language document). A computer program can be deployed to be executed on one computer, or located at one location, or distributed over multiple locations and executed on multiple computers interconnected by a communication network.

[0203] The processes and logic flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data to generate output. The processes and logic flows can also be executed by dedicated logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).

[0204] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes one or more mass storage devices for storing data, such as magnetic, magneto - optical disks, or optical disks, or is operatively coupled for receiving data from or transferring data to or both from one or more such mass storage devices. However, a computer need not have such devices. Computer - readable media suitable for storing computer program instructions and data include, by way of example, semiconductor memory devices, such as erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto - optical disks; and all forms of non - volatile memory, media, and memory devices including compact disk read only memory (CD ROM) and digital versatile disk read only memory (DVD - ROM) disks. The processor and the memory can be augmented or incorporated in a dedicated logic circuit.

[0205] This specification includes many details, but they should be construed as descriptions of features that may be specific to particular embodiments of a particular technology rather than as limitations on the scope of any subject or of what may be claimed. The particular features described in this specification in connection with separate embodiments may be implemented in combination with a single embodiment. Conversely, the various features described in connection with a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Further, features may be described above as acting in certain combinations and even initially claimed as such, but one or more features from a claimed combination may in some cases be excised from that combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0206] Similarly, operations are represented in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in that particular order or in a sequential order to achieve the desired result, or that all of the operations shown be performed. Further, the separation of various system components in the embodiments described in this specification should not be understood as requiring such separation in all embodiments.

[0207] Only a few implementations and examples have been described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.

[0208] This specification includes many details, but they should be construed as descriptions of features that may be specific to particular embodiments of a particular technology rather than as limitations on the scope of any subject or of what may be claimed. The particular features described herein in connection with separate embodiments may be implemented in combination with a single embodiment. Conversely, the various features described in connection with a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Further, features may have been previously described as operating in a particular combination and even initially claimed as such, but in some cases, one or more features from the claimed combination may be removable from that combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0209] Similarly, operations are represented in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in that particular order or in a sequential order to achieve the desired result, or that all of the operations shown be performed. Further, the separation of various system components in the embodiments described herein should not be understood as requiring such separation in all embodiments.

[0210] Only a few implementations and examples have been described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Explanation of Reference Numerals

[0211] 100 MPD 200 pictures 202 tiles 204 sub-pictures / slices 206 partitions in CTU 600 video processing system 602 input unit 604 coding component 700 Video Processing Device 702 Processor 704 Memory 706 Video Processing Hardware 800 Video Coding System 810 Source Device 812 Video Source 814,900 Video Encoder 816,826 I / O Interface 820 Destination Device 822 Display Device 824,1000 Video Decoder 830 Network 840 Storage Medium / Server 901 Partition Unit 902 Prediction Unit 903 Mode Selection Unit 904 Motion Estimation Unit 905 Motion Compensation Unit 906 Intra Prediction Unit 907 Residual Generation Unit 908 Transformation Unit 909 Quantization Unit 910 Inverse Quantization Unit 911 Inverse Transformation Unit 912 Reconstruction Unit 913,1007 Buffer 914 Entropy Coding Unit 1001 Entropy Decoding Unit 1002 Motion Compensation Unit 1003 Intra Prediction Unit 1004 Inverse Quantization Unit 1005 Inverse Transformation Unit 1006 Reconstruction Unit

Claims

1. A method for processing media data, comprising: determining that a preselected element includes a first indicator for conversion between the media data and a media data file, the first indicator indicating that the purpose of the preselected element is to provide a picture-in-picture experience in which supplementary video of the media data appears overlaid in a target picture-in-picture area within the main video of the media data; performing the conversion between the media data and the media data file based on the first indicator; and a method having the above steps.

2. The preselected element further includes a second indicator, the second indicator indicating that a video data unit representing the target picture-in-picture area of the main video can be replaced by a corresponding video data unit of the supplementary video; The method according to claim 1.

3. The second indicator is represented by dataUnitsReplaceable; The method according to claim 2.

4. When the second indicator is equal to 1, the client is allowed to select to replace the video data unit representing the target picture-in-picture area in the main video with the corresponding video data unit of the supplementary video; The method according to claim 2.

5. When the second indicator is equal to 1, the separate decoding of the main video and the supplementary video is avoided by allowing the client to select to replace the video data unit representing the target picture-in-picture area in the main video with the corresponding video data unit of the supplementary video; The method according to claim 2.

6. The corresponding video data unit of the supplementary video is all video data units within a decoding time synchronization sample in the representation of the supplementary video; The method according to claim 2.

7. The video data units within the decoding time synchronization sample correspond to a specific picture in the main video; The method according to claim 6.

8. The pre-selection element further includes a third indicator, and the third indicator indicates which video data unit of the main video represents the target picture-in-picture area. The method according to claim 1.

9. The third indicator has a list of region identifiers (IDs) represented by region IDs. The method according to claim 8.

10. The region ID designates an identifier (ID) of the video data unit representing the target picture-in-picture area of the main video. The method according to claim 9.

11. The region ID designates an identifier (ID) of the video data unit representing the target picture-in-picture area of the main video as a list separated by blanks. The method according to claim 9.

12. The region ID is a sub-picture ID, and the video data unit representing the target picture-in-picture area is a video coding layer (VCL) network abstraction layer (NAL) unit. The VCL NAL unit representing the target picture-in-picture area includes the sub-picture ID, and the sub-picture ID is the same as the sub-picture ID in the corresponding VCL NAL unit of the supplementary video. The method according to claim 9.

13. The pre-selection element is arranged in a media presentation description (MPD) file. The method according to claim 1.

14. The pre-selection element is arranged in a dynamic adaptive streaming (DASH) pre-selection element on the hypertext transfer protocol. The method according to claim 1.

15. The conversion includes encoding the media data into the media data file. The method according to any one of claims 1 to 14.

16. The conversion includes decoding the media data from the media data file. The method according to any one of claims 1 to 14.

17. An apparatus for processing media data, comprising a processor and a non-transitory memory having instructions, the instructions, when executed by the processor, cause the processor to A step of determining that a preselection element includes a first indicator for the conversion between the media data and the media data file, wherein the first indicator indicates that the purpose of the preselection element is to provide a picture-in-picture experience in which a supplementary video of the media data appears to be overlaid on a target picture-in-picture area within the main video of the media data, the step of determining; A step of performing the conversion between the media data and the media data file based on the first indicator; Causing; An apparatus. [

18. ] A method of storing a media data file, comprising: A step of determining that a preselection element includes a first indicator, wherein the first indicator indicates that the purpose of the preselection element is to provide a picture-in-picture experience in which a supplementary video of the media data appears to be overlaid on a target picture-in-picture area within the main video of the media data, the step of determining; A step of generating the media data file based on the first indicator; A step of storing the media data file in a non-transitory computer-readable recording medium; A method having the above.

Citation Information

Patent Citations

  • The process of transmitting data for multiplexing video components.

    JP2013535886A

  • Region-wise packing, content coverage, and signaling frame packing for media content

    JP2020526982A

  • Systems and methods for signaling sub-picture composition information for virtual reality applications

    WO2019194241A1