Method, apparatus, and medium for video processing
By introducing new track reference types in ISOBMFF-based media files to manage track pairs and signal picture-in-picture areas, the challenge of supporting picture-in-picture services is addressed, enabling seamless integration and improved user experience.
Patent Information
- Application Number
- JP2024519069
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-27
- Filing Date
- 2022-09-26
- Publication Date
- 2025-10-07
- Estimated Expiration
- 2042-09-26
AI Technical Summary
Existing media streaming technologies, particularly those based on ISO Base Media File Format (ISOBMFF), lack a mechanism to support picture-in-picture services effectively, as they do not provide a way to indicate track pairs for providing such services, nor do they allow for swapping or signaling of video data units representing target picture-in-picture areas.
Introduce new track reference types to indicate track pairs for providing picture-in-picture services, allowing for swapping or non-swapping of video data units, and signal the location and size of the target picture-in-picture area within ISOBMFF-based media files.
Enables effective support for picture-in-picture services by clearly identifying and managing track pairs, facilitating seamless integration and swapping of video data units, thereby enhancing user experience and content selection capabilities.
Smart Images

Figure 0007751092000001 
Figure 0007751092000002 
Figure 0007751092000003
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 63 / 248,832, filed September 27, 2021, the disclosure of which is incorporated herein in its entirety.
[0002] FIELD Embodiments of the present disclosure relate generally to video processing technology, and more particularly to file format design for picture-in-picture support. [Background technology]
[0003] Media streaming applications are typically based on IP, TCP, and HTTP transport methods and typically rely on file formats such as the ISO Base Media File Format (ISOBMFF). One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). In DASH, there may be multiple representations for the video and / or audio data of multimedia content, and the different representations may correspond to different coding characteristics (e.g., different profiles or levels of the video coding standard, different bit rates, different spatial resolutions, etc.). Additionally, a technology called "picture-in-picture" has been proposed. This makes it worthwhile to research file formats that support picture-in-picture services. Summary of the Invention
[0004] SUMMARY OF THE INVENTION Embodiments of the present disclosure provide a solution for video processing.
[0005] In a first aspect, there is provided a method for video processing, the method comprising performing a conversion between a media file of a first video and a bitstream of the first video, the media file including an indication indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing picture-in-picture services.
[0006] According to the proposed method, an indication is adopted to indicate that two tracks are a track pair for providing picture-in-picture service, thereby advantageously enabling picture-in-picture service support in ISOBMFF-based media files.
[0007] In a second aspect, an apparatus for processing video data is proposed, the apparatus comprising: a processor; and a non-transitory memory having instructions, the instructions, when executed by the processor, causing the processor to perform a method according to the first aspect of the present disclosure.
[0008] In a third aspect, a non-transitory computer-readable storage medium is proposed, which stores instructions for causing a processor to perform the method according to the first aspect of the present disclosure.
[0009] In a fourth aspect, another non-transitory computer-readable storage medium is proposed, which stores a bitstream of a first video generated by a method executed by a media data processing device, the method including converting between a media file of the first video and the bitstream, the media file including an indication indicating that a first track carrying the bitstream of the first video and a second track carrying the bitstream of a second video are a track pair for providing picture-in-picture services.
[0010] In a fifth aspect, a method for storing a bitstream of a first video is proposed, the method comprising: performing a conversion between a media file of the first video and the bitstream; and storing the bitstream on a non-transitory computer-readable recording medium, the media file including an indication indicating that a first track carrying the bitstream of the first video and a second track carrying the bitstream of a second video are a track pair for providing picture-in-picture services.
[0011] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a media file of a first video generated by a method executed by a video processing device, the method including converting between the media file and a bitstream of the first video. The media file includes an indication of the video processing method, the indication indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing picture-in-picture services.
[0012] In a seventh aspect, a method for storing a media file of a first video is proposed, the method comprising: performing a conversion between the media file and a bitstream of the first video; and storing the media file on a non-transitory computer-readable recording medium, the media file including an indication indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing picture-in-picture services.
[0013] This Summary is provided to introduce, in a simplified form, a collection of concepts that are further described in the Detailed Description. This Summary is not intended to identify essential features or essential points of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. [Brief explanation of the drawings]
[0014] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent through the following detailed description, which refers to the accompanying drawings, in which like reference numerals generally refer to the same components. [Figure 1] 1 illustrates an example block diagram of a video encoding system according to some embodiments of the present disclosure. [Figure 2] 1 illustrates a first example of a block diagram illustrating a video encoding device according to some embodiments of the present disclosure. [Figure 3] 1 illustrates an example block diagram of a video decoding device according to some embodiments of the present invention. [Figure 4] A picture is shown partitioned into 18 tiles, 24 slices, and 24 sub-pictures. [Figure 5] 1 illustrates a typical sub-picture-based viewport-dependent 360° video delivery scheme. [Figure 6] Illustrates the extraction of one subpicture from a bitstream containing two subpictures and four slices. [Figure 7] 1 illustrates an example of picture-in-picture support based on VVC subpictures. [Figure 8] 1 illustrates a flowchart of a method for video processing according to some embodiments of the present disclosure. [Figure 9] 1 illustrates a block diagram of a computing device in which various embodiments of the present disclosure may be implemented. Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION OF THE INVENTION
[0015] Next, the principles of the present disclosure will be described with reference to some embodiments. It should be understood that these embodiments are provided for illustrative purposes only, and are intended to help those skilled in the art to understand and practice the present disclosure, without implying any limitations on the scope of the present disclosure. The present disclosure may be implemented in various ways other than the following forms described below.
[0016] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0017] References in this disclosure to "one embodiment," "an embodiment," "exemplary embodiment," etc., indicate that the described embodiment may include a characteristic feature, structure, or characteristic, but not all embodiments need include that characteristic feature, structure, or characteristic. Moreover, such phrases are not necessarily limited to referring to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an exemplary embodiment, it is submitted that it is within the knowledge of one of ordinary skill in the art to affect such feature, structure, or characteristic in connection with other embodiments, even if not explicitly stated.
[0018] Terms such as "first" and "second" may be used herein to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes a combination of one or more of the listed terms.
[0019] The terminology used in the specification is for the purpose of describing particular aspects only and is not intended to limit example embodiments. As used in the specification, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should be further understood that the terms "comprises," "comprising," "includes," and / or "including," when used in the specification, specify the presence of stated features, elements, or components, but do not exclude the presence of one or more other features, elements, components, and / or groups thereof.
[0020] Example environment 1 is a block diagram illustrating an example video encoding system 100 that can utilize techniques of this disclosure. As shown, video encoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoder, and the destination device 120 may also be referred to as a video decoder. In operation, source device 110 may be configured to generate encoded video data, and destination device 120 may be configured to decode the encoded video data generated by source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0021] The video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or combinations thereof.
[0022] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modem and / or a transmitter. The coded video data may be transmitted directly to the destination device 120 via the I / O interface 116 over the network 130A. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0023] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, configured to interface with an external display device.
[0024] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.
[0025] FIG. 2 is a block diagram illustrating an example of a video encoding device 200, which may be the video encoding device 114 in the system 100 shown in FIG. 1, in accordance with some embodiments of this disclosure.
[0026] The video encoding device 200 may be configured to implement any or all of the techniques of this disclosure. In the example of FIG. 2, the video encoding device 200 includes multiple functional components. The techniques described in this disclosure may be shared among various components of the video encoding device 200. In some examples, a processor (processing unit) may be configured to perform any or all of the techniques described in this disclosure.
[0027] In some embodiments, the functional components of the video encoding device 200 may include a division unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0028] In other examples, the video encoding device 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. An IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0029] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be integrated, but are depicted separately in the example of FIG. 2 for purposes of illustration.
[0030] The division unit 201 may divide a picture into one or more video blocks. The video encoding device 200 and the video decoding device 300 may support various video block sizes.
[0031] The mode select unit 203 may select one of either intra or inter coding modes based on, for example, an error result, and provide the resulting intra-coded or inter-coded block to a residual generation unit 207 to generate residual block data and to a reconstruction unit 212 to reconstruct a coded block for use as a reference picture. In some examples, the mode select unit 203 may select a combined intra- and inter-prediction (CIIP) mode that performs prediction based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode select unit 203 may select a resolution of a motion vector for the block (e.g., sub-pixel or integer pixel precision).
[0032] To perform inter prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block to one or more reference frames from buffer 213. The motion compensation unit 205 may determine a predictive video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.
[0033] The motion estimation unit 204 and motion compensation unit 205 may perform different operations on a current video block based, for example, on whether the current video block is an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture made up of macroblocks that are all based on macroblocks in the same picture. Additionally, as used herein, in some aspects, "P slice" and "B slice" may refer to a portion of a picture made up of macroblocks that are independent of macroblocks in the same picture.
[0034] In some examples, the motion estimation unit 204 may perform unidirectional prediction on the current video block, and the motion estimation unit 204 may search reference pictures in list 0 or list 1 for a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index that indicates a reference picture in list 0 or list 1 that contains the reference video block, and a motion vector that indicates a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. The motion compensation unit 205 may generate a predictive video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0035] Alternatively, in other embodiments, the motion estimation unit 204 may bidirectionally predict the current video block, and the motion estimation unit 204 may search for a reference video block for the current video block from among the reference pictures in list 0 and may search for another reference video block for the current video block from among the reference pictures in list 1. The motion estimation unit 204 may then generate reference indexes that indicate the reference pictures in lists 0 and 1 that contain the reference video blocks, and motion vectors that indicate the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and the motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predictive video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0036] In some implementations, the motion estimation unit 204 may output a full set of motion information for a decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by reference to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0037] In one embodiment, the motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoding device 300 that the current video block has the same motion information as another video block.
[0038] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoding device 300 may determine the motion vector of the current video block using the motion vector of the indicated video block and the motion vector difference.
[0039] As mentioned above, video encoding device 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by video encoding device 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0040] The intra prediction unit 206 may perform intra prediction on the current video block. If the intra prediction unit 206 intra predicts the current video block, the intra prediction unit 206 may generate predictive data for the current video block based on decoded samples of other video blocks in the same picture. The predictive data for the current video block may include a predicted video block and various syntax elements.
[0041] Residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.
[0042] In other examples, for example, in skip mode, there may be no residual data for the current video block, and residual generation unit 207 may not perform the subtraction operation.
[0043] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0044] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0045] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in buffer 213.
[0046] After reconstruction unit 212 reconstructs a video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0047] The entropy encoding unit 214 may receive data from other functional components of the video encoding device 200. If the entropy encoding unit 214 receives data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0048] FIG. 3 is a block diagram illustrating an example of a video decoding device 300 according to some embodiments of the present disclosure, which may be an example of the video decoding device 124 in the system 100 shown in FIG. 1.
[0049] Video decoding device 300 may be configured to perform any or all of the techniques of this disclosure. In the embodiment of FIG. 3, video decoding device 300 includes multiple functional modules. The techniques described in this disclosure may be shared among various components of video decoding device 300. In some examples, a processor (processing unit) may be configured to perform any or all of the techniques described in this disclosure.
[0050] 3, video decoding device 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306, as well as a buffer 307. In some examples, video decoding device 300 may perform a decoding path that is approximately the reverse of the decoding path described with respect to video encoding device 200 (FIG. 5).
[0051] The entropy decoding unit 301 retrieves an encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 decodes the entropy-encoded video data, and from the entropy-decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information by, for example, performing AMVP and merge mode. AMVP is used, which includes deriving several best candidates based on data from neighboring PB and reference pictures. The motion information generally includes horizontal and vertical motion vector displacement values, one or two reference picture indexes, and, for prediction regions in B slices, identification of which reference picture list is associated with each index. As used herein, in some aspects, "merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.
[0052] The motion compensation unit 302 may generate motion-compensated blocks, possibly performing interpolation based on an interpolation filter, and an identifier for the interpolation filter to be used with sub-pixel accuracy may be included in the syntax element.
[0053] The motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using the interpolation filter as used by the video encoding device 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filter used by the video encoding device 200 based on the received syntax information and use the interpolation filter to generate a prediction block.
[0054] The motion compensation unit 302 may use at least some of the syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is divided, a mode indicating how each division is coded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information to decode the encoded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently from other slices of the same picture with respect to entropy coding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0055] The intra prediction unit 303 may form a prediction block from spatially neighboring blocks, for example, using an intra prediction mode received in the bitstream. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0056] The reconstruction unit 306 may obtain decoded blocks, for example, by adding residual blocks with corresponding prediction blocks generated by motion compensation unit 302 or intra prediction unit 303. If desired, a deblocking filter may be applied to filter the decoded blocks to remove blocking artifacts. The decoded video blocks are stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and generates decoded video for display on a display device.
[0057] Some exemplary embodiments of the present disclosure are described in detail below. Section headings are used herein for ease of understanding and should not be understood to limit the applicability of the techniques and embodiments described in a section to that section alone. Furthermore, while certain embodiments are described with respect to general-purpose video encoding or other specific video codecs, the disclosed techniques are also applicable to other video encoding technologies. Furthermore, while some embodiments describe video encoding steps in detail, it is understood that the decoding step, which undoes the encoding, is implemented by a decoder. Furthermore, the term video processing encompasses video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compressed format to another or at different compressed bit rates. 1. Summary This disclosure relates to video file formats. Specifically, this disclosure relates to media This concept may be applied individually or in various combinations to media file formats, for example based on the ISO Base Media File Format (ISOBMFF) or its extensions. 2. Background technology 2.1 Video Coding Standards Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC MPEG-1 and MPEG-4 Visual, and both organizations H.262 / MPEG-2 Video and H.264 / MPEG-4 They jointly developed the Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding architecture that utilizes temporal prediction and transform coding. In 2015, VCEG and MPEG jointly established the Joint Video Experts Team (JVET) to explore future video coding technologies beyond HEVC. Since then, many new methods have been adopted by JVET and incorporated into reference software called JEM (Joint Exploration Mode). JVET was later renamed the Joint Video Experts Team (JVET) when the Versatile Video Coding (VVC) project was officially launched. VVC is a new coding standard that aims to achieve a 50% bitrate reduction compared to HEVC and was finalized at the 19th JVET General Meeting, which ended on July 1, 2020. The Generic Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the Generic Supplementary Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) are designed for use in the widest range of applications, including traditional uses such as television broadcasting, videoconferencing, and playback from storage media, as well as newer, more advanced uses such as adaptive bitrate streaming, video region extraction, content synthesis and combining from multiple coded video bitstreams, multi-view video, scalable layered coding, and viewport-adaptive 360° immersive media. The Basic Video Coding (EVC) standard (ISO / IEC 23094-1) is another video coding standard recently developed by MPEG. 2.2 File Format Standards Media streaming applications are typically based on IP, TCP, and HTTP transport methods and typically rely on file formats such as the ISO Base Media File Format (ISOBMFF) [7]. One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). To use video formats with ISOBMFF and DASH, video format-specific file format specifications, such as the AVC and HEVC file formats, are required for encapsulation of video content in ISOBMFF tracks and DASH representations and segments. Important information about the video bitstream, such as profile, tier, level, and much more, needs to be exposed as file format-level metadata and / or a DASH Media Presentation Description (MPD) for content selection, both for initialization at the start of a streaming session and for stream adaptation during a streaming session. Similarly, to use a picture format with ISOBMFF, a specific file format specification is required for the picture format, such as the AVC picture file format and the HEVC picture file format. The VVC video file format, a file format for storing VVC video content based on ISOBMFF, is currently being developed by MPEG. The VVC picture file format, a file format for storing picture content coded using VVC, based on ISOBMFF, is currently being developed by MPEG. 2.3 Picture Division and Subpictures in VVC In VVC, a picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that covers a rectangular area of a picture. A tile is a sequence of CTUs that covers a rectangular area of a picture. The CTUs in a tile are scanned in raster scan order within the tile. A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a picture tile. Two modes of slicing are supported: the raster scan slice mode and the rectangular slice mode. In the raster scan slice mode, a slice contains a complete sequence of tiles in the tile raster scan of a picture. In the rectangular slice mode, a slice contains either the number of complete tiles that collectively form a rectangular area of the picture, or the number of consecutive complete CTU rows of a tile that collectively form a rectangular area of the picture. The tiles in a rectangular slice are scanned in tile raster scan order within the rectangular area corresponding to the slice. A subpicture contains one or more slices that collectively cover a rectangular area of a picture. 2.3.1 Subpicture Concept and Function In VVC, each subpicture consists of one and only one rectangular slice that collectively covers a rectangular area of the picture, as shown, for example, in Figure 4. Subpictures may be defined as extractable (i.e., coded independently of other subpictures of the same picture and previous pictures in decoding order) or non-extractable. Whether a subpicture is extractable or not, the encoder can individually control for each subpicture whether in-loop filtering (including deblocking, SAO, and ALF) is applied across subpicture boundaries. Functionally, subpictures are similar to motion constrained tile sets (MCTS) in HEVC. They both enable independent encoding and extraction of rectangular subsets of a sequence of coded pictures for use cases such as viewport-dependent 360° video streaming optimization and region of interest (ROI) applications. With streaming 360° video, and especially omnidirectional video, at any given moment, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) is visible to the user, while the user can turn their head and change their viewing direction, and therefore their current viewport, at any time. While it is desirable to have lower-quality representations of at least some of the areas not covered by the current viewport available at the client and ready to display to the user if the user suddenly changes their viewing direction somewhere in the sphere, a high-quality representation of the omnidirectional video is only needed for the current viewport being displayed to the user at any given time. Dividing the high-quality representation of the entire omnidirectional video into sub-pictures of appropriate granularity allows for an optimization such as that shown in Figure 4, with 12 high-resolution sub-pictures on the left and the remaining 12 sub-pictures of the omnidirectional video at lower resolution on the right. Another typical sub-picture-based viewport-dependent 360° video delivery scheme is shown in Figure 5, where only the full video of the higher resolution representation is composed of sub-pictures, while the full video of the lower resolution representation can be encoded without sub-pictures and with less frequent RAPs than the higher resolution representation. The client receives the full video of the lower resolution, while for the higher resolution video, the client receives and decodes only the sub-pictures that currently cover the viewport. 2.3.2 Differences between Subpicture and MCTS There are several important design differences between subpictures and MCTS. First, the subpicture feature in VVC allows motion vectors of coded blocks to point outside the subpicture, even if the subpicture is extractable, by applying sample padding at subpicture boundaries as well as at picture boundaries. Second, additional modifications are introduced for motion vector selection and derivation in merge mode and the decoder-side motion vector refinement process of VVC. This allows for higher coding efficiency compared to non-standard motion constraints applied at the encoder side for MCTS. Third, when extracting one or more extractable subpictures from a sequence of pictures to create a conforming subbitstream, rewriting of the SH (and PH NAL units, if present) is not required. Subbitstream extraction based on the HEVC MCTS requires rewriting of the SH. Note that both HEVC MCTS extraction and VVC subpicture extraction require rewriting of the SPS and PPS. However, typically, only a few parameter sets exist in a bitstream, while each picture has at least one slice, so rewriting the SH can be a significant burden for application systems. Fourth, slices of different subpictures within a picture are allowed to have different NAL unit types. This is a feature often referred to as mixed NAL unit types or mixed subpicture types within a picture, which will be explained in more detail below. Fifth, VVC specifies HRD and level definitions for subpicture sequences, so that the conformance of each sub-bitstream of an extractable subpicture sequence can be guaranteed by the encoder. 2.3.3 Mixing Subpicture Types Within a Picture In AVC and HEVC, all VCL NAL units in a picture are required to have the same NAL unit type. VVC introduces the option to mix pictures with subpictures that have specific different VCL NAL unit types, thus providing support for random access not only at the picture level but also at the subpicture level. In VVC, VCL NAL units within a subpicture are still required to have the same NAL unit type. The random access capability from IRAP subpictures is beneficial for 360° video applications. In a viewport-dependent 360° video distribution scheme similar to that shown in Figure 5, the contents of spatially adjacent viewports largely overlap, i.e., only a portion of the subpictures within a viewport are replaced by new subpictures during a viewport orientation change, while the majority of subpictures remain within the viewport. Although the subpicture sequence newly introduced to a viewport must start with an IRAP slice, a significant reduction in the overall transmission bitrate can be achieved when the remaining subpictures are allowed to perform inter-prediction upon viewport change. An indication of whether a picture contains only one type of NAL unit or more than one type is provided in the PPS referenced by that picture (i.e., using a flag called pps_mixed_nalu_types_in_pic_flag). A picture can simultaneously consist of sub-pictures containing IRAP slices and sub-pictures containing trailing slices. A few other combinations of different NAL unit types within a picture are also allowed, including leading picture slices of NAL unit types RASL and RADL, which makes it possible to merge sub-picture sequences with open and closed GOP coding structures extracted from different bitstreams into one bitstream. 2.3.4 Subpicture Layout and ID Signaling The layout of subpictures in VVC is signaled in the SPS and is therefore constant in CLVS. Each subpicture is signaled by the location of its top-left CTU and its width and height in CTU numbers, thus ensuring that the subpicture covers a rectangular area of the picture at CTU granularity. The order in which the subpictures are signaled in the SPS determines the index of each subpicture within the picture. To enable extraction and merging of subpicture sequences without rewriting the SH or PH, the slice addressing scheme in VVC is based on a subpicture ID and a subpicture-specific slice index to associate slices with subpictures. In SH, the subpicture ID and subpicture-level slice index of the subpicture containing the slice are signaled. Note that the value of the subpicture ID of a particular subpicture can be different from the value of its subpicture index. The mapping between the two is either signaled within the SPS or PPS (but never both), or implicitly inferred. If present, the subpicture ID mapping needs to be rewritten or added when rewriting the SPS and PPS in the subpicture sub-bitstream extraction process. Both the subpicture ID and the subpicture-level slice index indicate to the decoder the exact location of the first decoded CTU of a slice within the DPB slot of the decoded picture. After sub-bitstream extraction, the subpicture ID of a subpicture remains unchanged, but the subpicture index may change. Even when the raster scan CTU address of the first CTU in a slice in a subpicture changes compared to its value in the original bitstream, the unchanged values of the subpicture ID and subpicture level slice index in the respective SH still accurately determine the location of each CTU in the decoded picture of the extracted subbitstream. Figure 6 shows the use of the subpicture ID, subpicture index, and subpicture level slice index to enable subpicture extraction using an example including two subpictures and four slices. Similar to subpicture extraction, signaling about subpictures allows merging several subpictures from different bitstreams into one bitstream by simply rewriting the SPS and PPS, provided that the different bitstreams are generated cooperatively (e.g., using distinct subpicture IDs but otherwise with roughly aligned SPS, PPS, and PH parameters such as CTU size, chroma format, encoding tool, etc.). Although subpictures and slices are signaled independently within the SPS and PPS, respectively, there are inherent inter-constraints between the subpicture layout and the slice layout in order to form a conforming bitstream. First, the presence of subpictures requires the use of rectangular slices and prohibits raster-scan slices. Second, slices of a given subpicture are assumed to be consecutive NAL units in decoding order, which means that the subpicture layout constrains the order of coded slice NAL units in the bitstream. 2.4 Picture-in-Picture Services Picture-in-picture services provide the ability to include a picture with a smaller resolution within a picture with a larger resolution. Such a service can be useful for showing two videos to a user simultaneously, whereby the video with the larger resolution is considered the main video and the video with the smaller resolution is considered the auxiliary video. Such picture-in-picture services can be used to provide accessibility services in which the main video is complemented by signage video. VVC subpictures can be used for picture-in-picture services by utilizing both the extraction and merging properties of VVC subpictures. For such services, the primary video is coded using several subpictures, one of which is the same size as the auxiliary video, placed at the exact location where the auxiliary video is intended to be composited into the primary video, and coded independently to enable extraction. If a user chooses to watch a version of the service that includes auxiliary video, the subpicture corresponding to the picture-in-picture region of the primary video is extracted from the primary video bitstream, and the auxiliary video bitstream is merged with the primary video bitstream at that location, as shown in Figure 7. Figure 7 shows an example of picture-in-picture support based on VVC subpictures. In this case, the pictures in the primary and auxiliary video must share the same video characteristics, in particular the bit depth, sample aspect ratio, size, frame rate, color space and transfer characteristics, and chroma sample positions. The primary and auxiliary video bitstreams do not need to use the same NAL unit type within each picture. However, merging requires that the coding order of the pictures in the primary and auxiliary bitstreams be the same. Because subpicture merging is required here, the subpicture IDs used in the primary and auxiliary video cannot overlap. Even if the auxiliary video bitstream consists of only one subpicture without further tile or slice division, subpicture information, particularly the subpicture ID and subpicture ID length, needs to be signaled to enable merging of the auxiliary video bitstream with the primary video bitstream. The subpicture ID length used to signal the length of the subpicture ID syntax element in the slice NAL unit of the auxiliary video bitstream must be the same as the subpicture ID length used to signal the subpicture ID in the slice NAL unit of the primary video bitstream. Furthermore, to simplify merging the auxiliary video bitstream with the primary video bitstream without having to rewrite the PPS partition information, it may be beneficial to encode the auxiliary video using only one slice and one tile within the corresponding area of the primary video. The primary and auxiliary video bitstreams must signal the same coding tools in the SPS, PPS, and picture headers. This includes using the same maximum and minimum allowed sizes for block partitions and the same value of the initial quantization parameter as signaled in the PPS (the same value of the pps_init_qp_minus26 syntax element). The use of coding tools can be modified at the slice header level. When both primary and auxiliary bitstreams are available in an ISOBMFF-based media file, the primary and auxiliary bitstreams can be stored in two separate file format tracks. 3. Challenges The following challenges have been observed with picture-in-picture support in ISOBMFF-based media files: 1) Although it is possible to use different file format tracks to store picture-in-picture main and auxiliary bitstreams separately, there is a lack of a mechanism for indicating such a purpose for such a pair of tracks in an ISOBMFF-based media file. 2) For example, as mentioned above, it is possible to use VVC sub-pictures for a picture-in-picture experience, but it is also possible to use other codecs and methods that do not allow the coded video data units representing the target picture-in-picture region in the main video to be swapped with the corresponding video data units in the auxiliary video. Therefore, it is necessary to indicate in the ISOBMFF-based media file whether such a swap is possible. 3) If the above interchanging is possible, the client needs to know which coded video data unit in each picture of the primary video represents the target picture-in-picture area so that the interchanging can be performed. Therefore, in this case, this information needs to be signaled in the ISOBMFF-based media file. 4) For content selection purposes and possibly other purposes as well, it would be useful to signal in ISOBMFF-based media files the location and size of the target picture-in-picture area in the main video. 4. Exemplary Embodiments To solve the above-mentioned problems, the following methods are disclosed. The present embodiments should be regarded as examples for illustrating the general concept and should not be construed in a narrow sense. Furthermore, these embodiments may be applied individually or combined in any way. For convenience of description, a pair of tracks carrying main and auxiliary bitstreams that together provide a picture-in-picture experience will be referred to as a picture-in-picture track pair or picture-in-picture track pair. 1) To solve the first problem, a new type of track reference is defined to indicate that the track containing the track reference and the track referenced by the track reference are a pair of picture-in-picture tracks. a. In one example, this new type of track reference is indicated by a track reference type equal to a particular value, e.g., "pips" (meaning "referencing a picture-in-picture auxiliary bitstream"), and the track containing this track reference carries the main bitstream, and the track referenced by the track reference carries the auxiliary bitstream. b. In another example, this new type of track reference is indicated by a track reference type equal to a particular value, e.g., "pipm" (meaning "referencing picture-in-picture main bitstream"), and the track containing this track reference carries an auxiliary bitstream, and the track referenced by the track reference carries an auxiliary bitstream. c. In yet another example, both track reference types described above are defined. 2) To solve the first and second problems, two new types of track references are defined to be included in tracks carrying the main bitstream: one indicates a pair of picture-in-picture tracks that allows swapping of coded video data units representing target picture-in-picture areas in the main video with corresponding video data units in the auxiliary bitstream, and the other indicates a pair of picture-in-picture tracks that does not allow such video data unit swapping. a. In one example, these two new types of track references are indicated by track reference type values equal to "ppsr" (meaning "points to a picture-in-picture auxiliary bitstream with video data unit reordering enabled") and "ppsn" (meaning "points to a picture-in-picture auxiliary bitstream with video data unit reordering not enabled"), respectively. 3) Alternatively, to solve the first and second problems, two new types of track references are defined to be included in tracks carrying auxiliary bitstreams, one indicating a pair of picture-in-picture tracks that allows the interchange of coded video data units representing target picture-in-picture areas in the main video with corresponding video data units in the auxiliary video, and the other indicating a pair of picture-in-picture tracks that does not allow such video data unit interchange. a. In one example, these two new types of track references are indicated by track reference type values equal to "ppmr" (meaning "points to a picture-in-picture main bitstream with video data unit interleaving enabled") and "ppmn" (meaning "points to a picture-in-picture main bitstream without video data unit interleaving enabled"), respectively. 4) To solve the first and second problems, alternatively, four new types of track references are defined, as described in items 2 and 3 above. 5) To solve all of the fourth challenges, a new type of entity grouping is defined: a. The new type of entity grouping is named Picture-in-Picture entity grouping, with grouping_type equal to "pinp" (or a different name or a different grouping type value, but with similar characteristics as described below). b. In one example, it is specified that each entity in the entity group must be a video track. c. PicInPicEntityGroupBox has been defined to carry at least one or more of the following information by extending EntityToGroupBox: i. The number N of main bitstream tracks. The entities (i.e., tracks in this context) identified by the first N entity_id values in the EntityToGroupBox are main bitstream tracks, and the entities identified by other entity_id values in the EntityToGroupBox are auxiliary bitstream tracks. To play a picture-in-picture experience, one of the main bitstream tracks is selected and one of the auxiliary bitstream tracks is selected. 1. Alternatively, the main bitstream tracks are signaled by a list of indexes into a list of entity_id values in the EntityToGroupBox, and other entities / tracks in the entity group are auxiliary bitstream tracks. 2. Alternatively, the main bitstream tracks are signaled by a list of indices into a list of track_id values, and the other entities / tracks in the entity group are auxiliary bitstream tracks. ii. An indication of whether the coded video data units representing the target picture-in-picture area in the primary video can be replaced with corresponding video data units in the auxiliary video. 1. In one example, the indication is signaled by a one-bit flag called data_units_replacable, with values 1 and 0 indicating that such video data unit replacement is enabled and disabled, respectively. iii. A list of region IDs to indicate which coded video data units in each picture of the primary video represent the target picture-in-picture regions. 1. In one example, it is specified that the specific semantics of region IDs must be specified specifically for a particular video codec. a. In one example, for VVC, it is specified that the region IDs are sub-picture IDs and the coded video data units are VCL NAL units. The VCL NAL units that represent the target picture-in-picture region in the primary video are the ones that have sub-picture IDs that are the same as the sub-picture IDs in the corresponding VCL NAL units of the auxiliary video (generally, all VCL NAL units of one picture in the auxiliary video share the same sub-picture ID that is explicitly signaled, and in this case there is only one region ID in the list of region IDs). b. In one example, in the case of VVC, when the client selects to replace a coded video data unit (which is a VCL NAL unit) representing a target picture-in-picture area in the primary video with a corresponding VCL NAL unit in the auxiliary video before sending it to a video decoding device, it is specified that, for each sub-picture ID, the VCL NAL unit in the primary video is replaced with the corresponding VCL NAL unit having that sub-picture ID in the auxiliary video without changing the order of the corresponding VCL NAL units. iv. The position and size within the main video to embed / overlay the auxiliary video, which is smaller in size than the main video. 1. In one example, this is signaled by signaling four values (x, y, width, height), where x, y specify the top-left position of the region, and width and height specify the width and height of the region. The units may be luma samples / pixel. 2. In one example, it is specified that when data_units_replacable is equal to 1 and position and size information is present, the position and size shall accurately represent the target picture-in-picture area in the main video. 3. In one example, when data_units_replacable is equal to 0 and position and size information is present, the position and size indicate a preferred area for overlaying and embedding the auxiliary video (i.e., the client may choose to overlay the auxiliary video in a different area of the main video). 4. In one example, when data_units_replaable is equal to 0 and position and size information is present, the position and size indicate a preferred area for overlaying and embedding the auxiliary video (i.e., the client may choose to overlay the auxiliary video in a different area of the main video). 5. Implementation form Below are some example embodiments for example embodiment item 5 and some of its subitems summarized in section 4 above. These embodiments can be applied to ISOBMFF. 5.1 Picture-in-Picture Entity Grouping 5.1.1 Definition The picture-in-picture service provides the ability to include video with a smaller spatial resolution within video with a larger spatial resolution, referred to as auxiliary video and primary video, respectively. Tracks within the same entity group with grouping_type equal to "pinp" can be used to support the picture-in-picture service by selecting one of the tracks indicated to contain the primary video and one of the other tracks (containing the auxiliary video). All entities within a picture-in-picture entity group shall be video tracks. 5.1.2 Syntax aligned(8) class PicInPicEntityGroupBox extends EntityToGroupBox('pinp',0,0) { unsigned int(8) num_main_video_tracks; unsigned int(1) data_units_replacable; unsigned int(1) pinp_window_info_present; bit (6) reserved = 0; if(data_units_replacable) { unsigned int(8) num_region_ids; for(i=0; i <num_region_ids; i++) unsigned int(16) region_id[i]; } if(pinp_window_info_present) { unsigned int(16) x; unsigned int(16) y; unsigned int(16) width; unsigned int(16) height; } } 5.1.3 Semantics num_main_video_tracks specifies the number of tracks in this entity group that carry picture-in-picture main video. data_units_replacable indicates whether coded video data units representing the target picture-in-picture area in the primary video are replaced by corresponding video data units in the auxiliary video. A value of 1 indicates that such video data unit replacement is enabled, and a value of 0 indicates that such video data unit replacement is not enabled. When data_units_replacable is equal to 1, the player may choose to replace the decoded video data units that represent the target picture-in-picture area in the primary video with the corresponding coded video data units in the auxiliary video before sending them to the video decoder for decoding. In this case, for a particular picture in the primary video, the corresponding video data units in the auxiliary video are all coded video data units in the decoded time synchronization samples in the auxiliary video track. In the case of VVC, when the client chooses to replace, for each sub-picture ID, the coded video data units (that are VCL NAL units) that represent the target picture-in-picture area in the primary video with the corresponding VCL NAL units in the auxiliary video before sending them to the video decoder, the VCL NAL units in the primary video are replaced with the corresponding VCL NAL units with that sub-picture ID in the auxiliary video without changing the order of the corresponding VCL NAL units. pinp_window_info_present equal to 1 specifies that the fields x, y, width, and height are present. A value of 1 specifies that these fields are not present. num_region_ids specifies the number of subsequent region_id[i] fields. region_id[i] specifies the i-th ID for the coded video data unit that represents the target picture-in-picture region. The exact semantics of region IDs need to be clearly specified for a particular video codec. For VVC, region IDs are subpicture IDs, and coded video data units are VCL NAL units. The VCL NAL units that represent target picture-in-picture regions in the primary video are the ones that have these subpicture IDs that are the same as the subpicture IDs in the corresponding VCL NAL units of the auxiliary video. x specifies the horizontal position of the top-left coded video pixel (sample) of the target picture-in-picture area in the main video, in video pixels (samples). y specifies the vertical position of the top-left coded video pixel (sample) of the target picture-in-picture area in the main video, in video pixels (samples). width specifies the width of the target picture-in-picture area in the main video, in video pixels (samples). height specifies the height of the target picture-in-picture area in the main video, in video pixels (samples).
[0058] Embodiments of the present disclosure relate to file format designs for picture-in-picture support. As used herein, a "picture-in-picture (PiP) service" provides the ability to include video with a lower spatial resolution (also called "auxiliary video" or "PiP video") within video with a larger spatial resolution (also called "primary video").
[0059] FIG. 8 is a flowchart of a method 800 for video processing according to some embodiments of the present disclosure. Method 800 may be implemented in a client or a server. As used herein, the term "client" may refer to a piece of computer hardware or software that accesses services made available by a server as part of a client-server model of a computer network. As an example, a client may be a smartphone or a tablet. As used herein, the term "server" may refer to a computing-enabled device, where the client accesses services over a network. A server may be a physical or virtual computing device.
[0060] As shown in FIG. 8, method 800 begins at 802 with performing conversion between a media file of a first video and a bitstream of the first video. The media file includes an indication indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing picture-in-picture services. As an example, if the spatial resolution of the first video is smaller than the spatial resolution of the second video, the indication may be a "pipm" track reference. That is, the track carrying the "pipm" track reference carries the bitstream of the auxiliary video, and the track referenced by the track reference carries the bitstream of the main video. It should be understood that the above example is provided for illustrative purposes only. The scope of the present disclosure is not limited in this respect.
[0061] According to the method 800, an indication is adopted to indicate that two tracks are a track pair for providing picture-in-picture services, thereby the proposed method advantageously enables supporting picture-in-picture services in media files based on ISOBMFF.
[0062] In some embodiments, the instruction may include a first type of track reference. The track reference may also indicate a second track. In one example, as described above, if the spatial resolution of the first video is smaller than the spatial resolution of the second video, the first type of track reference is a "pipm" track reference. In another example, if the spatial resolution of the second video is smaller than the spatial resolution of the first video, the first type of track reference is a "pips" track reference. That is, a track bearing the "pips" track reference carries a bitstream of the main video, and a track referenced by the track reference carries a bitstream of the auxiliary video.
[0063] In some embodiments, the media file of the secondary video may also include a second type of track reference indicating that the first track and the second track are a track pair for providing picture-in-picture services. Illustratively, if the spatial resolution of the first video is smaller than the spatial resolution of the secondary video, the second type of track reference is a "pips" track reference.
[0064] In some embodiments, converting may include generating a media file and storing the bitstream in the media file, hi some additional or alternative embodiments, converting may include parsing the media file to reconstruct the bitstream.
[0065] In some embodiments, the bitstream of the first video may be stored on a non-transitory computer-readable recording medium. The bitstream of the first video may be generated by a method executed by a video processing device, which performs conversion between a media file of the first video and the bitstream, and the media file includes an indication indicating that a first track carrying the bitstream of the first video and a second track carrying the bitstream of a second video are a track pair for providing picture-in-picture services.
[0066] In some embodiments, a conversion is performed between a media file of a first video and a bitstream, the media file including an indication indicating that a first track carrying the bitstream of the first video and a second track carrying the bitstream of a second video are a track pair for providing picture-in-picture services, and the bitstream may be stored on a non-transitory computer-readable recording medium.
[0067] In some embodiments, a media file of a first video may be stored on a non-transitory computer-readable recording medium. The media file of the first video may be generated by a method executed by a video processing device, the method including converting between the media file of the first video and a bitstream, the media file including an indication indicating that a first track carrying the bitstream of the first video and a second track carrying the bitstream of a second video are a track pair for providing picture-in-picture services.
[0068] In some embodiments, a conversion between a media file of a first video and a bitstream is performed, the media file including an indication indicating that a first track carrying the bitstream of the first video and a second track carrying the bitstream of a second video are a track pair for providing picture-in-picture services, and the media file may be stored on a non-transitory computer-readable recording medium.
[0069] Embodiments of the present disclosure can be described in light of the following clauses, the features of which can be combined in any reasonable manner.
[0070] Item 1: A method for video processing, comprising a step of performing conversion between a media file of a first video and a bitstream of the first video, wherein the media file comprises an indication indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing picture-in-picture services; a video processing method.
[0071] Item 2: The method of item 1, wherein the instructions include a first type of track reference.
[0072] Item 3: The method of item 2, wherein the spatial resolution of the first video is smaller than the spatial resolution of the second video, and the first type of track reference is a "pipm" track reference.
[0073] Item 4: The method of item 2, wherein the spatial resolution of the second video is smaller than the spatial resolution of the first video and the first type of track reference is a "pips" track reference.
[0074] Item 5: The method according to any one of clauses 2 to 4, wherein the track reference indicates a second track.
[0075] Item 6: A method according to any one of items 2 to 5, wherein the media file of the second video includes a second type track reference, and the second type track reference indicates that the first track and the second track are a track pair for providing a picture-in-picture service.
[0076] Item 7: The method according to any one of items 1 to 6, wherein the conversion includes generating the media file and storing the bitstream in the media file.
[0077] Item 8: The method of any one of items 1 to 6, wherein the conversion includes analyzing the media file to reconstruct the bitstream.
[0078] Item 9: An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions that, when executed by the processor, cause the processor to perform the method of any one of items 1 to 8.
[0079] Item 10: A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method of any one of items 1 to 8.
[0080] Item 11: A non-transitory computer-readable recording medium for storing a bitstream of a first video generated by a method executed by a video processing device, the method including a step of performing a conversion between a media file of the first video and the bitstream, the media file including instructions indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing picture-in-picture services.
[0081] Item 12: A method for storing a bitstream of a first video, comprising: a step of performing a conversion between a media file of the first video and the bitstream; and a step of storing the bitstream on a non-transitory computer-readable recording medium, wherein the media file includes an instruction indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing picture-in-picture services.
[0082] Item 13: A video processing method for storing a media file of a first video generated by a method executed by a video processing device, the method including a step of performing a conversion between the media file and a bitstream of the first video, and the media file including an instruction indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing a picture-in-picture service.
[0083] Item 14: A method for storing a media file of a first video, comprising: a step of performing a conversion between the media file and a bitstream of the first video; and a step of storing the media file on a non-transitory computer-readable recording medium, wherein the media file includes an instruction indicating that a first track carrying the bitstream of the first video and a second track carrying the bitstream of a second video are a track pair for providing picture-in-picture services.
[0084] Equipment example 9 illustrates blocks of a computing device 900 in which various embodiments of the present disclosure may be implemented. The computing device 900 may be implemented as or included in the source device 110 (or video encoding device 114 or 200) or the destination device 120 (or video decoding device 124 or 300).
[0085] It will be appreciated that the computing device 900 illustrated in FIG. 9 is for illustrative purposes only and is not intended to suggest any limitation on the functionality and scope of the embodiments of the present disclosure.
[0086] 9, the computing device 900 includes a general-purpose computing device 900. The computing device 900 may include at least one or more processors or processing units 910, memory 920, a storage unit 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960.
[0087] In some embodiments, computing device 900 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large-scale computing device, etc. The user terminal may be any type of mobile, fixed, or portable terminal, including, for example, a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a PC, a notebook PC, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio receiver, an electronic book reader, a game console, or any combination thereof, including accessories and peripherals of these devices. Computing device 900 may be considered capable of supporting any type of interface to a user (e.g., "wearable" circuitry, etc.).
[0088] The processing unit 910 may be a physical or virtual processor and may perform various processes based on programs stored in memory 920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of the computing device 900. The processing unit 910 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0089] Computing device 900 typically includes a variety of computer storage media. Such media may be any medium accessible by computing device 900, including, but not limited to, volatile and nonvolatile media, or removable and non-removable media. Memory 920 may be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. Storage unit 930 may be any removable or non-removable medium and may include a machine-readable medium such as a memory, a flash memory drive, a magnetic disk, or another medium that can be used to store information and / or data and that can be accessed by computing device 900.
[0090] The computing device 900 may further include additional removable / non-removable, volatile / non-volatile memory media. Although not shown in FIG. 9, a magnetic disk drive that reads from and writes to a removable, non-volatile magnetic disk and an optical disk drive that reads from and writes to a removable, non-volatile optical disk may also be provided. In such cases, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0091] The communications unit 940 communicates with additional computing devices via a communications medium. Additionally, the functionality of the components within computing device 900 may be implemented by a single computing cluster or multiple computing machines that can communicate via communications connections. Thus, computing device 900 may operate in a networked environment using logical connections with one or more other servers, network personal computers (PCs), or additional general network nodes.
[0092] The input device(s) 950 may be one or more of a variety of input devices such as a mouse, keyboard, tracking ball, voice input device, etc. The output device(s) 960 may be one or more of a variety of output devices such as a display, speakers, printer, etc. The communication unit 940 may further enable the computing device 900 to communicate with one or more external devices (not shown), such as storage and display devices, one or more devices that allow a user to interact with the computing device 900, or, if desired, any devices (such as a network card, modem, etc.) that enable the computing device 900 to communicate with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).
[0093] In some embodiments, instead of being integrated into a single device, some or all of the components of computing device 900 may be further deployed in a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to be aware of the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services over a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider may provide applications over a LAN that can be accessed through a web browser or any other computing component. Software or components of a cloud computing architecture and corresponding data may be stored on servers in remote locations. Computing resources within a cloud computing environment may be merged or distributed across locations in remote data centers. A cloud computing infrastructure may provide services through a shared data center but act as a single access point for users. Thus, a cloud computing architecture may be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on the client device.
[0094] The computing device 900 may be used to implement video encoding / decoding in embodiments of the present disclosure. The memory 920 may include one or more video encoding modules 925 having one or more program instructions. These modules are accessible and executable by the processing unit 910 to perform the functions of the various embodiments described herein.
[0095] In an exemplary embodiment that performs video encoding, input device 950 may receive video data as input 970 to be encoded. The video data may be processed, for example, by video encoding module 925, to generate an encoded bitstream. The encoded bitstream may be provided via output 960 as output device 980.
[0096] In an exemplary embodiment that performs video decoding, output 950 may receive an encoded bitstream as input 970. The encoded bitstream may be processed, for example, by video encoding module 925 to generate decoded video data. The decoded video data may be provided via output 960 as output device 980.
[0097] While the present disclosure has been particularly shown and described with reference to preferred embodiments thereof, it will be apparent to those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present disclosure, as defined by the appended claims. Such modifications are intended to be included within the scope of the present application. Accordingly, the foregoing description of embodiments of the present application is not intended to be limiting.
Claims
1. A video processing method comprising: performing a conversion between a media file of a first video and a bitstream of the first video; the media file includes a first indication indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing a picture-in-picture service; 20. The video processing method of claim 19, wherein the media file further includes a second indication indicating whether a first set of coded video data units representing a target picture-in-picture region in the second video are replaceable by a second set of coded video data units associated with the first video.
2. The method of claim 1 , wherein the first indication includes a track reference of a first type.
3. 3. The method of claim 2, wherein the spatial resolution of the first video is less than the spatial resolution of the second video, and the first type of track references are "pipm" track references.
4. The method of claim 2 , wherein the spatial resolution of the second video is less than the spatial resolution of the first video, and the first type of track references are “pips” track references.
5. The method of claim 2 , wherein the track reference points to the second track.
6. 3. The method of claim 2, wherein the media file of the second video includes a second type of track reference, and the second type of track reference indicates that the first track and the second track are a track pair for providing picture-in-picture services.
7. The method of claim 1, wherein if the first set of coded video data units are replaceable by the second set of coded video data units, the first set of coded video data units are allowed to be replaced by the second set of coded video data units before being sent to a decoder for decoding.
8. The method described in claim 7, wherein for a picture in the second video, the corresponding coded video data units of the first video are all coded video data units within the decoded time-synchronized samples in the first track.
9. The method of claim 1 , wherein the converting comprises generating the media file and storing the bitstream in the media file.
10. The method of claim 1 , wherein the converting comprises parsing the media file to reconstruct the bitstream.
11. 11. Apparatus for processing video data, comprising a processor and a non-transitory memory having instructions that, when executed by the processor, cause the processor to perform a method according to any one of claims 1 to 10.
12. A non-transitory computer readable storage medium storing instructions that cause a processor to perform the method of any one of claims 1 to 10.
13. 1. A method for storing a bitstream of a primary video, comprising: performing a conversion between a media file of the first video and the bitstream; storing the bitstream on a non-transitory computer-readable recording medium; the media file includes a first indication indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing a picture-in-picture service; the media file further includes a second indication indicating whether a first set of coded video data units representing a target picture-in-picture region in the second video are replaceable by a second set of coded video data units associated with the first video.
14. 1. A method for storing a media file of a first video, comprising: performing a conversion between the media file and a bitstream of the first video; storing the media file on a non-transitory computer-readable recording medium; the media file includes a first indication indicating that a first track carrying the bitstream of the first video and a second track carrying a bitstream of a second video are a track pair for providing a picture-in-picture service; the media file further includes a second indication indicating whether a first set of coded video data units representing a target picture-in-picture region in the second video are replaceable by a second set of coded video data units associated with the first video.
Citation Information
Patent Citations
Picture-in-picture methods for multimedia applications
JP2013538482A
Information processing device and method
JP2019146183A
Information processing device and information processing method
WO2016199608A1
Method and apparatus for late binding in media content
WO2020183053A1