Method, apparatus, and medium for video processing

The introduction of new track references and entity groupings in ISOBMFF addresses the lack of support for picture-in-picture services by enabling the swapping of video data units, ensuring seamless integration and playback of main and supplementary videos.

JP7701130B2Active Publication Date: 2025-07-01BYTEDANCE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024519072
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-27
Filing Date
2022-09-26
Publication Date
2025-07-01
Estimated Expiration
2042-09-26

AI Technical Summary

Technical Problem

Existing media file formats like ISOBMFF lack mechanisms to support picture-in-picture services, specifically in handling and indicating the replacement of encoded/decoded video data units between main and supplementary videos, which is crucial for providing seamless picture-in-picture experiences.

Method used

Introduce new track reference types and entity groupings in ISOBMFF to enable the swapping of encoded/decoded video data units between main and supplementary videos, allowing for the creation of picture-in-picture services by defining specific track references and groupings that indicate whether and how these units can be replaced.

Benefits of technology

Enables efficient support for picture-in-picture services in media files based on ISOBMFF, ensuring seamless integration and playback of main and supplementary videos with accurate spatial resolution and format compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701130000001
    Figure 0007701130000001
  • Figure 0007701130000002
    Figure 0007701130000002
  • Figure 0007701130000003
    Figure 0007701130000003
Patent Text Reader

Abstract

An embodiment of the present disclosure provides a technique for processing video. A method for processing video includes performing a conversion between a media file of a first video and a bitstream of the first video, the media file including a first indication for indicating whether a first group of encoded and decoded video data units indicating a target picture-in-picture region in the first video can be replaced with a second group of encoded and decoded video data units associated with a second video. The method of the present disclosure can easily support picture-in-picture services in a media file based on the ISO Base Media File Format (ISOBMFF).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims priority based on U.S. Provisional Application No. 63 / 248,832, filed on September 27, 2021, and incorporates by reference all the descriptions set forth in the U.S. application into this application.

[0002] Embodiments of the present disclosure mainly relate to video processing technology, and more specifically, to the design of a file format for corresponding to picture - in - picture.

Background Art

[0003] Media stream applications generally comply with transmission methods such as Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP), and generally depend on file formats such as ISO - based Media File Format (ISOBMFF). One such streaming system is HTTP - based Dynamic Adaptive Streaming over HTTP (DASH). In DASH, there may be various representations for video and / or audio data of multimedia content, and different representations may correspond to different codec characteristics (for example, different profiles or levels of video codec standards, different bitrates, different spatial resolutions, etc.). Furthermore, a technology called "picture - in - picture" has also been proposed. Therefore, there is room for consideration of the file format corresponding to the picture - in - picture service.

Summary of the Invention

[0004] Embodiments of the present disclosure provide a technical solution for processing video.

[0005] According to a first aspect, a method for processing video is provided. The method includes performing a conversion between a media file of a first video and a bitstream of the first video, the media file including a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0006] According to the method described above, the indication is adopted to indicate whether a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video can be replaced with a second group of encoded / decoded video data units related to the second video. Thereby, the proposed method can easily enable support for picture-in-picture services in media files based on ISOBMFF.

[0007] According to a second aspect, an apparatus for processing video data is provided. The apparatus for processing the video data includes a processor and a non-transitory memory storing instructions. When the instructions are executed by the processor, the processor is caused to execute the method according to the first aspect of the present disclosure.

[0008] According to a third aspect, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium causes a processor to execute instructions of the method according to the first aspect of the present disclosure.

[0009] According to a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed in a video processing apparatus. The method includes performing a conversion between a media file of a first video and the bitstream of the first video. The media file includes a first instruction for indicating whether a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0010] According to a fifth aspect, a method for storing a bitstream of a first video is provided. The method includes performing a conversion between a media file of a video and the bitstream of the first video, and storing the bitstream in a non-transitory computer-readable recording medium. The media file includes a first instruction for indicating whether a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0011] According to a sixth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a media file of a first video generated by a method executed in a video processing apparatus. The method includes performing a conversion between the media file of the first video and the bitstream of the first video, and the media file includes a first instruction for indicating whether a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0012] According to the seventh aspect, a method for storing a media file of a first video is provided. The method includes performing a conversion between the media file of the first video and the bitstream of the first video, and storing the media file in a non-transitory computer-readable recording medium. The media file includes a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0013] The description of the summary of the invention is for explaining the selection of concepts in a simplified form, and these will be described in detail in the following embodiments for carrying out the invention. The summary of the invention is not intended to identify the important features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure.

Brief Description of the Drawings

[0014] According to the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0015] Next, the principles of the present disclosure will be explained with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of explanation and to assist those skilled in the art in understanding and implementing the present disclosure, and do not mean to limit the scope of the present disclosure. The disclosure described in this specification can be implemented in various ways in addition to those described below.

[0016] In the following description and claims, unless otherwise specified, the meanings of all scientific and technical terms used in this specification are the same as those generally understood by those skilled in the technical field to which the present disclosure belongs.

[0017] Expressions such as "one embodiment", "embodiment", and "exemplary embodiment" in the present disclosure indicate that the described embodiments may include specific features, structures, or characteristics, but not all embodiments must include the above specific features, structures, or characteristics. Further, these phrases do not necessarily refer to the same embodiment. Further, when explaining a specific feature, structure, or characteristic with reference to an exemplary embodiment, it is considered within the knowledge of those skilled in the art to affect the features, structures, or characteristics of other embodiments, whether explicitly described or not.

[0018] The terms such as "first" and "second" can be used to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. The term "and / or" used herein includes any and all combinations of one or more of the listed terms.

[0019] As used herein, the terms are used only for the purpose of describing a particular embodiment and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless clearly stated otherwise in the context. Also, the terms "comprising", "including", and / or "having" used herein indicate the presence of the described features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0020] Environment of the Embodiment FIG. 1 illustrates a block diagram of an exemplary video codec system 100 to which the technology of the present disclosure is applicable. Thus, the video codec system 100 may include a source device 110 and a target device 120. The source device 110 is also referred to as a video encoding device, and the target device 120 is also referred to as a video decoding device. During operation, the source device 110 may be arranged to generate encoded video data, and the target device 120 may be arranged to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0021] The video source 112 may include a source such as, for example, a video capture device. The video capture device may include, but is not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or combinations thereof.

[0022] The video data can include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a bit sequence that forms and encodes the video data. The bitstream may include encoded pictures and data related thereto. The encoded picture is a representation of an encoded picture. The data related thereto may include a sequence parameter set, a picture parameter set, and other syntax structures (syntax). The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data can be directly transmitted to the target device 120 via the I / O interface 116 by the network 130A. The encoded video data can be stored in the storage medium / server 130B so that the target device 120 can access it.

[0023] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modulator / demodulator. The I / O interface 126 can obtain the encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 can decode the encoded video data. The display device 122 can display the decoded video data to the user. The display device 122 may be integrally configured with the target device 120, or an external display device interface may be arranged to be connectable to the target device 120 outside the target device 120.

[0024] The video encoder 114 and the video decoder 124 can operate in accordance with video compression standards, such as, for example, the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other conventional and / or further standards.

[0025] FIG. 2 is an exemplary block diagram showing a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be an example of the video encoder 114 of the system 100 shown in FIG. 1.

[0026] The video encoder 200 can be configured to implement any or all of the techniques of the present disclosure. In the example of FIG. 2, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the respective components of the video encoder 200. In some examples, the processor may be configured to execute any or all of the techniques described in the present disclosure.

[0027] In some embodiments, the video encoder 200 may include a splitting unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding / decoding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0028] In another example, the video encoder 200 may include more, fewer, or different functional components. In one embodiment, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit can perform prediction in the IBC mode. At least one reference picture is the picture in which the current video block is located.

[0029] Furthermore, some components (e.g., motion estimation unit 204 and motion compensation unit 205) can be integrated, but for the sake of convenience of explanation, these components are shown separated in the example of FIG. 2.

[0030] The splitting unit 201 can split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can accommodate various video block sizes.

[0031] The mode selection unit 203 can select, for example, one of various codec modes (intra coding or inter coding) based on error results, provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra and inter prediction (CIIP) modes to perform prediction based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 can also select the resolution (e.g., sub-pixel accuracy or integer-pixel accuracy) for the motion vector for the block.

[0032] To perform inter prediction for the current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 can identify a predicted video block for the current video block according to the motion information and the decoded samples of the pictures excluding the picture related to the current video block from the buffer 213.

[0033] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block according to whether the current video block is located in an I slice, a P slice, or a B slice. For example, as used herein, an "I slice" may be a part of a picture composed of macroblocks, and all macroblocks are derived from macroblocks in the same picture. Further, as used herein, in some cases, a "P slice" and a "B slice" may be parts of a picture composed of macroblocks different from those in the same picture.

[0034] In some examples, the motion estimation unit 204 can perform a unidirectional prediction on the current video block. The motion estimation unit 204 can search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Then, the motion estimation unit 204 can generate a reference index indicating the reference picture including the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction flag, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block of the current video block according to the reference video block indicated by the motion information of the current video block.

[0035] In another example, the motion estimation unit 204 can perform bidirectional prediction for the current video block. The motion estimation unit 204 can search for a reference picture in list 0 to find a reference video block for the current video block, and further search for a reference picture in list 1 to find another reference video block for the current video block. Then, the motion estimation unit 204 can generate a plurality of reference indexes indicating a plurality of reference pictures including a plurality of reference video blocks in list 0 and list 1, and a plurality of motion vectors indicating a plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 can output the plurality of reference indexes and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block according to the plurality of reference video blocks indicated by the motion information of the current video block.

[0036] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can refer to the motion information of other video blocks and transmit the motion information of the current video block by signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0037] In one embodiment, the motion estimation unit 204 can indicate one value to the video decoder 300 in the syntax structure related to the current video block, and this value indicates that the current video block has the same motion information as other video blocks.

[0038] In another example, the motion estimation unit 204 can identify a motion vector difference (MVD) with another video block in the syntax structure related to the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can determine the motion vector of the current video block by using the motion vector of the indicated video block and the motion vector difference.

[0039] As described above, the video encoder 200 can transmit motion vectors in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0040] The intra prediction unit 206 can perform intra prediction for the current video block. When the intra prediction unit 206 performs intra prediction for the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and each syntax element.

[0041] The residual generation unit 207 can generate residual data for the current video block by subtracting the (multiple) predicted video blocks of the current video block from the current video block (e.g., indicated by a minus sign). The residual data of the current video block may include residual video blocks corresponding to different sample portions of the samples of the current video block.

[0042] In another example, for example, in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not need to perform the subtraction operation.

[0043] The transformation processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transformations to the residual video block associated with the current video block.

[0044] After the transformation processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0045] The inverse quantization unit 210 and the inverse transformation unit 211 can apply inverse quantization and inverse transformation to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block and store it in the buffer 213.

[0046] After reconstructing the video block, the reconstruction unit 212 can perform loop filter processing to reduce video block effect artifacts in the video block.

[0047] The entropy encoding / decoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding / decoding unit 214 receives data, the entropy encoding / decoding unit 214 can perform one or more entropy encoding processes to generate entropy encoding / decoding data and output a bitstream including the entropy encoding / decoding data.

[0048] FIG. 3 shows an exemplary block diagram of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be an example of the video decoder 124 in the system 100 shown in FIG. 1.

[0049] The video decoder 300 may be configured to execute any or all of the techniques of the present disclosure. In the example of FIG. 3, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the components of the video decoder 300. In some examples, the processor may be arranged to execute any or all of the techniques described in the present disclosure.

[0050] In the example of FIG. 3, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can generally execute a decoding process for the encoding process described for the video encoder 200.

[0051] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., an encoded video data block). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can identify motion information including a motion vector, a motion vector precision, a reference picture list index, and other motion information from the entropy-decoded video data. The motion compensation unit 302 can identify the information, for example, by executing AMVP and the merge mode. AMVP is used including obtaining several most likely candidates from the data of neighboring PBs and reference pictures. The motion information generally includes horizontal and vertical motion vector displacement values and one or two reference picture indexes, and in the case of a prediction area in a B slice, further includes an identification of which reference picture list is associated with each index. As used herein, in some aspects, "merge mode" may refer to deriving motion information from spatially or temporally adjacent blocks.

[0052] The motion compensation unit 302 can generate a motion compensation block to perform interpolation by an interpolation filter in some cases. An identifier of the interpolation filter used with sub-pixel precision may be included in the syntax element.

[0053] The motion compensation unit 302 can calculate an interpolation value for sub-integer pixels of a reference block by using the interpolation filter used by the video encoder 200 during the encoding of a video block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can generate a prediction block by using the interpolation filter.

[0054] The motion compensation unit 302 can determine, using at least some of the syntax information, the size of the blocks of one or more frames and / or slices for encoding an encoded video sequence, the partitioning information that describes how each macroblock of a picture of the encoded video sequence is partitioned, the mode that indicates how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may be a data construct that is independently decodable from other slices of the same picture for entropy encoding, signal prediction, and residual signal reconstruction. A slice may be the entire picture or a region of the picture.

[0055] The intra prediction unit 303 can form a predicted block from spatially adjacent blocks, for example, using an intra prediction pattern received in the bitstream. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients decoded by the entropy decoding unit 301 provided in the bitstream. The inverse transform unit 305 applies an inverse transform.

[0056] The reconstruction unit 306 can obtain a decoded block, for example, by adding a residual block and a corresponding predicted block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking effect filter may be applied to filter the decoded block to remove block effect artifacts. Then, the decoded video blocks are stored in the buffer 307, and the buffer 307 provides reference blocks for subsequent motion compensation / intra prediction and further generates the decoded video for presentation to a display device.

[0057] Hereinafter, some exemplary embodiments of the present disclosure will be described in detail. It should be noted that the chapter headings in this specification are used for ease of understanding and do not limit the embodiments disclosed in a certain chapter to that chapter. Furthermore, although some embodiments have been described with reference to a general-purpose video codec or other specific video codecs, the disclosed technology is also applicable to other video codec technologies. Additionally, in some embodiments, the video encoding steps have been described in detail, but it should be understood that the corresponding decoding steps for canceling the encoding are realized by a decoder. Moreover, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding that represents video pixels from one compression format to another compression format, or at another compression bitrate. 1. Overview The present disclosure relates to video file formats. Specifically, it relates to picture-in-picture correspondence in media files. These concepts can be applied individually or in various combinations to media file formats, such as ISO base media file format (ISOBMFF) or its extensions. 2. Background Art 2.1 Regarding Video Codec Standards Video codec standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture using temporal prediction and transform coding. To explore future video coding technologies other than HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new approaches and incorporated them into a reference software called the Joint Exploration Model (JEM). With the official launch of the Versatile Video Coding (VVC) project, JVET changed its name to the Joint Video Experts Team (JVET). VVC is a new codec standard that aims to reduce the bitrate by 50% compared to HEVC, which was finalized at the 19th meeting of JVET on July 1, 2020. The Versatile Video Coding (VVC) standard (ITU-T H.266|ISO / IEC 23090-3) and the related Multifunctional Supplemental Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IEC 23002-7) are designed to be used in the broadest range of applications, including both traditional applications (such as television broadcasting, video conferencing, or playback from storage media) and newer, more advanced applications (such as adaptive bitrate streaming, video region extraction, and combinations and merges of multiple encoded / decoded video bitstreams, multi-view video, scalable layer coding, and viewport-adaptive 360° immersive media). The Elementary Video Coding (EVC) standard (ISO / IEC 23094-1) is another video codec standard recently developed by MPEG. 2.2 File Format Standards Media stream applications typically rely on IP, TCP, and HTTP transmission methods and usually depend on file formats such as the ISO base media file format (ISOBMFF). One of these streaming systems is Dynamic Adaptive Streaming over HTTP (DASH). To use video formats with ISOBMFF and DASH, video format-specific file format standards are required to encapsulate video content into ISOBMFF tracks and DASH representations or clips, such as the AVC file format and the HEVC file format. Important information regarding the video bitstream, such as profile, layer, class, and many other information, needs to be disclosed as file format class metadata and / or DASH Media Presentation Description (MPD) for the purpose of selecting appropriate media clips for both the purpose of content selection, such as initialization at the time of streaming session disclosure and streaming adaptation during the streaming session period. Similarly, to use the ISOBMFF image format, image format-specific file format standards are required, such as the AVC image file format and the HEVC image file format. The VVC video file format, which is a file format for storing VVC video content based on ISOBMFF, is currently being developed by MPEG. The VVC image file format based on ISOBMFF for storing image content encoded by VVC is currently being developed by MPEG. 2.3 Picture Partitioning and Sub-pictures in VVC In VVC, a picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs in a rectangular area covering the picture. The CTUs in a tile are scanned in raster scan order within this tile. A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture. Two types of slice modes, namely, raster scan slice mode and rectangular slice mode are supported. In the raster scan slice mode, a slice contains a complete tile sequence in the tile raster scan of a picture. In the rectangular slice mode, a slice contains a plurality of complete tiles that jointly form a rectangular region of a picture, or a plurality of consecutive complete CTU rows of one tile that jointly form a rectangular region of a picture. Tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to the slice. A sub-picture contains one or more slices that jointly cover a rectangular region of a picture. 2.3.1 Regarding the concept and function of sub-pictures In VVC, for example, as shown in FIG. 4, each sub-picture consists of one or more complete rectangular slices that jointly cover a rectangular region of a picture. A sub-picture may be limited to being extractable (i.e., encoded and decoded independently of other sub-pictures of the same image and other sub-pictures of previous images in the decoding order) or non-extractable. Regardless of whether a sub-picture is extractable or non-extractable, the encoder can control whether to apply loop filtering (including deblocking, SAO, and ALF) across sub-picture boundaries for each sub-picture individually. Functionally, a sub-picture is similar to a motion-constrained tile set (MCTS) in HEVC. Both allow for independent codec operations and extraction of rectangular subsets of an encoded and decoded picture sequence, for application examples such as 360° video stream optimization regarding a viewport and region of interest (ROI) applications. In 360° video streams (also known as omnidirectional videos), at any given point in time, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) is presented to the user, and the user can rotate his / her head at any time to change the viewing direction and the current viewport. At the same time, for areas that are not covered by at least the current viewport of the client terminal, it is desirable to have some low-quality representations, and be prepared to be rendered to the user to prevent the user from suddenly changing his / her viewing direction to any position on the spherical surface. At any given point in time, the high-quality representation of the omnidirectional video on the spherical surface is only required for the current viewport that needs to be rendered to the user. Such optimization can be achieved by dividing the high-quality representation of the entire omnidirectional video into sub-pictures with appropriate granularity, as shown in Figure 4. Here, Figure 4 has 12 sub-pictures with high resolution located on the left side and the remaining 12 sub-pictures of the omnidirectional video with low resolution located on the right side. Another representative viewport-related 360° video delivery method using sub-pictures is, as shown in Figure 5, that only the high-resolution representation of the complete video is composed of sub-pictures, while the low-resolution representation of the complete video can be encoded and decoded with a lower frequency of RAP without using sub-pictures, compared to the high-resolution representation. At the client terminal, a low-resolution complete video is received, but for the high-resolution video, only the sub-pictures covering the current viewport are received and decoded. 2.3.2 Differences between Sub-Pictures and MCTS There are several important design differences between sub-pictures and MCTS. First, similar to the picture boundary, the sub-picture feature of VVC allows the motion vectors of the coding blocks to point outside the sub-picture. Even in this case, the sub-picture can be extracted by applying sample padding at the sub-picture boundary. Second, during the merge mode and the motion refinement period on the VVC decoder side, the selection and export of motion vectors introduce additional variations, enabling higher codec efficiency compared to the non-conventional motion constraints applied on the MCTS encoder side. Third, when extracting one or more extractable sub-pictures from an image sequence to create a sub-bitstream (which is a coherent bitstream), it is not necessary to rewrite the SH (and the PHNAL unit if it exists). In HEVC MCTS-based sub-bitstream extraction, rewriting of the SH is required. Note that the SPS and PPS need to be rewritten in both HEVC MCTS extraction and VVC sub-picture extraction. However, generally, there are only a few parameter sets in one bitstream, and each picture has at least one slice, so rewriting the SH is a significant burden on the application system. Fourth, the slices of different sub-pictures within one picture can have different NAL unit types. As will be explained in more detail below, this is generally a feature called the mixed NAL unit type or mixed sub-picture type in a picture. Fifth, since VVC defines the HRD and level definition of the sub-picture sequence, the encoder can ensure the consistency of the sub-bitstream of each extractable sub-picture sequence. 2.3.3 Regarding the Mixed Sub-Picture Type of a Picture In AVC and HEVC, all VCL NAL units in a picture must have the same NAL unit type. VVC introduces the option that a picture can contain sub-pictures with different VCL NAL unit types, and supports random access not only at the picture level but also at the sub-picture level. In VVC, the VCL NAL units in a sub-picture still need to have the same NAL unit type. The function of random access from IRAP sub-pictures is beneficial for 360° video applications. In a delivery method similar to the 360° video that depends on the viewport shown in FIG. 5, the content of spatially adjacent viewports largely overlaps, that is, while the viewport direction changes, only a very small part of the sub-pictures in the viewport are replaced with new sub-pictures, but most of the sub-pictures remain in the viewport. The sequence of sub-pictures newly introduced into the viewport must be disclosed from an IRAP slice. However, if the remaining sub-pictures are allowed to perform inter-frame prediction during the change of the viewport, a significant reduction in the overall transmission bitrate can be achieved. The indication of whether a picture contains only a single type of NAL unit or multiple types is provided in the PPS to which the picture refers (that is, the flag named pps_mixed_nalu_types_in_pic_flag is used). A picture can be composed of sub-pictures including IRAP slices and sub-pictures including subsequent slices. There can be several other combinations of different NAL unit types in a picture, including the previous picture slices of NAL unit types RASL and RADL. This enables the merging of sub-picture sequences with open-GOP and close-GOP codec structures extracted from different bitstreams into one bitstream. 2.3.4 Regarding the Layout and ID Signaling of Sub-Pictures Since the layout of the sub-pictures in VVC is signaled in the SPS, it is not changed in the CLVS. Since each sub-picture is represented by the position of the CTU at its upper left corner and the width and height of a plurality of CTUs, the sub-picture thus ensures a rectangular area that covers the picture at the CTU granularity. The index of each sub-picture within the picture is determined in the order of the sub-pictures signaled in the SPS. In order to be able to extract and merge the sub-picture sequence without rewriting the SH or PH, the slice addressing method in VVC associates the slice with the sub-picture based on the sub-picture ID and the specific slice index of the sub-picture. In SH, the sub-picture ID of the sub-picture containing the slice and the slice index at the sub-picture level are signaled. Note that the value of the sub-picture ID of a specific sub-picture may be different from the value of its sub-picture index. The mapping between the two is signaled in the SPS or PPS (not both), or is implicitly inferred. If there is a mapping of the sub-picture ID, when rewriting the SPS and PPS during the sub-picture sub-bitstream extraction process, it is necessary to rewrite or add the mapping of the sub-picture ID. The sub-picture ID, together with the slice index at the sub-picture level, instructs the decoder on the exact position of the first decoded CTU of the slice within the DPB time slit of the decoded picture. After the extraction of the sub-bitstream, the sub-picture ID of the sub-picture does not change, but the sub-picture index may change. Even if the raster scan CTU address of the first CTU in the slice of the sub-picture changes compared to the value in the original bitstream, the unchanged sub-picture ID and the slice index at the sub-picture level in the corresponding SH correctly determine the position of each CTU in the decoded picture of the extracted sub-bitstream. Figure 6 shows the extraction of sub-pictures using the sub-picture ID, sub-picture index, and slice index at the sub-picture level by way of an example including two sub-pictures and four slices. Similar to the extraction of sub-pictures, for the signaling of sub-pictures, assuming that different bitstreams are cooperatively generated (for example, different sub-picture IDs are used, but in other aspects, SPS, PPS, and PH parameters such as the CTU size, chroma format, and codec tools are often aligned), by only rewriting the SPS and PPS, multiple sub-pictures from different bitstreams can be merged into a single bitstream. Sub-pictures and slices are signaled independently in SPS and PPS respectively, but there are specific mutual constraints between the sub-picture and slice layouts to form a consistent bitstream. First, to have sub-pictures, it is necessary to use rectangular slices and prohibit raster scan slices. Second, the slices of a given sub-picture need to be consecutive NAL units in the decoding order. This means that the sub-picture layout restricts the order of the encoded and decoded slice NAL units in the bitstream. 2.4 Regarding Picture-in-Picture Services The Picture-in-Picture service provides the function of including a low-resolution picture in a high-resolution picture. Such a service allows the user to display two videos simultaneously. Therefore, the video with the higher resolution is used as the main video, and the video with the lower resolution is used as the supplementary video. Such a Picture-in-Picture service is used to provide an accessibility service by supplementing the main video with a signage video. VVC sub-pictures are used in picture-in-picture services by taking advantage of their extraction and merging characteristics. In such services, the main video is encoded using multiple sub-pictures, and one of the multiple sub-pictures is the same size as the supplementary video and is located at the exact position where the supplementary video is to be combined with the main video and is encoded independently and extractably. When a user views a version of the service that includes the supplementary video, as shown in FIG. 7, the sub-picture corresponding to the picture-in-picture area of the main video is extracted from the main video bitstream, and the supplementary video bitstream is merged into its position in the main video bitstream. FIG. 7 shows an example of picture-in-picture correspondence using VVC sub-pictures. In this case, the pictures of the main video and the supplementary video must share the same video characteristics, particularly the bit depth, sample aspect ratio, size, frame rate, color space, transmission characteristics, and chroma sample position must be the same. The main video bitstream and the supplementary video bitstream do not need to use NAL unit types in each picture. However, for merging, it is necessary that the pictures of the main bitstream and the supplementary bitstream are encoded and decoded in the same order. In the present disclosure, since merge sub-pictures are required, the sub-picture IDs used in the main video and the supplementary video must not overlap. The supplementary video bitstream, even if it has no other tile or slice partitioning and consists of only one sub-picture, needs to merge the supplementary video bitstream into the main video bitstream by signaling sub-picture information, especially the sub-picture ID and the sub-picture ID length. The sub-picture ID length of the sub-picture syntax element within the slice NAL unit for signaling the supplementary video bitstream must be the same as the sub-picture ID length of the sub-picture ID within the slice NAL unit for signaling the main video bitstream. Also, in order to simplify the merging of the supplementary video bitstream and the main video bitstream without rewriting the PPS partitioning information, it is beneficial to encode the supplementary video using only one slice and one tile in the corresponding area of the main video. The main video bitstream and the auxiliary video bitstream must signal the same codec tools used in the SPS, PPS, and picture header. This includes using the same maximum allowable size, minimum allowable size for block partitioning, and the same value as the initial quantization parameter indicated in the PPS (the same value as the pps_init_qp_minus26 syntax element). The use of codec tools can be changed at the slice header level. If the main bitstream and the supplementary bitstream are available within an ISOBMFF-based media file, they can be stored in two separate file format tracks. 3. Problems When supporting picture-in-picture in an ISOBMFF-based media file, the following problems have been pointed out: 1) It is possible to store each of the picture-in-picture main bitstream and the supplementary bitstream using different file format tracks, but there is a lack of a mechanism for indicating such a pair of tracks within a media file based on ISOBMFF. 2) It is possible to realize the picture-in-picture experience using VVC sub-pictures. However, for example, as described above, when it is not possible to swap the coded video data unit indicating the target picture-in-picture area in the main video with the corresponding video data unit of the supplementary video, other decoders and methods may be utilized. Therefore, in a media file based on ISOBMFF, it is necessary to indicate whether such a swap is possible. 3) When the above-described swap is possible, on the client terminal, if it does not know which of the encoded / decoded video data units in each picture of the main video represents the image area in the target picture, the swap cannot be performed. Therefore, in this case, in a media file based on ISOBMFF, it is necessary to signal this information. 4) For the purpose of content selection and other possible purposes, it is useful to signal the position and size of the target picture-in-picture area in the main video in a media file based on ISOBMFF. 4. Exemplary Embodiments To solve the above-described problems, the following method, which will be generally described, is disclosed. Note that this embodiment should not be understood as being limited as an example for explaining general concepts. Furthermore, the embodiments may be applied individually or in any combination. For convenience, a pair of tracks carrying the main bitstream and the supplementary bitstream that jointly provide the picture-in-picture experience is referred to as a pair of picture-in-picture tracks or a pair of picture-in-picture tracks. 1) To solve the initial problem, a new track reference type is defined such that a track contains a track reference and the track referred to by the track reference is a pair of picture-in-picture tracks. a. In one embodiment, such a new type of track reference is indicated by a track reference type equal to a specific value, for example, "pips" (meaning "refer to the picture-in-picture supplementary bitstream"), and the track containing the track reference carries the main bitstream, while the track referred to by the track reference carries the supplementary bitstream. b. In another example, such a new type of track reference is indicated by a track reference type equal to a specific value, for example, "pipm" (meaning "refer to the picture-in-picture main bitstream"), and the track containing the track reference carries the supplementary bitstream, while the track referred to by the track reference carries the main bitstream. c. In yet another example, two types of track reference types as described above are defined. 2) To solve the first and second problems, two new types of track references are defined for tracks carrying the main bitstream, one of which indicates a pair of picture-in-picture tracks that can replace the encoded video data unit indicating the target picture-in-picture area in the main video with the corresponding video data unit in the supplementary bitstream, and the other indicates a pair of picture-in-picture tracks that do not allow such replacement of the video data unit. a. In one embodiment, these two new types of track references are indicated by track reference type values equal to "ppsr" (meaning "refer to the picture-in-picture supplementary bitstream where replacement of the video data unit is possible") and "ppsn" (meaning "picture-in-picture supplementary bitstream that does not allow replacement of the video data unit"). 3) Alternatively, to solve the first and second problems, a new type definition of two types of track references included in the track carrying the supplementary bitstream is defined. One of them indicates a pair of picture-in-picture tracks that can replace the encoded / decoded video data unit indicating the target picture-in-picture area in the main video in the corresponding video data unit of the supplementary bitstream, and the other indicates a pair of picture-in-picture tracks that do not allow such replacement of video data units. a. In one embodiment, the new type of these two types of track references is indicated by track reference type values equal to "ppmr" (meaning "referring to the picture-in-picture main bitstream where replacement of video data units is possible") and "ppmn" (meaning "referring to the picture-in-picture main bitstream that does not allow replacement of video data units"). 4) Alternatively, to solve the first and second problems, as described in items 2 and 3 above, a new type of the four types of track references is defined. 5) To solve the four problems described above, a new type of entity grouping is defined. Details will be described below. a. The new type of entity grouping is named picture-in-picture entity grouping, and its grouping_type is equal to "pinp" (or a different name or different grouping type value, but having the same function as described below). b. In one embodiment, it is stipulated that each entity within the entity group must be a video track. c. The PicInPicEntityGroupBox is defined by extending the EntityToGroupBox and carries at least one or more of the following information: i. The number of main bitstream tracks is N. The entities (i.e., tracks here) identified by the first N entity_id values within EntityToGroupBox are main bitstream tracks, and those identified by other entity_id values within the entity EntityToGroupBox are supplementary bitstream tracks. To play the picture-in-picture experience, one main bitstream track in the main bitstream tracks is selected, and one supplementary bitstream track in the supplementary bitstream tracks is selected. 1. Alternatively, the main bitstream tracks are signaled to the list of Entity_id values within EntityToGroupBox via an index list, and the other entities / tracks within the entity group are supplementary bitstream tracks. 2. Alternatively, the main bitstream tracks are signaled via a list of track_id values, and the other entities / tracks within the entity group are supplementary bitstream tracks. ii. Regarding an indication for whether the encoded / decoded video data unit representing the target picture-in-picture area in the main video can be replaced with the corresponding video data unit in the supplementary video. 1. In one embodiment, the indication is signaled by a 1-bit flag called data_units_replacable, and the values 1 and 0 indicate enabling and disabling the replacement of this video data unit, respectively. iii. Regarding a list of region IDs for indicating which of the encoded / decoded video data units in each picture of the main video indicate the target picture-in-picture area. 1. In one embodiment, it is stipulated that for a specific video decoder, it is necessary to explicitly define the specific semantics of the region ID. a. In one embodiment, it is defined as follows: In the case of VVC, the region ID is the sub-picture ID, and the encoded / decoded video data unit is the VCL NAL unit. The VCL NAL unit indicating the target picture in-picture region in the main video is the VCL NAL unit having these sub-picture IDs, and these sub-picture IDs are the same as the sub-picture IDs in the corresponding VCL NAL unit of the supplementary video (generally, all VCL NAL units of one frame in the supplementary video share the same sub-picture ID that is clearly signaled, and in this case, there is only one region ID in the list of region IDs). b. In one embodiment, it is defined as follows: In the case of VVC, before being sent to the video decoder, at the client terminal, if it is selected to replace the encoded / decoded video data unit (i.e., the VCL NAL unit) indicating the target picture in-picture region in the main video using the corresponding VCL NAL unit of the supplementary video, then for each sub-picture ID, without changing the order of the corresponding VCL NAL units, the VCL NAL unit in the main video is replaced with the corresponding VCL NAL unit having the sub-picture ID in the supplementary video. iv. The position and size in the main video for embedding / overlaying the supplementary video are smaller than the main video in terms of size. 1. In one embodiment, it is signaled with four values (x, y, width, height), where x and y specify the position of the upper left corner of the region, and width and height specify the width and height of the region. The unit may be luminance samples / pixel. 2. In one embodiment, when data_units_replacable is equal to 1 and position and size information exist, it is stipulated that the position and size should correctly indicate the target picture in-picture region in the main video. 3. In one embodiment, when data_units_replacable is equal to 0 and location and size information exists, the location and size information is defined to indicate a suitable area for embedding and covering the supplementary video (i.e., on the client terminal, it is possible to select to overlay the supplementary video on a different area of the main video). 4. In one embodiment, when data_units_replacable is equal to 0 and location and size information does not exist, there is no information or recommendation on where to cover the supplementary video, and it is defined to completely depend on the client's selection. 5. Embodiment The following are embodiments of some exemplary embodiments generally described in Item 5 and Section 4 above. These embodiments are applicable to ISOBMFF. 5.1 Regarding Picture-in-Picture Entity Grouping 5.1.1 Regarding Definitions The Picture-in-Picture service provides a function of including a video with a lower spatial resolution in a video with a higher spatial resolution, which are respectively called the supplementary video and the main video. By selecting one track from the tracks indicated to include the main video and one track (including the supplementary video) from other tracks, the tracks within the same entity group where the grouping_type is equal to "pinp" are used to correspond to the Picture-in-Picture service. All entities within the Picture-in-Picture entity group should be video tracks. 5.1.2 Regarding Syntax aligned(8) class PicInPicEntityGroupBox extends SEntityToGroupBox("pinp", 0, 0) { unsigned int(8) num_main_video_tracks; unsigned int(1) data_units_replacable; unsigned int(1) pinp_window_info_present; bit(6) reserved = 0; if (data_units_replacable) { unsigned int(8) num_region_ids; for (i = 0; i < num_region_ids; i++) unsigned int(16) region_id[i]; } if (pinp_window_info_present) { unsigned int(16) x; unsigned int(16) y; unsigned int(16) width; unsigned int(16) height; } } 5.1.3 Regarding Semantics num_main_video_tracks specifies the number of tracks carrying picture-in-picture main video within its entity group. data_units_replacable indicates whether the encoded / decoded video data units representing the target picture-in-picture region within the main video can be replaced by the corresponding video data units of the supplementary video. A value of 1 indicates that such replacement of video data units is allowed, and a value of 0 indicates that such replacement of video data units is not allowed. When data_units_replacable is equal to 1, before being sent to the video decoder for decoding, the player can choose to replace the encoded and decoded video data unit indicating the target picture in picture area in the main video with the encoded and decoded video data unit corresponding to the supplementary video. In this case, for a specific picture in the main video, the corresponding video data units of the supplementary video are all the encoded and decoded video data units within the decoded time synchronization sample in the supplementary video track. In the case of VVC, before being sent to the video decoder, on the client side, if it is selected to replace the encoded and decoded video data unit (i.e., VCL NAL unit) indicating the target picture in picture area in the main video with the corresponding VCL NAL unit in the supplementary video, for each sub-picture ID, without changing the order of the corresponding VCL NAL units, the VCL NAL units in the main video can be replaced with the corresponding VCL NAL units having the sub-picture ID in the supplementary video. That pinp_window_info_present is equal to 1 specifies that the fields x, y, width, and height exist. A value of 0 specifies that these fields do not exist. num_region_ids specifies the number of subsequent fields Region_id[i]. region_id[i] specifies the i-th ID of the encoded and decoded video data unit indicating the target picture in picture area. For a specific video decoder, it is necessary to explicitly define the specific semantics of the region ID. In the case of VVC, the region ID is the sub-picture ID, and the encoded and decoded video data unit is the VCL NAL unit. The VCL NAL unit indicating the target picture in picture area in the main video is the VCL NAL unit having these sub-picture IDs, and these sub-picture IDs are the same as the sub-picture IDs within the corresponding VCL NAL unit of the supplementary video. x specifies the horizontal position of an encoded video pixel (sample) located at the upper left corner of the target picture-in-picture area in the main video. The unit is a video pixel (sample). y specifies the vertical position of an encoded video pixel (sample) located at the upper left corner of the target picture-in-picture area in the main video. The unit is a video pixel (sample). Width specifies the width of the target picture-in-picture area in the main video. The unit is a video pixel (sample). Height specifies the height of the target picture-in-picture area in the main video. The unit is a video pixel (sample).

[0058] Embodiments of the present disclosure relate to a file format design for corresponding to one type of picture-in-picture. As used herein, a "picture-in-picture (PiP) service" provides a function (also referred to as a "main video") of including a video with a lower spatial resolution (also referred to as a "supplementary video" or a "PiP video") in a video with a higher spatial resolution.

[0059] FIG. 8 shows a flowchart of a method 800 for processing video according to some embodiments of the present disclosure. Method 800 can be implemented on a client terminal or a server. As used herein, the term "client terminal" may refer to a part of computer hardware or software that accesses services provided by a server in a client-terminal / server model as a computer network. For example, the client terminal may be a smartphone or a tablet. As used herein, the term "server" may refer to a device having computing capabilities. In this case, the client terminal accesses the server via a network. The server may be a physical computer device or a virtual computer device.

[0060] As shown in FIG. 8, method 800 starts at block 802, where a conversion between a media file of a first video and a bitstream of the first video is performed. The media file includes a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture in picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video. For example, the first indication may include a 1-bit flag. In one example, when the flag is a first value (e.g., 1), the first group of encoded / decoded video data units can be replaced with the second group of encoded / decoded video data units, and when the flag is a second value (e.g., 0), the first group of encoded / decoded video data units cannot be replaced with the second group of encoded / decoded video data units. It should be understood that the above examples are for illustrative purposes only. The scope of the present disclosure is not limited thereto.

[0061] According to method 800, an indication is adopted for indicating whether an encoded / decoded video data unit indicating an image area in a target picture in the first video can be replaced with an encoded / decoded video data unit related to a second video. Thereby, the proposed method can easily enable correspondence to a picture in picture service in a media file based on ISOBMFF.

[0062] In some embodiments, the spatial resolution of the second video is lower than that of the first video. That is, the second video is a supplementary video and the first video is a main video.

[0063] In some embodiments, the first indication may include a first track reference that indicates a track carrying the bitstream of the second video. For example, when the first track reference has a first type, a first group of encoded / decoded video data units can be swapped with a second group of encoded / decoded video data units, and when the first track reference has a second type, a first group of encoded / decoded video data units cannot be swapped with a second group of encoded / decoded video data units. In one embodiment, the first track reference of the first type is a "ppsr" track reference, and the first track reference of the second type is a "ppsn" track reference. It should be understood that the above examples are for illustrative purposes only. The scope of the present disclosure is not limited thereto.

[0064] In some embodiments, the media file of the second video includes a second indication for indicating whether a first group of encoded / decoded video data units can be swapped with a second group of encoded / decoded video data units. For example, the second indication includes a second track reference that indicates a track carrying the bitstream of the first video. When the second track reference has a third type, a first group of encoded / decoded video data units can be swapped with a second group of encoded / decoded video data units, and when the second track reference has a fourth type, a first group of encoded / decoded video data units cannot be swapped with a second group of encoded / decoded video data units. In one embodiment, the second track reference of the third type is a "ppmr" track reference, and the second track reference of the fourth type is a "ppmn" track reference. It should be understood that the above examples are for illustrative purposes only. The scope of the present disclosure is not limited thereto.

[0065] In some embodiments, the first group of encoded / decoded video data units includes video coding layer network abstraction layer (VCL NAL) units, and the second group of encoded / decoded video data units includes VCL NAL units.

[0066] In some embodiments, the conversion includes creating a media file and storing the bitstream in the media file. Alternatively or additionally, in some embodiments, the conversion includes parsing the media file to reconstruct the bitstream.

[0067] In some embodiments, the bitstream of the first video can be stored in a non-transitory computer-readable recording medium. The bitstream of the first video can be generated by a method executed in a video processing device. According to the method, a conversion between the media file of the first video and the bitstream of the first video is performed. The media file includes a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture in picture area in the first video can be replaced with a second group of encoded / decoded video data units related to the second video.

[0068] In some embodiments, a conversion between the media file of the first video and the bitstream is performed. The media file includes a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture in picture area in the first video can be replaced with a second group of encoded / decoded video data units related to the second video. The bitstream can be stored in a non-transitory computer-readable recording medium.

[0069] In some embodiments, the media file of the first video can be stored in a non-transitory computer-readable recording medium. The media file of the first video can be generated by a method executed in a video processing device. According to the method, a conversion between the media file of the first video and the bitstream of the first video is performed. The media file includes a first instruction for indicating whether a first group of encoded / decoded video data units indicating a target picture in picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0070] In some embodiments, a conversion between the media file of the first video and the bitstream is performed. The media file includes a first instruction for indicating whether a first group of encoded / decoded video data units indicating a target picture in picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video. The media file can be stored in a non-transitory computer-readable recording medium.

[0071] Embodiments of the present disclosure can be described based on the following clauses, and the features related to these clauses can be combined in any reasonable form.

[0072] Clause 1. A method for processing video, comprising performing a conversion between a media file of a first video and a bitstream of the first video, wherein the media file includes a first instruction for indicating whether a first group of encoded / decoded video data units indicating a target picture in picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0073] Clause 2. The method according to Clause 1, wherein the spatial resolution of the second video is lower than the spatial resolution of the first video.

[0074] Clause 3. The first indication is the method described in any one of Clauses 1 to 2 and includes a 1-bit flag.

[0075] Clause 4. The method according to Clause 3, wherein when the flag is the first value, the encoded / decoded video data units of the first group are interchangeable with the encoded / decoded video data units of the second group, and when the flag is the second value, the encoded / decoded video data units of the first group are not interchangeable with the encoded / decoded video data units of the second group.

[0076] Clause 5. The first indication is the method described in any one of Clauses 1 to 2 and includes a first track reference indicating a track carrying the bitstream of the second video.

[0077] Clause 6. The method according to Clause 5, wherein when the first track reference has the first type, the encoded / decoded video data units of the first group are interchangeable with the encoded / decoded video data units of the second group, and when the first track reference has the second type, the encoded / decoded video data units of the first group are not interchangeable with the encoded / decoded video data units of the second group.

[0078] Clause 7. The method according to Clause 6, wherein the first track reference of the first type is a "ppsr" track reference, and the first track reference of the second type is a "ppsn" track reference.

[0079] Clause 8. The method according to any one of Clauses 5 to 7, wherein the media file of the second video includes a second indication for indicating whether the encoded / decoded video data units of the first group are interchangeable with the encoded / decoded video data units of the second group.

[0080] Clause 9. The second instruction is the method according to clause 8, including a second track reference that indicates a track carrying the bitstream of the first video.

[0081] Clause 10. When the second track reference has a third type, the first group of encoded / decoded video data units is interchangeable with the second group of encoded / decoded video data units, and when the second track reference has a fourth type, the first group of encoded / decoded video data units is not interchangeable with the second group of encoded / decoded video data units, the method according to clause 9.

[0082] Clause 11. The second track reference of the third type is a "ppmr" track reference, and the second track reference of the fourth type is a "ppmn" track reference, the method according to clause 10.

[0083] Clause 12. The first group of encoded / decoded video data units includes video codec layer network abstraction layer (VCL NAL) units, and the second group of encoded / decoded video data units includes VCL NAL units, the method according to any one of clauses 1 to 11.

[0084] Clause 13. The conversion includes creating a media file and storing the bitstream in the media file, the method according to any one of clauses 1 to 12.

[0085] Clause 14. The conversion includes analyzing the media file and reconstructing the bitstream, the method according to any one of clauses 1 to 12.

[0086] Clause 15. An apparatus for processing video data, comprising a processor and a non-transitory memory storing instructions, and when the instructions are executed by the processor, causing the processor to execute the method according to any one of clauses 1 to 14.

[0087] A non - transitory computer - readable storage medium for storing instructions that cause a processor to execute the method according to any one of clauses 1 to 14.

[0088] A non - transitory computer - readable recording medium that stores a bitstream of a first video generated by a method executed in a video processing apparatus, the method including performing a conversion between a media file of the first video and the bitstream, the media file including a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture - in - picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0089] A method for storing a bitstream of a first video, the method including performing a conversion between a media file of the first video and the bitstream, and storing the bitstream in a non - transitory computer - readable recording medium, the media file including a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture - in - picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0090] A non - transitory computer - readable recording medium that stores a media file of a first video generated by a method executed in a video processing apparatus, the method including performing a conversion between the media file and a bitstream of the first video, the media file including a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture - in - picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0091] A method for storing a media file of a first video, the method including performing a conversion between the media file and a bitstream of the first video, and storing the media file in a non-transitory computer-readable recording medium, wherein the media file includes a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture in-picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video.

[0092] Example of the Device FIG. 9 shows a block diagram of a computer device 900 capable of implementing each embodiment of the present disclosure. The computer device 900 may be configured as a source device 110 (or video encoder 114 or 200) or a target device 120 (or video decoder 124 or 300), or may be included in the source device 110 (or video encoder 114 or 200) or the target device 120 (or video decoder 124 or 300).

[0093] It should be understood that the computer device 900 shown in FIG. 9 is for illustrative purposes only and does not imply any limitation on the functions and scope of the embodiments of the present disclosure.

[0094] As shown in FIG. 9, the computer device 900 includes a general-purpose computer device 900. The computer device 900 may include at least one or a plurality of processors or processing units 910, a memory 920, a storage device 930, one or a plurality of communication units 940, one or a plurality of input devices 950, and one or a plurality of output devices 960.

[0095] In some embodiments, the computer device 900 may be configured as any user terminal or server terminal having computing functions. The server terminal may be a server provided by a service provider, a large computer device, or the like. The user terminal may be, for example, a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a locator device, a television receiver, a radio broadcast receiver, an electronic paper device, a game device, or any combination thereof, and any type of mobile terminal, fixed terminal, or portable terminal including attachments and peripheral devices of these devices or any combination thereof. The computer device 900 is assumed to be capable of corresponding to any type of user interface (for example, a "wearable" circuit device, etc.).

[0096] The processing unit 910 may be a physical processor or a virtual processor, and realizes various processes by a program stored in the memory 920. In a multiprocessor system, a plurality of processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computer device 900. The processing unit 910 is also referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0097] Computer device 900 generally includes various computer storage media. Such media can be any media accessible by computer device 900, such as, for example, volatile and non-volatile media, or removable and non-removable media. Memory 920 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory), or any combination thereof. Storage device 930 can be any removable or non-removable media, such as, for example, media readable by devices such as memory, flash drives, disks, or other media used to store information and / or data and accessible in computer device 900.

[0098] Computer device 900 may further include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in FIG. 9, a disk driver for a removable non-volatile disk for reading and / or writing from / to the removable non-volatile disk, and a disk driver for a removable non-volatile disk for reading and / or writing from / to the removable non-volatile disk can be provided. In this case, each driver can be connected to a bus (not shown) via one or more data media interfaces.

[0099] Communication unit 940 further communicates with other computer devices via a communication medium. Also, the functions of the components within computer device 900 can be realized by a single computing cluster or multiple computing machines communicable via a communication connection. Thus, computer device 900 can utilize a logical connection to one or more other servers, networked personal computers (PCs), or other general network nodes to operate in a networked environment.

[0100] The input device 950 may be one or more of various input devices such as, for example, a mouse, a keyboard, a trackball, a voice input device, etc. The output device 960 may be one or more of various output devices such as, for example, a display, a speaker, a printer, etc. The computer device 900 may communicate with one or more external devices (not shown) via the communication unit 940. The external devices are, for example, a storage device and a display device. The computer device 900 may, if necessary, communicate with one or more devices with which the user can interact with the computer device 900, or communicate with any device (such as, for example, a network card, a modem, etc.) that enables the computer device 900 to communicate with one or more other computer devices. Such communication is performed via an input / output (I / O) interface (not shown).

[0101] In some embodiments, some or all components of the computer device 900 are not integrated into a single device and may be arranged in a cloud computing architecture. In a cloud computing architecture, the components are provided remotely and can cooperate to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, but the end user does not need to recognize the physical location or arrangement of the system or hardware that provides these services. In each embodiment, cloud computing uses an appropriate protocol to provide services via a wide area network (e.g., the Internet). For example, a cloud computing provider provides an application accessible via a web browser or any other computing component via a wide area network. The software or components of the cloud computing architecture and the corresponding data can be stored on a remote server. The computing sources in a cloud computing environment can be aggregated or distributed in remote data center locations. The cloud computing infrastructure operates as a single access point for the user, but can provide services via a data center. Therefore, the cloud computing architecture can be used to provide the components and functions described herein from a service provider in a remote location. Alternatively, it may be provided by a conventional server or installed on a client terminal device.

[0102] In an embodiment of the present disclosure, the computer device 900 can be used to implement video encoding / decoding. The memory 920 may include one or more video codec modules 925 having one or more program instructions. These modules can be accessed and executed by the processing unit 910 to perform the functions of each embodiment described herein.

[0103] In an exemplary embodiment that performs video encoding, the input device 950 can receive video data as input 970 to be encoded. The video data can be processed, for example, by a video codec module 925 to generate an encoded bitstream. The encoded bitstream can be provided as output 980 via an output device 960.

[0104] In an exemplary embodiment that performs video decoding, the input device 950 can receive an encoded bitstream as input 970. The encoded bitstream can be processed, for example, by a video codec module 925 to generate decoded video data. The decoded video data can be provided as output 980 via an output device 960.

[0105] Although the present disclosure has been shown and described in detail with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and details can be made therein without departing from the spirit and scope of the present application as defined by the appended claims. These changes are intended to be covered by the scope of the present application. Accordingly, the foregoing description of embodiments of the present application is not intended to be limiting.

Claims

1. A method for processing video, comprising: performing a conversion between a media file of a first video and a bitstream of the first video, wherein the media file includes a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video, and the media file is based on an International Organization for Standardization base media file format (ISO BMFF).

2. The method according to claim 1, wherein a spatial resolution of the second video is lower than a spatial resolution of the first video.

3. The method according to claim 1, wherein the first indication includes a 1-bit flag.

4. When the flag is a first value, the first group of encoded / decoded video data units can be replaced with the second group of encoded / decoded video data units, and when the flag is a second value, the first group of encoded / decoded video data units cannot be replaced with the second group of encoded / decoded video data units. The method according to claim 3.

5. The method according to claim 1, wherein the first indication includes a first track reference indicating a track carrying the bitstream of the second video.

6. When the first track reference has a first type, the first group of encoded / decoded video data units can be replaced with the second group of encoded / decoded video data units, and when the first track reference has a second type, the first group of encoded / decoded video data units cannot be replaced with the second group of encoded / decoded video data units. The method according to claim 5. The method according to claim 5.

7. The method according to claim 6, wherein the first track reference of the first type is a "ppsr" track reference, and the first track reference of the first type is a "ppsn" track reference.

8. The method according to claim 5, wherein the media file of the second video includes a second instruction for indicating whether the encoded / decoded video data units of the first group can be replaced with the encoded / decoded video data units of the second group.

9. The method according to claim 8, wherein the second instruction includes a second track reference indicating a track carrying the bitstream of the first video.

10. When the second track reference has a third type, the encoded / decoded video data units of the first group can be replaced with the encoded / decoded video data units of the second group, and when the second track reference has a fourth type, the encoded / decoded video data units of the first group cannot be replaced with the encoded / decoded video data units of the second group, The method according to claim 9.

11. The method according to claim 10, wherein the second track reference of the third type is a "ppmr" track reference, and the second track reference of the fourth type is a "ppmn" track reference.

12. The encoded / decoded video data units of the first group include video codec layer network abstraction layer (VCL NAL) units, The method according to claim 1, wherein the encoded / decoded video data units of the second group include VCL NAL units.

13. The method according to claim 1, wherein when the encoded / decoded video data units of the first group can be replaced with the encoded / decoded video data units of the second group, the encoded / decoded video data units of the first group are allowed to be replaced with the encoded / decoded video data units of the second group before being sent to a decoder for decoding.

14. For a picture in the first video, the corresponding encoded / decoded video data units of the second video are all the encoded / decoded video data units within the decoded time synchronization samples in the track of the second video. The method according to claim 13.

15. The method according to claim 1, wherein the conversion includes creating the media file and storing the bitstream in the media file.

16. The method according to claim 1, wherein the conversion includes analyzing the media file to reconstruct the bitstream.

17. An apparatus for processing video data, comprising: a processor and a non-transitory memory storing instructions, wherein when the instructions are executed by the processor, the processor is caused to execute the method according to any one of claims 1 to 16.

18. A non-transitory computer-readable storage medium storing instructions for causing a processor to execute the method according to any one of claims 1 to 16.

19. A method for storing a bitstream of a first video, comprising: performing a conversion between the media file of the first video and the bitstream; storing the bitstream in a non-transitory computer-readable recording medium, wherein the media file includes a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video; wherein the media file is based on an International Organization for Standardization base media file format (ISO BMFF).

20. A method for storing a media file of a first video, comprising: performing a conversion between the media file and the bitstream of the first video; storing the media file in a non-transitory computer-readable recording medium, wherein the media file includes a first indication for indicating whether a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video can be replaced with a second group of encoded / decoded video data units related to a second video; wherein the media file is based on an International Organization for Standardization base media file format (ISO BMFF).

Citation Information

Patent Citations

  • Recording medium, playback device, recording device, encoding method, and decoding method related to higher image quality

    WO2012147350A1

  • Method and apparatus for encoding and decoding a video stream with subpictures

    WO2021052794A1

  • Processing of filler data units in video streams

    WO2021142372A1