Method, apparatus, and medium for video processing
The method addresses the lack of efficient picture-in-picture service support in ISOBMFF media files by enabling the replacement of video data units and signaling picture-in-picture area information within the media file, thereby enhancing video processing efficiency.
Patent Information
- Application Number
- JP2024519073
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-27
- Filing Date
- 2022-09-26
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2042-09-26
AI Technical Summary
Current media file formats, such as ISOBMFF, lack a mechanism to efficiently enable picture-in-picture services, particularly in indicating replaceable video data units and signaling the position and size of the target picture-in-picture area.
A method and apparatus for processing video that includes converting a media file of a first video into a bitstream, where the media file contains an indication for a group of encoded/decoded video data units indicating a target picture-in-picture area, and these data units can be replaced with corresponding units from a second video.
This solution enables efficient correspondence to picture-in-picture services in media files based on ISOBMFF by allowing replacement of video data units and signaling the necessary positional and size information, thereby improving transmission and decoding efficiency.
Smart Images

Figure 0007700373000001 
Figure 0007700373000002 
Figure 0007700373000003
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority based on U.S. Provisional Application No. 63 / 248,832 filed on September 27, 2021, and incorporates by reference all the descriptions set forth in the U.S. application into this application.
[0002] Embodiments of the present disclosure mainly relate to video processing technologies, and more specifically, to the design of file formats for picture - in - picture.
Background Art
[0003] Media stream applications generally conform to transmission methods such as Internet Protocol (IP), Transmission Control Protocol (TCP), or Hypertext Transfer Protocol (HTTP), and generally rely on file formats such as ISO - based Media File Format (ISOBMFF). One such streaming system is HTTP - based Dynamic Adaptive Streaming over HTTP (DASH). In DASH, there may be various representations for video and / or audio data of multimedia content, and different representations may correspond to different codec characteristics (e.g., different profiles or levels of video codec standards, different bitrates, different spatial resolutions, etc.). Furthermore, a technology called "picture - in - picture" has also been proposed. Therefore, there is room for consideration of file formats corresponding to picture - in - picture services.
Summary of the Invention
[0004] Embodiments of the present disclosure provide a technical solution for processing video.
[0005] According to a first aspect, a method for processing video is provided. The method includes performing a conversion between a media file of a first video and a bitstream of the first video. The media file includes a first indication for indicating a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video. The first group of encoded / decoded video data units is replaceable with a second group of encoded / decoded video data units related to a second video.
[0006] According to the above method, the indication is adopted to indicate a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video. The first group of encoded / decoded video data units is replaceable with a second group of encoded / decoded video data units related to a second video. Thereby, the proposed method can easily enable correspondence to a picture-in-picture service in a media file based on ISOBMFF.
[0007] According to a second aspect, an apparatus for processing video data is provided. The apparatus for processing the video data includes a processor and a non-transitory memory storing instructions. When the instructions are executed by the processor, the method according to the first aspect of the present disclosure is caused to be executed by the processor.
[0008] According to a third aspect, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium causes a processor to execute instructions of the method according to the first aspect of the present disclosure.
[0009] According to the fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed in a video processing apparatus. The method includes performing a conversion between a media file of a first video and the bitstream of the first video, where the media file includes a first instruction for indicating a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video. The first group of encoded / decoded video data units is replaceable with a second group of encoded / decoded video data units related to a second video.
[0010] According to the fifth aspect, a method for storing a bitstream of a first video is provided. The method includes performing a conversion between a media file of the video and the bitstream of the first video, where the media file includes a first instruction for indicating a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video. The first group of encoded / decoded video data units is replaceable with a second group of encoded / decoded video data units related to a second video.
[0011] According to the sixth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a media file of a first video generated by a method executed in a video processing apparatus. The method includes performing a conversion between the media file of the first video and the bitstream of the first video, and the media file includes a first instruction for indicating a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video. The first group of encoded / decoded video data units is replaceable with a second group of encoded / decoded video data units related to a second video.
[0012] According to the seventh aspect, a method for storing a media file of a first video is provided. The method includes performing a conversion between the media file of the first video and the bitstream of the first video, and storing the media file in a non-transitory computer-readable recording medium. The media file includes a first instruction for instructing a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video, and the first group of encoded / decoded video data units can be replaced with a second group of encoded / decoded video data units related to a second video.
[0013] The description of the summary of the invention is for explaining the selection of concepts in a simplified form, and these will be described in detail in the following embodiments for carrying out the invention. The summary of the invention is not intended to identify the important features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure.
Brief Description of the Drawings
[0014] According to the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Embodiments for Carrying Out the Invention
[0015] Next, the principles of the present disclosure will be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of explanation and to assist those skilled in the art in understanding and implementing the present disclosure, and do not mean to limit the scope of the present disclosure. The disclosure content described in this specification can be implemented in various ways in addition to those described below.
[0016] In the following description and claims, unless otherwise specified, the meanings of all scientific and technical terms used in this specification are the same as those generally understood by those skilled in the technical field to which the present disclosure belongs.
[0017] "One embodiment", "an embodiment", "an exemplary embodiment", etc. in the present disclosure indicate that the described embodiments may include specific features, structures, or characteristics, but not all embodiments must include the above specific features, structures, or characteristics. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when describing specific features, structures, or characteristics with reference to exemplary embodiments, it is considered within the knowledge of those skilled in the art that, whether explicitly described or not, the features, structures, or characteristics of other embodiments may be affected.
[0018] Terms such as "first" and "second" can be used to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiment, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. The term "and / or" used herein includes any and all combinations of one or more of the listed terms.
[0019] As used herein, the terms are used only for the purpose of describing a specific embodiment and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless clearly stated otherwise in the context. Also, the terms "comprising", "including", and / or "having" used herein indicate the presence of the described features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0020] Environment of the Embodiment FIG. 1 illustrates a block diagram of an exemplary video codec system 100 to which the techniques of the present disclosure are applicable. Thus, video codec system 100 may include a source device 110 and a target device 120. Source device 110 is also referred to as a video encoding device, and target device 120 is also referred to as a video decoding device. During operation, source device 110 may be arranged to generate encoded video data, and target device 120 may be arranged to decode the encoded video data generated by source device 110. Source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0021] Video source 112 may include a source such as, for example, a video capture device. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a provider of video content, a computer graphics system for generating video data, and / or combinations thereof.
[0022] Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a series of bits that form the video data. The bitstream may include encoded pictures and data associated therewith. An encoded picture is a representation of a picture encoded. The data associated therewith may include a sequence parameter set, a picture parameter set, and other syntax elements (syntax structure). I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be transmitted directly to target device 120 via I / O interface 116 by network 130A. The encoded video data may be stored in storage medium / server 130B for access by target device 120.
[0023] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a demodulator. The I / O interface 126 can obtain encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 can decode the encoded video data. The display device 122 can display the decoded video data to the user. The display device 122 may be integrally configured with the target device 120, or an external display device interface may be arranged outside the target device 120 to be connectable to this target device 120.
[0024] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other conventional and / or further standards.
[0025] FIG. 2 is an exemplary block diagram showing a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be an example of the video encoder 114 of the system 100 shown in FIG. 1.
[0026] The video encoder 200 can be configured to implement any or all of the techniques of the present disclosure. In the example of FIG. 2, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the components of the video encoder 200. In some examples, the processor may be configured to execute any or all of the techniques described in the present disclosure.
[0027] In some embodiments, the video encoder 200 may include a splitting unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding / decoding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0028] In another example, the video encoder 200 may include more, fewer, or different functional components. In one embodiment, the prediction unit 202 may include a block-in-copy (IBC) unit. The IBC unit can perform prediction in the IBC mode. At least one reference picture is the picture in which the current video block is located.
[0029] Furthermore, some components (e.g., the motion estimation unit 204 and the motion compensation unit 205) can be integrated, but for ease of interpretation, these components are shown separately in the example of FIG. 2.
[0030] The splitting unit 201 can split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can accommodate various video block sizes.
[0031] The mode selection unit 203 selects, for example, one of various codec modes (intra-coding or inter-coding) based on an error result, provides the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provides the residual block data to the reconstruction unit 212 to reconstruct the coded block as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction (CIIP) modes to perform prediction based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution (e.g., sub-pixel accuracy or integer-pixel accuracy) for the motion vector for the block.
[0032] To perform inter-frame prediction for the current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 can identify a predicted video block for the current video block according to the motion information and the decoded samples of the pictures excluding the picture associated with the current video block from the buffer 213.
[0033] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block according to whether the current video block is located in an I slice, a P slice, or a B slice. For example, as used herein, an "I slice" may be a part of a picture composed of macroblocks, and all macroblocks are derived from macroblocks in the same picture. Further, as used herein, in some aspects, a "P slice" and a "B slice" may be parts of a picture composed of macroblocks different from the macroblocks in the same picture.
[0034] In some examples, the motion estimation unit 204 can perform uni - directional prediction for the current video block. The motion estimation unit 204 can search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Then, the motion estimation unit 204 can generate a reference index indicating the reference picture including the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction flag, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block of the current video block according to the reference video block indicated by the motion information of the current video block.
[0035] Also, in another example, the motion estimation unit 204 can perform bi - directional prediction for the current video block. The motion estimation unit 204 can search the reference pictures in list 0 to find a reference video block for the current video block, and further search the reference pictures in list 1 to find another reference video block for the current video block. Then, the motion estimation unit 204 can generate a plurality of reference indexes indicating a plurality of reference pictures including a plurality of reference video blocks in list 0 and list 1, and a plurality of motion vectors indicating a plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 can output the plurality of reference indexes and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block according to the plurality of reference video blocks indicated by the motion information of the current video block.
[0036] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can transmit the motion information of the current video block in a signal by referring to the motion information of other video blocks. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0037] In one embodiment, the motion estimation unit 204 can instruct the video decoder 300 with one value in the syntax structure related to the current video block, and this value indicates that the current video block has the same motion information as other video blocks.
[0038] In another example, the motion estimation unit 204 can identify a motion vector difference (MVD) with other video blocks in the syntax structure related to the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0039] As described above, the video encoder 200 can transmit the motion vector in a signal in a predictive manner. Two examples of predictive signaling techniques that the video encoder 200 can implement include advanced motion vector prediction (AMVP) and merge mode signaling.
[0040] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and each syntax element.
[0041] The residual generation unit 207 can generate residual data for the current video block by subtracting the (multiple) predicted video blocks of the current video block from the current video block (for example, indicated by a minus sign). The residual data of the current video block may include residual video blocks corresponding to different sample portions of the samples of the current video block.
[0042] In another example, for example, in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0043] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0044] After the transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, the quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0045] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transform coefficient video block to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block related to the current video block and store it in the buffer 213.
[0046] After reconstructing the video block, the reconstruction unit 212 can perform loop filter processing to reduce video block effect artifacts in the video block.
[0047] The entropy encoding / decoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding / decoding unit 214 receives data, the entropy encoding / decoding unit 214 can perform one or more entropy encoding processes to generate entropy encoding / decoding data and output a bitstream including the entropy encoding / decoding data.
[0048] FIG. 3 shows an exemplary block diagram of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be an example of the video decoder 124 in the system 100 shown in FIG. 1.
[0049] The video decoder 300 may be configured to execute any or all of the techniques of the present disclosure. In the example of FIG. 3, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the components of the video decoder 300. In some examples, the processor may be arranged to execute any or all of the techniques described in the present disclosure.
[0050] In the example of FIG. 3, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can generally perform decoding processes for the encoding processes described for video encoder 200.
[0051] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., an encoded video data block). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can identify motion information including a motion vector, motion vector precision, a reference picture list index, and other motion information from the entropy-decoded video data. Motion compensation unit 302 can identify the information, for example, by performing AMVP and merge mode. AMVP is used including obtaining several most likely candidates from data of neighboring PBs and reference pictures. Motion information generally includes horizontal and vertical motion vector displacement values and one or two reference picture indexes, and in the case of a prediction area in a B slice, further includes identification of which reference picture list is associated with each index. As used herein, in some aspects, "merge mode" may refer to deriving motion information from spatially or temporally adjacent blocks.
[0052] Motion compensation unit 302 can generate a motion compensation block to perform interpolation by an interpolation filter in some cases. An identifier of the interpolation filter used with sub-pixel precision may be included in the syntax element.
[0053] During the encoding of a video block, the motion compensation unit 302 can calculate interpolation values for sub-pixel of a reference block by using an interpolation filter used by the video encoder 200. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can generate a prediction block by using the interpolation filter.
[0054] The motion compensation unit 302 can use at least some of the syntax information to determine the size of the blocks of one or more frames and / or one or more slices for encoding an encoded video sequence, the partitioning information explaining how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may be a data configuration that can be decoded independently from other slices of the same picture for entropy encoding, signal prediction, and residual signal reconstruction. A slice may be the entire picture or a region of the picture.
[0055] The intra prediction unit 303 can form a prediction block from spatially adjacent blocks, for example, by using an intra prediction pattern received in the bitstream. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients decoded by the entropy decoding unit 301 provided in the bitstream. The inverse transform unit 305 applies an inverse transform.
[0056] The reconstruction unit 306 can obtain a decoded block, for example, by adding a residual block and a corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, in order to remove block effect artifacts, a deblocking effect filter may be applied to filter the decoded block. Then, the decoded video block is stored in the buffer 307, and the buffer 307 provides a reference block for subsequent motion compensation / intra prediction, and the buffer 307 further generates a decoded video for presentation on a display device.
[0057] Hereinafter, some exemplary embodiments of the present disclosure will be described in detail. It should be noted that the chapter headings in this specification are used for ease of understanding and do not limit the embodiments disclosed in a certain chapter to that chapter. Further, although some embodiments have been described with reference to a general-purpose video codec or other specific video codecs, the disclosed technology is also applicable to other video codec technologies. Further, in some embodiments, the video encoding steps have been described in detail, but it should be understood that the corresponding decoding steps for canceling the encoding are realized by a decoder. Further, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding that represents video pixels from one compression format to another compression format or at another compression bit rate. 1. Overview The present disclosure relates to a video file format. Specifically, it relates to picture-in-picture correspondence in a media file. These concepts can be applied individually or in various combinations to media file formats, such as ISO-based media file formats (ISOBMFF) or extensions thereof. 2. Background Art 2.1 About Video Codec Standards Video codec standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture using temporal prediction and transform coding. To explore future video coding technologies other than HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new approaches and incorporated them into a reference software called the Joint Exploration Model (JEM). With the official launch of the Versatile Video Coding (VVC) project, JVET changed its name to the Joint Video Experts Team (JVET). VVC is a new codec standard that aims to reduce the bitrate by 50% compared to HEVC, which was finalized at the 19th meeting of JVET on July 1, 2020. The Versatile Video Coding (VVC) standard (ITU-T H.266|ISO / IEC 23090-3) and the related Multifunctional Supplemental Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IEC 23002-7) are designed to be used in the broadest range of applications, including both traditional applications (such as television broadcasting, video conferencing, or playback from storage media) and newer, more advanced applications (such as adaptive bitrate streaming, video region extraction, and combinations and merges of multiple encoded / decoded video bitstreams, multi-view video, scalable layer codec, and viewport-adaptive 360° immersive media content). The Elementary Video Coding (EVC) standard (ISO / IEC 23094-1) is another video codec standard recently developed by MPEG. 2.2 File Format Standards Media stream applications typically rely on IP, TCP, and HTTP transmission methods and usually depend on file formats such as the ISO Base Media File Format (ISOBMFF). One of these streaming systems is Dynamic Adaptive Streaming over HTTP (DASH). To use video formats with ISOBMFF and DASH, video format-specific file format standards are required to encapsulate video content into ISOBMFF tracks and DASH representations or clips, such as the AVC file format and the HEVC file format. Important information regarding the video bitstream, such as profile, layer, class, and many other information, needs to be disclosed as file format class metadata and / or DASH Media Presentation Description (MPD) for the purpose of content selection, such as the selection of appropriate media clips for both initialization at the time of streaming session disclosure and streaming adaptation during the streaming session period. Similarly, to use the ISOBMFF image format, image format-specific file format standards are required, such as the AVC image file format and the HEVC image file format. The VVC video file format, which is a file format for storing VVC video content based on ISOBMFF, is currently being developed by MPEG. The VVC image file format for storing image content encoded by VVC and based on ISOBMFF is currently being developed by MPEG. 2.3 Picture Partitioning and Sub-pictures in VVC In VVC, a picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs in a rectangular area covering the picture. The CTUs in a tile are scanned in raster scan order within this tile. A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture. Two types of slice modes, namely, raster scan slice mode and rectangular slice mode, are supported. In the raster scan slice mode, a slice contains a complete tile sequence in the tile raster scan of a picture. In the rectangular slice mode, a slice contains a plurality of complete tiles that jointly form a rectangular region of a picture, or a plurality of consecutive complete CTU rows of one tile that jointly form a rectangular region of a picture. Tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to the slice. A sub-picture contains one or more slices that jointly cover a rectangular region of a picture. 2.3.1 Regarding the concept and function of sub-pictures In VVC, for example, as shown in FIG. 4, each sub-picture consists of one or more complete rectangular slices that jointly cover a rectangular region of a picture. A sub-picture may be limited to being extractable (i.e., encoded and decoded independently of other sub-pictures of the same image and other sub-pictures of previous images in the decoding order) or non-extractable. Regardless of whether a sub-picture is extractable or non-extractable, the encoder can control whether to apply loop filtering (including deblocking, SAO, and ALF) across sub-picture boundaries for each sub-picture individually. Functionally, a sub-picture is similar to a motion-constrained tile set (MCTS) in HEVC. Both allow for independent codec operations and extraction of a rectangular subset of an encoded / decoded picture sequence, for application examples such as 360° video stream optimization regarding a viewport and region of interest (ROI) applications. In a 360° video stream (also known as an omnidirectional video), at any given point in time, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) is presented to the user, and the user can always turn his / her head to change the viewing direction and the current viewport. At the same time, for the areas that are at least not covered by the current viewport of the client terminal, it is desirable to have some low-quality representations, and be prepared to be rendered to the user to prevent the user from suddenly changing his / her viewing direction to any position on the spherical surface. At any given point in time, the high-quality representation of the omnidirectional video on the spherical surface is only required for the current viewport that needs to be rendered to the user. Such optimization can be achieved by dividing the high-quality representation of the entire omnidirectional video into sub-pictures with appropriate granularity as shown in Figure 4. Here, Figure 4 has 12 sub-pictures with high resolution located on the left side and the remaining 12 sub-pictures of the omnidirectional video with low resolution located on the right side. Another typical viewport-related 360° video delivery method using sub-pictures, as shown in Figure 5, is that only the high-resolution representation of the complete video is composed of sub-pictures, while the low-resolution representation of the complete video can be encoded and decoded with a lower frequency of RAP without using sub-pictures compared to the high-resolution representation. At the client terminal, the low-resolution complete video is received, but for the high-resolution video, only the sub-pictures covering the current viewport are received and decoded. 2.3.2 Differences between Sub-Pictures and MCTS There are several important design differences between sub-pictures and MCTS. First, similar to the picture boundary, the sub-picture feature of VVC allows the motion vectors of the coding blocks to point outside the sub-picture. Even in this case, the sub-picture can be extracted by applying sample padding at the sub-picture boundary. Second, during the merge mode and the motion refinement period on the VVC decoder side, the selection and export of motion vectors introduce additional variations, enabling higher codec efficiency compared to the non-conventional motion constraints applied on the MCTS encoder side. Third, when extracting one or more extractable sub-pictures from an image sequence to create a sub-bitstream (which is a coherent bitstream), there is no need to rewrite the SH (and the PHNAL unit if it exists). In HEVC MCTS-based sub-bitstream extraction, rewriting of the SH is required. Note that the SPS and PPS need to be rewritten in both HEVC MCTS extraction and VVC sub-picture extraction. However, generally, there are only a few parameter sets in one bitstream, and each picture has at least one slice, so rewriting the SH is a heavy burden on the application system. Fourth, the slices of different sub-pictures within one picture can have different NAL unit types. As will be explained in more detail below, this is generally a feature referred to as the mixed NAL unit type or mixed sub-picture type in a picture. Fifth, since VVC defines the HRD and level definition of the sub-picture sequence, the encoder can ensure the consistency of the sub-bitstream of each extractable sub-picture sequence. 2.3.3 Regarding the Mixed Sub-Picture Type of Pictures In AVC and HEVC, all VCL NAL units in a picture must have the same NAL unit type. VVC introduces the option that a picture can contain sub-pictures with different VCL NAL unit types, supporting random access not only at the picture level but also at the sub-picture level. In VVC, the VCL NAL units in a sub-picture still need to have the same NAL unit type. The function of random access from IRAP sub-pictures is beneficial for 360° video applications. In a delivery method similar to the 360° video that depends on the viewports shown in Figure 5, the content of spatially adjacent viewports overlaps significantly, that is, while the viewport direction changes, only a small part of the sub-pictures in the viewport are replaced with new sub-pictures, but most of the sub-pictures remain in the viewport. The sequence of sub-pictures newly introduced into the viewport must be disclosed from IRAP slices. However, if the remaining sub-pictures are allowed to perform inter-frame prediction during the viewport change, a significant reduction in the overall transmission bitrate can be achieved. The indication of whether a picture contains only a single type of NAL unit or multiple types is provided in the referenced PPS of the picture (i.e., using the flag named pps_mixed_nalu_types_in_pic_flag). A picture can be composed of sub-pictures containing IRAP slices and sub-pictures containing subsequent slices. There can be several other combinations of different NAL unit types in a picture, including the preceding picture slices of NAL unit types RASL and RADL. This allows merging sub-picture sequences with open-GOP and close-GOP codec structures extracted from different bitstreams into one bitstream. 2.3.4 Regarding the Layout and ID Signaling of Sub-Pictures Since the layout of VVC sub - pictures is signaled in the SPS, it is not changed in the CLVS. Since each sub - picture is represented by the position of its top - left - corner CTU and the width and height of a plurality of CTUs, the sub - picture thus ensures a rectangular area covering the picture at the CTU granularity. The index of each sub - picture within a picture is determined in the order of sub - pictures signaled in the SPS. In order to be able to extract and merge the sub - picture sequence without rewriting the SH or PH, the slice addressing method in VVC associates a slice with a sub - picture based on the sub - picture ID and a specific slice index of the sub - picture. In SH, the sub - picture ID of the sub - picture containing the slice and the slice index at the sub - picture level are signaled. Note that the value of the sub - picture ID of a specific sub - picture may be different from the value of its sub - picture index. The mapping between the two is signaled in the SPS or PPS (if neither is present), or is implicitly inferred. If there is a mapping of sub - picture IDs, when rewriting the SPS and PPS during the sub - picture sub - bitstream extraction process, it is necessary to rewrite or add the mapping of sub - picture IDs. The sub - picture ID, together with the slice index at the sub - picture level, instructs the decoder of the exact position of the first decoded CTU of the slice within the DPB time - slit of the decoded picture. After the extraction of the sub - bitstream, the sub - picture ID of the sub - picture does not change, but the sub - picture index may change. Even if the raster - scan CTU address of the first CTU in the slice of the sub - picture changes compared to the value in the original bitstream, the unchanged sub - picture ID and the slice index at the sub - picture level in the corresponding SH correctly determine the position of each CTU in the decoded picture of the extracted sub - bitstream. Figure 6 shows the extraction of sub - pictures using the sub - picture ID, sub - picture index, and slice index at the sub - picture level by an example including two sub - pictures and four slices. Similar to the extraction of sub-pictures, for the signaling of sub-pictures, assuming that different bitstreams are cooperatively generated (for example, different sub-picture IDs are used, but in other aspects, SPS, PPS, and PH parameters such as the CTU size, chroma format, and codec tools are often made consistent), by only rewriting the SPS and PPS, multiple sub-pictures from different bitstreams can be merged into a single bitstream. Sub-pictures and slices are signaled independently in SPS and PPS respectively, but there are inherent mutual constraints between the sub-picture and slice layouts to form a consistent bitstream. First, to have sub-pictures, it is necessary to use rectangular slices and prohibit raster scan slices. Second, the slices of a given sub-picture need to be consecutive NAL units in the decoding order. This means that the sub-picture layout restricts the order of the encoded / decoded slice NAL units in the bitstream. 2.4 Regarding Picture-in-Picture Service The Picture-in-Picture service provides the function of including a low-resolution picture in a high-resolution picture. Such a service allows the user to display two videos simultaneously. Therefore, the video with the higher resolution is used as the main video, and the video with the lower resolution is used as the supplementary video. Such a Picture-in-Picture service is used to provide an accessibility service by supplementing the main video with a signage video. VVC sub-pictures are used for picture-in-picture services by leveraging their extraction and merging characteristics. In such services, the main video is encoded using multiple sub-pictures, and one of the multiple sub-pictures is of the same size as the supplementary video and is located at the exact position where the supplementary video is to be combined with the main video, and is encoded independently and extractably. When a user views a version of a service that includes a supplementary video, as shown in Figure 7, the sub-picture corresponding to the picture-in-picture area of the main video is extracted from the main video bitstream, and the supplementary video bitstream is merged into its position in the main video bitstream. Figure 7 shows an example of picture-in-picture correspondence using VVC sub-pictures. In this case, the pictures of the main video and the supplementary video must share the same video characteristics, particularly the bit depth, sample aspect ratio, size, frame rate, color space, transmission characteristics, and chroma sample position must be the same. The main video bitstream and the supplementary video bitstream do not need to use NAL unit types in each picture. However, for merging, it is necessary that the pictures of the main bitstream and the supplementary bitstream are encoded and decoded in the same order. In the present disclosure, since a merged subpicture is required, the subpicture IDs used in the main video and the supplementary video must not overlap. The supplementary video bitstream, even if it has no other tile or slice partitioning and consists of only one subpicture, needs to signal subpicture information, particularly the subpicture ID and the subpicture ID length, in order to merge the supplementary video bitstream into the main video bitstream. The subpicture ID length of the subpicture syntax element within the slice NAL unit for signaling the supplementary video bitstream must be the same as the subpicture ID length of the subpicture ID within the slice NAL unit for signaling the main video bitstream. Also, in order to simplify the merging of the supplementary video bitstream and the main video bitstream without rewriting the PPS partitioning information, it is beneficial to encode the supplementary video using only one slice and one tile in the corresponding area of the main video. The main video bitstream and the auxiliary video bitstream must signal the same codec tools used in the SPS, PPS, and picture header. This includes using the same maximum allowable size, minimum allowable size for block partitioning, and the same value as the initial quantization parameter indicated in the PPS (the same value as the pps_init_qp_minus26 syntax element). The use of codec tools can be changed at the slice header level. If the main bitstream and the supplementary bitstream are available within a media file based on ISOBMFF, they can be stored in two separate file format tracks. 3. Problems When supporting picture-in-picture in a media file based on ISOBMFF, the following problems have been pointed out: 1) It is possible to store each of the picture-in-picture main bitstream and the supplementary bitstream using different file format tracks, but there is a lack of a mechanism for indicating such a pair of tracks within a media file based on ISOBMFF. 2) It is possible to realize the picture-in-picture experience using VVC sub-pictures. However, for example, as described above, when it is not possible to replace the coded video data unit indicating the target picture-in-picture area in the main video with the corresponding video data unit of the supplementary video, other decoders and methods may be utilized. Therefore, in a media file based on ISOBMFF, it is necessary to indicate whether such replacement is possible. 3) When the above-described replacement is possible, at the client terminal, if it does not know which of the encoded / decoded video data units in each picture of the main video represents the image area in the target picture, the replacement cannot be performed. Therefore, in this case, it is necessary to signal this information in a media file based on ISOBMFF. 4) For the purpose of content selection and other possible purposes, it is useful to signal the position and size of the target picture-in-picture area in the main video in a media file based on ISOBMFF. 4. Exemplary Embodiments To solve the above-described problems, the following method, which will be generally described, is disclosed. Note that this embodiment should not be understood as being limited as an example for explaining general concepts. Further, the embodiments may be applied individually or in any combination. For convenience, a pair of tracks carrying the main bitstream and the supplementary bitstream that jointly provide the picture-in-picture experience is referred to as a pair of picture-in-picture tracks or a pair of picture-in-picture tracks. 1) To solve the initial problem, a new track reference type is defined such that a track contains a track reference and the track referred to by the track reference is a pair of picture-in-picture tracks. a. In one embodiment, such a new type of track reference is indicated by a track reference type equal to a specific value, e.g., "pips" (meaning "refer to the picture-in-picture supplementary bitstream"), the track containing the track reference carries the main bitstream, and the track referred to by the track reference carries the supplementary bitstream. b. In another example, such a new type of track reference is indicated by a track reference type equal to a specific value, e.g., "pipm" (meaning "refer to the picture-in-picture main bitstream"), the track containing the track reference carries the supplementary bitstream, and the track referred to by the track reference carries the main bitstream. c. In yet another example, two types of track reference types as described above are defined. 2) To solve the first and second problems, two new types of track references are defined for tracks carrying the main bitstream, one of which indicates a pair of picture-in-picture tracks that can replace the encoded video data unit indicating the target picture-in-picture area in the main video with the corresponding video data unit in the supplementary bitstream, and the other indicates a pair of picture-in-picture tracks that do not allow such replacement of the video data unit. a. In one embodiment, these two new types of track references are indicated by track reference type values equal to "ppsr" (meaning "refer to the picture-in-picture supplementary bitstream where the video data unit can be replaced") and "ppsn" (meaning "picture-in-picture supplementary bitstream that does not allow replacement of the video data unit"). 3) Alternatively, to solve the first and second problems, a new type of two types of track references included in the track carrying the supplementary bit stream is defined. One of them indicates a pair of picture-in-picture tracks that can replace the encoded / decoded video data unit indicating the target picture-in-picture area in the main video with the corresponding video data unit of the supplementary bit stream, and the other indicates a pair of picture-in-picture tracks that do not allow such replacement of the video data unit. a. In one embodiment, this new type of two types of track references is indicated by track reference type values equal to "ppmr" (meaning "refer to the picture-in-picture main bit stream where replacement of video data units is possible") and "ppmn" (meaning "refer to the picture-in-picture main bit stream that does not allow replacement of video data units"). 4) Alternatively, to solve the first and second problems, as described in the second and third items above, a new type of the four types of track references is defined. 5) To solve the four problems described above, a new type of entity grouping is defined. This will be described in detail below. a. The new type of entity grouping is named picture-in-picture entity grouping, and its grouping_type is equal to "pinp" (or a different name or different grouping type value, but having a similar function as described below). b. In one embodiment, it is stipulated that each entity within the entity group must be a video track. c. PicInPicEntityGroupBox is defined by extending EntityToGroupBox and carries at least one or more of the following information: i. The number of main bitstream tracks is N. The entities (i.e., the tracks here) identified by the first N entity_id values within EntityToGroupBox are the main bitstream tracks, and those identified by other entity_id values within the entity EntityToGroupBox are the supplementary bitstream tracks. To play the picture-in-picture experience, one main bitstream track in the main bitstream tracks is selected, and one supplementary bitstream track in the supplementary bitstream tracks is selected. 1. Alternatively, the main bitstream tracks are signaled to the list of Entity_id values within EntityToGroupBox via an index list, and the other entities / tracks within the entity group are the supplementary bitstream tracks. 2. Alternatively, the main bitstream tracks are signaled via a list of track_id values, and the other entities / tracks within the entity group are the supplementary bitstream tracks. ii. Regarding the indication for whether the encoded / decoded video data unit representing the target picture-in-picture area in the main video can be replaced with the corresponding video data unit in the supplementary video. 1. In one embodiment, the indication is signaled by a 1-bit flag called data_units_replacable, and the values 1 and 0 indicate enabling and disabling the replacement of this video data unit, respectively. iii. Regarding the list of region IDs for indicating which of the encoded / decoded video data units in each picture of the main video indicates the target picture-in-picture area. 1. In one embodiment, it is stipulated that for a specific video decoder, it is necessary to explicitly define the specific semantics of the region ID. a. In one embodiment, it is defined as follows: In the case of VVC, the region ID is the sub-picture ID, and the encoded / decoded video data unit is the VCL NAL unit. The VCL NAL unit indicating the target picture in-picture region in the main video is the VCL NAL unit having these sub-picture IDs, and these sub-picture IDs are the same as the sub-picture IDs in the corresponding VCL NAL unit of the supplementary video (generally, all VCL NAL units of one frame in the supplementary video share the same sub-picture ID that is clearly signaled, and in this case, there is only one region ID in the list of region IDs). b. In one embodiment, it is defined as follows: In the case of VVC, before being sent to the video decoder, on the client terminal, if it is selected to replace the encoded / decoded video data unit (i.e., the VCL NAL unit) indicating the target picture in-picture region in the main video using the corresponding VCL NAL unit of the supplementary video, then for each sub-picture ID, without changing the order of the corresponding VCL NAL units, the VCL NAL unit in the main video is replaced with the corresponding VCL NAL unit having the sub-picture ID in the supplementary video. iv. The position and size in the main video for embedding / overlaying the supplementary video are smaller than the main video in terms of size. 1. In one embodiment, it is signaled with four values (x, y, width, height), where x and y specify the position of the upper left corner of the region, and width and height specify the width and height of the region. The unit may be luminance samples / pixel. 2. In one embodiment, when data_units_replacable is equal to 1 and the position and size information exist, it is defined that the position and size should correctly indicate the target picture in-picture region in the main video. 3. In one embodiment, when data_units_replacable is equal to 0 and position and size information exist, the position and size information is defined to indicate a suitable area for embedding and covering the supplementary video (i.e., on the client terminal, it is possible to select to overlay the supplementary video on a different area of the main video). 4. In one embodiment, when data_units_replacable is equal to 0 and position and size information does not exist, there is no information or recommendation on where to cover the supplementary video, and it is defined to completely depend on the client's selection. 5. Embodiment The following are embodiments of some exemplary embodiments generally described in Item 5 and Section 4 above. These embodiments are applicable to ISOBMFF. 5.1 Regarding picture-in-picture entity grouping 5.1.1 Regarding definitions The picture-in-picture service provides a function of including a video with a lower spatial resolution in a video with a higher spatial resolution, which are respectively called the supplementary video and the main video. By selecting one track from the tracks indicated to include the main video and one track (including the supplementary video) from other tracks, the tracks within the same entity group where the grouping_type is equal to "pinp" are used to correspond to the picture-in-picture service. All entities within the picture-in-picture entity group should be video tracks. 5.1.2 Regarding syntax aligned(8) class PicInPicEntityGroupBox extends SEntityToGroupBox("pinp", 0, 0) { unsigned int(8) num_main_video_tracks; unsigned int(1) data_units_replacable; unsigned int(1) pinp_window_info_present; bit(6) reserved = 0; if (data_units_replacable) { unsigned int(8) num_region_ids; for (i = 0; i < num_region_ids; i++) unsigned int(16) region_id[i]; } if (pinp_window_info_present) { unsigned int(16) x; unsigned int(16) y; unsigned int(16) width; unsigned int(16) height; } } 5.1.3 Semantics num_main_video_tracks specifies the number of tracks carrying the picture-in-picture main video within the entity group. data_units_replacable indicates whether the encoded / decoded video data units indicating the target picture-in-picture region in the main video can be replaced with the corresponding video data units of the supplementary video. A value of 1 indicates enabling such replacement of video data units, and a value of 0 indicates disabling such replacement of video data units. When data_units_replacable is equal to 1, before being sent to the video decoder for decoding, the player can choose to replace the encoded / decoded video data unit indicating the target picture in picture region in the main video with the encoded / decoded video data unit corresponding to the supplementary video. In this case, for a specific picture in the main video, the corresponding video data units of the supplementary video are all the encoded / decoded video data units within the decoded time synchronization samples in the supplementary video track. In the case of VVC, before being sent to the video decoder, on the client side, if the player chooses to replace the encoded / decoded video data unit (i.e., VCL NAL unit) indicating the target picture in picture region in the main video with the corresponding VCL NAL unit in the supplementary video, for each sub-picture ID, without changing the order of the corresponding VCL NAL units, the VCL NAL units in the main video can be replaced with the corresponding VCL NAL units having the sub-picture ID in the supplementary video. When pinp_window_info_present is equal to 1, it specifies that the fields x, y, width, and height exist. A value of 0 specifies that these fields do not exist. num_region_ids specifies the number of subsequent fields Region_id[i]. region_id[i] specifies the i-th ID of the encoded / decoded video data unit indicating the target picture in picture region. For a specific video decoder, it is necessary to explicitly define the specific semantics of the region ID. In the case of VVC, the region ID is the sub-picture ID, and the encoded / decoded video data unit is the VCL NAL unit. The VCL NAL unit indicating the target picture in picture region in the main video is the VCL NAL unit having these sub-picture IDs, and these sub-picture IDs are the same as the sub-picture IDs within the corresponding VCL NAL units of the supplementary video. x specifies the horizontal position of an encoded video pixel (sample) located at the upper left corner of the target picture-in-picture area in the main video. The unit is a video pixel (sample). y specifies the vertical position of an encoded video pixel (sample) located at the upper left corner of the target picture-in-picture area in the main video. The unit is a video pixel (sample). Width specifies the width of the target picture-in-picture area in the main video. The unit is a video pixel (sample). Height specifies the height of the target picture-in-picture area in the main video. The unit is a video pixel (sample).
[0058] Embodiments of the present disclosure relate to file format design for corresponding to one type of picture-in-picture. As used herein, a "picture-in-picture (PiP) service" provides a function of including a video with a lower spatial resolution (also referred to as a "supplementary video" or a "PiP video") in a video with a higher spatial resolution (also referred to as the "main video").
[0059] FIG. 8 shows a flowchart of a method 800 for processing video according to some embodiments of the present disclosure. Method 800 can be implemented on a client terminal or a server. As used herein, the term "client terminal" may refer to a part of computer hardware or software that accesses services provided by a server in a client terminal-server model as a computer network. For example, the client terminal may be a smartphone or a tablet. As used herein, the term "server" may refer to a device having computing capabilities. In this case, the client terminal accesses the server via a network. The server may be a physical computer device or a virtual computer device.
[0060] As shown in FIG. 8, method 800 starts at block 802, where a conversion between a media file of a first video and a bitstream of the first video is performed. The media file includes a first indication for indicating a first group of encoded / decoded video data units that indicate a target picture-in-picture area in the first video, and the first group of encoded / decoded video data units can be replaced with a second group of encoded / decoded video data units related to a second video. For example, the first indication includes a list of region identifiers (IDs) that identify regions in the first video. It should be understood that the above examples are for illustrative purposes only. The scope of the present disclosure is not limited thereto.
[0061] According to the proposed method, an indication is adopted to indicate a first group of encoded / decoded video data units that indicate a target picture-in-picture area in the first video. The first group of encoded / decoded video data units can be replaced with a second group of encoded / decoded video data units related to a second video. Thereby, the proposed method can easily enable correspondence to a picture-in-picture service in a media file based on ISOBMFF.
[0062] In some embodiments, the spatial resolution of the second video is lower than that of the first video. That is, the second video is a supplementary video and the first video is a main video.
[0063] In some embodiments, the first indication includes a list of region identifiers (IDs) that identify regions in the first video. In some embodiments, for one region ID in the list of region IDs, a second encoded and decoded video data unit having the region ID in the second group of encoded and decoded video data units replaces a first encoded and decoded video data unit having the region ID in the first group of encoded and decoded video data units. For example, FIG. 9 shows a schematic diagram providing a picture-in-picture. As shown in FIG. 9, the first video may include sub-pictures having sub-picture IDs 00, 01, 02, and 03. For example, if the list of region IDs includes sub-picture ID 00, a second encoded and decoded video data unit having sub-picture ID 00 in the second video 920 can replace the first encoded and decoded video data unit having sub-picture ID 00 in the first video 910.
[0064] In this way, the bitstream of the supplementary video can be merged into the bitstream of the main video. Only the merged bitstream needs to be transmitted or decoded, rather than both the bitstream of the supplementary video and the bitstream of the main video. Thereby, the transmission efficiency and / or decoding efficiency can be advantageously improved.
[0065] In some embodiments, the first video is encoded and decoded using a Versatile Video Coding (VVC) method, and the region IDs in the list of region identifications are sub-picture IDs that identify sub-pictures in the first video. In some embodiments, the first group of encoded and decoded video data units includes video decoder layer network abstraction layer (VCL NAL) units, and the second group of encoded and decoded video data units includes VCL NAL units.
[0066] In some embodiments, the first indication may be included in the data configuration of the media file. For example, the data configuration may be a "pinp" entity group. That is, the data configuration is a new type of entity group, the name of which is the picture-in-picture entity group, and its attribute grouping_type is equal to "pinp". In some embodiments, the entities in the "pinp" entity group are the tracks carrying the bitstream of the first video. It should be understood that the above examples are for illustrative purposes only. The scope of the present disclosure is not limited thereto.
[0067] In some embodiments, the data configuration may further include a second indication for indicating a group of tracks carrying the bitstream of the first video. In one example, the second indication may include a value equal to the number of tracks in the group of tracks. For example, when the number of tracks carrying the bitstream of the first video is N, the indication may be a value N indicating that the tracks identified by the first N entity IDs in the data configuration are the tracks carrying the bitstream of the first video, and the tracks identified by the remaining entity IDs are the tracks carrying the bitstream of the second video. Alternatively, the second indication may include an index list indicating the identification (ID) of the tracks in the group of tracks. In another example, the second indication may include a list of track IDs of the tracks in the group of tracks. It should be understood that the above examples are for illustrative purposes only. The scope of the present disclosure is not limited thereto.
[0068] In some embodiments, the size of the target picture in the picture area may be smaller than the size of the first video. The data configuration may further include position information and size information of the target picture in the picture area. In one embodiment, the position information may indicate the horizontal position and the vertical position of the upper left corner of the target picture in the picture area. Alternatively or additionally, the size information may indicate the width and height of the target picture in the picture area. For example, FIG. 10 shows the position information and size information of the target picture in the picture area 1001. As shown in FIG. 10, the position information may indicate the horizontal position X and the vertical position Y of the target picture in the picture area 1001 in the first video 1010. The size information may include the width 1002 and height 1003 of the target picture in the picture area 1001.
[0069] In some embodiments, if the media file includes a third indication for instructing that the first group of encoded / decoded video data units cannot be replaced by the second group of encoded / decoded video data units, the first area in the first video can be determined for the second video. The second video can be superimposed on the first video in the first area.
[0070] In some embodiments, the media file may further include position information and size information of the target picture in the picture area. The first area is determined based on the target picture in the picture area. In some embodiments, the position information may indicate the horizontal position and the vertical position of the upper left corner of the target picture in the picture area. The size information may indicate the width and height of the target picture in the picture area.
[0071] In some embodiments, the conversion includes creating a media file and storing the bitstream in the media file. Alternatively or additionally, in some embodiments, the conversion includes parsing the media file and reconstructing the bitstream.
[0072] In some embodiments, the bitstream of the first video can be stored in a non-transitory computer-readable recording medium. The bitstream of the first video can be generated by a method executed in a video processing device. According to the method, a conversion between the media file of the first video and the bitstream of the first video is performed. The media file includes a first instruction for instructing a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video. The first group of encoded / decoded video data units can be replaced with a second group of encoded / decoded video data units related to the second video.
[0073] In some embodiments, a conversion between the media file of the first video and the bitstream is performed. The media file includes a first instruction for instructing a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video. The first group of encoded / decoded video data units can be replaced with a second group of encoded / decoded video data units related to the second video.
[0074] In some embodiments, the media file of the first video can be stored in a non-transitory computer-readable recording medium. The media file of the first video can be generated by a method executed in a video processing device. According to the method, a conversion between the media file of the first video and the bitstream of the first video is performed. The media file includes a first instruction for instructing a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video. The first group of encoded / decoded video data units can be replaced with a second group of encoded / decoded video data units related to the second video.
[0075] In some embodiments, a conversion between a media file of a first video and a bitstream is performed. The media file includes a first indication for indicating a first group of encoded / decoded video data units that indicate a target picture-in-picture area in the first video. The first group of encoded / decoded video data units is replaceable with a second group of encoded / decoded video data units related to a second video.
[0076] Embodiments of the present disclosure can be described based on the following clauses, and the features related to these clauses can be combined in any reasonable form.
[0077] Clause 1. A method comprising the step of performing a conversion between a media file of a first video and a bitstream of the first video, the media file including a first indication for indicating a first group of encoded / decoded video data units that indicate a target picture-in-picture area in the first video, the first group of encoded / decoded video data units being replaceable with a second group of encoded / decoded video data units related to a second video.
[0078] Clause 2. The method according to Clause 1, wherein the spatial resolution of the second video is lower than the spatial resolution of the first video.
[0079] Clause 3. The method according to any one of Clauses 1 to 2, wherein the first indication includes a list of region identifiers (IDs) for identifying regions in the first video.
[0080] Clause 4. For one region ID in the list of region IDs, the method according to Clause 3, further comprising the step of replacing a first encoded / decoded video data unit having the region ID in the first group of encoded / decoded video data units with a second encoded / decoded video data unit having the region ID in the second group of encoded / decoded video data units.
[0081] Clause 5. The first video is encoded and decoded using the Versatile Video Coding (VVC) method. The region ID in the region identification list is the sub-picture ID that identifies the sub-picture in the first video, and is the method described in any one of Clauses 3 to 4.
[0082] Clause 6. The first group of encoded and decoded video data units includes a Video Coding Layer Network Abstraction Layer (VCL NAL) unit. The second group of encoded and decoded video data units includes a VCL NAL unit, and is the method described in any one of Clauses 1 to 5.
[0083] Clause 7. The first indication is included in the data structure of the media file, and is the method described in any one of Clauses 1 to 6.
[0084] Clause 8. The data structure is the "pinp" entity group, and is the method described in Clause 7.
[0085] Clause 9. The entity in the "pinp" entity group is the track that carries the bitstream of the first video, and is the method described in Clause 8.
[0086] Clause 10. The data structure further includes a second indication for indicating a group of tracks that carry the bitstream of the first video, and is the method described in any one of Clauses 7 to 9.
[0087] Clause 11. The second indication includes one of a value equal to the number of tracks in a group of tracks, an index list indicating the identification (ID) of the tracks in a group of tracks, and a list of track IDs of the tracks in a group of tracks, and is the method described in Clause 10.
[0088] Clause 12. The size of the target picture in-picture region is smaller than the size of the first video, and the data structure further includes the position information and size information of the target picture in-picture region, and is the method described in any one of Clauses 7 to 11.
[0089] Clause 13. The method described in Clause 12, where the position information indicates the horizontal and vertical positions of the upper left corner of the target picture-in-picture area, and the size information indicates the width and height of the target picture-in-picture area.
[0090] Clause 14. The method described in any one of Clauses 1 to 2, where when the media file includes a third instruction indicating that the first group of encoded / decoded video data units cannot be replaced by the second group of encoded / decoded video data units, further including the step of determining a first area in the first video for the second video, and the step of superimposing the second video on the first video in the first area.
[0091] Clause 15. The method described in Clause 14, where the media file further includes position information and size information of the target picture-in-picture area, and determining the first area includes the step of determining the first area based on the target picture-in-picture area.
[0092] Clause 16. The method described in Clause 15, where the position information indicates the horizontal and vertical positions of the upper left corner of the target picture-in-picture area, and the size information indicates the width and height of the target picture-in-picture area.
[0093] Clause 17. The method described in any one of Clauses 1 to 16, where the conversion includes creating a media file and storing the bitstream in the media file.
[0094] Clause 18. The method described in any one of Clauses 1 to 16, where the conversion includes analyzing the media file and reconstructing the bitstream.
[0095] Clause 19. An apparatus for processing video data, comprising a processor and a non-transitory memory storing instructions, where when the instructions are executed by the processor, the processor is caused to execute the method described in any one of Clauses 1 to 18.
[0096] A non - transitory computer - readable storage medium for storing a command to cause a processor to execute the method according to any one of Items 1 to 18 of Item 20.
[0097] A non - transitory computer - readable recording medium that stores a bitstream of a first video generated by a method executed in a video processing apparatus, the method including the step of performing a conversion between a media file of the first video and the bitstream of the first video, the media file including a first instruction for indicating a first group of encoded / decoded video data units indicating a target picture - in - picture area in the first video, and the first group of encoded / decoded video data units being replaceable with a second group of encoded / decoded video data units related to a second video.
[0098] A method of storing a bitstream of a video, the method including the steps of performing a conversion between a media file of a first video and the bitstream of the first video, and storing the bitstream in a non - transitory computer - readable recording medium, the media file including a first instruction for indicating a first group of encoded / decoded video data units indicating a target picture - in - picture area in the first video, and the first group of encoded / decoded video data units being replaceable with a second group of encoded / decoded video data units related to a second video.
[0099] A non - transitory computer - readable recording medium that stores a media file of a first video generated by a method executed in a video processing device, the method including performing a conversion between the media file and a bitstream of the first video, the media file including a first instruction for indicating a first group of encoded / decoded video data units indicating a target picture - in - picture area in the first video, and the first group of encoded / decoded video data units being replaceable with a second group of encoded / decoded video data units related to a second video.
[0100] A method for storing a media file of a first video, the method including performing a conversion between the media file and a bitstream of the first video, and storing the media file in a non - transitory computer - readable recording medium, the media file including a first instruction for indicating a first group of encoded / decoded video data units indicating a target picture - in - picture area in the first video, and the first group of encoded / decoded video data units being replaceable with a second group of encoded / decoded video data units related to a second video.
[0101] Example of the Device FIG. 11 shows a block diagram of a computer device 1100 capable of implementing each embodiment of the present disclosure. The computer device 1100 may be configured as a source device 110 (or video encoder 114 or 200) or a target device 120 (or video decoder 124 or 300), or may be included in the source device 110 (or video encoder 114 or 200) or the target device 120 (or video decoder 124 or 300).
[0102] It should be understood that the computer device 1100 shown in FIG. 11 is for illustrative purposes only and does not imply any limitation on the functions and scope of the embodiments of the present disclosure.
[0103] As shown in FIG. 11, the computer device 1100 includes a general-purpose computer device 1100. The computer device 1100 may include at least one or a plurality of processors or processing units 1110, a memory 1120, a storage device 1130, one or a plurality of communication units 1140, one or a plurality of input devices 1150, and one or a plurality of output devices 1160.
[0104] In some embodiments, the computer device 1100 may be configured as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a mainframe computer device, or the like. The user terminal may be, for example, a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a locator device, a television receiver, a radio broadcast receiver, an electronic paper device, a game device, or any combination thereof, as well as attachments and peripheral devices of these devices or any combination thereof, and may be any type of mobile terminal, fixed terminal, or portable terminal. It is assumed that the computer device 1100 can be compatible with any type of user interface (for example, a "wearable" circuit device, etc.).
[0105] The processing unit 1110 may be a physical processor or a virtual processor, and various processes are realized by a program stored in the memory 1120. In a multiprocessor system, a plurality of processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computer device 1100. The processing unit 1110 is also referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0106] Computer device 1100 generally includes various computer storage media. Such media may be any media accessible by computer device 1100, such as, for example, volatile and non-volatile media, or removable and non-removable media. Memory 1120 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory), or any combination thereof. Storage device 1130 may be any removable or non-removable media, such as, for example, media readable by devices such as memory, flash drives, disks, or other media used for storing information and / or data and accessible in computer device 1100.
[0107] Computer device 1100 may further include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in FIG. 11, a disk driver for a removable non-volatile disk for reading and / or writing from / to the removable non-volatile disk, and a disk driver for a removable non-volatile disk for reading and / or writing from / to the removable non-volatile disk can be provided. In this case, each driver can be connected to a bus (not shown) via one or more data media interfaces.
[0108] Communication unit 1140 communicates with other computer devices via a communication medium. Also, the functions of the components within computer device 1100 can be realized by a single computing cluster or multiple computing machines communicable via a communication connection. Thus, computer device 1100 can utilize a logical connection to one or more other servers, networked personal computers (PCs), or other general network nodes to operate in a networked environment.
[0109] The input device 1150 may be one or more of various input devices such as, for example, a mouse, a keyboard, a trackball, a voice input device, etc. The output device 1160 may be one or more of various output devices such as, for example, a display, a speaker, a printer, etc. The computer device 1100 may communicate with one or more external devices (not shown) via the communication unit 1140. The external devices are, for example, a storage device and a display device. The computer device 1100 may, if necessary, communicate with one or more devices with which the user can interact with the computer device 1100, or communicate with any device (such as, for example, a network card, a modem, etc.) that enables the computer device 1100 to communicate with one or more other computer devices. Such communication is performed via an input / output (I / O) interface (not shown).
[0110] In some embodiments, some or all components of computer device 1100 are not integrated into a single device and may be arranged in a cloud computing architecture. In a cloud computing architecture, components are provided remotely and can cooperate to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, but the end user does not need to be aware of the physical location or arrangement of the system or hardware that provides these services. In each embodiment, cloud computing uses an appropriate protocol to provide services over a wide area network (e.g., the Internet). For example, a cloud computing provider provides an application accessible via a web browser or any other computing component over a wide area network. The software or components of the cloud computing architecture and the corresponding data can be stored on a remote server. Computing resources in a cloud computing environment can be aggregated or distributed across remote data center locations. The cloud computing infrastructure operates as a single access point for the user, but can provide services via a data center. Thus, a cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, it may be provided by a conventional server or installed on a client terminal device.
[0111] In embodiments of the present disclosure, computer device 1100 can be used to implement video encoding / decoding. Memory 1120 may include one or more video codec modules 1125 having one or more program instructions. These modules can be accessed and executed by processing unit 1110 to perform the functions of each embodiment described herein.
[0112] In an exemplary embodiment of performing video encoding, the input device 1150 can receive video data as input 1170 to be encoded. The video data can be processed, for example, by the video codec module 1125 to generate an encoded bitstream. The encoded bitstream can be provided as output 1180 via the output device 1160.
[0113] In an exemplary embodiment of performing video decoding, the input device 1150 can receive an encoded bitstream as input 1170. The encoded bitstream can be processed, for example, by the video codec module 1125 to generate decoded video data. The decoded video data can be provided as output 1180 via the output device 1160.
[0114] Although the present disclosure has been shown and described in detail with reference to the preferred embodiments thereof, those skilled in the art will understand that various changes in form and details can be made thereto without departing from the spirit and scope of the present application as defined by the appended claims. These changes are intended to be covered by the scope of the present application. Therefore, the above description of the embodiments of the present application is not intended to be limiting.
Claims
1. A video processing method, comprising: executing a conversion between a media file of a first video and a bitstream of the first video, wherein the media file includes a first indication for indicating a first group of encoded / decoded video data units that indicate a target picture-in-picture area in the first video, and the first group of encoded / decoded video data units can be replaced with a second group of encoded / decoded video data units related to a second video, wherein the media file is based on an International Organization for Standardization Base Media File Format (ISO BMFF).
2. The method according to claim 1, wherein a spatial resolution of the second video is lower than a spatial resolution of the first video.
3. The method according to claim 1, wherein the first indication includes a list of region identifications (IDs) that identify regions in the first video.
4. The method further includes: for one region ID in the list of region IDs, replacing a first encoded / decoded video data unit having the region ID in the first group of encoded / decoded video data units with a second encoded / decoded video data unit having the region ID in the second group of encoded / decoded video data units.
5. The first video is encoded / decoded by a Versatile Video Coding (VVC) method, and the region ID in the list of region identifications is a sub-picture ID that identifies a sub-picture in the first video.
6. The first group of encoded / decoded video data units includes Video Coding Layer Network Abstraction Layer (VCL NAL) units, and the second group of encoded / decoded video data units includes VCL NAL units.
7. The first indication is included in a data structure of the media file.
8. The data structure is a "pinp" entity group.
9. The method according to claim 8, wherein the entity in the "pinp" entity group is a track carrying the bitstream of the first video.
10. The method according to claim 7, wherein the data configuration further includes a second instruction for instructing a group of tracks carrying the bitstream of the first video.
11. The second instruction is a value equal to the number of tracks in the group of tracks, an index list indicating the identification (ID) of the tracks in the group of tracks, The method according to claim 10, including one of a list of track IDs of the tracks in the group of tracks.
12. The method according to claim 7, wherein the size of the target picture in-picture area is smaller than the size of the first video, and the data configuration further includes position information and size information of the target picture in-picture area.
13. The position information indicates the horizontal position and the vertical position of the upper left corner of the target picture in-picture area, The method according to claim 12, wherein the size information indicates the width and the height of the target picture in-picture area.
14. The method is When the media file includes a third instruction for instructing that the encoded / decoded video data units of the first group cannot be replaced by the encoded / decoded video data units of the second group, determining a first area in the first video for the second video; The method according to claim 1, further including superimposing the second video on the first video in the first area.
15. The media file further includes position information and size information of the target picture in-picture area, The method according to claim 14, wherein determining the first area includes determining the first area based on the target picture in-picture area.
16. The position information indicates the horizontal position and the vertical position of the upper left corner of the target picture in-picture area, The method according to claim 15, wherein the size information indicates the width and the height of the target picture in-picture area.
17. The method according to claim 1, wherein the encoded / decoded video data units of the first group are allowed to be replaced by the encoded / decoded video data units of the second group before being sent to a decoder for decoding.
18. The method according to claim 17, wherein for a picture in the first video, the corresponding encoded / decoded video data units of the second video are all the encoded / decoded video data units within a decoded time synchronization sample in a track of the second video.
19. The method according to claim 1, wherein the conversion includes creating the media file and storing the bitstream in the media file.
20. The method according to claim 1, wherein the conversion includes analyzing the media file to reconstruct the bitstream.
21. A video data processing apparatus, comprising a processor and a non-transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of claims 1 to 20.
22. A non-transitory computer-readable storage medium storing instructions for causing a processor to execute the method according to any one of claims 1 to 20.
23. A method for storing a bitstream of a video, comprising performing a conversion between a media file of a first video and the bitstream of the first video, and storing the bitstream in a non-transitory computer-readable recording medium, wherein the media file includes a first indication for indicating a first group of encoded / decoded video data units indicating a target picture in-picture area in the first video, and the first group of encoded / decoded video data units are replaceable with a second group of encoded / decoded video data units related to a second video, and the media file is based on an International Organization for Standardization base media file format (ISO BMFF).
24. A method for storing a media file of a first video, comprising performing a conversion between the media file and the bitstream of the first video, and including the step of storing the media file in a non-transitory computer-readable recording medium, the media file includes a first instruction for instructing a first group of encoded / decoded video data units indicating a target picture-in-picture area in the first video, and the first group of encoded / decoded video data units is replaceable with a second group of encoded / decoded video data units related to a second video, the media file is based on an International Organization for Standardization base media file format (ISO BMFF), method.
Citation Information
Patent Citations
Recording medium, playback device, recording device, encoding method, and decoding method related to higher image quality
WO2012147350A1
Method and apparatus for encoding and decoding a video stream with subpictures
WO2021052794A1
Processing of filler data units in video streams
WO2021142372A1