Enhancing picture-in-picture signaling in media files
By introducing the ‘pipm’ track reference in the video track, the definition of PiP video and main video is clarified, and the conversion of visual media data and bitstream is performed, multiple problems in the design of picture-in-picture service in the prior art are solved, and a more flexible and compatible video storage and usage methods are realized.
Patent Information
- Application Number
- CN202380068913.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-26
- Filing Date
- 2023-09-25
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, there are multiple problems in the design of picture-in-picture services in NALFF, including unclear definitions of PiP videos and main videos, unsupported VVC video storage methods, and incompatible multi-track storage of HEVC or L-HEVC videos.
By specifying a video track containing a reference to the 'pipm track, it is indicated that the video in the track or the video in the backup group can be used as PiP video, and the corresponding primary video can be included in the referenced track or the backup group to which it belongs. The conversion between the visual media data and the bitstream is performed based on the 'pipm' track reference.
It solves the problem of unclear definition of PiP video and main video, supports the storage and use of VVC videos in multiple tracks, and is compatible with the multi-track storage method of HEVC or L-HEVC video, improving the flexibility and compatibility of picture-in-picture services.
Smart Images

Figure CN119968853A_ABST
Abstract
Description
[0001] Cross-references to related patent applications
[0002] This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 409,952, filed on September 26, 2022, the teachings and disclosures of the foregoing patent application are hereby incorporated by reference in their entirety. Technical Field
[0003] This patent document relates to the generation, storage and use of digital audio and video media information in file format. Background Art
[0004] Digital video consumes the largest amount of bandwidth used on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth requirements for digital video usage are likely to continue to grow. Summary of the invention
[0005] A first aspect relates to a method for processing video data, comprising: determining that a reference track contains a 'supm' track reference, the 'supm' track reference indicating that a video in the reference track can be used as a supplementary video and indicating that a corresponding main video can be contained in the referenced track; and performing conversion between visual media data and a bitstream based on the 'supm' track reference.
[0006] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any of the aforementioned aspects.
[0007] A third aspect relates to a non-transitory computer-readable medium, comprising a computer program product for use with a video codec device, the computer program product comprising computer executable instructions stored on the non-transitory computer-readable medium, so that when the computer program product is executed by a processor, the video codec device performs the method of any of the aforementioned aspects.
[0008] A fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: determining that a reference track contains a 'supm' track reference, the 'supm' track reference indicating that a video in the reference track can be used as a supplementary video and indicating that a corresponding main video can be included in the referenced track; and generating the bitstream based on the determination.
[0009] A fifth aspect relates to a method for storing a bitstream of a video, comprising: determining that a reference track contains a 'supm' track reference, the 'supm' track reference indicating that the video in the reference track can be used as a supplementary video and indicating that the corresponding main video can be included in the referenced track; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] For clarity, any of the foregoing embodiments may be combined with one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0011] These and other features will become more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.
[0013] Figure 1 is a block diagram illustrating an example video processing system.
[0014] Figure 2 is a block diagram of an example video processing device.
[0015] Figure 3 is a flow chart of an example method for video processing.
[0016] Figure 4 is a block diagram illustrating an example video codec system.
[0017] Figure 5 is a block diagram illustrating an example encoder.
[0018] Figure 6 is a block diagram illustrating an example decoder.
[0019] Figure 7 is a schematic diagram of an example encoder.
[0020] Figure 8 is a flow chart of an example method for video processing. DETAILED DESCRIPTION
[0021] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of technologies, whether currently known or yet to be developed. The present disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0022] The section headings used in this document are intended to facilitate understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 terminology in some descriptions is only for ease of understanding and is not intended to limit the scope of the disclosed technology. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, editorial changes to text for the Versatile Video Codec (VVC) specification and / or the International Organization for Standardization (ISO) Base Media File Format (ISOBMFF) standard are shown in bold italics to indicate deleted text and in bold to indicate added text.
[0023] 1. Preliminary Discussion
[0024] This document relates to media file formats. In particular, the present disclosure relates to signaling of picture-in-picture services in media files. For media file formats, these examples can be applied alone or in various combinations, for example, based on ISOBMFF or its extensions, for example, carrying network abstraction layer (NAL) unit structured video in ISOBMFF.
[0025] 2. Further Discussion
[0026] 2.1 Video Codec Standards
[0027] Video codec standards have evolved primarily through the development of standards by the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T developed H.261 and H.263, ISO / IEC developed Moving Picture Experts Group (MPEG)-1 and MPEG-4 Vision, and the two organizations jointly developed the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / High Efficiency Video Codec (HEVC) [1] standards. Since H.262, video codec standards have been based on a hybrid video codec structure that utilizes temporal prediction plus transform coding. Recently, the Versatile Video Codec (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) [3] and the associated Versatile Supplementary Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) [4] have been designed for the widest range of applications, including both traditional uses (such as television broadcasting, video conferencing or playback from storage media), and also newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, content synthesis and merging of video bitstreams from multiple codecs, multi-view video, scalable layered codecs, and viewport adaptive 360° immersive media. The Essential Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard that has been recently developed by MPEG.
[0028] 2.2 File Format Standards
[0029] Media streaming applications are typically based on Internet Protocol (IP), Transmission Control Protocol (TCP) and Hypertext Transfer Protocol (HTTP) transport methods and often rely on file formats such as ISOBMFF [5]. One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH) [6]. In order to use video formats with ISOBMFF and DASH, a video format-specific file format specification is required, also known as the Network Abstraction Layer File Format (NALFF) [7], which includes file format specifications for all NAL unit-based video codecs (such as AVC, HEVC, VVC and their extensions) to encapsulate video content in ISOBMFF tracks and in DASH representations and segments. Important information about the video bitstream (e.g., profile, tier, level, etc.) needs to be exposed as file format-level metadata and / or DASH Media Presentation Description (MPD) to facilitate content selection, e.g., selecting appropriate media segments for initialization at the start of a streaming session and for stream adaptation during a streaming session. Similarly, in order to use an image format with ISOBMFF, a file format specification specific to the image format (such as the AVC image file format and the HEVC image file format in [8]) is required.
[0030] 2.2 Picture-in-Picture Signaling in NALFF
[0031] A picture-in-picture service provides the ability to place a small resolution picture into a larger resolution picture. Such a service may be useful for showing two videos to a user simultaneously, where the larger resolution video is considered the primary video and the smaller resolution video is considered the supplementary video. Such a picture-in-picture service may be used to provide an accessibility service, where the primary video is supplemented by a signage video.
[0032] [9] includes a design for picture-in-picture signaling in NALFF. The design is as follows.
[0033] 4.16 Picture-in-Picture Area Replacement Sample Group
[0034] 4.16.1 Definition
[0035] A picture-in-picture (PiP) service provides the ability to include a video of smaller spatial resolution within a video of larger spatial resolution (referred to as PiP video and primary video, respectively). A video track containing a 'pipm' track reference indicates that the track contains the PiP video, and indicates that the primary video is contained in the referenced track, or in any track in the alternate group to which the referenced track belongs (if any). It should be noted that the 'pipm' track may also be referred to as a supplementary main ('supm') track. Furthermore, in some examples, the PiP video may also be referred to as a supplementary video.
[0036] For each pair of PiP video and main video, using the tools defined in ISO / IEC 14496-12, the window in the main video for embedding / overlaying the PiP video (whose size is smaller than the main video) is indicated by the values of the matrix fields of the TrackHeaderBox of the PiP video track and the main video track, and the value of the layer field of the TrackHeaderBox of the PiP video track should be smaller than the value of the layer field of the TrackHeaderBox of the main video track to overlay the PiP video in front of the main video.
[0037] When PicInPicRegionReplacementEntry is present in the PiP video track, it indicates that the NAL unit representing the target PiP region in the main video can be replaced with the corresponding NAL unit of the PiP video track. In this case, the PiP video and the main video are required to be encoded and decoded using the same video codec. The absence of this sample group indicates that it is not known whether such a replacement is possible.
[0038] When this sample group exists, the player can choose to replace the NAL unit representing the target PiP area in the main video with the corresponding NAL unit of the PiP video before sending it to the video decoder for decoding. In this case, for a specific picture in the main video, the corresponding NAL unit of the PiP video is all the NAL units in the samples synchronized with the decoding time in the PiP video track.
[0039] 4.16.2 Syntax
[0040] class PicInPicRegionReplacementEntry()extends VisualSampleGroupEntry('pprr'){
[0041] bit(5)reserved=0;
[0042] unsigned int(3)region_id_type;
[0043] unsigned int(8)num_region_ids_minus1;
[0044] for(i=0;i<=num_region_ids_minus1;i++)
[0045] unsigned int(16)region_id[i];
[0046] }
[0047] 4.16.3 Semantics
[0048] region_id_type indicates the type of value that region_id takes. If the video codec used for the main video track is VVC (i.e., the sample entry type is 'vvc1', 'vvi1', or 'vvs1'), in which case the video codec used for the PiP video track is also VVC, then region_id_type equal to 0 specifies that the region ID is a VVC sub-picture ID. Otherwise, region_id_type values equal to 0 are reserved. When region_id_type is equal to 1, the region ID is the groupID value in the NAL unit mapping sample group of a NAL unit that can be replaced by the NAL unit of the PiP track. region_id_type values greater than 1 are reserved.
[0049] num_region_ids_minus1 plus 1 specifies the number of region_id[i] fields below.
[0050] region_id[i] specifies the i-th ID of the NAL unit representing the target picture-in-picture region.
[0051] When region_id_type is equal to 1, the main video track must have a 'nalm' sample group with grouping_type_parameter equal to 'pprr', indicating NAL units in the main track that can be replaced by NAL units with the same groupID value in the PiP track.
[0052] When region_id_type is equal to 1 and num_region_ids is equal to 1, the 'nalm' sample group shall not be present in the PiP track, and all NAL units of the PiP track are inferred to have groupID equal to region_id[0].
[0053] When region_id_type is equal to 1 and num_region_ids is greater than 1, the 'nalm' sample group with grouping_type_parameter equal to 'pprr' shall be present in the PiP track and a mapping of groupID values to NAL units shall be provided.
[0054] 3. Technical problems solved by the disclosed technical solutions
[0055] The example design of picture-in-picture signaling in NALFF-based media files has the following problems.
[0056] First, the PiP video is contained in a track that contains a 'pipm' track reference, and that track may belong to an alternate group consisting of multiple video tracks. Video in any of those other video tracks in the same alternate group may also be used as PiP video. However, this is not currently allowed.
[0057] Second, like any track reference, a 'pipm' track reference may reference multiple tracks or groups of tracks unless otherwise permitted. Therefore, the term "referenced track" is unclear.
[0058] Third, since the video contained in any of the referenced tracks referenced by the 'pipm' track or other tracks in the same alternative group may be primary video, the term "primary video" or "primary video track" is unclear without reference to the context.
[0059] Fourth, similarly, if the video in any of those other video tracks in the same alternating group as the PiP video track can also be used as PiP video, then the term "PiP video" or "PiP video track" is unclear without reference to the context.
[0060] Fifth, VVC video can be stored in a file in different ways. One way is to store VVC video in a single track. Another way is to store VVC video in multiple tracks, which consist of a Merge base track and multiple sub-picture tracks referenced by the Merge base track through a 'subp' track reference. However, the current design does not support using VVC video stored in multiple tracks as the main video used in the PiP service.
[0061] Sixth, the semantics of region_id_type contains the following statement: If the video codec for the main video track is VVC (ie, the sample entry type is 'vvc1', 'vvi1' or 'vvs1'). However, the track containing the main video of the PiP service cannot be stored in only one track as a sub-picture track.
[0062] Seventh, in one aspect, the primary video track can be the referenced track referenced by the 'pipm' track or any track in the alternate group to which the referenced track belongs. In the second aspect, the presence of a PicInPicRegionReplacementEntry in the PiP video track indicates that the NAL unit representing the target PiP region in the primary video can be replaced by the corresponding NAL unit of the PiP video track. However, video tracks within the same alternate group may have different spatial resolutions, and when they do have different spatial resolutions, it is difficult for them to be in the same alternate group while being associated with a PiP video track containing a PicInPicRegionReplacementEntry.
[0063] Eighth, HEVC or layered HEVC (L-HEVC) video can be stored in a file in different ways. One way is to store the HEVC or L-HEVC video in a single track. Another way is to store the HEVC or L-HEVC video in multiple tracks, which include a HEVC or L-HEVC slice base track and multiple HEVC or L-HEVC slice tracks, where each slice track has a 'tbas' track reference that references the slice base track. However, the example design does not support using HEVC or L-HEVC video stored in multiple tracks as a main video used in a PiP service.
[0064] 4. List of solutions and implementation examples
[0065] In order to solve the above problems, the method as outlined below is disclosed. These examples should be regarded as examples to explain the general concept and should not be interpreted narrowly. In addition, these examples can be applied individually or in any combination.
[0066] Example 1
[0067] To address the first problem, it is specified that a video track containing a 'pipm' track reference indicates that the video in that track or the video in any track (if any) in the alternate group to which the track belongs can be used as PiP video.
[0068] Furthermore, the corresponding main video may be a video contained in the referenced track, or a video contained in any track (if any) in the backup group to which the referenced track belongs.
[0069] In one example, the corresponding primary video may be a video contained in the referenced track, or a video contained in any track (if any) in the alternate group to which the referenced track belongs.
[0070] In one example, specifying a video track containing a 'pipm' track reference indicates that the video in the track can be used as a PiP video, and the corresponding primary video can be the video contained in the referenced track or in any track (if any) in the alternate group to which the referenced track belongs.
[0071] Example 2
[0072] To address the second issue, specify one of the following, and the referenced track is the track whose track ID is the first entry in the TrackReferenceTypeBox with reference_type equal to 'pipm':
[0073] When present, a TrackReferenceTypeBox with reference_type equal to 'pipm' shall contain only the track identifier and shall not contain any track group identifiers.
[0074] When present, a TrackReferenceTypeBox with reference_type equal to 'pipm' shall contain the track identifier in the first entry, and the other entries (if any) shall contain the track group identifiers.
[0075] When present, the first entry in the TrackReferenceTypeBox with reference_type equal to 'pipm' shall be the track identifier.
[0076] Example 3
[0077] In order to solve the third problem, it is stipulated that for each pair of PiP video and main video, the track containing the PiP video is also called PiP video track.
[0078] Example 4
[0079] To solve the fourth problem, it is specified that for each pair of PiP video and main video, the main video track is a track that is the referenced track referenced by the 'pipm' track or any track in the alternate group to which the referenced track belongs (if any).
[0080] Example 5
[0081] In order to solve the fifth problem, the following aspects are stipulated.
[0082] A video track containing a 'pipm' track reference indicates that the video in the track can be used as a PiP video, and the corresponding primary video can be a video contained at least in the referenced track or in any track (if any) in the alternate group to which the referenced track belongs.
[0083] When the video codec for the main video is VVC, the main video can be contained in a single track or in multiple tracks consisting of a Merge base track and multiple sub-picture tracks referenced by the Merge base track through a 'subp' track reference. In the former case, the main video track is the single track. In the latter case, the main video track is the Merge base track.
[0084] Example 6
[0085] To solve the sixth problem, the following statement in the semantics of region_id_type: If the video codec used for the main video track is VVC (i.e., the sample entry type is 'vvc1', 'vvi1', or 'vvs1') is changed to the following: If the video codec used for the main video track is VVC (i.e., the sample entry type is 'vvc1' or 'vvi1').
[0086] Example 7
[0087] In order to solve the seventh problem, it is stipulated that when PicInPicRegionReplacementEntry exists in the PiP video track, it indicates that the NAL unit representing the target PiP area in the corresponding main video with the same coded picture width and coded picture height as the video contained in the referenced track referenced by the 'pipm' track can be replaced with the corresponding NAL unit of the PiP video.
[0088] Example 8
[0089] In order to solve the eighth problem, the following aspects are stipulated.
[0090] A video track containing a 'pipm' track reference indicates that the video in the track can be used as a PiP video, and the corresponding primary video may be a video contained at least in the referenced track or in any track (if any) in the alternate group to which the referenced track belongs.
[0091] When the video codec used for the main video is HEVC or L-HEVC, the main video can be contained in a single track or multiple tracks consisting of an HEVC or L-HEVC slice base track and multiple HEVC or L-HEVC slice tracks containing a 'tbas' track reference that references the slice base track. In the former case, the main video track is the single track. In the latter case, the main video track is the slice base track.
[0092] 5. Examples
[0093] The following are some example embodiments of the various aspects outlined in Section 5. Most relevant parts that have been added or modified are shown in bold font, while some deleted parts are shown in italic bold font. There may also be some other changes that are editorial in nature and are therefore not highlighted.
[0094] 5.1 First Embodiment
[0095] The following is a first embodiment for Examples 1, 2, 3, 4, 5, 6 and 8b outlined above. The text changes shown are relative to the design of picture-in-picture signalling in NALFF in [9].
[0096] 4.17 Picture-in-Picture Track Reference
[0097] A Picture-in-Picture (PiP) service provides the ability to include a video of smaller spatial resolution within a video of larger spatial resolution (referred to as PiP video and main video, respectively).
[0098] A video track containing a 'pipm' track reference indicates that the track contains PiP video, and the primary video is contained in the referenced track or in any track (if any) in the alternate group to which the referenced track belongs. A video track containing a 'pipm' track reference indicates that the video in the track or in any track (if any) in the alternate group to which the track belongs can be used as PiP video, and the corresponding primary video can be a video contained at least in the referenced track or in any track (if any) in the alternate group to which the referenced track belongs.
[0099] When present, a TrackReferenceTypeBox with reference_type equal to 'pipm' shall contain only the track identifier and shall not contain any track group identifiers.
[0100] For each pair of PiP video and main video, the following applies:
[0101] - A track containing PiP video is also called a PiP video track.
[0102] - The primary video track is the track of the referenced track that is referenced by the 'pipm' track or any track in the alternate group to which the referenced track belongs (if any).
[0103] -When the video codec for the main video is VVC, the main video may be contained in a single track or in multiple tracks consisting of a Merge base track and multiple sub-picture tracks referenced by the Merge base track through a 'subp' track reference. In the former case, the main video track is the single track. In the latter case, the main video track is the Merge base track.
[0104] -The window used to embed / overlay PiP video (whose size is smaller than main video) in the main video is indicated by the values of the matrix fields of the PiP video track and the TrackHeaderBox of the main video track, and the value of the layer field of the TrackHeaderBox of the PiP video track should be smaller than the value of the layer field of the TrackHeaderBox of the main video track to overlay the PiP video in front of the main video.
[0105] When the video codec used for the main video is HEVC or L-HEVC, the main video can be contained in a single track or multiple tracks consisting of an HEVC or L-HEVC slice base track and multiple HEVC or L-HEVC slice tracks containing a 'tbas' track reference that references the slice base track. In the former case, the main video track is the single track. In the latter case, the main video track is the slice base track.
[0106] When the video codec for the main video is VVC, the main video can be contained in a single track or in multiple tracks consisting of a Merge base track and multiple sub-picture tracks referenced by the Merge base track through a 'subp' track reference. In the former case, the main video track is the single track. In the latter case, the main video track is the Merge base track.
[0107] 4.18 Picture-in-Picture Area Replacement Sample Group
[0108] 4.18.1 Definition
[0109] When PicInPicRegionReplacementEntry is present in the PiP video track, it indicates that the NAL units representing the target PiP region in the main video can be replaced with the corresponding NAL units of the PiP video. In this case, the PiP video and the main video are required to be encoded and decoded using the same video codec. The absence of this sample group indicates that it is not known whether such a replacement is possible.
[0110] When this sample group exists, the player can choose to replace the NAL unit representing the target PiP area in the main video with the corresponding NAL unit of the PiP video before sending it to the video decoder for decoding. In this case, for a specific picture in the main video, the corresponding NAL unit of the PiP video is all the NAL units in the samples synchronized with the decoding time in the PiP video track.
[0111] 4.18.2 Syntax
[0112] class PicInPicRegionReplacementEntry()extends VisualSampleGroupEntry('pprr'){
[0113] bit(5)reserved=0;
[0114] unsigned int(3)region_id_type;
[0115] unsigned int(8)num_region_ids_minus1;
[0116] for(i=0;i<=num_region_ids_minus1;i++)
[0117] unsigned int(16)region_id[i];
[0118] }
[0119] 4.18.3 Semantics
[0120] region_id_type indicates the type of value that region_id takes. If the video codec used for the main video track is VVC (i.e., the sample entry type is 'vvc1', 'vvi1', or 'vvs1' 'vvc1' or 'vvi1'), in which case the video codec used for the PiP video track is also VVC, then region_id_type equal to 0 specifies that the region ID is a VVC sub-picture ID. Otherwise, region_id_type values equal to 0 are reserved. When region_id_type is equal to 1, the region ID is the groupID value in the NAL unit mapping sample group of the NAL unit, which can be replaced by the NAL unit of the PiP video track. region_id_type values greater than 1 are reserved.
[0121] num_region_ids_minus1 plus 1 specifies the number of region_id[i] fields below.
[0122] region_id[i] specifies the i-th ID of the NAL unit representing the target picture-in-picture region.
[0123] When region_id_type is equal to 1, the main video track must have a 'nalm' sample group with grouping_type_parameter equal to 'pprr', indicating NAL units in the main video that can be replaced by NAL units with the same groupID value in the PiP video track.
[0124] When region_id_type is equal to 1 and num_region_ids is equal to 1, the 'nalm' sample group shall not be present in the PiP video track, and all NAL units of the PiP video track are inferred to have groupID equal to region_id[0].
[0125] When region_id_type is equal to 1 and num_region_ids is greater than 1, the 'nalm' sample group with grouping_type_parameter equal to 'pprr' must be present in the PiP video track and provide a mapping of groupID values to NAL units.
[0126] 6. References
[0127] [1] ITU-T and ISO / IEC, “High Efficiency Video Codecs”, Rec. ITU-T H.265 | ISO / IEC 23008-2
[0128] (Currently valid version).
[0129] [2] J. Chen, E. Alshina, G. J. Sullivan, J.-R. Ohm, and J. Boyce, “Algorithmic Description of the Joint Exploration Test Model 7 (JEM7),” JVET-G1001, August 2017.
[0130] [3] Rec. ITU-T H.266 | ISO / IEC 23090-3, “Versatile Video Codec”.
[0131] [4] Rec. ITU-T Rec. H.274 | ISO / IEC 23002-7, “Multifunctional supplementary enhancement information message for encoding and decoding video bitstreams”.
[0132] [5] ISO / IEC 14496-12: “Information technology – Codecs of audio and video objects – Part 12: ISO base media file format”.
[0133] [6] ISO / IEC 23009-1: "Information technology – Dynamic adaptive streaming over HTTP (DASH) – Part 1: Media presentation description and fragment format".
[0134] [7] ISO / IEC 14496-15: "Information technology — Coding and decoding of audio and video objects — Part 15: Structured video carried as Network Abstraction Layer (NAL) units in the ISO base media file format."
[0135] [8] ISO / IEC 23008-12: “Information technology – Efficient coding and media transport in heterogeneous environments – Part 12: Image file formats”.
[0136] [9] ISO / IEC 14496-15 Edition 6 CDAM 2, “Text of ISO / IEC 14496-15 Edition 6 CDAM 2 Picture-in-Picture Support and Other Extensions”, MPEG WG 03 Output Document N0652, July 2022.
[0137] Figure 1 4000 is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.
[0138] System 4000 may include a codec component 4004, which may implement various codecs or coding methods described in this document. Codec component 4004 may reduce the average bit rate of the video from the input 4002 to the codec component 4004 to the output, to generate the codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 4004 may be stored or transmitted via a communication connection as represented by component 4006. The stored or transmitted bitstream (or codec) representation of the video received at input 4002 may be used by component 4008 to generate pixel values or displayable video, which is sent to display interface 4010. The process of generating a user-viewable video from a bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used in encoders, and decoders will perform corresponding decoding tools or operations that reverse the encoding results.
[0139] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0140] Figure 2 4106 is a block diagram of an example video processing device 4100. Device 4100 can be used to implement one or more methods described herein. Device 4100 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. Device 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. (Multiple) processors 4102 can be configured to implement one or more methods described in this document. Memory (multiple memories) 4104 can be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 4106 can be used to implement some of the techniques described in this document in hardware circuits. In some embodiments, video processing circuitry 4106 can be at least partially included in processor 4102 (e.g., a graphics coprocessor).
[0141] Figure 3 4200 is a flow chart of an example method 4200 for video processing. The method 4200 determines at step 4202 that a video track contains a 'pipm' track reference indicating that video in the video track or any other track in a spare group including the video track can be used as a PiP video. At step 4204, conversion between visual media data and a bitstream is performed based on the 'pipm' track reference.
[0142] It should be noted that the method 4200 may be implemented in an apparatus for processing video data, the apparatus comprising a processor and a non-transitory memory having instructions thereon, such as the video encoder 4400, the video decoder 4500, and / or the encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform the method 4200. In addition, the method 4200 may be performed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer executable instructions stored on a non-transitory computer-readable medium, such that when executed by the processor, the video codec device performs the method 4200.
[0143] Figure 443 is a block diagram illustrating an example video codec system 4300 that can utilize the techniques of the present disclosure. The video codec system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data, which may be referred to as a video encoding device. The target device 4320 may decode the encoded video data generated by the source device 4310, which may be referred to as a video decoding device.
[0144] Source device 4310 may include video source 4312, video encoder 4314 and input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider and / or a computer graphics system for generating video data, or a combination of these sources. Video data may include one or more pictures. Video encoder 4314 encodes video data from video source 4312 to generate a bitstream. The bitstream may include a series of bits that form a codec representation of video data. The bitstream may include codec pictures and associated data. Codec pictures are coded representations of pictures. Associated data may include sequence parameter sets, picture parameter sets and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly transmitted to target device 4320 via network 4330 via I / O interface 4316. The encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0145] Target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain encoded video data from source device 4310 or storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the target device 4320, or may be external to the target device 4320, and may be configured to be connected to an external display device via an interface.
[0146] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or later standards.
[0147] Figure 5 is a block diagram showing an example of a video encoder 4400, which may be Figure 4Video encoder 4314 in system 4300 shown. Video encoder 4400 can be configured to perform any or all of the techniques of the present disclosure. Video encoder 4400 includes multiple functional components. The techniques described in the present disclosure can be shared between the various components of video encoder 4400. In some examples, the processor can be configured to perform any or all of the techniques described in the present disclosure.
[0148] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 (which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, an intra-frame prediction unit 4406), a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer area 4413 and an entropy coding unit 4414.
[0149] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, in which at least one reference picture is a picture in which the current video block is located.
[0150] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405 , may be highly integrated but are represented separately in the example of the video encoder 4400 for purposes of explanation.
[0151] The segmentation unit 4401 may segment the picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.
[0152] The mode selection unit 4403 may select one of the coding modes (e.g., intra or inter) based on the error result, for example, and provide the resulting intra or inter coded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 may select a combined intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 4403 may also select a resolution of motion vectors for the block in the case of inter prediction (e.g., sub-pixel or integer pixel precision).
[0153] In order to perform inter-frame prediction on the current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing one or more reference frames from the buffer 4413 with the current video block. The motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 4413 other than the picture associated with the current video block.
[0154] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0155] In some examples, the motion estimation unit 4404 may perform unidirectional prediction on the current video block, and the motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures of list 0 or list 1. The motion estimation unit 4404 may then generate a reference index indicating the reference picture containing the reference video block in list 0 or list 1, and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 may output the reference index, the prediction direction indicator, and the motion vector as motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0156] In other examples, the motion estimation unit 4404 may perform bidirectional prediction on the current video block, and the motion estimation unit 4404 may search for a reference video block of the current video block in the reference pictures in list 0, and may also search for another reference video block of the current video block in the reference pictures in list 1. The motion estimation unit 4404 may then generate a reference index and a motion vector, the reference index indicating the reference pictures containing the reference video block in list 0 and list 1, and the motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0157] In some examples, motion estimation unit 4404 may output a complete set of motion information for use in a decoding process by a decoder. In some examples, motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, motion estimation unit 4404 may reference motion information of another video block to signal motion information of the current video block. For example, motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0158] In one example, the motion estimation unit 4404 may indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 4500 that the current video block has the same motion information as another video block.
[0159] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0160] As discussed above, the video encoder 4400 can predictively signal motion vectors.Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0161] The intra prediction unit 4406 may perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0162] The residual generation unit 4407 may generate residual data for the current video block by subtracting the prediction video block(s) of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0163] In other examples, for the current video block, there may be no residual data for the current video block, such as in skip mode, and the residual generation unit 4407 may not perform a subtraction operation.
[0164] The transform processing unit 4408 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0165] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0166] The inverse quantization unit 4410 and the inverse transform unit 4411 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.
[0167] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0168] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives the data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy coded data, and output a bitstream including the entropy coded data.
[0169] Figure 6 is a block diagram showing an example of a video decoder 4500, which may be Figure 4 Video decoder 4324 in the system 4300 shown. Video decoder 4500 can be configured to perform any or all of the techniques of the present disclosure. In the example shown, video decoder 4500 includes multiple functional components. The techniques described in the present disclosure can be shared between the various components of video decoder 4500. In some examples, the processor can be configured to perform any or all of the techniques described in the present disclosure.
[0170] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 may perform a decoding process that is generally substantially the inverse of the encoding process described with reference to the video encoder 4400.
[0171] The entropy decoding unit 4501 may retrieve an encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 4501 may decode the entropy-encoded video data, and from the entropy-decoded video data, the motion compensation unit 4502 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 4502 may determine such information, for example, by performing AMVP and Merge modes.
[0172] The motion compensation unit 4502 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter used with sub-pixel precision may be included in the syntax element.
[0173] The motion compensation unit 4502 may calculate interpolation of sub-integer pixels of a reference block using an interpolation filter used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filter used by the video encoder 4400 according to received syntax information and use the interpolation filter to generate a prediction block.
[0174] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) slices of the encoded video sequence, partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame codec block, and other information used to decode the encoded video sequence.
[0175] The intra prediction unit 4503 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 4504 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.
[0176] The reconstruction unit 4506 may add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. If necessary, a deblocking filter may be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 4507 to provide a reference block for subsequent motion compensation / intra prediction, and also to generate a decoded video for presentation on a display device.
[0177] Figure 74600 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing VVC technology. The encoder 4600 includes three loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602 that uses a predefined filter, the SAO 4604 and the ALF 4606 use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, where the offset and filter coefficients are transmitted by signal using auxiliary information of the codec. The ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool that attempts to capture and repair artifacts produced by previous stages.
[0178] The encoder 4600 also includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive an input video. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction using a reference picture obtained from a reference picture buffer 4612. The residual block from the inter prediction or intra prediction is fed to a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed to an entropy coding component 4618. The entropy coding component 4618 entropy codes and decodes the prediction result and the quantized transform coefficients and sends them to a video decoder (not shown). The quantized component output from the quantization component 4616 can be fed to an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting images to the DF 4602 , SAO 4604 , and ALF 4606 for filtering before the images are stored in the reference picture buffer 4612 .
[0179] Figure 8 4700 is a flow chart of an example method 4700 for video processing. The method 4700 includes determining at step 4702 that a reference track contains a 'supm' track reference, the 'supm' track reference indicating that a video in the reference track can be used as a supplementary video and indicating that a corresponding primary video can be contained in the referenced track. At step 4704, conversion is performed between visual media data and a bitstream based on the 'supm' track reference.
[0180] It should be noted that method 4700 may be implemented in an apparatus for processing video data, the apparatus comprising a processor and a non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4700. In addition, method 4700 may be performed by a non-transitory computer-readable medium, the medium comprising a computer program product for use by a video codec device. The computer program product comprises computer executable instructions stored on a non-transitory computer-readable medium, such that when executed by the processor, the video codec device performs method 4700.
[0181] A list of solutions preferred by some examples is provided next.
[0182] The following solutions illustrate examples of the techniques discussed herein.
[0183] 1. A method for processing media data, comprising: determining that a video track contains a 'pipm' track reference, the 'pipm' track reference indicating that video in the video track or in any other track in an alternate group including the video track can be used as a picture-in-picture (PiP) video; and performing conversion between visual media data and a visual media data file based on the 'pipm' track reference.
[0184] 2. The method according to solution 1, wherein the corresponding main video is a video contained in the referenced track or a video contained in any other track in the backup group including the referenced track.
[0185] 3. A method according to any one of solutions 1 to 2, wherein the video track containing the 'pipm' track reference indicates that the video in the referenced track can be used as a PiP video, and the corresponding primary video can be the video contained in the referenced track or the video contained in any track in the alternative group including the referenced track.
[0186] 4. A method according to any of solutions 1 to 3, wherein, when present, TrackReferenceTypeBox with reference_type equal to 'pipm' shall contain only track identifiers and shall not contain any track group identifiers.
[0187] 5. A method according to any of solutions 1 to 4, wherein, when present, a TrackReferenceTypeBox with reference_type equal to 'pipm' shall contain the track identifier in the first entry and the other entries contain track group identifiers.
[0188] 6. A method according to any of solutions 1 to 5, wherein the first entry in the TrackReferenceTypeBox with reference_type equal to 'pipm' shall be the track identifier.
[0189] 7. A method according to any one of solutions 1 to 6, wherein, for each pair of PiP video and main video, the track containing the PiP video is also called PiP video track.
[0190] 8. The method according to any one of solutions 1 to 7, wherein, for each pair of PiP video and main video, the main video track is the referenced track referenced by the 'pipm' track or any track in the alternate group including the referenced track.
[0191] 9. A method according to any one of solutions 1 to 8, wherein the video track containing the 'pipm' track reference indicates that the video in the video track can be used as a PiP video, and the corresponding primary video can be at least the video contained in the video track or the video in any track (if any) in the alternate group to which the referenced track belongs.
[0192] 10. A method according to any one of solutions 1 to 9, wherein, when the video codec used for the main video is Versatile Video Codec (VVC), the main video can be contained in a single track or in multiple tracks, the multiple tracks including a Merge base track and multiple sub-picture tracks referenced by the Merge base track through a 'subp' track reference, and wherein the main video track is a single track or the Merge base track.
[0193] 11. The method according to any one of solutions 1 to 10, wherein when the video codec for the primary video track is VVC, the sample entry type is 'vvc1' or 'vvi1'.
[0194] 12. A method according to any one of solutions 1 to 11, wherein, when PicInPicRegionReplacementEntry is present in a PiP video track, PicInPicRegionReplacementEntry indicates that a network abstraction layer (NAL) unit representing a target PiP region in a corresponding main video can be replaced with a corresponding NAL unit of the PiP video, the corresponding main video having the same coded picture width and coded picture height as a video contained in a referenced track referenced by the 'pipm' track.
[0195] 13. A method according to any one of solutions 1 to 12, wherein a video track containing a 'pipm' track reference indicates that the video in the track can be used as a PiP video, and the corresponding primary video can be a video contained at least in the referenced track or in any track in an alternate group including the referenced track.
[0196] 14. A method according to any one of solutions 1 to 13, wherein, when the video codec used for the main video is HEVC or L-HEVC, the main video can be contained in a single track or multiple tracks, the multiple tracks including a HEVC or L-HEVC slice base track and multiple HEVC or L-HEVC slice tracks containing a 'tbas' track reference referring to the slice base track, wherein in the former case, the main video track is the single track, and in the latter case, the main video track is the slice base track.
[0197] 15. A device for processing video data, comprising: a processor; and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of solutions 1 to 14.
[0198] 16. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs a method according to any one of solutions 1 to 14.
[0199] 17. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: determining that a video track contains a 'pipm' track reference, which indicates that a video in the video track or a video in any other track in a spare group including the video track can be used as a picture-in-picture (PiP) video; and generating a bitstream based on the determination.
[0200] 18. A method for storing a bitstream of a video, comprising: determining that a video track contains a 'pipm' track reference indicating that video in the video track or video in any other track in an alternate group including the video track can be used as a picture-in-picture (PiP) video; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0201] 19. A method, apparatus or system as described in this document.
[0202] The following solutions illustrate further examples of the techniques discussed herein.
[0203] 1. A method for processing media data, comprising: determining that a reference track contains a 'supm' track reference, the 'supm' track reference indicating that a video in the reference track can be used as a supplementary video and indicating that a corresponding main video can be contained in the referenced track; and performing conversion between visual media data and a visual media data file based on the 'supm' track reference.
[0204] 2. The method according to solution 1, wherein the 'supm' track reference further indicates that video in any track in the alternate set including the reference track can be used as supplementary video.
[0205] 3. The method according to any one of solutions 1 to 2, wherein the 'supm' track reference further indicates that the corresponding primary video can be contained in any track in the alternate group including the referenced track.
[0206] 4. A method according to any one of solutions 1 to 3, wherein, for each pair of a supplementary video and a main video, the track containing the supplementary video is called a supplementary video track.
[0207] 5. A method according to any one of solutions 1 to 4, wherein for each pair of supplementary video and main video, the main video track is a track of a referenced track referenced by the 'supm' track or a track of any track in an alternate group including the referenced track.
[0208] 6. A method according to any one of solutions 1 to 5, wherein, when the video codec used for the main video is Versatile Video Codec (VVC), the main video can be contained in a single track, in which case the main video track is the single track.
[0209] 7. A method according to any one of solutions 1 to 6, wherein, when the video codec used for the main video is VVC, the main video can be contained in multiple tracks, the multiple tracks including a VVC Merge base track and multiple VVC sub-picture tracks referenced by the VVCMerge base track through a 'subp' track reference, in which case the main video track is the VVC Merge base track.
[0210] 8. The method according to any one of solutions 1 to 7, wherein when the video codec for the primary video track is VVC, the sample entry type is 'vvc1' or 'vvi1'.
[0211] 9. A method according to any one of Solutions 1 to 8, wherein, when the video codec used for the main video is High Efficiency Video Codec (HEVC) or Layered HEVC (L-HEVC), the main video can be contained in a single track, in which case the main video track is the single track.
[0212] 10. A method according to any one of solutions 1 to 9, wherein, when the video codec used for the main video is HEVC or L-HEVC, the main video can be contained in multiple tracks, the multiple tracks including a HEVC or L-HEVC slice base track and multiple HEVC or L-HEVC slice tracks containing a 'tbas' track reference referring to the slice base track, in which case the main video track is the slice base track.
[0213] 11. A method according to any one of solutions 1 to 10, wherein the reference video track containing the 'supm' track reference indicates that the video in the referenced track can be used as a supplementary video, and indicates that the corresponding primary video can be the video contained in the referenced track or the video contained in any track in the backup group including the referenced track.
[0214] 12. A method according to any of solutions 1 to 11, wherein, when present, TrackReferenceTypeBox with reference_type equal to 'supm' shall contain only track identifiers and shall not contain any track group identifiers.
[0215] 13. A method according to any of solutions 1 to 12, wherein, when present, a TrackReferenceTypeBox with reference_type equal to 'supm' shall contain a track identifier in the first entry and the other entries contain track group identifiers.
[0216] 14. A method according to any of solutions 1 to 13, wherein the first entry in the TrackReferenceTypeBox with reference_type equal to 'supm' shall be the track identifier.
[0217] 15. A method according to any one of solutions 1 to 14, wherein the reference video track containing the 'supm' track reference indicates that the video in the referenced video track can be used as a supplementary video, and indicates that the corresponding primary video can be a video contained in at least the referenced video track or any track in the alternate group including the referenced video track.
[0218] 16. A method according to any one of solutions 1 to 15, wherein, when PicInPicRegionReplacementEntry is present in a supplementary video track, PicInPicRegionReplacementEntry indicates that a network abstraction layer (NAL) unit representing a target picture-in-picture (PiP) region in a corresponding main video having the same coded picture width and coded picture height as a video contained in a referenced track referenced by a 'supm' track can be replaced with a corresponding NAL unit of the supplementary video.
[0219] 17. A method according to any one of solutions 1 to 16, wherein the reference video track containing the 'supm' track reference indicates that the video in the reference track can be used as a supplementary video, and indicates that the corresponding primary video can be a video contained in at least the referenced track or any track in the alternate group including the referenced track.
[0220] 18. The method according to any one of solutions 1 to 17, wherein the converting comprises encoding the visual media data into the visual media data file.
[0221] 19. The method according to any one of solutions 1 to 17, wherein the converting comprises decoding the visual media data from the visual media data file.
[0222] 20. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of solutions 1 to 19.
[0223] 21. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer executable instructions stored on the non-transitory computer-readable medium, so that when the computer program product is executed by a processor, the video codec device performs a method according to any one of solutions 1 to 19.
[0224] 22. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: determining that a reference track contains a 'supm' track reference, the 'supm' track reference indicating that a video in the reference track can be used as a supplementary video and indicating that a corresponding main video can be included in the referenced track; and generating a bitstream based on the determination.
[0225] 23. A method for storing a bitstream of a video, comprising: determining that a reference track contains a 'supm' track reference, the 'supm' track reference indicating that a video in the reference track can be used as a supplementary video and indicating that a corresponding main video can be included in the referenced track; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0226] In the solution described herein, an encoder may comply with the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder may parse the syntax elements in the codec representation using the format rules to generate decoded video, knowing the presence and absence of syntax elements according to the format rules.
[0227] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. As defined by the syntax, the bitstream representation of the current video block may correspond, for example, to bits that are co-located or distributed at different locations in the bitstream. For example, a macroblock may be encoded based on transformed and coded error residual values, and the macroblock may also be encoded using bits in the header and other fields in the bitstream. In addition, during conversion, the decoder may parse the bitstream based on the determination and known presence or absence of certain fields, as described in the above solution. Similarly, the encoder may determine whether to include certain syntax fields and generate the codec representation accordingly by including or excluding the syntax fields in the codec representation.
[0228] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document may be implemented in digital electronic circuits, or in computer software, firmware or hardware (including the structures disclosed in this document and their equivalents), or in a combination of one or more of them. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, for execution by a data processing device or for controlling the operation of the data processing device. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances that affect a machine-readable propagation signal, or a combination of one or more of them. The term "data processing device" includes all devices, equipment and machines for processing data, including, for example, a programmable processor, a computer or multiple processors or computers. In addition to hardware, a device may also include code that creates an execution environment for the computer program in question, for example, code that constitutes a processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagation signal is an artificially generated signal, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0229] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language (including compiled or interpreted languages) and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that contains other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing portions of one or more modules, subroutines, or code). A computer program may be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communications network.
[0230] The processes and logic flows described in this document may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may also be implemented as, special purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).
[0231] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data or be operably coupled to these mass storage devices to receive data from them or to transfer data to them, or both. However, a computer does not necessarily require such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, for example, including semiconductor storage devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disk read-only memory (CD ROM) and digital versatile disk read-only memory (DVD-ROM) disks. The processor and memory can be supplemented by or incorporated into a dedicated logic circuit.
[0232] Although this patent document contains many details, these details should not be interpreted as limitations on the scope of any subject matter or what may be claimed, but rather as descriptions of features that may be specific to a particular embodiment of a particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments, either individually or in any suitable subcombination. In addition, although the above features may be described as working in certain combinations or even initially claimed as such, in some cases one or more features in the combination may be removed from the claimed combination, and the claimed combination may be directed to a subcombination or a variation of a subcombination.
[0233] Similarly, although operations are depicted in a particular order in the drawings, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed, in order to achieve the desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0234] Only a few implementations and examples are described, and other implementations, enhancements, and variations may be made based on what is described and shown in this patent document.
[0235] A first component is directly coupled to a second component when there is no intermediate component between the first component and the second component other than a line, a trace, or another medium. A first component is indirectly coupled to a second component when there is an intermediate component between the first component and the second component other than a line, a trace, or another medium. The term "coupled" and its variations include both direct coupling and indirect coupling. Unless otherwise specified, the use of the term "approximately" is intended to include a range of ±10% of the subsequent value.
[0236] Although several embodiments are provided in the present disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples should be considered illustrative rather than restrictive, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0237] In addition, without departing from the scope of the present disclosure, the techniques, systems, subsystems and methods described and shown as discrete or separate in various embodiments can be combined or integrated with other systems, modules, techniques or methods. Other items shown or discussed as coupled can be directly connected, or can be indirectly coupled or communicated by electrical, mechanical or other means through some interface, device or intermediate component. Other examples of changes, substitutions and alterations can be determined by those skilled in the art, and these changes, substitutions and alterations can be made without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing media data, comprising: Determining that a reference track includes a 'supm' track reference, the 'supm' track reference indicating that a video in the reference track can be used as a supplementary video and indicating that a corresponding primary video can be included in the referenced track; as well as Conversion between visual media data and visual media data files is performed based on the 'supm' track reference.
2. The method according to claim 1, wherein: The 'supm' track reference also indicates that video in any track in the alternate set including the reference track can be used as supplementary video.
3. The method according to any one of claims 1 to 2, wherein: The 'supm' track reference also indicates that the corresponding primary video may be contained in any track in the alternate set including the referenced track.
4. The method according to any one of claims 1 to 3, wherein: For each pair of complementary video and main video, the track containing the complementary video is called complementary video track.
5. The method according to any one of claims 1 to 4, wherein: For each pair of complementary video and main video, the main video track is the track of the referenced track referenced by the 'supm' track or any track in the alternate set including the referenced track.
6. The method according to any one of claims 1 to 5, wherein: When the video codec used for the main video is Versatile Video Codec (VVC), the main video may be contained in a single track, in which case the main video track is the single track.
7. The method according to any one of claims 1 to 6, wherein: When the video codec for the main video is VVC, the main video may be contained in multiple tracks including a VVC Merge base track and multiple VVC sub-picture tracks referenced by the VVC Merge base track through a 'subp' track reference, in which case the main video track is the VVC Merge base track.
8. The method according to any one of claims 1 to 7, wherein: When the video codec for the primary video track is VVC, the sample entry type is 'vvc1' or 'vvi1'.
9. The method according to any one of claims 1 to 8, wherein: When a video codec for a main video is High Efficiency Video Codec (HEVC) or Layered HEVC (L-HEVC), the main video may be contained in a single track, in which case the main video track is the single track.
10. The method according to any one of claims 1 to 9, wherein: When the video codec for the main video is HEVC or L-HEVC, the main video may be contained in a plurality of tracks including a HEVC or L-HEVC slice base track and a plurality of HEVC or L-HEVC slice tracks containing a 'tbas' track reference to the slice base track, in which case the main video track is the slice base track.
11. The method according to any one of claims 1 to 10, wherein: The reference video track containing the 'supm' track reference indicates that the video in the referenced track can be used as supplementary video and that the corresponding primary video can be the video contained in the referenced track or the video contained in any track in the alternate set including the referenced track.
12. The method according to any one of claims 1 to 11, wherein: When present, a TrackReferenceTypeBox with reference_type equal to 'supm' shall contain only the track identifier and shall not contain any track group identifiers.
13. The method according to any one of claims 1 to 12, wherein: When present, a TrackReferenceTypeBox with reference_type equal to 'supm' shall contain the track identifier in the first entry, and the track group identifiers in the other entries.
14. The method according to any one of claims 1 to 13, wherein: The first entry in a TrackReferenceTypeBox with reference_type equal to 'supm' shall be the track identifier.
15. The method according to any one of claims 1 to 14, wherein: The reference video track containing the 'supm' track reference indicates that the video in the referenced video track can be used as supplementary video, and indicates that the corresponding primary video can be a video contained in at least the referenced video track or any track in an alternate group including the referenced video track.
16. The method according to any one of claims 1 to 15, wherein: When PicInPicRegionReplacementEntry is present in a supplementary video track, PicInPicRegionReplacementEntry indicates that a network abstraction layer (NAL) unit representing a target picture-in-picture (PiP) region in the corresponding primary video having the same coded picture width and coded picture height as the video contained in the referenced track referenced by the 'supm' track may be replaced with the corresponding NAL unit of the supplementary video.
17. The method according to any one of claims 1 to 16, wherein: The reference video track containing the 'supm' track reference indicates that the video in the reference track can be used as supplementary video, and indicates that the corresponding primary video can be a video contained in at least the referenced track or any track in the alternate group including the referenced track.
18. The method according to any one of claims 1 to 17, wherein: The converting includes encoding the visual media data into the visual media data file.
19. The method according to any one of claims 1 to 17, wherein: The converting includes decoding the visual media data from the visual media data file.
20. An apparatus for processing video data, comprising: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 19.
21. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer executable instructions stored on the non-transitory computer-readable medium, so that when the computer program product is executed by a processor, the video codec device performs the method according to any one of claims 1 to 19.
22. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing device, wherein: The method comprises: determining that a reference track includes a 'supm' track reference, the 'supm' track reference indicating that a video in the reference track can be used as a supplementary video and indicating that a corresponding primary video can be included in the referenced track; and A bitstream is generated based on the determination.
23. A method for storing a bit stream of a video, comprising: Determining that a reference track includes a 'supm' track reference, the 'supm' track reference indicating that a video in the reference track can be used as a supplementary video and indicating that a corresponding primary video can be included in the referenced track; generating a bitstream based on the determination; as well as The bit stream is stored in a non-transitory computer-readable recording medium.
Citation Information
Patent Citations
Method for content presentation during trick mode operations
CN103181164A
A device and method for video media recording / playingwith multiple tracks
KR1020050121345A
Methods and apparatus for immersive media content overlays
US20200014906A1
Coded Picture with Mixed VCL NAL Unit Type
US20220109861A1
A method, an apparatus and a computer program product for video encoding and video decoding
WO2021136880A1