Picture-in-Picture Signal Transmission in Media Files

By optimizing the coding of region ID types and ensuring valid region ID counts, the method enhances the efficiency and compatibility of picture-in-picture signal transmission in ISOBMFF-based media files, addressing existing inefficiencies and limitations.

JP2025517164AInactive Publication Date: 2025-06-03BYTEDANCE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024566344
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-09
Filing Date
2023-05-09
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies for transmitting picture-in-picture signals in ISOBMFF-based media files face issues such as inefficient use of bits in the region identifier type field, meaningless values for the number of region IDs, and limited compatibility with video codecs other than VVC.

Method used

The proposed method involves coding the region ID type field using 4 bits or less, ensuring the number of region IDs is greater than 0, and specifying that a region ID type of 0 indicates a VVC sub-picture ID only when both main and PiP video tracks use VVC, thereby optimizing bit usage and enhancing compatibility.

Benefits of technology

This approach reduces unnecessary bit usage, eliminates meaningless region ID values, and allows for compatible use with various video codecs, thereby improving the efficiency and versatility of picture-in-picture signal transmission in ISOBMFF-based media files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517164000001_ABST
    Figure 2025517164000001_ABST
Patent Text Reader

Abstract

A mechanism for processing image data is disclosed. In the region ID type field (region_id_type), the type of value taken by the region identifier (ID) is determined. In the region ID type field, the type of value taken by the region ID is coded in 4 bits or less. The conversion between visual media data and the media data file is performed based on the type of value taken by the region ID of the region ID type field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the priority and benefit of U.S. Patent Application No. 63 / 339,676, filed on May 9, 2022, the entire disclosure of which is incorporated herein by reference.

[0002] This disclosure relates to the generation, storage, and consumption of digital audio - visual media information in file formats.

Background Art

[0003] Digital video occupies the largest bandwidth used in the Internet and other digital communication networks. As the number of user devices capable of receiving and displaying video increases, the bandwidth demand for digital video utilization is likely to continue to increase.

Summary of the Invention

[0004] A first embodiment is a method for processing video data, which includes determining a value of a region identifier (ID) type (region_id_type) within a region identifier (ID) type field, where the value of the region ID type is coded using 4 bits or less within the region ID type field, and performing a conversion between visual media data and a media data file based on the value of the region ID type within the region ID type field.

[0005] Optionally, in any of the foregoing embodiments, in other implementations, it is provided that the remaining 4 or more or multiple bit groups of the region ID type field are reserved.

[0006] Optionally, in any of the foregoing embodiments, in other implementations, it is provided that the region ID type being equal to 0 specifies that the region ID is a general video coding (VVC) sub - picture ID.

[0007] Optionally, in any of the foregoing embodiments, in other implementations, the region ID type provides an indication of the type of value taken by the region ID.

[0008] Optionally, in any of the foregoing embodiments, in other implementations, the region ID is provided to be used for the transmission of picture-in-picture signals in the International Organization for Standardization (ISO) Base Media File Format (ISOBMFF).

[0009] Optionally, in any of the foregoing embodiments, in other implementations, the media data file is provided to include the International Organization for Standardization (ISO) Base Media File Format (ISOBMFF).

[0010] A second embodiment is a method for processing video data, including determining a value of the number of region identifiers, where the value of the number of region identifiers is constrained to be greater than 0, and performing a conversion between visual media data and a media data file based on the value of the number of region identifiers.

[0011] Optionally, in any of the foregoing embodiments, in other implementations, the number of region identifiers is provided to be represented as a number (num_region_ids_minus1) that is 1 less than the number of region identifiers.

[0012] Optionally, in any of the foregoing embodiments, in other implementations, adding 1 to num_region_ids_minus1 provides the specification of a region identifier (region_id[i]).

[0013] Optionally, in any of the foregoing embodiments, in other implementations, the region identifier is provided to specify the i-th ID of a coded video data unit representing a target picture-in-picture region.

[0014] Optionally, in any of the foregoing embodiments, in other implementations, it is provided that the region identifier is determined according to a loop specified by for(i=0;i<=num_region_ids_minus1;i++).

[0015] Optionally, in any of the foregoing embodiments, in other implementations, it is provided that the number of region identifiers is represented by num_region_ids.

[0016] Optionally, in any of the foregoing embodiments, in other implementations, it is provided that the number of region identifiers is represented as num_region_ids_minus1.

[0017] A third embodiment is a method for processing video data, including determining a value of a region identifier type (region_id_type), and specifying that when the value of the region identifier type is equal to 0 and the video coding used for coding the main video track and the picture-in-picture (PiP) video track is VVC, the region identifier (ID) is a general-purpose video coding (VVC) sub-picture ID, and performing a conversion between visual media data and a media data file based on the value of the region identifier type.

[0018] Optionally, in any of the foregoing embodiments, in other implementations, it is provided that when the video codec of the main video track is not VVC, a value of the region ID type equal to 0 is reserved.

[0019] Optionally, in any of the foregoing embodiments, in other implementations, it is provided that when the video codec used for coding the main video track and the picture-in-picture video track is not VVC, a value of the region ID type equal to 0 is reserved.

[0020] Optionally, in any of the foregoing embodiments, in other implementations, when the video codec used for the main video track is VVC, it is provided that the sample entry type is vvc1, vvi1, or vvs1.

[0021] A fourth embodiment relates to a video data processing apparatus including a processor and a non-transitory memory having instructions that cause the processor to execute any of the foregoing embodiments when executed by the processor.

[0022] A fifth embodiment relates to a non-transitory computer-readable medium including a computer program product for use by a video coding apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that cause a processor to execute any of the foregoing methods when executed by the processor.

[0023] A sixth embodiment relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method executed by a video processing apparatus, the method determining a value of a region identifier (ID) type (region_id_type) in a region identifier (ID) type field, the value of the region ID type being coded in the region ID type field using 4 bits or less, and performing a conversion between visual media data and a media data file based on the value of the region ID type in the region ID type field.

[0024] A seventh embodiment is a method for storing a video bitstream, which includes determining a value of a region identifier (ID) type in a region identifier (ID) type field, where the value of the region ID type is coded in the region ID type field using 4 bits or less, generating a bitstream based on the determination, and storing the bitstream in a non-transitory computer-readable recording medium.

[0025] An eighth embodiment relates to a non-transitory computer-readable recording medium storing a video bitstream generated by a method executed by a video processing apparatus, the method including determining a value of a number of region identifiers, where the value of the number of region identifiers is constrained to be greater than zero, and performing a conversion between visual media data and a media data file based on the value of the number of region identifiers.

[0026] A ninth embodiment relates to a non-transitory computer-readable recording medium storing a method for storing a video bitstream, the method including determining a value of a number of region identifiers, where the value of the number of region identifiers is constrained to be greater than zero, generating a bitstream based on the determination, and storing the bitstream in a non-transitory computer-readable recording medium.

[0027] A non-transitory computer-readable recording medium storing a bitstream of video generated by a method executed by a video processing apparatus, the method determining a value of a region identifier type (region_id_type), wherein a value of the region identifier type equal to 0 specifies that, when the video codec used for coding a main video track and a picture-in-picture (PiP) video track is VVC, the region identifier (ID) is a general video coding (VVC) sub-picture ID, and performing conversion between visual media data and a media data file based on the value of the region identifier type.

[0028] A method of storing a bitstream of video, the method determining a value of a region identifier type (region_id_type), wherein a value of the region identifier type equal to 0 specifies that, when the video codec used for coding a main video track and a picture-in-picture (PiP) video track is VVC, the region identifier (ID) is a general video coding (VVC) sub-picture ID, generating a bitstream based on the determination, and storing the bitstream in a non-transitory computer-readable recording medium.

[0029] Embodiment 12 relates to a method, apparatus, or system described in the present disclosure.

[0030] For the purpose of clarification, any one of the embodiments can be combined with any one or more of the other embodiments to create a new embodiment within the scope of the present disclosure.

[0031] These features and other features will be more clearly understood from the following detailed description, which is detailed in conjunction with the accompanying drawings and the claims.

[0032] Next, for a more complete understanding of the present disclosure, reference should be made to the following brief description in connection with the accompanying drawings and detailed description.

Brief Description of the Drawings

[0033]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Modes for Carrying Out the Invention

[0034] Exemplary examples of one or more embodiments are provided below. However, the disclosed system and / or method can be implemented using any number of techniques, whether currently known or not yet developed. The present disclosure should in no way be limited to the examples, drawings, and techniques illustrated below, but includes the exemplary designs and implementations illustrated and described herein, and can be modified within the full scope of the appended claims and their equivalents.

[0035] In this disclosure, section headings are used for ease of understanding, and the applicability of the technologies and embodiments disclosed in each section is not limited to that section only. Further, the terms of H.266 are used only in some explanations for ease of understanding and do not limit the scope of the disclosed technologies. Thus, the technologies described herein are also applicable to other video codec protocols and designs. In this disclosure, with respect to drafts of the VVC specification or the ISOBMFF file format specification, editorial changes are indicated by strikethrough italic text for text changed by editing and underlined bold text for added text.

[0036] 1. Initial Discussion This disclosure relates to media file formats. Specifically, this disclosure relates to the signal transmission of picture-in-picture services within a media file. This mechanism can be applied individually or in various combinations, for example, based on the International Organization for Standardization (ISO) base media file format (ISOBMFF) or its extensions for a media file format, and / or the conveyance of structured video in network abstraction layer (NAL) units within ISOBMFF.

[0037] 2. Introduction to Video Coding 2.1 Video Coding Standards Video coding standards have mainly evolved through the development of standards by the Telecommunication Standardization Sector (ITU-T) of the International Telecommunication Union (ITU) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T has developed H.261 and H.263, ISO / IEC has developed Motion Picture Experts Group (MPEG)-1 and MPEG-4 Visual, and both organizations have jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / High Efficiency Video Coding (HEVC) [1] standards. Since H.262, video coding standards have been based on a hybrid video coding structure of temporal prediction + transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET). Many techniques have been adopted in JVET and incorporated into the reference software named Joint Exploration Model (JEM) [2]. JVET was renamed the Joint Video Experts Team (JVET) when the Versatile Video Coding (VVC) project was officially launched. VVC [3] is a coding standard aiming for a 50% bitrate reduction compared to HEVC.

[0038] The Versatile Video Coding (VVC) standard (ITU-T H.266|ISO / IEC 23090-3) [3][4] and the related Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IEC 23002-7) [5][6] are designed for use in a maximum wide range of applications, including not only conventional applications such as television broadcasting, video conferencing, or playback from storage media, but also more recent and advanced use cases such as adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple coded video bitstreams, multi-view video, scalable layering, viewport-adaptive 360° immersive media.

[0039] The EVC (Essential Video Coding) standard (ISO / IEC 23094-1) is another video coding standard being developed by MPEG.

[0040] 2.2 Standard of File Format Media streaming applications are based on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP) transport methods and rely on file formats such as the ISO Base Media File Format (ISOBMFF) [7]. One such streaming system is DASH (dynamic adaptive streaming over HTTP) [8]. To use video formats with ISOBMFF and DASH, file format specifications specific to the video format, such as the AVC file format and HEVC file format in [9], are required for encapsulation of video content in ISOBMFF tracks, DASH representations, and segments. Important information regarding the video bitstream, such as profile, tier, level, and many other pieces of information, needs to be exposed as file format-level metadata and / or DASH Media Presentation Description (MPD) for the purpose of content selection, e.g., for the selection of appropriate media segments both at the initialization of a streaming session and for stream adaptation during the streaming session. Similarly, to use picture formats with ISOBMFF, file format specifications specific to the picture format, such as the AVC picture file format and HEVC picture file format in

[10] , are required.

[0041] 2.3 Conversion of Video Pictures for Presentation in ISOBMFF In ISOBMFF, the movie header box and the track header box contain matrix fields as follows. template int(32)[9] matrix= {0x00010000,0,0,0x00010000,0,0,0x40000000}; / / Unity matrix

[0042] These matrix values specify the transformation for the display of video pictures. Not all derived specifications use matrices. If no matrix is used, the matrix is set to the identity matrix. When a matrix is used, a point (p, q) is transformed to (p´, q´) using the matrix as follows. (p q 1)*|a b u|=(m n z) |c d v| |x y w| m = ap + cq + x; n = bp + dq + y; z = up + vq + w; p´ = m / z; q´ = n / z

[0043] The coordinates {p, q} are on the decoded frame, and {p´, q´} are on the rendered output. Thus, for example, the matrix {2, 0, 0, 0, 2, 0, 0, 0, 1} exactly doubles the pixel dimensions of the picture. Coordinates transformed by the matrix may not be normalized and the coordinates represent the actual sample positions. Thus {x, y} can be regarded as, for example, the translation vector of the picture. The origin of the coordinates is located at the upper left, the X value increases towards the right, and the Y value increases downwards. {p, q} and {p´, q´} can be regarded as the upper left corner of the original picture (after scaling to the size determined by the width and height of the track header) and the absolute pixel position relative to the transformed (rendered) surface, respectively.

[0044] Each track is composited into the overall picture using the specified corresponding matrix. This result is then transformed and composited according to the movie-level matrix within the MovieHeaderBox. Whether the resulting picture is clipped to remove pixels depends on the application. Such removed pixels are not displayed, for example, in a vertical rectangular area within a window. Thus, for example, if only one video track is being displayed, the video track has a translation to {20,30}, and the identity matrix is within the MovieHeaderBox, the application can choose not to display the empty L-shaped area between the picture and the origin. All values within the matrix are stored as 16.16 fixed-point values, except for u, v, and w which are stored as 2.30 fixed-point values. The values within the matrix are stored in the order {a,b,u,c,d,v,x,y,w}.

[0045] 2.4 Pixel Data Processing in ISOBMFF The processing of pixel data from the decoder output to the drawing of pixel data on the screen may not need to conform to the ISOBMFF. However, some structures enable the signaling of such drawing capabilities. Conformance is based on the following assumptions.

[0046] First, for the ‘iso3’ brand, or brands sharing the corresponding requirements, the width and height of the TrackHeaderBox are set assuming that the pixels being drawn are square so that the pixel aspect ratio is 1:1. For other brands, the use of related structures for drawing is undetermined.

[0047] Second, VisualSampleEntry records the expected size of the pixel buffer required to receive the codec output. The expected size may be truncated by the in-stream structure. For example, if a video codec operates only on multiples of 16 pixels per row or column, when the width and height of the video supplied to the encoder are 1000×500, the encoder typically uses internal blocks of 64×32 pixels. However, the encoder is expected to use codec-specific truncation structures not disclosed at the ISOBMFF level to output a 1000×500 video. In this case, the width and height of the VisualSampleEntry fields will be 1000×500.

[0048] Third, the PixelAspectRatioBox records the aspect ratio to be applied to the pixels output by the decoder, but does not imply the corresponding adjustment. The adjustment is made by setting the scaling value in the TrackHeaderBox. For example, if the pixels are horizontally scaled by a factor of 1 / 2 before the encoder, it is necessary to apply the inverse scaling to display a distortion-free video in the decoder output. This is done by setting the height of the TrackHeaderBox equal to the height of the VisualSampleEntry and the width of the TrackHeaderBox equal to twice the width of the VisualSampleEntry. Additionally, if the in-stream structure does not have such information, a PixelAspectRatioBox with hSpacing greater than twice the corresponding vSpacing field can be used.

[0049] Fourth, a CleanApertureBox may be provided to further crop the video.

[0050] The processing of decoded pixels is assumed as follows. Any cropping recorded by the CleanApertureBox is applied to the pixels output by the decoder. If the CleanApertureBox exists, the cropped picture is scaled horizontally by the factor TrackHeaderBox.width / CleanAperture.width and vertically by the factor TrackHeaderBox.height / CleanAperture.height. Otherwise, if the CleanApertureBox does not exist, the decoded picture is scaled horizontally by the factor TrackHeaderBox.width / SampleEntry.width and vertically by the factor TrackHeaderBox.height / SampleEntry.height. This operation is sometimes referred to as normalization to the track dimensions. Then, the TrackHeaderBox matrix is applied. All visual tracks are overlaid in ascending order of the TrackHeaderBox.layer value. Then, the MovieHeaderBox matrix is applied to the composition. This is an example of a processing model, and in a specific implementation following this model, especially when the combination of the above operations results in an identity transformation, it is necessary to avoid resampling the picture.

[0051] 2.5 Grouping of Tracks and Grouping of Entities ISOBMFF specifies both track grouping and entity grouping. Track grouping is signaled based on track group boxes contained in the track box. Therefore, track grouping is track-level signaling. This track group box enables indication of a group of tracks where each group shares specific characteristics or the tracks within the group have specific relationships. The specific characteristics or relationships are indicated by the box type of the boxes contained in the track group box. The boxes contained in the track group box include identifiers (IDs) that can be used to specify that tracks belong to the same track group. Tracks for which the box types of the boxes contained within the track group box are the same and the values of the identifiers of the boxes contained within these boxes are the same belong to the same track group.

[0052] An entity group is a grouping of items and may also group tracks. Entities in an entity group share specific characteristics or have specific relationships as indicated by the grouping type. The entity group is displayed in the GroupsListBox. The entity group specified in the GroupsListBox of the file-level MetaBox refers to track or file-level items. The entity group specified in the GroupsListBox of the movie-level MetaBox refers to movie-level items. The entity group specified in the GroupsListBox of the track-level MetaBox refers to the track-level items of that track. The GroupsListBox contains an EntityToGroupBox that each specifies one entity group. The GroupsListBox contains the entity groups specified for the file. This box is called the EntityToGroupBox and contains a set of full boxes with 4-character codes indicating the defined group types. If the GroupsListBox exists in the file-level MetaBox, the item_ID value equal to the track_ID value of any TrackHeaderBox shall not exist in the ItemInfoBox of any file-level MetaBox.

[0053] 2.6 Picture-in-Picture Signaling in ISOBMFF The picture-in-picture service provides a function of including a picture with a small resolution in a picture with a large resolution. Such a service is effective for simultaneously displaying two videos to the user, where the video with a large resolution is regarded as the main video and the video with a small resolution is regarded as the auxiliary video. Such a picture-in-picture service can be used to provide an accessibility service in which the main video is complemented by a signage video. The design of the picture-in-picture signal in ISOBMFF is included in

[11] . The design is as follows.

[0054] 2.6.1 Picture-in-Picture Sample Group 2.6.1.1 Definition The Picture-in-Picture (PiP) service provides the ability to include a video with a small spatial resolution within a video with a large spatial resolution. A video track containing a'subt' track reference indicates that the track contains a PiP video and that the main video is included in the referenced track or, if any, any track within the alternative group to which the referenced track belongs. For each pair of PiP video and main video, the window within the main video for embedding / overlaying a PiP video that is smaller in size than the main video is indicated by the value of the matrix field of the TrackHeaderBox of the PiP video track and the main video track. In order to overlay the PiP video in front of the main video, it is required that the value of the layer field of the TrackHeaderBox of the auxiliary video track is smaller than the value of the main video track. If a PicInPicInfoEntry exists in the PiP video track, it indicates that the coded video data unit representing the target PiP region of the main video is replaced by the corresponding coded video data unit of the PiP video. In this case, it is required to use the same video codec for the coding of the PiP video and the main video. The absence of this sample group indicates that it is unknown whether such replacement is possible.

[0055] 2.6.1.2 Syntax class PicInPicInfoEntry() extends VisualSampleGroupEntry('pinp') { unsigned int(8) region_id_type; unsigned int(8) num_region_ids; for(i = 0; i < num_region_ids; i++) unsigned int(16) region_id[i]; }

[0056] 2.6.1.3 Semantics If the PicInPicInfoEntry exists, the player can choose to replace the coded video data unit representing the target PiP area within the main video with the corresponding coded video data unit of the PiP video before sending it to the video decoder for decoding. In this case, the corresponding video data units of the PiP video corresponding to a specific picture of the main video are all the coded video data units within the decoding time synchronization samples of the PiP video track.

[0057] The region_id_type indicates the type of value that region_id takes. When region_id_type is 0, the region ID is the VVC sub-picture ID. When region_id_type is 1, the region ID is the value of the group ID of the NAL unit map sample group of the NAL unit that may be replaced by the NAL unit of the PiP track. Values of region_id_type greater than 1 are reserved. num_region_ids specifies the number of the following region_id[i] fields. region_id[i] specifies the i-th ID of the coded video data unit representing the picture-in-picture region of interest. When region_id_type is 1, the main video track has a 'nalm' sample group where the grouping_type_parameter is equal to 'pinp', indicating the NAL unit of the main track that can be replaced with the NAL unit of the PiP track having the same group ID value. When region_id_type is equal to 1 and num_region_ids is equal to 1, the 'nalm' sample group shall not exist in the PiP track, and all NAL units of the PIP track are implicitly considered to have a group ID equal to region_id[0]. When region_id_type is equal to 1 and num_region_ids is greater than 1, the 'nalm' sample group where the grouping_type_parameter is equal to 'pinp' shall exist in the PiP track and shall provide a mapping to the NAL unit of the group ID value.

[0058] 3. Technical problems solved by the disclosed technical solutions Design examples for picture-in-picture signal transmission in ISOBMFF-based media files have the following problems. First, the region identifier (region_id) type (region_id_type) field is coded in 8 bits. However, less than 8 bits may be sufficient. Second, the value of the number of region IDs (num_region_ids) is allowed to be equal to 0. However, a value of 0 for num_region_ids is meaningless. Third, the specification defines that when the region ID type (region_id_type) is 0, the region ID is the VVC sub-picture ID. However, this does not allow the use of a region_id_type equal to 0 for other video codecs that support the coding of regions such as VVC sub-pictures.

[0059] 4. List of Solutions and Embodiments To solve the above problems, a method as summarized below is disclosed. Embodiments should be considered as examples for explaining general concepts and should not be interpreted narrowly. Furthermore, these embodiments can be applied individually or combined in any way.

[0060] Example 1 To solve the first problem, the region_id_type field may be coded in less than 8 bits. For example, the region_id_type field may be coded using 2 bits, 3 bits, or 4 bits, and the remaining 8 bits can be reserved for further use in a backward-compatible manner.

[0061] Example 2 To solve the second problem, the present disclosure can specify that the value of num_region_ids must be greater than 0. In another example, the name of the num_region_ids field can be changed to a value (num_region_ids_minus1) that is 1 less than the number of region IDs. Further, the present disclosure specifies that the value obtained by adding 1 to num_region_ids_minus1 specifies the number of the region_id[i] fields of the next region ID, and the loop "for(i = 0; i < num_region_ids; i++)" can also be changed to "for(i = 0; i <= num_region_ids_minus1; i++)".

[0062] Example 3 To solve the third problem, the present disclosure provides that when the video codec used for the main video track is VVC (indicated by the sample entry type being equal to 'vvc1', 'vvi1', or 'vvs1'), the video codec used for the picture-in-picture (PiP) video track is also VVC, and a value of region_id_type equal to 0 can specify that the region ID is a VVC sub-picture ID. Further, the present disclosure may provide that when the video codec used for the main video track is not VVC, the video codec used for the PiP video track is also not VVC, and when the value of region_id_type is equal to 0, it is reserved.

[0063] 5. Embodiments Examples of embodiments related to all of the above disclosure items and most of their sub-items are shown below. The most relevant parts that are added or changed are shown in underlined bold, and some of the deleted parts are shown in italic bold. Additionally, there may be other editorial changes, but their descriptions are omitted. Parts with no changes are not included.

[0064] 5.1.1 Picture-in-Picture Sample Group 5.1.1.1 Definition The Picture-in-Picture (PiP) service provides the ability to include a video with a smaller spatial resolution within a video with a larger spatial resolution. A video track containing a 'subt' track reference indicates that the track contains a PiP video and that the main video is included in the referenced track or, if any, any track within the alternative group to which the referenced track belongs. For each pair of PiP video and main video, the window within the main video for embedding / overlaying a PiP video smaller in size than the main video is indicated by the value of the matrix field of the TrackHeaderBox of the PiP video track and the main video track. To layer the PiP video in front of the main video, the value of the layer field of the TrackHeaderBox of the auxiliary video track is required to be smaller than the value of the main video track. If a PicInPicInfoEntry exists in the PiP video track, it indicates that the coded video data unit representing the target PiP region of the main video can be replaced with the corresponding coded video data unit of the PiP video. In this case, it is required to use the same video codec for the coding of the PiP video and the main video. The absence of a sample group indicates that it is not clear whether such replacement is possible.

[0065] 5.1.1.2 Syntax

[0066]

Chem.

[0067] Semantics If a PicInPicInfoEntry exists, the player can choose to replace the coded video data unit representing the target PiP region within the main video with the corresponding coded video data unit of the PiP video before sending it to the video decoder for decoding. In this case, for a specific picture of the main video, the corresponding video data unit of the PiP video is all the coded video data units within the decode time synchronization sample of the PiP video track. The region_id_type indicates the type of the value of region_id.

[0068]

Chem.

[0069]

Chem.

[0070] When region_id_type is 1, the main video track has a 'nalm' sample group where the grouping_type_parameter is equal to 'pinp', indicating the NAL units of the main track that can be replaced with the NAL units of the PiP track having the same group ID value. When region_id_type is equal to 1 and num_region_ids is equal to 1, the 'nalm' sample group shall not exist in the PiP track, and all NAL units of the PIP track are implicitly considered to have a groupID equal to region_id[0]. When region_id_type is equal to 1 and num_region_ids is greater than 1, a 'nalm' sample group where the grouping_type_parameter is equal to 'pinp' exists in the PiP track, providing the mapping of the group ID value to the NAL units.

[0071] 6. References [1] ITU-T and ISO / IEC, "High efficiency video coding", Rec.ITU-T H.265|ISO / IEC 23008-2 (current version). [2] J.Chen, E.Alshina, G.J.Sullivan, J.-R.Ohm, J.Boyce, "Algorithm description of Joint Exploration Test Model 7 (JEM7)", JVET-G1001, Aug. 2017. [3] Rec.ITU-T H.266|ISO / IEC 23090-3, "Versatile Video Coding", 2020. [4] B.B.Bross, J.Chen, S.Liu, Y.-K.Wang (editors), "Versatile Video Coding (Draft 10)", JVET-S2001. [5] Rec.ITU-T Rec.H.274|ISO / IEC 23002-7, "Versatile Supplemental Enhancement Information Messages for Coded Video Bitstream", 2020. [6] J.Boyce, V.Drugeon, G.J.Sullivan, Y.-K.Wang (editors), "Versatile supplemental enhancement information messages for coded video bitstreams (Draft 5)", JVET-S2007. [7] ISO / IEC 14496-12: "Information technology - Coding of audio-visual objects - Part 12: ISO base media file format". [8] ISO / IEC 23009-1: "Information technology - Dynamic adaptive streaming over HTTP (DASH) - Part 1: Media presentation description and segment format". The fourth edition text of the DASH standard specification is published in MPEG input document m52458. [9] ISO / IEC 14496-15: "Information Technology - Coding of Audio-Visual Objects - Part 15: Transmission of Structured Video in Network Abstraction Layer (NAL) Units in the ISO Base Media File Format".

[10] ISO / IEC 23008-12: "Information Technology - High-Efficiency Coding and Media Delivery in Heterogeneous Environments - Part 12: Picture File Format".

[11] K.K. Sreedhar, M.M. Hannuksela, Lukasz Kondrad, and Lauri Ilola, "On picture-in-picture signalling in ISOBMFF", MPEG input document m59497, April 2022.

[0072] Figure 1 is a block diagram showing an example of a video processing system 4000 in which various techniques disclosed in the present disclosure can be implemented. Various embodiments can include some or all of the components of system 4000. System 4000 can include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 4002 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wireless LAN (Wi-Fi) or cellular interfaces.

[0073] System 4000 may include a coding component 4004 that can implement various coding or encoding methods disclosed in this disclosure. The coding component 4004 may reduce the average bitrate of the video from the input 4002 to the output of the coding component 4004 to generate a coded representation of the video. Thus, the coding technology may also be referred to as video compression technology or video transcoding technology. The output of the coding component 4004 is either stored or transmitted via a communication connection as represented by the component 4006. The stored or communicated bitstream (or coded) representation of the video received at the input 4002 may be used by the component 4008 to generate pixel values or a displayable video that is transmitted to the display interface 4010. The process of generating a video viewable by a user from the bitstream representation may be referred to as video decompression. Further, certain video processing operations are called "coding" operations or tools, while coding tools or operations are used in an encoder, and it will be understood that the corresponding decoding tools or operations that reverse the result of coding are executed in a decoder.

[0074] Examples of a peripheral bus interface or a display interface include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI (registered trademark)), DisplayPort, and the like. Examples of a storage interface include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, and the like. The technology described in this disclosure may be implemented in various electronic devices such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0075] FIG. 2 is a block diagram of an example of a video processing apparatus 4100. The apparatus 4100 may be used to implement one or more of the methods described herein. The apparatus 4100 may be implemented in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, or the like. The apparatus 4100 may include one or more processors 4102, one or more memories 4104, and a video processing circuit 4106. The processor(s) 4102 may be configured to implement one or more of the methods described in this disclosure. The memory (storage device) 4104 may be used to store data and coding used to implement the methods and techniques described in this disclosure. The video processing circuit 4106 may be used to implement some of the techniques described in this disclosure in a hardware circuit. In some embodiments, at least a part of the video processing circuit 4106 may be included in the processor 4102, for example, a graphics co-processor.

[0076] FIG. 3 is a flowchart of an example of a method 4200 for processing video data. In block 4202, the method 4200 includes determining a value of a region identifier (ID) type (region_id_type) in a region identifier (ID) type field. The value of the region ID type is coded in the region ID type field using 4 bits or less. For example, the region ID may be coded with 4 bits, 3 bits, 2 bits, or 1 bit. In block 4204, the method 4200 includes performing a conversion between visual media data and a media data file based on the value of the region ID type in the region ID type field. The conversion in step 4204 may include encoding by an encoder or decoding by a decoder in some embodiments.

[0077] In an embodiment, the remaining four or more bits of the group within the region ID type field are reserved. In an embodiment, the region ID type being equal to 0 specifies that the region ID is a general-purpose video coding (VVC) sub-picture ID. In an embodiment, the region ID type indicates the type of value that the region ID takes. In an embodiment, the region ID is used for the signal transmission of picture-in-picture in the International Organization for Standardization base media file format (ISOBMFF). In an embodiment, the media data file is composed of the International Organization for Standardization base media file format (ISOBMFF).

[0078] Figure 4 is a flowchart of an example of a method 4220 for processing video data. In block 4222, method 4220 includes determining a value of the number of region identifiers, where the value of the number of region identifiers is constrained to be greater than 0. In block 4224, method 4220 includes performing a conversion between visual media data and a media data file based on the value of the number of region identifiers. The conversion in step 4224 may include, in some examples, encoding in an encoder or decoding in a decoder.

[0079] In an embodiment, the number of region identifiers is represented by a number (num_region_ids_minus1) that is one less than the number of region identifiers. In an embodiment, adding 1 to num_region_ids_minus1 is to specify a region identifier (region_id[i]). In an embodiment, the region identifier specifies the i-th ID of a coded video data unit representing a target picture-in-picture region. In an embodiment, the region identifier is determined according to a loop specified by for(i = 0; i <= num_region_ids_minus1; i++). In an embodiment, the number of region identifiers is represented as num_region_ids. In an embodiment, the number of region identifiers is represented by num_region_ids_minus1.

[0080] FIG. 5 is a flowchart of an example of a method 4240 for processing video data. In block 4242, method 4240 includes determining a value of a region identifier type (region_id_type). The fact that the value of the region identifier type is equal to 0 specifies that, when the video codec used for coding the main video track and the picture-in-picture (PiP) video track is VVC, the region identifier (ID) is a versatile video coding (VVC) sub-picture ID. In block 4244, method 4240 includes performing a conversion between visual media data and a media data file based on the value of the region identifier type. The conversion in step 4244 may include encoding at an encoder or decoding at a decoder, depending on the embodiment.

[0081] In an embodiment, when the video codec of the main video track is not VVC, a value of the region ID type equal to zero is reserved. In an embodiment, the fact that the value of the region ID type is equal to 0 is reserved when the video codec used for coding the main video track and the picture-in-picture video track is not VVC. In an embodiment, when the video codec used for the main video track is VVC, the sample entry type is vvc1, vvi1, or vvs1.

[0082] Note that methods 4200, 4220, and 4240 may be implemented in an apparatus for processing video data that includes a processor and a non-transitory memory having instructions, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In such a case, the instructions, when executed by the processor, cause the processor to execute method 4200. Further, method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video coding device. The computer program product stores computer-executable instructions stored on a non-transitory computer-readable medium that, when executed by a processor, cause a video coding device to execute method 4200.

[0083] FIG. 6 is a block diagram illustrating an example of a video coding system 4300 that may utilize the techniques of the present disclosure. The video coding system 4300 may include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, which may be referred to as a video encoding device. The destination device 4320 may decode the encoded video data generated by the source device 4310, which may be referred to as a video decoding device.

[0084] The source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. The video source 4312 may include a video capture device, an interface for receiving video data from a video content provider, and / or a source such as a computer graphics system for generating video data, or a combination of such sources. The video data may be composed of one or more pictures. The video encoder 4314 encodes the video data from the video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded picture is a coded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to the destination device 4320 via the I / O interface 4316 through the network 4330. The encoded video data may also be stored in the storage medium / server 4340 for access by the destination device 4320.

[0085] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 can obtain the encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to the user. The display device 4322 may be integrated with the destination device 4320 or may be external to the destination device 4320 configured to interface with an external display device.

[0086] The video encoder 4314 and the video decoder 4324 can operate in accordance with video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVM) standard, and other current and / or further standards.

[0087] FIG. 7 is a block diagram showing an example of a video encoder 4400, which may be the video encoder 4314 in the system 4300 shown in FIG. 4. The video encoder 4400 may be configured to execute some or all of the techniques of the present disclosure. The video encoder 4400 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video encoder 4400. In some examples, the processor may be configured to execute some or all of the techniques described in the present disclosure.

[0088] The functional components of the video encoder 4400 may include a splitting unit 4401, a prediction unit 4402 that may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, an intra prediction unit 4406, a residual generation unit 4407, a transformation processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transformation unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy encoding unit 4414.

[0089] In other examples, the video encoder 4400 can include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is the picture in which the current video block is located.

[0090] Furthermore, some components such as motion estimation unit 4404 and motion compensation unit 4405 may be highly integrated, but for the sake of explanation, they are shown separately in the example of video encoder 4400.

[0091] Splitting unit 4401 can split a picture into one or more video blocks. Video encoder 4400 and video decoder 4500 can support various video block sizes.

[0092] Mode selection unit 4403 may select, for example, either an intra or inter coding mode based on an error result, and provide the resulting intra or inter coded block to residual generation unit 4407 that generates residual block data and reconstruction unit 4412 that reconstructs the coded block for use as a reference picture. In some examples, mode selection unit 4403 may select a combined intra and inter prediction (CIIP) mode where the prediction is based on a combination of an inter prediction signal and an intra prediction signal. Similarly, mode selection unit 4403 may also select the resolution of the motion vector of a block (e.g., sub-pixel accuracy or integer pixel accuracy) in the case of inter prediction.

[0093] To perform inter prediction for the current video block, motion estimation unit 4404 may generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 may determine a predicted video block of the current video block based on the motion information and decoded picture samples of a picture from buffer 4413 other than the picture associated with the current video block.

[0094] The motion estimation unit 4404 and the motion compensation unit 4405 can execute different processes on the current video block according to, for example, whether the current video block is an I slice, a P slice, or a B slice.

[0095] In some examples, the motion estimation unit 4404 may perform a unidirectional prediction on the current video block. The motion estimation unit 4404 may search for a reference picture in list 0 or list 1 for a reference block of the current video block. Then, the motion estimation unit 4404 may generate a reference index indicating the reference picture in list 0 or list 1 including the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.

[0096] In other examples, the motion estimation unit 4404 may perform a bidirectional prediction on the current video block. The motion estimation unit 4404 may search for a reference picture in list 0 for a reference video block of the current video block, and may also search for a reference picture in list 1 for another reference video block of the current video block. Next, the motion estimation unit 4404 may generate a reference index indicating the reference pictures in list 0 and list 1 including the reference video blocks, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.

[0097] In some examples, the motion estimation unit 4404 may output a full set of motion information for the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a full set of motion information for the video. Rather, the motion estimation unit 4404 may signal the motion information of a video block by referring to the motion information of another video block. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.

[0098] In one example, the motion estimation unit 4404 may indicate to the video decoder 4500, in the syntax structure related to the current video block, a value indicating that the current video block has the same motion information as another video block.

[0099] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in the syntax structure related to the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 can use the difference between the motion vectors of the video blocks to determine the motion vector of the current video block.

[0100] As described above, the video encoder 4400 can transmit motion vectors predictively. Two examples of predictive transmission techniques that can be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and merge mode transmission.

[0101] The intra prediction unit 4406 may perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data for the current video block based on the decoded samples of other video blocks within the same picture. The prediction data of the current video block may include a predicted video block and various syntax elements.

[0102] The residual generation unit 4407 may generate residual data of the current video block by subtracting the predicted video block of the current video block from the current video block. The residual video block of the current video block may include residual video blocks corresponding to different sample components of the samples of the current video block.

[0103] In other examples, for example, in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform the subtraction operation.

[0104] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block related to the current video block.

[0105] After the transform processing unit 4408 generates the transform coefficient video block related to the current video block, the quantization unit 4409 can quantize the transform coefficient video block related to the current video block based on one or more quantization parameter (QP) values related to the current video block.

[0106] The inverse quantization unit 4410 and the inverse transform unit 4411 can each apply inverse quantization and inverse transform to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.

[0107] After the reconstruction unit 4412 reconstructs the video block, a loop filtering process can be performed to reduce video blocking artifacts within the video block.

[0108] The entropy encoding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, the entropy encoding unit 4414 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0109] FIG. 8 is a block diagram illustrating an example of a video decoder 4500 that can be the video decoder 4324 in the system 4300 shown in FIG. 4. The video decoder 4500 can be configured to perform some or all of the techniques of the present disclosure. In the example shown, the video decoder 4500 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video decoder 4500. In some examples, a processor may be configured to perform some or all of the techniques described in the present disclosure.

[0110] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. The video decoder 4500 can execute a decoding path that is generally the reverse of the encoding path described for the video encoder 4400 in some examples.

[0111] The entropy decoding unit 4501 can obtain an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., an encoded block of video data). The entropy decoding unit 4501 may decode the entropy-coded video data, and from the entropy-decoded video data, the motion compensation unit 4502 may determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 4502 may determine the information, for example, by executing AMVP and merge modes.

[0112] The motion compensation unit 4502 can generate a motion-compensated block and, optionally, perform interpolation based on an interpolation filter. An identifier of the interpolation filter used with sub-pixel precision may be included in the syntax element.

[0113] The motion compensation unit 4502 can calculate interpolation values of sub-integer pixels of a reference block using the interpolation filter used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 according to the received syntax information and generate a prediction block using the interpolation filter.

[0114] The motion compensation unit 4502 can determine, using part of the syntax information, the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, the partition information that describes how each macroblock of the picture of the encoded video sequence is divided, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) between the inter-coded blocks, and other information for decoding the encoded video sequence.

[0115] The intra prediction unit 4503 can form a prediction block from spatially adjacent blocks, for example, using the intra prediction mode received in the bitstream. The inverse quantization unit 4504 inverse quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.

[0116] The reconstruction unit 4506 can sum the residual block and the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form the decoded block. Optionally, a deblocking filter can be applied to filter the decoded block to remove blocking artifacts. The decoded video block is stored in the buffer 4507, provides a reference block for subsequent motion compensation / intra prediction, and generates the decoded video for display on a display device.

[0117] FIG. 9 is a schematic diagram of an example of the encoder 4600. The encoder 4600 is suitable for implementing the technology of VVC. The encoder 4600 includes three in-loop filters, namely, a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Different from the DF 4602 that uses a pre-defined filter, the SAO 4604 and the ALF 4606 utilize the original samples of the picture to reduce the mean squared error between the original samples and the reconstructed samples by respectively applying a finite impulse response (FIR) filter having coded side information for adding an offset and signaling the offset and filter coefficients. The ALF 4606 is located at the final processing stage of each picture and can be regarded as a tool that tries to capture and correct the artifacts created in the previous stage.

[0118] The encoder 4600 further includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive an input video. The intra prediction component 4608 is configured to perform intra prediction, and the ME / MC component 4610 is configured to perform inter prediction using reference pictures obtained from the reference picture buffer 4612. Residual blocks from inter prediction or intra prediction are supplied to a transform (T) component 4614 and a quantization (Q) component 4616, where quantized residual transform coefficients are generated and supplied to an entropy coding component 4618. The entropy coding component 4618 entropy-codes the prediction results and the quantized transform coefficients and transmits them towards a video decoder (not shown). The quantization component output from the quantization component 4616 may be supplied to an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 can output the picture to the DF 4602, SAO 4604, and ALF 4606 for filtering before the picture is stored in the reference picture buffer 4612.

[0119] Next, several preferred solutions are enumerated in some examples.

[0120] The following solutions show examples of the technologies discussed in the present disclosure.

[0121] The following solutions show examples of embodiments of the technologies described in the previous section (e.g., item 1).

[0122] 1. A method for processing video data (e.g., the method 4200 shown in FIG. 3), comprising: determining (4202) a value of a region identifier (ID) type (region_id_type) field, where the region_id_type field is coded using 4 bits or less; and performing (4204) a conversion between visual media data and a media data file based on the value of the region_id_type field.

[0123] 2. The method of solution 1, wherein the remaining groups of 4 bits or more associated with the region_id_type field are reserved.

[0124] The following solutions show embodiments of the technology described in the previous section (e.g., item 2).

[0125] 3. A method for processing video data, comprising: determining a value of a number of region identifiers (num_region_ids) field, where the num_region_ids field is restricted to be greater than 0; and performing a conversion between visual media data and a media data file based on the value of the num_region_ids field.

[0126] 4. The method according to any one of solutions 1 to 3, wherein the num_region_ids field is represented as a num_region_ids minus 1 (num_region_ids_minus1) field.

[0127] 5. The method according to any one of solutions 1 to 4, wherein adding 1 to the num_region_ids_minus1 field specifies the number of region identifier fields (region_id[i]).

[0128] 6. The method according to any one of solutions 1 to 5, wherein region_id[i] is determined according to a loop specified by for(i = 0; i <= num_region_ids_minus1; i++).

[0129] The following solution shows an example of an embodiment of the technology described in the previous section (e.g., item 3).

[0130] 7. A method for processing video data, comprising determining a value of a region identifier type (region_id_type) field, where the region_id_type field is equal to 0 to specify that the region identifier (ID) is a general video coding (VVC) sub-picture ID when the VVC codec is used to code a video main track, and performing a conversion between visual media data and a media data file based on the value of the region_id_type field.

[0131] 8. The method according to any one of Solutions 1 to 7, wherein when the VVC codec is used for coding a video main track, the sample entry type is set to vvc1, vvi1, or vvs1.

[0132] 9. The method according to any one of Solutions 1 to 8, wherein when the VVC codec is not used for coding a video main track and the VVC codec is not used for coding a picture-in-picture (PiP) video track, the value 0 of the region_id_type field is reserved.

[0133] 10. An apparatus for processing video data, comprising a processor and a memory having instructions thereon, wherein the instructions are executed by the processor to cause the processor to perform the method according to any one of Solutions 1 to 9.

[0134] 11. A non-transitory computer-readable medium including a computer program product for use in a video coding apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, which, when executed by a processor, cause the video coding apparatus to execute the method according to any one of Solutions 1 to 9.

[0135] 12. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing apparatus, the method determining a value of a region identifier (ID) type (region_id_type) field, where the region_id_type field is coded using 4 bits or less, and generating a bitstream based on the determination.

[0136] 13. A method of storing a bitstream of a video, the method determining a value of a region identifier (ID) type (region_id_type) field, where the region_id_type field is coded with 4 bits or less, generating a bitstream based on the determination, and storing the bitstream in a non-transitory computer-readable recording medium.

[0137] 14. The method, apparatus or system described in the present disclosure.

[0138] In the solution, the encoder can conform to the format rule by generating a representation coded according to the format rule. In the solution described in the present disclosure, the decoder can analyze the syntax elements in the coded representation using the knowledge of the presence or absence of the syntax elements according to the format rule and generate a decoded video using the format rule.

[0139] In the present disclosure, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation or vice versa. The bitstream representation of a current video block corresponds to bits that are spread at the same or different locations within the bitstream, as defined by the syntax, for example. For example, a macroblock is encoded from the perspective of the transformed and coded error residual values and using the bits of the headers and other fields within the bitstream. Further, during the conversion, the decoder can analyze the bitstream with the knowledge that some fields may or may not be present, based on a determination, as described in the solution means. Similarly, the encoder can generate a coded representation accordingly by determining whether a particular syntax field is included or not and including or excluding the syntax field in the coded representation.

[0140] The present disclosure and other solutions, examples, embodiments, modules, and functional operations described in the present disclosure can be implemented in digital electronic circuits, or in computer software, firmware, or hardware including the structures disclosed in the present disclosure and their structural equivalents, or in one or more combinations thereof. Embodiments of the present disclosure and other embodiments can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or one or more combinations thereof. The term "data processing apparatus" includes, by way of example, a programmable processor, a computer, or multiple processors or computers, and encompasses all apparatus, devices, and machines for processing data. The apparatus can include, in addition to hardware, code that constructs an execution environment for the computer program, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagated signal is an artificially generated signal, e.g., a mechanically generated electrical signal, an optical signal, an electromagnetic wave signal, generated to encode information for transmission to an appropriate receiver device.

[0141] A computer program (also referred to as a program, software, software application, script, or coding) can be written in any form of programming language, including compiled languages or interpreted languages, and can be deployed in any form, such as as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program in question, or multiple coordinated files (e.g., files that hold one or more modules, subprograms, or portions of code). A computer program can be arranged to be executed on one computer, on multiple computers located at one site, or on multiple computers distributed across multiple sites and interconnected by a communication network.

[0142] The processes and logical flows described in this disclosure can be executed by one or more programmable processors executing one or more computer programs, operating on input data and functioning by generating output. The processes and logical flows can also be executed by special-purpose logic circuits, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), and the apparatus can also be implemented as special-purpose logic circuits.

[0143] Processors suitable for the execution of a computer program include, by way of example, both general-purpose and special-purpose microprocessors, as well as one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read-only memory or a random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing the instructions and data. In general, a computer also operates operatively coupled to receive data from, transfer data to, or both, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, and optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks; and compact disk read-only memory (CD ROM) and digital versatile disk read-only memory (DVD-ROM) disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.

[0144] Although this disclosure includes many specific details, these should not be construed as limitations on the scope of any subject matter or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. The specific features described in the context of individual embodiments of this disclosure can also be implemented in combination within a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately, or in any suitable partial combination, in multiple embodiments. Further, features may be described and initially even claimed as acting in a particular combination, but one or more features from the claimed combination may in some cases be excluded from the combination, and the claimed combination may be directed to a sub - combination or variation of a sub - combination.

[0145] Similarly, operations are depicted in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown, or in a sequential order, or that all of the illustrated operations be performed, in order to achieve desirable results. Further, the separation of various system components in the embodiments described in this disclosure should not be understood as requiring such separation in all embodiments.

[0146] Although only some implementations and examples are described, other implementations, enhancements, and variations can be made based on what is described and illustrated in this disclosure.

[0147] The first component is directly coupled to the second component if no component other than a line, trace, or other medium intervenes between the first and second components. If a component other than a line, trace, or another medium intervenes between the first and second components, the first component is indirectly coupled to the second component. The term "coupled" and its variations include both direct and indirect coupling. The use of the term "about" means a range that includes ±10% of the subsequent numerical value, unless otherwise specified.

[0148] Although several embodiments are provided in the present disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The examples are illustrative and not restrictive, and the intention is not limited to the details given herein. For example, various elements or components may be combined or integrated in other systems, or certain features may be omitted or not implemented.

[0149] Furthermore, in each embodiment, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate may be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of the present disclosure. Other items shown or discussed as being combined may be directly connected or may be indirectly coupled or communicated through any interface, device, or intermediate component, whether electrical, mechanical, or otherwise. Other examples of changes, substitutions, and modifications are ascertainable by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

1. A method for processing visual media data, comprising: determining, in a region identifier (ID) type field (region_id_type), a type of a value taken by the region identifier (ID); wherein the type of the value taken by the region ID is coded in the region ID type field using 4 bits or less; performing a conversion between the visual media data and a media data file based on the type of the value taken by the region ID in the region ID type field; a method.

2. The region ID type field is coded using 3 bits, and 5 bits are reserved; The method according to claim 1.

3. The region ID type field equal to 0 designates that the region ID is a general video coding (VVC) sub-picture ID; The method according to any one of claims 1 to 2.

4. The region ID is used for signal transmission of a picture-in-picture in an International Organization for Standardization base media file format (ISO BMFF); The method according to any one of claims 1 to 3.

5. The media data file includes an International Organization for Standardization base media file format (ISO BMFF); The method according to any one of claims 1 to 4.

6. A method for processing visual media data, comprising: determining a value of a number of the region identifiers, wherein the value of the number of the region identifiers is restricted to be greater than 0; performing a conversion between the visual media data and a media data file based on the value of the number of the region identifiers; a method.

7. The number of the region identifiers is represented as a number obtained by subtracting 1 from the number of the region identifiers (num_region_ids_minus1); The method according to claim 6.

8. Adding 1 to the num_region_ids_minus1 designates a region identifier (region_id[i]); The method according to claim 7.

9. The region identifier designates an i-th ID of a coded video data unit representing a target picture-in-picture region; The method according to claim 8.

10. The region identifier is determined according to a loop specified by for (i = 0; i <= num_region_ids_minus1; i++); The method according to any one of claims 8 to 9.

11. The number of the region identifiers is represented by num_region_ids, The method according to any one of claims 8 to 10.

12. The number of the region identifiers is represented by num_region_ids_minus1, The method according to any one of claims 8 to 10.

13. A method for processing visual media data, Determining a type (region_id_type) of a value taken by a region identifier (ID), wherein, when a video codec used for coding a main video track and a picture-in-picture (PiP) video track is VVC, a value of region_id_type equal to 0 specifies that the region identifier (ID) is a versatile video coding (VVC) sub-picture ID, Performing conversion between the visual media data and a media data file based on the value of the type of the region identifier. Method.

14. That the value of the region_id_type is equal to 0 is reserved when the video codec of the main video track is not VVC, The method according to claim 13.

15. That the value of the region_id_type is equal to 0 is reserved when the video codecs used for coding the main video track and the picture-in-picture video track are not VVC, The method according to any one of claims 13 to 14.

16. When the video codec used for the main video track is VVC, the sample entry type is vvc1, vvi1, or vvs1, The method according to any one of claims 13 to 15.

17. The conversion encodes the visual media data into the visual media data file of the bitstream, The method according to any one of claims 1 to 16.

18. The conversion decodes the visual media data file from the bitstream to obtain the visual media data, The method according to any one of claims 1 to 16.

19. A video data processing apparatus, One or more processors, A non-transitory memory having instructions, When executed by the one or more processors, the instructions cause the one or more processors to execute the method according to any one of claims 1 to 18. Apparatus.

20. A non-transitory computer-readable medium including a computer program product for use by an image coding apparatus, wherein the computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium to cause the image coding apparatus to execute the method according to any one of claims 1 to 18 when executed by a processor. Non-transitory computer-readable medium. **Claim 21** A non-transitory computer-readable recording medium storing a bitstream of an image generated by a method executed by an image processing apparatus, wherein the method determines a value of a region identifier (ID) type (region_id_type) in the region identifier (ID) type field, wherein the value of the region ID type is coded in the region ID type field using 4 bits or less, and performs conversion between visual media data and a media data file based on the value of the region ID type in the region ID type field. Non-transitory computer-readable recording medium. **Claim 22** A method of storing a bitstream of an image, the method determining a value of a region identifier (ID) type (region_id_type) in the region identifier (ID) type field, wherein the value of the region ID type is coded in the region ID type field using 4 bits or less, generating the bitstream based on the determination, and storing the bitstream in a non-transitory computer-readable recording medium. Method. **Claim 23** A non-transitory computer-readable recording medium storing a bitstream of an image generated by a method executed by an image processing apparatus, wherein the method determines a value of the number of region identifiers, wherein the value of the number of region identifiers is constrained to be greater than 0, and performs conversion between visual media data and a media data file based on the value of the number of region identifiers. Non-transitory computer-readable recording medium. **Claim 24** A method of storing a bitstream of an image, the method determining a value of the number of region identifiers, wherein the value of the number of region identifiers is constrained to be greater than 0, generating the bitstream based on the determination, and storing the bitstream in a non-transitory computer-readable recording medium. Method.

25. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method executed by a video processing apparatus, wherein the method comprises: determining a value of a region identifier type (region_id_type), wherein when the video codec used for coding a main video track and a picture-in-picture (PiP) video track is VVC and the value of the region identifier type is equal to 0, it is specified that the region identifier (ID) is a general video coding (VVC) sub-picture ID, performing conversion between visual media data and a media data file based on the value of the type of the region identifier, a non-transitory computer-readable recording medium.

26. A method of storing a bitstream of video, the method comprising: determining a value of a region identifier type (region_id_type), wherein when the video codec used for coding a main video track and a picture-in-picture (PiP) video track is VVC and the value of the region identifier type is equal to 0, it is specified that the region identifier (ID) is a general video coding (VVC) sub-picture ID, generating the bitstream based on the determination, storing the bitstream in a non-transitory computer-readable recording medium, a method.

27. The method, apparatus, or system according to the present disclosure.

Citation Information

Patent Citations

  • Method and apparatus for encoding, decoding, or displaying picture-in-picture

    WO2023203423A1