Extended Signaling of Extended Dependent Random Access Point Supplementary Enhancement Information
By determining the EDRAP leading pictures decodable flag and applying specific constraints, the method addresses the challenges of video data processing, enhancing efficiency and optimizing bitstream management.
Patent Information
- Application Number
- JP2023580378
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-28
- Filing Date
- 2022-06-27
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2042-06-27
AI Technical Summary
Existing digital video technologies face challenges in efficiently processing and managing video data due to increasing bandwidth demands and complexities in picture order and reference constraints.
The method involves determining the value of an extended-dependent random-access point (EDRAP) leading pictures decodable flag syntax element and performing a conversion between visual media data and a bitstream based on this flag, imposing specific ordering constraints on EDRAP pictures.
This approach enhances the efficiency of video data processing by optimizing bitstream generation and decoding, ensuring proper decoding and output ordering of EDRAP pictures, and reducing bandwidth requirements.
Smart Images

Figure 0007697065000013 
Figure 0007697065000014 
Figure 0007697065000015
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application This application 、 was filed on June 28, 2021 is based on International Patent Application No. PCT / CN2022 / 101412, filed on June 27, 2022. All of the foregoing patent applications are and claims the benefit of International Application No. PCT / CN2021 / 102636, which is incorporated herein by reference in its entirety. entirely by reference
Figure 1
[0002] [Technical Field] This patent document relates to the generation, storage, and consumption of digital audio - video media information in file format.
Background Art
[0003] Digital video consumes the most bandwidth on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use is likely to continue to increase.
Summary of the Invention
[0004] A first aspect relates to a method for processing video data, including determining a value of an extended - dependency random - access point (EDRAP) leading - pictures - decodable - flag (edrap_leading_pictures_decodable_flag) syntax element, and performing a conversion between visual media data and a bitstream based on the edrap_leading_pictures_decodable_flag syntax element.
[0005] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the bitstream includes at least one EDRAP picture, and the value of the edrap_leading_pictures_decodable_flag indicates whether an ordering constraint is imposed on the EDRAP picture.
[0006] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that when the value of edrap_leading_pictures_decodable_flag is zero, no ordering constraint is imposed on the EDRAP pictures.
[0007] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that when the value of edrap_leading_pictures_decodable_flag is 1, an ordering constraint is imposed on the EDRAP pictures.
[0008] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that a constraint specifies that any picture that is in the same layer as the EDRAP picture and follows the EDRAP picture in decoding order shall be in the same layer as the EDRAP picture in output order and follow another picture that precedes the EDRAP picture in decoding order.
[0009] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that a constraint specifies that any picture that is in the same layer as the EDRAP picture, follows the EDRAP picture in decoding order, and precedes the EDRAP picture in output order shall be in the active entries of the picture reference picture list of the picture and, except for the list of referenceable pictures, shall not include other pictures that are in the same layer as the EDRAP picture and precede the EDRAP picture in decoding order.
[0010] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the edrap_leading_pictures_decodable_flag syntax element is included in the EDRAP SEI message.
[0011] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the list of referenceable pictures includes IRAP (Intra Random Access Point) or EDRAP pictures in decoding order that are within the same coded layer video sequence (CLVS).
[0012] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each picture in the list of referenceable pictures is identified by the i-th EDRAP reference access point identifier (edrap_ref_rap_id[i]) syntax element.
[0013] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the edrap_leading_pictures_decodable_flag syntax element is u(v)-coded.
[0014] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each EDRAP picture is a trailing picture.
[0015] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each EDRAP picture has a time sublayer identifier equal to zero.
[0016] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each EDRAP picture, excluding the list of referenceable pictures, does not include pictures in the same layer in the active entries of the reference picture list of the EDRAP picture.
[0017] Optionally, in any of the foregoing aspects, another implementation of the aspect is that any picture that is in the same layer as the EDRAP picture and follows the EDRAP picture in both the decoding order and the output order, except for the list of referenceable pictures, is in the active entry of the picture's reference picture list in the same layer and the bitstream is constrained to not include other pictures that precede the EDRAP picture in the decoding order or the output order.
[0018] Optionally, in any of the foregoing aspects, another implementation of the aspect is that the transformation includes encoding visual media data into a bitstream.
[0019] Optionally, in any of the foregoing aspects, another implementation of the aspect is that the transformation includes decoding visual media data from a bitstream.
[0020] A second aspect relates to an apparatus for processing video data, comprising a processor and a non-transitory memory having instructions, which, when executed by the processor, cause the processor to execute any of the methods of the foregoing aspects.
[0021] A third aspect relates to a non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to execute any of the methods of the foregoing aspects.
[0022] A fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method executed by a video processing apparatus, the method including determining a value of an extended dependent random access point (EDRAP) leading pictures decodable flag (edrap_leading_pictures_decodable_flag) syntax element, and generating a bitstream based on the determining step.
[0023] A fifth aspect relates to a method for storing a bitstream of video, the method including determining a value of an extended dependent random access point (EDRAP) leading pictures decodable flag (edrap_leading_pictures_decodable_flag) syntax element, generating a bitstream based on the determining step, and storing the bitstream in a non-transitory computer-readable recording medium.
[0024] For clarity, any one of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0025] These and other features will be more clearly understood from the following detailed description in conjunction with the accompanying drawings and the claims.
Brief Description of the Drawings
[0026] For a more complete understanding of the present disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
[0027] First, exemplary implementations of one or more embodiments are provided below, but it should be understood that the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or not yet developed. The present disclosure should in no way be limited to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the full scope of the appended claims and their equivalents.
[0028] This patent document relates to image and / or video coding techniques. Specifically, this document relates to the annotation area, depth representation information, and signaling of extended dependent random access point (EDRAP) indications in supplementary enhancement information (SEI) messages. These examples can be applied individually or in various combinations to video bitstreams coded by any codec, such as the versatile video coding (VVC) standard and the versatile SEI message (VSEI) standard for coded video bitstreams.
[0029] The present disclosure includes the following abbreviations: Alpha Channel Information (ACI), Adaptive Parameter Set (APS), Access Unit (AU), Coded Layer Video Sequence (CLVS), Coded Layer Video Sequence Start (CLVSS), Cyclic Redundancy Check (CRC), Color Transformation Information (CTI), Coded Video Sequence (CVS), Dependent Random Access Point (DRAP), Depth Representation Information (DRI), Extended Dependent Random Access Point (EDRAP), Finite Impulse Response (FIR), Intra Random Access Point (IRAP), Multi-View Acquisition Information (MAI), Network Abstraction Layer (NAL), Picture Parameter Set (PPS), Picture Unit (PU), Random Access Skip Reading (RASL), Region-Wise Packing (RWP), Sample Aspect Ratio (SAR), Sample Aspect Ratio Information (SARI), Scalability Dimension Information (SDI), Supplemental Enhancement Information (SEI), Step-Wise Time Sub-Layer Access (STSA), Video Coding Layer (VCL), General Supplemental Enhancement Information (VSEI) also known as Rec. ITU-T H.274|ISO / IEC 23002-7, Video User Utility Information (VUI), and Versatile Video Coding (VVC) also known as Rec. ITU-T H.266|ISO / IEC 23090-3.
[0030] Video coding standards have mainly evolved through the development of International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and ISO / International Electrotechnical Commission (IEC) standards. ITU-T produced H.261 and H.263, ISO / IEC produced Motion Picture Experts Group (MPEG)-1 and MPEG-4 Visual, and the two organizations jointly produced H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / High Efficiency Video Coding (HEVC) standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore further video coding technologies beyond HEVC, the JVET (Joint Video Exploration Team) was jointly established by the VCEG (Video Coding Experts Group) and MPEG. Many methods were adopted by the JVET and incorporated into a reference software named JEM (Joint Exploration Model). The JVET was later renamed JVET (Joint Video Experts Team) when the Versatile Video Coding (VVC) project was officially launched. VVC is a coding standard targeting a 50% bitrate reduction compared to HEVC. VVC has been completed by the JVET.
[0031] The VVC standard, also known as ITU-T H.266|ISO / IEC 23090-3, and the related Versatile Supplementary Enhancement Information (VSEI) standard, also known as ITU-T H.274|ISO / IEC 23002-7, are designed for use in a wide range of applications such as television broadcasting, video conferencing, playback from storage media, adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple coded video bitstreams, multi-view video, scalable hierarchical coding, and viewport-adaptive 360-degree (360°) immersive media. The Essential Video Coding (EVC) standard (ISO / IEC 23094-1) is another video coding standard developed by MPEG.
[0032] Exemplary modifications to the VSEI standard include specifications for additional SEI messages, including annotated region SEI messages, alpha channel information SEI messages, depth representation information SEI messages, multi-view acquisition information SEI messages, scalability dimension information SEI messages, Extended Dependent Random Access Point (DRAP) indication SEI messages, display orientation SEI messages, and color conversion information SEI messages.
[0033] An exemplary annotated region SEI message syntax is as follows. [Table 1] TIFF0007697065000002.tif236160TIFF0007697065000003.tif76161
[0034] Exemplary annotated region SEI message semantics are as follows. The annotated region SEI message conveys parameters that identify the annotated region using a bounding box that represents the size and location of the identified object. The use of this SEI message may require the definition of the following variables. Such variables include the cropped picture width and height in luma samples, indicated by CroppedWidth and CroppedHeight respectively herein, the chroma subsampling width and height, indicated by SubWidthC and SubHeightC respectively, the conforming cropping window left offset, indicated by ConfWinLeftOffset, and the conforming cropping window top offset, indicated by ConfWinTopOffset.
[0035] An ar_cancel_flag set equal to 1 indicates that the annotated region SEI message cancels the persistence of any previous annotated region SEI message associated with one or more layers to which the annotated region SEI message applies. An ar_cancel_flag set equal to 0 indicates that the annotated region information continues. When ar_cancel_flag is equal to 1 or when a new CVS of the current layer starts, the variables LabelAssigned[i], ObjectTracked[i], and ObjectBoundingBoxAvail are set equal to 0 for i in the range of 0 to 255, inclusive.
[0036] The ar_not_optimized_for_viewing_flag set equal to 1 indicates that the decoded picture to which the annotated region SEI message is applied is not optimized for user viewing, but rather optimized for some other purpose such as algorithm object classification performance. The ar_not_optimized_for_viewing_flag set equal to 0 indicates that the decoded picture to which the annotated region SEI message is applied may or may not be optimized for user viewing.
[0037] The ar_true_motion_flag set equal to 1 indicates that the motion information in the coded picture to which the annotated region SEI message is applied is selected for the purpose of accurately representing the object motion of the objects in the annotated region. The ar_true_motion_flag set equal to 0 indicates that the motion information in the coded picture to which the annotated region SEI message is applied may or may not be selected for the purpose of accurately representing the object motion for the objects in the annotated region.
[0038] The ar_occluded_object_flag set equal to 1 indicates that the ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] syntax elements may each represent the size and location of an object or a part of an object that may not be visible or may be only partially visible within the cropped decoded picture. The ar_occluded_object_flag set equal to 0 indicates that the ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] syntax elements represent the size and location of an object that is fully visible within the cropped decoded picture. Bitstream conformance may require that the value of the ar_occluded_object_flag be the same for all annotated_regions() syntax structures within the CVS.
[0039] The ar_partial_object_flag_present_flag set equal to 1 indicates the presence of the ar_partial_object_flag[ar_object_idx[i]] syntax element. The ar_partial_object_flag_present_flag set equal to 0 indicates the absence of the ar_partial_object_flag[ar_object_idx[i]] syntax element. Bitstream conformity may require that the value of ar_partial_object_flag_present_flag be the same for all annotated_regions() syntax structures within the CVS.
[0040] The ar_object_label_present_flag set equal to 1 indicates the presence of label information corresponding to the object in the annotated region. The ar_object_label_present_flag set equal to 0 indicates the absence of label information corresponding to the object in the annotated region.
[0041] The ar_object_confidence_info_present_flag set equal to 1 indicates the presence of the ar_object_confidence[ar_object_idx[i]] syntax element. The ar_object_confidence_info_present_flag set equal to 0 indicates the absence of the ar_object_confidence[ar_object_idx[i]] syntax element. Bitstream conformity may require that the value of ar_object_confidence_present_flag be the same for all annotated_regions() syntax structures within the CVS.
[0042] ar_object_confidence_length_minus1 + 1 specifies the length of the ar_object_confidence[ar_object_idx[i]] syntax element in bits. Bitstream conformance may require that the value of ar_object_confidence_length_minus1 be the same for all annotated_regions() syntax structures within the CVS.
[0043] The ar_object_label_language_present_flag set equal to 1 indicates the presence of the ar_object_label_language syntax element. The ar_object_label_language_present_flag set equal to 0 indicates the absence of the ar_object_label_language syntax element. Let ar_bit_equal_to_zero be equal to zero.
[0044] ar_object_label_language contains a language tag, followed by a null-terminating byte equal to 0x00. The length of the ar_object_label_language syntax element shall be 255 bytes or less, not including the null-terminating byte. If not present, the language of the label is not specified.
[0045] ar_num_label_updates indicates the total number of labels associated with the annotated regions being signaled. The value of ar_num_label_updates shall be in the range of 0 to 255, inclusive. ar_label_idx[i] indicates the index of the signaled label. The value of ar_label_idx[i] shall be in the range of 0 to 255, inclusive.
[0046] The ar_label_cancel_flag set equal to 1 cancels the duration range of the ar_label_idx[i]-th label. The ar_label_cancel_flag set equal to 0 indicates that the signaled value can be assigned to the ar_label_idx[i]-th label. ar_label[ar_label_idx[i]] specifies the content of the ar_label_idx[i]-th label. It is assumed that the length of the ar_label[ar_label_idx[i]] syntax element is 255 bytes or less excluding the null-terminating byte.
[0047] ar_num_object_updates indicates the number of object updates to be signaled. It is assumed that ar_num_object_updates is within the range of 0 to 255 including both end values. ar_object_idx[i] is the index of the object parameter to be signaled. It is assumed that ar_object_idx[i] is within the range of 0 to 255 including both end values. The ar_object_cancel_flag set equal to 1 cancels the duration range of the ar_object_idx[i]-th object. The ar_object_cancel_flag set equal to 0 indicates that the parameter associated with the ar_object_idx[i]-th tracked object is to be signaled. The ar_object_label_update_flag set equal to 1 indicates that the object label is to be signaled. The ar_object_label_update_flag equal to 0 indicates that the object label is not to be signaled.
[0048] ar_object_label_idx[ar_object_idx[i]] indicates the index of the label corresponding to the object at ar_object_idx[i]. When ar_object_label_idx[ar_object_idx[i]] does not exist, its value is inferred, if any, from the previous annotated region SEI message in the same CVS in the output order. ar_bounding_box_update_flag set equal to 1 indicates that the object bounding box parameters are signaled. ar_bounding_box_update_flag set equal to 0 indicates that the object bounding box parameters are not signaled.
[0049] ar_bounding_box_cancel_flag set equal to 1 cancels the persistence range of ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], ar_bounding_box_height[ar_object_idx[i]], ar_partial_object_flag[ar_object_idx[i]], and ar_object_confidence[ar_object_idx[i]]. ar_bounding_box_cancel_flag set equal to 0 indicates that the ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], ar_bounding_box_height[ar_object_idx[i]], ar_partial_object_flag[ar_object_idx[i]], and ar_object_confidence[ar_object_idx[i]] syntax elements are signaled.
[0050] ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] respectively specify the coordinates of the upper left corner of the bounding box of the ar_object_idx[i]-th object in the cropped decoded picture with respect to the adaptive SPS-specified compliant cropping window, the width, and the height.
[0051] The value of ar_bounding_box_left[ar_object_idx[i]] shall be within the range from 0 to CroppedWidth / SubWidthC - 1, inclusive of both end values. The value of ar_bounding_box_top[ar_object_idx[i]] shall be within the range from 0 to CroppedHeight / SubHeightC - 1, inclusive of both end values. The value of ar_bounding_box_width[ar_object_idx[i]] shall be within the range from 0 to CroppedWidth / SubWidthC - ar_bounding_box_left[ar_object_idx[i]], inclusive of both end values. The value of ar_bounding_box_height[ar_object_idx[i]] shall be within the range from 0 to CroppedHeight / SubHeightC - ar_bounding_box_top[ar_object_idx[i]], inclusive of both end values. The identified object rectangle includes luma samples having horizontal picture coordinates from SubWidthC*(ConfWinLeftOffset + ar_bounding_box_left[ar_object_idx[i]]) to SubWidthC*(ConfWinLeftOffset + ar_bounding_box_left[ar_object_idx[i]] + ar_bounding_box_width[ar_object_idx[i]]) - 1, inclusive of both end values, and vertical picture coordinates from SubHeightC*(ConfWinTopOffset + ar_bounding_box_top[ar_object_idx[i]]) to SubHeightC*(ConfWinTopOffset + ar_bounding_box_top[ar_object_idx[i]] + ar_bounding_box_height[ar_object_idx[i]]) - 1, inclusive of both end values.The values of ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] persist in the output order within the CVS for each value of ar_object_idx[i]. When not present, the values of ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], or ar_bounding_box_height[ar_object_idx[i]] are inferred from the previous annotated region SEI message in the output order in the CVS, if any.
[0052] The ar_partial_object_flag[ar_object_idx[i]] set equal to 1 indicates that the ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] syntax elements represent the size and location of an object that is only partially visible within the cropped decoded picture. The ar_partial_object_flag[ar_object_idx[i]] set equal to 0 indicates that the ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] syntax elements represent the size and location of an object that may or may not be partially visible within the cropped decoded picture. When not present, the value of ar_partial_object_flag[ar_object_idx[i]] is inferred, if possible, from the previous annotated region SEI message in output order in the CVS.
[0053] ar_object_confidence[ar_object_idx[i]] indicates the confidence associated with the ar_object_idx[i]-th object in units of 2-(ar_object_confidence_length_minus1+1). The higher the value of ar_object_confidence[ar_object_idx[i]], the higher the confidence. The length of the ar_object_confidence[ar_object_idx[i]] syntax element is ar_object_confidence_length_minus1+1 bits. When not present, the value of _object_confidence[ar_object_idx[i]] is inferred from the previous annotated region SEI message in the output order in CVS, if any.
[0054] Next, the depth representation information SEI message will be described. An exemplary depth representation information SEI message syntax is as follows.
Table 2
[0055] An exemplary depth representation information element syntax is as follows.
Table 3
[0056] An exemplary depth representation information SEI message semantics is as follows. The syntax elements in the depth representation information (DRI) SEI message specify various parameters for auxiliary pictures of type AUX_DEPTH for the purpose of rendering on a three-dimensional (3D) display such as view synthesis after processing the decoded primary and auxiliary pictures. For example, the depth or parallax range of the depth picture is specified.
[0057] The use of this SEI message may require the definition of the following variables. In this specification, the bit depth for samples of the luma component, indicated by BitDepthY. When CVS does not contain an SDI SEI message for which sdi_aux_id[i] is equal to 2 for at least one value of i, no picture in CVS should be associated with a DRI SEI message. When an access unit (AU) contains both an SDI SEI message for which sdi_aux_id[i] is equal to 2 for at least one value of i and a DRI SEI message, the SDI SEI message shall precede the DRI SEI message in decoding order. When present, the DRI SEI message shall be associated with one or more layers indicated as depth-auxiliary layers by the SDI SEI message. The following semantics apply separately to each nuh_layer_id targetLayerId among the nuh_layer_id values to which the DRI SEI message applies. When present, the DRI SEI message may be included in any access unit. When present, it is recommended that an SEI message be included for random access purposes in an access unit in which the coded picture with nuh_layer_id equal to targetLayerId is an IRAP picture. The information indicated in the DRI SEI message applies to all pictures with nuh_layer_id equal to targetLayerId from the access unit containing the SEI message, except for the next picture in decoding order up to but excluding the end of the CLVS of targetLayerId or of the nuh_layer_id equal to targetLayerId, associated with the DRI SEI message that is applicable earlier in decoding order.
[0058] The z_near_flag set equal to 0 specifies that the syntax element specifying the nearest depth value does not exist in the syntax structure. The z_near_flag set equal to 1 specifies that the syntax element specifying the nearest depth value exists in the syntax structure. The z_far_flag set equal to 0 specifies that the syntax element specifying the farthest depth value does not exist in the syntax structure. The z_far_flag set equal to 1 specifies that the syntax element specifying the farthest depth value exists in the syntax structure. The d_min_flag set equal to 0 specifies that the syntax element specifying the minimum disparity value does not exist in the syntax structure. The d_min_flag set equal to 1 specifies that the syntax element specifying the minimum disparity value exists in the syntax structure. The d_max_flag set equal to 0 specifies that the syntax element specifying the maximum disparity value does not exist in the syntax structure. The d_max_flag set equal to 1 specifies that the syntax element specifying the maximum disparity value exists in the syntax structure. The depth_representation_type specifies the representation definition of the decoded luma samples of the auxiliary picture, as specified in Table 1. In Table 1, the disparity specifies the horizontal displacement between two texture views, and the Z value specifies the distance from the camera. The variable maxVal is set equal to (1<<BitDepthY)-1.
Table 4
[0059] The disparity_ref_view_id specifies the ViewId value from which the disparity value is derived. Note that the disparity_ref_view_id exists only when d_min_flag is equal to 1 or d_max_flag is equal to 1, and is useful for depth_representation_type values equal to 1 and 3. The variables in the x column of Table 2 are derived as follows from the respective variables in the s, e, n, and v columns of Table 2. When the value of e is in the range of 0 to 127 (excluding 0), x is set equal to (-1)s * 2e - 31 * (1 + n ÷ 2v). Otherwise (when e is equal to 0), x is set equal to (-1)s * 2-(30 + v) * n.
Table 5
[0060] The DMin value and the DMax value, when they exist, are specified in units of the luma sample width of the coded picture whose ViewId is equal to the ViewId of the auxiliary picture. The units of the ZNear and ZFar values, when they exist, are the same but are not specified. depth_nonlinear_representation_num_minus1 + 2 specifies the number of piecewise linear segments for mapping depth values to a scale uniformly quantized with respect to disparity. For i ranging from 0 to depth_nonlinear_representation_num_minus1 + 2 inclusive, depth_nonlinear_representation_model[i] specifies a piecewise linear segment for mapping the decoded luma sample values of the auxiliary picture to a scale uniformly quantized with respect to disparity. The values of depth_nonlinear_representation_model[0] and depth_nonlinear_representation_model[depth_nonlinear_representation_num_minus1 + 2] are both inferred to be equal to 0.
[0061] When the depth_representation_type is equal to 3, the auxiliary picture contains non-linearly transformed depth samples. The variable DepthLUT[i] is used to convert the decoded depth sample value from a non-linear representation to a linear representation, for example, a uniformly quantized disparity value, as specified below. The shape of this conversion is defined by a line segment approximation in the two-dimensional linear disparity-non-linear disparity space. The first (0,0) and the last (maxVal,maxVal) nodes of the curve are predefined. The positions of the additional nodes are sent in the form of the deviation (depth_nonlinear_representation_model[i]) from the straight line curve. These deviations are uniformly distributed along the entire range from 0 to maxVal, including the end values, at intervals corresponding to the value of nonlinear_depth_representation_num_minus1.
[0062] For variable DepthLUT[i] for i in the range from 0 to maxVal, including the end values, is specified as follows:
Number
[0063] When the depth_representation_type is equal to 3, for all decoded luma sample values dS of the auxiliary picture within the range from 0 to maxVal, including the end values, DepthLUT[dS] represents a disparity uniformly quantized in the range from 0 to maxVal, including the end values.
[0064] The depth representation information element semantics are as follows. The syntax structure specifies the values of the elements in the DRI SEI message. The syntax structure sets the values of the OutSign, OutExp, OutMantissa, and OutManLen variables representing floating-point values. When a syntax structure is included in another syntax structure, the variable names OutSign, OutExp, OutMantissa, and OutManLen must be interpreted as replaced by the variable names used when the syntax structure is included.
[0065] The da_sign_flag set to 0 indicates that the sign of the floating-point value is positive. The da_sign_flag set to 1 indicates that the sign is negative. The variable OutSign is set equal to da_sign_flag. The da_exponent specifies the exponent of the floating-point value. The value of da_exponent shall be in the range of 0 to 27 - 2, including both end values. The value 27 - 1 is reserved. The decoder shall treat the value 27 - 1 as an indication of an unspecified value. The variable OutExp is set equal to da_exponent. The da_mantissa_len_minus1 + 1 specifies the number of bits within the da_mantissa syntax element. The value of da_mantissa_len_minus1 shall be in the range of 0 to 31, including both end values. The variable OutManLen is set equal to da_mantissa_len_minus1 + 1. The da_mantissa specifies the mantissa of the floating-point value. The variable OutMantissa is set equal to da_mantissa.
[0066] The extended DRAP indication SEI message is as follows. An exemplary extended DRAP indication SEI message syntax is as follows. [Table 6]
[0067] An example of the semantics of the extended DRAP indication SEI message is as follows. The picture associated with the extended DRAP (EDRAP) indication SEI message is called an EDRAP picture. The presence of the EDRAP indication SEI message indicates that the constraints on the picture order and picture references specified in this sub-closure apply. Due to these constraints, the decoder can properly decode the EDRAP picture and the pictures in the same layer, and can follow it in both the decoding order and the output order without having to decode the other pictures in the same layer except for the list of pictures referenceablePictures. This includes a list of IRAP or EDRAP pictures in decoding order that are within the same CLVS and are identified by the edrap_ref_rap_id[i] syntax element.
[0068] The constraints indicated by the presence of the EDRAP indication SEI message are all considered to apply and are as follows. The EDRAP picture is a trailing picture. The EDRAP picture has a temporal sublayer identifier equal to 0. The EDRAP picture does not include pictures in the same layer in the active entries of its reference picture list, except for referenceablePictures. Any picture in the same layer that follows the EDRAP picture in both the decoding order and the output order does not include pictures in the same layer in the active entries of its reference picture list, except for referenceablePictures, that precede the EDRAP picture in either the decoding order or the output order.
[0069] When edrap_leading_pictures_decodable_flag is equal to 1, the following applies. Any picture in the same layer that follows an EDRAP picture in the decoding order shall follow, in the output order, any picture in the same layer that precedes the EDRAP picture in the decoding order. Any picture in the same layer that follows an EDRAP picture in the decoding order and precedes the EDRAP picture in the output order, except for referenceablePictures, shall not include, in the active entries of its reference picture list, any picture in the same layer that precedes the EDRAP picture in the decoding order. Any picture in the list referenceablePictures shall not include, in the active entries of its reference picture list, any picture in the same layer that is not a picture at a position earlier than that of the picture in the list referenceablePictures. Thus, the first picture in referenceablePictures shall not include, in the active entries of its reference picture list, any picture from the same layer, even when it is an EDRAP picture rather than an IRAP picture.
[0070] edrap_rap_id_minus1+1 specifies the random access point (RAP) picture identifier, indicated as RapPicId, of the EDRAP picture. Each IRAP or EDRAP picture is associated with a RapPicId value. The RapPicId value for an IRAP picture is inferred to be equal to 0. The RapPicId values for any two EDRAP pictures associated with the same IRAP picture shall be different. edrap_reserved_zero_12bits shall be equal to 0 in a bitstream compliant with this disclosure. Other values of edrap_reserved_zero_12bits are reserved. The decoder may ignore the value of edrap_reserved_zero_12bits. edrap_num_ref_rap_pics_minus1+1 indicates the number of IRAP or EDRAP pictures that are within the same CLVS as the EDRAP picture and can be included in the active entries of the reference picture list of the EDRAP picture. edrap_ref_rap_id[i] indicates the RapPicId of the i-th RAP picture that can be included in the active entries of the reference picture list of the EDRAP picture. The i-th RAP picture shall be either an IRAP picture associated with the current EDRAP picture or an EDRAP picture associated with the same IRAP picture as the current EDRAP picture.
[0071] The following are exemplary technical problems solved by the disclosed technical solutions. Exemplary designs for the annotated region SEI message, depth representation information SEI message, and EDRAP indication SEI message have at least the following problems. In the case of the annotated region SEI message, the value range (ar_object_label_idx[ar_object_idx[i]]) of the AR object label index of the ue(v) coding syntax element of the i-th annotated region object index is missing. One practical problem associated with not having a specified value range for the ue(v) coding syntax element is that it may be uncertain how many bits can be used for the corresponding variable in the implementation form. If the maximum number of bits used in the implementation form is not sufficient, the decoder may crash when encountering a value larger than the maximum value allowed by the number of bits used. In the case of the depth representation information SEI message, the descriptor (e.g., coding method) of the i-th depth non-linear representation model (depth_nonlinear_representation_model[i]) syntax element is not specified. Without specifying the coding method, the decoder may not be able to determine how to parse the syntax element. In the case of the depth representation information SEI message, the value ranges of the depth representation type (a ue(v) coding syntax element), disparity reference view identifier, depth non-linear representation num - 1, and depth_nonlinear_representation_model[i] are not specified. In the case of the EDRAP indication SEI message, the semantics of the edrap_leading_pictures_decodable_flag syntax element are missing.
[0072] This specification discloses mechanisms for addressing one or more of the problems listed above. For example, the present disclosure specifies an exemplary value range for ar_object_label_idx[ar_object_idx[i]]. Further, the present disclosure specifies an exemplary descriptor for depth_nonlinear_representation_model[i]. Additionally, the present disclosure specifies exemplary value ranges for depth_representation_type, disparity_ref_view_id, depth_nonlinear_representation_num_minus1, and depth_nonlinear_representation_model[i]. Additionally, the present disclosure specifies exemplary semantics for the edrap_leading_pictures_decodable_flag.
[0073] FIG. 1 is a schematic diagram showing an exemplary bitstream 100. The bitstream 100 may include compressed video and associated syntax. For example, the bitstream 100 may be encoded by an encoder, transmitted over one or more networks, and decoded by a decoder for display to a user. For example, the bitstream 100 may be defined as a sequence of bits forming a representation of a sequence of access units (AUs) that form one or more coded video sequences (CVSs). An AU is a set of one or more pictures associated with a corresponding output time in the video sequence. The bitstream may take the form of a network abstraction layer (NAL) unit stream or a byte stream.
[0074] The bitstream 100 includes one or more sequence parameter sets (SPSs) 113, a plurality of picture parameter sets (PPSs) 115, a plurality of slices 125, an annotation region (AR) SEI message 131, a DRI SEI message 133, and an EDRAP indication SEI message 135. The SPS 113 includes sequence data-related parameters common to all pictures in the encoded video sequence included in the bitstream 100. The parameters in the SPS 113 can include, for example, picture sizing, bit depth, coding tool parameters, bitrate limits, and the like. Note that each sequence points to an SPS 113, but in some examples, a single SPS 113 can include data for multiple sequences. The PPS 115 includes parameters that apply to an entire picture. Thus, each picture in the video sequence can refer to a PPS 115. Note that each picture refers to a PPS 115, but in some examples, a single PPS 115 can include data for multiple pictures. For example, multiple similar pictures can be coded according to similar parameters. In such a case, a single PPS 115 can include data for such similar pictures. The PPS 115 can indicate coding tools, quantization parameters, offsets, etc. available for slices in the corresponding picture.
[0075] Each slice includes a slice header and image data from an area in the picture. The slice header contains parameters unique to each slice. Thus, there can be one slice header for each slice in a video sequence. The slice header can include slice type information, picture order count (POC), reference picture list, prediction weight, tile entry point, deblocking parameters, and so on. In some examples, note that the bitstream 100 can also include a picture header, which is a syntax structure containing parameters applied to all slices in a single picture. For this reason, the picture header and the slice header can be used interchangeably in some contexts. For example, certain parameters can be moved between the slice header and the picture header depending on whether such parameters are common to all slices in the picture. The image data in slice 125 includes video data encoded according to inter prediction and / or intra prediction, as well as corresponding transformed and quantized residual data. Video data from one or more slices can be coded from a picture by an encoder and decoded in a decoder to reconstruct the picture.
[0076] Slice 125 can be defined as an integer number of complete tiles (e.g., within a tile) or an integer number of consecutive complete coding tree units (CTUs) of a picture, where a tile or CTU row is exclusively included in a single NAL unit. Thus, slice 125 is also included in a single NAL unit. Slice 125 is further divided into CTUs and / or coding tree blocks (CTBs). A CTU is a group of samples of a predefined size that can be partitioned by a coding tree. A CTB is a subset of a CTU and includes the luma or chroma components of the CTU. The CTU / CTB is further divided into coding blocks based on a coding tree. The coding blocks can then be encoded / decoded according to a prediction mechanism.
[0077] The bitstream 100 can include one or more SEI messages. An SEI message is a syntax structure with specified semantics that conveys information not required by the decoding process to determine the values of samples in the decoded picture. The bitstream 100 can include many different SEI messages for different functions. In this example, the bitstream includes an AR SEI message 131, a DRI SEI message 133, and an EDRAP indication SEI message 135.
[0078] The AR SEI message 131 is an SEI message that conveys parameters for identifying annotated regions in one or more pictures by adopting a bounding box. The bounding box represents the size and location of the annotated region and identifies one or more objects included in the annotated region. Thus, the AR SEI message 131 includes metadata describing regions in the picture. The decoder can use the AR SEI message 131 to determine whether and / or how such regions should be decoded during the display process. The AR SEI message 131 includes the ar_object_label_idx[ar_object_idx[i]] 141 syntax element. The ar_object_label_idx[ar_object_idx[i]] 141 indicates the index of the label corresponding to the i-th indexed AR object (the ar_object_idx[i]-th object). For example, the AR objects are indexed, and any i-th AR object can be determined by ar_object_idx[i]. Further, the AR object labels are indexed, and any i-th AR object label can be determined by ar_object_label_idx[i]. Thus, the AR SEI ar_object_label_idx[ar_object_idx[i]] 141 obtains the index of the label of the i-th AR object.
[0079] DRI SEI message 133 is an SEI message that conveys parameters for a picture containing depth and / or disparity information for rendering on a three-dimensional (3D) display. Depth is the location of a pixel / sample in 3D space. Disparity is the displacement between the locations of two features (e.g., two pixels) within the image plane. DRI SEI message 133 includes a depth_nonlinear_representation_model[i] 142 syntax element, a depth_nonlinear_representation_num_minus1 143 syntax element, a depth_representation_type 144 syntax element, and a disparity_ref_view_id 145 syntax element. depth_nonlinear_representation_model[i] 142 specifies each of i piecewise linear segments for mapping the decoded luma sample values (e.g., depth values) of an auxiliary picture to a scale that is uniformly quantized with respect to disparity. depth_nonlinear_representation_num_minus1(143) + 2 specifies the number of piecewise linear segments for mapping depth values to a scale that is uniformly quantized with respect to disparity. Thus, depth_nonlinear_representation_num_minus1(143) + 2 specifies the number of i segments in depth_nonlinear_representation_model[i] 142. depth_representation_type 144 specifies the representation definition of the decoded luma samples of an auxiliary picture. The allowed values of depth_representation_type 144 and the corresponding interpretations of each value are included in Table 1 above and / or Table Y1 below. disparity_ref_view_id 145 specifies the ViewId value from which the disparity value is derived. Thus, disparity_ref_view_id 145 indicates the ViewId value that is used as a reference when determining the disparity (e.g., displacement and / or difference between locations) for samples in an auxiliary picture.
[0080] The EDRAP indication SEI message 135 indicates the use of an EDRAP picture. An EDRAP picture is a random access picture that is coded by inter prediction based on one or more reference pictures. For example, an EDRAP picture can be coded by referring to a preceding EDRAP picture and / or a preceding IRAP picture. An IRAP picture is coded by intra prediction and can be decoded without referring to other pictures. The EDRAP method may employ an external bitstream that includes a timing set of reference pictures for each of the EDRAP pictures. In this way, an EDRAP picture can be selected for random access to the main bitstream, and the reference pictures used to decode the EDRAP picture can be obtained from the external bitstream. The EDRAP indication SEI message 135 is an SEI message that indicates constraints on picture order and picture reference, used to ensure that the decoder can select any EDRAP picture for random access and correctly decode the selected EDRAP picture using only the corresponding reference pictures (e.g., in the external bitstream). The EDRAP indication SEI message 135 includes the edrap_leading_pictures_decodable_flag146 syntax element, which contains a value indicating whether a set of ordering constraints applies to the EDRAP pictures corresponding to the EDRAP indication SEI message 135.
[0081] The presence of an EDRAP indicating SEI message 135 imposes certain constraints on the EDRAP picture order in the bitstream. For example, each EDRAP picture is a trailing picture. Further, each EDRAP picture has a temporal sublayer identifier equal to zero. The temporal sublayer divides pictures into a base layer and one or more enhancement layers. A decoder with lower capabilities can decode and display the base layer for a lower frame rate, while a decoder with higher capabilities can decode an increased number of enhancement layers to achieve a higher frame rate. Limiting the temporal sublayer identifier to zero ensures that the EDRAP pictures are in the base layer and thus available to all decoders. Another constraint is that each EDRAP picture, except for the list of referenceable pictures, must not contain pictures in the same layer within the active entries of the reference picture list of the EDRAP picture. The list of referenceable pictures includes IRAP pictures and / or EDRAP pictures in decode order. Thus, this constraint limits EDRAP pictures to reference only preceding IRAP pictures and EDRAP pictures. Yet another constraint is that any picture in the same layer as the EDRAP picture and following the EDRAP picture in both decode order and output order, except for the list of referenceable pictures, must not contain other pictures in the same layer within the active entries of the reference picture list of the picture that precede the EDRAP picture in decode order or output order. This constraint prevents pictures following an EDRAP picture from referencing pictures that precede the EDRAP picture. In the case of random access at an EDRAP picture, the preceding pictures are not available, and thus, if such pictures are referenced by a trailing picture, it results in an error due to the unavailability of the reference picture.
[0082] The edrap_leading_pictures_decodable_flag146 syntax element can impose additional constraints on EDRAP pictures. The first of such additional constraints is that any picture in the same layer as the EDRAP picture and following the EDRAP picture in decoding order shall be in the same layer as the EDRAP picture in output order and follow any other picture that precedes the EDRAP picture in decoding order. In some cases, the coding order and the output order are different. This allows for better compression in some cases but requires pictures to be reordered before display. A picture that follows a random access point in decoding order and precedes it in output order is known as a leading picture. This constraint ensures that the leading pictures of an EDRAP picture are not placed before the trailing pictures from the previous EDRAP picture in output order.
[0083] The second of such additional constraints is that any picture in the same layer as the EDRAP picture, following the EDRAP picture in decoding order, and preceding it in output order shall not include any other picture in the same layer as the EDRAP picture and preceding the EDRAP picture in decoding order, except in the active entries of the reference picture list of the picture, in the list of referenceable pictures. This constraint ensures that leading pictures reference only pictures that follow the EDRAP picture and IRAP and / or EDRAP pictures in the list of referenceable pictures. This ensures that leading pictures can be decoded when the corresponding EDRAP picture is used for random access.
[0084] As described above, the value ranges, descriptors, and / or semantics of ar_object_label_idx[ar_object_idx[i]] 141, depth_nonlinear_representation_model[i] 142, depth_nonlinear_representation_num_minus1 143, depth_representation_type 144, disparity_ref_view_id 145, and edrap_leading_pictures_decodable_flag 146 are not specified in some exemplary systems. Accordingly, the present disclosure includes such value ranges, descriptors, and / or semantics for the foregoing parameter / syntax elements. This enables the decoder to correctly interpret these values without experiencing undefined behavior such as glitches and / or crashes.
[0085] ar_object_label_idx[ar_object_idx[i]], 141, depth_nonlinear_representation_model[i], 142, depth_nonlinear_representation_num_minus1, 143, depth_representation_type, 144, disparity_ref_view_id, 145, and edrap_leading_pictures_decodable_flag146 may be associated with a descriptor indicating a coding mechanism used to code the corresponding syntax element. Such descriptors may include ue(v), u(N), se(v), and u(v). ue(v) indicates that the syntax element value is coded as an unsigned integer exponential-Golomb coded syntax element, with the left-most bit first and using a variable number of bits. The exponential-Golomb code syntax includes representing the value as a plus-one binary and representing the leading zeros as a minus-one format with the leading value. u(N) indicates that the syntax element value is coded as an unsigned integer using N bits. se(v) indicates that the syntax element value is coded as a signed integer exponential-Golomb coded syntax element, with the left-most bit first and using a variable number of bits. u(v) indicates that the syntax element value is coded as an unsigned integer using a variable number of bits.
[0086] To solve the above problems and other problems, a method summarized below is disclosed. The items should be regarded as examples for explaining general concepts and should not be interpreted narrowly. Furthermore, these items can be applied individually or in any combination.
[0087] Example 1
[0088] In one example, to solve at least one of the problems listed above, the value of the ar_object_label_idx[ar_object_idx[i]] 141 syntax element is specified to be within the range of N to M, inclusive of both end values, where N and M are integer values and N is less than M. In one example, N = 0 and M = 255. In one example, the value of ar_object_label_idx[ar_object_idx[i]] 141 is specified to be within different ranges, such as 0 to 3, inclusive of both end values, 0 to 7, inclusive of both end values, 0 to 15, inclusive of both end values, 0 to 31, inclusive of both end values, 0 to 63, inclusive of both end values, etc.
[0089] Example 2
[0090] In one example, to solve at least one of the problems listed above, the depth_nonlinear_representation_model[i] 142 syntax element is specified to be coded as ue(v).
[0091] Example 3
[0092] In one example, the value of depth_nonlinear_representation_model[i] 142 is specified to be within the range of N to M, inclusive of both end values, such as N = 0 and M = 65535. In one example, the value of depth_nonlinear_representation_model[i] 142 is specified to be within different ranges, such as 0 to 3, inclusive of both end values, 0 to 7, inclusive of both end values, 0 to 15, inclusive of both end values, 0 to 31, inclusive of both end values, 0 to 63, inclusive of both end values, 0 to 127, inclusive of both end values, 0 to 255, inclusive of both end values, 0 to 511, inclusive of both end values, 0 to 1023, inclusive of both end values, 0 to 2047, inclusive of both end values, 0 to 4095, inclusive of both end values, 0 to 8191, inclusive of both end values, 0 to 16383, inclusive of both end values, etc.
[0093] Example 4
[0094] In one example, the depth_nonlinear_representation_model[i] 142 syntax element is specified to be coded using different coding methods. In one example, the depth_nonlinear_representation_model[i] 142 syntax element is specified to be coded using u(N) coding, where N is equal to a positive integer value such as a value in the range of 2 to 16 including the end values. In another example, the depth_nonlinear_representation_model[i] 142 syntax element is specified to be coded using se(v) coding. In another example, the depth_nonlinear_representation_model[i] 142 syntax element is specified to be coded using u(v) coding, and the length is specified to be equal to, for example, Log2(MaxNumModes) in units of bits, where the variable MaxNumModes indicates the maximum number of modes, and the function Log2(x) returns the base-2 logarithm of x.
[0095] Example 5
[0096] In one example, the depth_nonlinear_representation_num_minus1 143 syntax element is specified to be coded using a coding method different from the ue(v) coding method such as u(N) and u(v).
[0097] Example 6
[0098] In one example, to solve at least one of the problems listed above, the value of depth_representation_type144 is specified to be within the range of N to M, inclusive of both end values, where N and M are integer values and N is smaller than M. In one example, N = 0 and M = 15. In one example, the value of depth_representation_type144 is specified to be within different ranges, such as within the range of 0 to 3, inclusive of both end values, within the range of 0 to 7, inclusive of both end values, within the range of 0 to 31, inclusive of both end values, within the range of 0 to 63, inclusive of both end values, within the range of 0 to 127, inclusive of both end values, within the range of 0 to 255, inclusive of both end values, etc.
[0099] Example 7
[0100] In one example, to solve at least one of the problems listed above, the value of depth_nonlinear_representation_num_minus1 143 is specified to be within the range of 0 to 62, inclusive of both end values. In one example, the value of depth_nonlinear_representation_num_minus 143 is specified to be within different ranges, such as within the range of 0 to 6, inclusive of both end values, within the range of 0 to 14, inclusive of both end values, within the range of 0 to 30, inclusive of both end values, within the range of 0 to 126, inclusive of both end values, within the range of 0 to 254, inclusive of both end values, etc.
[0101] Example 8
[0102] In one example, to solve at least one of the problems listed above, the value of disparity_ref_view_id145 is specified to be within the range of 0 to 1023, inclusive of both end values. In one example, the value of disparity_ref_view_id145 is specified to be within different ranges, such as within the range of 0 to 63, inclusive of both end values, within the range of 0 to 127, inclusive of both end values, within the range of 0 to 255, inclusive of both end values, within the range of 0 to 511, inclusive of both end values, within the range of 0 to 2047, inclusive of both end values, within the range of 0 to 4095, inclusive of both end values, within the range of 0 to 8191, inclusive of both end values, within the range of 0 to 16383, inclusive of both end values, within the range of 0 to 32767, inclusive of both end values, within the range of 0 to 65535, inclusive of both end values, etc.
[0103] Example 9
[0104] In one example, to solve at least one of the problems listed above, the semantics of the edrap_leading_pictures_decodable_flag146 syntax element are specified as follows. An edrap_leading_pictures_decodable_flag146 equal to 1 specifies that both of the following constraints apply. Any picture in the same layer that follows an EDRAP picture in the decoding order shall follow, in the output order, a picture in the same layer that precedes the EDRAP picture in the decoding order. Any picture in the same layer that follows an EDRAP picture in the decoding order and precedes the EDRAP picture in the output order shall, except for referenceablePictures, not include, in the active entries of its reference picture list, a picture in the same layer that precedes the EDRAP picture in the decoding order. An edrap_leading_pictures_decodable_flag146 equal to 0 does not impose such constraints.
[0105] Next, an embodiment of the above example will be described. This embodiment can be applied to VSEI. Regarding the VSEI specification, the most relevant parts that are added or modified are shown in bold underlined font (underlined font only in this specification (excluding tables)), and some of the deleted parts are shown in bold italic font ([[ ]] in this specification (excluding tables)). There may be some other changes that are essentially editorial and thus not emphasized.
[0106] The semantics of an exemplary annotated region SEI message are as follows. The annotated region SEI message carries parameters that identify the annotated region using a bounding box that represents the size and location of the identified object.
[0107] ...
[0108] ar_object_label_idx[ar_object_idx[i]] indicates the index of the label corresponding to the object at the ar_object_idx[i] -th position. When ar_object_label_idx[ar_object_idx[i]] does not exist, its value is inferred, if possible, from the previous annotated region SEI message in the same CVS in the output order. The value of depth_representation_type shall be within the range of 0 to 15, inclusive of both end values. ...
[0109] Depth Representation Information SEI Message Syntax
Table 7
[0110] Exemplary depth representation information SEI message semantics are as follows. The syntax elements in the depth representation information (DRI) SEI message specify various parameters for auxiliary pictures of type AUX_DEPTH for the purpose of rendering on a 3D display such as view synthesis after processing the decoded primary and auxiliary pictures. Specifically, the depth or disparity range of the depth picture is specified. ...
[0111] depth_representation_type specifies the representation definition of the decoded lumasamples of the auxiliary picture as specified in Table Y1. In Table Y1, disparity specifies the horizontal displacement between two texture views, and the Z - value specifies the distance from the camera. The value of disparity_ref_view_id shall be within the range of 0 to 1023, inclusive of both end values. The variable maxVal is set equal to (1<<BitDepthY)-1.
Table 8
[0112] disparity_ref_view_id specifies the ViewId value from which the disparity value is derived. The value of depth_nonlinear_representation_num_minus1 shall be within the range of 0 to 62, inclusive of both end values.The disparity_ref_view_id exists only when d_min_flag is equal to 1 or d_max_flag is equal to 1, and is useful for depth_representation_type values equal to 1 and 3. The variables in the x column of Table Y2 are derived as follows from the respective variables in the s, e, n, and v columns of Table Y2. When the value of e is in the range of 0 to 127 (excluding 0), x is set equal to (-1)s * 2e - 31 * (1 + n ÷ 2v). Otherwise (when e is equal to 0), x is set equal to (-1)s * 2-(30 + v) * n.
Table 9
[0113] The DMin value and the DMax value, when they exist, are specified in units of the luma sample width of the coded picture whose ViewId is equal to the ViewId of the auxiliary picture. The units of the ZNear and ZFar values, when they exist, are the same but are not specified.
[0114] depth_nonlinear_representation_num_minus1 + 2 specifies the number of piecewise linear segments for mapping depth values to a scale uniformly quantized with respect to disparity. The value of depth_nonlinear_representation_model[i] shall be within the range of 0 to 65535, inclusive of both end values. For i ranging from 0 to depth_nonlinear_representation_num_minus1 + 2 inclusive, depth_nonlinear_representation_model[i] specifies a piecewise linear segment for mapping the decoded luma sample values of the auxiliary picture to a scale uniformly quantized with respect to disparity. The edrap_leading_pictures_decodable_flag equal to 1 specifies that both of the following constraints apply. Any picture in the same layer that follows the EDRAP picture in the decoding order shall follow, in the output order, any picture in the same layer that precedes the EDRAP picture in the decoding order. Any picture in the same layer that follows the EDRAP picture in the decoding order and precedes the EDRAP picture in the output order shall, except for referenceablePictures, not include, in the active entries of its reference picture list, any picture in the same layer that precedes the EDRAP picture in the decoding order. The edrap_leading_pictures_decodable_flag equal to 0 does not impose such constraints. The values of depth_nonlinear_representation_model[0] and depth_nonlinear_representation_model[depth_nonlinear_representation_num_minus1 + 2] are both inferred to be equal to 0.
[0115] ...
[0116] An example of the semantics of the extended DRAP indication SEI message is as follows. The picture associated with the extended DRAP (EDRAP) indication SEI message is called an EDRAP picture. The presence of the EDRAP indication SEI message indicates that the constraints on the picture order and picture references specified in this sub-closure apply. Due to these constraints, the decoder can properly decode the EDRAP picture and the pictures in the same layer, and can follow it in both the decoding order and the output order without having to decode other pictures in the same layer except for the list of pictures referenceablePictures, which contains a list of IRAP or EDRAP pictures in the decoding order that are within the same CLVS and identified by the edrap_ref_rap_id[i] syntax element.
[0117] The constraints indicated by the presence of the EDRAP indication SEI message are all assumed to apply and are as follows. The EDRAP picture is a trailing picture. The EDRAP picture has a temporal sublayer identifier equal to 0. The EDRAP picture does not contain pictures in the same layer in the active entries of its reference picture list, except for referenceablePictures. Any picture in the same layer that follows the EDRAP picture in both the decoding order and the output order does not contain pictures in the same layer in the active entries of its reference picture list, except for referenceablePictures, that precede the EDRAP picture in the decoding order or the output order.
[0118] When edrap_leading_pictures_decodable_flag is equal to 1, the following applies. Any picture in the same layer that follows an EDRAP picture in the decoding order shall follow, in the output order, any picture in the same layer that precedes the EDRAP picture in the decoding order. Any picture in the same layer that follows an EDRAP picture in the decoding order and precedes the EDRAP picture in the output order shall, except for referenceablePictures, not include in the active entries of its reference picture list any picture in the same layer that precedes the EDRAP picture in the decoding order.
[0119] Any picture in the list referenceablePictures shall not include in the active entries of its reference picture list any picture in the same layer that is not a picture at a position earlier in the list referenceablePictures. Thus, the first picture in referenceablePictures shall not include in the active entries of its reference picture list any picture from the same layer, even when it is an EDRAP picture rather than an IRAP picture.
[0120] edrap_rap_id_minus1+1 specifies the RAP picture identifier, shown as RapPicId, of the EDRAP picture. Each IRAP or EDRAP picture is associated with a RapPicId value. It is inferred that the RapPicId value for an IRAP picture is equal to 0. The RapPicId values for any two EDRAP pictures associated with the same IRAP picture shall be different.
[0121]
[0122] The edrap_reserved_zero_12bits shall be equal to 0 in the bitstream conforming to this version of this specification. Other values of edrap_reserved_zero_12bits are reserved. The decoder shall ignore the value of edrap_reserved_zero_12bits. edrap_num_ref_rap_pics_minus1 + 1 indicates the number of IRAP or EDRAP pictures that are in the same CLVS as the EDRAP picture and can be included in the active entries of the reference picture list of the EDRAP picture. edrap_ref_rap_id[i] indicates the RapPicId of the i-th RAP picture that can be included in the active entries of the reference picture list of the EDRAP picture. The i-th RAP picture shall be either an IRAP picture associated with the current EDRAP picture or an EDRAP picture associated with the same IRAP picture as the current EDRAP picture. ...
[0123] Figure 2 is a block diagram showing an exemplary video processing system 4000 in which various techniques disclosed in this specification can be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interface connections include wired interfaces such as Ethernet®, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0124] System 4000 may include a coding component 4004 that can implement various coding or encoding methods described in this document. The coding component 4004 may reduce the average bit rate of the video from input 4002 to the output of the coding component 4004 in order to generate a coded representation of the video. Thus, the coding technique may be referred to as a video compression or video transcoding technique. The output of the coding component 4004 may either be stored or transmitted via connected communication as represented by component 4006. The stored or communicated bitstream (or coded) representation of the video received at input 4002 may be used by component 4008 to generate pixel values or displayable video that is sent to the display interface 4010. The process of generating user-viewable video from the bitstream representation may be referred to as video restoration. Further, certain video processing operations are referred to as "coding" operations or tools, but it should be understood that coding tools or operations are used in an encoder and the corresponding decoding tools or operations that reverse the result of coding are performed by a decoder.
[0125] Examples of a peripheral bus interface or a display interface may include a USB (universal serial bus), HDMI (registered trademark) (high definition multimedia interface), a display port, etc. Examples of a storage interface include SATA (serial advanced technology attachment), PCI, an IDE interface, etc. The techniques described in this document may be embodied in various electronic devices such as a cellular phone, a laptop, a smartphone, or other devices that can perform digital data processing and / or video display.
[0126] FIG. 3 is a block diagram of an exemplary video processing apparatus 4100. The apparatus 4100 may be used to implement one or more of the methods described herein. The apparatus 4100 may be embodied in, for example, a smartphone, a tablet, a computer, a monolithic Internet of Things (IoT) receiver, etc. The apparatus 4100 may include one or more processors 4102, one or more memories 4104, and a video processing circuit 4106. The processor(s) 4102 may be configured to implement one or more of the methods described in this document. The memory(ies) 4104 may be used to store data and code used to implement the methods and techniques described herein. The video processing circuit 4106 may be used in a hardware circuit to implement some of the techniques described in this document. In some embodiments, the video processing circuit 4106 may be at least partially included in the processor 4102, for example, a graphics coprocessor.
[0127] FIG. 4 is a flowchart of an exemplary method 4200 of video processing. Method 4200 includes, at step 4202, determining the value of the edrap_leading_pictures_decodable_flag syntax element. The bitstream includes at least one EDRAP picture. In one example, the value of the edrap_leading_pictures_decodable_flag indicates whether an ordering constraint is imposed on the EDRAP pictures. The ordering constraint may not be imposed on the EDRAP pictures when the value of the edrap_leading_pictures_decodable_flag is zero. The ordering constraint may be imposed on the EDRAP pictures when the value of the edrap_leading_pictures_decodable_flag is 1. The constraint may specify that any picture in the same layer as the EDRAP picture and following the EDRAP picture in decoding order shall be in the same layer as the EDRAP picture in output order and follow any other picture preceding the EDRAP picture in decoding order. The constraint may further specify that any picture in the same layer as the EDRAP picture, following the EDRAP picture in decoding order and preceding the EDRAP picture in output order shall not include any other picture in the same layer as the EDRAP picture and preceding the EDRAP picture in decoding order, except in the active entries of the reference picture list of the picture in the list of referenceable pictures.
[0128] In one example, the edrap_leading_pictures_decodable_flag syntax element is included in the EDRAP SEI message. In one example, the list of referenceable pictures includes IRAP or EDRAP pictures in decoding order that are within the same CLVS. In one example, each picture in the list of referenceable pictures is identified by the edrap_ref_rap_id[i] syntax element. In one example, the edrap_leading_pictures_decodable_flag syntax element is u(v)-coded.
[0129] In one example, the EDRAP SEI message specifies additional constraints. In one example, the additional constraints include that each EDRAP picture is a trailing picture. In one example, the additional constraints include that each EDRAP picture has a time sublayer identifier equal to zero. In one example, the additional constraints include that each EDRAP picture, except for the list of referenceable pictures, does not include pictures in the same layer in the active entries of the reference picture list of the EDRAP picture. In one example, the additional constraints include that the bitstream is constrained such that any picture in the same layer as the EDRAP picture and following the EDRAP picture in both the decoding order and the output order, except for the list of referenceable pictures, does not include other pictures in the same layer in the active entries of the reference picture list of the picture and preceding the EDRAP picture in the decoding order or the output order.
[0130] In step 4204, a conversion is performed between the visual media data and the bitstream based on the edrap_leading_pictures_decodable_flag syntax element. When method 4200 is executed on an encoder, the conversion includes encoding the visual media data into a bitstream. When method 4200 is executed on a decoder, the conversion includes decoding the visual media data from the bitstream.
[0131] Note that method 4200 may be implemented in an apparatus for processing video data, such as video encoder 4400, video decoder 4500, and / or encoder 4600, comprising a processor and a non-transitory memory having instructions. In such a case, the instructions, when executed by the processor, cause the processor to execute method 4200. Further, method 4200 can be executed by a non-transitory computer-readable medium comprising a computer program product for use by a video coding device. The computer program product comprises computer-executable instructions stored in the non-transitory computer-readable medium that, when executed by the processor, cause the video coding device to execute method 4200.
[0132] FIG. 5 is a block diagram showing an exemplary video coding system 4300 that may utilize the techniques of the present disclosure. Video coding system 4300 may include a source device 4310 and a destination device 4320. Source device 4310 generates encoded video data, which may sometimes be referred to as a video coding device. Destination device 4320 may decode the encoded video data generated by source device 4310, which may sometimes be referred to as a video decoding device.
[0133] The source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. The video source 4312 may include a video capture device, an interface for receiving video data from a video content provider, and / or a source such as a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 4314 encodes the video data from the video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that forms a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded picture is a coded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to the destination device 4320 via the I / O interface 4316 through the network 4330. The encoded video data may also be stored on the storage medium / server 4340 for access by the destination device 4320.
[0134] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain the encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to the user. The display device 4322 may be integrated with the destination device 4320 or may be external to the destination device 4320, which can be configured to interface with an external display device.
[0135] Video encoder 4314 and video decoder 4324 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVM) standard, and other current and / or future standards.
[0136] FIG. 6 is a block diagram showing an example of a video encoder 4400 that may be the video encoder 4314 in the system 4300 shown in FIG. 5. The video encoder 4400 may be configured to execute any or all of the techniques of the present disclosure. The video encoder 4400 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video encoder 4400. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.
[0137] The functional components of the video encoder 4400 may include a partitioning unit 4401, a prediction unit 4402 that may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, an intra prediction unit 4406, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy encoding unit 4414.
[0138] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode where at least one reference picture is the picture in which the current video block is located.
[0139] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 may be highly integrated, but are shown separately in the example of the video encoder 4400 for illustration purposes.
[0140] The segmentation unit 4401 can segment a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0141] The mode selection unit 4403 can select, for example, one of the intra or inter coding modes based on an error result, provide the obtained intra or inter coded block to the residual generation unit 4407 to generate residual block data, and provide it to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select a combination of intra and inter prediction (CIIP) modes where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 4403 can also select the resolution of the motion vector for a block (e.g., sub-pixel or integer pixel accuracy) in the case of inter prediction.
[0142] To perform inter prediction on the current video block, the motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from the buffer 4413 with the current video block. The motion compensation unit 4405 can determine a predicted video block for the current video block based on the motion information and the decoded samples of the pictures from the buffer 4413 other than the picture associated with the current video block.
[0143] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0144] In some examples, the motion estimation unit 4404 may perform unidirectional prediction for the current video block, and the motion estimation unit 4404 may search for reference pictures in list 0 or list 1 for the reference video block for the current video block. Then, the motion estimation unit 4404 may generate a reference index indicating the reference picture in list 0 or list 1 that includes the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0145] In other examples, the motion estimation unit 4404 may perform bidirectional prediction for the current video block, and the motion estimation unit 4404 may search for a reference picture in list 0 for the reference video block for the current video block, and may also search for a reference picture in list 1 for another reference video block for the current video block. Then, the motion estimation unit 4404 may generate a reference index indicating the reference pictures in list 0 and list 1 that include the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0146] In some examples, motion estimation unit 4404 may output a full set of motion information for decoder decoding processing. In some examples, motion estimation unit 4404 may not output a full set of motion information for the current video. Instead, motion estimation unit 4404 may signal the motion information of the current video block by referring to the motion information of another video block. For example, motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0147] In one example, motion estimation unit 4404 may indicate a value to video decoder 4500 indicating that in the syntax structure associated with the current video block, the current video block has the same motion information as another video block.
[0148] In another example, motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0149] As described above, video encoder 4400 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by video encoder 4400 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0150] The intra prediction unit 4406 can perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 can generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data of the current video block can include a predicted video block and various syntax elements.
[0151] The residual generation unit 4407 can generate residual data of the current video block by subtracting the predicted video block(s) of the current video block from the current video block. The residual data of the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0152] In other examples, for example, in skip mode, there may be no residual data for the current video block for the current video block, and the residual generation unit 4407 may not perform this subtraction operation.
[0153] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0154] After the transform processing unit 4408 generates the transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0155] Inverse quantization unit 4410 and inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 4412 can add the reconstructed residual video block to the corresponding samples from one or more prediction video blocks generated by prediction unit 4402 to generate the reconstructed video block associated with the current block for storage in buffer 4413.
[0156] After reconstruction unit 4412 reconstructs the video block, a loop filter processing operation can be performed to reduce video blocking artifacts in the video block.
[0157] Entropy encoding unit 4414 can receive data from other functional components of video encoder 4400. When entropy encoding unit 4414 receives the data, entropy encoding unit 4414 can perform one or more entropy encoding operations to generate the entropy encoded data and output a bitstream containing the entropy encoded data.
[0158] FIG. 7 is a block diagram illustrating an example of video decoder 4500 that can be video decoder 4324 in system 4300 shown in FIG. 5. Video decoder 4500 can be configured to execute any or all of the techniques of the present disclosure. In the illustrated example, video decoder 4500 includes a plurality of functional components. The techniques described in the present disclosure can be shared among various components of video decoder 4500. In some examples, a processor can be configured to execute any or all of the techniques described in the present disclosure.
[0159] In the illustrated example, video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, video decoder 4500 may perform a decoding path generally inverse to the encoding path described with respect to video encoder 4400.
[0160] Entropy decoding unit 4501 may extract the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). Entropy decoding unit 4501 may decode the entropy-coded video data, and from the entropy-decoded video data, motion compensation unit 4502 may determine motion information including a motion vector, motion vector precision, a reference picture list index, and other motion information. Motion compensation unit 4502 may determine such information, for example, by performing AMVP and merge mode.
[0161] Motion compensation unit 4502 may, in some cases, perform interpolation based on an interpolation filter to generate a motion-compensated block. An identifier of the interpolation filter to be used at sub-pixel precision may be included in the syntax element.
[0162] Motion compensation unit 4502 may use the interpolation filter used by video encoder 4400 during encoding of the video block to calculate interpolation values for sub-integer pixels of the reference block. Motion compensation unit 4502 may determine the interpolation filter used by video encoder 4400 according to the received syntax information and use the interpolation filter to generate a prediction block.
[0163] The motion compensation unit 4502 may use a part of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, the partitioning information that describes how each macroblock of the picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0164] The intra prediction unit 4503 may use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 4504 inverse quantizes, i.e., de-quantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.
[0165] The reconstruction unit 4506 may add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. Optionally, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 4507, which provides reference blocks for subsequent motion compensation / intra prediction and generates the decoded video for presentation on a display device.
[0166] FIG. 8 is a schematic diagram of an exemplary encoder 4600. The encoder 4600 is suitable for implementing the technology of VVC. The encoder 4600 includes three in-loop filters, namely, a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Different from the DF 4602 that uses a predefined filter, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture to reduce the mean squared error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter together with coded side information that signals the offset and the filter coefficients, respectively. The ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool that tries to capture and correct the artifacts generated in the previous stage.
[0167] The encoder 4600 further includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive an input video. The intra prediction component 4608 is configured to perform intra prediction, and the ME / MC component 4610 is configured to perform inter prediction using a reference picture obtained from the reference picture buffer 4612. A residual block from inter prediction or intra prediction is supplied to a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are supplied to an entropy coding component 4618. The entropy coding component 4618 entropy-codes the prediction result and the quantized transform coefficients and transmits them to a video decoder (not shown). The quantization component output from the quantization component 4616 may be supplied to an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 can output the image to the DF 4602, SAO 4604, and ALF 4606 for filtering before the image is stored in the reference picture buffer 4612.
[0168] A list of preferred solutions for several examples is provided below.
[0169] The following solutions illustrate examples of the techniques described herein.
[0170] 1. A method of media data processing (e.g., method 4200 shown in FIG. 4), comprising performing a conversion between a video and a bitstream of the video according to rules, the rules specifying a range of values of one or more syntax fields indicating an index to a label for a corresponding object in an annotated region of the video, the range being between N and M, where N and M are integers.
[0171] 2. The method according to solution 1, wherein N = 0 and M = 255.
[0172] 3. The method according to Solution 1, where N = 0 and M = 3.
[0173] The following solutions show exemplary embodiments of the techniques described in the previous section (e.g., item 2).
[0174] 4. A method for processing video data, comprising the step of performing a conversion between a video and a bitstream of the video according to rules, the rules specifying that the type of coding used to code a syntax element indicating the number of non-linear segments used in a piecewise non-linear mapping of depth information for an object in the video is coded in the bitstream.
[0175] 5. The method according to Solution 1, where the rules specify that the syntax element is coded as a variable-length unsigned integer zero-order exponential Golomb code syntax element in left-bit-first coding.
[0176] 6. The method according to Solution 1, where the rules specify that the syntax element is coded as u(N’), and N’ is a positive integer.
[0177] 7. The method according to Solution 1, where the rules specify that the syntax element is coded as a signed integer zero-order exponential Golomb code syntax element in left-bit-first coding.
[0178] 8. The method according to any one of Solutions 1 to 4, where the rules specify that the value of the syntax element is constrained to be within the ranges of N and M, and N and M are integers.
[0179] 9. The method according to Solution 5, where N = 0 and M = 65535.
[0180] 10. The method according to Solution 2, where N = 0 and M = 3.
[0181] The following solutions show exemplary embodiments of the techniques described in the previous section (e.g., item 3).
[0182] 11. A method for processing video data, comprising the step of performing a conversion between a video and a bitstream of the video according to rules, the rules specifying constraints on syntax elements in a supplementary enhancement information syntax structure representing depth information of one or more objects in the video.
[0183] 12. The method according to solution 11, wherein the syntax element includes a depth representation type, and the rules specify that the value of the syntax element is constrained to be within the range of N to M, where N and M are integers.
[0184] 13. The method according to solution 11, wherein the syntax element indicates the number of non-linear mapping models for depth information, and the rules specify that the syntax element is within the range of 0 to M, where M is an integer.
[0185] 14. The method according to solution 11, wherein the syntax element indicates an identifier of a disparity reference view, and the rules specify that the value of the syntax element is between 0 and M, where M is an integer.
[0186] The following solutions show exemplary embodiments of the techniques described in the previous section (e.g., item 4).
[0187] 15. A method for processing video data, comprising the step of performing a conversion between a video and a bitstream of the video according to rules, the rules specifying that the value of a flag indicating a picture that is an extended dependent random access point controls (1) a first order constraint on pictures in the same layer as the picture and following the picture in decoding order and output order, and (2) a second order constraint on pictures in the same layer as the picture and following the picture in decoding order and preceding the picture in output order.
[0188] 16. The method according to solution 15, wherein the value is equal to 1.
[0189] 17. The method according to any one of solutions 1 to 16, wherein the conversion includes generating a bitstream from the video.
[0190] 18. The method according to any one of solutions 1 to 16, wherein the conversion includes generating a video from the bitstream.
[0191] 19. A video decoding apparatus comprising a processor configured to perform the method according to one or more of solutions 1 to 18.
[0192] 20. A video encoding apparatus comprising a processor configured to perform the method according to one or more of solutions 1 to 18.
[0193] 21. A computer program product storing computer code, which, when executed by a processor, causes the processor to perform the method according to any one of solutions 1 to 18.
[0194] 22. A video processing method including generating a bitstream according to the method according to any one of solutions 1 to 8, and storing the bitstream in a computer-readable medium.
[0195] 23. The method, apparatus or system described in this document.
[0196] As used herein, the term "video processing" may refer to video encoding, video decoding, video compression, or video restoration. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may correspond to bits that are either collocated within the bitstream, as defined by the syntax, or are spread across different locations within the bitstream. For example, a macroblock may be encoded with respect to the transformed and coded error residual values, as well as bits in headers and other fields in the bitstream. Further, during conversion, the decoder may parse the bitstream using knowledge that some fields may or may not be present, based on decisions as explained in the above solutions. Similarly, the encoder may determine whether a particular syntax field should or should not be included and, accordingly, generate the coded representation by including or excluding the syntax field from the coded representation.
[0197] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, or in combinations of one or more of them, including the structures disclosed in this document and their structural equivalents. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition that provides a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" includes, by way of example, all apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer programs in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., an electrical, optical, or electromagnetic signal generated by a machine, generated to encode information for transmission to an appropriate receiver device.
[0198] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, either as a stand-alone program or as part of a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer, or located at one site, or distributed across multiple sites and executed on multiple computers interconnected by a communication network.
[0199] The processes and logical flows described in this document can be performed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, and the apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0200] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions, and one or more memory devices for storing instructions and data. In general, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto - optical disks, or optical disks, or be operatively coupled to perform data reception or transfer to or from these devices. However, a computer need not have such devices. Computer - readable media suitable for storing computer program instructions and data include all forms of non - semiconductor memory media and memory devices including, by way of example, volatile memory devices such as erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto - optical disks, compact disk read only memory (CDROM) and digital versatile disk read only memory (DVDROM) disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0201] This patent document contains many details, but these should not be construed as limitations on any subject matter or what may be claimed. Rather, they should be construed as descriptions of features that may be specific to particular embodiments of particular techniques. The specific features described in this patent document in the context of separate embodiments may be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Further, features may be described and even initially claimed as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be deleted from that combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.
[0202] Similarly, operations are shown in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order to achieve the desired result, or that all of the illustrated operations be performed. Further, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0203] Only some implementations and examples are described, and other implementations, extensions, and variations may be made based on what is described and shown in this patent document.
[0204] When there are no intervening components other than a line, trace, or other medium between a first component and a second component, the first component is directly coupled to the second component. When there are intervening components other than a line, trace, or other medium between the first component and the second component, the first component is indirectly coupled to the second component. The term "coupled" and its variations include both being directly coupled and being indirectly coupled. The use of the term "about" means a range that includes ±10% of the subsequent number, unless otherwise specified.
[0205] Although several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. This example should be considered illustrative and not restrictive, and its intent should not be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or some features may be omitted or not implemented.
[0206] In addition, the techniques, systems, subsystems, and methods described and illustrated in various embodiments as discrete or separate may be combined with or integrated into other systems, modules, techniques, or methods without departing from the scope of this disclosure. Other items shown or discussed as being coupled may be directly connected or may be indirectly coupled or communicated through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of changes, substitutions, and modifications will be apparent to those skilled in the art and may be made without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing video data, comprising: determining a value of a first syntax element included in an Extended Dependency Random Access Point (EDRAP) indication supplementary enhancement information (SEI) message for conversion between a video including a current picture and a bitstream of the video; performing the conversion based on the determining step and rules; wherein the value of the first syntax element indicates whether an ordering constraint is imposed on the current picture; the rules specify that a second syntax element in a Depth Representation Information (DRI) indication supplementary enhancement information (SEI) message is within a range of 0 to 15 including both end values; the second syntax element in the DRI SEI message specifies a representation definition of decoded luma samples of auxiliary pictures; a method.
2. The method according to claim 1, wherein the ordering constraint is not imposed on the current picture when the value of the first syntax element is zero.
3. The method according to claim 1, wherein the ordering constraint is imposed on the current picture when the value of the first syntax element is one.
4. The first syntax element equal to 1 has the following ordering constraint: 1) Any picture in the same layer and subsequent to the current picture in the decoding order shall be subsequent to any picture in the same layer and preceding the current picture in the decoding order in the output order, and 2) Any picture that is in the same layer, follows the current picture in the decoding order, and precedes the current picture in the output order, except for the list of referenceable pictures, shall not contain a picture in the same layer that precedes the current picture in the decoding order in the active entries of the reference picture list of the picture. The method according to claim 1, specifying that both of are applicable.
5. The method according to claim 4, wherein the list of referenceable pictures includes an Intra Random Access Point (IRAP) picture or an EDRAP picture in decoding order within the same Coded Layer Video Sequence (CLVS).
6. Each picture in the list of referenceable pictures is identified by a second syntax element in the EDRAP indication SEI message. The second syntax element indicates the picture identifier of the Random Access Point (RAP) picture included in the active entry of the reference picture list of the current picture. The method according to claim 4.
7. The method according to claim 6, wherein the second syntax element is coded using an unsigned integer using 16 bits.
8. The method according to claim 1, wherein the first syntax element is coded using an unsigned integer using 1 bit.
9. The method according to claim 1, wherein the current picture is an EDRAP picture.
10. The method according to claim 1, wherein the current picture is a trailing picture.
11. The method according to claim 1, wherein the current picture has a time sublayer identifier equal to zero.
12. The method according to claim 1, wherein the current picture does not include pictures in the same layer in the active entries of the reference picture list of the current picture, except for the list of referenceable pictures.
13. The method according to claim 1, wherein the bitstream is in the same layer as the current picture, and any picture subsequent to the current picture in both the decoding order and the output order does not include pictures in the same layer in the active entries of the reference picture list of the picture, except for the list of referenceable pictures, and is restricted not to include pictures in the same layer and preceding the current picture in the decoding order or the output order.
14. The method according to any one of claims 1 to 13, wherein the conversion includes encoding the video into the bitstream.
15. The method according to any one of claims 1 to 13, wherein the conversion includes decoding the video from the bitstream.
16. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to: determine a value of a first syntax element included in an Extended Dependency Random Access Point (EDRAP) indication Supplemental Enhancement Information (SEI) message for conversion between a video including a current picture and a bitstream of the video; execute the conversion based on the determining step and rules; and the value of the first syntax element indicates whether an ordering constraint is imposed on the current picture; the rules specify that a second syntax element in a Depth Representation Information (DRI) indication Supplemental Enhancement Information (SEI) message is within the range of 0 to 15 including both end values; The second syntax element in the DRI SEI message specifies the representation definition of the decoded luma samples of the auxiliary picture. Device.
17. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to determine a value of a first syntax element included in an extended dependent random access point (EDRAP) indication supplementary enhancement information (SEI) message for conversion between a video including a current picture and a bitstream of the video; execute the conversion based on the determining step and rules; and cause to perform the value of the first syntax element indicates whether an ordering constraint is imposed on the current picture; the rules specify that a second syntax element in a depth representation information (DRI) indication supplementary enhancement information (SEI) message is within a range of 0 to 15 including both end values; the second syntax element in the DRI SEI message specifies the representation definition of the decoded luma samples of the auxiliary picture; non-transitory computer-readable storage medium.
18. A method for storing a bitstream of a video, comprising: determining a value of a first syntax element included in an extended dependent random access point (EDRAP) indication supplementary enhancement information (SEI) message for a video including a current picture; generating the bitstream of the video based on the determining step and rules; storing the bitstream in a non-transitory computer-readable recording medium; and including the value of the first syntax element indicates whether an ordering constraint is imposed on the current picture; The rule specifies that the second syntax element in the depth representation information (DRI) indication supplementary enhancement information (SEI) message is within the range of 0 to 15 including both end values, and the second syntax element in the DRI SEI message specifies the representation definition of the decoded luma samples of the auxiliary picture. Method.
Citation Information
Cited By
Enhanced signaling of supplemental enhancement information
US12501052B2