High-level syntax in image headers
By optimizing the signaling constraints and parameter settings in the VVC standard, the problems of unclear stripe type signaling, unclear semantics of adaptive loop filter disabling, and no limit on the number of repetitions of non-VCL NAL units have been solved, improving the efficiency and consistency of video encoding and decoding, and adapting to multi-layer video encoding and decoding standards.
Patent Information
- Application Number
- CN202180025544.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-30
- Filing Date
- 2021-03-29
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2041-03-29
AI Technical Summary
The existing VVC standard has problems such as unclear stripe type signaling constraints, unclear semantics for enabling and disabling adaptive loop filters, redundant signaling for image width and height and consistency window parameters, and no limit on the number of repetitions of non-VCL NAL units, which affect the efficiency and consistency of video encoding and decoding.
By adjusting the signaling constraints of slice_type, updating the semantics of ph_alf_enabled_flag, optimizing the consistency window parameters, and limiting the repetition times of non-VCL NAL units, the clarity and efficiency of signaling are ensured. This includes inferring the slice type under specific conditions and disabling the adaptive loop filter, optimizing the consistency window parameters, and limiting the repetition time of VPS, SPS, PPS, APS, and DCI NAL units.
It improves the efficiency and consistency of the video encoding and decoding process, reduces redundant signaling, enhances the performance and compatibility of video processing, and adapts to multi-layer video encoding and decoding standards.
Smart Images

Figure CN115380525B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application, pursuant to applicable patent law and / or the rules of the Paris Convention, consequently claims priority and benefit to U.S. Provisional Patent Application No. 63 / 002,064, filed March 30, 2020. For all purposes under the law, the entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video processing. Background Technology
[0004] Digital video accounts for the largest share of bandwidth usage on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders to process codec representations using control information useful for decoding the codec representations of video.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images containing one or more stripes and a video bitstream, wherein the bitstream conforms to a format rule; wherein the format rule specifies whether or how the stripe type in the bitstream indicates one or more stripes depends on a condition, wherein the condition is based on a general constraint flag, a network abstraction layer unit type, or whether the stripe is in the first image of the accessed unit.
[0007] In another example, a video processing method is disclosed. This method includes performing a conversion between a video comprising an image containing multiple stripes and a bitstream of the video, wherein the bitstream conforms to format rules that specify the applicability of adaptive loop filtering of all stripes in the image, controlled by flags in a specified image header.
[0008] In another example, a video processing method is disclosed. This method includes performing a conversion between a video comprising one or more images containing one or more stripes and a bitstream of the video, according to format rules specifying the repetition time of a set of parameters associated with the video.
[0009] In another example, a video processing method is disclosed. This method includes performing a conversion between a video comprising images within a video unit and a video bitstream, according to format rules, wherein the format rules specify that, in response to the image width being equal to the maximum allowed image width within the video unit and the image height being equal to the maximum allowed image height within the video unit, a consistency window flag corresponding to the image in the image parameter set is set to a value of 0.
[0010] In another example, a video processing method is disclosed. This method includes performing a conversion between a video comprising one or more images containing one or more slices and a codec representation of the video, wherein the codec representation conforms to a format rule specifying whether a field in the codec representation controls constraints on the slice type or whether the slice type is included in the codec representation, wherein the field includes a general constraint flag, a network abstraction layer unit type, or whether the video slice is in the first video image of the accessed unit.
[0011] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more images containing one or more stripes and a codec representation of the video, wherein the codec representation conforms to a format rule specifying values of flags in the image header of the video image to disable adaptive loop filtering of all stripes in the video image.
[0012] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more images containing one or more stripes and a codec representation of the video, wherein the codec representation conforms to a format rule specifying that a consistency window flag is set to a disabled mode when the height and width of the current image are equal to the maximum height and maximum width in the video.
[0013] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more images containing one or more stripes and a codec representation of the video, wherein the codec representation conforms to a format rule for repetition time of a specified set of parameters.
[0014] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.
[0015] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.
[0016] In yet another example, a computer-readable medium storing code is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.
[0017] These and other features will be described in this document. Attached Figure Description
[0018] Figure 1 This is a block diagram of an example video processing system.
[0019] Figure 2 This is a block diagram of a video processing device.
[0020] Figure 3 This is a flowchart of an example method for video processing.
[0021] Figure 4 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.
[0022] Figure 5 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0023] Figure 6 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0024] Figure 7 An example of an adaptive loop filter (ALF) filter shape (chroma: 5×5 rhombus, luminance: 7×7 rhombus) is shown.
[0025] Figure 8 Examples of ALF and CC-ALF diagrams are shown.
[0026] Figures 9-11 A flowchart illustrating an example method for video processing is shown. Detailed Implementation
[0027] Chapter headings are used in this document for ease of understanding, and the applicability of the technologies and embodiments disclosed in each chapter is not limited to that chapter alone. Furthermore, the use of H.266 technical terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs.
[0028] 1. Introduction
[0029] This document relates to video codec technologies. Specifically, it concerns improvements to stripe-type signaling, ALF and conformance windows, and the repetition of certain non-VCL NAL units (including VPS, SPS, PPS, APS, and DCI NAL units). These ideas can be applied individually or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video codecs, such as the Multi-Functional Video Codec (VVC) under development.
[0030] 2. Abbreviation
[0031] ALF Adaptive Loop Filter
[0032] APS Adaptive Parameter Set
[0033] AU Access Unit
[0034] AUD Access Unit Separator
[0035] AVC Advanced Video Codec
[0036] CLVS codec layer video sequence
[0037] CPB image buffer
[0038] CRA Fully Random Access
[0039] CTU (Codec Tree Unit)
[0040] CVS codec video sequence
[0041] DCI decoding capability information
[0042] DPB Decoding Image Buffer
[0043] DU decoding unit
[0044] End of EOB bitstream
[0045] End of EOS sequence
[0046] GDR gradually decoded and refreshed
[0047] HEVC High-Efficiency Video Encoding and Decoding
[0048] HRD Assumption Reference Decoder
[0049] IDR Instant Decoding and Refresh
[0050] JEM Joint Exploration Model
[0051] LMCS Luminance Mapping and Chroma Scaling
[0052] MCTS Motion Restraint Piece Set
[0053] NAL Network Abstraction Layer
[0054] OLS Output Layer Set
[0055] PH image header
[0056] PPS Image Parameter Set
[0057] PTL levels, tiers, and grades
[0058] PU Image Unit
[0059] RADL random access decodeable front-end (image)
[0060] RAP Random Access Point
[0061] RASL random access skips prerequisites (image)
[0062] RBSP raw byte sequence payload
[0063] RPL Reference Image List
[0064] SAO Sample Adaptive Offset
[0065] SEI Assist Enhancement Information
[0066] SPS Sequence Parameter Set
[0067] STSA Stepwise Temporal Sublayer Access
[0068] SVC Scalable Video Codec
[0069] VCL (Video Codec Layer)
[0070] VPS Video Parameter Set
[0071] VTM VVC Test Model
[0072] VUI Video Availability Information
[0073] VVC Multi-Functional Video Encoding and Decoding
[0074] 3. Preliminary Discussion
[0075] Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC[1] standards. Since H.262, video coding standards have been based on a hybrid video coding architecture, which employs temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new approaches and incorporated them into a reference software called the Joint Exploration Model (JEM)[2]. JVET meetings are held concurrently every quarter, and the goal of the new coding standards is to reduce the bit rate by 50% compared to HEVC. The new video codec standard was officially named Multifunctional Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at every JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The current goal of the VVC project is to achieve Technical Finalization (FDIS) at the meeting in July 2020.
[0076] 3.1. Parameter Set
[0077] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. SPS and PPS are supported in all of AVC, HEVC, and VVC. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.
[0078] SPS is designed to carry sequence-level header information, and PPS is designed to carry infrequently changing image-level header information. Using SPS and PPS, infrequently changing information does not need to be repeated for each sequence or image, thus avoiding redundant signaling. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission but also improving fault tolerance.
[0079] A VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0080] APS is introduced to carry such image-level or strip-level information, which requires a considerable number of bits to encode and decode, can be shared by multiple images, and can have a considerable number of different variations in the sequence.
[0081] 3.2. Strip headers and image headers in VVC
[0082] Similar to HEVC, the stripe header in VVC conveys information about a specific stripe. This includes the stripe address, stripe type, stripe QP, picture sequence count (POC) least significant bit (LSB), RPS and RPL information, weighted prediction parameters, loop filter parameters, slice and WPP entry offsets, etc.
[0083] VVC introduces a Picture Header (PH), which contains header parameters for a specific picture. Each picture must have one or only one PH. The PH essentially carries the parameters that would be present in the slice header if no PH were introduced, but each parameter has the same value for all slices of the picture. These include IRAP / GDR picture indication, inter-frame / intra-frame slice allow flags, POCLSB and optionally POC MSB, information about RPL, deblocking, SAO, ALF, QP increment and weighted prediction, codec block segmentation information, virtual boundaries, juxtaposed picture information, etc. Often, each picture in the entire picture sequence contains only one slice. To allow for at least two NAL units per picture in this case, a PH syntax structure can be included in the PH NAL unit or the slice header.
[0084] In VVC, signaling information about the juxtaposed images is provided in the image header or strip header for temporal motion vector prediction.
[0085] 3.3. Changes in image resolution within a sequence
[0086] In AVC and HEVC, the spatial resolution of an image cannot be changed unless a new sequence with a new SPS begins with an IRAP image. VVC enables intra-sequence image resolution changes at locations where an IRAP image is not encoded; this IRAP image is always intra-coded. This feature is sometimes called Reference Image Resampling (RPR) because it requires resampling of the reference image used for inter-frame prediction when the reference image has a different resolution than the current image being decoded.
[0087] The scaling factor is limited to greater than or equal to 1 / 2 (2x downsampling from the reference image to the current image) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling factors between the reference and current images. The three sets of resampling filters are applied to scaling factors ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, similar to motion-compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process where the scaling factor ranges from 1 / 1.25 to 8. The horizontal and vertical scaling factors are derived based on the image width and height, as well as the left, right, top, and bottom scaling offsets specified for the reference and current images.
[0088] Other aspects of the VVC design that support this feature differ from HEVC include: i) Picture resolution and the corresponding consistency window are signaled in the PPS instead of the SPS, where the maximum picture resolution is signaled in the SPS. ii) For a single-layer bitstream, each picture storage (the slot in the DPB used to store one decoded picture) occupies the buffer size required to store the decoded picture with the maximum picture resolution.
[0089] 3.4. Adaptive Loop Filter (ALF)
[0090] Two diamond filter shapes (such as) Figure 7 (As shown) is used for block-based ALF. 7×7 rhombuses are applied to the luma component, and 5×5 rhombuses are applied to the chroma component. One of up to 25 filters is selected for each 4×4 block based on the orientation and activity of the local gradient. Each 4×4 block in the image is categorized according to orientation and activity. Before filtering each 4×4 block, simple geometric transformations, such as rotation or diagonal and vertical flips, can be applied to the filter coefficients based on the gradient values calculated for that block. This is equivalent to applying these transformations to samples in the filter's support region. The idea is to make different blocks more similar by aligning the orientation of the different blocks to which the ALF is applied. Block-based categorization is not applied to the chroma component.
[0091] ALF filter parameters are signaled in the Adaptive Parameter Set (APS). Within an APS, a set of up to 25 luma filter coefficients and clipping value indices, and a set of up to 8 chroma filter coefficients and clipping value indices can be signaled. To reduce bit overhead, filter coefficients from different categories of luma components can be merged. In the picture or stripe header, up to 7 APS IDs can be signaled to specify the luma filter set used for the current picture or stripe. The filtering process is further controlled at the CTB level. The luma CTB can select a filter set from 16 fixed filter sets and the filter sets signaled in the APS. For the chroma component, the APSID is signaled in the picture or stripe header to indicate the chroma filter set used for the current picture or stripe. At the CTB level, if there is more than one chroma filter set in the APS, the filter index is signaled for each chroma CTB. When ALF is enabled for a CTB, for each sample within the CTB, a diamond filter with weights provided by signaling is executed, where a clipping operation is applied to clip the difference between neighboring samples and the current sample. The clipping operation introduces non-linearity to make ALF more efficient by reducing the influence of neighboring sample values that differ too much from the current sample value.
[0092] The Cross-Component Adaptive Loop Filter (CC-ALF) can further enhance each chroma component on top of the previously described ALF. The goal of CC-ALF is to refine each chroma component using luminance sample values. This is achieved by applying a diamond-shaped high-pass linear filter and then using the output of this filtering operation for chroma refinement. Figure 8 A system-level diagram of the CC-ALF process for other loop filters is provided. For example... Figure 8 As shown, CC-ALF uses the same input as the luminance ALF to avoid additional steps in the entire loop filtering process.
[0093] 4. Technical problems solved by the disclosed solutions
[0094] The existing design in the latest VVC documentation (in JVET-Q2001-vE / v15) has the following issues:
[0095] 1) The value of slice_type is constrained as follows:
[0096] When nal_unit_type is within the range of IDR_W_RADL to CRA_NUT (inclusive) and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, slice_type should be equal to 2.
[0097] However, the value of slice_type must also be equal to 2 under the following two conditions: i) when intra_only_constraint_flag is equal to 1; and ii) when the NAL cell type is IRAP NAL cell type and the current image is the first image in the current AU.
[0098] 2) The semantics of ph_alf_enabled_flag defined below are ambiguous.
[0099] `ph_alf_enabled_flag` equal to 1 specifies that the adaptive loop filter is enabled for all stripes associated with `PH`, and this adaptive loop filter can be applied to the Y, Cb, or Cr color components in the stripes. `ph_alf_enabled_flag` equal to 0 specifies that the adaptive loop filter can be disabled for one or more or all stripes associated with `PH`. When it does not exist, `ph_alf_enabled_flag` is inferred to be equal to 0.
[0100] 3) Consistency window parameters are always communicated in the PPS signaling, including when the image width and height are the same as the maximum image width and height communicated in the SPS referenced by the PPS. Conversely, consistency window parameters for images with the maximum image width and height are also communicated in the SPS signaling. The signaling for consistency window parameters of images with the maximum image width and height in the PPS is redundant.
[0101] 4) Most SEI message repetitions are limited to a maximum of 4 times within a PU or DU. Repetition of PH, AUD, EOS, and EOBNAL cells is not allowed. Data filling NAL must be permitted.
[0102] The number of times a unit is repeated (e.g., to achieve a constant bit rate). However, there is no limit to the number of repetitions for other non-VCL NAL units (i.e., VPS, SPS, PPS, APS, and DCI NAL units).
[0103] 5. List of technical solutions
[0104] To address the above and other issues, the following summarized methods are disclosed. This invention should be considered as an example of explaining general concepts and not interpreted in a narrow sense. Furthermore, these inventions can be applied individually or combined in any way.
[0105] 1) To address problem 1, constraints on slice_type and / or the signaling of slice_type can depend on conditions relating to general constraint flags / NAL cell type / whether the current image is the first image in the current AU.
[0106] a. In one example, the condition could include:
[0107] i. When intra_only_constraint_flag equals 1.
[0108] ii. When the NAL cell type is IRAP NAL cell type and the current image is the first image in the current AU.
[0109] iii. When an instruction (e.g., the SPS flag) tells that only intra-frame stripes are allowed in a picture (or the CLVS containing the current picture, or any other set of pictures containing the current picture).
[0110] b. The constraint on the slice_type value can be updated such that, additionally, the slice_type value is also required to be equal to 2 if one or more of the first two conditions are true.
[0111] c. Alternatively, when one or more of the first two conditions are true, the signaling for slice_type can be skipped and inferred as an I-slice (i.e., slice_type is 2).
[0112] d. In addition, when the NAL cell type is IRAP NAL cell type and the current layer is an independent layer, the signaling of slice_type can also be skipped and inferred as an I-slice.
[0113] 2) To solve problem 2, you can specify ph_alf_enabled_flag equal to 0 to disable ALF for all stripes of the current image.
[0114] 3) To solve problem 3, we can require that when the image width and height are the maximum image width and height, the value of pps_conformance_window_flag should be equal to 0.
[0115] a. Additionally, it can be specified that if the image width and height are the maximum image width and height, the value of the PPS consistency window syntax element is inferred to be the same as the value of the signaling notification in SPS; otherwise, it is inferred to be equal to 0.
[0116] 4) To address issue 4, one or more of the following constraints can be specified to provide some limitations on the repetition time of VPS, SPS, PPS, APS, and DCI NAL units, while not...
[0117] This affects features such as random access.
[0118] For VPS
[0119] a. When a VPS NAL unit with a specific value of vps_video_parameter_set_id exists in a CVS, the VPS NAL unit should exist in the first AU of the CVS, may exist in any AU with at least one VCL NAL unit having a nal_unit_type in the range of IDR_W_RADL to GDR_NUT (inclusive), and should not exist in any other AU.
[0120] i. Alternatively, the above "IDR_W_RADL to GDR_NUT" is changed to "IDR_W_RADL to RSV_IRAP_12".
[0121] b. The number of VPS NAL units in the PU with a specific value of vps_video_parameter_set_id should not be greater than 1.
[0122] For SPS
[0123] c. Suppose that the associated AU set of CLVS is the first CLVS contained in the decoding order. picture The set of AUs starting from the AU and ending with the AU containing the last image of CLVS in decoding order (including both AUs).
[0124] d. When an SPS NAL unit with a specific value of sps_seq_parameter_set_id exists in the associated AU set associatedAuSet of the CLVS of the reference SPS, the SPS NAL unit should exist in the first AU of associatedAuSet and may exist in any AU of associatedAuSet that has at least one VCL NAL unit with nal_unit_type in the range of IDR_W_RADL to GDR_NUT (inclusive), but should not exist in any other AU.
[0125] i. Alternatively, when an SPS NAL unit with a specific value of sps_seq_parameter_set_id exists in a CLVS, it should exist in the first PU of the CLVS and may exist in any PU with at least one codec stripe NAL unit having a nal_unit_type in the range from IDR_W_RADL to GDR_NUT (inclusive of IDR_W_RADL and GDR_NUT), but should not exist in any other PU.
[0126] ii. Alternatively, in section 4.d or 4.di, change "IDR_W_RADL to..."
[0127] "GDR_NUT" was changed to "IDR_W_RADL to RSV_IRAP_12".
[0128] The number of SPS NAL cells in e.PU with a specific value of sps_seq_parameter_set_id should not be greater than 1.
[0129] For PPS
[0130] The number of PPS NAL cells in f.PU with a specific value of pps_pic_parameter_set_id should not be greater than 1.
[0131] For APS
[0132] The number of APS NAL cells in g.PU with specific values for adaptation_parameter_set_id and aps_params_type should not be greater than 1.
[0133] i. Alternatively, the number of APS NAL cells in the DU with specific values for adaptation_parameter_set_id and aps_params_type should not be greater than 1.
[0134] For DCI
[0135] h. when DCI When a NAL cell exists in a bitstream, it should exist in the first CVS of the bitstream.
[0136] i. When a DCI NAL unit exists in a CVS, it should exist in the first AU of the CVS, may exist in any AU that has at least one VCL NAL unit with a nal_unit_type in the range of IDR_W_RADL to GDR_NUT (inclusive of IDR_W_RADL and GDR_NUT), and should not exist in any other AU.
[0137] The number of DCI NAL units in j.PU should not be greater than 1.
[0138] 6. Example of an implementation plan
[0139] The following are some example embodiments of aspects of the invention summarized in Section 5 above, which can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-Q2001-vE / v15. The most relevant parts that have been added or modified are highlighted in bold italics, and some of the deleted parts are highlighted in bold brackets. Some other changes are editorial in nature or not part of this invention and are therefore not highlighted.
[0140] 6.1. First Embodiment
[0141] This embodiment pertains to item 1.
[0142] The following constraints:
[0143] When nal_unit_type is within the range of IDR_W_RADL to CRA_NUT (inclusive) and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, slice_type should be equal to 2.
[0144] The changes are as follows:
[0145] The value of slice_type should be equal to 2 when intra_only_constraint_flag equals 1 or both of the following conditions are true:
[0146] The value of -nal_unit_type is in the range from IDR_W_RADL to CRA_NUT (inclusive of IDR_W_RADL and CRA_NUT).
[0147] The value of -vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1 or the current image is the first image in the current AU.
[0148] 6.2. Second Embodiment
[0149] This embodiment pertains to item 2.
[0150] The semantics of ph_alf_enabled_flag have been updated as follows:
[0151] A value of 0 for ph_alf_enabled_flag indicates that the adaptive loop filter can be disabled for all stripes associated with PH, including one, more, or more stripes.
[0152] 6.3. Third Embodiment
[0153] This embodiment pertains to item 3.
[0154] 7.4.3.4 Image Parameter Set RBSP Semantics ...
[0156] A value of 1 for `pps_conformance_window_flag` indicates that the consistency clipping window offset parameter follows in PPS. A value of 0 for `pps_conformance_window_flag` indicates that the consistency clipping window offset parameter does not exist in PPS. The value of `pps_conformance_window_flag` should be 0 when `pic_width_in_luma_samples` equals `pic_width_max_in_luma_samples` and `pic_height_in_luma_samples` equals `pic_height_max_in_luma_samples`.
[0157] pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset specify the sample points of the image output from the decoding process in CLVS, according to the rectangular area specified in the coordinates of the output image.
[0158] When pps_conformance_window_flag equals 0, the following applies:
[0159] If pic_width_in_luma_samples equals pic_width_max_in_luma_samples, and pic_height_in_luma_samples equals pic_height_max_in_luma_samples, then the values of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset are inferred to be equal to sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset, respectively.
[0160] Otherwise, the values of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset are inferred to be equal to 0.
[0161] The consistent cropping window contains luminance samples, where the horizontal image coordinates are from SubWidthC*pps_conf_win_left_offset to pic_width_in_luma_samples-(SubWidthC*pps_conf_win_right_offset+1) (including SubWidthC*pps_conf_win_left_offset and pic_width_in_luma_samples-(SubWidthC*pps_conf_win_right_offset+1)), and the vertical image coordinates are from SubHeightC*pps_conf_win_top_offset to pic_height_in_luma_samples-(SubHeightC*pps_conf_win_bottom_offset+1) (including SubHeightC*pps_conf_win_top_offset and pic_height_in_luma_samples-(SubHeightC*pps_conf_win_bottom_offset+1)).
[0162] The value of SubWidthC*(pps_conf_win_left_offset+pps_conf_win_right_offset) should be less than pic_width_in_luma_samples, and the value of SubHeightC*(pps_conf_win_top_offset+pps_conf_win_bottom_offset) should be less than pic_height_in_luma_samples.
[0163] When ChromaArrayType is not equal to 0, the corresponding specified sample points of the two chroma arrays are sample points with image coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the image coordinates of the specified luminance sample point.
[0164] Note 2 - The consistent cropping window offset parameter is applied only to the output. All internal decoding processes are applied to the uncropped image size.
[0165] Let ppsA and ppsB be any two PPSs referencing the same SPS. The requirement for bitstream consistency is that when ppsA and ppsB have the same values for pic_width_in_luma_samples and pic_height_in_luma_samples, respectively, ppsA and ppsB should also have the same values for pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset, respectively.
[0166] When pic_width_in_luma_samples equals pic_width_max_in_luma_samples and pic_height_in_luma_samples equals pic_height_max_in_luma_samples, the bitstream consistency requirement is that pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset are equal to sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset, respectively.
[0167] Figure 1 This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0168] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Codec techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 may be stored or transmitted via a communication connection as indicated by component 1906. The bitstream (or codec) representation of the video received at input 1902, whether stored or communicated, can be used by component 1908 to generate pixel values or transmit as displayable video to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it will be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that inversely represent the codec results will be performed by the decoder.
[0169] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0170] Figure 2 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors (multiple) 3602 can be configured to implement one or more methods described in this document. The memories (multiple) 3604 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described in this document in a hardware circuit system.
[0171] Figure 4 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.
[0172] like Figure 4As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data, and this source device 110 may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110, and this target device 120 may be referred to as a video decoding device.
[0173] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0174] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and related data. A codec picture is a codec representation of a picture. Related data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.
[0175] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0176] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120 configured to interface with an external display device.
[0177] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other current and / or additional standards.
[0178] Figure 5 This is a block diagram illustrating an example of a video encoder 200, which may be... Figure 4 The video encoder 114 in the system 100 shown.
[0179] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 5 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0180] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0181] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[0182] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretive purposes, in Figure 5 The example is represented separately.
[0183] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0184] The mode selection unit 203 can select one of the encoding / decoding modes (e.g., intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction modes (CIIP), where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the block's motion vector (e.g., sub-pixel or integer pixel precision).
[0185] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0186] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip.
[0187] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for reference images in list 0 or list 1 for reference video blocks of the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image in list 0 or list 1, which contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0188] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in list 1. Motion estimation unit 204 can then generate a reference index indicating the reference images in lists 0 and 1 containing the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0189] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process.
[0190] In some examples, the motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, the motion estimation unit 204 may refer to motion information signaling from another video block to inform the motion information of the current video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0191] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0192] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0193] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling Notification.
[0194] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0195] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0196] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0197] Transform unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0198] After the transform unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0199] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block, which is stored in buffer 213.
[0200] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce the video block effect in the video block.
[0201] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.
[0202] Figure 6 This is a block diagram illustrating an example of a video decoder 300, which may be... Figure 4 The video decoder 124 in the system 100 shown.
[0203] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 6 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0204] exist Figure 6 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 5 The encoding process described is the opposite of the decoding process.
[0205] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and from the entropy-coded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by executing AMVP and Merge modes.
[0206] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.
[0207] The motion compensation unit 302 can use an interpolation filter, such as that used by the video encoder 200 during the encoding of a video block, to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.
[0208] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0209] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 304 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0210] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove block artifacts. The decoded video block is then stored in the buffer 307 to provide a reference block for subsequent motion compensation / intra-frame prediction, and also generates decoded video for presentation on the display device.
[0211] The following is a list of preferred solutions for some embodiments.
[0212] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0213] 1. A video processing method (e.g., Figure 3The method 3000 shown includes performing a conversion between a video and a video codec representation comprising one or more images containing one or more stripes (3002), wherein the codec representation conforms to a format rule for controlling the constraints on the stripe type or whether the stripe type is included in the codec representation for a specified field in the codec representation, wherein the field includes a general constraint flag, a network abstraction layer unit type, or whether the video stripe is in the first video image of the accessed unit.
[0214] 2. The method according to Solution 1, wherein the format rule specifies that intra-frame constraints have been enabled for the video stripe.
[0215] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).
[0216] 3. A video processing method, comprising: performing a conversion between a video comprising one or more images containing one or more stripes and a codec representation of the video, wherein the codec representation conforms to a format rule specifying a value of a flag in an image header of the video image to disable adaptive loop filtering of all stripes in the video image.
[0217] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 3).
[0218] 4. A video processing method, comprising: performing a conversion between a video comprising one or more images containing one or more stripes and a codec representation of the video, wherein the codec representation conforms to a format rule specifying that a consistency window flag is set to a disabled mode when the height and width of the current image are equal to the maximum height and maximum width in the video.
[0219] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 4).
[0220] 5. A video processing method, comprising: performing a conversion between a video comprising one or more images containing one or more stripes and a codec representation of the video, wherein the codec representation conforms to a format rule of repetition time of a specified parameter set.
[0221] 6. The method according to Solution 5, wherein the parameter set is an adaptive parameter set, a video parameter set, a sequence parameter set, or an image parameter set.
[0222] 7. The method according to Solution 5, wherein the parameter set is a Decoding Capability Information Network Abstraction Layer Unit (DCINAL).
[0223] 8. The method according to Solution 6, wherein the parameter set is a video parameter set, and wherein a format rule specifies that, where the video parameter set includes a specific value of an identifier field, the video parameter set is included in the first access unit of the encoded / decoded video representation.
[0224] 9. The method according to Solution 8, wherein the format rule further specifies that a set of video parameters with a specific value of an identifier field is included if and only if another access unit has a network abstraction layer type within a range of two pre-specified values.
[0225] 10. The method according to Solution 6, wherein the parameter set is a sequence parameter set, and wherein the format rules specify one or more access units of one or more codec layers whose codec representations are organized as a video sequence, and wherein the format rules specify that a network abstraction layer including a sequence parameter set with a specific identifier value is included in the first access unit in the set of access units of the reference sequence parameter set.
[0226] 11. The method according to Solution 7, wherein the format rules specify that, where DCINAL is included in the codec representation of the video, DCI NAL is included in the first codec video sequence of the video.
[0227] 12. The method according to solution 7 or 11, wherein the format rule further specifies that the number of DCINAL units in the prediction unit is limited to one.
[0228] 13. The method according to any one of solutions 1 to 12, wherein the conversion includes encoding the video into a codec representation.
[0229] 14. The method according to any one of solutions 1 to 12, wherein the conversion includes decoding the codec representation to generate pixel values of the video.
[0230] 15. A video decoding apparatus, comprising a processor configured to implement the method according to one or more of solutions 1 to 14.
[0231] 16. A video encoding apparatus, comprising a processor configured to implement the method according to one or more of solutions 1 to 14.
[0232] 17. A computer program product storing computer code, which, when executed by a processor, causes the processor to perform the method according to any one of solutions 1 to 14.
[0233] 18. A method, apparatus or system described in this document.
[0234] The following list provides a set of second preferred solutions implemented through some embodiments.
[0235] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0236] 1. A video processing method (e.g., Figure 3 The method 300 described in the text includes: performing a conversion between a video and a video bitstream comprising one or more pictures containing one or more stripes (302), wherein the bitstream conforms to a format rule; wherein the format rule specifies whether or how the stripe type of one or more stripes in the bitstream is indicated in a condition that depends on a general constraint flag, a network abstraction layer unit type, or whether the stripe is in the first picture of the accessed unit.
[0237] 2. The method according to Solution 1, wherein the condition includes a general constraint flag indicating intra-frame-only constraints on the stripe.
[0238] 3. The method according to Solution 1, wherein the condition includes that the stripe is in the first picture of the access unit and the network abstraction layer unit type has a specific type, wherein the specific type indicates the intra-frame random access point type.
[0239] 4. The method according to Solution 1, wherein the condition includes a bitstream indication that only intra-frame stripes are permitted in a set of pictures including pictures.
[0240] 5. The method described in Solution 4, wherein the image set corresponds to the image.
[0241] 6. The method according to Solution 4, wherein the set of images corresponds to a codec layer video sequence (CLVS) including the images.
[0242] 7. The method according to any one of solutions 1-6, wherein the format rule specifies that in response to (a) a general constraint flag or a network abstraction layer unit type satisfying the condition, or (b) a stripe in the first picture of the accessed unit, the stripe type value 2 is indicated in the bitstream.
[0243] 8. The method according to Solution 1, wherein the format rule specifies that the stripe type has a value of 2, and in response to (a) a general constraint flag or a network abstraction layer unit type satisfying the condition, or (b) the stripe is in the first picture of the accessed unit, the indication of the stripe type is omitted from the bitstream.
[0244] 9. The method according to Solution 1, wherein the format rule specifies that the stripe type has a value of 2, and in response to (a) the network abstraction layer unit type is an intra-frame random access point type, and (b) the layer to which the picture containing the stripe belongs is an independently decodable layer, the indication of the stripe type is omitted from the bitstream.
[0245] The following solutions illustrate additional examples of exemplary embodiments of the techniques discussed in the previous section (e.g., items 2 and 4).
[0246] 1. A video processing method (e.g., Figure 9 The method described in the image (900) includes: performing a conversion between a video containing multiple stripes and a video bitstream (902), wherein the bitstream conforms to the format rules for the applicability of adaptive loop filtering of all stripes in the image, as specified in the image header.
[0247] 2. The method according to Solution 1, wherein a value of 0 for the flag indicates that adaptive loop filtering is disabled for all stripes in the image.
[0248] 3. The method according to Solution 1, wherein a value of 1 for the flag indicates that adaptive loop filtering is enabled for all stripes in the image.
[0249] 4. A video processing method (e.g., Figure 10 The method 1000 described in the text includes: performing a conversion between a video and a bitstream of a video containing one or more images with one or more stripes according to a format rule (1002), wherein the format rule specifies the repetition time of a set of parameters associated with the video.
[0250] 5. The method described in Solution 4, wherein the parameter set is a video parameter set (VPS).
[0251] 6. The method according to Solution 5, wherein the format rule specifies that, in response to a VPS Network Abstraction Layer (NAL) unit containing a codec video sequence (CVS) of a VPS with a specific identifier value, the VPS NAL unit is included in the first access unit (AU) of the CVS, and the value of another VPS NAL unit based on another AU of the CVS is selectively included in another AU and excluded from the remaining AUs of the CVS.
[0252] 7. The method described in Solution 6, wherein the value of another VPS NAL unit in another AU is in the range from IDR_W_RADL to GDR_NUT.
[0253] 8. The method according to Solution 6, wherein the value of another VPS NAL unit in another AU is in the range from IDR_W_RADL to RSV_IRAP_12.
[0254] 9. The method according to any one of solutions 5-8, wherein the format rule further specifies that no more than one VPS NAL unit with a given identifier value is included in a picture unit (PU) in the bitstream.
[0255] 10. The method according to Solution 4, wherein the parameter set is a sequence parameter set (SPS).
[0256] 11. The method according to solution 10, wherein the codec layer video sequence (CLVS) in the bitstream includes an associated set of access units (AUs), the associated set of AUs including the first AU containing the first picture of the CLVS in decoding order and the last AU containing the last picture of the CLVS in decoding order.
[0257] 12. The method according to solution 11, wherein the format rules specify that, in response to an SPS Network Abstraction Layer (NAL) cell containing an SPS with a specific identifier value, the SPS NAL cell is included in the first AU of the associated AU set, and is selectively included in another AU based on the value of another SPS NAL cell in another AU of the associated set, and excluded from the remaining AUs of the CVS.
[0258] 13. The method according to solution 11, wherein the format rule specifies that, in response to the codec layer video sequence (CVLS) containing an SPS with a specific identifier value in the first picture unit (PU) of the CVLS, the SPS NAL unit is selectively included in the other PU based on the value of the strip NAL unit in another PU of the CVLS, and excluded from the remaining PUs of the CVLS.
[0259] 14. The method according to solution 12, wherein the value of another SPS NAL unit in another AU is in the range from IDR_W_RADL to GDR_NUT.
[0260] 15. The method according to solution 13, wherein the value of the strip NAL cell in another PU is in the range from IDR_W_RADL to GDR_NUT.
[0261] 16. The method according to solution 12, wherein the value of another SPS NAL unit in another AU is in the range from IDR_W_RADL to RSV_IRAP_12.
[0262] 17. The method according to solution 13, wherein the value of the strip NAL cell in another PU is in the range from IDR_W_RADL to RSV_IRAP_12.
[0263] 18. The method according to any one of solutions 11-17, wherein the format rule further specifies that no more than one SPS NAL unit with a given identifier value is included in a picture unit (PU) in the bitstream.
[0264] 19. The method according to Solution 4, wherein the parameter set is a picture parameter set (PPS), and wherein the format rules further specify that no more than one PPS network abstraction layer (NAL) unit with a given identifier value is included in the picture unit (PU) in the bitstream.
[0265] 20. The method described in Solution 4, wherein the parameter set is an adaptive parameter set (APS).
[0266] 21. The method according to solution 20, wherein the format rules further specify that no more than one APS Network Abstraction Layer (NAL) unit of a given identifier value and parameter type is included in a picture unit (PU) in the bitstream.
[0267] 22. The method according to solution 20, wherein the format rules further specify that no more than one APS Network Abstraction Layer (NAL) unit of a given identifier value and parameter type is included in the decoding unit (DU) of the bitstream.
[0268] 23. The method according to Solution 4, wherein the parameter set is a Decoding Capability Information Network Abstraction Layer Unit (DCI NAL).
[0269] 24. The method according to solution 23, wherein the format rules specify that, when present, DCI NAL units are not allowed to be included in a CVS that is not the first codec video sequence (CVS) in the bitstream.
[0270] 25. The method according to solution 23, wherein the format rules specify that, in response to the codec video sequence (CVS) including DCINAL units, the DCINAL units are in the first access unit (AU) of the CVS, and are selectively present in another AU based on whether the other AU includes a video codec layer (VCL) NAL unit with a specific NAL unit identifier value, and are excluded from the remaining AUs of the CVS.
[0271] 26. The method according to solution 21, wherein the specific identifier value is in the range from IDR_W_RADL to GDR_NUT.
[0272] 27. The method according to any one of solutions 23-26, wherein the format rule specifies that the picture unit (PU) includes at most one DCINAL unit.
[0273] The following solutions illustrate additional examples of preferred embodiments of the techniques discussed in the previous section (e.g., item 3).
[0274] 1. A video processing method (e.g., Figure 11 The method described in the text (1100) includes: performing a conversion between a video containing images in a video unit and a video bitstream according to a format rule (1102), wherein the format rule specifies that, in response to the image width being equal to the maximum permissible image width in the video unit and the image height being equal to the maximum permissible image height in the video unit, a consistency window flag corresponding to the image in the image parameter set is set to a value of 0.
[0275] 2. The method according to Solution 1, wherein the maximum permissible image width and the maximum permissible image height are indicated in the sequence parameter set referenced by the video unit.
[0276] 3. According to the method of Solution 2, wherein the format rules specify that, in response to the image width being equal to the maximum allowed image width in the video unit and the image height being equal to the maximum allowed image height in the video unit, the consistency window syntax element is excluded from the image parameter set and is inferred to have the same value as indicated in the sequence parameter set.
[0277] 4. According to the method described in Solution 2, wherein the formatting rules specify that, in response to the image width not being equal to the maximum allowed image width in the video unit or the image height not being equal to the maximum allowed image height in the video unit, the consistency window syntax element is inferred to have a value of 0.
[0278] In the solutions listed above, the conversion involves encoding the video into a bitstream.
[0279] In the solutions listed above, the conversion involves generating video from a bitstream.
[0280] In some embodiments, a video decoding apparatus may include a processor configured to implement the method according to one or more of the solutions described above.
[0281] In some embodiments, a video encoding apparatus including a processor may be configured to implement one or more of the methods described above.
[0282] In some embodiments, a computer-readable medium may store code thereon that, when executed by a processor, causes the processor to implement the method according to any one of the above solutions.
[0283] In some embodiments, a video processing method includes generating a bitstream according to any one or more of the methods described above, and storing the bitstream on a computer-readable medium.
[0284] In some embodiments, a computer-readable medium may store a bitstream thereon, the bitstream being generated from video according to any one or more of the methods described above.
[0285] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block can, for example, correspond to bits juxtaposed or scattered in different places within the bitstream. For example, a macroblock can be encoded according to the error residual values of the transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.
[0286] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances affecting machine-readable propagation signals, or a combination of one or more of them. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device.
[0287] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file storing one or more modules, subroutines, or code sections). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a communications network.
[0288] The processes and logic described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic can also be executed by dedicated logic circuits, and the devices can be implemented as dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0289] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from, transfer data to, or receive data from and transfer data to such mass storage devices. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0290] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or potentially claimed scope, but rather as descriptions of features specific to particular embodiments of a particular art. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be excluded from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[0291] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential manner, or as performing all shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0292] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.
Claims
1. A video processing method, comprising: Perform conversion between a video containing multiple stripes of images and the bitstream of that video. The bitstream conforms to a first format rule, which specifies that the flags in the image header control the applicability of adaptive loop filtering for all stripes in the image. Wherein, a value of 0 for the flag indicates that the adaptive loop filtering is disabled for all stripes in the image, and a value of 1 for the flag indicates that the adaptive loop filtering is enabled for all stripes in the image.
2. The video processing method according to claim 1, further comprising: The conversion between a video containing one or more images with one or more stripes and the bitstream of the video is performed according to the second format rules. The second format rule specifies the repetition time of the parameter set associated with the video.
3. The method according to claim 2, wherein, The parameter set is a video parameter set (VPS).
4. The method according to claim 3, wherein, The second format rule specifies that, in response to a VPS Network Abstraction Layer (NAL) unit containing a codec video sequence (CVS) of a VPS with a specific identifier value, the VPS NAL unit is included in the first access unit (AU) of the CVS, and is allowed to be selectively included in another AU of the CVS, and excluded from the remaining AUs of the CVS.
5. The method according to claim 4, wherein, The other AU has at least one VCL NAL unit with a value in the range from IDR_W_RADL to GDR_NUT.
6. The method according to claim 4, wherein, The other AU has at least one VCL NAL unit with a value in the range of IDR_W_RADL to RSV_IRAP_12.
7. The method according to claim 3, wherein, The second format rule also specifies that no more than one VPS NAL unit with a given identifier value is included in the picture unit (PU) in the bitstream.
8. The method according to claim 2, wherein, The parameter set is the sequence parameter set (SPS).
9. The method according to claim 8, wherein, The codec layer video sequence (CLVS) in the bitstream includes an associated set of access units (AUs), which includes the first AU containing the first picture of the CLVS in the decoding order and the last AU containing the last picture of the CLVS in the decoding order.
10. The method according to claim 9, wherein, The second format rule specifies that, in response to an SPS Network Abstraction Layer (NAL) unit containing an SPS with a specific identifier value, the SPS NAL unit is included in the first AU of the associated AU set, and is allowed to be selectively included in another AU of the associated AU set, and excluded from the remaining AUs of the codec video sequence (CVS).
11. The method according to claim 9, wherein, The second format rule specifies that, in response to an SPS Network Abstraction Layer (NAL) unit containing an SPS with a specific identifier value, the SPS NAL unit is included in the first picture unit (PU) of the codec layer video sequence (CVLS), and is allowed to be selectively included in another PU of the CVLS, and excluded from the remaining PUs of the CVLS.
12. The method according to claim 10, wherein, The other AU has at least one VCL NAL unit with a value in the range from IDR_W_RADL to GDR_NUT.
13. The method according to claim 11, wherein, The other PU has striped NAL cells with values ranging from IDR_W_RADL to GDR_NUT.
14. The method of claim 10, wherein, The other AU has VCL NAL units with values ranging from IDR_W_RADL to RSV_IRAP_12.
15. The method according to claim 11, wherein, The other PU has striped NAL cells with values ranging from IDR_W_RADL to RSV_IRAP_12.
16. The method according to claim 9, wherein, The second format rule also specifies that no more than one SPS NAL unit with a given identifier value is included in the picture unit (PU) of the bitstream.
17. The method according to claim 2, wherein, The parameter set is a picture parameter set (PPS), and the second format rule further specifies that no more than one PPS network abstraction layer (NAL) unit with a given identifier value is included in the picture unit (PU) in the bitstream.
18. The method according to claim 2, wherein, The parameter set is an adaptive parameter set (APS).
19. The method according to claim 18, wherein, The second format rule also specifies that no more than one APS Network Abstraction Layer (NAL) unit with a given identifier value and parameter type is included in the picture unit (PU) in the bitstream.
20. The method according to claim 18, wherein, The second format rule also specifies that no more than one APS Network Abstraction Layer (NAL) unit with a given identifier value and parameter type is included in the decoding unit (DU) in the bitstream.
21. The method according to claim 2, wherein, The parameter set is a Decoding Capability Information Network Abstraction Layer (DCINAL) unit.
22. The method according to claim 21, wherein, The second format rule specifies that, when present, the DCI NAL unit is not allowed to be included in a CVS that is not the first codec video sequence (CVS) in the bitstream.
23. The method according to claim 21, wherein, The second format rule specifies that, in response to the codec video sequence (CVS) including the DCI NAL unit, the DCI NAL unit is in the first access unit (AU) of the CVS, and is selectively present in another AU based on whether the other AU includes a video codec layer (VCL) NAL unit with a specific NAL unit identifier value, and is excluded from the remaining AUs of the CVS.
24. The method according to claim 23, wherein, The specific NAL unit identifier value is in the range from IDR_W_RADL to GDR_NUT.
25. The method according to claim 21, wherein, The second format rule specifies that a picture unit (PU) includes at most one DCI NAL unit.
26. The method according to any one of claims 1 to 25, wherein, The conversion includes generating the video from the bitstream.
27. The method according to any one of claims 1 to 25, wherein, The conversion includes encoding the video into the bitstream.
28. A video decoding apparatus comprising a processor configured to implement the method according to any one of claims 1 to 26.
29. A video encoding apparatus comprising a processor configured to implement the method according to any one of claims 1 to 25, 27.
30. A computer-readable medium storing code, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 27.
31. A video processing method, comprising: The method according to any one of claims 1 to 27 generates a bitstream, and The bitstream is stored on a computer-readable medium.
32. A computer-readable medium storing a bitstream, said bitstream being generated from video by a processor according to any one of claims 1 to 27.