Limiting of time domain reference in GDR-based video coding
By introducing Progressive Decoding Refresh (GDR) and Adaptive Parameter Set (APS) reference constraints into video encoding and decoding, the problem of excessive bandwidth requirements for high-resolution and high-frame-rate videos is solved, achieving more efficient video encoding, decoding, and storage.
Patent Information
- Application Number
- CN202480027088.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-28
- Filing Date
- 2024-04-26
- Publication Date
- 2025-12-12
AI Technical Summary
Existing video encoding and decoding technologies require ever-increasing bandwidth when processing high-resolution and high-frame-rate videos, leading to excessive network load, and existing standards are unable to effectively reduce bitrates to improve compression efficiency.
A video encoding and decoding method based on Progressive Decoding Refresh (GDR) is adopted. By determining and applying reference constraints in the Adaptive Parameter Set (APS), the video encoding and decoding process is optimized, and efficient conversion between video data and bitstream is achieved.
It effectively reduces the bit rate in the video encoding and decoding process, improves compression efficiency, reduces network bandwidth requirements, and adapts to the video decoding needs of different devices.
Smart Images

Figure CN121128162A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority and benefit to International Patent Application No. PCT / CN2023 / 091532, filed on April 28, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the generation, storage, and use of digital audio and video media information in file formats. Background Technology
[0004] Digital video accounts for the largest share of bandwidth used on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the Invention
[0005] The first aspect relates to a method for processing video data, comprising: when performing progressive decode refresh (GDR) based video encoding and decoding, determining a reference to apply one or more constraints to an adaptive parameter set (APS); and performing a conversion between visual media data and a bitstream based on the APS.
[0006] The second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any of the aspects described above.
[0007] The third aspect relates to a non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that when the computer-executable instructions are executed by a processor, the video codec device performs the method of any of the preceding aspects.
[0008] The fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining, when performing progressive decode refresh (GDR) based video encoding and decoding, to apply one or more constraints to a reference in an adaptive parameter set (APS); and generating the bitstream based on the determination.
[0009] The fifth aspect relates to a method for storing a bitstream of video, comprising: when performing progressive decode refresh (GDR) based video encoding and decoding, determining a reference to apply one or more constraints to an adaptive parameter set (APS); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] The sixth aspect relates to the methods, apparatus, or systems described in this disclosure.
[0011] For clarity, any of the above embodiments may be combined with one or more other above embodiments to create new embodiments within the scope of this disclosure.
[0012] These and other features will become clearer through the following detailed description in conjunction with the accompanying drawings and claims. Attached Figure Description
[0013] To gain a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals denote like parts.
[0014] Figure 1 An example of the nominal vertical and horizontal positions of 4:2:2 luminance and chrominance samples in an image is shown.
[0015] Figure 2 A sample encoder block diagram is shown.
[0016] Figure 3 An example image is shown, segmented into raster scan strips.
[0017] Figure 4 An example image is shown, segmented into rectangular scan strips.
[0018] Figure 5 An example image showing the structure divided into bricks is shown.
[0019] Figures 6A-6C An example of a codec tree block (CTB) spanning image boundaries is shown.
[0020] Figure 7 An example of an intra-frame prediction mode is shown.
[0021] Figure 8 An example of a block boundary is shown in the image.
[0022] Figure 9 An example of pixels used in a filter is shown.
[0023] Figure 10 An example of the filter shape for an adaptive loop filter (ALF) is shown.
[0024] Figure 11 An example of the transform coefficients supported by a 5×5 diamond filter is shown.
[0025] Figure 12 An example of relative coordinates supported by a 5×5 rhombus filter is shown.
[0026] Figure 13 This is a block diagram illustrating an example video processing system.
[0027] Figure 14 This is a block diagram of an example video processing device.
[0028] Figure 15 This is a flowchart of an example method for video processing.
[0029] Figure 16 This is a block diagram illustrating an example video codec system.
[0030] Figure 17 This is a block diagram showing an example encoder.
[0031] Figure 18 This is a block diagram showing an example decoder.
[0032] Figure 19 This is a schematic diagram of an example encoder. Detailed Implementation
[0033] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or yet to be developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and embodiments shown below, including the exemplary designs and implementations shown and described herein, but modifications can be made within the scope of the appended claims and all their equivalents.
[0034] Chapter headings are used in this disclosure for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the embodiments described herein are applicable to other video codec protocols and designs.
[0035] 1. Preliminary Discussion
[0036] This disclosure relates to video encoding and decoding techniques. Specifically, it relates to loop filters and other encoding and decoding tools in image / video encoding and decoding. These ideas can be applied individually or in various combinations to video codecs, such as High Efficiency Video Codec (HEVC), Multi-Functional Video Codec (VVC), or other video encoding and decoding techniques.
[0037] 2. Abbreviation
[0038] This disclosure includes the following abbreviations: Advanced Video Coding (ITU-T H.264 | ISO / IEC 14496-10) (AVC), Cocoded Picture Buffer (CPB), Pure Random Access (CRA), Codec Tree Unit (CTU), Coded Video Sequence (CVS), Decoded Picture Buffer (DPB), Decoded Parameter Set (DPS), General Constraint Information (GCI), High-Efficiency Video Coding, also known as ITU-T H.265 | ISO / IEC 23008-2 (HEVC), Joint Exploration Model (JEM), Motion Constraint Piece Set (MCTS), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Header (PH), Picture Parameter Set (PPS), Grade, Layer and Level (PTL), Picture Unit (PU), Reference Picture Resampling (RPR), Raw Byte Sequence Payload (RBSP), Supplementary Enhancement Information (SEI), Strip Header (SH), Sequence Parameter Set (SPS), Video Coding Layer (VCL), Video Parameter Set (VPS), Multifunctional Video Coding, also known as ITU-T H.265 (HEVC). H.266 | ISO / IEC 23090-3 (VVC), VVC Test Model (VTM), Video Availability Information (VUI), Transform Unit (TU), Codec Unit (CU), Deblocking Filter (DF), Sample Adaptive Compensation (SAO), Adaptive Loop Filter (ALF), Codec Block Flag (CBF), Quantization Parameter (QP), Rate Distortion Optimization (RDO), Bilateral Filter (BF), Progressive Decode Refresh (GDR).
[0039] 3. Video codec standards
[0040] Video coding standards have evolved primarily through the development of standards by the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T produced the H.261 and H.263 standards, ISO / IEC produced the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Vision, and the two organizations jointly produced the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / HEVC standard[1]. Starting with H.262, video coding standards are based on a hybrid video coding architecture, which utilizes temporal prediction plus transform coding. In order to explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by the Video Coding Experts Group (VCEG) and MPEG. JVET adopted many methods and incorporated them into a reference software called the Joint Exploration Model (JEM)[2]. When the Multifunctional Video Coding (VVC) project was officially launched, JVET was renamed the Joint Video Experts Group (JVET). VVC is a codec standard that aims to reduce the bit rate by 50% compared to HEVC. The VVC working draft and VVC Test Model (VTM) are constantly being updated.
[0041] A sample version of the VVC draft, namely the Multi-Functional Video Codec (Draft 10), can be found at: https: / / jvet-experts.org / doc_end_user / documents / 19_Teleconference / wg11 / JVET-S2001-v17.zip. A sample version of the VVC reference software, named VTM, can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet-u-ee2 / VVCSoftware_VTM / - / tree / VTM-11.2.
[0042] ITU-T VCEG and ISO / IEC MPEG Joint Technical Committee (JTC) 1 / Subcommittee (SC) 29 / Working Group (WG) 11 are investigating the potential need for standardization of future video codec technologies with compression capabilities significantly exceeding the current VVC standard. This future standardization effort could take the form of extended forms of VVC (multiple extensions) or entirely new standards. These groups are conducting this exploratory activity through a joint collaborative effort called JVET to evaluate compression technology designs proposed by experts in the field. JVET has established the first exploratory experiment (EE), and reference software called the Enhanced Compression Model (ECM) is in use. The test model ECM is continuously being updated.
[0043] 3.1 Color Space and Chromaticity Downsampling
[0044] A color space, also known as a color model (or color system), is a mathematical model that describes a range of colors as tuples of numbers, such as 3 or 4 values or color components (e.g., RGB). In general, a color space is a refinement of a coordinate system and its subspaces. For video compression, the most commonly used color spaces are Luminance, Blue Difference, and Red Difference (YCbCr) and Red, Green, and Blue (RGB).
[0045] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luminance component, and CB and CR are the blue and red difference chromaticity components. Y' (with an apostrophe) is distinguished from Y, which is luminance, meaning that light intensity is non-linearly encoded based on gamma-corrected RGB primary colors.
[0046] Chromaticity downsampling is a practice of encoding images by applying a lower resolution to chromaticity information compared to luminance information. It takes advantage of the fact that the human visual system is less sensitive to color differences than to luminance differences. 3.1.1 4:4:4
[0048] In a 4:4:4 scheme, each of the three Y'CbCr components has the same sample rate, thus eliminating chromaticity downsampling. This scheme is sometimes used in high-end cinema scanners and film post-production. 3.1.2 4:2:2
[0050] In 4:2:2, both chroma components are sampled at half the luminance sampling rate. The horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third with little to no visual difference. An example of the nominal vertical and horizontal positions of the 4:2:2 color format is shown in... Figure 1 It is depicted in the middle. 3.1.3 4:2:0
[0052] In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved because the Cb and Cr channels are sampled only on each alternating row, so the data rate is the same. Cb and Cr are downsampled by a factor of 2 in both the horizontal and vertical directions. There are three variations of the 4:2:0 scheme with different horizontal and vertical positions.
[0053] In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are located between pixels vertically (at gap positions). In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located at gap positions, in the middle of alternating luminance samples. In 4:2:0 DV, Cb and Cr are co-located horizontally. Vertically, they are co-located on alternating lines.
[0054] Table 1. SubWidthC and SubHeightC values derived from chroma_format_idc and separate_colour_plane_flag
[0055] chroma_format_idc separate_colour_plane_flag Color format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1
[0056] 3.2 Example codec stream for video codecs
[0057] Figure 2 An example of a VVC encoder block diagram is shown, comprising three loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Compensation (SAO), and ALF. Unlike DF, which uses predefined filters, SAO and ALF utilize the original samples of the current image, reducing the mean square error between the original and reconstructed samples by adding an offset and by applying a Finite Impulse Response (FIR) filter, respectively, and by utilizing the encoded / decoded side information through signal transmission offset and filter coefficients. ALF is located in the last processing stage of each image and can be viewed as a tool attempting to capture and repair artifacts caused by previous stages.
[0058] 3.3 Definition of Video / Encoding / Decoding Unit
[0059] An image is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of the image. A slice can be divided into one or more bricks, each brick comprising multiple CTU rows within the slice. A slice that is not divided into multiple bricks can also be called a brick. However, a brick that is a proper subset of a slice cannot be called a slice. A strip contains several slices of an image or several bricks of a slice.
[0060] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, the stripe contains a sequence of slices from a raster scan of the image. In rectangular stripe mode, the stripe contains multiple tiles that together form a rectangular area of the image. The tiles within the rectangular stripe are arranged in the order of the stripe's raster scan. Figure 3 An example of raster scan strip segmentation of an image (with 18×12 luminance CTUs) is shown, in which the image is divided into 12 slices and 3 raster scan strips.
[0061] Figure 4 An example of rectangular strip segmentation of an image (with 18×12 luminance CTUs) is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0062] Figure 5 An example of an image divided into slices, bricks, and rectangular strips is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows), 11 bricks (the top left slice contains 1 brick, the top right slice contains 5 bricks, the bottom left slice contains 2 bricks, and the bottom right slice contains 3 bricks) and 4 rectangular strips.
[0063] 3.3.1 CTU / CTB Dimensions
[0064] In VVC, the CTU size transmitted via signaling in the Sequence Parameter Set (SPS) by the syntax element log2_ctu_size_minus2 can be as small as 4×4.
[0065] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax
[0066] seq_parameter_set_rbsp() { descriptor sps_decoding_parameter_set_id u(4) sps_video_parameter_set_id u(4) sps_max_sub_layers_minus1 u(3) sps_reserved_zero_5bits u(5) profile_tier_level( sps_max_sub_layers_minus1 ) gra_enabled_flag u(1) sps_seq_parameter_set_id ue(v) chroma_format_idc ue(v) if (chroma_format_idc == 3) separate_colour_plane_flag u(1) pic_width_in_luma_samples ue(v) pic_height_in_luma_samples ue(v) conformance_window_flag u(1) if( conformance_window_flag ) { conf_win_left_offset ue(v) conf_win_right_offset ue(v) conf_win_top_offset ue(v) conf_win_bottom_offset ue(v) } bit_depth_luma_minus8 ue(v) bit_depth_chroma_minus8 ue(v) log2_max_pic_order_cnt_lsb_minus4 ue(v) sps_sub_layer_ordering_info_present_flag u(1) for( i = ( sps_sub_layer_ordering_info_present_flag? 0 : sps_max_sub_layers_minus1 ); i <= sps_max_sub_layers_minus1; i++ ) { sps_max_dec_pic_buffering_minus1[ i ] ue(v) sps_max_num_reorder_pics[i] ue(v) sps_max_latency_increase_plus1[i] ue(v) } long_term_ref_pics_flag u(1) sps_idr_rpl_present_flag u(1) rpl1_same_as_rpl0_flag u(1) for( i = 0; i <!rpl1_same_as_rpl0_flag? 2 : 1; i++ ) { num_ref_pic_lists_in_sps[i] ue(v) for( j = 0; j < num_ref_pic_lists_in_sps[ i ]; j++) ref_pic_list_struct( i, j ) } qtbtt_dual_tree_intra_flag u(1) log2_ctu_size_minus2 ue(v) log2_min_luma_coding_block_size_minus2 ue(v) partition_constraints_override_enabled_flag u(1) sps_log2_diff_min_qt_min_cb_intra_slice_luma ue(v) sps_log2_diff_min_qt_min_cb_inter_slice ue(v) sps_max_mtt_hierarchy_depth_inter_slice ue(v) sps_max_mtt_hierarchy_depth_intra_slice_luma ue(v) if( sps_max_mtt_hierarchy_depth_intra_slice_luma != 0 ) { sps_log2_diff_max_bt_min_qt_intra_slice_luma ue(v) sps_log2_diff_max_tt_min_qt_intra_slice_luma ue(v) } if( sps_max_mtt_hierarchy_depth_inter_slices != 0 ) { sps_log2_diff_max_bt_min_qt_inter_slice ue(v) sps_log2_diff_max_tt_min_qt_inter_slice ue(v) } if( qtbtt_dual_tree_intra_flag ) { sps_log2_diff_min_qt_min_cb_intra_slice_chroma ue(v) sps_max_mtt_hierarchy_depth_intra_slice_chroma ue(v) if ( sps_max_mtt_hierarchy_depth_intra_slice_chroma != 0 ) { [[ID= } } … }
[0067] `log2_ctu_size_minus2` plus 2 specifies the luma codec block size for each CTU. `log2_min_luma_coding_block_size_minus2` plus 2 specifies the minimum luma codec block size. The variables `CtbLog2SizeY`, `CtbSizeY`, `MinCbLog2SizeY`, `MinCbSizeY`, `MinTbLog2SizeY`, `MaxTbLog2SizeY`, `MinTbSizeY`, `MaxTbSizeY`, `PicWidthInCtbsY`, `PicHeightInCtbsY`, `PicSizeInCtbsY`, `PicWidthInMinCbsY`, `PicHeightInMinCbsY`, `PicSizeInMinCbsY`, `PicSizeInSamplesY`, `PicWidthInSamplesC`, and `PicHeightInSamplesC` are derived as follows:
[0068] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9)
[0069] CtbSizeY = 1 << CtbLog2SizeY (7-10)
[0070] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-11)
[0071] MinCbSizeY = 1 << MinCbLog2SizeY (7-12)
[0072] MinTbLog2SizeY = 2 (7-13)
[0073] MaxTbLog2SizeY = 6 (7-14)
[0074] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)
[0075] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)
[0076] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (7-17)
[0077] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY )(7-18)
[0078] PicSizeInCtbsY = PicWidthInCtbsY PicHeightInCtbsY (7-19)
[0079] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)
[0080] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)
[0081] PicSizeInMinCbsY = PicWidthInMinCbsY PicHeightInMinCbsY (7-22)
[0082] PicSizeInSamplesY = pic_width_in_luma_samples pic_height_in_luma_samples (7-23)
[0083] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)
[0084] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)
[0085] 3.3.2 CTUs in a picture
[0086] Assume that the CTB / LCU size is indicated by M×N (usually M equals N), and for a CTB located at the picture boundary (or slice or strip or other type of boundary, the picture boundary is taken as an example), K×L samples are within the picture boundary, where K < M or L < N. For those CTBs depicted as in , the CTB size still equals M×N. However, as shown in , the lower boundary of the CTB is outside the picture; as shown in , the right boundary of the CTB is outside the picture; or as shown in , the lower boundary / right boundary of the CTB is outside the picture.
[0087] 3.4 Intra prediction
[0088] To capture any edge direction presented in natural videos, the number of directional intra modes is extended from 33 as used in HEVC to 65. The extended directional modes are depicted in , and the planar and DC modes remain the same. These denser directional intra prediction modes apply to all block sizes and apply to both luminance and chrominance intra prediction.
[0089] As shown in , the angular intra prediction directions can be defined as clockwise from 45 degrees to -135 degrees. In VTM, for non-square blocks, several angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled and remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes remains unchanged, e.g., 67, and the intra mode coding and decoding remain unchanged.
[0090] In HEVC, each intra-coded block has a square shape, and the length of each side of the block is a power of 2. Therefore, no division is needed to generate intra-prediction values using DC mode. In VVC, blocks can have rectangular shapes, which typically requires division for each block. To avoid division for DC prediction, only the longer side is used to calculate the average of non-square blocks.
[0091] 3.5 Inter-frame prediction
[0092] For each inter-frame predicted CU, motion parameters include motion vectors, reference picture indexes, reference picture list usage indexes, and extended information for new encoding / decoding features used in VVC, which will be used for sample generation in the inter-frame predicted CU. Motion parameters can be transmitted via signaling in an explicit or implicit manner. When a CU is encoded / decoded in skip mode, the CU is associated with a PU and does not have significant residual coefficients, encoded / decoded motion vector increments, and / or reference picture indexes. A Merge mode is defined, whereby motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates and extended scheduling introduced in VVC. The Merge mode can be applied to any inter-frame predicted CU, not just skip mode. An alternative to the Merge mode is explicit transmission of motion parameters, where motion vectors, corresponding reference picture indexes for each reference picture list, reference picture list usage flags, and other useful information are explicitly transmitted via signaling for each CU.
[0093] 3.6 Deblocking Filter
[0094] Deblocking filtering is an example of a loop filter in a video codec. In VVC, the deblocking filtering process is applied to CU boundaries, transform subblock boundaries, and predictive subblock boundaries. Predictive subblock boundaries include prediction unit boundaries introduced by subblock-based temporal motion vector prediction (SbTMVP) and affine modes. Transform subblock boundaries include transform unit boundaries introduced by subblock transform (SBT) and intra-fractional sub-segmentation (ISP) modes, as well as transforms resulting from the implicit partitioning of large CUs. The processing order of the deblocking filter is defined as first performing horizontal filtering on the vertical edges of the entire image, and then performing vertical filtering on the horizontal edges. This specific order allows multiple horizontal or vertical filtering processes to be applied in parallel threads. The filtering process can also be implemented on a CTB-by-CTB basis with only a small processing latency.
[0095] The vertical edges in the image are first filtered. Then, the horizontal edges in the image are filtered using samples modified by the vertical edge filtering process as input. Vertical and horizontal edges in the CTB of each CTU are processed individually on a per-code / decode unit basis. The vertical edges of the codec blocks in a codec unit are filtered, starting from the edge on the left-hand side of the codec block and proceeding geometrically through the edges towards the right-hand side of the codec block. The horizontal edges of the codec blocks in a codec unit are filtered, starting from the edge on the top of the codec block and proceeding geometrically through the edges towards the bottom of the codec block.
[0096] This is a schematic diagram 800 of sample point 802 within an 8×8 block of sample point 804. As shown, schematic diagram 800 includes the horizontal and vertical block boundaries on the 8×8 grids 806 and 808, respectively. In addition, schematic diagram 800 depicts the non-overlapping blocks of 8×8 sample points 810, which can be de-blocked in parallel.
[0097] 3.6.1 Boundary Decision
[0098] The filter is applied to the 8×8 block boundaries. Additionally, such boundaries must be transform block boundaries or codec sub-block boundaries, for example, due to the use of Affine Motion Prediction (ATMVP). For other boundaries, deblocking filtering is disabled.
[0099] 3.6.2 Boundary Strength Calculation
[0100] For transform block boundaries / encoder / decoder sub-block boundaries, if the boundary is located in an 8×8 grid, the boundary can be filtered, and the settings of bS[xDi][yDj] (where [xDi][yDj] represents the coordinates) of the edge are defined as in Tables 2 and 3, respectively.
[0101] Table 2. Boundary Strength (when SPS IBC is disabled)
[0102] Y U V 5 At least one of the adjacent blocks is an intra-frame codec. 2 2 2 4 At least one of the TU boundary and adjacent blocks has non-zero transformation coefficients. 1 1 1 3 The number of reference images or MVs for adjacent blocks differs (1 for one-way prediction and 2 for two-way prediction). 1 N / A N / A 2 The absolute difference between motion vectors of adjacent blocks belonging to the same reference image is greater than or equal to an integer brightness sample. 1 N / A N / A 1 other 0 0 0
[0103] Table 3. Boundary Strength (when SPS IBC is enabled)
[0104] Priority condition Y U V 8 At least one of the adjacent blocks is an intra-frame codec. 2 2 2 7 At least one of the TU boundary and adjacent blocks has non-zero transformation coefficients. 1 1 1 6 Adjacent blocks use different prediction modes (e.g., one uses IBC, and the other uses inter-frame encoding / decoding). 1 5 The absolute difference between the IBC and the motion vector belonging to the adjacent block is greater than or equal to an integer brightness sample. 1 N / A N / A 4 The number of reference images or MVs for adjacent blocks differs (1 for one-way prediction and 2 for two-way prediction). 1 N / A N / A 3 The absolute difference between motion vectors of adjacent blocks belonging to the same reference image is greater than or equal to an integer brightness sample. 1 N / A N / A 1 other 0 0 0
[0105] 3.6.3 Deblocking decision for the luminance component
[0106] Figure 9This is an example of pixels involved in filter on / off decisions and strong / weak filter selection. A wider and stronger luminance filter is used only if conditions 1, 2, and 3 are all true. Condition 1 is the "bulk condition." This condition detects whether samples on the P-side and Q-side belong to a bulk, represented by the variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.
[0107] bSidePisLargeBlk = ((edge type is vertical and p0 belongs to CU with width >= 32) || (edge type is horizontal and p0 belongs to CU with height >= 32))? TRUE: FALSE
[0108] bSideQisLargeBlk = ((edge type is vertical and q0 belongs to CU with width >= 32) || (edge type is horizontal and q0 belongs to CU with height >= 32))? TRUE: FALSE
[0109] Based on bSidePisLargeBlk and bSideQisLargeBlk, condition 1 is defined as follows:
[0110] Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk) ? TRUE: FALSE
[0111] Next, if condition 1 is true, condition 2 will be further examined. First, the following variables are derived:
[0112] First, derive dp0, dp3, dq0, and dq3 using the HEVC method.
[0113] if (p side is greater than or equal to 32)
[0114] dp0 = (dp0 + Abs(p50 - 2)) p40 + p30 + 1) >> 1
[0115] dp3 = (dp3 + Abs(p53 - 2)) p43 + p33 + 1) >> 1
[0116] if (q side is greater than or equal to 32)
[0117] dq0 = (dq0 + Abs(q50 - 2)) q40 + q30 + 1) >> 1
[0118] dq3 = (dq3 + Abs(q53 - 2)) q43 + q33 + 1) >> 1
[0119] Condition 2 = (d < β) ? TRUE: FALSE
[0120] Where d = dp0 + dq0 + dp3 + dq3.
[0121] If conditions 1 and 2 are valid, then further check whether any of the blocks uses a sub-block:
[0122] If (bSidePisLargeBlk)
[0123] {
[0124] If (block P's mode == SUBBLOCKMODE)
[0125] Sp = 5
[0126] else
[0127] Sp = 7
[0128] }
[0129] else
[0130] Sp = 3
[0131] If (bSideQisLargeBlk)
[0132] {
[0133] If (block Q's mode == SUBBLOCKMODE)
[0134] Sq = 5
[0135] else
[0136] Sq = 7
[0137] }
[0138] else
[0139] Sq = 3
[0140] Finally, if both conditions 1 and 2 are valid, the deblocking method will check condition 3 (the strong filter condition), which is defined as follows. In condition 3, StrongFilterCondition, the following variables are derived:
[0141] Derive dpq using the HEVC method.
[0142] Following the HEVC derivation method, sp3 = Abs( p3 - p0 )
[0143] if (p side is greater than or equal to 32)
[0144] if (Sp == 5)
[0145] sp3 = ( sp3 + Abs( p5 - p3 ) + 1) >> 1
[0146] else
[0147] sp3 = ( sp3 + Abs( p7 - p3 ) + 1) >> 1
[0148] Following the HEVC derivation method, sq3 = Abs( q0 - q3 )
[0149] if (q side is greater than or equal to 32)
[0150] If (Sq == 5)
[0151] sq3 = ( sq3 + Abs( q5 - q3 ) + 1) >> 1
[0152] else
[0153] sq3 = ( sq3 + Abs( q7 - q3 ) + 1) >> 1
[0154] According to HEVC, StrongFilterCondition = (dpq < (β >> 2), sp3 + sq3 < (3)). β >> 5), and Abs(p0 - q0) is less than (5). tC + 1 ) >> 1) ? TRUE : FALSE.
[0155] 3.6.4 A more robust deblocking filter for luminance
[0156] When samples on either side of the boundary belong to a large block, a bilinear filter is used. Samples belonging to a large block are defined as those with a vertical edge width >= 32 and a horizontal edge height >= 32. The bilinear filter is listed below. Then, in the above HEVC deblocking, the block boundary samples pi (i = 0 to Sp-1) and qi (i = 0 to Sq-1) (pi and qi are the i-th samples in the row used for filtering the vertical edge, or the i-th samples in the column used for filtering the horizontal edge) are replaced by the following linear interpolation:
[0157]
[0158]
[0159] in and The item is the amplitude limit related to the above-mentioned location, and , , , and It is given below.
[0160] 3.6.5 Color Deblocking Decision
[0161] A strong chroma filter is used on both sides of the block boundary. Here, the chroma filter is selected when both sides of the chroma edge are greater than or equal to 8 (chroma position) and the following decision with three conditions is met: The first is for the boundary strength and the decision on the size of the block. The filter can be applied when the block width or height orthogonally spanning the block edge in the chroma sample domain is equal to or greater than 8. The second and third are essentially the same as those used for HEVC luminance deblocking decisions, which are the on / off decision and the strong filter decision, respectively.
[0162] In the first decision, the boundary strength (bS) is modified for chroma filtering, and the conditions are checked sequentially. If a condition is met, the remaining conditions with lower priority are skipped. Chroma deblocking is performed when bS equals 2, or when bS equals 1 when a large boundary is detected. The second and third conditions are essentially the same as the HEVC luma strong filter decision below.
[0163] In the second condition, d is then derived using the HEVC luminance deblocking method. The second condition will be true when d is less than β. In the third condition, StrongFilterCondition is derived as follows:
[0164] Derive dpq using the HEVC method;
[0165] Derive sp3 = Abs( p3 - p0 ) using the HEVC method; and
[0166] The derivation of sq3 = Abs( q0 - q3 ) follows the HEVC method.
[0167] According to the HEVC design, StrongFilterCondition = (dpq < (β >> 2), sp3 + sq3 < (β >> 3), and Abs(p0 - q0) < (5). tC + 1) >> 1)
[0168] 3.6.6 Strong Deblocking Filter for Chroma
[0169] The following strong deblocking filter is defined for chroma:
[0170] p2′= (3 p3+2 p2+p1+p0+q0+4) >> 3
[0171] p1′= (2 p3+p2+2 p1+p0+q0+q1+4) >> 3
[0172] p0′= (p3+p2+p1+2 p0+q0+q1+q2+4) >> 3
[0173] The example chroma filter performs deblocking on a 4×4 chroma sample grid.
[0174] 3.6.7 Location-related amplitude limiting
[0175] Position-dependent limiting (tcPD) is applied to the output samples of strong and long filters that involve modifying 7, 5, and 3 samples at the boundaries. Assuming a quantization error distribution, the limiting value can be increased for samples expected to have higher quantization noise, thus expecting a higher deviation between the reconstructed sample value and the true sample value.
[0176] For each P or Q boundary filtered by an asymmetric filter, based on the result of the decision process, the position-related threshold table is selected from two tables (e.g., Tc7 and Tc3 listed below) that are provided to the decoder as edge information:
[0177] Tc7 = { 6, 5, 4, 3, 2, 1, 1};Tc3 = { 6, 4, 2};
[0178] tcPD = (Sp == 3) ? Tc3 : Tc7;
[0179] tcQD = (Sq == 3) ? Tc3 : Tc7;
[0180] For P or Q boundaries filtered by a short symmetric filter, apply a lower-amplitude position correlation threshold:
[0181] Tc3 = { 3, 2, 1};
[0182] After defining the threshold, the filtered p'i and q'i sample values are clipped according to the tcP and tcQ clipping values:
[0183] p''i = Clip3(p'i + tcPi, p'i – tcPi, p'i );
[0184] q''j = Clip3(q'j + tcQj, q'j – tcQ j, q'j );
[0185] Where p'i and q'i are the filtered sample values, p''i and q''j are the output sample values after clipping, and tcPi is the clipping threshold derived from the VVC tc parameters, tcPD, and tcQD. The function Clip3 is the clipping function specified in VVC.
[0186] 3.6.8 Sub-block Removal and Adjustment
[0187] To enable parallel-friendly deblocking using both long filters and sub-block deblocking, the long filter is restricted to modifying a maximum of 5 samples on the side using sub-block deblocking (AFFINE or ATMVP or decoder-side motion vector refinement (DMVR)), as shown in the long filter's brightness control. Extending this, sub-block deblocking is adjusted such that sub-block boundaries on the 8×8 grid near the CU or implicit TU boundaries are restricted to modifying a maximum of two samples on each side.
[0188] The following applies to sub-block boundaries that are not aligned with the CU boundary.
[0189] If (block Q's mode == SUBBLOCKMODE && edge != 0) {
[0190] if (!(implicitTU && (edge == (64 / 4))))
[0191] if (edge == 2 || edge == (orthogonalLength - 2) || edge== (56 / 4) || edge == (72 / 4))
[0192] Sp = Sq = 2;
[0193] else
[0194] Sp = Sq = 3;
[0195] else
[0196] Sp = Sq = bSideQisLargeBlk ? 5:3
[0197] }
[0198] Where edge = 0 corresponds to the CU boundary, edge = 2 or orthogonalLength-2 corresponds to the sub-block boundary 8 samples away from the CU boundary, etc. If implicit partitioning of TU is used, then implicit TU is true.
[0199] 3.7 Sample point adaptive compensation
[0200] Sample Adaptive Compensation (SAO) is applied to the reconstructed signal after the deblocking filter using an offset specified by the encoder for each CTB. The video encoder first determines whether the SAO process will be applied to the current slice. If SAO is applied to a slice, each CTB is classified into one of five SAO types as shown in Table 4. The concept of SAO is to classify pixels into multiple categories and reduce distortion by adding an offset to the pixels in each category. SAO operations include Edge Offset (EO) and Band Offset (BO), where EO uses edge attributes to classify pixels in SAO types 1 through 4, and BO uses pixel intensity to classify pixels in SAO type 5. Each applicable CTB has SAO parameters including sao_merge_left_flag, sao_merge_up_flag, SAO type, and four offsets. If sao_merge_left_flag equals 1, the current CTB will reuse the SAO type and offset of the left CTB. If sao_merge_up_flag equals 1, the current CTB will reuse the SAO type and offset of the upper CTB.
[0201] Table 4. Specifications for SAO Types
[0202] SAO type The type of sample adaptive compensation to be used Number of categories 0 none 0 1 One-dimensional 0-degree pattern edge offset 4 2 One-dimensional 90-degree pattern edge offset 4 3 One-dimensional 135-degree pattern edge offset 4 4 One-dimensional 45-degree pattern edge offset 4 5 With offset 4
[0203] 3.8 Adaptive Loop Filter
[0204] Adaptive Loop Filtering (ALF) for video encoding and decoding minimizes the mean square error between the original and decoded samples using Wiener-based adaptive filters. ALF is located at the last processing stage of each picture and can be considered a tool for capturing and repairing artifacts from previous stages. Appropriate filter coefficients are determined by the encoder and explicitly transmitted to the decoder via the signal. To achieve better encoding and decoding efficiency, especially for high-resolution video, local adaptation is used for the luminance signal by applying different filters to different regions or blocks in the picture. In addition to filter adaptation, filter on / off control at the codec tree unit (CTU) level also contributes to improving encoding and decoding efficiency. Syntactically, filter coefficients are sent in a picture-level header called the adaptive parameter set, and the filter on / off flags of the CTUs are interleaved at the CTU level in the stripe data. This syntax design not only supports picture-level optimization but also achieves low encoding latency.
[0205] 3.8.1 Signaling of Parameters
[0206] According to the ALF design in VTM, filter coefficients and clipping indices are carried in the ALF Adaptive Parameter Set (APS). An ALF APS can include up to eight chroma filters and a luma filter set with up to 25 filters. An index is also included for each of the 25 luma categories. Categories with the same index share the same filters. By merging different categories, the number of bits required to represent the filter coefficients is reduced. The absolute values of the filter coefficients are represented using zero-order exponent Golomb code, followed by a sign bit for non-zero coefficients. When clipping is enabled, a two-bit fixed-length code is also used to transmit the clipping index via signal transmission for each filter coefficient. Up to eight ALF APSs can be used simultaneously by the decoder.
[0207] The filter control syntax elements of ALF in VTM include two types of information. First, the ALF on / off flag is transmitted via signaling at the sequence, picture, strip, and CTB levels. Chroma ALF can only be enabled at the picture and strip levels if the luma ALF is enabled at the corresponding level. Second, if the ALF is enabled at that level, filter usage information is transmitted via signaling at the picture, strip, and CTB levels. If all strips within a picture use the same APS, the referenced ALF APSID is encoded / decoded at the strip or picture level. The luma component can reference up to 7 ALF APSs, and the chroma component can reference 1 ALF APS. For the luma CTB, an index is transmitted via signaling to indicate which ALF APS or offline-trained luma filter set is used. For the chroma CTB, an index indicates which filter in the referenced APS is used.
[0208] The data syntax elements of the ALF associated with the luminance component in VTM are listed below:
[0209] alf_data() { descriptor alf_luma_filter_signal_flag u(1) if( alf_luma_filter_signal_flag ) { alf_luma_clip_flag u(1) alf_luma_num_filters_signalled_minus1 ue(v) if( alf_luma_num_filters_signalled_minus1 > 0 ) for( filtIdx = 0; filtIdx < NumAlfFilters; filtIdx++ ) alf_luma_coeff_delta_idx[filtIdx] u(v) for( sfldx = 0; sfldx <= alf_luma_num_filters_signalled_minus1; sfldx++ ) for( j = 0; j < 12; j++ ) { alf_luma_coeff_abs[ sfIdx ][ j ] ue(v) if( alf_luma_coeff_abs[ sfIdx ][ j ] ) alf_luma_coeff_sign[ sfIdx ][ j ] u(1) } if(alf_luma_clip_flag) for( sfldx = 0; sfldx <= alf_luma_num_filters_signalled_minus1; sfldx++ ) for( j = 0; j < 12; j++ ) alf_luma_clip_idx[ sfIdx ][ j ] u(2) }
[0210] `alf_luma_filter_signal_flag` equal to 1 specifies the set of luminance filters transmitted via signaling. `alf_luma_filter_signal_flag` equal to 0 specifies that the set of luminance filters is not transmitted via signaling. `alf_luma_clip_flag` equal to 0 specifies that linear adaptive loop filtering is applied to the luminance component. `alf_luma_clip_flag` equal to 1 specifies that nonlinear adaptive loop filtering can be applied to the luminance component. `alf_luma_num_filters_signalled_minus1` plus 1 specifies the number of adaptive loop filter classes that the luminance coefficients can be transmitted via signaling. The value of `alf_luma_num_filters_signalled_minus1` must be in the range of 0 to `NumAlfFilters - 1` (inclusive). `alf_luma_coeff_delta_idx[filtIdx]` specifies the index of the adaptive loop filter luminance coefficient increment for the filter class indicated by `filtIdx`, ranging from 0 to `NumAlfFilters - 1`. `alf_luma_coeff_delta_idx[filtIdx]` is presumed to be 0 if it does not exist. The length of `alf_luma_coeff_delta_idx[filtIdx]` is Ceil(Log2(alf_luma_num_filters_signalled_minus1 + 1)) bits. The value of `alf_luma_coeff_delta_idx[filtIdx]` must be in the range from 0 to `alf_luma_num_filters_signalled_minus1` (inclusive).
[0211] `alf_luma_coeff_abs[ sfIdx ][ j ]` specifies the absolute value of the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx`. When `alf_luma_coeff_abs[ sfIdx ][ j ]` does not exist, it is presumed to be equal to 0. The value of `alf_luma_coeff_abs[ sfIdx ][ j ]` must be in the range of 0 to 128 (inclusive). `alf_luma_coeff_sign[ sfIdx ][ j ]` specifies the sign of the j-th luminance coefficient of the filter indicated by `sfIdx`, as follows:
[0212] If alf_luma_coeff_sign[ sfIdx ][ j ] equals 0, then the corresponding luminance filter coefficient has a positive value.
[0213] Otherwise (alf_luma_coeff_sign[ sfIdx ][ j ] equals 1), the corresponding luminance filter coefficient has a negative value.
[0214] When alf_luma_coeff_sign[ sfIdx ][ j ] does not exist, it is presumed to be equal to 0.
[0215] `alf_luma_clip_idx[sfIdx][j]` specifies the limiting index to be used before multiplying by the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx`. When `alf_luma_clip_idx[sfIdx][j]` does not exist, it is presumed to be equal to 0. The codec tree unit syntax elements of the ALF associated with the luminance component in the VTM are listed below:
[0216] coding_tree_unit() { descriptor xCtb = CtbAddrX « CtbLog2SizeY yCtb = CtbAddrY « CtbLog2SizeY if ( sh_alf_enabled_flag ){ alf_ctb_flag[ 0 ][ CtbAddrX ][ CtbAddrY ] ae(v) if( alf_ctb_flag[ 0 ][ CtbAddrX ][ CtbAddrY ] ) { if( sh_num_alf_aps_ids_luma > 0 ) alf_use_aps_flag ae(v) if(alf_use_aps_flag) { if(sh_num_alf_aps_ids_luma > 1) alf_luma_prev_filter_idx ae(v) } else alf_luma_fixed_filter_idx ae(v) } }
[0217] `alf_ctb_flag[ cIdx ][ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]` equal to 1 specifies that the adaptive loop filter is applied to the codec tree block of the codec tree unit at the luma position (xCtb, yCtb) for the color component indicated by `cIdx`. `alf_ctb_flag[ cIdx ][ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]` equal to 0 specifies that the adaptive loop filter is not applied to the codec tree block of the codec tree unit at the luma position (xCtb, yCtb) for the color component indicated by `cIdx`.
[0218] When `alf_ctb_flag[ cIdx ][ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY]` does not exist, it is presumed to be equal to 0. `alf_use_aps_flag` equal to 0 specifies that a fixed set of filters from the fixed filter set is applied to the luma CTB. `alf_use_aps_flag` equal to 1 specifies that a set of filters from the APS is applied to the luma CTB. When `alf_use_aps_flag` does not exist, it is presumed to be equal to 0. `alf_luma_prev_filter_idx` specifies the previous filter applied to the luma CTB. The value of `alf_luma_prev_filter_idx` must be in the range of 0 to `sh_num_alf_aps_ids_luma - 1` (inclusive). When `alf_luma_prev_filter_idx` does not exist, it is presumed to be equal to 0.
[0219] The variable AlfCtbFiltSetIdxY[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ], which specifies the filter set index of the luminance CTB at position (xCtb, yCtb), is derived as follows:
[0220] If alf_use_aps_flag equals 0, then AlfCtbFiltSetIdxY[ xCtb >> CtbLog2SizeY][ yCtb >> CtbLog2SizeY ] is set to equal alf_luma_fixed_filter_idx.
[0221] Otherwise, AlfCtbFiltSetIdxY[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY] is set to equal to 16 + alf_luma_prev_filter_idx.
[0222] alf_luma_fixed_filter_idx specifies the fixed filter applied to the lumen CTB. The value of alf_luma_fixed_filter_idx must be in the range of 0 to 15 (inclusive).
[0223] Based on the VTM-based ALF design, the ECM-based ALF design further introduces the concept of alternative filter sets into the luminance filter. The luminance filter is trained for multiple alternatives / rounds based on the updated luminance CTU ALF on / off decisions for each alternative / round. In this way, multiple filter sets will be associated with each trained alternative, and the class merging results for each filter set can be different. Each CTU can select the optimal filter set via RDO, and the relevant alternative information will be transmitted via signal transmission. The data syntax elements of the ALF associated with the luminance component in the ECM are listed below:
[0224] alf_data() { Descriptor alf_luma_filter_signal_flag u(1) if(alf_luma_filter_signal_flag) { alf_luma_num_alts_minus1 ue(v) for(altIdx = 0; altIdx < alf_luma_num_alts_minus1 +1; altIdx++){ alf_luma_clip_flag[altIdx] u(1) alf_luma_num_filters_signalled_minus1[altIdx] ue(v) if(alf_luma_num_filters_signalled_minus1[altIdx] > 0){ for( filtIdx = 0; filtIdx < NumAlfFilters; filtIdx++ ) alf_luma_coeff_delta_idx[altIdx][filtIdx] u(v) } for (sfldx = 0; sfldx <= alf_luma_num_filters_signalled_minus1[altldx]; sfldx++) { for(j = 0; j < 19; j++){ alf_luma_coeff_abs[altIdx][sfIdx][j] ue(v) if(alf_luma_coeff_abs[altIdx][sfIdx][j]) alf_luma_coeff_sign[altIdx][sfIdx][j] u(1) } } if(alf_luma_clip_flag[altIdx]) for( sfldx = 0; sfldx <= alf_luma_num_filters_signalled_minus1[altldx]; sfldx++ ) for( j = 0; j <19; j++ ) alf_luma_clip_idx[altIdx][sfIdx][j] u(2) } }
[0225] The increment of `alf_luma_num_alts_minus1` by 1 specifies the number of alternative filter sets for the luminance component. The value of `alf_luma_num_alts_minus1` must be in the range of 0 to 3 (inclusive). `alf_luma_clip_flag[altIdx]` equal to 0 specifies that linear adaptive loop filtering is applied to the luminance component of the alternative luminance filter set at index `altIdx`. `alf_luma_clip_flag[altIdx]` equal to 1 specifies that nonlinear adaptive loop filtering can be applied to the luminance component of the alternative luminance filter set at index `altIdx`. The increment of `alf_luma_num_filters_signalled_minus1[altIdx]` by 1 specifies the number of adaptive loop filter categories that can transmit the luminance coefficients via signal transmission for the alternative luminance filter set at index `altIdx`. The value of alf_luma_num_filters_signalled_minus1[altIdx] must be in the range of 0 to NumAlfFilters - 1 (inclusive).
[0226] `alf_luma_coeff_delta_idx[altIdx][filtIdx]` specifies the index of the adaptive loop filter luminance coefficient increments through signal transmission for the alternative luminance filter set with index `altIdx`, indicated by `filtIdx` ranging from 0 to `NumAlfFilters - 1`. `alf_luma_coeff_delta_idx[filtIdx][altIdx]` is presumed to be 0 if it does not exist. The length of `alf_luma_coeff_delta_idx[altIdx][filtIdx]` is Ceil(Log2(alf_luma_num_filters_signalled_minus1[altIdx] + 1)) bits. The value of `alf_luma_coeff_delta_idx[altIdx][filtIdx]` must be within the range of 0 to `alf_luma_num_filters_signalled_minus1[altIdx]` (inclusive). `alf_luma_coeff_abs[altIdx][sfIdx][j]` specifies the absolute value of the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx` of the alternative luminance filter set indexed at `altIdx`. When `alf_luma_coeff_abs[altIdx][sfIdx][j]` does not exist, it is presumed to be equal to 0. The value of alf_luma_coeff_abs[altIdx][sfIdx][j] must be in the range of 0 to 128 (inclusive).
[0227] alf_luma_coeff_sign[altIdx][sfIdx][j] specifies the sign of the j-th luminance coefficient of the filter indicated by sfIdx of the alternative luminance filter set with index altIdx, as follows:
[0228] If alf_luma_coeff_sign[altIdx][sfIdx][j] equals 0, then the corresponding luminance filter coefficient has a positive value.
[0229] Otherwise (alf_luma_coeff_sign[altIdx][sfIdx][j] equals 1), the corresponding luminance filter coefficient has a negative value.
[0230] When alf_luma_coeff_sign[altIdx][sfIdx][j] does not exist, it is presumed to be equal to 0.
[0231] `alf_luma_clip_idx[altIdx][sfIdx][j]` specifies the limiting index to be used before multiplying the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx` of the alternative luminance filter set indexed at `altIdx`. When `alf_luma_clip_idx[altIdx][sfIdx][j]` does not exist, it is presumed to be equal to 0. The codec tree unit syntax elements of the ALF associated with the luminance component in the ECM are listed below:
[0232] coding_tree_unit() { Descriptor xCtb = CtbAddrX « CtbLog2SizeY yCtb = CtbAddrY « CtbLog2SizeY if(sh_alf_enabled_flag){ alf_ctb_flag[0][CtbAddrX][CtbAddrY] ae(v) if(alf_ctb_flag[0][CtbAddrX][CtbAddrY]) { if(sh_num_alf_aps_ids_luma > 0) alf_use_aps_flag ae(v) if(alf_use_aps_flag) { if(sh_num_alf_aps_ids_luma > 1) alt_ctb_luma_filter_alt_idx[CtbAddrX][CtbAddrY] ae(v) alf_luma_prev_filter_idx ae(v) } else alf_luma_fixed_filter_idx ae(v) } }
[0233] `alf_ctb_luma_filter_alt_idx[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]` specifies the index of the alternative luma filter for the codec tree block of the luma component applied to the luma location (xCtb, yCtb). If `alf_ctb_luma_filter_alt_idx[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]` does not exist, it is presumed to be equal to 0.
[0234] 3.8.2 Filter Shape
[0235] In JEM, there are a maximum of three diamond filter shapes (e.g. Figure 10 The square (shown) can be selected for the luma component. Indices are transmitted via signaling at the image level to indicate the filter shape used for the luma component. Each square represents a sample, and Ci (i = 0~6 (left), 0~12 (middle), 0~20 (right)) represents the coefficient to be applied to the sample. For the chroma component in the image, a 5×5 rhombus shape is always used. In VVC, a 7×7 rhombus shape is always used for luma, while a 5×5 rhombus shape is always used for chroma.
[0236] 3.8.3 Classification of ALF
[0237] Each 2×2 (or 4×4) block is classified into one of 25 categories. The classification index C is based on its directionality. and activity level The quantization value is derived as follows:
[0238] .
[0239] In order to calculate and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using the one-dimensional Laplacian operator:
[0240] ,
[0241] ,
[0242] ,
[0243] .
[0244] index and Referencing the coordinates of the top left sample point in the 2×2 block, and Indicator coordinates The reconstructed sample points are then used. Then, the gradients in the horizontal and vertical directions are... The maximum and minimum values are set as follows:
[0245] , ;
[0246] Furthermore, the maximum and minimum values of the gradients in the two diagonal directions are set as follows:
[0247] , .
[0248] In order to derive directionality The values are compared with each other and with two thresholds. and Compare:
[0249] Step 1. If and If both are true, then Set as ;
[0250] Step 2. If If yes, continue from step 3; otherwise, continue from step 4.
[0251] Step 3. If ,but Set as ,otherwise Set as ;
[0252] Step 4. If ,but Set as ,otherwise Set as .
[0253] Activity value Calculated as:
[0254] .
[0255] It is further quantized to the range of 0 to 4 (inclusive), and the quantized value is represented as For the two chromaticity components in the image, no classification method is applied; that is, a single set of ALF coefficients is applied for each chromaticity component.
[0256] 3.8.4 Geometric Transformation of Filter Coefficients
[0257] Before filtering each 2×2 block, geometric transformations (such as rotation or diagonal and vertical flips) are applied to the filter coefficients associated with coordinates (k, l). This depends on the gradient values computed for that block. This is equivalent to applying these transformations to samples in the filter's support region. The idea is to make the different blocks more similar by aligning the directions of the different blocks to which the ALF is applied.
[0258] Three geometric transformations are introduced, including diagonal flip, vertical flip, and rotation:
[0259] diagonal: ,
[0260] Vertical Flip: ,
[0261] Rotation: ;
[0262] in It is the size of the filter, and These are coefficient coordinates, which make the position... In the top left corner, and in position In the bottom right corner, the transform is applied to the filter coefficients f(k, l), depending on the gradient values calculated for that block. The relationship between the transform and the four gradients in the four directions is summarized in Table 5. Figure 11 The transformation coefficients for each position based on a 5×5 rhombus are shown.
[0263] Table 5. Mapping of gradients and transformations computed for a block
[0264] Gradient value Transformation g d2 g d1 and g h g v ]]> No transformation g d2 g d1 and g v g h ]]> Diagonal g d1 g d2 and g h g v ]]> Vertical flip <![CDATA[g d1 <g d2 And g v <g h ]]> Rotation
[0265] 3.8.5 Filtering Process
[0266] On the decoder side, when ALF is enabled for a block, each sample within the block... Filtering causes sample values As shown below, where L represents the filter length, Represents the filter coefficients. This represents the decoded filter coefficients.
[0267]
[0268] Figure 12 This example illustrates the relative coordinates used in support of a 5×5 diamond filter, assuming the current sample coordinates (i, j) are (0, 0). Samples at different coordinates filled with the same color are multiplied by the same filter coefficients.
[0269] 3.8.6 Nonlinear Filtering Reconstruction
[0270] Linear filtering can be reconstructed into the following expression without affecting encoding / decoding efficiency:
[0271]
[0272] in They are the same filter coefficients.
[0273] VVC introduces nonlinearity by using a simple limiting function to measure the value at neighboring sample points ( Compared with the current sample value being filtered ( When the difference is too large, the influence of neighboring sample values is reduced, thus making ALF more efficient. More specifically, the ALF filter is modified as follows:
[0274]
[0275] in It is a limiting function. It is the limiting parameter, which depends on Filter coefficients. The encoder performs optimization to find the optimal values. .
[0276] A limiting parameter was specified for each ALF filter. Each filter coefficient transmits a limiting value via signal transmission. This means that up to 12 limiting values for each luminance filter and up to 6 limiting values for each chrominance filter can be transmitted via signal transmission in the bitstream. To limit signaling costs and encoder complexity, only 4 fixed values are used, which are the same for both inter-frame and intra-frame stripes.
[0277] Because the variance of local differences in luminance is typically higher than that in chrominance, two different sets of filters are applied for luminance and chrominance. The maximum sample value in each set (here, for a 10-bit bit depth of 1024) is also introduced so that clipping can be disabled if unnecessary. The four values are selected by roughly equally dividing the entire range of luminance sample values (encoded and decoded on 10 bits) and the chrominance range from 4 to 1024 in the logarithmic domain. More precisely, the luminance table of clipping values has been obtained using the following formula:
[0278] AlfClip L Where M=2 10 And N=4
[0279] Similarly, the colorimetric table for the limiting values is obtained according to the following formula:
[0280] AlfClip C Where M=2 10 N=4 and A=4
[0281] 3.9 Bilateral Loop Filter
[0282] 3.9.1 Bilateral Image Filter
[0283] Bilateral image filters are nonlinear filters that smooth noise while preserving edge structure. Bilateral filtering is a technique that makes the filter weights decrease not only with the distance between samples but also with increasing intensity difference. This improves the smoothing of overly smoothed edges. The weights are defined as follows:
[0284]
[0285] in and These are the distances in the horizontal and vertical directions, respectively. It is the intensity difference between sample points.
[0286] The edge-preserving denoising bilateral filter employs low-pass Gaussian filters for both the domain and range filters. The domain low-pass Gaussian filter assigns higher weights to pixels spatially close to the center pixel. The range low-pass Gaussian filter assigns higher weights to pixels similar to the center pixel. Combining the range and domain filters, the bilateral filter at edge pixels becomes a thin Gaussian filter, oriented along the edge and significantly reduced in the gradient direction. This is why the bilateral filter can smooth noise while preserving edge structure.
[0287] 3.9.2 Bilateral Filters in Video Encoding and Decoding
[0288] Bilateral filters in video encoding and decoding are encoding and decoding tools for VVC [2]. The filter acts as a loop filter in parallel with the Sample Adaptive Compensation (SAO) filter. Both the bilateral filter and the SAO operate on the same input samples, each filter produces an offset, and these offsets are then added to the input samples to produce output samples, which are then clipped before proceeding to the next stage. Spatial filtering intensity Determined by the block size, smaller blocks are filtered more strongly, and the intensity of the filter is... Determined by the quantization parameters, with stronger filtering used for higher QP values. Only the four nearest samples are used, therefore the filtered sample intensity... It can be calculated as
[0289]
[0290] in Indicates the intensity of the central sample point. This indicates the intensity difference between the center sample point and the sample point above it. These represent the intensity differences between the central sample point and the sample points below, to the left, and to the right, respectively.
[0291] 4. Progressive decoding refresh in video encoding and decoding
[0292] The GDR (Generic Refresh) method alleviates the latency issue of intra-frame encoded images. Instead of encoding and decoding intra-frame images at random access points, GDR progressively refreshes the images by spreading the intra-frame encoded regions across multiple images. The GDR concept was introduced as the Recovery Point SEI message in several earlier video codec standards (e.g., AVC, HEVC). GDR-related syntax and GDR units are included in the VVC specification. For example, VVC more effectively supports GDR functionality by using virtual boundaries to allow for finer-grained progressive intra-frame refresh.
[0293] Typically, a GDR image consists of clean (or refreshed) regions and dirty (or non-refreshed) regions. The clean regions may include a forced intra-frame region adjacent to the dirty regions for Progressive Intra-Frame Refresh (PIR). The boundary between the clean and dirty regions is transmitted via signaling through the virtual boundary syntax in the image header.
[0294] 5. Technical problems solved by the disclosed embodiments
[0295] The example design of GDR-based temporal information reference in video encoding and decoding has the following problems:
[0296] First, in the example Adaptive Parameter Set (APS) design, clean regions may reference APS generated from dirty regions, leading to leaks or mismatches.
[0297] Second, in the example of motion-compensated infill design, clean areas may reference sample points / motion information in dirty areas, leading to leakage or mismatch.
[0298] Third, in the example time-domain information reference design, clean regions may reference or inherit time-domain information generated from dirty regions.
[0299] 6. List of solutions and implementation examples
[0300] To address the aforementioned problems, methods outlined below are disclosed. These embodiments should be considered as examples for explaining general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments can be applied individually or in any combination.
[0301] In this disclosure, a video unit can refer to a sequence, picture, subpicture, strip, CTU, block, and / or region. A video unit may include one color component or multiple color components.
[0302] In this disclosure, clean regions and dirty regions can refer to regions encoded and decoded by GDR and regions not encoded and decoded by GDR, respectively. A GDR image can refer to an encoded image having one or more GDR regions.
[0303] In the following description, temporal information refers to any information derived / inherited / borrowed from video units (e.g., pictures) that are different from the current video unit.
[0304] 1) An APS reference implementation constraint is proposed in GDR-based video coding and decoding.
[0305] a. In one example, the ALF-APS reference can be restricted in GDR-based video codecs.
[0306] a) In one example, one or more ALF-APS cannot be referenced by GDR images / clean areas / dirty areas.
[0307] b) In one example, when encoding a GDR image / clean area / dirty area, one or more ALF-APS can be reset / initialized.
[0308] c) In one example, when decoding a GDR image / clean area / dirty area, one or more ALF-APS can be reset / initialized.
[0309] d) In one example, one or more ALF-APS generated from a GDR image / clean area / dirty area can be transmitted / derived / predefined via signal transmission.
[0310] e) In one example, one or more ALF-APS generated from a GDR image / clean area / dirty area can be stored at the encoder / decoder.
[0311] f) In one example, one or more ALF-APS generated from a GDR image / clean area / dirty area can be referenced by a non-GDR image.
[0312] g) In one example, ALF-APS may include different parameters for ALF-luminance / ALF-chrominance / cross-component ALF (CCALF).
[0313] 1. In one example, a syntax element may be included to represent a new parameter used for transmitting ALF-luminance via a signal.
[0314] 2. In one example, a syntax element may be included to represent a new parameter used for transmitting ALF-chroma over a signal.
[0315] 3. In one example, a syntax element may be included to represent a new parameter for transmitting CCALF-Cb / Cr via signaling.
[0316] 4. In one example, the number of substitutions in ALF-luminance can be included.
[0317] 5. In one example, the filter shape / type of ALF-luminance may be included.
[0318] 6. In one example, for each alternative to ALF-brightness, a classifier index may be included.
[0319] 7. In one example, for each alternative to ALF-luminance, a syntax element representing the use of non-linear functions may be included.
[0320] 8. In one example, for each alternative to ALF-luminance, the number of filter / combining categories transmitted through the signal can be included.
[0321] 9. In one example, for each alternative to ALF-luminance, the category merging result can be included.
[0322] 10. In one example, each alternative to ALF-luminance may include filter coefficients with a symmetrical design.
[0323] 11. In one example, each alternative to ALF-luminance may include filter coefficients that do not have a symmetrical design.
[0324] 12. In one example, for each alternative to ALF-luminance, a filter nonlinear limiting parameter with a symmetrical design may be included.
[0325] 13. In one example, for each alternative to ALF-luminance, nonlinear limiting parameters of the filter without a symmetrical design may be included.
[0326] 14. In one example, the filter shape / type of ALF-chroma may be included.
[0327] 15. In one example, the number of substitutions in ALF-chroma may be included.
[0328] 16. In one example, for each alternative to ALF-chroma, a syntax element representing enabling non-linear functions may be included.
[0329] 17. In one example, each alternative to ALF-chromaticity may include filter coefficients with a symmetrical design.
[0330] 18. In one example, each alternative to ALF-chromaticity may include filter coefficients that do not have a symmetrical design.
[0331] 19. In one example, for each alternative to ALF-chromaticity, filter nonlinear parameters with a symmetrical design may be included.
[0332] 20. In one example, each alternative to ALF-chromaticity may include filter nonlinear parameters that do not have a symmetrical design.
[0333] 21. In one example, the filter shape / type of CCALF can be included.
[0334] 22. In one example, the number of filters for Cb / Cr passing through the signal for CCALF may be included.
[0335] 23. In one example, each filter for Cb / Cr of CCALF that passes through the signal may include the absolute value of the filter coefficients.
[0336] 24. In one example, for each filter that transmits Cb / Cr through the signal for CCALF, the sign of the filter coefficients may be included.
[0337] 25. In one example, any other parameters related to ALF / CCALF may be included.
[0338] b. In one example, the luminance mapping (LMCS)-APS reference with chroma scaling can be restricted in GDR-based video codecs.
[0339] a) In one example, one or more LMCS-APS cannot be referenced by GDR image / clean area / dirty area.
[0340] b) In one example, when encoding a GDR image / clean area / dirty area, one or more LMCS-APS can be reset / initialized.
[0341] c) In one example, when decoding a GDR image / clean area / dirty area, one or more LMCS-APS can be reset / initialized.
[0342] d) In one example, one or more LMCS-APS generated from a GDR image / clean area / dirty area can be transmitted / derived / predefined via signal transmission.
[0343] e) In one example, one or more LMCS-APS generated from a GDR image / clean area / dirty area can be stored at the encoder / decoder.
[0344] f) In one example, one or more LMCS-APS generated from a GDR image / clean area / dirty area can be referenced by a non-GDR image.
[0345] g) In one example, LMCS-APS may include different parameters for luminance mapping / chrominance scaling.
[0346] 1. In one example, the index of the first valid codeword segment of the luminance map may be included.
[0347] 2. In one example, the index of the last valid codeword segment of the luminance map may be included.
[0348] 3. In one example, the maximum offset between the unmapped codeword and the mapped codeword of the luminance mapping can be included.
[0349] 4. In one example, for each valid codeword segment, the absolute value of the offset between the unmapped codeword and the mapped codeword may be included.
[0350] 5. In one example, for each valid codeword segment, the symbol may include the offset between the unmapped codeword and the mapped codeword.
[0351] 6. In one example, the absolute value of the chroma scaling offset may be included.
[0352] 7. In one example, a sign for the chroma scaling offset may be included.
[0353] 8. In one example, any other parameters related to brightness mapping may be included.
[0354] 9. In one example, any other parameters related to chroma scaling may be included.
[0355] c. In one example, the SAO / Cross-Component SAO (CCSAO) / BF-APS reference can be restricted in GDR-based video coding and decoding.
[0356] a) In one example, one or more SAO / CCSAO / BF-APS cannot be referenced by GDR images / clean areas / dirty areas.
[0357] b) In one example, when encoding a GDR image / clean area / dirty area, one or more SAO / CCSAO / BF-APS can be reset / initialized.
[0358] c) In one example, when decoding a GDR image / clean area / dirty area, one or more SAO / CCSAO / BF-APS can be reset / initialized.
[0359] d) In one example, one or more SAO / CCSAO / BF-APS generated from a GDR image / clean area / dirty area can be transmitted / derived / predefined via signal transmission.
[0360] e) In one example, one or more SAO / CCSAO / BF-APS generated from a GDR image / clean area / dirty area can be stored at the encoder / decoder.
[0361] f) In one example, one or more SAO / CCSAO / BF-APS generated from a GDR image / clean area / dirty area can be referenced by a non-GDR image.
[0362] g) In one example, SAO / CCSAO / BF-APS may include different parameters for SAO / CCSAO / BF.
[0363] 1. In one example, the classifier index of SAO may be included.
[0364] 2. In one example, the number of offsets of the SAO transmitted via signal may be included.
[0365] 3. In one example, the absolute value of the offset of the SAO may be included.
[0366] 4. In one example, the sign of the SAO offset may be included.
[0367] 5. In one example, a valid indexed SAO may be included.
[0368] 6. In one example, the classifier index of CCSAO may be included.
[0369] 7. In one example, the co-position of the CCSAO brightness may be included.
[0370] 8. In one example, the number of offsets transmitted via signaling for CCSAO may be included.
[0371] 9. In one example, the absolute value of the offset of CCSAO may be included.
[0372] 10. In one example, the sign of the offset of CCSAO may be included.
[0373] 11. In one example, a valid indexed CCSAO may be included.
[0374] 12. In one example, a valid edge pattern index of CCSAO may be included.
[0375] 13. In one example, the classifier index of BF may be included.
[0376] 14. In one example, the filter strength index of BF may be included.
[0377] 15. In one example, the coefficient of BF may be included.
[0378] 16. In one example, the nonlinear limiting parameter of BF may be included.
[0379] 17. In one example, any other parameters related to SAO may be included.
[0380] 18. In one example, any other parameters related to CCSAO may be included.
[0381] 19. In one example, any other parameters related to BF may be included.
[0382] d. In one example, in GDR-based video encoding and decoding, temporal information generated from a deblocking filter (DBF) / BF / SAO / CCSAO or any other tool can be restricted.
[0383] a) In one example, for GDR images / clean areas / dirty areas, the temporal information generated from DBF / BF / SAO / CCSAO cannot be referenced.
[0384] b) In one example, when encoding GDR images / clean areas / dirty areas, the temporal information generated from DBF / BF / SAO / CCSAO can be reset / initialized.
[0385] c) In one example, when decoding a GDR image / clean region / dirty region, the temporal information generated from DBF / BF / SAO / CCSAO can be reset / initialized.
[0386] d) In one example, temporal information generated from DBF / BF / SAO / CCSAO from GDR images / clean areas / dirty areas can be transmitted / derived / predefined via signal transmission.
[0387] e) In one example, the temporal information generated from DBF / BF / SAO / CCSAO from the GDR image / clean area / dirty area can be stored at the encoder / decoder.
[0388] f) In one example, temporal information generated from DBF / BF / SAO / CCSAO from GDR images / clean regions / dirty regions can be referenced by non-GDR images.
[0389] g) In one example, the time-domain information generated from DBF / BF / SAO / CCSAO may include different parameters.
[0390] 1. In one example, the classifier index of SAO may be included.
[0391] 2. In one example, the number of offsets of the SAO transmitted via signal may be included.
[0392] 3. In one example, the absolute value of the offset of the SAO may be included.
[0393] 4. In one example, the sign of the SAO offset may be included.
[0394] 5. In one example, a valid indexed SAO may be included.
[0395] 6. In one example, the classifier index of CCSAO may be included.
[0396] 7. In one example, the co-position of the CCSAO brightness may be included.
[0397] 8. In one example, the number of offsets transmitted via signaling for CCSAO may be included.
[0398] 9. In one example, the absolute value of the offset of CCSAO may be included.
[0399] 10. In one example, the sign of the offset of CCSAO may be included.
[0400] 11. In one example, a valid indexed CCSAO may be included.
[0401] 12. In one example, a valid edge pattern index of CCSAO may be included.
[0402] 13. In one example, the classifier index of BF may be included.
[0403] 14. In one example, the filter strength index of BF may be included.
[0404] 15. In one example, the coefficient of BF may be included.
[0405] 16. In one example, the nonlinear limiting parameter of BF may be included.
[0406] 17. In one example, any other parameters related to SAO may be included.
[0407] 18. In one example, any other parameters related to CCSAO may be included.
[0408] 19. In one example, any other parameters related to BF may be included.
[0409] 20. In one example, the boundary strength of the DBF may be included.
[0410] 21. In one example, the filter strength of the DBF may be included.
[0411] 22. In one example, the filter length of the DBF may be included.
[0412] 23. In one example, any other parameters related to DBF may be included.
[0413] 2) A method for imposing restrictions on motion information in GDR-based video encoding and decoding is proposed.
[0414] a. In one example, motion-compensated padding can be limited in GDR-based video codecs.
[0415] a) In one example, motion-compensated fill can be forcibly disabled in GDR images / clean areas / dirty areas.
[0416] b) In one example, motion-compensated padding can be reset in the GDR image / clean area / dirty area.
[0417] c) In one example, motion-compensated fill can be applied to a GDR image / clean area / dirty area using samples from the current image / region.
[0418] d) In one example, motion-compensated padding can be applied to a non-GDR image by using samples from GDR images / clean regions / dirty regions.
[0419] b. In one example, in GDR-based video encoding and decoding, motion information generated from intra-block copy (IBC) or any other tool can be restricted.
[0420] a) In one example, IBC motion information across non-GDR images and across GDR images cannot be referenced by GDR image / clean region / dirty region.
[0421] b) In one example, IBC motion information across clean and dirty regions cannot be referenced by GDR image / clean region / dirty region.
[0422] c) In one example, IBC motion information generated from GDR images / clean areas / dirty areas can be stored.
[0423] d) In one example, IBC motion information generated from a GDR image / clean region / dirty region can be referenced by a non-GDR image.
[0424] e) In one example, IBC motion information can be referenced from local motion information.
[0425] f) In one example, IBC motion information can be referenced from time-domain motion information.
[0426] 3) Restrictions on parameter reference / inheritance are proposed in GDR-based video encoding and decoding.
[0427] a. In one example, temporal parameter reference / inheritance can be restricted in GDR-based video codecs.
[0428] a) In one example, temporal context adaptive binary arithmetic codec (CABAC) parameters cannot be referenced by GDR image / clean region / dirty region.
[0429] b) In one example, the temporal CABAC parameters cannot be referenced for GDR images / clean areas / dirty areas.
[0430] c) In one example, when encoding a GDR image / clean region / dirty region, the temporal CABAC parameters can be reset / initialized.
[0431] d) In one example, when decoding a GDR image / clean region / dirty region, the temporal CABAC parameters can be reset / initialized.
[0432] e) In one example, the temporal CABAC parameters generated from the GDR image / clean area / dirty area can be transmitted / derived / predefined via signal transmission.
[0433] f) In one example, the temporal CABAC parameters generated from the GDR image / clean region / dirty region can be stored at the encoder / decoder.
[0434] g) In one example, temporal CABAC parameters generated from a GDR image / clean region / dirty region can be referenced by a non-GDR image.
[0435] h) In one example, temporal parameters cannot be referenced by GDR images / clean areas / dirty areas.
[0436] i) In one example, temporal parameters cannot be referenced for GDR images / clean areas / dirty areas.
[0437] j) In one example, when encoding a GDR image / clean area / dirty area, the temporal parameters can be reset / initialized.
[0438] k) In one example, the temporal parameters can be reset / initialized when decoding a GDR image / clean area / dirty area.
[0439] l) In one example, the temporal parameters generated from the GDR image / clean area / dirty area can be transmitted / derived / predefined via signal transmission.
[0440] m) In one example, the temporal parameters generated from the GDR image / clean area / dirty area can be stored at the encoder / decoder.
[0441] n) In one example, temporal parameters generated from a GDR image / clean region / dirty region can be referenced by a non-GDR image.
[0442] o) In one example, temporal parameters can refer to information generated from different inter-frame / intra-frame codecs.
[0443] 1. In one example, the temporal parameters can be generated from a convolutional cross-component model (CCCM).
[0444] 2. In one example, the time-domain parameters can be generated from a cross-component linear model (CCLM).
[0445] 3. In one example, the time-domain parameters can be generated from any other cross-component codec tool.
[0446] 4. In one example, the time-domain parameters can be generated from a history-based affine codec.
[0447] 5. In one example, temporal parameters can be generated from history-based Local Illumination Compensation (LIC) codecs.
[0448] 6. In one example, the temporal parameters can be generated from any history-based intra-frame codec tool.
[0449] 7. In one example, the temporal parameters can be generated from any history-based inter-frame codec tool.
[0450] 4) In one example, the disclosed methods may be used in post-processing and / or pre-processing.
[0451] 5) In one example, the above methods can be used in combination.
[0452] 6) Alternatively, the above methods can be used alone.
[0453] 7) In one example, the constraints proposed / described for GDR-based video coding and decoding methods can be applied to any loop filtering tool, prediction tool, preprocessing or postprocessing filtering method in video coding and decoding.
[0454] a. In one example, the constraints imposed on GDR-based video encoding and decoding methods can be applied to loop filtering methods.
[0455] b. In one example, the constraints proposed for GDR-based video coding and decoding methods can be applied to intra-frame prediction methods.
[0456] c. In one example, the constraints proposed for GDR-based video encoding and decoding methods can be applied to inter-frame prediction methods.
[0457] d. In one example, the constraints proposed for GDR-based video encoding and decoding methods can be applied to preprocessing filtering methods.
[0458] e. In one example, the constraints imposed on GDR-based video encoding and decoding methods can be applied to post-processing filtering methods.
[0459] 8) In the above examples, a video unit can refer to a sequence / picture / subpicture / strip / piece / code-decode tree unit (CTU) / CTU line / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / any other region containing more than one luminance or chrominance sample / pixel.
[0460] 9) Whether and / or how the methods disclosed above can be used to transmit signals in a bitstream.
[0461] a. In one example, they can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / decoder capability information (DCI) / PPS / APS / strip header / piece group header.
[0462] b. In one example, they can be transmitted via signal at PB / TB / CB / PU / TU / CU / Virtual Pipeline Data Unit (VPDU) / CTU / CTU line / strip / piece / sub-picture / other types of areas containing more than one sample or pixel.
[0463] 10) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / double tree segmentation, color components, and stripe / picture type.
[0464] 6. References
[0465] [1] J. Strom, P. Wennersten, J. Enhorn, D. Liu, K. Andersson and R.Sjoberg, “Bilateral Loop Filter in Combination with SAO,” in proceeding of IEEE Picture Coding Symposium (PCS), Nov. 2019.
[0466] Figure 13This is a block diagram illustrating an example video processing system 4000 in which various embodiments disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be received in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Network (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0467] System 4000 may include an encoding component 4004 capable of implementing the various encoding / decoding or encoding methods described in this disclosure. Encoding component 4004 may reduce the average bit rate from the video input 4002 to the output of encoding component 4004 to produce an encoded representation of the video. Encoding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of encoding component 4004 may be stored or transmitted via a communication connection such as that represented by component 4006. The bitstream (or encoded) representation of the video received at input 4002, whether stored or transmitted via communication, may be used by component 4008 to generate pixel values or displayable video that is sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding” operations or tools, it is understood that encoding tools or operations are used by encoders, and corresponding decoding tools or operations that reverse the encoded result will be performed by decoders.
[0468] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE), etc. The embodiments described in this disclosure can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0469] Figure 14This is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processors(multiple) 4102 can be configured to implement one or more methods described herein. The memories(multiple) 4104 can be used to store data and code for implementing the methods and embodiments described herein. The video processing circuitry 4106 can be used to implement some embodiments described herein in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, such as a graphics coprocessor.
[0470] Figure 15 This is a flowchart of an example method 4200 for video processing. Method 4200 includes: in step 4202, when performing progressive decode refresh (GDR) based video encoding / decoding, determining a reference to apply one or more constraints to an adaptive parameter set (APS); and in step 4204, performing a conversion between visual media data and a bitstream based on the APS. According to the example, the conversion in step 4204 may include encoding at the encoder or decoding at the decoder.
[0471] It should be noted that method 4200 can be implemented in a means of processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions cause the processor to execute method 4200 when executed by the processor. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium, which includes a computer program product for use by a video encoding / decoding device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that when the computer-executable instructions are executed by a processor, the video encoding / decoding device executes method 4200.
[0472] Figure 16 This is a block diagram illustrating an example video encoding / decoding system 4300 that can utilize embodiments of the present disclosure. The video encoding / decoding system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data, and this source device 4310 may be referred to as a video encoding device. The target device 4320 can decode the encoded video data generated by the source device 4310, and this target device 4320 may be referred to as a video decoding device.
[0473] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include encoded pictures and associated data. Encoded pictures are coded representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0474] Target device 4320 may include I / O interface 4326, video decoder 4324, and display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may acquire encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320 or may be external to target device 4320, wherein target device 4320 may be configured to interface with an external display device.
[0475] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as HEVC, VVC and other existing and / or further standards.
[0476] Figure 17 This is a block diagram illustrating an example of a video encoder 4400, which can be... Figure 16 The system 4300 shown includes a video encoder 4314. The video encoder 4400 can be configured to perform any or all embodiments of this disclosure. The video encoder 4400 includes multiple functional components. The embodiments described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all embodiments described in this disclosure.
[0477] The functional components of the video encoder 4400 may include a segmentation unit 4401; a prediction unit 4402, which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406; a residual generation unit 4407; a transform processing unit 4408; a quantization unit 4409; an inverse quantization unit 4410; an inverse transform unit 4411; a reconstruction unit 4412; a buffer 4413; and an entropy coding unit 4414.
[0478] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0479] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, these components are represented separately in the example of the video encoder 4400.
[0480] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0481] The mode selection unit 4403 can select one of several encoding / decoding modes (intra-frame encoding / decoding or inter-frame encoding / decoding), for example, based on error results, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the block based on motion vectors (e.g., sub-pixel precision or integer pixel precision).
[0482] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 4413 other than the image associated with the current video block.
[0483] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0484] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 4404 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0485] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images of list 0, and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 4404 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0486] In some examples, the motion estimation unit 4404 can output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0487] In one example, the motion estimation unit 4404 may indicate a value to the video decoder 4500 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0488] In another example, motion estimation unit 4404 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0489] As discussed above, the video encoder 4400 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0490] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples of other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0491] The residual generation unit 4407 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0492] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform subtraction operations.
[0493] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0494] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0495] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding sample points of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block and store it in the buffer 4413.
[0496] After the video block is reconstructed by reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0497] The entropy coding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, it can perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0498] Figure 18 This is a block diagram illustrating an example of a video decoder 4500. The video decoder 4500 can be... Figure 16 The system 4300 shown includes a video decoder 4324. The video decoder 4500 can be configured to perform any or all embodiments of this disclosure. In the example shown, the video decoder 4500 includes multiple functional components. The embodiments described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all embodiments described in this disclosure.
[0499] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 4400.
[0500] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by executing AMVP and Merge modes.
[0501] The motion compensation unit 4502 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.
[0502] The motion compensation unit 4502 can use the interpolation filter used by the video encoder 4400 during the encoding of a video block to calculate the interpolation for sub-integer pixels of the reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate a prediction block.
[0503] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.
[0504] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 4504 performs inverse quantization (i.e., dequantization) on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies an inverse transform.
[0505] The reconstruction unit 4506 can sum the residual block with the corresponding predicted block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 4507, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0506] Figure 19 This is a schematic diagram of an example encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive compensation (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image, respectively, by adding an offset and by applying a finite impulse response (FIR) filter, and by utilizing the encoded / decoded side information through signal transmission offset and filter coefficients to reduce the mean square error between the original and reconstructed samples. ALF 4606 is located in the last processing stage of each image and can be considered as a tool to attempt to capture and repair artifacts caused by previous stages.
[0507] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using reference images obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed into an entropy coding component 4618. The entropy coding component 4618 entropy-codes the prediction results and the quantized transform coefficients and transmits them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 is able to output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.
[0508] The following is a list of some preferred solutions.
[0509] The following solutions illustrate examples of the embodiments discussed herein.
[0510] 1. A method for processing video data, comprising: when performing progressive decode refresh (GDR) based video encoding / decoding, determining a reference to apply one or more constraints to an adaptive parameter set (APS); and performing a conversion between visual media data and a bitstream based on said APS.
[0511] 2. The method according to claim 1, wherein the APS cannot be referenced by a GDR image, a clean area, or a dirty area.
[0512] 3. The method according to claim 1 or 2, wherein the APS is allowed to be reset or initialized when encoding or decoding the GDR image, clean area or dirty area.
[0513] 4. The method according to any one of claims 1-3, wherein the APS generated from the GDR image, clean area or dirty area is permitted to be transmitted, derived, predefined or stored at the encoder or decoder.
[0514] 5. The method according to any one of claims 1-4, wherein the APS generated from a GDR image, a clean region, or a dirty region is allowed to be referenced by a non-GDR image.
[0515] 6. The method according to any one of claims 1-5, wherein the APS is an adaptive loop filter (ALF) APS, and the ALF APS includes different parameters when associated with ALF-luminance, ALF-chrominance, or cross-component ALF (CCALF).
[0516] 7. The method according to any one of claims 1-6, wherein a syntax element for transmitting ALF-luminance parameters via signal transmission is included in the APS, or a syntax element for transmitting ALF-chroma parameters via signal transmission is included in the APS, or a syntax element for transmitting CCALF-blue chromatic aberration (Cb) or red chromatic aberration (Cr) parameters via signal transmission is included in the APS, or the number of alternatives in ALF-luminance is included in the APS, or the filter shape or type of ALF-luminance is included in the APS, or a classifier index for each alternative of ALF-luminance is included in the APS, or an indication of enabling each ALF-luminance alternative. The syntax elements of the alternative nonlinear function are included in the APS, or for each alternative of ALF-luminance; the number of filters or combining categories transmitted through the signal are included in the APS, or for each alternative of ALF-luminance; the category combining result is included in the APS, or for each alternative of ALF-luminance; filter coefficients with symmetrical design are included in the APS, or for each alternative of ALF-luminance; filter coefficients without symmetrical design are included in the APS, or for each alternative of ALF-luminance; filter nonlinear limiting parameters with symmetrical design are included in the APS, or for each alternative of ALF-luminance. The filter nonlinear limiting parameter without a symmetrical design is included in the APS, or the filter shape or type of ALF-chroma is included in the APS, or the number of substitutions in ALF-chroma is included in the APS, or for each substitution of ALF-chroma, the syntax element indicating enabling the nonlinear function is included in the APS, or for each substitution of ALF-chroma, filter coefficients with a symmetrical design are included in the APS, or for each substitution of ALF-chroma, filter coefficients without a symmetrical design are included in the APS, or for each substitution of ALF-chroma, filter nonlinear parameters with a symmetrical design are included in the APS. In the APS, or for each alternative of ALF-chromaticity, the nonlinear parameters of the filter without a symmetrical design are included in the APS, or the filter shape or type of CCALF is included in the APS, or the number of filters transmitted through the signal for Cb or Cr of CCALF is included in the APS, or for each Cb / Cr filter transmitted through the signal for CCALF, the absolute value of the filter coefficients is included in the APS, or for each Cb / Cr filter transmitted through the signal for CCALF, the sign of the filter coefficients is included in the APS, or parameters related to ALF or CCALF are included in the APS.
[0517] 8. The method according to any one of claims 1-7, wherein the APS is a luminance mapping (LMCS) APS with chroma scaling, and the LMCS APS includes different parameters for luminance mapping and chroma scaling.
[0518] 9. The method according to any one of claims 1-8, wherein the index of the first valid codeword segment of the luminance mapping is included in the APS, or the index of the last valid codeword segment of the luminance mapping is included in the APS, or the maximum offset between the unmapped codeword and the mapped codeword of the luminance mapping is included in the APS, or for each valid codeword segment, the absolute value of the offset between the unmapped codeword and the mapped codeword is included in the APS, or for each valid codeword segment, the sign of the offset between the unmapped codeword and the mapped codeword is included in the APS, or the absolute value of the chroma scaling offset is included in the APS, or the sign of the chroma scaling offset is included in the APS, or any other parameter related to luminance mapping or chroma scaling is included in the APS.
[0519] 10. The method according to any one of claims 1-9, wherein the APS is a sample adaptive compensation (SAO) APS, a cross-component SAO (CCSAO) APS, a bilateral filter (BF) APS, and the APS includes different parameters for SAO, CCSAO, or BF.
[0520] 11. The method according to any one of claims 1-10, wherein the classifier index of the SAO is included in the APS, or the number of offsets of the SAO transmitted through the signal is included in the APS, or the absolute value of the offset of the SAO is included in the APS, or the sign of the offset of the SAO is included in the APS, or the effective band index of the SAO is included in the APS, or the classifier index of the CCSAO is included in the APS, or the iso-position luminance position of the CCSAO is included in the APS, or the number of offsets of the CCSAO transmitted through the signal is included in the APS. Alternatively, the absolute value of the CCSAO offset is included in the APS, or the sign of the CCSAO offset is included in the APS, or the effective band index of the CCSAO is included in the APS, or the effective edge pattern index of the CCSAO is included in the APS, or the classifier index of the BF is included in the APS, or the filter strength index of the BF is included in the APS, or the coefficients of the BF are included in the APS, or the nonlinear limiting parameter of the BF is included in the APS, or any other parameter related to SAO, CCSAO, or BF is included in the APS.
[0521] 12. The method according to any one of claims 1-11, wherein the temporal information generated from the deblocking filter (DBF), block filter (BF), scalar array (SAO), or scalar array (CCSAO) cannot be referenced from the GDR image, clean region, or dirty region; or when encoding the GDR image, clean region, or dirty region, the temporal information generated from the DBF, BF, SAO, or CCSAO is allowed to be reset or initialized; or when decoding the GDR image, clean region, or dirty region, the temporal information generated from the DBF, BF, SAO, or CCSAO is allowed to be reset or initialized; or the temporal information generated from the GDR image, clean region, or dirty region, or generated from the DBF, BF, SAO, or CCSAO, is allowed to be transmitted, derived, predefined, or stored at the encoder or decoder; or the temporal information generated from the GDR image, clean region, or dirty region, or generated from the DBF, BF, SAO, or CCSAO, can be referenced by a non-GDR image.
[0522] 13. The method according to any one of claims 1-12, wherein the time-domain information generated from DBF, BF, SAO or CCSAO includes different parameters.
[0523] 14. The method according to any one of claims 1-13, wherein the classifier index of the SAO is included in the time-domain information, or the number of offsets of the SAO transmitted by signal is included in the time-domain information, or the absolute value of the offset of the SAO is included in the time-domain information, or the sign of the offset of the SAO is included in the time-domain information, or the effective band index of the SAO is included in the time-domain information, or the classifier index of the CCSAO is included in the time-domain information, or the co-position luminance position of the CCSAO is included in the time-domain information, or the number of offsets of the CCSAO transmitted by signal is included in the time-domain information, or the absolute value of the offset of the CCSAO is included in the time-domain information, or the sign of the offset of the CCSAO is included in the time-domain information. The following parameters are included in the time-domain information: SAO, CCSAO, BF, or DBF.
[0524] 15. The method according to any one of claims 1-14, wherein, in GDR-based video encoding and decoding, motion-compensated padding, motion information generated from intra-block copy (IBC), or any other tool is limited.
[0525] 16. The method according to any one of claims 1-15, wherein motion-compensated fill is allowed to be forcibly disabled in a GDR image, clean region, or dirty region; or motion-compensated fill is allowed to be reset in a GDR image, clean region, or dirty region; or motion-compensated fill is allowed to be applied to a GDR image, clean region, or dirty region by using samples within the current image or region; or motion-compensated fill is allowed to be applied to a non-GDR image by using samples within a GDR image, clean region, or dirty region; or IBC motion information across non-GDR images and across GDR images cannot be referenced by a GDR image, clean region, or dirty region; or IBC motion information across clean region and dirty region cannot be referenced by a GDR image, clean region, or dirty region; or IBC motion information generated from a GDR image, clean region, or dirty region is allowed to be stored; or IBC motion information generated from a GDR image, clean region, or dirty region can be referenced by a non-GDR image; or IBC motion information can reference local motion information or temporal motion information.
[0526] 17. The method according to any one of claims 1-16, wherein, in GDR-based video encoding and decoding, temporal parameter reference or inheritance is restricted.
[0527] 18. The method according to any one of claims 1-17, wherein the temporal context adaptive binary arithmetic codec (CABAC) parameters cannot be referenced by the GDR image, clean region, or dirty region, or the temporal CABAC parameters cannot be referenced for the GDR image, clean region, or dirty region, or the temporal CABAC parameters are allowed to be reset or initialized when encoding the GDR image, clean region, or dirty region, or the temporal CABAC parameters are allowed to be reset or initialized when decoding the GDR image, clean region, or dirty region, or the temporal CABAC parameters generated from the GDR image, clean region, or dirty region are allowed to be transmitted through signal transmission, derived, predefined, or stored at the encoder or decoder, or the temporal CABAC parameters generated from the GDR image, clean region, or dirty region are allowed to be transmitted through signal transmission, derived, predefined, or stored at the encoder or decoder, or the temporal CABAC parameters generated from the GDR image, clean region, or dirty region are allowed to be generated from the GDR image, clean region, or dirty region. The temporal CABAC parameters generated from an image, clean region, or dirty region are referenced by a non-GDR image, or the temporal parameters cannot be referenced by a GDR image, clean region, or dirty region, or the temporal parameters cannot be referenced for a GDR image, clean region, or dirty region, or the temporal parameters are allowed to be reset or initialized when encoding a GDR image, clean region, or dirty region, or the temporal parameters are allowed to be reset or initialized when decoding a GDR image, clean region, or dirty region, or the temporal parameters generated from a GDR image, clean region, or dirty region are allowed to be transmitted, derived, predefined, or stored at the encoder or decoder, or the temporal parameters generated from a GDR image, clean region, or dirty region are allowed to be referenced by a non-GDR image.
[0528] 19. The method according to any one of claims 1-18, wherein temporal parameter references are allowed to be generated from information generated by different inter-frame or intra-frame codec tools.
[0529] 20. The method of any one of claims 1-19, wherein temporal parameters are allowed to be generated from a convolutional cross-component model (CCCM), or from a cross-component linear model (CCLM), or from any other cross-component encoding / decoding tool, or from history-based affine encoding / decoding, or from history-based local illumination compensation (LIC) encoding / decoding, or from any history-based intra-frame encoding / decoding tool or inter-frame encoding / decoding tool.
[0530] 21. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1-20.
[0531] 22. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that when the computer-executable instructions are executed by a processor, the video codec apparatus performs the method of any one of claims 1-20.
[0532] 23. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining, while performing progressive decode refresh (GDR) based video encoding and decoding, a reference to be applied to an adaptive parameter set (APS) with one or more constraints; and generating the bitstream based on the determination.
[0533] 24. A method for storing a bitstream of video, comprising: determining, while performing progressive decode refresh (GDR) based video encoding and decoding, a reference to which one or more constraints are applied in an adaptive parameter set (APS); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0534] 25. A method, apparatus or system described in this disclosure.
[0535] In the described solution, the encoder conforms to the format rules by generating an encoded representation based on those rules. In the described solution, the decoder parses the syntax elements in the encoded representation using known information about their presence or absence, based on the format rules, to generate the decoded video.
[0536] In this disclosure, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block can correspond to bits at co-positions or propagated at different positions in the bitstream defined by the syntax. For example, a macroblock can be encoded based on the error residual value after transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream based on this determination, knowing whether some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields, and generate encoded / decoded representations accordingly by including or excluding syntax fields from the encoded / decoded representation.
[0537] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition affecting machine-readable propagation signals, or a combination thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for an associated computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, which are generated to encode information to be transmitted to a suitable receiver device.
[0538] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the related program, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communication network.
[0539] The processing and logic flows described in this disclosure can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).
[0540] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or receive data from or transfer data to one or more mass storage devices via operative coupling, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0541] While this disclosure contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of this disclosure. Certain features described in the context of individual embodiments in this disclosure may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, although features may function in certain combinations as described above, and even were originally claimed in this manner, in some cases one or more features in a claimed combination may be removed from that combination, and the claimed combination may be for sub-combinations or variations thereof.
[0542] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed sequentially in the particular order or sequence shown, or requiring all shown operations to be performed in order to achieve the desired result. Furthermore, the partitioning of various system components in the embodiments described in this disclosure should not be construed as requiring such partitioning in all embodiments.
[0543] Only a few implementations and examples are described, and other implementations, improvements and variations may be made based on what is described and shown in this disclosure.
[0544] When there is no intermediate component (other than a line, trace, or other medium between the first and second components), the first component is directly coupled to the second component. When there is an intermediate component between the first and second components other than a line, trace, or other medium, the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. The use of the term "about" means including a range of ±10% of the following figures, unless otherwise specified.
[0545] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The present examples are to be considered illustrative rather than restrictive and are not intended to be limited to the details set forth herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0546] Furthermore, the technologies, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate may be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as couplings may be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component (whether electrical, mechanical, or other). Other examples of variations, substitutions, and modifications can be determined by those skilled in the art and may be made without departing from the spirit and scope of this disclosure.
Claims
1. A method for processing video data, comprising: When performing progressive decode refresh (GDR) based video encoding and decoding, determine which one or more constraints to apply to the reference in the adaptive parameter set (APS); as well as The conversion between visual media data and bitstream is performed based on the APS.
2. The method according to claim 1, wherein, The APS cannot be referenced by GDR images, clean areas, or dirty areas.
3. The method according to claim 1 or 2, wherein, When encoding or decoding GDR images, clean areas, or dirty areas, the APS can be reset or initialized.
4. The method according to any one of claims 1-3, wherein, The APS generated from GDR images, clean areas, or dirty areas is allowed to be transmitted, derived, predefined, or stored at the encoder or decoder.
5. The method according to any one of claims 1-4, wherein, The APS generated from a GDR image, a clean area, or a dirty area can be referenced by a non-GDR image.
6. The method according to any one of claims 1-5, wherein, The APS is an Adaptive Loop Filter (ALF) APS, and the ALF APS includes different parameters when associated with ALF-luminance, ALF-chrominance, or Cross Component ALF (CCALF).
7. The method according to any one of claims 1-6, wherein: Syntax elements for transmitting ALF-luminance parameters via signal transmission are included in the APS, or Syntax elements for transmitting ALF-chromatic parameters via signals are included in the APS, or Syntax elements for transmitting parameters of CCALF-blue chromatic aberration (Cb) or CCALF-red chromatic aberration (Cr) via signals are included in the APS, or The number of substitutions in ALF-luminance is included in the APS, or The ALF-luminance filter shape or type is included in the APS, or Each alternative classifier index for ALF-luminance is included in the APS, or The syntax element indicating that each alternative nonlinear function for ALF-luminance is enabled is included in the APS, or For each alternative to ALF-luminance, the number of filters or combining categories transmitted through the signal is included in the APS, or For each alternative to ALF-luminance, the category merging result is included in the APS, or For each alternative to ALF-luminance, filter coefficients with a symmetrical design are included in the APS, or For each alternative to ALF-luminance, filter coefficients without a symmetrical design are included in the APS, or For each alternative to ALF-luminance, the nonlinear limiting parameters of the filter with a symmetrical design are included in the APS, or For each alternative to ALF-luminance, the nonlinear limiting parameter of the filter, which does not have a symmetrical design, is included in the APS, or The ALF-chromatic filter shape or type is included in the APS, or The number of substitutions in ALF-chromaticity is included in the APS, or For each alternative to ALF-chroma, the syntax element indicating that non-linear functions are enabled is included in the APS, or For each substitution of ALF-chromaticity, filter coefficients with a symmetrical design are included in the APS, or For each alternative to ALF-chromaticity, filter coefficients without a symmetrical design are included in the APS, or For each substitution of ALF-chromaticity, the nonlinear parameters of the filter with a symmetrical design are included in the APS, or For each alternative to ALF-chromaticity, the nonlinear parameters of the filter without a symmetrical design are included in the APS, or The filter shape or type of CCALF is included in the APS, or The number of filters for Cb or Cr in CCALF transmitted through the signal is included in the APS, or For each filter transmitting the signal for Cb / Cr of CCALF, the absolute value of the filter coefficients is included in the APS, or For each filter transmitting the signal for Cb / Cr of CCALF, the sign of the filter coefficients is included in the APS, or Parameters related to ALF or CCALF are included in the APS.
8. The method according to any one of claims 1-7, wherein, The APS is a Luminance Mapping (LMCS) APS with Chroma Scaling, and the LMCS APS includes different parameters for luminance mapping and chroma scaling.
9. The method according to any one of claims 1-8, wherein: The index of the first valid codeword segment of the luminance map is included in the APS, or The index of the last valid codeword segment of the luminance map is included in the APS, or The maximum offset between the unmapped codeword and the mapped codeword in the luminance mapping is included in the APS, or For each valid codeword segment, the absolute value of the offset between the unmapped codeword and the mapped codeword is included in the APS, or For each valid codeword segment, the sign of the offset between the unmapped codeword and the mapped codeword is included in the APS, or The absolute value of the chroma scaling offset is included in the APS, or The sign of the chroma scaling offset is included in the APS, or Any other parameters related to luminance mapping or chrominance scaling are included in the APS.
10. The method according to any one of claims 1-9, wherein, The APS is a Sample Adaptive Compensation (SAO) APS, a Cross-Component SAO (CCSAO) APS, or a Bilateral Filter (BF) APS, and the APS includes different parameters for SAO, CCSAO, or BF.
11. The method according to any one of claims 1-10, wherein: The classifier index of SAO is included in the APS, or The amount of offset transmitted by the SAO through the signal is included in the APS, or The absolute value of the offset of SAO is included in the APS, or The sign of the offset of SAO is included in the APS, or The valid band index of SAO is included in the APS, or The classifier index of CCSAO is included in the APS, or The co-positional luminance location of CCSAO is included in the APS, or The amount of offset transmitted via signal transmission in CCSAO is included in the APS, or The absolute value of the offset of CCSAO is included in the APS, or The sign of the offset of CCSAO is included in the APS, or The valid band index of CCSAO is included in the APS, or The valid edge pattern index of CCSAO is included in the APS, or The classifier index of BF is included in the APS, or The filter strength index of BF is included in the APS, or The coefficient of BF is included in the APS, or The nonlinear limiting parameter of BF is included in the APS, or Any other parameters related to SAO, CCSAO, or BF are included in the APS.
12. The method according to any one of claims 1-11, wherein: Temporal information generated from deblocking filter (DBF), block filter (BF), scalar autoclave (SAO), or cCSAO cannot be referenced from GDR images, clean regions, or dirty regions, or When encoding GDR images, clean regions, or dirty regions, the temporal information generated from DBF, BF, SAO, or CCSAO is allowed to be reset or initialized, or When decoding GDR images, clean regions, or dirty regions, the temporal information generated from DBF, BF, SAO, or CCSAO can be reset or initialized, or Temporal information generated from GDR images, clean or dirty regions, or from DBF, BF, SAO, or CCSAO, is permitted to be transmitted via signal transmission, derived, predefined, or stored at the encoder or decoder. Temporal information generated from GDR images, clean or dirty regions, or from DBF, BF, SAO, or CCSAO can be referenced by non-GDR images.
13. The method according to any one of claims 1-12, wherein, The time-domain information generated from DBF, BF, SAO, or CCSAO includes different parameters.
14. The method according to any one of claims 1-13, wherein: The classifier index of SAO is included in the time-domain information, or The amount of offset transmitted by the signal in the SAO is included in the time-domain information, or The absolute value of the offset of SAO is included in the time-domain information, or The sign of the offset of SAO is included in the time-domain information, or The effective band index of SAO is included in the time-domain information, or The classifier index of CCSAO is included in the time-domain information, or The co-positional luminance location of CCSAO is included in the time-domain information, or The amount of offset transmitted through the signal in CCSAO is included in the time-domain information, or The absolute value of the offset of CCSAO is included in the time-domain information, or The sign of the offset in CCSAO is included in the time-domain information, or The effective band index of CCSAO is included in the time-domain information, or The effective edge pattern index of CCSAO is included in the time-domain information, or The classifier index of BF is included in the time-domain information, or The filter strength index of BF is included in the time-domain information, or The coefficients of BF are included in the time-domain information, or The nonlinear limiting parameter of BF is included in the time-domain information, or The boundary strength of the DBF is included in the time-domain information, or The filter strength of the DBF is included in the time-domain information, or The filter length of the DBF is included in the time-domain information, or Any other parameters related to SAO, CCSAO, BF, or DBF are included in the time-domain information.
15. The method according to any one of claims 1-14, wherein, In GDR-based video coding and decoding, motion-compensated padding, motion information generated from intra-block copy (IBC), or any other tools are limited.
16. The method according to any one of claims 1-15, wherein: Allow motion-compensated fill to be forcibly disabled in GDR images, clean areas, or dirty areas, or Allows motion-compensated fill to be reset in GDR images, clean areas, or dirty areas, or Motion-compensated fill can be applied to GDR images, clean or dirty regions using samples from the current image or region, or... Motion-compensated filling can be applied to non-GDR images by using samples from GDR images, clean regions, or dirty regions.
17. The method according to any one of claims 1-16, wherein: IBC motion information across non-GDR images and across GDR images cannot be referenced by GDR images, clean regions, or dirty regions, or IBC motion information across clean and dirty regions cannot be referenced by GDR images, clean regions, or dirty regions, or Allows IBC motion information generated from GDR images, clean or dirty regions to be stored, or IBC motion information generated from GDR images, clean regions, or dirty regions can be referenced by non-GDR images, or IBC motion information can reference local motion information or temporal motion information.
18. The method according to any one of claims 1-17, wherein, In GDR-based video encoding and decoding, the ability to reference or inherit temporal parameters is restricted.
19. The method according to any one of claims 1-18, wherein: Temporal context adaptive binary arithmetic codec (CABAC) parameters cannot be referenced by GDR images, clean regions, or dirty regions, or For GDR images, clean or dirty regions, temporal CABAC parameters cannot be referenced, or When encoding GDR images, clean regions, or dirty regions, temporal CABAC parameters can be reset or initialized, or When decoding GDR images, clean regions, or dirty regions, temporal CABAC parameters can be reset or initialized, or Temporal CABAC parameters generated from GDR images, clean regions, or dirty regions are allowed to be transmitted via signal transmission, derived, predefined, or stored at the encoder or decoder, or Temporal CABAC parameters generated from GDR images, clean regions, or dirty regions can be referenced by non-GDR images, or Temporal parameters cannot be referenced by GDR images, clean areas, or dirty areas, or For GDR images, whether clean or dirty, temporal parameters cannot be referenced, or When encoding GDR images, clean regions, or dirty regions, temporal parameters can be reset or initialized, or When decoding GDR images, clean regions, or dirty regions, temporal parameters can be reset or initialized, or Temporal parameters generated from GDR images, clean regions, or dirty regions are allowed to be transmitted via signal transmission, derived, predefined, or stored at the encoder or decoder, or Temporal parameters generated from GDR images, clean regions, or dirty regions can be referenced by non-GDR images.
20. The method according to any one of claims 1-19, wherein, Temporal parameters can be referenced from information generated by different inter-frame or intra-frame codec tools.
21. The method according to any one of claims 1-20, wherein: Temporal parameters can be generated from the convolutional cross-component model (CCCM), or Allows time-domain parameters to be generated from a cross-component linear model (CCLM), or Allow time-domain parameters to be generated from any other cross-component codec tool, or Allows time-domain parameters to be generated from history-based affine encoding / decoding, or Allows temporal parameters to be generated from history-based local illumination compensation (LIC) encoding / decoding, or Temporal parameters can be generated from any history-based intra-frame or inter-frame codec tool.
22. An apparatus for processing video data, comprising: processor; and a non-transitory memory thereon having instructions, wherein, when executed by the processor, the instructions cause the processor to perform the method of any one of claims 1-21.
23. A non-transitory computer-readable medium comprising a computer program product for use with a video encoding / decoding device, wherein, The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that when the computer-executable instructions are executed by a processor, the video codec device performs the method of any one of claims 1-21.
24. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein, The method includes: When performing progressive decode refresh (GDR) based video encoding and decoding, determine which one or more constraints to apply to the reference in the adaptive parameter set (APS); and The bit stream is generated based on the APS.
25. A method for storing a video bitstream, comprising: When performing progressive decode refresh (GDR) based video encoding and decoding, determine which one or more constraints to apply to the reference in the adaptive parameter set (APS); The bit stream is generated based on the APS; as well as The bit stream is stored in a non-transitory computer-readable recording medium.