Interaction between loop shaping and interframe codec tools
By converting video blocks and codec representations in video processing through loop shaping technology, the problem of video compression technology occupying a large amount of bandwidth is solved, the codec efficiency and video quality are improved, and the sensitivity to data loss and errors is reduced.
Patent Information
- Application Number
- CN202080012232.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-01
- Filing Date
- 2020-02-01
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2040-02-01
AI Technical Summary
Existing video compression technology occupies a large amount of bandwidth on the Internet and digital communication networks. As the number of connected user devices increases, bandwidth demand continues to grow, and existing video processing standards are unable to effectively improve encoding and decoding efficiency.
Loop shaping technology is used to optimize the video processing process by converting video blocks and codec representations between the first domain and the second domain, using methods such as motion information refinement, chroma residual scaling, codec mode parameters and filtering operations.
It improves video encoding and decoding efficiency, reduces bandwidth requirements, enhances video quality and the flexibility of encoding algorithms, and reduces sensitivity to data loss and errors.
Smart Images

Figure CN113383547B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is based on international patent application No. PCT / CN2020 / 074136 filed on February 1, 2020, which claims priority to and the benefit of international patent application No. PCT / CN2019 / 074437 filed on February 1, 2019. The entire disclosure of the aforementioned application is incorporated by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to video processing technology, devices and systems. Background Art
[0004] Despite advances in video compression technology, digital video still accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of networked user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0005] Devices, systems, and methods related to digital video processing, and more particularly, to in-loop reshaping (ILR) for video processing, are described. The methods described can be applied to both existing video processing standards (e.g., High Efficiency Video Coding (HEVC)) and future video processing standards, or to video processors that include video codecs.
[0006] In one representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a motion information refinement process based on samples in a first domain or a second domain for a conversion between a current video block of a video and a codec representation of the video; and performing the conversion based on a result of the motion information refinement process, wherein during the conversion, samples are obtained from a first prediction block in the first domain using unrefined motion information for the current video block, at least a second prediction block is generated in the second domain using refined motion information for determining a reconstructed block, and reconstructed samples for the current video block are generated based on at least the second prediction block.
[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein during the conversion, the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, wherein a codec tool is applied during the conversion using parameters derived based on at least a first set of samples in a video region of the video and a second set of samples in a reference picture of the current video block, and wherein the domain of the first sample and the domain of the second sample are aligned.
[0008] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining parameters of a codec mode for a current video block of a current video region of a video based on one or more parameters of a codec mode of a previous video region; and performing codecs on the current video block to generate a codec representation of the video based on the determination, wherein the parameters of the codec mode are included in a parameter set in the codec representation of the video, wherein performing the codecs comprises transforming a representation of the current video block in a first domain to a representation of the current video block in a second domain, and wherein, during the codecs performed using the codec mode, the current video block is constructed based on the first domain and the second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0009] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes receiving a codec representation of a video including a parameter set, wherein the parameter set includes parameter information for a codec mode; and performing decoding of the codec representation using the parameter information to generate a current video block of a current video region of the video from the codec representation, wherein the parameter information for the codec mode is based on one or more parameters of a codec mode of a previous video region, wherein, in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0010] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion comprises applying a filtering operation to a prediction block in a first domain or in a second domain different from the first domain.
[0011] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current video block of a video and a codec representation of the video, wherein during the conversion, a final reconstructed block is determined for the current video block, and wherein the temporary reconstructed block is generated using a prediction method and represented in a second domain.
[0012] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein, in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein a parameter set in the codec representation includes parameter information for the codec mode.
[0013] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video block of the video being a chroma block and a codec representation of the video, wherein during the conversion, the current video block is constructed based on a first domain and a second domain, and wherein the conversion further comprises applying a forward shaping process and / or an inverse shaping process to one or more chroma components of the current video block.
[0014] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video chroma block of a video and a codec representation of the video, wherein performing the conversion comprises: determining based on a rule whether luma-dependent chroma residual scaling (LCRS) is enabled or disabled, and reconstructing the current video chroma block based on the determination.
[0015] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining, for a conversion between a current video block of a video and a codec representation of the video, based on one or more coefficient values of the current video block, whether to disable use of a codec mode; and performing the conversion based on the determination, wherein during the conversion using the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0016] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: for conversion between a current video block exceeding a virtual pipe data unit (VPDU) of a video, dividing the current video block into regions; and performing the conversion by separately applying a codec mode to each region, wherein during the conversion by applying the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0017] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining, for a conversion between a current video block of a video and a codec representation of the video, whether to disable use of a codec mode based on a size or a color format of the current video block; and performing the conversion based on the determination, wherein during the conversion using the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0018] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein, in the codec mode, the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein at least one syntax element in the codec representation provides an indication of the use of the codec mode and an indication of a shaper model.
[0019] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining that a codec mode is disabled for conversion between a current video block of a video and a codec representation of the video, wherein the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner; and conditionally skipping forward shaping and / or inverse shaping based on the determination.
[0020] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein, in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein a plurality of forward shapings and / or a plurality of inverse shapings are applied in the shaping mode for the video region.
[0021] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining a codec mode that enables conversion between a current video block of a video and a codec representation of the video; and performing the conversion using a palette mode, wherein in the palette mode, at least a palette of representative sample values is used for the current video block, and wherein, in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0022] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: for converting between a current video block of a video and a codec representation of the video, determining that the current video block is encoded in a palette mode, wherein in the palette mode, a palette of at least representative sample values is used to encode the current video block; and, due to the determination, performing the conversion by disabling a codec mode, wherein, when the codec mode is applied to the video block, the video block is constructed based on chroma residuals scaled in a luma-dependent manner.
[0023] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion uses a first codec mode and a palette codec mode, wherein in the palette codec mode, a palette of at least representative pixel values is used to encode and decode the current video block; and performing a conversion between a second video block of the video that was encoded without using the palette codec mode and the codec representation of the video, wherein the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein the first codec mode is applied differently to the first video block and the second video block.
[0024] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining a codec mode enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using an intra block copy mode, wherein the intra block copy mode generates a prediction block using at least a block vector pointing to a picture including the current video block, and wherein, in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0025] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: for a conversion between a current video block of a video and a codec representation of the video, determining that the current video block is encoded in an intra block copy (IBC) mode, wherein the intra block copy mode uses at least a block vector pointing to a video frame containing the current video block to generate a prediction block for encoding and decoding the current video block; and, due to the determination, performing the conversion by disabling the codec mode, wherein, when the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0026] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion uses an intra block copy mode and a first codec mode, wherein the intra block copy mode uses at least a block vector pointing to a video frame containing the current video block to generate a prediction block; and performing a conversion between a second video block of the video that was encoded and decoded without using the intra block copy mode and a codec representation of the video, wherein the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein the first codec mode is applied differently to the first video block and the second video block.
[0027] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining a codec mode enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using a block-based delta pulse codec modulation (BDPCM) mode, wherein in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0028] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: for converting between a current video block of a video and a codec representation of the video, determining that the current video block is coded using a block-based delta pulse codec modulation (BDPCM) mode; and, due to the determination, performing the conversion by disabling a codec mode, wherein, when the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0029] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and a block-based delta pulse codec modulation (BDPCM) mode; and performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is coded without using the BDPCM mode and the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein the first codec mode is applied differently to the first video block and the second video block.
[0030] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining a codec mode enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using a transform skip mode, wherein in the transform skip mode, transforming a prediction residual is skipped when encoding the current video block, wherein in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0031] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: for a conversion between a current video block of a video and a codec representation of the video, determining that the current video block is coded in a transform skip mode, wherein in the transform skip mode, a transform on a prediction residual is skipped when coding the current video block; and, due to the determination, performing the conversion by disabling a codec mode, wherein, when the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0032] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and a transform skip mode, wherein in the transform skip mode, a transform of a prediction residual is skipped when encoding and decoding the current video block; and performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded and decoded without using the transform skip mode and the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein the first codec mode is applied differently to the first video block and the second video block.
[0033] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining a codec mode enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using an intra pulse codec modulation mode, wherein the current video block is coded without applying a transform and transform-domain quantization, wherein the current video block is constructed based on samples in a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0034] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: for a conversion between a current video block of a video and a codec representation of the video, determining that the current video block is coded in an intra pulse codec modulation mode, wherein the current video block is coded in the intra pulse codec modulation mode without applying a transform and a transform domain quantization; and, due to the determination, performing the conversion by disabling a codec mode, wherein, when the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0035] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and an intra-frame pulse codec modulation mode, wherein the video block is encoded without applying a transform and a transform-domain quantization in the intra-frame pulse codec modulation mode; and performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded without using the intra-frame pulse codec modulation mode and the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on the first domain and the second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein the first codec mode is applied differently to the first video block and the second video block.
[0036] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining a codec mode enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using a modified transform and quantization bypass mode, wherein in the modified transform and quantization bypass mode, the current video block is losslessly coded without transform and quantization, wherein in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0037] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: for a conversion between a current video block of a video and a codec representation of the video, determining that the current video block is coded in a transform and quantization bypass mode, wherein in the transform and quantization bypass mode, the current video block is losslessly coded without transform and quantization; and, based on the determination, performing the conversion by disabling a codec mode, wherein, when the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0038] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and a transform and quantization bypass mode, wherein in the transform and quantization bypass mode, the current video block is losslessly encoded and decoded without transform and quantization; and performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded and decoded without using the transform and quantization bypass mode, and the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein the first codec mode is applied differently to the first video block and the second video block.
[0039] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein, in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein information for the codec mode is signaled in a parameter set other than a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), or an adaptation parameter set (APS) for carrying adaptive loop filtering (ALF) parameters.
[0040] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein, in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein information for the codec mode is signaled in an adaptation parameter set (APS) along with adaptive loop filtering (ALF) information, wherein the information for the codec mode and the ALF information are included in one NAL unit.
[0041] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode in which the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein information for the codec mode is signaled in an adaptive parameter set (APS) of a first type that is different from a second type that is used to signal adaptive loop filtering (ALF) information.
[0042] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode in which the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, wherein the video region is not allowed to reference an adaptation parameter set or a parameter set signaled prior to a data structure of a specified type for processing the video, and wherein the data structure of the specified type is signaled prior to the video region.
[0043] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode in which the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein syntax elements of a parameter set comprising parameters for processing the video have predefined values in a conforming bitstream.
[0044] In another representative aspect, the above-described method is embodied in the form of processor-executable code and stored in a computer-readable program medium.
[0045] In another representative aspect, a device configured or operable to perform the above method is disclosed. The device may include a processor programmed to implement the method.
[0046] In another representative aspect, a video decoder device may implement a method as described herein.
[0047] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the description, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 An example of constructing a Merge candidate list is shown.
[0049] Figure 2 Examples of locations of spatial candidates are shown.
[0050] Figure 3 An example of candidate pairs for which redundancy checking of spatial merge candidates is performed is shown.
[0051] Figure 4A and Figure 4B An example of the position of the second prediction unit (PU) based on the size and shape of the current block is shown.
[0052] Figure 5 An example of motion vector scaling for temporal merge candidates is shown.
[0053] Figure 6 An example of candidate positions of time-domain Merge candidates is shown.
[0054] Figure 7 An example of generating a combined bi-predictive Merge candidate is shown.
[0055] Figure 8 An example of constructing motion vector prediction candidates is shown.
[0056] Figure 9 An example of motion vector scaling for spatial motion vector candidates is shown.
[0057] Figure 10 An example of motion prediction using an Alternative Temporal Motion Vector Prediction (ATMVP) algorithm for a Coding Unit (CU) is shown.
[0058] Figure 11 An example of a codec unit (CU) with sub-blocks and neighboring blocks used by a Spatial-Temporal Motion Vector Prediction (STMVP) algorithm is shown.
[0059] Figure 12 An example of neighboring sample points used to derive illumination compensation (IC) parameters is shown.
[0060] Figure 13A and Figure 13B Examples of a simplified 4-parameter affine model and a simplified 6-parameter affine model are shown, respectively.
[0061] Figure 14An example of the affine motion vector field (MVF) of each sub-block is shown.
[0062] Figure 15A and Figure 15B Examples of a 4-parameter affine model and a 6-parameter affine model are shown respectively.
[0063] Figure 16 An example of motion vector prediction of AF_INTER of inherited affine candidates is shown.
[0064] Figure 17 An example of motion vector prediction of AF_INTER of constructed affine candidates is shown.
[0065] Figure 18A and Figure 18B Example candidate blocks and CPMV prediction value derivation for AF_MERGE mode are shown respectively.
[0066] Figure 19 An example of candidate positions for the affine merge mode is shown.
[0067] Figure 20 An example of the UMVE search process is shown.
[0068] Figure 21 An example of a UMVE search point is shown.
[0069] Figure 22 An example of Decoder Side Motion Vector Refinement (DMVR) based on bilateral template matching is shown.
[0070] Figure 23 An exemplary flow diagram of a decoding flow with shaping is shown.
[0071] Figure 24 An example of neighboring samples used in a bilateral filter is shown.
[0072] Figure 25 An example of a window covering two samples used in the weight calculation is shown.
[0073] Figure 26 An example of a scan pattern is shown.
[0074] Figure 27 An example of an inter-mode decoding process is shown.
[0075] Figure 28 Another example of an inter-mode decoding process is shown.
[0076] Figure 29An example of an inter-mode decoding process using a post-reconstruction filter is shown.
[0077] Figure 30 Another example of an inter-mode decoding process using a post-reconstruction filter is shown.
[0078] Figure 31A and Figure 31B A flow chart illustrating an example method for video processing is shown.
[0079] Figures 32A to 32D A flow chart illustrating an example method for video processing is shown.
[0080] Figure 33 A flow chart illustrating an example method for video processing is shown.
[0081] Figure 34A and Figure 34B A flow chart illustrating an example method for video processing is shown.
[0082] Figures 35A to 35F A flow chart illustrating an example method for video processing is shown.
[0083] Figures 36A to 36C A flow chart illustrating an example method for video processing is shown.
[0084] Figures 37A to 37C A flow chart illustrating an example method for video processing is shown.
[0085] Figures 38A to 38L A flow chart illustrating an example method for video processing is shown.
[0086] Figures 39A to 39E A flow chart illustrating an example method for video processing is shown.
[0087] Figure 40A and Figure 40B An example of a hardware platform for implementing the visual media decoding or visual media encoding techniques described in this document is shown. DETAILED DESCRIPTION
[0088] Due to the growing demand for higher resolution video, video processing methods and techniques are ubiquitous in modern technology. Video codecs typically include electronic circuits or software that compress or decompress digital video and are constantly being improved to provide higher coding and decoding efficiency. Video codecs convert uncompressed video to a compressed format and vice versa. There is a complex relationship between video quality, the amount of data used to represent the video (determined by the bit rate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, and end-to-end delay (latency). The compression format typically conforms to a standard video compression specification, such as the High Efficiency Video Codec (HEVC) standard (also known as H.265 or MPEG-H Part 2), the to-be-completed Versatile Video Coding standard, or other current and / or future video codec standards.
[0089] Embodiments of the disclosed technology can be applied to existing video codec standards (e.g., HEVC, H.265) and future standards to improve compression performance. Section headings are used in this document to improve the readability of the description and do not in any way limit the discussion or embodiments (and / or implementation methods) to the corresponding section.
[0090] 1 Example of inter-frame prediction in HEVC / H.265
[0091] Video codec standards have improved significantly in recent years and now partially provide high codec efficiency and support for higher resolutions. Recent standards such as HEVC and H.265 are based on a hybrid video codec structure that utilizes temporal prediction plus transform coding.
[0092] 1.1 Example of Prediction Model
[0093] Each inter-prediction PU (prediction unit) has motion parameters for one or two reference picture lists. In some embodiments, the motion parameters include motion vectors and reference picture indices. In other embodiments, the use of one of the two reference picture lists can also be signaled using inter_pred_idc. In other embodiments, motion vectors can be explicitly encoded and decoded as deltas relative to the predicted value.
[0094] When a CU is encoded or decoded in skip mode, one PU is associated with the CU and there are no significant residual coefficients, no coded motion vector increments or reference picture indices. A Merge mode is specified, whereby the motion parameters of the current PU are obtained from neighboring PUs including spatial and temporal candidates. Merge mode can be applied to any inter-predicted PU, not only for skip mode. An alternative to Merge mode is the explicit transmission of motion parameters, where the motion vector (more precisely, the motion vector difference (MVD) compared to the motion vector prediction value), the corresponding reference picture index for each reference picture list, and the reference picture list usage are explicitly signaled per PU. This type of mode is named Advanced Motion Vector Prediction (AMVP) in this document.
[0095] When signaling indicates that one of the two reference picture lists is to be used, a PU is generated from a block of samples. This is called "unidirectional prediction." Unidirectional prediction is applicable to both P slices and B slices.
[0096] When signaling indicates that two reference picture lists are to be used, a PU is generated from two sample blocks. This is called "bi-prediction." Bi-prediction is only applicable to B slices.
[0097] Reference Image List
[0098] In HEVC, the term inter prediction is used to refer to predictions derived from data elements (e.g., sample values or motion vectors) of reference pictures other than the currently decoded picture. As in H.264 / AVC, a picture can be predicted from multiple reference pictures. Reference pictures used for inter prediction are organized in one or more reference picture lists. A reference index identifies which reference picture in the list should be used to create the prediction signal.
[0099] A single reference picture list (list 0) is used for P slices, and two reference picture lists (list 0 and list 1) are used for B slices. It should be noted that the reference pictures included in lists 0 / 1 can be based on past and future pictures in terms of capture / display order.
[0100] 1.1.1 Implementation of Candidates for Merge Mode
[0101] When predicting a PU using Merge mode, the index pointing to the entry in the Merge candidate list is parsed from the bitstream and used to retrieve the motion information. The construction of this list can be summarized according to the following sequence of steps:
[0102] Step 1: Initial candidate derivation
[0103] Step 1.1: Spatial Candidate Derivation
[0104] Step 1.2: Redundancy check of spatial candidates
[0105] Step 1.3: Time Domain Candidate Derivation
[0106] Step 2: Additional candidate insertions
[0107] Step 2.1: Create bidirectional prediction candidates
[0108] Step 2.2: Insert zero motion candidates
[0109] Figure 1 An example of constructing a Merge candidate list based on the sequence of steps summarized above is shown. For spatial Merge candidate derivation, a maximum of four Merge candidates are selected from candidates located at five different positions. For temporal Merge candidate derivation, a maximum of one Merge candidate is selected from two candidates. Since the number of candidates for each PU is assumed to be constant at the decoder, additional candidates are generated when the number of candidates does not reach the maximum number of Merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is constant, the index of the best Merge candidate is encoded using truncated unary (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single Merge candidate list, which is the same as the Merge candidate list of a 2N×2N prediction unit.
[0110] 1.1.2 Constructing Spatial Merge Candidates
[0111] In the derivation of spatial Merge candidates, from Figure 2 Up to four Merge candidates are selected from the candidates at the positions depicted in . The order of derivation is A1, B1, B0, A0, and B2. Position B2 is considered only when any PU at position A1, B1, B0, A0 is unavailable (for example, because it belongs to another strip or slice) or is intra-coded. After the candidate at position A1 is added, a redundancy check is performed on the addition of the remaining candidates. The redundancy check ensures that candidates with the same motion information are excluded from the list, thereby improving coding efficiency.
[0112] In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs in Figure 3 The pairs linked by arrows in , and only when the candidates used for redundancy check do not have the same motion information, the corresponding candidate is added to the list. Another source of duplicate motion information is a "second PU" associated with a partition other than 2N×2N. As an example, Figure 4A and Figure 4B The second PU is depicted for the N×2N and 2N×N cases, respectively. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In some embodiments, adding this candidate may result in two prediction units having the same motion information, which is redundant for having only one PU in the codec unit. Similarly, position B1 is not considered when the current PU is partitioned into 2N×N.
[0113] 1.1.3 Building Time Domain Merge Candidates
[0114] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the collocated PU belonging to the picture with the smallest POC difference with the current picture in a given reference picture list. The reference picture list to be used for the derivation of the collocated PU is explicitly signaled in the slice header.
[0115] Figure 5 An example of the derivation of a scaled motion vector for a temporal merge candidate (shown as a dashed line) is shown, where the motion vector is scaled from the motion vector of the collocated PU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal merge candidate is set to zero. For B slices, two motion vectors are obtained, one for reference picture list 0 and the other for reference picture list 1, and are combined to form a bi-predictive merge candidate.
[0116] like Figure 6 As depicted, in a collocated PU (Y) belonging to a reference frame, the position of the temporal candidate is selected between candidates C0 and C1. If the PU at position C0 is not available, is intra-coded, or is outside the current CTU, position C1 is used. Otherwise, position C0 is used in the derivation of the temporal merge candidate.
[0117] 1.1.4 Building Merge Candidates of Additional Types
[0118] In addition to spatiotemporal merge candidates, there are two additional types of merge candidates: combined bi-predictive merge candidates and zero merge candidates. Combined bi-predictive merge candidates are generated by utilizing spatiotemporal merge candidates. Combined bi-predictive merge candidates are used only for B slices. Combined bi-predictive candidates are generated by combining the first reference picture list motion parameters of an initial candidate with the second reference picture list motion parameters of another candidate. If the two tuples provide different motion hypotheses, they will form a new bi-predictive candidate.
[0119] Figure 7 An example of this process is shown where two candidates in the original list (710, on the left) with mvL0 and refIdxL0 or mvL1 and refIdxL1 are used to create combined bi-predictive Merge candidates that are added to the final list (720, on the right). There are many rules about combining that are considered to generate these additional Merge candidates.
[0120] Zero-motion candidates are inserted to fill the remaining entries in the Merge candidate list and thus reach the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference picture index that starts at zero and increments each time a new zero-motion candidate is added to the list. The number of reference frames used by these candidates is one for unidirectional prediction and two for bidirectional prediction, respectively. In some embodiments, no redundancy check is performed on these candidates.
[0121] 1.2 Embodiment of Advanced Motion Vector Prediction (AMVP)
[0122] AMVP utilizes the spatiotemporal correlation of motion vectors with neighboring PUs for explicit transmission of motion parameters. A motion vector candidate list is constructed by first checking the availability of the left and upper temporal neighboring PU positions, removing redundant candidates, and adding zero vectors to make the candidate list length constant. The encoder can then select the best prediction value from the candidate list and send the corresponding index indicating the selected candidate. Similar to Merge index signaling, the index of the best motion vector candidate is truncated unary coding. In this case, the maximum value to be encoded is 2 (see Figure 8 ). In the following sections, details on the derivation process of motion vector prediction candidates are provided.
[0123] 1.2.1 Example of deriving AMVP candidates
[0124] Figure 8 The derivation process of motion vector prediction candidates is summarized and can be performed for each reference picture list with refidx as input.
[0125] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. Figure 2 The motion vectors of each PU at the five different positions shown in are used to finally derive two motion vector candidates.
[0126] For temporal motion vector candidate derivation, a motion vector candidate is selected from two candidates derived based on two different collocated positions. After generating the first list of spatiotemporal candidates, duplicate motion vector candidates are removed from the list. If the number of potential candidates is greater than two, motion vector candidates with a reference picture index greater than 1 within the list are removed from the associated reference picture list. If the number of spatiotemporal motion vector candidates is less than two, an additional zero motion vector candidate is added to the list.
[0127] 1.2.2 Constructing Spatial Motion Vector Candidates
[0128] In the derivation of spatial motion vector candidates, a maximum of two candidates are considered among five potential candidates, which are selected from the candidate located as previously Figure 2 The derivation order of the left side of the current PU is defined as A0, A1 and scaled A0, scaled A1. The derivation order of the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases that can be used as motion vector candidates, two of which do not require spatial scaling and two use spatial scaling. These four different cases are summarized as follows:
[0129] --No airspace scaling
[0130] (1) Same reference picture list and same reference picture index (same POC)
[0131] (2) Different reference picture lists but same reference pictures (same POC)
[0132] --Airspace scaling
[0133] (3) Same reference picture list but different reference pictures (different POC)
[0134] (4) Different reference picture lists and different reference pictures (different POCs)
[0135] First, the case where no spatial scaling is performed is checked, followed by the case where spatial scaling is allowed. Regardless of the reference picture list, spatial scaling is considered when the POC between the reference picture of the neighboring PU and the reference picture of the current PU is different. If all PUs of the left candidate are unavailable or intra-coded, scaling of the upper motion vector is allowed to facilitate the parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling of the upper motion vector is not allowed.
[0136] like Figure 9As shown in the example in , for the spatial scaling case, the motion vectors of the neighboring PUs are scaled in a similar way to the temporal scaling. One difference is that the reference picture list and the index of the current PU are given as input; the actual scaling process is the same as the scaling process for temporal scaling.
[0137] 1.2.3 Constructing Temporal Motion Vector Candidates
[0138] Except for the reference picture index derivation, all the processes for deriving the temporal Merge candidate are the same as those for deriving the spatial motion vector candidate (e.g. Figure 6 In some embodiments, the reference picture index is signaled to the decoder.
[0139] 2. Example of inter-frame prediction method in Joint Exploration Model (JEM)
[0140] In some embodiments, a reference software called the Joint Exploration Model (JEM) is used to explore future video codec technologies. In JEM, sub-block-based prediction is adopted in several codec tools, such as affine prediction, optional temporal motion vector prediction, spatio-temporal motion vector prediction, bi-directional optical flow (BIO), frame-rate up conversion (FRUC), locally adaptive motion vector resolution (LAMVR), overlapped block motion compensation (OBMC), local illumination compensation (LIC), and decoder-side motion vector refinement (DMVR).
[0141] 2.1 Example of motion vector prediction based on sub-CU
[0142] In JEM with QuadTrees plus Binary Trees (QTBT), each CU can have at most one motion parameter set for each prediction direction. In some embodiments, two sub-CU level motion vector prediction methods are considered in the encoder by dividing the large CU into sub-CUs and deriving motion information for all sub-CUs of the large CU. The optional temporal motion vector prediction (ATMVP) method allows each CU to obtain multiple motion information sets from multiple blocks smaller than the current CU in the collocated reference picture. In the spatial-temporal motion vector prediction (STMVP) method, the motion vector of the sub-CU is recursively derived by using the temporal motion vector prediction value and the spatial neighboring motion vector. In some embodiments, in order to preserve a more accurate motion field for sub-CU motion prediction, motion compression of the reference frame can be disabled.
[0143] 2.1.1 Example of Optional Temporal Motion Vector Prediction (ATMVP)
[0144] In the ATMVP method, the temporal motion vector prediction (TMVP) method is modified by obtaining multiple motion information sets (including motion vectors and reference indices) from blocks smaller than the current CU.
[0145] Figure 10 An example of the ATMVP motion prediction process for CU 1000 is shown. The ATMVP method predicts the motion vector of sub-CU 1001 within CU 1000 in two steps. The first step is to use the time domain vector to identify the corresponding block 1051 in reference picture 1050. Reference picture 1050 is also called the motion source picture. The second step is to divide the current CU 1000 into sub-CUs 1001 and obtain the motion vector and reference index for each sub-CU from the block corresponding to each sub-CU.
[0146] In the first step, the reference picture 1050 and the corresponding block are determined based on the motion information of the spatially neighboring blocks of the current CU 1000. To avoid repeated scanning of neighboring blocks, the first merge candidate in the merge candidate list of the current CU 1000 is used. The first available motion vector and its associated reference index are set to the temporal vector and index of the motion source picture. This allows for more accurate identification of corresponding blocks compared to TMVP, where the corresponding block (sometimes referred to as a collocated block) is always located to the lower right or center relative to the current CU.
[0147] In the second step, the corresponding blocks of the sub-CU 1051 are identified by the time domain vector in the motion source picture 1050 by adding the time domain vector to the coordinates of the current CU. For each sub-CU, the motion information of its corresponding block (e.g., the minimum motion grid covering the center sample) is used to derive the motion information of the sub-CU. After the motion information of the corresponding N×N block is identified, it is converted into the motion vector and reference index of the current sub-CU in the same way as the TMVP of HEVC, where motion scaling and other processes are applied. For example, the decoder checks whether the low latency condition is met (e.g., the POC of all reference pictures of the current picture is less than the POC of the current picture), and may use the motion vector MVx (e.g., the motion vector corresponding to reference picture list X) to predict the motion vector MVy of each sub-CU (e.g., where X is equal to 0 or 1, and Y is equal to 1-X).
[0148] 2.1.2 Example of Spatio-Temporal Motion Vector Prediction (STMVP)
[0149] In the STMVP method, the motion vector of a sub-CU is recursively derived in raster scan order. Figure 11 An example of a CU with four sub-blocks and neighboring blocks is shown. Consider an 8×8 CU 1100, which includes four 4×4 sub-CUs A (1101), B (1102), C (1103), and D (1104). The neighboring 4×4 blocks in the current frame are labeled a (1111), b (1112), c (1113), and d (1114).
[0150] The motion derivation of sub-CU A starts with identifying its two spatial neighbors. The first neighbor is the N×N block (block c 1113) on the upper side of sub-CU A 1101. If block c (1113) is not available or is intra-coded, check the other N×N blocks on the upper side of sub-CU A (1101) (from left to right, starting from block c 1113). The second neighbor is the block on the left side of sub-CU A1101 (block b1112). If block b (1112) is not available or is intra-coded, check the other blocks on the left side of sub-CU A1101 (from top to bottom, starting from block b 1112). The motion information obtained from the neighboring blocks of each list is scaled to the first reference frame of the given list. Next, the temporal motion vector prediction value (TMVP) of sub-block A1101 is derived by following the same process as the TMVP derivation specified in HEVC. The motion information of the collocated block at block D 1104 is obtained and scaled accordingly. Finally, after retrieving and scaling the motion information, all available motion vectors are averaged separately for each reference list. The average motion vector is designated as the motion vector of the current sub-CU.
[0151] 2.1.3 Example of Sub-CU Motion Prediction Mode Signaling
[0152] In some embodiments, sub-CU modes are enabled as additional Merge candidates, and no additional syntax elements are required to signal these modes. Two additional Merge candidates are added to the Merge candidate list of each CU to represent the ATMVP mode and the STMVP mode. In other embodiments, if the sequence parameter set indicates that ATMVP and STMVP are enabled, up to seven Merge candidates can be used. The encoding logic of the additional Merge candidates is the same as that of the Merge candidates in the HM, which means that for each CU in a P slice or a B slice, two additional Merge candidates may also require two RD checks. In some embodiments, such as JEM, all bins of the Merge index are context-coded by CABAC (Context-based Adaptive Binary Arithmetic Coding). In other embodiments, such as HEVC, only the first bin is context-coded, and the remaining bins are context-bypass coded.
[0153] 2.2 Example of Local Illumination Compensation (LIC) in JEM
[0154] Local Illumination Compensation (LIC) is based on a linear model of illumination variations using a scaling factor a and an offset b, and is adaptively enabled or disabled for each inter-mode coded codec unit (CU).
[0155] When LIC is applied to a CU, the least square method is used to derive parameters a and b by using the neighboring samples of the current CU and its corresponding reference samples. More specifically, Figure 12 As shown, the subsampled (2:1 subsampled) neighboring samples and corresponding samples (identified by the motion information of the current CU or sub-CU) of the CU in the reference picture are used.
[0156] 2.2.1 Derivation of prediction blocks
[0157] IC parameters are derived and applied separately for each prediction direction. For each prediction direction, the decoded motion information is used to generate the first prediction block, and then a temporary prediction block is obtained by applying the LIC model. The two temporary prediction blocks are then used to derive the final prediction block.
[0158] When a CU is encoded or decoded in Merge mode, the LIC flag is copied from the neighboring blocks in a manner similar to the motion information copying in Merge mode; otherwise, the LIC flag is signaled to the CU to indicate whether LIC is applicable.
[0159] When LIC is enabled for a picture, an additional CU-level RD check is required to determine whether LIC is applicable to the CU. When LIC is enabled for a CU, the Mean-Removed Sum of Absolute Difference (MR-SAD) and the Mean-Removed Sum of Absolute Hadamard-Transformed Difference (MR-SATD) are used for integer-pixel motion search and fractional-pixel motion search, respectively, instead of SAD and SATD.
[0160] To reduce coding complexity, the following coding scheme is applied in JEM: When there is no significant illumination change between the current picture and its reference pictures, LIC is disabled for the entire picture. To identify this situation, the encoder calculates the histogram of the current picture and each of its reference pictures. If the histogram difference between the current picture and each of its reference pictures is less than a given threshold, LIC is disabled for the current picture; otherwise, LIC is enabled for the current picture.
[0161] 2.3 Example of inter-frame prediction method in VVC
[0162] There are several new codec tools for inter-frame prediction improvements, such as Adaptive Motion Vector Difference Resolution (AMVR) for signaling MVD, affine prediction mode, triangular prediction mode (TPM), ATMVP, generalized bi-prediction (GBI), and bidirectional optical flow (BIO).
[0163] 2.3.1 Example of Codec Block Structure in VVC
[0164] In VVC, a quadtree / binarytree / multitree (QT / BT / TT) structure is used to divide the picture into square or rectangular blocks. In addition to QT / BT / TT, independent trees (also known as dual codec trees) are also used for I frames in VVC. For independent trees, the codec block structure is signaled separately for luma and chroma components.
[0165] 2.3.2 Example of Adaptive Motion Vector Difference Resolution
[0166] In some embodiments, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the PU's motion vector and the predicted motion vector) is signaled in units of quarter luma samples. In JEM, Local Adaptive Motion Vector Resolution (LAMVR) is introduced. In JEM, MVD can be coded or decoded in units of quarter luma samples, integer luma samples, or four luma samples. The MVD resolution is controlled at the codec unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU with at least one non-zero MVD component.
[0167] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter luma sample MV precision is not used, another flag is signaled to indicate whether integer luma sample MV precision or quad luma sample MV precision is used.
[0168] When the first MVD resolution flag of the CU is zero, or the CU is not coded (meaning that all MVDs in the CU are zero), the CU uses a quarter-luminance sample MV resolution. When the CU uses integer luminance sample MV precision or four-luminance sample MV precision, the MVP in the CU's AMVP candidate list is rounded to the corresponding precision.
[0169] 2.3.3 Example of Affine Motion Compensated Prediction
[0170] In HEVC, only the translation motion model is applied to Motion Compensation Prediction (MCP). However, the camera and the object may have various motions, such as zooming in / out, rotation, perspective motion, and / or other irregular motions. In VVC, a simplified affine transformation motion compensation prediction is applied using a 4-parameter affine model and a 6-parameter affine model. Figure 13A and Figure 13B As shown, the affine motion field of a block is described by two (in a 4-parameter affine model using variables a, b, e and f) or three (in a 6-parameter affine model using variables a, b, c, d, e and f) control point motion vectors, respectively.
[0171] The motion vector field (MVF) of a block is described by the following equations with a 4-parameter affine model and a 6-parameter affine model, respectively:
[0172]
[0173]
[0174] In this article, (mvh 0,mv h 0) is the motion vector of the upper left control point (CP), and (mv h 1, mv h 1) is the motion vector of the upper right control point, and (mv h 2, mv h 2) is the motion vector of the lower left control point, and (x, y) represents the coordinates of the representative point relative to the upper left sample point in the current block. The CP motion vector can be signaled (such as in affine AMVP mode) or derived on the fly (such as in affine Merge mode). w and h are the width and height of the current block. In practice, division is implemented by right shift and rounding operations. In VTM, the representative point is defined as the center position of the sub-block. For example, when the coordinates of the upper left corner of the sub-block relative to the upper left sample point in the current block are (xs, ys), the coordinates of the representative point are defined as (xs+2, ys+2). For each sub-block (for example, 4×4 in VTM), the motion vector of the entire sub-block is derived using the representative point.
[0175] Figure 14 An example of an affine MVF for each sub-block of block 1300 is shown, where a sub-block based affine transform prediction is applied to further simplify motion compensated prediction. To derive the motion vector for each M×N sub-block, the motion vector of the center sample of each sub-block can be calculated according to equations (1) and (2) and rounded to the motion vector fractional accuracy (e.g., 1 / 16 in JEM). A motion compensated interpolation filter can then be applied to generate a prediction for each sub-block with the derived motion vector. The affine mode introduces an interpolation filter of 1 / 16 pixels. After MCP, the high-precision motion vector of each sub-block is rounded and saved to the same accuracy as the standard motion vector.
[0176] 2.3.3.1 Example of Signaling of Affine Prediction
[0177] Similar to the translational motion model, due to affine prediction, there are also two modes for signaling side information. They are AFFINE_INTER and AFFINE_MERGE modes.
[0178] 2.3.3.2 Example of AF_INTER mode
[0179] AF_INTER mode can be applied to CUs with width and height greater than 8. A CU-level affine flag is signaled in the bitstream to indicate whether AF_INTER mode is used.
[0180] In this mode, for each reference picture list (list 0 or list 1), an affine AMVP candidate list is constructed with three types of affine motion predictors in the following order, where each candidate includes the estimated CPMV of the current block. The best CPMV found at the encoder side (such as Figure 17 The difference between mv0 mv1 mv2) and the estimated CPMV is signaled. In addition, the index of the affine AMVP candidate from which the estimated CPMV is derived is further signaled.
[0181] 1) Inherited affine motion prediction value
[0182] The checking order is similar to that of the spatial MVP in the HEVC AMVP list. First, the inherited affine motion prediction value on the left is derived from the first block in {A1, A0} that is affine-coded and has the same reference picture as the current block. Second, the inherited affine motion prediction value on the top is derived from the first block in {B1, B0, B2} that is affine-coded and has the same reference picture as the current block. Figure 16 The five blocks A1, A0, B1, B0, B2 are depicted in FIG.
[0183] Once a neighboring block is found to be coded in affine mode, the CPMV of the codec covering the neighboring block is used to derive the CPMV prediction value of the current block. For example, if A1 is coded in non-affine mode and A0 is coded in 4-parameter affine mode, the inherited affine MV prediction value on the left will be derived from A0. In this case, the CPMV of the CU covering A0 (as in Figure 18B Zhongyou The upper left CPMV represented by The upper right CPMV represented by is used to derive the estimated CPMV of the current block, which is given by Indicates the upper left (coordinates (x0, y0)), upper right (coordinates (x1, y1)), and lower right positions (coordinates (x2, y2)) of the current block.
[0184] 2) Constructed affine motion prediction value
[0185] like Figure 17 As shown, the constructed affine motion prediction value contains the control point motion vector (CPMV) derived from the adjacent inter-frame codec blocks with the same reference picture. If the current affine motion model is 4-parameter affine, the number of CPMVs is 2, otherwise if the current affine motion model is 6-parameter affine, the number of CPMVs is 3. CPMV in the upper left corner Derived from the MV of the first block in group {A, B, C} that is inter-coded and has the same reference picture as the current block. Derived from the MV of the first block in group {D, E} that is inter-coded and has the same reference picture as the current block. Derived from the MV at the first block in group {F, G} that is inter-coded and has the same reference picture as the current block.
[0186] - If the current affine motion model is 4-parameter affine, only if and When both are established, the constructed affine motion prediction value is inserted into the candidate list, that is, and Used as the estimated CPMV of the upper left (coordinates (x0, y0)) and upper right (coordinates (x1, y1)) positions of the current block.
[0187] - If the current affine motion model is 6-parameter affine, only if and When all are established, the constructed affine motion prediction value is inserted into the candidate list, that is, and Used as the estimated CPMV for the top left (coordinates (x0, y0)), top right (coordinates (x1, y1)), and bottom right (coordinates (x2, y2)) positions of the current block.
[0188] When inserting the constructed affine motion predictors into the candidate list, no pruning process is applied.
[0189] 3) Ordinary AMVP motion prediction value
[0190] The following applies until the number of affine motion predictors reaches a maximum.
[0191] 1) If available, by setting all CPMVs equal to To derive the affine motion prediction value.
[0192] 2) If available, by setting all CPMVs equal to To derive the affine motion prediction value.
[0193] 3) If available, by setting all CPMVs equal to To derive the affine motion prediction value.
[0194] 4) If available, derive affine motion prediction values by setting all CPMVs equal to HEVC TMVPs.
[0195] 5) Derive the affine motion prediction value by setting all CPMVs to zero MV.
[0196] Please note that It has been derived in the constructed affine motion predictor.
[0197] In AF_INTER mode, when using 4 / 6 parameter affine mode, 2 / 3 control points are required, so 2 / 3 MVDs need to be encoded and decoded for these control points, such as Figure 15A and Figure 15B In the existing embodiment, MV can be derived as follows, for example, it predicts mvd1 and mvd2 from mvd0.
[0198]
[0199]
[0200]
[0201] In this article, mvd i and mv1 are the predicted motion vector, motion vector difference, and motion vector of the upper left pixel (i=0), upper right pixel (i=1), or lower left pixel (i=2), respectively. Figure 15B In some embodiments, the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the sum of the two components. For example, new MV = mvA + mvB means that the two components of the new MV are set to (xA + xB) and (yA + yB), respectively.
[0202] 2.3.3.3 Example of AF_Merge Mode
[0203] When the CU is applied to the AF_MERGE mode, it obtains the first block encoded and decoded in the affine mode from the valid neighboring reconstructed blocks. And the selection order of the candidate blocks is from the left, top, top right, bottom left to top left, as shown in Figure 18A As shown (indicated by A, B, C, D, E in order). For example, if the adjacent lower left block is Figure 18B If A0 in the block is encoded and decoded in affine mode, the control point (CP) motion vector mv0 of the upper left corner, upper right corner and lower left corner of the adjacent CU / PU containing block A is obtained N 、mv1 N and mv2 N . And based on mv0 N 、mv1 N and mv2 N Calculate the motion vector mv0 of the upper left corner / upper right / lower left corner of the current CU / PUC 、mv1 C and mv2 C (For 6-parameter affine models only.) Note that in VTM-2.0, the upper-left subblock (e.g., a 4×4 block in VTM) stores mv0, and if the current block is affine-encoded, the upper-right subblock stores mv1. If the current block is affine-encoded with a 6-parameter affine model, the lower-left subblock stores mv2; otherwise (with a 4-parameter affine model), LB stores mv2'. The other subblocks store the MV for MC.
[0204] After calculating the CPMV of the current CU v0 and v1 according to the affine motion model in equations (1) and (2), the MVF of the current CU can be generated. In order to identify whether the current CU is coded or decoded in AF_MERGE mode, when there is at least one neighboring block coded or decoded in affine mode, an affine flag can be signaled in the bitstream.
[0205] In some embodiments (e.g., JVET-L0142 and JVET-L0632), the affine merge candidate list can be constructed using the following steps:
[0206] 1) Insert inherited affine candidates
[0207] An inherited affine candidate is one that is derived from the affine motion model of its valid neighboring affine codec blocks. A maximum of two inherited affine candidates are derived from the affine motion models of the neighboring blocks and inserted into the candidate list. For the left predictor, the scan order is {A0, A1}; for the top predictor, the scan order is {B0, B1, B2}.
[0208] 2) Insert the constructed affine candidate
[0209] If the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (set to 5 in this paper), the constructed affine candidate is inserted into the candidate list. The constructed affine candidate refers to a candidate constructed by combining the neighboring motion information of each control point.
[0210] a) The motion information of the control point is first obtained from Figure 19 The spatial and temporal neighbors are specified for derivation. CPk (k = 1, 2, 3, 4) represents the kth control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial locations used to predict CPk (k = 1, 2, 3); T is the temporal location used to predict CP4.
[0211] The coordinates of CP1, CP2, CP3 and CP4 are (0, 0), (W, 0), (H, 0) and (W, H), respectively, where W and H are the width and height of the current block.
[0212] The motion information for each control point is obtained according to the following priority order:
[0213] For CP1, the priority is B2 → B3 → A2. If B2 is available, use B2. Otherwise, if B2 is available, use B3. If neither B2 nor B3 is available, use A2. If all three candidates are unavailable, motion information for CP1 cannot be obtained.
[0214] For CP2, the inspection priority is B1→B0.
[0215] For CP3, the inspection priority is A1→A0.
[0216] For CP4, use T.
[0217] b) Secondly, the combination of control points is used to construct affine merge candidates.
[0218] I. Motion information of three control points is required to construct a 6-parameter affine candidate. The three control points can be selected from one of the following four combinations: {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}. The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, and {CP1, CP3, CP4} will be converted into a 6-parameter motion model represented by the top left, top right, and bottom left control points.
[0219] II. Motion information of two control points is required to construct a 4-parameter affine candidate. The two control points can be selected from one of the following six combinations: {CP1, CP4}, {CP2, CP3}, {CP1, CP2}, {CP2, CP4}, {CP1, CP3}, {CP3, CP4}. The combination {CP1, CP4}, {CP2, CP3}, {CP2, CP4}, {CP1, CP3}, {CP3, CP4} will be converted into a 4-parameter motion model represented by the top left and top right control points.
[0220] III. The constructed combinations of affine candidates are inserted into the candidate list in the following order:
[0221] {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4}, {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3}, {CP2, CP3}, {CP1, CP4}, {CP2, CP4}, {CP3, CP4}
[0222] i. For a combined reference list X (X is 0 or 1), the most used reference index in the control point is selected as the reference index of list X, and the motion vector pointing to the difference reference picture is scaled.
[0223] After a candidate is derived, a complete pruning process is performed to check whether the same candidate has been inserted into the list. If there is an identical candidate, the derived candidate is discarded.
[0224] 3) Fill with zero motion vectors
[0225] If the number of candidates in the affine merge candidate list is less than 5, a zero motion vector with a zero reference index is inserted into the candidate list until the list is full.
[0226] More specifically, for the sub-block Merge candidate list, MV is set to (0, 0) and the prediction direction is set to the 4-parameter Merge candidate of unidirectional prediction (for P slices) and bidirectional prediction (for B slices) from list 0.
[0227] 2.3.4 Example of Merge with Motion Vector Difference (MMVD)
[0228] In JVET-L0054, the Ultimate Motion Vector Expression (UMVE, also known as MMVD) is given. UMVE and a proposed motion vector expression method are used in Skip or Merge mode.
[0229] UMVE reuses the same Merge candidates as those included in the conventional Merge candidate list in VVC. Among these Merge candidates, a basic candidate can be selected and further extended by the proposed motion vector expression method.
[0230] UMVE provides a new method for representing motion vector difference (MVD), in which the MVD is represented by the starting point, motion amplitude and motion direction.
[0231] The proposed technique uses the Merge candidate list as is, but only candidates of the default Merge type (MRG_TYPE_DEFAULT_N) are considered for UMVE expansion.
[0232] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list, as shown below.
[0233] Table 1: Basic candidate IDX
[0234] Basic Candidate IDX 0 1 2 3 Nth MVP First MVP Second MVP Third MVP Fourth MVP
[0235] If the number of basic candidates is equal to 1, the basic candidate IDX is not signaled.
[0236] The distance index is the motion magnitude information. The distance index indicates the predefined distance from the starting point information. The predefined distances are as follows:
[0237] Table 2: Distance IDX
[0238] Distance from IDX 0 1 2 3 4 5 6 7 Pixel distance 1 / 4 pixel 1 / 2 pixel 1 pixel 2 pixels 4 pixels 8 pixels 16 pixels 32 pixels
[0239] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions as shown below.
[0240] Table 3: Direction IDX
[0241]
[0242]
[0243] In some embodiments, the UMVE flag is signaled immediately after the Skip flag or Merge flag is transmitted. If the Skip or Merge flag is true, the UMVE flag is parsed. If the UMVE flag is equal to 1, the UMVE syntax is parsed. If not, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, AFFINE mode is used. If not, the Skip / Merge index is parsed as the VTM's Skip / Merge mode.
[0244] No additional line buffer is required for UMVE candidates because the software's skip / merge candidates are used directly as base candidates. The MV complement is determined immediately before motion compensation using the input UMVE index. There's no need to maintain a long line buffer for this purpose.
[0245] Under the current common test conditions, the first or second merge candidate in the merge candidate list can be selected as the base candidate.
[0246] 2.3.5 Example of Decoder-Side Motion Vector Refinement (DMVR)
[0247] In bidirectional prediction, for the prediction of a block region, two prediction blocks formed using motion vectors (MVs) from list 0 and MVs from list 1, respectively, are combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two motion vectors of the bidirectional prediction are further refined.
[0248] In the JEM design, motion vectors are refined through a bilateral template matching process. Bilateral template matching is applied in the decoder to perform a distortion-based search between the bilateral template and the reconstructed samples in the reference picture to obtain refined MV values without the transmission of additional motion information. Figure 22 An example is depicted in . The bilateral template is generated as a weighted combination (i.e., average) of two prediction blocks from the initial MV0 of list 0 and the initial MV1 of list 1, respectively, as Figure 22 As shown. The template matching operation involves calculating the cost metric between the generated template and the sample area in the reference picture (around the initial prediction block). For each of the two reference pictures, the MV that produces the minimum template cost is considered to be the updated MV of the list to replace the original MV. In JEM, nine MV candidates are searched for each list. The nine MV candidates include the original MV and 8 surrounding MVs that are offset by one luminance sample from the original MV in the horizontal direction, vertical direction, or both directions. Finally, as Figure 22 The two new MVs shown (i.e., MV0' and MV1') are used to generate the final bidirectional prediction result. The sum of absolute differences (SAD) is used as the cost metric. Note that when calculating the cost of a prediction block generated by a surrounding MV, the MV rounded to integer pixels is actually used to obtain the prediction block, rather than the actual MV.
[0249] To further simplify the DMVR process, JVET-M0147 proposes several changes to the JEM design. More specifically, the DMVR design adopted by VTM-4.0 (to be released soon) has the following key features:
[0250] o Early termination w / (0,0) position SAD between list 0 and list 1
[0251] οDMVR block size, W*H>=64&&H>=8
[0252] ο Divide the CU into multiple DMVR 16×16 sub-blocks with CU size > 16*16
[0253] o Reference block size (W+7)*(H+7) (for luma)
[0254] o 25-point SAD-based integer pixel search (i.e., (+-2) refined search range, single stage)
[0255] οDMVR based on bilinear interpolation
[0256] ο MVD mirroring between list 0 and list 1, allowing bidirectional matching
[0257] οSub-pixel refinement based on “parametric error surface equation”
[0258] Luma / Chroma MC with reference block padding (if needed)
[0259] ο MV for refinement of MC and TMVP only
[0260] 2.3.6 Example of Combined Intra and Inter Prediction (CIIR)
[0261] In JVET-L0100, multi-hypothesis prediction is proposed, where combined intra-frame and inter-frame prediction is a way to generate multiple hypotheses.
[0262] When multi-hypothesis prediction is applied to improve intra mode, multi-hypothesis prediction combines an intra prediction and a Merge index prediction. In a Merge CU, when the flag is true, a flag is signaled for the Merge mode to select the intra mode from the intra candidate list. For the luma component, the intra candidate list is derived from 4 intra prediction modes including DC, planar, horizontal and vertical modes, and the size of the intra candidate list can be 3 or 4 depending on the block shape. When the CU width is greater than twice the CU height, the horizontal mode is not included in the intra mode list, and when the CU height is greater than twice the CU width, the vertical mode is removed from the intra mode list. Weighted averaging is used to combine one intra prediction mode selected by the intra mode index and one merge index prediction selected by the merge index. For chroma components, DM is always applied without additional signaling. The weights used for combined prediction are described as follows. When DC or planar mode is selected, or the CB width or height is less than 4, equal weights are applied. For those CBs whose width and height are greater than or equal to 4, when horizontal / vertical mode is selected, one CB is first divided vertically / horizontally into four equal-area regions. Each weight set (denoted as (w_intra i ,w_inter i ), where i is from 1 to 4 and (w_intra1, w_inter1) = (6, 2), (w_intra2, w_inter2) = (5, 3), (w_intra3, w_inter3) = (3, 5) and (w_intra4, w_inter4) = (2, 6)) will be applied to the corresponding areas. (w_intra1, w_inter1) is used for the area closest to the reference sample, while (w_intra4, w_inter4) is used for the area farthest from the reference sample. The combined prediction can then be calculated by adding the two weighted predictions and shifting them right by 3 bits. In addition, the intra prediction mode of the intra hypothesis of the predicted value can be saved for subsequent reference by neighboring CUs.
[0263] 2.4 Inter-Loop Reshaping (ILR) in JVET-M 0427
[0264] The basic idea of in-loop reshaping (ILR) is to convert the original (in the first domain) signal (prediction / reconstruction signal) into the second domain (the shaped domain).
[0265] The loop luminance shaper is implemented as a pair of look-up tables (LUTs), but only one of the two LUTs needs to be signaled because the other LUT can be calculated from the signaled LUT. Each LUT is a one-dimensional, 10-bit, 1024-entry mapping table (1D-LUT). One LUT is the forward LUT, FwdLUT, which takes the input luminance code value Y i Mapped to the change value Y r :Y r =FwdLUT[Y i ]. The other LUT is the inverse LUT, InvLUT, which changes the code value Y r Map to ( Represents Y i reconstructed value).
[0266] 2.4.1 Piecewise Linear (PWL) Model
[0267] In some embodiments, piecewise linear (PWL) is implemented as follows:
[0268] Let x1, x2 be the two input pivot points, and y1, y2 be their corresponding output pivot points for a piece. The output value y for any input value x between x1 and x2 can be interpolated by the following equation:
[0269] y=((y2-y1) / (x2-x1))*(x-x1)+y1
[0270] In a fixed-point implementation, the equation can be rewritten as:
[0271] y=((m*x+2FP_PREC-1)>>FP_PREC)+c
[0272] Where m is a scalar, c is an offset, and FP_PREC is a constant value specifying the precision.
[0273] Note that in the CE-12 software, the PWL model is used to pre-compute the 1024-entry FwdLUT and InvLUT mapping tables; however, the PWL model also allows implementations that compute the same mapping values on the fly without pre-computing the LUTs.
[0274] 2.4.2 Test CE12-2
[0275] 2.4.2.1 Brightness Shaping
[0276] Test 2 of in-loop luma shaping (ie, proposed CE12-2) provides a lower complexity pipeline that also eliminates the decoding delay of block-wise intra prediction in inter slice reconstruction. Intra prediction is performed in the shaping domain for both inter and intra slices.
[0277] Regardless of the slice type, intra prediction is always performed in the shaped domain. In this arrangement, intra prediction can start immediately after the previous TU reconstruction is completed. This arrangement can also provide a unified process for intra mode instead of being slice-dependent. Figure 23 A block diagram of the mode-based CE12-2 decoding process is shown.
[0278] CE12-2 also tests a 16-segment piecewise linear (PWL) model for luma and chroma residual scaling, instead of the 32-segment PWL model of CE12-1.
[0279] Inter-strip reconstruction using the in-loop luma shaper in CE12-2 (light green shaded blocks indicate signals in the shaping domain: luma residual; intra-luma predicted; and intra-luma reconstructed).
[0280] 2.4.2.2 Luma-dependent Chroma Residue Scaling (LCRS)
[0281] Luma-dependent chroma residual scaling is a multiplication process implemented using fixed-point integer arithmetic. Chroma residual scaling compensates for the interaction between luma and chroma signals. Chroma residual scaling is applied at the TU level. More specifically, the following applies:
[0282] o For intra-frame, the reconstructed luminance is averaged.
[0283] οFor inter-frames, the predicted luminance is averaged.
[0284] The average is used to identify the index in the PWL model. This index identifies the scaling factor cScaleInv. The chroma residual is multiplied by this number.
[0285] Note that the chroma scaling factors are computed based on the predicted luma values from the forward map rather than the reconstructed luma values.
[0286] 2.4.2.3 Signaling Notification of ILR Side Information
[0287] Parameters are (currently) sent in a slice group header (similar to ALF). These are reported to require 40-100 bits. Slice groups can be another way to represent pictures. The following table is based on version 9 of JVET-L1001. Added syntax is highlighted in italics.
[0288] 7.3.2.1 Sequence Parameter Set RBSP Syntax
[0289]
[0290]
[0291] 7.3.3.1 Generalized slice group header syntax
[0292]
[0293] Add new syntax slice group shaper model:
[0294]
[0295] In the generalized sequence parameter set RBSP semantics, the following semantics are added:
[0296] sps_reshaper_enabled_flag equal to 1 specifies that the reshaper is used in the Coded Video Sequence (CVS). sps_reshaper_enabled_flag equal to 0 specifies that the reshaper is not used in the CVS.
[0297] In the slice group header syntax, add the following semantics
[0298] tile_group_reshaper_model_present_flag equal to 1 specifies that tile_group_reshaper_model() is present in the slice group header. tile_group_reshaper_model_present_flag equal to 0 specifies that tile_group_reshaper_model() is not present in the slice group header. When tile_group_reshaper_model_present_flag is not present, it is inferred to be equal to 0.
[0299] tile_group_reshaper_enabled_flag equal to 1 specifies that the reshaper is enabled for the current slice group. tile_group_reshaper_enabled_flag equal to 0 specifies that the reshaper is not enabled for the current slice group. When tile_group_reshaper_enable_flag is not present, it is inferred to be equal to 0.
[0300] tile_group_reshaper_chroma_residual_scale_flag equal to specifies that chroma residual scaling is enabled for the current slice group. tile_group_reshaper_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling is not enabled for the current slice group. When tile_group_reshaper_chroma_residual_scale_flag is not present, it is inferred to be equal to 0.
[0301] Add tile_group_reshaper_model() syntax
[0302] reshape_model_min_bin_idx specifies the minimum bin (or segment) index to be used during reshape construction. The value of reshape_model_min_bin_idx should be in the range of 0 to MaxBinIdx, inclusive. The value of MaxBinIdx should be equal to 15.
[0303] reshape_model_delta_max_bin_idx specifies the maximum allowed bin (or segment) index MaxBinIdx minus the maximum bin index to be used during reshaper construction. The value of reshape_model_max_bin_idx is set equal to MaxBinIdx – reshape_model_delta_max_bin_idx.
[0304] reshaper_model_bin_delta_abs_cw_prec_minus1 plus 1 specifies the number of bits used to represent the syntax reshape_model_bin_delta_abs_CW[i].
[0305] reshape_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value for the i-th bin.
[0306] reshaper_model_bin_delta_sign_CW_flag[i] specifies the sign of reshape_model_bin_delta_abs_CW[i] as follows:
[0307] –If reshape_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is positive.
[0308] –Otherwise (reshape_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] is negative.
[0309] When reshape_model_bin_delta_sign_CW_flag[i] is not present, it is inferred to be equal to 0.
[0310] Variable RspDeltaCW[i]=(1 2*reshape_model_bin_delta_sign_CW[i])*reshape_model_bin_delta_abs_CW[i];
[0311] The variable RspCW[i] is derived as follows:
[0312] The variable OrgCW is set equal to (1 < <BitDepth Y ) / (MaxBinIdx+1).
[0313] – If reshaper_model_min_bin_idx<=i<=reshaper_model_max_bin_idx
[0314] RspCW[i]=OrgCW+RspDeltaCW[i].
[0315] –Otherwise, RspCW[i]=0.
[0316] If BitDepth Y If the value of is equal to 10, the value of RspCW[i] should be in the range of 32 to 2*OrgCW-1.
[0317] The variable InputPivot[i] (i is in the range of 0 to MaxBinIdx+1, inclusive) is derived as follows
[0318] InputPivot[i]=i*OrgCW
[0319] The variables ReshapePivot[i] (i is in the range of 0 to MaxBinIdx+1, inclusive), ScaleCoef[i] and InvScaleCoeff[i] (i is in the range of 0 to MaxBinIdx, inclusive) are derived as follows:
[0320]
[0321] The variable ChromaScaleCoef[i] (i is in the range of 0 to MaxBinIdx, inclusive) is derived as follows:
[0322] ChromaResidualScaleLut
[64] ={16384,16384,16384,16384,16384,16384,16384,8192,8192,8192,819 2,5461,5461,5461,5461,4096,4096,4096,4096,3277,3277,3277,3277,2731,2731,2731,2731,2341,23 41,2341,2048,2048,2048,1820,1820,1820,1638,1638,1638,1489,1489,1489,1365,1365,1365,1365,1260,1260,1260,1260,1170,1170,1170,1092,1092,1092,1024,1024,1024,1024}
[0323] shiftC=11
[0324] –if(RspCW[i]==0)
[0325] ChromaScaleCoef[i]=(1< <shiftC)
[0326] – Otherwise (RspCW[i]!=0), ChromaScaleCoef[i]=ChromaResidualScaleLut[RspCW[i]>>1]
[0327] 2.4.2.4 Use of ILR
[0328] On the encoder side, each picture (or slice group) is first converted to the shaped domain. All encoding and decoding processes are performed in the shaped domain. For intra prediction, neighboring blocks are in the shaped domain; for inter prediction, reference blocks (generated from the original domain from the decoded picture buffer) are first converted to the shaped domain. The residual is then generated and encoded and decoded into the bitstream.
[0329] After the entire picture (or slice group) is encoded / decoded, the samples in the shaped domain are converted to the original domain and then the deblocking filter and other filters are applied.
[0330] Disable forward shaping of the prediction signal for the following cases:
[0331] o The current block is intra-coded o The current block is coded as CPR (Current Picture Referencing, also known as Intra Block Copy, IBC)
[0332] o The current block is coded in Combined Inter-Intra mode (CIIP) and forward shaping is disabled for intra-predicted blocks
[0333] JVET-N0805
[0334] In JVET-N0805, in order to avoid signaling ILR side information in the slice group header, it is proposed to signal them in the APS. It includes the following main ideas:
[0335] – Optionally, LMCS parameters may be transmitted in the SPS. LMCS refers to the Luma Mapping With Chroma Scaling (LMCS) technology defined in relevant video codec standards.
[0336] – Defines APS type for ALF and LMCS parameters. Each APS has only one type.
[0337] –Transmit LMCS parameters in APS
[0338] – If LMCS tool is enabled, set a flag in TGH to indicate whether LMCS aps_id is present. If not signalled, use SPS parameters.
[0339] *Semantic constraints need to be added so that there is always a valid thing being referenced when the tool is enabled.
[0340] 2.5.2.5.1 Implementation of the Recommended Design in JVET-M1001 (VVC Working Draft 4)
[0341] Below, suggested changes are shown in italics.
[0342]
[0343] ...
[0345] sps_lmcs_enabled_flag equal to 1 specifies that luma mapping and chroma scaling are used in the Coded Video Sequence (CVS). sps_lmcs_enabled_flag equal to 0 specifies that luma mapping and chroma scaling are not used in the CVS.
[0346] sps_lmcs_default_model_present_flag equal to 1 specifies that default LMCS data is present in this SPS. sps_lmcs_default_model_flag equal to 0 specifies that default LMCS data is not present in this SPS. When not present, the value of sps_lmcs_default_model_present_flag is inferred to be equal to 0. ...
[0348]
[0349] aps_params_type specifies the type of APS parameters carried in the APS, as specified in the following table:
[0350] Table 7-x – APS parameter type codes and APS parameter types
[0351]
[0352]
[0353] Add the following definition to Clause 3:
[0354] ALF APS: APS with aps_params_type equal to ALF_APS.
[0355] LMCS APS: APS with aps_params_type equal to LMCS_APS.
[0356] Make the following semantic changes: ...
[0358] tile_group_alf_aps_id specifies the adaptation_parameter_set_id of the ALF APS referenced by the slice group. The TemporalId of the ALF APS NAL unit with adaptation_parameter_set_id equal to tile_group_alf_aps_id should be less than or equal to the TemporalId of the NAL unit of the codec slice group.
[0359] When multiple ALF APSs having the same adaptation_parameter_set_id value are referenced by two or more slice groups of the same picture, the multiple ALF APSs having the same adaptation_parameter_set_id value should have the same content. ...
[0361]
[0362]
[0363] tile_group_lmcs_enabled_flag equal to 1 specifies that luma mapping and chroma scaling are enabled for the current slice group. tile_group_lmcs_enabled_flag equal to 0 specifies that luma mapping and chroma scaling are not enabled for the current slice group. When tile_group_lmcs_enable_flag is not present, it is inferred to be equal to 0.
[0364] tile_group_lmcs_use_default_model_flag equal to 1 specifies that the default lmcs model is used for luma mapping and chroma scaling for the slice group. tile_group_lmcs_use_default_model_flag equal to 0 specifies that the lmcs model in the LMCS APS referenced by tile_group_lmcs_aps_id is used for luma mapping and chroma scaling for the slice group. When tile_group_reshaper_use_default_model_flag is not present, it is inferred to be equal to 0.
[0365] tile_group_lmcs_aps_id specifies the adaptation_parameter_set_id of the LMCS APS referenced by the slice group. The TemporalId of the LMCS APS NAL unit with adaptation_parameter_set_id equal to tile_group_lmcs_aps_id should be less than or equal to the TemporalId of the codec slice group NAL unit.
[0366] When multiple LMCS APSs having the same adaptation_parameter_set_id value are referenced by two or more slice groups of the same picture, the multiple LMCS APSs having the same adaptation_parameter_set_id value should have the same content.
[0367] tile_group_chroma_residual_scale_flag equal to 1 specifies that chroma residual scaling is enabled for the current slice group. tile_group_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling is not enabled for the current slice group. When tile_group_chroma_residual_scale_flag is not present, it is inferred to be equal to 0. ...
[0369] 2.4.2.6 JVET-N0138
[0370] This contribution proposes the extended use of the Adaptive Parameter Set (APS) to carry shaper model parameters as well as ALF parameters. In a recent meeting, it was decided that the APS should carry the ALF parameters instead of the slice group header to improve codec efficiency by avoiding unnecessary redundant signaling of parameters across multiple slice groups. For the same reason, it was proposed to use the APS instead of the slice group header to carry the shaper model parameters. To identify the parameter type in the APS (at least whether it is the ALF or the shaper model), the APS type information and the APS ID are required in the APS syntax.
[0371] Adaptive Parameter Set Syntax and Semantics
[0372] Below, suggested changes are shown in italics.
[0373]
[0374]
[0375] adaptation_parameter_set_type identifies the parameter type in the APS. The value of adaptation_parameter_set_type shall be in the range of 0 to 1, inclusive. If adaptation_parameter_set_type is equal to 0, ALF parameters are signaled. Otherwise, shaper model parameters are signaled.
[0376] Generalized slice group header syntax and semantics
[0377]
[0378] 2.5 Virtual Pipeline Data Unit
[0379] A Virtual Pipeline Data Unit (VPDU) is defined as a non-overlapping MxM-luma (L) / NxN-chroma (C) unit in a picture. In a hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages; different stages process different VPDUs simultaneously. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so it is said that keeping the VPDU size small is important. In the HEVC hardware decoder, the VPDU size is set to the maximum transform block (TB) size. Increasing the maximum TB size from 32×32-L / 16×16-C (as in HEVC) to 64×64-L / 32×32-C (as in the current VVC) can bring codec gains, and is expected to result in a 4x smaller VPDU size (64×64-L / 32×32-C) compared to HEVC. However, in addition to the quadtree (QT) codec unit (CU) partitioning, ternary tree (TT) and binary tree (BT) are also adopted in VVC to achieve additional codec gains, and TT and BT partitioning can be recursively applied to the 128×128-L / 64×64-C coding tree block (CTU), which is said to result in a 16-fold VPDU size (128×128-L / 64×64-C) compared to HEVC.
[0380] In the current design of VVC, the VPDU size is defined as 64×64-L / 32×32-C.
[0381] 2.6 Adaptive Parameter Set
[0382] Adaptive Parameter Set (APS) is used in VVC to carry ALF parameters. The slice group header contains aps_id, which is conditionally present when ALF is enabled. The APS contains aps_id and ALF parameters. (From JVET-M0132) New NUT (NAL unit type, as in AVC and HEVC) values are assigned to APS. For common test conditions in VTM-4.0 (coming soon), it is recommended to only use aps_id=0 and transmit the APS with each picture. Currently, the range of APS ID values will be 0..31 and APS can be shared across pictures (and can be different in different slice groups within a picture). When present, the ID value should be fixed length encoded. ID values cannot be reused for different content within the same picture.
[0383] 2.7 Related Tools
[0384] 2.7.1 Diffusion filter (DF)
[0385] In JVET-L0157, a diffusion filter is proposed, where the intra / inter prediction signal of a CU can be further modified by a diffusion filter.
[0386] Uniform diffusion filter The uniform diffusion filter is implemented by convolving the prediction signal with a fixed mask, which can be given as h I or h IV , defined as follows.
[0387] In addition to the prediction signal itself, a row of reconstructed samples to the left and above the block, which can be avoided on inter blocks, is used as input to the filtered signal.
[0388] Let pred be the prediction signal on a given block obtained by intra-frame or motion compensated prediction. In order to process the boundary points of the filter, the prediction signal needs to be expanded to the prediction signal pred ext Such extended forecasts can be formed in two ways:
[0389] Alternatively, as an intermediate step, a row of reconstructed samples from the left and top sides of the block is added to the prediction signal, and the resulting signal is then mirrored in all directions. Alternatively, only the prediction signal itself is mirrored in all directions. The latter extension is used for inter-frame blocks. In this case, only the prediction signal itself includes the extended prediction signal pred ext input.
[0390] If you want to use the filter h I , it is proposed to replace the prediction signal pred with the following:
[0391] hI *pred,
[0392] Use the above boundary extension. Here, the filter mask h I is given as:
[0393]
[0394] If you want to use the filter h IV , it is proposed to replace the prediction signal pred with the following:
[0395] h IV *pred.
[0396] Here, the filter h IV is given as:
[0397] h IV =h I *h I *h I *h I .
[0398] Directed diffusion filter Instead of using a signal-adaptive diffusion filter, a directional filter, a horizontal filter h, which still has a fixed mask, is used. hor and vertical filter h ver More precisely, the mask h corresponding to the previous part I The uniform diffusion filtering is simply restricted to be applied only in the vertical direction or only in the horizontal direction. The vertical filter is implemented by applying a fixed filter mask to the prediction signal as follows,
[0399]
[0400] And the horizontal filter is obtained by using the transposed mask And realize.
[0401] 2.7.2 Bilateral Filter (BF)
[0402] The bilateral filter was proposed in JVET-L0406 and is always applied to luma blocks with non-zero transform coefficients and a slice quantization parameter greater than 17. Therefore, no signaling is required to indicate the use of the bilateral filter. If a bilateral filter is applied, it is performed on the decoded samples immediately after the inverse transform. Furthermore, the filter parameters (i.e., weights) are explicitly derived from codec information.
[0403] The filtering process is defined as:
[0404]
[0405] Here, P 0,0is the intensity of the current sample point, P′ 0,0 is the correction strength of the current sample point, P k,0 and W k are the intensity and weighting parameters of the kth neighboring sample point respectively. Figure 24 An example of a current sample point and its four neighboring sample points (ie, K=4) is depicted in FIG.
[0406] More specifically, the weight W associated with the kth neighboring sample point k (x) is defined as follows:
[0407] W k (x) = Distance k ×Range k (x). (2)
[0408] In this article,
[0409] and
[0410] Here, σ d Depends on the codec mode and codec block size. When the TU is further divided, the described filtering process is applied to the intra-frame codec block and the inter-frame codec block to achieve parallel processing.
[0411] In order to better capture the statistical characteristics of the video signal and improve the performance of the filter, the weight function obtained by equation (2) is replaced by σ d The parameters (depending on the codec mode and the parameters of the block partition (minimum size), as listed in Table 4) are adjusted.
[0412] Table 4: σ for different block sizes and codec modes d Value
[0413] Min(block width, block height) Intra-frame mode Interframe mode 4 82 62 8 72 52 other 52 32
[0414] In order to further improve the coding performance, for inter-frame coding blocks when TU is not divided, the intensity difference between the current sample and one of its neighboring samples is replaced by the representative intensity difference between two windows covering the current sample and the neighboring samples. Therefore, the equation of the filtering process is modified to:
[0415]
[0416] Here, P k,m and P 0,m Respectively expressed in P k,0 and P 0,0 The mth sample value in the window centered at . In this proposal, the window size is set to 3×3. Figure 25The coverage P is depicted in 2,0 and P 0,0 Example of two windows.
[0417] 2.7.3 Hadamard Transform Domain Filter (HF)
[0418] In JVET-K0068, a loop filter in the 1D Hadamard transform domain is applied at the CU level after reconstruction with a multiplication-free implementation. The proposed filter is applied to all CU blocks that meet predefined conditions, and the filter parameters are derived from codec information.
[0419] The proposed filter is always applied to luma reconstructed blocks with non-zero transform coefficients, excluding 4x4 blocks, and if the slice quantization parameter is greater than 17. The filter parameters are explicitly derived from codec information. If the proposed filter is applied, it is performed on the decoded samples immediately after the inverse transform.
[0420] For each pixel from the reconstructed block, pixel processing consists of the following steps:
[0421] o Scan the 4 neighboring pixels around the processing pixel including the current pixel according to the scanning pattern
[0422] οRead the 4-point Hadamard transform of the pixel
[0423] ο Spectral filtering based on the following formula:
[0424]
[0425] Here, (i) is the index of the spectral component in the Hadamard spectrum, R(i) is the spectral component of the reconstructed pixel corresponding to the index, and σ is the filtering parameter derived from the codec quantization parameter QP using the following equation:
[0426] σ=2 (1+0.126*(QP-27)) .
[0427] Examples of scan patterns are shown in Figure 26 , where A is the current pixel and {B, C, D} are the surrounding pixels.
[0428] For pixels located on a CU boundary, the scanning pattern is adjusted to ensure that all required pixels are within the current CU.
[0429] 3 Disadvantages of existing implementation methods
[0430] In existing ILR implementations, the following disadvantages may exist:
[0431] 1) Sending ILR side information in the slice group header is inappropriate because it requires too many bits. In addition, prediction between different pictures / slice groups is not allowed. Therefore, for each slice group, ILR side information needs to be sent, which may cause codec loss at low bit rates, especially for low resolutions.
[0432] 2) The interaction between ILR and DMVR (or other newly introduced codec tools) is unclear. For example, ILR is applied to the inter-frame prediction signal to convert the original signal to the shaped domain, and the decoded residual is in the shaped domain. DMVR also relies on the prediction signal to refine the motion vector of a block. Whether DMVR is applied in the original or shaped domain is unclear.
[0433] 3) The interaction between ILR and screen content codec tools (e.g., palette, B-DPCM, IBC, transform skip, transquant-bypass, I-PCM mode) is unclear.
[0434] 4) Luma-dependent chroma residual scaling for ILR. Therefore, additional latency is introduced (due to the dependency between luma and chroma), which is detrimental to hardware design.
[0435] 5) The goal of VPDU is to ensure that the processing of one 64×64 square area is completed before starting to process other 64×64 square areas. However, according to the design of ILR, there is no restriction on the use of ILR, which may lead to violations of VPDU because chrominance depends on the prediction signal of luma.
[0436] 6) When all zero coefficients appear in a CU, the prediction block and the reconstructed block still perform the forward and inverse shaping processes, which wastes computational complexity.
[0437] 7) JVET-N0138 proposes signaling ILR information in the APS. This solution may introduce several new issues. For example, two APSs are designed. However, the adaptation_parameter_set_id signaled for ILR can reference an APS that does not contain ILR information. Similarly, the adaptation_parameter_set_id signaled for adaptive loop filtering (ALF) can reference an APS that does not contain ALF information.
[0438] 4 Example Methods for Loop Shaping in Video Codecs
[0439] Embodiments of the presently disclosed technology overcome the shortcomings of existing implementations, thereby providing a video codec with higher codec efficiency. The loop shaping method based on the disclosed technology can enhance both existing and future video codec standards, as illustrated in the following examples described with respect to various implementations. The examples of the disclosed technology provided below illustrate the general concepts and are not meant to be construed as limiting. In the examples, the various features described in these examples can be combined unless explicitly indicated to the contrary. It should be noted that some of the proposed techniques can be applied to existing candidate list construction processes.
[0440] In this document, Decoder Side Motion Vector Derivation (DMVD) includes methods such as DMVR and FRUC that perform motion estimation to derive or refine block / subblock motion information, and BIO that performs sample-wise motion refinement. Various examples (Examples 1 to 42) are provided in the numbered list below.
[0441] 1. The motion information refinement process in DMVD techniques, such as DMVR, can depend on information in the shaping domain.
[0442] a. In one example, the prediction block generated from the reference picture in the original domain can be first converted to the reshaped domain before being used for motion information refinement.
[0443] i. Additionally, alternatively, cost calculation (eg, SAD, MR-SAD) / gradient calculation is performed in the shaped domain.
[0444] ii. Furthermore, alternatively, after the motion information is refined, the shaping process is disabled for the prediction block generated using the refined motion information.
[0445] b. Alternatively, the motion information refinement process in DMVD techniques, such as DMVR, can depend on information in the original domain.
[0446] i. The DMVD process can be called with the prediction block in the original domain.
[0447] ii. In one example, after the motion information is refined, the prediction block obtained using the refined motion information or the final prediction block (eg, a weighted average of two prediction blocks) may be further converted to the shaping domain to generate a final reconstructed block.
[0448] iii. Furthermore, alternatively, after the motion information is refined, the shaping process is disabled for the prediction block generated using the refined motion information.
[0449] 2. It is proposed to align the domain of samples in the current slice / slice group / picture with the domain of samples derived from the reference picture used to derive local illumination compensation (LIC) parameters (either in the original domain or in the shaped domain).
[0450] a. In one example, the shaping domain is used to derive LIC parameters.
[0451] i. Alternatively, samples (e.g., reference samples in a reference picture (with or without interpolation) and neighboring / non-neighboring samples of the reference sample (with or without interpolation)) may first be converted to the shaping domain before being used to derive LIC parameters.
[0452] b. In one example, the original domain is used to derive LIC parameters.
[0453] i. Additionally, alternatively, spatially neighboring / non-neighboring samples of the current block (eg, in the current slice group / picture / slice) may first be converted to the original domain before being used to derive LIC parameters.
[0454] c. It is proposed that when the LIC parameters are derived in one domain, the same domain of the prediction block should be used when applying the LIC parameters to the prediction block.
[0455] i. In one example, when invoking bullet a., the reference block may be converted to the shaped domain, and the LIC model is applied to the shaped reference block.
[0456] ii. In one example, when calling bullet b., the reference block is kept in the original domain, and the LIC model is applied to the reference block in the original domain.
[0457] d. In one example, the LIC model is applied to the prediction block in the shaped domain (eg, the prediction block is first converted to the shaped domain via forward shaping).
[0458] e. In one example, the LIC model is first applied to the prediction block in the original domain, after which the final prediction block that depends on the prediction block to which the LIC is applied can then be converted to the shaped domain (e.g., via forward shaping) and used to derive the reconstructed block.
[0459] f. The above method can be extended to other codecs that rely on both spatially adjacent / non-adjacent samples and reference samples in reference pictures.
[0460] 3. For filters applied to the prediction signal, such as the diffusion filter (DF), the filter is applied to the prediction block in the original domain.
[0461] a. Additionally, alternatively, shaping is then applied to the filtered prediction signal to generate a reconstructed block.
[0462] b. Figure 27 An example of a process for inter-frame coding and decoding is depicted in .
[0463] c. Alternatively, a filter is applied to the prediction signal in the shaped domain.
[0464] i. Furthermore, alternatively, shaping is first applied to the prediction block; thereafter, a filtering method may be further applied to the shaped prediction block to generate a reconstructed block.
[0465] ii. Figure 28 An example of a process for inter-frame coding and decoding is depicted in .
[0466] d. Filter parameters may depend on whether ILR is enabled.
[0467] 4. For filters applied to the reconstructed block (eg, bilateral filter (BF), Hadamard transform domain filter (HF)), the filter is applied to the reconstructed block in the original domain rather than the reshaped domain.
[0468] a. Furthermore, alternatively, the reconstructed block in the shaped domain is first converted to the original domain, after which the filter can be applied and used to generate the reconstructed block.
[0469] b. Figure 29 An example of a process for inter-frame coding and decoding is depicted in .
[0470] c. Alternatively, a filter can be applied to the reconstructed block in the shaped domain.
[0471] i. Additionally, alternatively, a filter may be applied first before applying the inverse shaping. Afterwards, the filtered reconstructed block may then be converted to the original domain.
[0472] ii. Figure 30 An example of a process for inter-frame coding and decoding is depicted in .
[0473] d. Filter parameters may depend on whether ILR is enabled.
[0474] 5. It is proposed to apply a filtering process in the reshaped domain that can be applied to reconstructed blocks (eg after intra / inter or other kinds of prediction methods).
[0475] a. In one example, the deblocking filter (DBF) process is performed in the plastic domain.
[0476] In this case, no inverse reshaping is applied before DBF.
[0477] i. In this case, the DBF parameters may be different depending on whether shaping is applied or not.
[0478] ii. In one example, the DBF process may depend on whether shaping is enabled.
[0479] 1. In one example, this method is applied when DBF is called in the original domain.
[0480] 2. Alternatively, this method is applied when calling DBF in the shaping domain.
[0481] b. In one example, the Sample Adaptive Offset (SAO) filtering process is performed in the reshaped domain. In this case, no inverse reshaping is applied before SAO.
[0482] c. In one example, the adaptive loop filter (ALF) filtering process is performed in the shaping domain. In this case, no inverse shaping is applied before the ALF.
[0483] d. Additionally, alternatively, inverse reshaping can be applied to the block after DBF.
[0484] e. Additionally, alternatively, inverse reshaping can be applied to the block after SAO.
[0485] f. Additionally, alternatively, inverse reshaping can be applied to the block after ALF.
[0486] g. The above filtering methods can be replaced by other types of filtering methods.
[0487] 6. It is proposed to signal ILR parameters in a new parameter set (such as ILR APS) instead of slice group header.
[0488] a. In one example, the slice group header may contain aps_id. Alternatively, aps_id may be conditionally present when ILR is enabled.
[0489] b. In one example, the ILR APS contains aps_id and ILR parameters.
[0490] c. In one example, a new NUT (NAL Unit Type, as in AVC and HEVC) value is assigned to the ILR APS.
[0491] d. In one example, the range of ILR APS ID values will be 0...M (eg, M=2K-1).
[0492] e. In one example, the ILR APS may be shared across pictures (and may be different in different slice groups within a picture).
[0493] f. In one example, when an ID value is present, it may be encoded using a fixed length encoding scheme. Alternatively, it may be encoded using Exponential-Golomb (EG) encoding, truncated unary encoding, or other binarization schemes.
[0494] g. In one example, ID values cannot be reused with different content within the same image.
[0495] h. In one example, the ILR APS and the APS of the ALF parameters can share the same NUT.
[0496] i. Alternatively, the ILR parameters can be carried with the current APS of the ALF parameters. In this case, the above method mentioning the ILR APS can be replaced by the current APS.
[0497] j. Alternatively, the ILR parameters can be carried in the SPS / VPS / PPS / sequence header / picture header.
[0498] k. In one example, the ILR parameters may include shaper model information, use of the ILR method, and a chroma residual scaling factor.
[0499] 1. Additionally, alternatively, the ILR parameters may be signaled in one level (such as in the APS), and / or the use of ILR may be further signaled in a second level (such as in the slice group header).
[0500] m. Additionally, alternatively, predictive coding can be applied to codec ILR parameters with different APS indices.
[0501] 7. Instead of applying luma-dependent chroma residual scaling (LCRS) to chroma blocks, it is proposed to apply a forward / inverse shaping process to chroma blocks to remove the dependency between luma and chroma.
[0502] In one example, one M-segment piecewise linear (PWL) model and / or forward / backward lookup table can be used for one chrominance component. Alternatively, two PWL models and / or forward / backward lookup tables can be used to encode and decode two chrominance components, respectively.
[0503] b. In one example, the PWL model and / or forward / backward lookup tables for chrominance can be derived from the PWL model and / or forward / backward lookup tables for luma.
[0504] i. In one example, no further signaling of the PWL model / lookup table for chrominance is required.
[0505] c. In one example, the PWL model and / or forward / backward lookup table for chroma may be signaled in SPS / VPS / APS / PPS / sequence header / picture header / slice group header / slice header / CTU row / CTU group / region.
[0506] 8. In one example, how to signal the ILR parameters of a picture / slice group may depend on the ILR parameters of previously coded pictures / slice groups.
[0507] a. For example, the ILR parameter of a picture / slice group can be predicted by the ILR parameters of one or more previously coded pictures / slice groups.
[0508] 9. Propose to disable luma-dependent chroma residual scaling (LCRS) for specific block sizes / temporal layers / slice group types / picture types / codec modes / specific types of motion information.
[0509] a. In one example, even when the forward / inverse shaping process is applied to the luma block, LCRS may not be applied to the corresponding chroma blocks.
[0510] b. Alternatively, even when the forward / inverse shaping process is not applied to the luma block, LCRS can still be applied to the corresponding chroma block.
[0511] c. In one example, when the Cross-Component Linear Model (CCLM) mode is applied, LCRS is not used. The CCLM mode includes LM, LM-A, and LM-L.
[0512] d. In one example, when the Cross-Component Linear Model (CCLM) mode is not applied, LCRS is not used. The CCLM modes include LM, LM-A, and LM-L.
[0513] e. In one example, when encoding and decoding luma blocks exceeds one VPDU (eg, 64×64).
[0514] i. In one example, when the luma block size contains less than M*H samples (eg, 16 or 32 or 64 luma samples), LCRS is not allowed.
[0515] ii. Alternatively, LCRS is not allowed when the minimum dimension of the width or / and height of the luma block is less than or not greater than X. In one example, X is set to 8.
[0516] iii. Alternatively, LCRS is not allowed when the minimum size of the width or / and height of the luma block is not less than X. In one example, X is set to 8.
[0517] iv. Alternatively, when the width of the block > th1 or >= th1 and / or the height of the luminance block > th2 or >= th2, LCRS is not allowed. In one example, th1 and / or th2 is set to 8.
[0518] 1. In one example, th1 and / or th2 is set to 128.
[0519] 2. In one example, th1 and / or th2 is set to 64.
[0520] v. Alternatively, when the width of the luminance block < th1 or <= th1 and / or the height of the luminance block < th2 or <= th2, LCRS is not allowed. In one example, th1 and / or th2 is set to 8.
[0521] 10. Whether to disable ILR (forward shaping process and / or reverse shaping process) can depend on coefficients.
[0522] a. In one example, when a block is encoded and decoded with all-zero coefficients, the forward shaping process applied to the prediction block is skipped.
[0523] b. In one example, when a block is encoded and decoded with all-zero coefficients, the reverse shaping process applied to the reconstructed block is skipped.
[0524] c. In one example, when a block is encoded and decoded with only one non-zero coefficient located at a specific position (e.g., the DC coefficient at the upper left position of a block, the coefficients at the upper left coding group within a block), the forward shaping process applied to the prediction block and / or the reverse shaping process applied to the reconstructed block is skipped.
[0525] d. In one example, when a block is encoded and decoded with only M (e.g., M = 1) non-zero coefficients, the forward shaping process applied to the prediction block and / or the reverse shaping process applied to the reconstructed block is skipped.
[0526] 11. It is proposed that if the encoded and decoded block exceeds one virtual pipeline data unit (VPDU), the ILR application area is divided into VPDU units. Each application area (e.g., with a maximum size of 64×64) is regarded as a separate CU for ILR operations.
[0527] a. In one example, when the width of the block > th1 or >= th1 and / or the height of the block > th2 or >= th2, it can be divided into sub-blocks with a width < th1 or <= th1 and / or a height < th2 or <= th2, and ILR can be performed on each sub-block.
[0528] i. In one example, the sub-blocks can have the same width or / and height.
[0529] ii. In one example, the sub - blocks other than those located at the right boundary or / and the bottom boundary may have the same width or / and height.
[0530] iii. In one example, the sub - blocks other than those located at the left boundary or / and the top boundary may have the same width or / and height.
[0531] b. In one example, when the size of the block (i.e., width * height)>th3 or >= th3, it can be divided into sub - blocks with size <th3 or <= th3, and ILR can be performed on each sub - block.
[0532] i. In one example, the sub - blocks may have the same size. <00013-- 11>
[0533] ii. In one example, the sub - blocks other than those located at the right boundary or / and the bottom boundary may have the same size.
[0534] iii. In one example, the sub - blocks other than those located at the left boundary or / and the top boundary may have the same size.
[0535] c. Alternatively, the use of ILR is limited to a specific block size.
[0536] i. In one example, when the codec block exceeds one VPDU (e.g., 64×64), ILR is not allowed.
[0537] ii. In one example, when the block size contains less than M * H samples (e.g., 16 or 32 or 64 luminance samples), ILR is not allowed.
[0538] iii. Alternatively, when the minimum size of the width or / and height of the block is less than or not greater than X, ILR is not allowed. In one example, X is set to 8. <-- 01323>
[0539] iv. Alternatively, when the minimum size of the width or / and height of the block is not less than X, ILR is not allowed. In one example, X is set to 8.
[0540] v. Alternatively, when the width of the block>th1 or >= th1 and / or the height of the block>th2 or >= th2, ILR is not allowed. In one example, th1 and / or th2 are set to 8.
[0541] 1. In one example, th1 and / or th2 are set to 128.
[0542] 2. In one example, th1 and / or th2 are set to 64.
[0543] Alternatively, when the width of the block < th1 or <= th1 and / or the height of the block < th2 or <= th2, ILR is not allowed. In one example, th1 and / or th2 are set to 8.
[0544] 12. The above methods (e.g., for chroma encoding / decoding, whether to disable ILR and / or whether to disable LCRS and / or whether to signal PWL / lookup table) may depend on the color format, such as 4:4:4 / 4:2:0.
[0545] 13. The indication for enabling ILR (e.g., tile_group_reshaper_enable_flag) may be decoded conditional on the indication of the existing reshaper model (e.g., tile_group_reshaper_model_present_flag).
[0546] a. Alternatively, tile_group_reshaper_model_present_flag may be decoded conditional on tile_group_reshaper_enable_flag.
[0547] b. Alternatively, only one of the two syntax elements including tile_group_reshaper_model_present_flag and tile_group_reshaper_enable_flag may be decoded. The value of the other syntax element is set to be equal to the one syntax element that may be signaled.
[0548] 14. Different clipping methods may be applied to the prediction signal and the reconstruction process.
[0549] a. In one example, an adaptive clipping method may be applied, and the maximum and minimum values to be clipped may be defined in the reshaping domain.
[0550] b. In one example, adaptive clipping may be applied to the prediction signal in the reshaping domain.
[0551] c. Further, alternatively, fixed clipping (e.g., according to bit depth) may be applied to the reconstructed block.
[0552] 15. Filter parameters (such as the parameters used in DF, BF, HF) may depend on whether ILR is enabled.
[0553] 16. It is proposed that for blocks encoded / decoded in Palette mode, ILR is disabled or applied differently.
[0554] a. In one example, when a block is encoded or decoded in palette mode, shaping and inverse shaping are skipped.
[0555] b. Alternatively, when the block is encoded or decoded in palette mode, different shaping and inverse shaping functions may be applied.
[0556] 17. Alternatively, when ILR is applied, the palette mode can be encoded and decoded differently.
[0557] a. In one example, when ILR is applied, the palette mode can be encoded and decoded in the original domain.
[0558] b. Alternatively, when ILR is applied, the palette mode can be encoded or decoded in the shaped domain.
[0559] c. In one example, when ILR is applied, the palette prediction value can be signaled in the original domain.
[0560] d. Alternatively, the palette prediction value can be signaled in the shaped domain.
[0561] 18. It is proposed that for blocks encoded or decoded in IBC mode, ILR be disabled or applied differently.
[0562] a. In one example, when a block is encoded or decoded in IBC mode, shaping and inverse shaping are skipped.
[0563] b. Alternatively, when the block is encoded or decoded in IBC mode, different shaping and inverse shaping are applied.
[0564] 19. Alternatively, when ILR is applied, the IBC mode may be encoded or decoded differently.
[0565] a. In one example, when ILR is applied, IBC can be performed in the original domain.
[0566] b. Alternatively, when ILR is applied, IBC can be performed in the shaping domain.
[0567] 20. It is proposed that for blocks coded in B-DPCM mode, ILR be disabled or applied differently.
[0568] a. In one example, when the block is coded in B-DPCM mode, shaping and inverse shaping are skipped.
[0569] b. Alternatively, when the block is coded in B-DPCM mode, different shaping and inverse shaping are applied.
[0570] 21. Alternatively, when ILR is applied, the B-DPCM mode can be coded differently.
[0571] a. In one example, when ILR is applied, B-DPCM can be performed in the original domain.
[0572] b. Alternatively, when ILR is applied, B-DPCM can be performed in the shaping domain.
[0573] 22. It is proposed that for blocks coded in transform skip mode, ILR be disabled or applied differently.
[0574] a. In one example, when a block is encoded in transform skip mode, shaping and inverse shaping are skipped.
[0575] b. Alternatively, when a block is encoded in transform skip mode, different shaping and inverse shaping may be applied.
[0576] 23. Alternatively, when ILR is applied, transform skip mode may be encoded and decoded differently.
[0577] a. In one example, when ILR is applied, transform skipping can be performed in the original domain.
[0578] b. Alternatively, when ILR is applied, transform skipping can be performed in the shaped domain.
[0579] 24. It is proposed that for blocks encoded in I-PCM mode, ILR be disabled or applied differently.
[0580] a. In one example, when a block is encoded or decoded in palette mode, shaping and inverse shaping are skipped.
[0581] b. Alternatively, when the block is encoded or decoded in palette mode, different shaping and inverse shaping functions may be applied.
[0582] 25. Alternatively, when ILR is applied, the I-PCM mode may be encoded or decoded differently.
[0583] a. In one example, when ILR is applied, the I-PCM mode can be encoded and decoded in the original domain.
[0584] b. Alternatively, when ILR is applied, the I-PCM mode can be encoded and decoded in the shaped domain.
[0585] 26. It is proposed that for blocks encoded in transform quantization bypass mode, ILR be disabled or applied differently.
[0586] a. In one example, when a block is encoded in transform quantization bypass mode, shaping and inverse shaping are skipped.
[0587] 27. Alternatively, when the block is encoded or decoded in transform and quantization bypass mode, different shaping and inverse shaping functions are applied.
[0588] 28. For the above bullets, when ILR is disabled, the forward shaping and / or reverse shaping process can be skipped.
[0589] a. Alternatively, the prediction and / or reconstruction and / or residual signal is in the original domain.
[0590] b. Optionally, the prediction and / or reconstruction and / or residual signal is in the shaped domain.
[0591] 29. Multiple shaping / inverse shaping functions (such as multiple PWL models) can be allowed to be used to encode and decode a picture / a slice group / a VPDU / a region / a CTU row / multiple CUs.
[0592] a. How to select from multiple functions may depend on block size / codec mode / picture type / low latency check flag / motion information / reference pictures / video content, etc.
[0593] b. In one example, multiple ILR side information sets (e.g., shaping / inverse shaping functions) may be signaled per SPS / VPS / PPS / sequence header / picture header / slice group header / slice header / region / VPDU / etc.
[0594] i. Additionally, alternatively, predictive coding of ILR side information can be utilized.
[0595] c. In one example, more than one aps_idx may be signaled in the PPS / picture header / slice group header / slice header / region / VPDU / etc.
[0596] 30. In one example, the shaping information is signaled in a new syntax set instead of VPS, SPS, PPS or APS. For example, the shaping information is signaled in a set denoted as inloop_reshaping_parameter_set() (IRPS, or any other name).
[0597] a. An example grammar is designed as follows. Added grammar is highlighted in italics.
[0598]
[0599] inloop_reshaping_parameter_set_id provides an identifier of the IRPS for reference by other syntax elements.
[0600] NOTE - IRPS may be shared across pictures and may be different in different slice groups within a picture.
[0601] irps_extension_flag equal to 0 specifies that the irps_extension_data_flag syntax element is not present in the IRPS RBSP syntax structure. irps_extension_flag equal to 1 specifies that the irps_extension_data_flag syntax element is present in the IRPS RBSP syntax structure.
[0602] The irps_extension_data_flag can have any value. Its presence and value do not affect the decoder's compliance with the profile specified in this version of this specification. Decoders conforming to this version of this specification should ignore all irps_extension_data_flag syntax elements.
[0603] b. Example syntax is designed as follows. Added syntax is highlighted in italics.
[0604] Generalized slice group header syntax and semantics
[0605]
[0606]
[0607] tile_group_irps_id specifies the inloop_reshaping_parameter_set_id of the IRPS referenced by the slice group. The TemporalId of the IRPS NAL unit whose inloop_reshaping_parameter_set_id is equal to tile_group_irps_id should be less than or equal to the TemporalId of the codec slice group NAL unit.
[0608] 31. In one example, IRL information is signaled in APS along with ALF information.
[0609] a. An example grammar is designed as follows. Added grammar is highlighted in italics.
[0610] Adaptive Parameter Set Syntax and Semantics
[0611]
[0612] b. In one example, a tile_group_aps_id is signaled in the slice group header to specify the adaptation_parameter_set_id of the APS referenced by the slice group. The ALF information and ILR information of the current slice group are signaled in the specified APS.
[0613] i. Example syntax is designed as follows. Added syntax is highlighted in italics.
[0614]
[0615] 32. In one example, ILR information and ALF information are signaled in different APSs.
[0616] a. The first ID (which can be named tile_group_aps_id_alf) is signaled in the slice group header to specify the first adaptation_parameter_set_id of the first APS referenced by the slice group. The ALF information of the current slice group is signaled in the specified first APS.
[0617] b. A second ID (named tile_group_aps_id_irps) is signaled in the slice group header to specify the second adaptation_parameter_set_id of the second APS referenced by the slice group. The ILR information of the current slice group is signaled in the specified second APS.
[0618] c. In one example, the first APS must have ALF information in the conformance bitstream;
[0619] d. In one example, the second APS must have ILR information in the conforming bitstream;
[0620] e. Example syntax is shown below. Added syntax is highlighted in italics.
[0621]
[0622]
[0623] 33. In one example, some APSs with a specified adaptation_parameter_set_id must have ALF information. In another example, some APSs with a specified adaptation_parameter_set_id must have ILR information.
[0624] a. For example, an APS with adaptation_parameter_set_id equal to 2N must have ALF information. N is any integer;
[0625] b. For example, the APS with adaptation_parameter_set_id equal to 2N+1 must have ILR information. N is any integer;
[0626] c. An example grammar is designed as follows. Added grammar is highlighted in italics.
[0627]
[0628]
[0629] i. For example, 2*tile_group_aps_id_alf specifies the first adaptation_parameter_set_id of the first APS referenced by the tile group. The ALF information of the current tile group is signaled in the specified first APS.
[0630] ii. For example, 2*tile_group_aps_id_irps+1 specifies the second adaptation_parameter_set_id of the second APS referenced by the tile group. The ILR information of the current tile group is signaled in the specified second APS.
[0631] 34. In one example, a slice group cannot reference an APS (or IRPS) signaled before a specified type of Network Abstraction Layer (NAL) unit that is signaled before the current slice group.
[0632] a. In one example, a slice group cannot reference an APS (or IRPS) signaled before a slice group of a specified type that is signaled before the current slice group.
[0633] b. For example, a slice group cannot reference an APS (or IRPS) signaled before an SPS that is signaled before the current slice group.
[0634] c. For example, a slice group cannot reference an APS (or IRPS) signaled before a PPS that is signaled before the current slice group.
[0635] d. For example, a slice group cannot reference an APS (or IRPS) signaled before an access unit delimiter NAL (Access Unit Delimiter, AUD) that is signaled before the current slice group.
[0636] e. For example, a slice group cannot reference an APS (or IRPS) signaled before the End of Bitstream (EoB) NAL, which is signaled before the current slice group.
[0637] f. For example, a slice group cannot reference an APS (or IRPS) signaled before the End of Sequence (EoS) NAL, which is signaled before the current slice group.
[0638] g. For example, a slice group cannot reference an APS (or IRPS) signaled before an Instantaneous Decoding Refresh (IDR) NAL that is signaled before the current slice group.
[0639] h. For example, a slice group cannot reference an APS (or IRPS) signaled before a Clean Random Access (CRA) NAL that is signaled before the current slice group.
[0640] i. For example, a slice group cannot reference an APS (or IRPS) signaled before an Intra Random Access Point (IRAP) access unit that is signaled before the current slice group.
[0641] j. For example, a slice group cannot reference an APS (or IRPS) signaled before an I slice group (or picture, or slice) that is signaled before the current slice group.
[0642] k. The methods disclosed in IDF-P1903237401H and IDF-P1903234501H may also be applied when carrying ILR information in APS or IRPS.
[0643] 35. The conforming bitstream should satisfy: when the loop shaping method is enabled for a video data unit (such as a sequence), the default ILR parameters should be defined, such as a default model.
[0644] a. When sps_lmcs_enabled_flag is set to 1, sps_lmcs_default_model_present_flag should be set to 1.
[0645] b. Default parameters may be signaled conditional on an ILR enabled flag instead of a default model present flag (such as sps_lmcs_default_model_present_flag).
[0646] c. For each slice group, a default model usage flag (such as tile_group_lmcs_use_default_model_flag) may be signaled without referencing the SPS default model usage flag.
[0647] d. The conforming bitstream should satisfy: When there is no ILR information in the corresponding APS type of ILR and a video data unit (such as a slice group) is forced to use the ILR technology, the default model should be used.
[0648] e. Alternatively, the conforming bitstream should satisfy: when there is no ILR information in the corresponding APS type of ILR and a video data unit (such as a slice group) is forced to use the ILR technology (such as tile_group_lmcs_enable_flag is equal to 1), the indication of using the default model should be true, for example, tile_group_lmcs_use_default_model_flag should be 1.
[0649] f. The restriction is that default ILR parameters (such as a default model) should be transmitted in a video data unit (such as an SPS).
[0650] i. Additionally, alternatively, when the SPS flag indicating the use of ILR is true, the default ILR parameters should be transmitted.
[0651] g. The restriction is that at least one ILR APS is transmitted in a video data unit (such as an SPS).
[0652] i. In one example, at least one ILR APS includes default ILR parameters (such as a default model).
[0653] 36. The default ILR parameters can be indicated by a flag. When the flag indicates that the default ILR parameters are used, no further signaling of ILR data is required.
[0654] 37. The default ILR parameters may be predefined when no default ILR parameters are signaled. For example, the default ILR parameters may correspond to identity mapping.
[0655] 38. The time domain layer information may be signaled together with the ILR parameters (such as in the ILR APS).
[0656] a. In one example, the time domain layer index can be signaled in lmcs_data().
[0657] b. In one example, the time domain layer index minus 1 can be signaled in lmcs_data().
[0658] c. Furthermore, alternatively, when encoding / decoding a slice group / slice, it is restricted to referencing those ILR APSs associated with smaller or equal time-domain layer indices.
[0659] d. Alternatively, when encoding / decoding a slice group / slice, it is restricted to referencing those ILR APSs associated with smaller time-domain layer indices.
[0660] e. Alternatively, when encoding / decoding a slice group / slice, it is restricted to referencing those ILR APSs associated with a larger temporal layer index.
[0661] f. Alternatively, when encoding / decoding a slice group / slice, it is restricted to referencing those ILR APSs associated with greater or equal temporal layer indices.
[0662] g. Alternatively, when encoding / decoding a slice group / slice, it is restricted to referencing those ILR APSs associated with equal temporal layer indices.
[0663] h. In one example, whether the above restrictions apply may depend on a piece of information that may be signaled to the decoder or inferred by the decoder.
[0664] 39. Time domain layer information may be signaled together with ALF parameters (such as in ALF APS).
[0665] a. In one example, the time domain layer index can be signaled in alf_data().
[0666] b. In one example, the time domain layer index minus 1 may be signaled in alf_data().
[0667] c. Furthermore, alternatively, when encoding / decoding a slice group / slice or a CTU within a slice / slice group, it is restricted to referencing those ALF APSs associated with smaller or equal temporal layer indices.
[0668] d. Alternatively, when encoding / decoding a slice group / slice, it is restricted to referencing those ALF APSs associated with smaller time-domain layer indices.
[0669] e. Alternatively, when encoding / decoding a slice group / slice, it is restricted to referencing those ALF APSs associated with a larger temporal layer index.
[0670] f. Alternatively, when encoding / decoding a slice group / slice, it is restricted to referencing those ALF APSs associated with greater or equal temporal layer indices.
[0671] g. Alternatively, when encoding / decoding a slice group / slice, it is restricted to referencing those ALF APSs associated with equal temporal layer indices.
[0672] h. In one example, whether the above restrictions apply may depend on a piece of information that may be signaled to the decoder or inferred by the decoder.
[0673] 40. In one example, the shaping mapping between the original sample points and the shaped sample points may not be a positive relationship, that is, a larger value is not allowed to be mapped to a smaller value.
[0674] a. For example, the shaping mapping between the original samples and the shaped samples can be a negative relationship, where for two values, the larger value in the original domain can be mapped to the smaller value in the shaped domain.
[0675] 41. In conforming bitstreams, the syntax element aps_params_type is only allowed to have a few predefined values, such as 0 and 1.
[0676] a. In another example, it is only allowed to be 0 and 7.
[0677] 42. In one example, if ILR can be applied (eg, sps_lmcs_enabled_flag is true), then default ILR information must be signaled.
[0678] 5 Example Implementations of the Disclosed Technology
[0679] In some embodiments, tile_group_reshaper_enable_flag is conditionally present when tile_group_reshaper_model_present_flag is enabled.Added syntax is highlighted in italics.
[0680] 7.3.3.1 Generalized slice group header syntax
[0681]
[0682] Alternatively, tile_group_reshaper_model_present_flag is conditionally present when tile_group_reshaper_enable_flag is enabled.
[0683]
[0684] Alternatively, only one of the two syntax elements tile_group_reshaper_model_present_flag or tile_group_reshaper_enable_flag can be signaled. The one syntax element that is not signaled is inferred to be equal to the one syntax element that can be signaled. In this case, one syntax element controls the use of ILR.
[0685] Alternatively, the conforming bitstream requires that tile_group_reshaper_model_present_flag should be equal to tile_group_reshaper_enable_flag.Alternatively, tile_group_reshaper_model_present_flag and / or tile_group_reshaper_enable_flag and / or tile_group_reshaper_model() and / or tile_group_reshaper_chroma_residual_scale_flag may be signaled in the APS instead of the slice group header.
[0686] Example #2 on top of JVET-N0805. Added syntax highlighted in italics.
[0687] ...
[0689] sps_lmcs_enabled_flag equal to 1 specifies that luma mapping and chroma scaling are used in the Coded Video Sequence (CVS). sps_lmcs_enabled_flag equal to 0 specifies that luma mapping and chroma scaling are not used in the CVS.
[0690] sps_lmcs_default_model_present_flag equal to 1 specifies that default LMCS data is present in this SPS. sps_lmcs_default_model_flag equal to 0 specifies that default LMCS data is not present in this SPS. When not present, the value of sps_lmcs_default_model_present_flag is inferred to be equal to 0. ...
[0692]
[0693] aps_params_type specifies the type of APS parameters carried in APS, as specified in the following table:
[0694] Table 7-x – APS parameter type codes and APS parameter types
[0695]
[0696]
[0697] ALF APS: APS with aps_params_type equal to ALF_APS.
[0698] LMCS APS: APS with aps_params_type equal to LMCS_APS.
[0699] Make the following semantic changes: ...
[0701] tile_group_alf_aps_id specifies the adaptation_parameter_set_id of the ALF APS referenced by the slice group. The TemporalId of the ALF APS NAL unit with adaptation_parameter_set_id equal to tile_group_alf_aps_id should be less than or equal to the TemporalId of the NAL unit of the codec slice group.
[0702] When multiple ALF APSs having the same adaptation_parameter_set_id value are referenced by two or more slice groups of the same picture, the multiple ALF APSs having the same adaptation_parameter_set_id value should have the same content. ...
[0704] tile_group_lmcs_enabled_flag equal to 1 specifies that luma mapping and chroma scaling are enabled for the current slice group. tile_group_lmcs_enabled_flag equal to 0 specifies that luma mapping and chroma scaling are not enabled for the current slice group. When tile_group_lmcs_enable_flag is not present, it is inferred to be equal to 0.
[0705] tile_group_lmcs_use_default_model_flag equal to 1 specifies that the default lmcs model is used for luma mapping and chroma scaling for the slice group. tile_group_lmcs_use_default_model_flag equal to 0 specifies that the lmcs model in the LMCS APS referenced by tile_group_lmcs_aps_id is used for luma mapping and chroma scaling for the slice group. When tile_group_reshaper_use_default_model_flag is not present, it is inferred to be equal to 0.
[0706] tile_group_lmcs_aps_id specifies the adaptation_parameter_set_id of the LMCS APS referenced by the slice group. The TemporalId of the LMCS APS NAL unit with adaptation_parameter_set_id equal to tile_group_lmcs_aps_id should be less than or equal to the TemporalId of the codec slice group NAL unit.
[0707] When multiple LMCS APSs having the same adaptation_parameter_set_id value are referenced by two or more slice groups of the same picture, the multiple LMCS APSs having the same adaptation_parameter_set_id value should have the same content.
[0708] tile_group_chroma_residual_scale_flag equal to 1 specifies that chroma residual scaling is enabled for the current slice group. tile_group_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling is not enabled for the current slice group. When tile_group_chroma_residual_scale_flag is not present, it is inferred to be equal to 0. ...
[0710] Luma Mapping and Chroma Scaling Data Syntax
[0711]
[0712] The examples described above can be incorporated into the methods described below (e.g., Figures 31A to 39E The method shown is in the context of a video decoder or a video encoder).
[0713] Figure 31AA flowchart of an exemplary method for video processing is shown. The method 3100 includes, at step 3110, performing a motion information refinement process based on samples in a first domain or a second domain for a conversion between a current video block of a video and a codec representation of the video. The method 3100 includes, at step 3120, performing the conversion based on a result of the motion information refinement process. In some embodiments, during the conversion, samples are obtained from a first prediction block in the first domain using unrefined motion information for the current video block, at least a second prediction block is generated in the second domain using refined motion information for determining a reconstructed block, and reconstructed samples for the current video block are generated based on at least the second prediction block.
[0714] Figure 31B A flowchart of an exemplary method for video processing is shown. Method 3120 includes, at step 3122, reconstructing a current video block based on at least one prediction block in a second domain. In some embodiments, during the conversion, the current video block is constructed based on the first domain and the second domain, and / or the chroma residual is scaled in a luma-dependent manner. In some embodiments, the codec is applied during the conversion using parameters derived based on at least a first set of samples in a video region of the video and a second set of samples in a reference picture of the current video block. In some embodiments, the domain of the first sample and the domain of the second sample are aligned.
[0715] Figure 32A A flowchart of an exemplary method for video processing is shown. Method 3210 includes, at step 3212, determining parameters of a codec mode for a current video block of a current video region of a video based on one or more parameters of a codec mode of a previous video region. Method 3210 also includes, at step 3214, performing codecs on the current video block based on the determination to generate a codec representation of the video. In some embodiments, the parameters of the codec mode are included in a parameter set in the codec representation of the video. In some embodiments, performing codecs includes transforming a representation of the current video block in a first domain into a representation of the current video block in a second domain. In some embodiments, during codecs performed using the codec mode, the current video block is constructed based on the first domain and the second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0716] Figure 32BA flow chart of an exemplary method for video processing is shown. Method 3220 includes, at step 3222, receiving a codec representation of a video including a parameter set, wherein the parameter set includes parameter information for a codec mode. Method 3220 also includes, at step 3224, performing decoding of the codec representation using the parameter information to generate a current video block for a current video region of the video from the codec representation. In some embodiments, the parameter information for the codec mode is based on one or more parameters of a codec mode for a previous video region. In some embodiments, in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0717] Figure 32C A flowchart of an exemplary method for video processing is shown. Method 3230 includes, at step 3232, performing a conversion between a current video block of a video and a codec representation of the video. In some embodiments, the conversion includes applying a filtering operation to a prediction block in a first domain or in a second domain different from the first domain.
[0718] Figure 32D A flow chart of an exemplary method for video processing is shown. Method 3240 includes, at step 3242, performing a conversion between a current video block of a video and a codec representation of the video. In some embodiments, during the conversion, a final reconstructed block is determined for the current video block. In some embodiments, the temporary reconstructed block is generated using a prediction method and represented in a second domain.
[0719] Figure 33 A flowchart of an exemplary method for video processing is shown. The method 3300 includes, at step 3302, performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein, in the codec mode, the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein a parameter set in the codec representation includes parameter information of the codec mode.
[0720] Figure 34A A flowchart of an exemplary method for video processing is shown. The method 3410 includes, at step 3412, performing a conversion between a current video block of the video that is a chroma block and a codec representation of the video, wherein during the conversion, the current video block is constructed based on a first domain and a second domain, and wherein the conversion further includes applying a forward shaping process and / or an inverse shaping process to one or more chroma components of the current video block.
[0721] Figure 34BA flowchart of an exemplary method for video processing is shown. The method 3420 includes, at step 3422, performing a conversion between a current video chroma block of a video and a codec representation of the video, wherein performing the conversion includes determining, based on a rule, whether luma-dependent chroma residual scaling (LCRS) is enabled or disabled, and reconstructing the current video chroma block based on the determination.
[0722] Figure 35A A flow chart of an exemplary method for video processing is shown. Method 3510 includes, at step 3512, determining whether to disable use of a codec mode for a conversion between a current video block of a video and a codec representation of the video based on one or more coefficient values of the current video block. Method 3510 also includes, at step 3514, performing the conversion based on the determination. In some embodiments, during the conversion using the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0723] Figure 35B A flow chart of an exemplary method for video processing is shown. Method 3520 includes, at step 3522, for converting between current video blocks of a video exceeding a virtual pipe data unit (VPDU) of the video, dividing the current video block into regions. Method 3520 also includes, at step 3524, performing the conversion by applying a codec mode separately to each region. In some embodiments, during the conversion by applying the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0724] Figure 35C A flow chart of an exemplary method for video processing is shown. The method 3530 includes, at step 3532, determining whether to disable use of a codec mode for a conversion between a current video block of a video and a codec representation of the video based on a size or color format of the current video block. The method 3530 also includes, at step 3534, performing the conversion based on the determination. In some embodiments, during the conversion using the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0725] Figure 35DA flowchart of an exemplary method for video processing is shown. The method 3540 includes, at step 3542, performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein in the codec mode, the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein at least one syntax element in the codec representation provides an indication of use of the codec mode and an indication of a shaper model.
[0726] Figure 35E A flow chart of an exemplary method for video processing is shown. Method 3550 includes, at step 3552, determining that a codec mode is disabled for conversion between a current video block of a video and a codec representation of the video. Method 3550 also includes, at step 3554, conditionally skipping forward shaping and / or inverse shaping based on the determination. In some embodiments, in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0727] Figure 35F A flowchart of an exemplary method for video processing is shown. The method 3560 includes, at step 3562, performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode in which the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein a plurality of forward shapings and / or a plurality of inverse shapings are applied in the shaping mode for the video region.
[0728] Figure 36A A flow chart of an exemplary method for video processing is shown. Method 3610 includes, at step 3612, determining whether a codec mode is enabled for converting between a current video block of a video and a codec representation of the video. Method 3610 also includes, at step 3614, performing the conversion using a palette mode, wherein in the palette mode, at least a palette of representative sample values is used for the current video block. In some embodiments, in the codec mode, the current video block is constructed based on samples in the first domain and the second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0729] Figure 36BA flow chart of an exemplary method for video processing is shown. Method 3620 includes, at step 3622, for converting between a current video block of a video and a codec representation of the video, determining that the current video block is encoded in a palette mode, wherein in the palette mode, a palette of at least representative sample values is used to encode and decode the current video block. Method 3620 also includes, at step 2624, performing the conversion by disabling the codec mode as a result of the determination. In some embodiments, when the codec mode is applied to the video block, the video block is constructed based on chroma residuals that are scaled in a luma-dependent manner.
[0730] Figure 36C A flowchart of an exemplary method for video processing is shown. Method 3630 includes, at step 3632, performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion uses a first codec mode and a palette codec mode, wherein in the palette codec mode, a palette of at least representative pixel values is used to encode and decode the current video block. Method 3630 also includes, at step 3634, performing a conversion between a second video block of the video that was encoded without using the palette codec mode and a codec representation of the video, wherein the conversion of the second video block uses the first codec mode. When the first codec mode is applied to the video block, the video block is constructed based on the first domain and the second domain, and / or the chroma residual is scaled in a luma-dependent manner. In some embodiments, the first codec mode is applied to the first video block and the second video block in different manners.
[0731] Figure 37A A flow chart of an exemplary method for video processing is shown. The method 3710 includes, at step 3712, determining whether a codec mode is enabled for converting between a current video block of a video and a codec representation of the video. The method 3710 also includes, at step 3714, performing the conversion using an intra block copy mode, wherein the intra block copy mode uses at least a block vector pointing to a picture including the current video block to generate a prediction block. In this codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0732] Figure 37BA flow chart of an exemplary method for video processing is shown. Method 3720 includes, at step 3722, for converting between a current video block of a video and a codec representation of the video, determining that the current video block is encoded in intra block copy (IBC) mode, wherein the intra block copy mode uses at least a block vector pointing to a video frame containing the current video block to generate a prediction block for encoding and decoding the current video block. Method 3720 also includes, at step 3724, performing the conversion by disabling the codec mode due to the determination. When the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0733] Figure 37C A flowchart of an exemplary method for video processing is shown. Method 3730 includes, at step 3732, performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion uses an intra block copy mode and a first codec mode, wherein the intra block copy mode uses at least a block vector pointing to a video frame containing the current video block to generate a prediction block. Method 3730 also includes, at step 3734, performing a conversion between a second video block of the video that was encoded and decoded without using the intra block copy mode and a codec representation of the video, wherein the conversion of the second video block uses the first codec mode. When the first codec mode is applied to the video block, the video block is constructed based on the first domain and the second domain, and / or the chroma residual is scaled in a luma-dependent manner, and the first codec mode is applied differently to the first video block and the second video block.
[0734] Figure 38A A flow chart of an exemplary method for video processing is shown. The method 3810 includes, at step 3812, determining that a codec mode is enabled for converting between a current video block of a video and a codec representation of the video. The method 3810 also includes, at step 3814, performing the conversion using a block-based delta pulse codec modulation (BDPCM) mode. In this codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0735] Figure 38BA flow chart of an exemplary method for video processing is shown. Method 3820 includes, at step 3822, determining, for conversion between a current video block of a video and a codec representation of the video, whether the current video block is coded using a block-based delta pulse codec modulation (BDPCM) mode. Method 3820 also includes, at step 3824, performing the conversion by disabling a codec mode due to the determination. When the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0736] Figure 38C A flowchart of an exemplary method for video processing is shown. Method 3830 includes, at step 3832, performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and a block-based delta pulse codec modulation (BDPCM) mode. Method 3830 also includes, at step 3834, performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded without using the BDPCM mode and the conversion of the second video block uses the first codec mode. When the first codec mode is applied to the video block, the video block is constructed based on the first domain and the second domain, and / or the chroma residual is scaled in a luma-dependent manner, and the first codec mode is applied differently to the first video block and the second video block.
[0737] Figure 38D A flow chart of an exemplary method for video processing is shown. Method 3840 includes, at step 3842, determining a codec mode enabled for conversion between a current video block of a video and a codec representation of the video. The method further includes, at step 3844, performing the conversion using a transform skip mode, wherein in the transform skip mode, a transform on a prediction residual is skipped when encoding and decoding the current video block. In the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0738] Figure 38E A flow chart of an exemplary method for video processing is shown. Method 3850 includes, at step 3852, determining, for a conversion between a current video block of a video and a codec representation of the video, whether the current video block is encoded in a transform skip mode, wherein in the transform skip mode, a transform on a prediction residual is skipped when encoding and decoding the current video block. Method 3850 also includes, at step 3854, performing the conversion by disabling a codec mode as a result of the determination. When the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0739] Figure 38F A flowchart of an exemplary method for video processing is shown. Method 3860 includes, at step 3862, performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and a transform skip mode, wherein in the transform skip mode, a transform of a prediction residual is skipped when encoding and decoding the current video block. Method 3860 also includes, at step 3864, performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded and decoded without using the transform skip mode, and the conversion of the second video block uses the first codec mode. When the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner, and the first codec mode is applied differently to the first video block and the second video block.
[0740] Figure 38G A flow chart of an exemplary method for video processing is shown. The method 3870 includes, at step 3872, determining that a codec mode is enabled for converting between a current video block of a video and a codec representation of the video. The method 3870 also includes, at step 3874, performing the conversion using an intra pulse codec modulation mode, wherein the current video block is encoded without applying a transform and transform domain quantization in the intra pulse codec modulation mode. In the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0741] Figure 38H A flow chart of an exemplary method for video processing is shown. Method 3880 includes, at step 3882, determining, for a conversion between a current video block of a video and a codec representation of the video, whether the current video block is coded in an intra pulse codec modulation mode, wherein in the intra pulse codec modulation mode, the current video block is coded without applying a transform and transform domain quantization. Method 3880 also includes, at step 3884, performing the conversion by disabling the codec mode due to the determination. When the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0742] Figure 38IA flowchart of an exemplary method for video processing is shown. Method 3890 includes, at step 3892, performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and an intra-frame pulse codec modulation mode, wherein the current video block is encoded without applying a transform and transform-domain quantization in the intra-frame pulse codec modulation mode. Method 3890 also includes, at step 3894, performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded without using the intra-frame pulse codec modulation mode and the conversion of the second video block uses the first codec mode. When the first codec mode is applied to the video block, the video block is constructed based on the first domain and the second domain, and / or the chroma residual is scaled in a luma-dependent manner, and the first codec mode is applied differently to the first video block and the second video block.
[0743] Figure 38J A flow chart of an exemplary method for video processing is shown. Method 3910 includes, at step 3912, determining a codec mode enabled for conversion between a current video block of a video and a codec representation of the video. Method 3910 also includes, at step 3914, performing the conversion using a modified transform and quantization bypass mode, wherein in the modified transform and quantization bypass mode, the current video block is losslessly coded without transform and quantization. In the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0744] Figure 38K A flowchart of an exemplary method for video processing is shown. Method 3920 includes, at step 3922, determining, for a conversion between a current video block of a video and a codec representation of the video, whether the current video block is coded in a transform and quantization bypass mode, wherein in the transform and quantization bypass mode, the current video block is losslessly coded without transform and quantization. Method 3920 also includes, at step 3924, performing the conversion by disabling the codec mode due to the determination. When the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0745] Figure 38LA flowchart of an exemplary method for video processing is shown. Method 3930 includes, at step 3932, performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and a transform quantization bypass mode, wherein in the transform quantization bypass mode, the current video block is losslessly encoded and decoded without transform and quantization. Method 3930 also includes, at step 3934, performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded and decoded without using the transform quantization bypass mode, and the conversion of the second video block uses the first codec mode. When the first codec mode is applied to the video block, the video block is constructed based on the first domain and the second domain, and / or the chroma residual is scaled in a luma-dependent manner, and the first codec mode is applied differently to the first video block and the second video block.
[0746] Figure 39A A flowchart of an exemplary method for video processing is shown. The method 3940 includes, at step 3942, performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode in which the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein information for the codec mode is signaled in a parameter set other than a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), or an adaptive parameter set (APS) for carrying adaptive loop filtering (ALF) parameters.
[0747] Figure 39B A flowchart of an exemplary method for video processing is shown. The method 3950 includes, at step 3952, performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein information for the codec mode is signaled in an adaptation parameter set (APS) along with adaptive loop filtering (ALF) information, wherein the information for the codec mode and the ALF information are included in one NAL unit.
[0748] Figure 39CA flowchart of an exemplary method for video processing is shown. The method 3960 includes, at step 3962, performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode in which the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein information for the codec mode is signaled in an adaptive parameter set (APS) of a first type that is different from a second type of APS used to signal adaptive loop filtering (ALF) information.
[0749] Figure 39D A flowchart of an exemplary method for video processing is shown. The method 3970 includes, at step 3972, performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode in which the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein the video region is not allowed to reference an adaptation parameter set or a parameter set signaled prior to a data structure of a specified type for processing the video, and wherein the data structure of the specified type is signaled prior to the video region.
[0750] Figure 39E A flowchart of an exemplary method for video processing is shown. The method 3980 includes, at step 3982, performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode in which the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein syntax elements of a parameter set including parameters for processing the video have predefined values in a conforming bitstream.
[0751] Figure 40A 4 is a block diagram of a video processing apparatus 4000. The apparatus 4000 may be configured to implement one or more of the methods described herein. The apparatus 4000 may be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 4000 may include one or more processors 4002, one or more memories 4004, and video processing hardware 4006. The processor(s) 4002 may be configured to implement one or more of the methods described herein (including but not limited to, Figures 31A to 39E Memory(ies) 4004 may be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 4006 may be used to implement some of the techniques described in this document in hardware circuitry.
[0752] Figure 40B is another example of a block diagram of a video processing system in which the disclosed technology may be implemented. Figure 40B is a block diagram illustrating an example video processing system 4100 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4100. System 4100 may include an input 4102 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 4102 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.
[0753] System 4100 may include a codec component 4104 that can implement the various codecs or encoding methods described in this document. Codec component 4104 can reduce the average bit rate of the video from input 4102 to the output of codec component 4104 to produce a coded representation of the video. Codec technology is therefore sometimes referred to as video compression or video transcoding technology. The output of codec component 4104 can be stored or sent via a communication connection such as represented by component 4106. The bitstream (or codec) representation of the video received at input 4102 or the communication transmission can be used by component 4108 to generate pixel values or transmit to a displayable video of display interface 4110. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it will be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.
[0754] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, and IDE interfaces, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0755] In some embodiments, the video encoding and decoding method may use Figure 40Aor Figure 40B The device is implemented on the hardware platform.
[0756] Various techniques and embodiments may be described using the following clause-based format.
[0757] The first set of clauses describes certain features and aspects of the disclosed technology listed in previous sections, including, for example, Examples 1 and 2.
[0758] 1. A method for video processing, comprising: performing a motion information refinement process based on samples in a first domain or a second domain for a conversion between a current video block of a video and a codec representation of the video; and performing the conversion based on a result of the motion information refinement process, wherein during the conversion, samples are obtained from a first prediction block in the first domain using unrefined motion information for the current video block, at least a second prediction block is generated in the second domain using refined motion information for determining a reconstructed block, and reconstructed samples for the current video block are generated based on at least the second prediction block.
[0759] 2. A method according to clause 1, wherein at least the second prediction block is generated from samples in the reference picture in the first domain using the refined motion information, and a reshaping process that converts the first domain to the second domain is further applied to at least the second prediction block.
[0760] 3. The method of clause 2, wherein after the shaping process, the second prediction block is converted to a representation in the second domain before being used to generate the reconstructed samples of the current video block.
[0761] 4. The method of clause 1, wherein performing the motion information refinement process is based on a decoder-side motion vector derivation (DMVD) method.
[0762] 5. The method of clause 4, wherein the DMVD method comprises decoder-side motion vector refinement (DMVR) or frame rate up-conversion (FRUC) or bidirectional optical flow (BIO).
[0763] 6. The method of clause 4, wherein cost calculation or gradient calculation in the DMVD process is performed based on samples in the first domain.
[0764] 7. The method of clause 6, wherein the cost calculation comprises sum of absolute differences (SAD) or mean removed sum of absolute differences (MR-SAD).
[0765] 8. A method according to clause 1, wherein the motion information refinement process is performed based on samples in at least a first prediction block in a first domain being converted to samples in a second domain, and wherein, after obtaining the refined motion information, a codec mode is disabled for at least the second prediction block, wherein in the codec mode the current video block is constructed based on the first domain and the second domain and / or the chroma residual is scaled in a luma-dependent manner.
[0766] 9. The method of clause 4, wherein the motion information refinement process is performed based on at least a first prediction block in the first domain, and wherein the motion information refinement process is invoked using the first prediction block in the first domain.
[0767] 10. The method of clause 1, wherein the final prediction block is generated as a weighted average of the two second prediction blocks, and the reconstructed samples of the current video block are generated based on the final prediction block.
[0768] 11. A method according to clause 1, wherein the motion information refinement process is performed based on the prediction block in the first domain, and wherein, after performing the motion information refinement process, a codec mode is disabled for at least the second prediction block, wherein in the codec mode, the current video block is constructed based on the first domain and the second domain, and / or the chroma residual is scaled in a luma-dependent manner.
[0769] 12. A method for video processing, comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein during the conversion, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, wherein a codec tool is applied during the conversion using parameters derived based on at least a first set of samples in a video region of the video and a second set of samples in a reference picture of the current video block, and wherein the domain of the first sample and the domain of the second sample are aligned.
[0770] 13. The method of clause 12, wherein the codec tool comprises a local illumination compensation (LIC) model that uses a linear model of illumination variations in the current video block during the conversion, and the LIC model is applied based on parameters.
[0771] 14. The method of clause 12, wherein the video region comprises a current slice, slice group, or picture.
[0772] 15. The method of clause 13, wherein the LIC model is applied to the prediction block in the second domain, and wherein the first set of sample points and the second set of sample points are in the second domain.
[0773] 16. The method of clause 13, wherein the reference block is converted to a second domain and the LIC model is applied to the prediction block in the second domain.
[0774] 17. The method of clause 15, wherein the first set of sample points and the second set of sample points are converted to the second domain before being used to derive the parameters.
[0775] 18. The method of clause 17, wherein the second set of samples comprises a reference sample in a reference picture and adjacent and / or non-adjacent samples of the reference sample.
[0776] 19. The method of clause 13, wherein the LIC model is applied to the prediction block in a first domain, and wherein the first set of sample points and the second set of sample points are in the first domain.
[0777] 20. The method of clause 13, wherein the reference block is maintained in the first domain and the LIC model is applied to the prediction block in the first domain.
[0778] 21. The method of clause 19, wherein the first set of sample points is converted to the first domain before being used to derive the parameters.
[0779] 22. The method of clause 21, wherein the first set of samples comprises spatially adjacent and / or non-adjacent samples of the current video block.
[0780] 23. A method according to clause 12, wherein the fields used to derive the parameters are used to apply the parameters to the prediction block.
[0781] 24. The method of clause 13, wherein the LIC model is applied to the prediction block in the second domain.
[0782] 25. A method according to clause 20 or 21, wherein after the LIC model is applied to the prediction block in the first domain, a final prediction block depending on the prediction block is converted to the second domain.
[0783] 26. A method according to any of clauses 1-25, wherein the first domain is a raw domain and the second domain is a shaped domain using a Luma Mapping and Chroma Scaling (LMCS) method that maps luma samples to specific values.
[0784] 27. The method of clause 26, wherein the LMCS uses a piecewise linear model to map luma samples to specific values.
[0785] 28. A method according to any of clauses 1-27, wherein performing the conversion comprises generating a codec representation from the current block.
[0786] 29. A method according to any of clauses 1-27, wherein performing the conversion comprises generating the current block from a codec representation.
[0787] 30. An apparatus in a video system comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of clauses 1 to 29.
[0788] 31. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any of clauses 1 to 29.
[0789] The second set of clauses describes certain features and aspects of the disclosed technology listed in previous sections, including, for example, Examples 3-5, 8, and 15.
[0790] 1. A method for video processing, comprising: determining, for a current video block of a current video region of a video, parameters of a codec mode of a current video block based on one or more parameters of a codec mode of a previous video region; and performing codecs on the current video block to generate a codec representation of the video based on the determination, wherein the parameters of the codec mode are included in a parameter set in the codec representation of the video, and wherein performing the codecs comprises transforming a representation of the current video block in a first domain to a representation of the current video block in a second domain, and wherein, during performing the codecs using the codec mode, the current video block is constructed based on the first domain and the second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0791] 2. A method for video processing, comprising: receiving a codec representation of a video including a parameter set, wherein the parameter set includes parameter information of a codec mode; and performing decoding of the codec representation by using the parameter information to generate a current video block of a current video region of the video from the codec representation, and wherein the parameter information of the codec mode is based on one or more parameters of a codec mode of a previous video region, wherein, in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0792] 3. A method according to clause 1 or 2, wherein the parameter set is different from the slice group header.
[0793] 4. A method according to clause 1 or 2, wherein the parameter set is an adaptive parameter set (APS).
[0794] 5. The method according to clause 1 or 2, wherein the current video region comprises a slice of a video picture of the video or a video picture of the video;
[0795] 6. A method according to clause 1 or 2, wherein the previous video region comprises one or more slices of a picture.
[0796] 7. The method of clause 1 or 2, wherein the previous video region comprises one or more video pictures of the video.
[0797] 8. A method for video processing, comprising: performing a conversion between a current video block of a video and a codec representation of the video, and wherein the conversion comprises applying a filtering operation to a prediction block in a first domain or in a second domain different from the first domain.
[0798] 9. A method according to clause 8, wherein a filtering operation is performed on the prediction block in the first domain to generate a filtered prediction signal, a codec mode is applied to the filtered prediction signal to generate a shaped prediction signal in the second domain, and the current video block is constructed using the shaped prediction signal.
[0799] 10. A method according to clause 8, wherein the coding mode is applied to the prediction block prior to application of the filtering operation to generate a shaped prediction signal in the second domain, and the filtering operation is performed using the shaped prediction signal to generate a filtered prediction signal, and the current video block is constructed using the filtered prediction signal.
[0800] 11. Method according to clause 9 or 10, wherein, in the coding mode, the current video block is constructed based on the first domain and the second domain, and / or the chroma residual is scaled in a luma-dependent manner.
[0801] 12. A method according to any of clauses 8-11, wherein the filtering operation comprises a diffusion filter.
[0802] 13. A method according to any of clauses 8-11, wherein parameters associated with the filtering operation depend on whether the filtering operation is applied to the block in the first domain or the second domain.
[0803] 14. A method according to clause 8, wherein the conversion further comprises: applying motion compensated prediction to the current video block to obtain a prediction signal before applying the filtering operation; applying the codec mode to the filtered prediction signal after applying the filtering operation to generate a shaped prediction signal, wherein the filtered prediction signal is generated by applying the filtering operation to the prediction signal; and constructing the current video block using the shaped prediction signal.
[0804] 15. A method according to clause 8, wherein the conversion further comprises: applying motion compensated prediction to the current video block to obtain a prediction signal before applying the filtering operation; applying a codec mode to the prediction signal to generate a shaped prediction signal; and after applying the filtering operation, constructing the current video block using the filtered shaped prediction signal, wherein the filtered shaped prediction signal is generated by applying the filtering operation to the shaped prediction signal.
[0805] 16. A method for video processing, comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein during the conversion a final reconstructed block is determined for the current video block, and wherein the temporary reconstructed block is generated using a prediction method and represented in a second domain.
[0806] 17. A method according to clause 16, wherein the converting further comprises: applying motion compensated prediction to the current video block to obtain a prediction signal; applying forward shaping to the prediction signal to generate a shaped prediction signal for generating a temporary reconstructed block; and applying inverse shaping to the temporary reconstructed block to obtain an inverse reconstructed block, and wherein filtering is applied to the inverse reconstructed block to generate the final reconstructed block.
[0807] 18. A method according to clause 16, wherein the converting further comprises: applying motion compensated prediction to the current video block to obtain a prediction signal; applying forward shaping to the prediction signal to generate a shaped prediction signal for generating a temporary reconstructed block; applying inverse shaping to a filtered reconstructed block to obtain a final reconstructed block, and wherein the filtered reconstructed block is generated by applying filtering to the temporary reconstructed block.
[0808] 19. A method according to any of clauses 16 to 18, wherein the converting further comprises applying a luma-dependent residual chroma scaling (LMCS) process that maps luma samples to specific values.
[0809] 20. A method according to clause 16, wherein the filter is applied to a temporarily reconstructed block in the first domain, the temporarily reconstructed block in the second domain is first converted to the first domain using an inverse shaping process before application of the filter, and the final reconstructed block depends on the filtered temporarily reconstructed block.
[0810] 21. A method according to clause 16, wherein the filter is applied directly to the temporary reconstructed block in the second domain and thereafter an inverse reshaping operation is applied to generate the final reconstructed block.
[0811] 22. The method of clause 16, wherein the filter comprises a bilateral filter (BF) or a Hadamard transform domain filter (HF).
[0812] 23. The method of clause 16, wherein the filter comprises a deblocking filter (DBF) process, a sample adaptive offset (SAO) filtering process, or an adaptive loop filter (ALF) filtering process.
[0813] 24. A method according to any of clauses 1-23, wherein filter parameters for the filtering operation or filter depend on whether a codec mode is enabled for the current video block, wherein, in the codec mode, the current video block is constructed based on the first domain and the second domain, and / or the chroma residual is scaled in a luma-dependent manner.
[0814] 25. A method according to any of clauses 1-25, wherein the first domain is a raw domain and the second domain is a shaped domain using a Luma Mapping and Chroma Scaling (LMCS) method that maps luma samples to specific values.
[0815] 26. The method of clause 25, wherein the LMCS uses a piecewise linear model to map luma samples to specific values.
[0816] 27. A method according to any of clauses 8-26, wherein performing the conversion comprises generating a codec representation from the current block.
[0817] 28. A method according to any of clauses 8-26, wherein performing the conversion comprises generating the current block from a codec representation.
[0818] 29. An apparatus in a video system comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of clauses 1 to 28.
[0819] 30. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any of clauses 1 to 28.
[0820] The third set of clauses describes certain features and aspects of the disclosed technology listed in previous sections, including, for example, Example 6.
[0821] 1. A video processing method, comprising: performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein, in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein a parameter set in the codec representation includes parameter information of the codec mode.
[0822] 2. The method of clause 1, wherein the parameter set is distinct from a slice group header.
[0823] 3. The method of clause 2, wherein the parameter set is an adaptive parameter set (APS).
[0824] 4. The method of clause 3, wherein the APS of the codec mode information is named Luma Mapping and Chroma Scaling (LMCS) APS.
[0825] 5. A method as described in clause 3, wherein an identifier of the APS to be used for the current video block is contained in a codec representation of the video.
[0826] 6. The method of clause 5, wherein whether the identifier is present in the codec representation of the video depends on whether a codec mode is enabled for the video region.
[0827] 7. The method of clause 3, wherein the parameter set comprises an identifier of the APS.
[0828] 8. The method of clause 1 , wherein the parameter sets are assigned NAL unit type values.
[0829] 9. The method of clause 1, wherein the parameter set identifier ranges from 0 to M, where M is 2 K -1.
[0830] 10. The method of clause 1, wherein the parameter sets are shared across pictures of the video.
[0831] 11. The method of clause 1, wherein the identifier of the parameter set has a value that is a fixed-length codec.
[0832] 12. The method of clause 1, wherein the parameter set identifier is encoded or decoded using an Exponential Golomb (EG) code, a truncated unary code, or a binarized code.
[0833] 13. A method according to clause 1, wherein for two sub-regions within the same picture, the parameter set has an identifier with two different values.
[0834] 14. The method of clause 3, wherein the parameter set and the APS of adaptive loop filter (ALF) information share the same network abstraction layer (NAL) unit type (NUT).
[0835] 15. The method of clause 1, wherein the parameter information is carried with a current APS of adaptive loop filter (ALF) information.
[0836] 16. The method of clause 1, wherein the parameter information is carried in a sequence parameter set (SPS), video parameter set (VPS), picture parameter set (PPS), sequence, header, or picture header.
[0837] 17. The method of clause 1, wherein the parameter information comprises at least one of an indication of shaper model information, use of a codec mode, or a chroma residual scaling factor.
[0838] 18. The method of clause 1, wherein parameter information is signaled in one level.
[0839] 19. The method of clause 1, wherein the parameter information includes the use of a codec mode signaled in the second level.
[0840] 20. A method according to clauses 18 and 19, wherein parameter information is signalled in the APS and usage of the codec mode is signalled in the video region level.
[0841] 21. The method of clause 1, wherein parameter information is parsed in one level.
[0842] 22. The method of clause 1, wherein the parameter information includes usage of a codec mode parsed in the second level.
[0843] 23. A method according to clause 21 or 22, wherein parameter information is parsed in an APS and usage of a codec mode is parsed in a video region level.
[0844] 24. The method of clause 1, wherein predictive coding is applied to codec parameter information having different APS indices.
[0845] 25. A method according to any of clauses 1-24, wherein the first domain is a raw domain and the second domain is a shaped domain using a Luma Mapping and Chroma Scaling (LMCS) method that maps luma samples to specific values.
[0846] 26. The method of clause 25, wherein the LMCS uses a piecewise linear model to map luma samples to specific values.
[0847] 27. A method according to any of clauses 1-26, wherein the video region is a picture or a slice group.
[0848] 28. A method according to any of clauses 1-26, wherein the video region level is a picture header or a slice group header.
[0849] 29. A method according to any of clauses 1-28, wherein the first domain is the original domain and the second domain is a shaped domain using a Luma Mapping and Chroma Scaling (LMCS) method that maps luma samples to specific values according to a shaping model.
[0850] 30. The method of clause 29, wherein the LMCS uses a piecewise linear model to map luma samples to specific values.
[0851] 31. A method according to any of clauses 1-30, wherein performing the conversion comprises generating a codec representation from the current block.
[0852] 32. A method according to any of clauses 1-30, wherein performing the conversion comprises generating the current block from a codec representation.
[0853] 33. An apparatus in a video system comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of clauses 1 to 32.
[0854] 34. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any of clauses 1 to 32.
[0855] The fourth set of clauses describes certain features and aspects of the disclosed technology listed in previous sections, including, for example, Examples 7 and 9.
[0856] 1. A method for video processing, comprising: performing a conversion between a current video block of the video being a chroma block and a codec representation of the video, wherein during the conversion, the current video block is constructed based on a first domain and a second domain, and wherein the conversion further comprises applying a forward shaping process and / or an inverse shaping process to one or more chroma components of the current video block.
[0857] 2. The method of clause 1, further comprising: refraining from applying luma-dependent chroma residual scaling (LCRS) to one or more chroma components of the current video block.
[0858] 3. The method of clause 1, wherein at least one of a piecewise linear (PWL) model, a forward lookup table, or a backward lookup table is used for chroma components.
[0859] 4. The method of clause 3, wherein the PWL model, forward lookup table, and backward lookup table for the chroma components are derived from the PWL model, forward lookup table, and backward lookup table of the corresponding luma component, respectively.
[0860] 5. The method of clause 3, wherein the PWL model is signaled in a sequence parameter set (SPS), video parameter set (VPS), adaptation parameter set (APS), picture parameter set (PPS), sequence header, picture header, slice group header, slice header, codec tree unit (CTU) row, CTU group, or region.
[0861] 6. The method of clause 3, wherein the forward lookup table and the backward lookup table are signaled in a sequence parameter set (SPS), a video parameter set (VPS), an adaptation parameter set (APS), a picture parameter set (PPS), a sequence header, a picture header, a slice group header, a slice header, a codec tree unit (CTU) row, a CTU group, or a region.
[0862] 7. A method for video processing, comprising: performing a conversion between a current video chroma block of a video and a codec representation of the video, wherein performing the conversion comprises: determining based on a rule whether luma-dependent chroma residual scaling (LCRS) is enabled or disabled, and reconstructing the current video chroma block based on the determination.
[0863] 8. The method of clause 7, wherein the rules specify disabling of LCRS for certain block sizes, temporal layers, slice group types, picture types, codec modes, and certain types of motion information.
[0864] 9. The method of clause 7, wherein the rules specify that LCRS is disabled for chroma blocks, and the forward and / or inverse shaping process is applied to the corresponding luma blocks.
[0865] 10. The method of clause 7, wherein the rule specifies that LCRS is applied to chroma blocks and that forward and / or inverse shaping processes are not applied to corresponding luma blocks.
[0866] 11. The method of clause 7, wherein the rule specifies disabling LCRS for a current video chroma block encoded using a Cross Component Linear Model (CCLM).
[0867] 12. The method of clause 7, wherein the rule specifies disabling LCRS for current video chroma blocks that are not coded using Cross Component Linear Model (CCLM).
[0868] 13. The method of clause 7, wherein the rule specifies that disabling LCRS is based on a size of a video block exceeding a virtual pipe data unit (VPDU).
[0869] 14. The method of clause 13, wherein LCRS is not allowed if the video block contains fewer than M*H samples of video samples.
[0870] 15. The method of clause 13, wherein LCRS is not allowed if the minimum dimension of the width and / or height of the video block is less than or equal to a specific value.
[0871] 16. The method of clause 13, wherein LCRS is not allowed if the minimum dimension of the width and / or height of the video block is not less than a specific value.
[0872] 17. The method according to clause 15 or 16, wherein the specific value is 8.
[0873] 18. The method of clause 13, wherein LCRS is not allowed if the width of the video block is equal to or greater than a first value and / or the height of the video block is equal to or greater than a second value.
[0874] 19. The method of clause 13, wherein LCRS is not allowed if the width of the video block is equal to or less than a first value and / or the height of the video block is equal to or less than a second value.
[0875] 20. The method of clause 18 or 19, wherein at least one of the first value or the second value is 8, 64 or 128.
[0876] 21. A method according to any of clauses 13-20, wherein the video block is a luma block or a chroma block.
[0877] 22. A method according to any of clauses 1-21, wherein the first domain is the original domain and the second domain is a shaped domain using a Luma Mapping and Chroma Scaling (LMCS) method that maps luma samples to specific values according to a shaping model.
[0878] 23. The method of clause 22, wherein the LMCS uses a piecewise linear model to map luma samples to specific values.
[0879] 24. A method according to any of clauses 1-23, wherein the chroma residual is scaled in a luma-dependent manner by performing a luma-dependent chroma residual scaling operation, wherein the luma-dependent chroma residual scaling operation comprises scaling the chroma residual before being used to derive the reconstruction of the video chroma block, and the scaling parameters are derived from the luma samples.
[0880] 25. A method according to any of clauses 1-24, wherein performing the conversion comprises generating a codec representation from the current block.
[0881] 26. A method according to any of clauses 1-24, wherein performing the conversion comprises generating the current block from a codec representation.
[0882] 27. An apparatus in a video system comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of clauses 1 to 26.
[0883] 28. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any of clauses 1 to 26.
[0884] The fifth set of clauses describes certain features and aspects of the disclosed technology listed in previous sections, including, for example, Examples 10-14, 28, 29, and 40.
[0885] 1. A method for video processing, comprising: for a conversion between a current video block of a video and a codec representation of the video, determining, based on one or more coefficient values of the current video block, whether to disable use of a codec mode; and performing the conversion based on the determination, wherein during the conversion using the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0886] 2. A method according to clause 1, wherein the shaping process comprises: selectively applying at least one of a forward shaping process to samples in a first domain, wherein the samples in the first domain are then converted to samples in a second domain; and selectively applying an inverse shaping process to samples in the second domain, wherein the samples in the second domain are then converted to a representation in the first domain.
[0887] 3. The method of clause 1 or 2, wherein the shaping process further comprises selectively applying a luma-dependent chroma residual scaling process.
[0888] 4. A method according to any of clauses 1-3, wherein the determination is based on whether the current video block is encoded with all-zero coefficients.
[0889] 5. The method of clause 2, wherein the forward shaping process is skipped based on whether the current video block is encoded with all-zero coefficients.
[0890] 6. The method of clause 2, wherein the current video block is coded with all-zero coefficients, and wherein the inverse shaping process is skipped.
[0891] 7. The method of clause 2, wherein the current video block is coded with all-zero coefficients, and wherein a luma-dependent chroma residual scaling process is skipped.
[0892] 8. The method of clause 2, wherein the determination is based on whether the current video block is encoded with only one non-zero coefficient located at a particular position.
[0893] 9. The method of clause 2, wherein the current video block is encoded with only one non-zero coefficient located at a specific position, and at least one of a forward shaping process, an inverse shaping process, or a luma-dependent chroma residual scaling process is skipped.
[0894] 10. The method of clause 2, wherein the determination is based on whether the current video block is encoded with M non-zero coefficients.
[0895] 11. The method of clause 2, wherein the current video block is encoded with M non-zero coefficients and at least one of a forward shaping process, an inverse shaping process, or a luma-dependent chroma residual scaling process is skipped.
[0896] 12. The method of clause 11, wherein M is 1.
[0897] 13. A method of video processing, comprising: for conversion between current video blocks of a video exceeding a virtual pipe data unit (VPDU) of the video, dividing the current video block into regions; and performing the conversion by separately applying a codec mode to each region, wherein, during the conversion by applying the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0898] 14. The method of clause 13, wherein each region corresponds to a separate codec unit (CU) of a codec mode.
[0899] 15. The method of clause 13, wherein a width of the current video block is equal to or greater than a first value, the current video block is divided into sub-blocks having one or more widths equal to or less than the first value, and a codec mode is enabled for each sub-block.
[0900] 16. The method of clause 13, wherein the height of the current video block is equal to or greater than a second value, the current video block is divided into sub-blocks having one or more heights equal to or less than the second value, and the codec mode is enabled for each sub-block.
[0901] 17. The method of clause 13, wherein a size of the current video block is equal to or greater than a third value, the current video block is divided into sub-blocks having one or more sizes equal to or less than the third value, and a codec mode is enabled for each sub-block.
[0902] 18. A method according to any of clauses 15-17, wherein the sub-blocks have the same width or the same height.
[0903] 19. A method for video processing, comprising: for a conversion between a current video block of a video and a codec representation of the video, determining whether to disable use of a codec mode based on a size or color format of the current video block; and performing the conversion based on the determination, wherein, during the conversion using the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0904] 20. The method of clause 19, wherein the determining determines to disable a codec mode for a current video block that exceeds a virtual pipe data unit (VPDU).
[0905] 21. The method of clause 19, wherein the determining determines to disable the codec mode for a current video block having a size containing a number of samples less than M*H.
[0906] 22. The method of clause 19, wherein the determining determines to disable the codec mode for the current video block if a minimum dimension of the width and / or height of the current video block is equal to or less than X as an integer.
[0907] 23. The method of clause 19, wherein the determining determines to disable the codec mode for the current video block if a minimum dimension of the width and / or height of the current video block is not less than X which is an integer.
[0908] 24. The method according to clause 22 or 23, wherein X is 8.
[0909] 25. The method of clause 19, wherein the determining determines to disable the codec mode for the current video block if the current video block has a width equal to or greater than a first value and / or a height equal to or greater than a second value.
[0910] 26. The method of clause 19, wherein the determining determines to disable the codec mode for the current video block if the current video block has a width equal to or less than a first value and / or a height equal to or less than a second value.
[0911] 27. The method of clause 25 or 26, wherein at least one of the first value or the second value is 8.
[0912] 28. A method according to any one of clauses 19 to 27, wherein disabling the codec mode includes disabling at least one of: 1) forward shaping that converts samples in the first domain to the second domain; 2) backward shaping that converts samples in the second domain to the first domain; 3) luminance-dependent chroma residual scaling.
[0913] 29. A method for video processing, comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein at least one syntax element in the codec representation provides an indication of use of the codec mode and an indication of a shaper model.
[0914] 30. The method of clause 29, wherein the indication of use of the codec mode is encoded based on an indication of a shaper model.
[0915] 31. The method of clause 29, wherein the indication of the shaper model is encoded based on the indication of the encoding mode.
[0916] 32. A method as described in clause 29, wherein only one of the syntax elements is encoded or decoded.
[0917] 33. A method according to any of clauses 1-32, wherein different pruning methods are applied to the predicted signal and the reconstructed signal.
[0918] 34. A method according to clause 33, wherein adaptive cropping allowing different cropping parameters within the video is applied to the prediction signal.
[0919] 35. The method of clause 34, wherein the maximum and minimum values of the adaptive clipping are defined in the second domain.
[0920] 36. The method of clause 33, wherein fixed clipping is applied to the reconstructed signal.
[0921] 37. A method for video processing, comprising: determining that a codec mode is disabled for conversion between a current video block of a video and a codec representation of the video; and conditionally skipping forward shaping and / or inverse shaping based on the determination, wherein, in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0922] 38. A method according to clause 37, wherein at least one of the prediction signal, the reconstruction signal or the residual signal is in the first domain.
[0923] 39. The method of clause 37, wherein at least one of the prediction signal, the reconstruction signal or the residual signal is in the second domain.
[0924] 40. A method for video processing, comprising: performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein in the codec mode the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein a plurality of forward shapings and / or a plurality of inverse shapings are applied in the shaping mode for the video region.
[0925] 41. The method of clause 40, wherein the video region comprises a picture, a slice group, a virtual pipe data unit (VPDU), a codec tree unit (CTU), a line, or a plurality of codec units.
[0926] 42. A method according to clause 40 or 41, wherein how to select multiple forward reshaping and / or multiple backward reshaping depends on at least one of the following: i) block size or video region size, ii) codec mode of the current video block or video region, iii) picture type of the current video block or video region, iv) low delay check flag of the current video block or video region, v) motion information of the current video block or video region, vi) reference picture of the current video block or video region, or vii) video content of the current video block or video region.
[0927] 43. A method according to any of clauses 1 to 42, wherein during the conversion, samples in the first domain are mapped to samples in the second domain, wherein the values of the samples in the second domain are smaller than the values of the samples in the first domain.
[0928] 44. A method according to any of clauses 1 to 43, wherein the first domain is a raw domain and the second domain is a shaped domain using a Luma Mapping and Chroma Scaling (LMCS) method to map luma samples to specific values.
[0929] 45. The method of clause 44, wherein the LMCS uses a piecewise linear model to map luma samples to specific values.
[0930] 46. A method according to any of clauses 1-45, wherein performing the conversion comprises generating a codec representation from the current block.
[0931] 47. A method according to any of clauses 1-45, wherein performing the conversion comprises generating the current block from a codec representation.
[0932] 48. An apparatus in a video system comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of clauses 1 to 47.
[0933] 49. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any of clauses 1 to 47.
[0934] The sixth set of clauses describes certain features and aspects of the disclosed technology listed in previous sections, including, for example, Examples 16 and 17.
[0935] 1. A method of video processing, comprising: determining a codec mode that is enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using a palette mode, wherein in the palette mode, a palette of at least representative sample values is used for the current video block, and wherein, in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or chroma residuals are scaled in a luma-dependent manner.
[0936] 2. The method of clause 1, wherein the palette of representative sample values comprises at least one of: 1) palette prediction values, or 2) escape samples.
[0937] 3. The method of clause 1, wherein the representative sample value represents a value in the first domain.
[0938] 4. The method of clause 1, wherein the representative sample value represents a value in the second domain.
[0939] 5. A method according to clause 1 or 2, wherein the palette prediction value used in palette mode and included in the codec representation is in the first domain or in the second domain.
[0940] 6. A method according to clause 1 or 2, wherein the escape samples used in palette mode and included in the codec representation are in the first domain or in the second domain.
[0941] 7. A method according to clause 1 or 2, wherein, when the palette prediction values and / or escape samples used in palette mode and included in the codec representation are in the second domain, a first reconstructed block in the second domain is first generated and used for encoding and decoding the subsequent block.
[0942] 8. A method according to clause 7, wherein when the palette prediction values and / or escape samples used in the modified palette mode and included in the codec representation are in the second domain, a final reconstructed block in the first domain is generated using the first reconstructed block and the inverse shaping process.
[0943] 9. The method of clause 8, wherein the inverse reshaping process is invoked just before the deblocking filter process.
[0944] 10. A method according to any of clauses 1-9, wherein the conversion is performed based on color components of the current video block.
[0945] 11. The method of clause 10, wherein the color component is a luminance component.
[0946] 12. A method of video processing, comprising: for a conversion between a current video block of a video and a codec representation of the video, determining that the current video block is encoded in a palette mode, wherein in the palette mode, a palette of at least representative sample values is used to encode and decode the current video block; and performing the conversion by disabling the codec mode due to the determination, wherein, when the codec mode is applied to the video block, the video block is constructed based on chroma residuals scaled in a luma-dependent manner.
[0947] 13. The method of clause 12, wherein the encoding mode is disabled when the current video block is encoded in palette mode.
[0948] 14. A method of video processing, comprising: performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion uses a first codec mode and a palette codec mode, wherein in the palette codec mode, a palette of at least representative pixel values is used to encode and decode the current video block; and performing a conversion between a second video block of the video that was encoded and decoded without using the palette codec mode and the codec representation of the video, and wherein the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a manner that depends on luminance, and wherein the first codec mode is applied to the first video block and the second video block in different manners.
[0949] 15. The method of clause 14, wherein the first codec mode applied to the first video block differs from the first codec mode applied to the second video block by disabling use of forward shaping and inverse shaping for converting samples between the first domain and the second domain.
[0950] 16. A method according to clause 14, wherein the first codec mode applied to the first video block is different from the first codec mode applied to the second video block due to the use of different reshaping and / or different inverse reshaping functions for converting samples between the first domain and the second domain.
[0951] 17. A method according to any of clauses 1-11 and 14-16, wherein the first domain is a raw domain and the second domain is a shaped domain using a Luma Mapping and Chroma Scaling (LMCS) method that maps luma samples to specific values.
[0952] 18. The method of clause 17, wherein the LMCS uses a piecewise linear model to map luma samples to specific values.
[0953] 19. A method according to any of clauses 1-18, wherein performing the conversion comprises generating a codec representation from the current block.
[0954] 20. A method according to any of clauses 1-18, wherein performing the conversion comprises generating the current block from a codec representation.
[0955] 21. An apparatus in a video system comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of clauses 1 to 21.
[0956] 22. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any of clauses 1 to 21.
[0957] This seventh set of clauses describes certain features and aspects of the disclosed technology listed in previous sections, including, for example, Examples 18 and 19.
[0958] 1. A method of video processing, comprising: determining a codec mode that is enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using an intra block copy mode, wherein the intra block copy mode uses at least a block vector pointing to a picture including the current video block to generate a prediction block, and wherein, in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0959] 2. The method of clause 1, wherein the prediction block is generated in the first domain.
[0960] 3. A method according to clause 1, wherein the residual block is represented in a codec representation in the first domain.
[0961] 4. The method of clause 1, wherein the prediction block is generated in the second domain.
[0962] 5. A method according to clause 1, wherein the residual block is represented in a codec representation in the second domain.
[0963] 6. A method according to clause 4 or 5, wherein a first building block of the current video block is obtained based on a sum of a residual block and a prediction block in the second domain, and the first building block is used for conversion between subsequent video blocks and codec representations of the video.
[0964] 7. A method according to clause 4 or 5, wherein the final building block of the current video block is obtained based on an inverse reshaping applied to the first building block to convert the first building block from the second domain to the first domain.
[0965] 8. A method according to any of clauses 1-7, wherein the conversion is performed based on color components of the current video block.
[0966] 9. The method of clause 8, wherein the color component is a luminance component.
[0967] 10. A method for video processing, comprising: for a conversion between a current video block of a video and a codec representation of the video, determining that the current video block is encoded in an intra block copy (IBC) mode, wherein the intra block copy mode uses at least a block vector pointing to a video frame containing the current video block to generate a prediction block for encoding and decoding the current video block; and performing the conversion by disabling the codec mode due to the determination, wherein, when the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0968] 11. The method of clause 10, wherein the codec mode is disabled when the current video block is coded in IBC mode.
[0969] 12. A method for video processing, comprising: performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion uses an intra block copy mode and a first codec mode, wherein the intra block copy mode uses at least a block vector pointing to a video frame containing a current video block to generate a prediction block; and performing a conversion between a second video block of the video that was encoded and decoded without using the intra block copy mode and a codec representation of the video, wherein the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein the first codec mode is applied differently to the first video block and the second video block.
[0970] 13. The method of clause 12, wherein the first codec mode applied to the first video block is different from the first codec mode applied to the second video block by disabling use of forward shaping and inverse shaping for converting samples between the first domain and the second domain.
[0971] 14. The method of clause 12, wherein the first codec mode applied to the first video block is different from the first codec mode applied to the second video block due to the use of different forward shaping and / or different inverse shaping for converting samples between the first domain and the second domain.
[0972] 15. A method according to any of clauses 1-14, wherein the first domain is a raw domain and the second domain is a shaped domain using a Luma Mapping and Chroma Scaling (LMCS) method that maps luma samples to specific values.
[0973] 16. The method of clause 15, wherein the LMCS uses a piecewise linear model to map luma samples to specific values.
[0974] 17. A method according to any of clauses 1-16, wherein performing the conversion comprises generating a codec representation from the current block.
[0975] 18. A method according to any of clauses 1-16, wherein performing the conversion comprises generating the current block from a codec representation.
[0976] 19. An apparatus in a video system comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of clauses 1 to 18.
[0977] 20. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any of clauses 1 to 18.
[0978] The eighth set of clauses describes certain features and aspects of the disclosed technology listed in the previous sections, including, for example, Examples 20-27.
[0979] 1. A method of video processing, comprising: determining a codec mode enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using a block-based delta pulse codec modulation (BDPCM) mode, wherein in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0980] 2. The method of clause 1, wherein the prediction block for the current video block is generated in the first domain.
[0981] 3. The method of clause 1, wherein the residual block of the current video block is represented in a codec representation in the first domain.
[0982] 4. The method of clause 1, wherein the prediction block for the current video block is generated in the second domain.
[0983] 5. The method of clause 1, wherein the residual block of the current video block is represented in a codec representation in the second domain.
[0984] 6. A method according to clause 4 or 5, wherein a first building block of the current video block is obtained based on a sum of a residual block and a prediction block in the second domain, and the first building block is used for conversion between subsequent video blocks and codec representations of the video.
[0985] 7. A method according to clause 4 or 5, wherein the final building block of the current video block is obtained based on an inverse reshaping applied to the first building block to convert the first building block from the second domain to the first domain.
[0986] 8. A method of video processing, comprising: for a conversion between a current video block of a video and a codec representation of the video, determining that the current video block is coded using a block-based delta pulse codec modulation (BDPCM) mode; and performing the conversion by disabling a codec mode due to the determination, wherein, when the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0987] 9. The method of clause 8, wherein the codec mode is disabled when the current video block is encoded in BDPCM mode.
[0988] 10. A method of video processing, comprising: performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and a block-based delta pulse codec modulation (BDPCM) mode; and performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded and decoded without using the BDPCM mode, and the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein the first codec mode is applied differently to the first video block and the second video block.
[0989] 11. The method of clause 10, wherein the first codec mode applied to the first video block is different from the first codec mode applied to the second video block by disabling use of forward shaping and inverse shaping for converting samples between the first domain and the second domain.
[0990] 12. The method of clause 10, wherein the first codec mode applied to the first video block is different from the first codec mode applied to the second video block due to using different forward shaping and / or different inverse shaping for converting samples between the first domain and the second domain.
[0991] 13. A method of video processing, comprising: determining a codec mode that is enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using a transform skip mode, wherein in the transform skip mode, a transform on a prediction residual is skipped when encoding and decoding the current video block, wherein in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0992] 14. The method of clause 13, wherein the prediction block for the current video block is generated in the first domain.
[0993] 15. The method of clause 13, wherein the residual block of the current video block is represented in a codec representation in the first domain.
[0994] 16. The method of clause 13, wherein the prediction block for the current video block is generated in the second domain.
[0995] 17. The method of clause 13, wherein the residual block is represented in a codec representation in the second domain.
[0996] 18. A method according to clause 16 or 17, wherein the first building block of the current video block is obtained by summing a residual block and a prediction block based on the second domain, and the first building block is used for conversion between subsequent video blocks and codec representations of the video.
[0997] 19. The method of clause 16 or 17, wherein the final building block of the current video block is obtained based on an inverse reshaping applied to the first building block to convert the first building block from the second domain to the first domain.
[0998] 20. A method of video processing, comprising: for a conversion between a current video block of a video and a codec representation of the video, determining that the current video block is encoded in a transform skip mode, wherein in the transform skip mode, a transform on a prediction residual is skipped when encoding and decoding the current video block; and performing the conversion by disabling a codec mode due to the determination, wherein, when the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[0999] 21. The method of clause 20, wherein the codec mode is disabled when the current video block is coded in transform skip mode.
[1000] 22. A method of video processing, comprising: performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and a transform skip mode, wherein in the transform skip mode, a transform of a prediction residual is skipped when encoding and decoding a current video block; and performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded and decoded without using the transform skip mode, and the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein the first codec mode is applied differently to the first video block and the second video block.
[1001] 23. The method of clause 22, wherein the first codec mode applied to the first video block is different from the first codec mode applied to the second video block by disabling use of forward shaping and inverse shaping for converting samples between the first domain and the second domain.
[1002] 24. The method of clause 22, wherein the first codec mode applied to the first video block differs from the first codec mode applied to the second video block due to using different forward shaping and / or different inverse shaping for converting samples between the first domain and the second domain.
[1003] 25. A method of video processing, comprising: determining that a codec mode is enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using an intra-frame pulse codec modulation mode, wherein in the intra-frame pulse codec modulation mode, the current video block is encoded without applying a transform and a transform domain quantization, wherein in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[1004] 26. The method of clause 25, wherein the prediction block for the current video block is generated in the first domain.
[1005] 27. The method of clause 25, wherein the residual block of the current video block is represented in a codec representation in the first domain.
[1006] 28. The method of clause 25, wherein the prediction block for the current video block is generated in the second domain.
[1007] 29. The method of clause 25, wherein the residual block is represented in a codec representation in the second domain.
[1008] 30. A method according to clause 28 or 29, wherein a first building block of the current video block is obtained based on a sum of a residual block and a prediction block in the second domain, and the first building block is used for conversion between subsequent video blocks and codec representations of the video.
[1009] 31. A method according to clause 28 or 29, wherein the final building block of the current video block is obtained based on an inverse reshaping applied to the first building block to convert the first building block from the second domain to the first domain.
[1010] 32. A method of video processing, comprising: for a conversion between a current video block of a video and a codec representation of the video, determining that the current video block is coded in an intra-frame pulse codec modulation mode, wherein in the intra-frame pulse codec modulation mode, the current video block is coded without applying a transform and a transform domain quantization; and due to the determination, performing the conversion by disabling the codec mode, wherein, when the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[1011] 33. The method of clause 32, wherein the codec mode is disabled when the current video block is coded in intra pulse codec modulation mode.
[1012] 34. A method of video processing, comprising: performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and an intra-frame pulse codec modulation mode, wherein in the intra-frame pulse codec modulation mode, the current video block is encoded and decoded without applying a transform and a transform domain quantization; and performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded and decoded without using the intra-frame pulse codec modulation mode, and the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on the first domain and the second domain, and / or the chroma residual is scaled in a manner dependent on the luminance, and wherein the first codec mode is applied to the first video block and the second video block in a different manner.
[1013] 35. The method of clause 34, wherein the first codec mode applied to the first video block differs from the first codec mode applied to the second video block by disabling use of forward shaping and inverse shaping for converting samples between the first domain and the second domain.
[1014] 36. A method as recited in clause 34, wherein the first codec mode applied to the first video block differs from the first codec mode applied to the second video block due to the use of different forward shaping and / or different inverse shaping for converting samples between the first domain and the second domain.
[1015] 37. A method of video processing, comprising: determining that a codec mode is enabled for conversion between a current video block of a video and a codec representation of the video; and performing the conversion using a modified transform and quantization bypass mode, wherein in the modified transform and quantization bypass mode, the current video block is losslessly coded without transform and quantization, wherein in the codec mode, the current video block is constructed based on samples in a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[1016] 38. The method of clause 37, wherein the prediction block for the current video block is generated in the first domain.
[1017] 39. The method of clause 37, wherein the residual block of the current video block is represented in a codec representation in the first domain.
[1018] 40. The method of clause 37, wherein the prediction block for the current video block is generated in the second domain.
[1019] 41. The method of clause 37, wherein the residual block is represented in a codec representation in the second domain.
[1020] 42. A method according to clause 40 or 41, wherein a first building block of the current video block is obtained based on a sum of a residual block and a prediction block in the second domain, and the first building block is used for conversion between subsequent video blocks and codec representations of the video.
[1021] 43. A method according to clause 40 or 41, wherein the final building block of the current video block is obtained based on an inverse reshaping applied to the first building block to convert the first building block from the second domain to the first domain.
[1022] 44. A method of video processing, comprising: for a conversion between a current video block of a video and a codec representation of the video, determining that the current video block is coded in a transform and quantization bypass mode, wherein in the transform and quantization bypass mode, the current video block is losslessly coded without transform and quantization; and performing the conversion by disabling a codec mode due to the determination, wherein, when the codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner.
[1023] 45. The method of clause 44, wherein the codec mode is disabled when the current video block is coded in intra pulse codec modulation mode.
[1024] 46. A method of video processing, comprising: performing a conversion between a first video block of a video and a codec representation of the video, wherein the conversion of the first video block uses a first codec mode and a transform quantization bypass mode, wherein in the transform quantization bypass mode, the current video block is losslessly encoded and decoded without transform and quantization; and performing a conversion between a second video block of the video and a codec representation of the video, wherein the second video block is encoded and decoded without using the transform quantization bypass mode, and the conversion of the second video block uses the first codec mode, wherein when the first codec mode is applied to the video block, the video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein the first codec mode is applied to the first video block and the second video block in different manners.
[1025] 47. The method of clause 46, wherein the first codec mode applied to the first video block differs from the first codec mode applied to the second video block by disabling use of forward shaping and inverse shaping for converting samples between the first domain and the second domain.
[1026] 48. A method as recited in clause 46, wherein the first codec mode applied to the first video block differs from the first codec mode applied to the second video block due to the use of different forward shaping and / or different inverse shaping for converting samples between the first and second domains.
[1027] 49. A method according to any of clauses 1 to 48, wherein the conversion is performed based on the colour components of the current video block.
[1028] 50. The method of clause 49, wherein the color component is a luminance component.
[1029] 51. A method according to any of clauses 1-50, wherein the first domain is a raw domain and the second domain is a shaped domain using a Luma Mapping and Chroma Scaling (LMCS) method that maps luma samples to specific values.
[1030] 52. The method of clause 51, wherein the LMCS maps luma samples to specific values using a piecewise linear model.
[1031] 53. A method according to any of clauses 1 to 52, wherein performing the conversion comprises generating a codec representation from the current block.
[1032] 54. A method according to any of clauses 1 to 52, wherein performing the conversion comprises generating the current block from a codec representation.
[1033] 55. An apparatus in a video system comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of clauses 1 to 54.
[1034] 56. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any of clauses 1 to 54.
[1035] The ninth clause aggregates descriptions of certain features and aspects of the disclosed technology listed in the previous sections, including, for example, Examples 30-34 and 41.
[1036] 1. A method of video processing, comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein information for the codec mode is signaled in a parameter set other than a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), or an adaptation parameter set (APS) for carrying adaptive loop filtering (ALF) parameters.
[1037] 2. The method of clause 1 , wherein the parameter sets are shared across pictures.
[1038] 3. The method of clause 1 , wherein the parameter set comprises one or more syntax elements, wherein the one or more syntax elements comprise at least one of an identifier of the parameter set or a flag indicating the presence of extension data for the parameter set.
[1039] 4. The method of clause 1, wherein the parameter set is specific to a slice group within a picture.
[1040] 5. A method of video processing, comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein information for the codec mode is signaled together with adaptive loop filtering (ALF) information in an adaptation parameter set (APS), wherein the information for the codec mode and the ALF information are included in one NAL unit.
[1041] 6. The method of clause 5, wherein the identifier of the APS is signaled in a slice group header.
[1042] 7. A method of video processing, comprising: performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein information for the codec mode is signaled in an adaptive parameter set (APS) of a first type that is different from a second type of APS used to signal adaptive loop filtering (ALF) information.
[1043] 8. The method of clause 7, wherein the identifier of the second type of APS is signaled in the video region level.
[1044] 9. The method of clause 7, wherein the identifier of the first type of APS is signaled in the video region level.
[1045] 10. The method of clause 7, wherein the first type of APS contained in the codec representation comprises a second type of APS, wherein the second type of APS comprises ALF information in the conforming bitstream.
[1046] 11. The method of clause 7, wherein the second type of APS contained in the codec representation comprises an APS of the first type, wherein the APS of the first type comprises information for a codec mode in the conforming bitstream.
[1047] 12. The method of clause 7, wherein the first type of APS and the second type of APS are associated with different identifiers.
[1048] 13. The method of clause 12, wherein the APS of the second type has an identifier equal to 2N, N being an integer.
[1049] 14. The method of clause 13, wherein the APS of the first type has an identifier equal to 2N+1, N being an integer.
[1050] 15. A method of video processing, comprising: performing a conversion between a current video block of a video region of a video and a codec representation of the video, wherein the conversion uses a codec mode in which the current video block is constructed based on a first domain and a second domain and / or a chroma residual is scaled in a luma-dependent manner, and wherein the video region is not allowed to reference an adaptation parameter set or a parameter set signaled prior to a data structure of a specified type for processing the video, and wherein the data structure of the specified type is signaled prior to the video region.
[1051] 16. The method of clause 15, wherein the data structure comprises at least one of a network abstraction layer (NAL) unit, a slice group, a sequence parameter set (SPS), a picture parameter set (PPS), an access unit delimiter NAL (AUD), an end of bitstream NAL (EoB), an end of sequence NAL (NAL), an instantaneous decoding refresh (IDR) NAL, a completely random access (CRA) NAL, an intra random access point (IRAP) access unit, an I-slice group, a picture, or a slice.
[1052] 17. A method of video processing, comprising: performing a conversion between a current video block of a video and a codec representation of the video, wherein the conversion uses a codec mode, wherein in the codec mode, the current video block is constructed based on a first domain and a second domain, and / or a chroma residual is scaled in a luma-dependent manner, and wherein syntax elements of a parameter set comprising parameters for processing the video have predefined values in a conforming bitstream.
[1053] 18. The method of clause 17, wherein the predefined values are 0 and 1.
[1054] 19. The method of clause 17, wherein the predefined values are 0 and 7.
[1055] 20. The method of any of clauses 1 to 19, wherein a video region comprises at least one of a slice group, a picture, a slice, or a slice.
[1056] 21. A method according to any of clauses 1 to 20, wherein the first domain is a raw domain and the second domain is a shaped domain using a Luma Mapping and Chroma Scaling (LMCS) method that maps luma samples to specific values.
[1057] 22. The method of clause 21, wherein the LMCS uses a piecewise linear model to map luma samples to specific values.
[1058] 23. A method according to any of clauses 1 to 22, wherein performing the conversion comprises generating a codec representation from the current block.
[1059] 24. A method according to any of clauses 1 to 22, wherein performing the conversion comprises generating the current block from a codec representation.
[1060] 25. An apparatus in a video system comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of clauses 1 to 24.
[1061] 26. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for performing the method according to any of clauses 1 to 24.
[1062] It will be appreciated from the foregoing that specific embodiments of the presently disclosed technology have been described herein for purposes of illustration, but that various modifications may be made without departing from the scope of the invention. Accordingly, the presently disclosed technology is not to be limited except as set forth in the appended claims.
[1063] The embodiments of the subject matter and functional operations described in this patent document can be implemented in various systems, digital electronic circuits, or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. The embodiments of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium, which are used to be executed by a data processing device or to control the operation of the data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances that affect machine-readable propagation signals, or a combination of one or more of them. The term "data processing unit" or "data processing device" includes all devices, equipment and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the device can also include code that creates an execution environment for the computer program in question, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[1064] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[1065] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and the apparatus can also be implemented as, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[1066] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more large-capacity storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or be operably coupled to receive data from the one or more large-capacity storage devices or to transfer data to the one or more large-capacity storage devices, or to receive data from the one or more large-capacity storage devices and to transfer data thereto. However, a computer does not require such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices. The processor and memory can be supplemented by or incorporated into a dedicated logic circuit.
[1067] It is intended that the specification and drawings be considered exemplary only, where exemplary means example. As used herein, the use of "or" is intended to include "and / or" unless the context clearly dictates otherwise.
[1068] Although this patent document contains many details, these details should not be interpreted as limitations on any invention or the scope that may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. Furthermore, although features may be described above as working in certain combinations and even initially claimed as such, one or more features from the claimed combination may be excluded from the combination in some cases, and the claimed combination may be directed to subcombinations or variations of subcombinations.
[1069] Similarly, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown, or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[1070] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for processing video data, comprising: For conversion between a current video block of a video and a bitstream of the video, determining a first prediction block in a first domain from at least one reference picture using unrefined motion information, wherein the current video block is a luma block; applying a motion information refinement process to the unrefined motion information based on at least the first prediction block to obtain motion offset information; generating a second prediction block from samples in the at least one reference picture using the motion offset information; converting the second prediction block from the first domain to a second domain where samples of the second prediction block are mapped to specific values based on a forward mapping process; as well as generating, in the second domain, reconstructed samples of the current video block based on the second prediction block, wherein a piecewise linear model is used to map the samples of the second prediction block to the specific value during the forward mapping process, and The scaling coefficient of the piecewise linear model is determined based on a first variable and a second variable, wherein the first variable is determined based on a syntax element included in an adaptive parameter set, and the second variable is determined based on a bit depth.
2. The method according to claim 1, wherein In the motion information refinement process, the motion offset information is generated based on a bilinear interpolation process and cost calculation, and the motion offset information is also used for temporal motion vector prediction of subsequent blocks.
3. The method according to claim 1, wherein In the motion information refinement process, the motion offset information is generated based on gradient calculations in different directions, wherein the gradient calculations are based on sample points in the first domain.
4. The method according to claim 2, wherein: The cost calculation includes the sum of absolute differences (SAD) or the mean of absolute differences (MR-SAD).
5. The method according to claim 1, wherein The reconstructed samples of the current video block are converted to the first domain based on an inverse mapping process.
6. The method according to claim 5, wherein: A filtering process is applied to the converted reconstructed samples in the first domain.
7. The method according to claim 1, wherein Residual samples of a chroma block corresponding to the current video block are derived by applying a scaling process based on reconstructed samples in the second domain of a luma component of a picture to which the current video block belongs.
8. The method according to claim 1, wherein The motion information refinement process includes at least one of a decoder-side motion vector refinement process, a frame rate up-conversion process, or a bidirectional optical flow process.
9. The method according to claim 1, wherein Applying the motion information refinement process to the unrefined motion information to obtain the motion offset information comprises: A plurality of motion vectors derived based on the unrefined motion information are checked, and one motion vector having a lowest cost is selected to derive the motion offset information.
10. The method according to claim 1, wherein The converting includes encoding the current video block into the bitstream.
11. The method according to claim 1, wherein The converting includes decoding the current video block from the bitstream.
12. An apparatus for processing video data, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: For conversion between a current video block of a video and a bitstream of the video, determining a first prediction block in a first domain from at least one reference picture using unrefined motion information, wherein the current video block is a luma block; applying a motion information refinement process to the unrefined motion information based on at least the first prediction block to obtain motion offset information; generating a second prediction block from samples in the at least one reference picture using the motion offset information; converting the second prediction block from the first domain to a second domain where samples of the second prediction block are mapped to specific values based on a forward mapping process; as well as generating, in the second domain, reconstructed samples of the current video block based on the second prediction block, wherein a piecewise linear model is used to map the samples of the second prediction block to the specific value during the forward mapping process, and The scaling coefficient of the piecewise linear model is determined based on a first variable and a second variable, wherein the first variable is determined based on a syntax element included in an adaptive parameter set, and the second variable is determined based on a bit depth.
13. The device according to claim 12, wherein In the motion information refinement process, the motion offset information is generated based on a bilinear interpolation process and cost calculation, and the motion offset information is also used for temporal motion vector prediction of subsequent blocks.
14. The device according to claim 12, wherein In the motion information refinement process, the motion offset information is generated based on gradient calculations in different directions, wherein the gradient calculations are based on sample points in the first domain.
15. A non-transitory computer-readable storage medium storing instructions that cause a processor to: For conversion between a current video block of a video and a bitstream of the video, determining a first prediction block in a first domain from at least one reference picture using unrefined motion information, wherein the current video block is a luma block; applying a motion information refinement process to the unrefined motion information based on at least the first prediction block to obtain motion offset information; generating a second prediction block from samples in the at least one reference picture using the motion offset information; converting the second prediction block from the first domain to a second domain where samples of the second prediction block are mapped to specific values based on a forward mapping process; as well as generating, in the second domain, reconstructed samples of the current video block based on the second prediction block, wherein a piecewise linear model is used to map the samples of the second prediction block to the specific value during the forward mapping process, and The scaling coefficient of the piecewise linear model is determined based on a first variable and a second variable, wherein the first variable is determined based on a syntax element included in an adaptive parameter set, and the second variable is determined based on a bit depth.
16. The non-transitory computer-readable storage medium of claim 15, wherein: In the motion information refinement process, the motion offset information is generated based on a bilinear interpolation process and cost calculation, and the motion offset information is also used for temporal motion vector prediction of subsequent blocks.
17. A method for storing a bitstream of a video, comprising: determining, for a current video block of a video, a first prediction block in a first domain from at least one reference picture using unrefined motion information, wherein the current video block is a luma block; applying a motion information refinement process to the unrefined motion information based on at least the first prediction block to obtain motion offset information; generating a second prediction block from samples in the at least one reference picture using the motion offset information; converting the second prediction block from the first domain to a second domain where samples of the second prediction block are mapped to specific values based on a forward mapping process; generating the bitstream based on the second prediction block in the second domain; as well as storing the bitstream in a non-transitory computer-readable recording medium, The reconstructed samples of the current video block are generated in the second domain based on the second prediction block, wherein a piecewise linear model is used to map the samples of the second prediction block to the specific value during the forward mapping process, and The scaling coefficient of the piecewise linear model is determined based on a first variable and a second variable, wherein the first variable is determined based on a syntax element included in an adaptive parameter set, and the second variable is determined based on a bit depth.
18. The method according to claim 17, wherein In the motion information refinement process, the motion offset information is generated based on a bilinear interpolation process and cost calculation, and the motion offset information is also used for temporal motion vector prediction of subsequent blocks.
Citation Information
Patent Citations
Motion vector prediction
US20180359483A1
Integrated image reshaping and video coding
WO2019006300A1