Smoothing Model in Video Processing

By introducing quantitative step signaling notification and block-based loop shaping in video encoding and decoding, the challenges of improving compression ratio and reducing complexity in the prior art are solved, and efficient video encoding and decoding performance and parallel implementation capabilities are achieved.

CN113545044BActive Publication Date: 2025-06-13DOUYIN VISION CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080018617.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-08
Filing Date
2020-03-09
Publication Date
2025-06-13
Estimated Expiration
2040-03-09

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have challenges in improving compression ratios and reducing complexity, especially in parallel implementations and low-complexity video encoding and decoding solutions.

Method used

A method is proposed to interact with other tools in video codecs, such as HEVC or pending multifunctional video codec standards, by using quantitative step signaling notification and block-based loop shaping in video codecs. The method includes determining shared plastic shaping model information for a plurality of video units of the video area and using this information in a codec representation and conversion between video.

Benefits of technology

Through this method, the performance of video encoding and decoding can be improved, better compression ratio can be provided, and video encoding and decoding solutions can be supported with low complexity or parallel implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113545044B_ABST
    Figure CN113545044B_ABST
Patent Text Reader

Abstract

A video processing method is provided, including: determining shaping model information shared by a plurality of video units for conversion between the plurality of video units in a video region of a video and coded / decoded representations of the plurality of video units; and performing conversion between the coded / decoded representation of the video and the video, wherein the shaping model information provides information for constructing video samples in a first domain and a second domain and / or scaling chrominance residuals of chrominance video units.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] In accordance with the provisions of the applicable Patent Law and / or the Paris Convention, this application timely claims the priority and benefits of International Patent Application No. PCT / CN2019 / 077429, filed on March 8, 2019. By law, the entire disclosure of the above - mentioned application is incorporated herein by reference as part of the disclosure of this application. Applicable to video coding and decoding. Technical field

[0003] This patent document relates to video coding and decoding technologies, devices, and systems. Background art

[0004] Currently, efforts are being made to improve the performance of current video codec technologies to provide a better compression ratio, or to provide video coding and decoding schemes that allow for lower complexity or parallel implementation. Industry experts have recently proposed several new video coding and decoding tools, which are currently being tested to determine their effectiveness. Summary of the invention

[0005] Devices, systems, and methods related to digital video coding and decoding, particularly related to quantization step signaling notification and the interaction of block - based loop filtering with other tools in video coding and decoding. It can be applied to existing video coding standards (such as HEVC), or pending standards (Versatile Video Coding). It may also be applicable to future video coding standards or video codecs.

[0006] In a representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: determining shaping model information shared by a plurality of video units for a conversion between the plurality of video units of a video region of a video and the coded - decoded representation of the plurality of video units; and performing a conversion between the coded - decoded representation of the video and the video, wherein the shaping model information provides information for constructing video samples in a first domain and a second domain, and / or scaling the chrominance residuals of chrominance video units.

[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: determining values of variables in shaping model information as a function of the bit depth of the video for a conversion between the coded - decoded representation of a video including one or more video regions and the video; and performing the conversion based on the determination, wherein the shaping information is applicable to in - loop filtering (ILR) of some of the one or more video regions, and wherein the shaping information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling the chrominance residuals of chrominance video units.

[0008] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between a coded representation of a video that includes one or more video regions and the video, wherein the coded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations of video units in a first domain and a second domain, and / or for scaling chrominance residuals of chrominance video units, and wherein the shaping model information has been initialized based on an initialization rule.

[0009] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: determining to enable or disable in-loop reshaping (ILR) for a conversion between a coded representation of a video that includes one or more video regions and the video; and performing the conversion based on the determination, and wherein the coded representation includes the shaping model information for the ILR applicable to some of the one or more video regions, and wherein the shaping model information provides information for reconstructing a video region based on a first domain and a second domain, and / or for scaling chrominance residuals of chrominance video units, and wherein in the case where the shaping model information has not been initialized, the determination determines to disable the ILR.

[0010] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between a coded representation of a video that includes one or more video regions and the video, wherein the coded representation includes the shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on a first domain and a second domain, and / or for scaling chrominance residuals of chrominance video units, and wherein the shaping model information is included in the coded representation only if the video region is coded using a specific coding type.

[0011] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: determining, based on a rule, whether shaping information from a second video region can be used for a conversion between a first video region of a video and the coded representation of the first video region; and performing the conversion according to the determination.

[0012] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between a video region of a video and an encoded / decoded representation of the video region such that the current video region is encoded / decoded using intra coding, wherein the encoded / decoded representation conforms to formatting rules that specify that the shaping model information in the encoded / decoded representation is conditionally based on the value of a flag in the encoded / decoded representation at the video region level.

[0013] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between an encoded / decoded representation of a video that includes one or more video regions and the video, wherein the encoded / decoded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling chrominance residuals of chrominance video units, and wherein the shaping model information includes a parameter set that includes a syntax element that specifies a difference between a maximum binary number index allowed and a maximum binary number index to be used in the reconstruction, and wherein the parameter is within a range.

[0014] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between an encoded / decoded representation of a video that includes one or more video regions and the video, wherein the encoded / decoded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling chrominance residuals of chrominance video units, wherein the shaping model information includes a parameter set that includes a maximum binary number index to be used in the reconstruction, and wherein the maximum binary number index is derived as a first value that is equal to a sum of a minimum binary number index to be used in the reconstruction and a syntax element that is an unsigned integer and is signaled after the minimum binary number index.

[0015] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between a coded representation of a video that includes one or more video regions and the video, wherein the coded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling chrominance residuals of chrominance video units, and wherein the shaping model information includes a parameter set that includes a first syntax element that derives the number of bits used to represent a second syntax element, the second syntax element specifying an absolute delta codeword value from a corresponding binary number, and wherein the value of the first syntax element is less than a threshold.

[0016] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between a coded representation of a video that includes one or more video regions and the video, wherein the coded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling chrominance residuals of chrominance video units, and wherein the shaping model information includes a parameter set that includes an i-th parameter that represents the slope of an i-th binary number used in the ILR and has a value based on the (i - 1)-th parameter, where i is a positive integer.

[0017] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between a coded representation of a video that includes one or more video regions and the video, wherein the coded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling chrominance residuals of chrominance video units, and wherein the shaping model information for the ILR includes a parameter set that includes reshape_model_bin_delta_sign_CW[i], does not signal the reshape_model_bin_delta_sign_CW[i] and RspDeltaCW[i] = reshape_model_bin_delta_abs_CW[i] is always positive.

[0018] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between a coded representation of a video that includes one or more video regions and the video, wherein the coded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling a chrominance residual of a chrominance video unit, and wherein the shaping model information includes a parameter set that includes a parameter invAvgLuma that uses a luminance value for the scaling depending on a color format of the video region.

[0019] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between a current video block of a video and a coded representation of the video, wherein the conversion includes a picture inverse mapping process that transforms reconstructed picture luminance samples to modified reconstructed picture luminance samples, wherein the picture inverse mapping process includes shearing, and wherein upper and lower boundaries are set to be separated from each other.

[0020] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between a coded representation of a video that includes one or more video regions and the video, wherein the coded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling a chrominance residual of a chrominance video unit, and wherein the shaping model information includes a parameter set that includes a pivot amount constrained such that Pivot[i] <= T.

[0021] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes: performing a conversion between a coded representation of a video that includes one or more video regions and the video, wherein the coded representation includes information for in-loop reshaping (ILR) and provides parameters for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling a chrominance residual of a chrominance video unit, and wherein a chrominance quantization parameter (QP) has an offset for deriving its value for each block or transform unit.

[0022] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: performing a conversion between a codec representation of a video including one or more video regions and the video, wherein the codec representation includes information applicable to in-loop filtering (ILR), and providing parameters for reconstructing video units of the video regions based on representations in a first domain and a second domain, and / or scaling the chrominance residuals of chrominance video units, and wherein a luminance quantization parameter (QP) has an offset for deriving its value for each block or transform unit.

[0023] One or more of the above-disclosed methods can be an encoder-side implementation or a decoder-side implementation.

[0024] Furthermore, in a representative aspect, a device in a video system is disclosed, which includes a processor and a non-transitory memory having instructions thereon. When the processor executes the instructions, the instructions cause the processor to implement any one or more of the disclosed methods.

[0025] Furthermore, a computer program product stored on a non-transitory computer-readable medium is also disclosed, the computer program product including program code for implementing any one or more of the disclosed methods.

[0026] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the description, and the claims. Description of the Drawings

[0027] Figure 1 An example of constructing a Merge candidate list is shown.

[0028] Figure 2 An example of the position of a spatial candidate is shown.

[0029] Figure 3 An example of a candidate pair subjected to redundancy check of a spatial Merge candidate is shown.

[0030] Figure 4A and 4B An example of the position of a second prediction unit (PU) based on the size and shape of the current block is shown.

[0031] Figure 5 An example of motion vector scaling of a temporal Merge candidate is shown.

[0032] Figure 6 An example of the candidate position of a temporal Merge candidate is shown.

[0033] Figure 7 An example of generating a combined bi-predictive Merge candidate is shown.

[0034] Figure 8 Shows an example of constructing motion vector prediction candidates.

[0035] Figure 9 Shows an example of motion vector scaling of spatial domain motion vector candidates.

[0036] Figure 10 Shows an example of optional temporal motion vector prediction (ATMVP).

[0037] Figure 11 Shows an example of spatio-temporal motion vector prediction.

[0038] Figure 12 Shows an example of neighboring samples for deriving local illumination compensation parameters.

[0039] Figure 13A and 13B Shows diagrams related to the 4-parameter affine model and the 6-parameter affine model respectively.

[0040] Figure 14 Shows an example of the affine motion vector field for each sub-block.

[0041] Figure 15A and 15B Shows examples of the 4-parameter affine model and the 6-parameter affine model respectively.

[0042] Figure 16 Shows an example of motion vector prediction for the affine inter-frame mode that inherits affine candidates.

[0043] Figure 17 Shows an example of motion vector prediction for the affine inter-frame mode that constructs affine candidates.

[0044] Figure 18A and 18B Shows a diagram related to the affine Merge mode.

[0045] Figure 19 Shows an example of candidate positions for the affine Merge mode

[0046] Figure 20 Shows an example of the final vector representation (UMVE) search process.

[0047] Figure 21 Shows an example of UMVE search points.

[0048] Figure 22 Shows an example of decoder-side motion video refinement (DMVR).

[0049] Figure 23 Shows a block diagram flow chart of decoding using the shaping step.

[0050] Figure 24 Shows an example of a sample in a bilateral filter.

[0051] Figure 25 Shows an example of a window sample for weight calculation.

[0052] Figure 26 Shows an example of a scan pattern.

[0053] Figure 27A and 27B is a block diagram of an example of a hardware platform for implementing the visual media processing described herein.

[0054] Figures 28A to 28E Shows a flowchart of an example method for video processing based on some implementations of the disclosed technology. Detailed Description

[0055] 1. Video Coding and Decoding in HEVC / H.265

[0056] Video coding and decoding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding and decoding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding and decoding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the Versatile Video Coding (VVC) standard with the goal of reducing the bitrate by 50% compared to HEVC. The latest version of the VVC draft (i.e., Versatile Video Coding (Draft 2)) can be found at:

[0057] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 11_Ljubljana / wg11 / JVET-K1001-v7.zip

[0058] The latest reference software for VVC (named VTM) can be found at:

[0059] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-2.1

[0060] 2.1 Inter - frame prediction in HEVC / H.265

[0061] Each inter - frame predicted PU has motion parameters for one or two reference picture lists. The motion parameters include motion vectors and reference picture indices. The use of one of the two reference picture lists can also be signaled using inter_pred_idc. The motion vector can be explicitly encoded and decoded as an increment relative to the predictor.

[0062] When a CU is coded using the skip mode, a PU is associated with the CU and there are no significant residual coefficients, no coded motion vector increments or reference picture indices. A Merge mode is specified by which the motion parameters of the current PU can be obtained from neighboring PUs (including spatial and temporal candidates). The Merge mode can be applied to any inter - frame predicted PU, not just the skip mode. Another option for the Merge mode is the explicit transmission of motion parameters, where the motion vector (more precisely, the motion vector difference (MVD) compared to the motion vector predictor), the corresponding reference picture indices for each reference picture list, and the use of the reference picture list are signaled explicitly for each PU. In the present disclosure, such a mode is named Advanced Motion Vector Prediction (AMVP).

[0063] When the signaling indicates that one of the two reference picture lists is to be used, a PU is generated from a single sample block. This is called "unidirectional prediction". Unidirectional prediction is available for both P - slices and B - slices.

[0064] When the signaling indicates that both reference picture lists are to be used, a PU is generated from two sample blocks. This is called "bidirectional prediction". Bidirectional prediction is only available for B - slices.

[0065] Details of the inter - frame prediction modes specified in HEVC are provided below. The description will start with the Merge mode.

[0066] 2.1.1 Reference picture lists

[0067] In HEVC, the term inter - prediction is used to denote prediction derived from data elements (e.g., sample values or motion vectors) of reference pictures other than the current decoded picture. As in H.264 / AVC, a picture can be predicted from multiple reference pictures. The reference pictures used for inter - prediction are organized in one or more reference picture lists. A reference index identifies which reference pictures in the list are applied to create the prediction signal.

[0068] A single reference picture list (list 0) is used for P - slices, and two reference picture lists (list 0 and list 1) are used for B - slices. It should be noted that the reference pictures included in list 0 / 1 can come from past and future pictures in terms of the capture / display order.

[0069] 2.1.2 Merge Mode

[0070] 2.1.2.1 Derivation of Merge Mode Candidates

[0071] When predicting a PU using the Merge mode, an index pointing to an entry in the Merge candidate list is parsed from the bitstream, and the motion information is retrieved using this index. The construction of this list is specified in the HEVC standard and can be summarized in the following step - by - step order:

[0072] Step 1: Initial candidate derivation

[0073] Step 1.1: Spatial candidate derivation

[0074] Step 1.2: Spatial candidate redundancy check

[0075] Step 1.3: Temporal candidate derivation

[0076] Step 2: Additional candidate insertion

[0077] Step 2.1: Creation of bi - directional prediction candidates

[0078] Step 2.2: Insertion of zero - motion candidates

[0079] In Figure 1These steps are also schematically shown. For spatial domain Merge candidate derivation, at most four Merge candidates are selected from candidates located at five different positions. For temporal domain Merge candidate derivation, at most one Merge candidate is selected from two candidates. Since the number of candidates for each PU is assumed to be constant at the decoder, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number of Merge candidates (maxNumMergeCand) signaled in the slice header. Since the number of candidates is constant, the index of the best Merge candidate is encoded using truncated unary binary (TU). If the size of the CU is equal to 8, all PUs of the current CU share a Merge candidate list, which is the same as the Merge candidate list of 2N×2N prediction units.

[0080] In the following, the operations associated with the foregoing steps are described in detail.

[0081] 2.1.2.2 Spatial Domain Candidate Derivation

[0082] In the derivation of spatial domain Merge candidates, at most four Merge candidates are selected from candidates located at Figure 2 the positions shown. The derivation order is A 1 , B 1 , B 0 , A 0 and B 2 . Only when any PU at positions A 1 , B 1 , B 0 , A 0 is unavailable (e.g., because it belongs to another slice or tile) or is intra-coded, position B 2 is considered. After adding candidates at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving coding and decoding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only pairs linked by the arrows in Figure 3 are considered, and a candidate is added to the list only when the corresponding candidates used for the redundancy check do not have the same motion information. Another source of copied motion information is the "second PU" associated with a partition different from 2Nx2N. For example, Figure 4A and 4BThe second PU in the cases of N×2N and 2N×N are described separately. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate may result in two prediction units with the same motion information, which is redundant for having only one PU in the codec unit. Similarly, when the current PU is partitioned into 2N×N, position B is not considered 1 .

[0083] 2.1.2.3 Temporal candidate derivation

[0084] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal Merge candidate, a scaled motion vector is derived based on the collocated PU belonging to the picture with the smallest POC difference from the current picture within the given reference picture list. The reference picture list used to derive the collocated PU is signaled explicitly in the slice header. Obtain the scaled motion vector of the temporal Merge candidate (as shown by the dashed line in Figure 5 ), which is scaled from the motion vector of the collocated PU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal Merge candidate is set to zero. The actual implementation of the scaling process is described in the HEVC specification. For B slices, two motion vectors (one for reference picture list 0 and the other for reference picture list 1) are obtained and combined to form a bi-predictive Merge candidate

[0085] In the collocated PU (Y) belonging to the reference frame, select the position of the temporal candidate between candidate C 0 and C 1 , as shown in Figure 6 . If the PU at position C 0 is unavailable, intra-coded, or outside the current coding tree unit (CTU, also known as LCU, the largest coding unit) row, then position C 1 is used. Otherwise, position C 0 is used for the derivation of the temporal Merge candidate

[0086] 2.1.2.4 Additional candidate insertion

[0087] In addition to the spatial and temporal Merge candidates, there are two additional types of Merge candidates: combined bi-predictive Merge candidates and zero Merge candidates. The combined bi-predictive Merge candidates are generated using the spatial and temporal Merge candidates. The combined bi-predictive Merge candidates are only used for B slices. The combined bi-predictive candidates are generated by combining the motion parameters of the first reference picture list of the initial candidate with the motion parameters of the second reference picture list of another candidate. If these two tuples provide different motion hypotheses, they form a new bi-predictive candidate. As an example, Figure 7 illustrates the situation where two candidates with MVL0 and refIdxL0 or MVL1 and refIdxL1 in the original list (on the left) are used to create a combined bi-predictive Merge candidate added to the final list (on the right). Many rules regarding the combinations considered to generate these additional Merge candidates are defined.

[0088] Zero motion candidates are inserted to fill the remaining entries in the Merge candidate list to reach the capacity of MaxNumMergeCand. These candidates have zero spatial displacement and a reference picture index that starts from zero and increases each time a new zero motion candidate is added to the list. Finally, no redundancy checks are performed on these candidates.

[0089] 2.1.3 AMVP

[0090] AMVP exploits the spatio-temporal correlation of motion vectors with neighboring PUs, which is used for the explicit transmission of motion parameters. For each reference picture list, a list of motion vector candidates is first constructed by checking the availability of the spatio-temporal neighboring PU positions in the upper left, removing redundant candidates, and adding zero vectors to make the candidate list length constant. Then, the encoder can select the best predictor from the candidate list and send the corresponding index indicating the selected candidate. Similar to the Merge index signaling, the index of the best motion vector candidate is encoded using truncated unary. In this case, the maximum value to be encoded is 2 (e.g., see Figure 8 ). In the following sections, details of the derivation process of motion vector prediction candidates are provided.

[0091] 2.1.3.1 Derivation of AMVP Candidates

[0092] Figure 8 Summarizes the derivation process of motion vector prediction candidates.

[0093] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For the derivation of spatial motion vector candidates, based on the location at Figure 2The motion vectors of each PU at the five different positions shown ultimately derive two motion vector candidates.

[0094] For the derivation of time-domain motion vector candidates, one motion vector candidate is selected from two candidates, which are derived based on two different collocated positions. After making the first candidate list, the duplicate motion vector candidates in the list are removed. If the number of potential candidates is greater than two, the motion vector candidates with a reference picture index greater than 1 in the associated reference picture list are removed from the list. If the number of spatio-temporal motion vector candidates is less than two, additional zero motion vector candidates are added to the list.

[0095] 2.1.3.2 Spatial Motion Vector Candidates

[0096] When deriving spatial motion vector candidates, at most two candidates are considered among the five potential candidates, which are derived from the Figure 2 PU at the positions shown, and these positions are the same as the positions of motion Merge. The derivation order for the left side of the current PU is defined as A 0 、A 1 、and scaled A 0 、scaled A 1 . The derivation order for the upper side of the current PU is defined as B 0 、B 1 、B 2 、scaled B 0 、scaled B 1 、scaled B 2 . Therefore, there are four cases on each side that can be used as motion vector candidates, where two cases do not require the use of spatial scaling, and two cases use spatial scaling. The four different cases are summarized as follows:

[0097] --No spatial scaling

[0098] (1) The same reference picture list and the same reference picture index (the same POC)

[0099] (2) Different reference picture lists, but the same reference picture (the same POC)

[0100] --Spatial scaling

[0101] (3) The same reference picture list, but different reference pictures (different POCs)

[0102] (4) Different reference picture lists and different reference pictures (different POCs)

[0103] First, check the case without spatial domain scaling, and then check the spatial domain scaling. When the POC is different between the reference picture of the neighboring PU and the reference picture of the current PU, spatial domain scaling is considered regardless of the reference picture list. If all PUs of the left candidate are unavailable or are intra-coded, scaling of the above-mentioned motion vector is allowed to assist in the parallel derivation of the left and upper MV candidates. Otherwise, spatial domain scaling of the above-mentioned motion vector is not allowed.

[0104] In the spatial domain scaling process, the motion vectors of neighboring PUs are scaled in a similar way to the temporal domain scaling, as Figure 9 shown. The main difference is that the reference picture list and index of the current PU are given as inputs; the actual scaling process is the same as the temporal domain scaling process.

[0105] 2.1.3.3 Temporal Motion Vector Candidates

[0106] Except for the derivation of the reference picture index, all derivation processes of the temporal Merge candidates are the same as those of the spatial domain motion vector candidates (see Figure 6 ). The reference picture index is signaled to the decoder.

[0107] 2.2 Sub-CU Based Motion Vector Prediction Method in JEM

[0108] In JEM with QTBT, each CU can have at most one set of motion parameters for each prediction direction. By dividing a large CU into sub-CUs and deriving the motion information of all sub-CUs of the large CU, two sub-CU level motion vector prediction methods are considered in the encoder. The Adaptive Temporal Motion Vector Prediction (ATMVP) method allows each CU to extract multiple sets of motion information from multiple blocks smaller than the current CU in the collocated reference pictures. In the Spatio-Temporal Motion Vector Prediction (STMVP) method, the motion vectors of the sub-CUs are derived recursively by using the temporal motion vector predictor and the spatial neighboring motion vectors.

[0109] To preserve a more accurate motion field for sub-CU motion prediction, motion compensation of the reference frames is currently disabled.

[0110] 2.2.1 Adaptive Temporal Motion Vector Prediction

[0111] Figure 10 An example of the Adaptive Temporal Motion Vector Prediction (ATMVP) is shown. In the Adaptive Temporal Motion Vector Prediction (ATMVP) method, the Temporal Motion Vector Prediction (TMVP) of the motion vector is modified by extracting multiple sets of motion information (including motion vectors and reference indices) from blocks smaller than the current CU. The sub-CU is a square N×N block (the default N is set to 4).

[0112] ATMVP predicts the motion vectors of sub-CUs within a CU in two steps. The first step is to identify the corresponding block in the reference picture using the so-called temporal vector. The reference picture is called the motion source picture. The second step is to divide the current CU into sub-CUs and obtain the motion vectors and the reference index for each sub-CU from the corresponding blocks.

[0113] In the first step, the reference picture and the corresponding block are determined by the motion information of the spatial neighboring blocks of the current CU. To avoid repeated scanning of neighboring blocks, the first Merge candidate in the Merge candidate list of the current CU is used. The first available motion vector and its associated reference index are set as the temporal vector and the index of the motion source picture. In this way, in ATMVP, compared with TMVP, the corresponding block can be identified more accurately, where the corresponding block (sometimes called the collocated block) is always located at the lower right or the center position relative to the current CU.

[0114] In the second step, the corresponding block of the sub-CU is identified by adding the temporal vector to the coordinates of the current CU using the temporal vector in the motion source picture. For each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the central sample) is used to derive the motion information of the sub-CU. After the motion information of the corresponding N×N block is identified, it is converted into the motion vector and the reference index of the current sub-CU, the same as the TMVP method of HEVC, where motion scaling and other processes are applied. For example, the decoder checks whether the low-latency condition is satisfied (e.g., the POC of all reference pictures of the current picture is less than the POC of the current picture), and may use the motion vector MV x (the motion vector corresponding to the reference picture list X) to predict the motion vector MV for each sub-CU y (X is equal to 0 or 1 and Y is equal to 1 - X).

[0115] 2.2.2 Space-Time Motion Vector Prediction (STMVP)

[0116] In this method, the motion vectors of sub-CUs are derived recursively in raster scan order. Figure 11 This concept is illustrated. Consider an 8×8 CU that contains four 4×4 sub-CUs A, B, C, and D. The neighboring 4×4 blocks in the current frame are labeled a, b, c, and d.

[0117] The motion derivation of sub-CU A starts with identifying its two spatial neighbors. The first neighbor is the N×N block (block c) above sub-CU A. If this block c is not available or is intra-coded / decoded, then other N×N blocks above sub-CU A are checked (from left to right, starting from block c). The second neighbor is the block (block b) to the left of sub-CU A. If block b is not available or is intra-coded / decoded, then other blocks to the left of sub-CU A are checked (from top to bottom, starting from block b). The motion information obtained from each neighboring block for each list is scaled to the first reference frame of the given list. Next, the temporal motion vector predictor (TMVP) of sub-block A is derived following the same process as specified in HEVC for TMVP derivation. The motion information of the collocated block at position D is extracted and scaled accordingly. Finally, after retrieving and scaling the motion information, all available motion vectors (up to 3) for each reference list are averaged separately. The averaged motion vector is designated as the motion vector of the current sub-CU.

[0118] 2.2.3 Sub-CU Motion Prediction Mode Signaling

[0119] The sub-CU mode is enabled as an additional Merge candidate and no additional syntax elements are required to signal this mode. Two additional Merge candidates are added to the Merge candidate list of each CU to represent the ATMVP mode and the STMVP mode. If the sequence parameter set indicates that ATMVP and STMVP are enabled, up to seven Merge candidates are used. The encoding logic for the additional Merge candidates is the same as that for the Merge candidates in HM, which means that for each CU in a P-slice or B-slice, more than two additional RD checks are required for the two additional Merge candidates.

[0120] In JEM, all binary digits of the Merge index are context-coded by CABAC. However, in HEVC, only the first binary digit is context-coded and the remaining binary digits are context-bypass-coded.

[0121] 2.3 Local Luminance Compensation in JEM

[0122] Local luminance compensation (LIC) is based on a linear model for luminance variation, using a scaling factor a and an offset b. And it is adaptively enabled or disabled for each inter-mode coded coding unit (CU).

[0123] When applying LIC to a CU, the parameters a and b are derived using the least squares error method by using the neighboring samples of the current CU and their corresponding reference samples. More specifically, as Figure 12 shown, the subsampled (2:1 subsampling) neighboring samples of the CU in the reference picture and the corresponding samples (identified by the motion information of the current CU or sub-CU) are used.

[0124] 2.3.1 Derivation of Prediction Blocks

[0125] IC parameters are derived and applied to each prediction direction respectively. For each prediction direction, a first prediction block is generated using decoded motion information, and then a temporary prediction block is obtained by applying the LIC model. Then, the final prediction block is derived using these two temporary prediction blocks.

[0126] When the CU is coded / decoded in Merge mode, the LIC flag is copied from neighboring blocks in a manner similar to the copying of motion information in Merge mode; otherwise, the LIC flag is signaled for the CU to indicate whether to apply LIC.

[0127] When LIC is enabled for a picture, additional CU-level RD checks are required to determine whether to apply LIC to the CU. When LIC is enabled for the CU, Mean Removed Sum of Absolute Differences (MR-SAD) and Mean Removed Sum of Absolute Hadamard Transform Differences (MR-SATD) are used for integer-pixel motion search and fractional-pixel motion search respectively, instead of SAD and SATD.

[0128] To reduce the coding complexity, the following coding scheme is adopted in JEM.

[0129] · When there is no significant luminance change between the current picture and its reference pictures, LIC is disabled for the entire picture. To identify this situation, the histogram of the current picture and each reference picture of the current picture are calculated at the encoder. If the histogram difference between the current picture and each reference picture of the current picture is less than a given threshold, LIC is disabled for the current picture; otherwise, LIC is enabled for the current picture.

[0130] 2.4 Inter-Frame Prediction Methods in VVC

[0131] There are several new coding / decoding tools for inter-frame prediction improvement, such as Adaptive Motion Vector Difference Resolution (AMVR) for signaling MVD, affine prediction mode, Triangular Prediction Mode (TPM), ATMVP, Generalized Bi-Directional Prediction (GBI), and Bidirectional Optical Flow (BIO).

[0132] 2.4.1 Coding / Decoding Block Structure

[0133] In VVC, a Quadtree / Binary Tree / Ternary Tree (QT / BT / TT) structure is adopted to divide the image into square or rectangular blocks.

[0134] In addition to QT / BT / TT, for I-frames, VVC also adopts a split tree (also known as a dual coding / decoding tree). Using the split tree, the coding / decoding block structure is signaled separately for the luminance and chrominance components.

[0135] 2.4.2 Adaptive Motion Vector Difference Resolution

[0136] In HEVC, when use_integer_mv_flag equals 0 in the slice header, the Motion Vector Difference (MVD) (between the motion vector of the PU and the predicted motion vector) is signaled in units of quarter luminance samples. In VVC, Local Adaptive Motion Vector Resolution (AMVR) is introduced. In VVC, the MVD can be encoded and decoded in units of quarter luminance samples, integer luminance samples, or four luminance samples (i.e., 1 / 4 pixel, 1 pixel, 4 pixels). The MVD resolution control is at the Coding Unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU having at least one non-zero MVD component.

[0137] For a CU having at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luminance sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter luminance sample MV precision is not used, another flag is signaled to indicate whether integer luminance sample MV precision or four luminance sample MV precision is used.

[0138] When the first MVD resolution flag of a CU is zero or the CU is not decoded (meaning all MVDs in the CU are zero), the CU uses quarter luminance sample MV resolution. When the CU uses integer luminance sample MV precision or four luminance sample MV precision, the MVP in the AMVP candidate list of the CU is rounded to the corresponding precision.

[0139] 2.4.3 Affine Motion Compensation Prediction

[0140] In HEVC, only the translational motion model is applied to Motion Compensation Prediction (MCP). However, there are various motions in the real world, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, a simplified affine transform motion compensation prediction using 4-parameter affine model and 6-parameter affine model is adopted. As Figure 13A and 13B shown, for the 4-parameter affine model, the affine motion field of a block is described by two Control Point Motion Vectors (CPMVs), and for the 6-parameter affine model, the affine motion field of a block is described by three CPMVs.

[0141] The Motion Vector Field (MVF) of a block is described by the following equations respectively. Equation (1) is for the 4-parameter affine model (where the 4 parameters are defined as variables a, b, e, and f) and Equation (2) is for the 6-parameter affine model (where the 4 parameters are defined as variables a, b, c, d, e, and f):

[0142]

[0143]

[0144] where (mv h 0 , mv h 0 ) is the motion vector of the top-left control point (CP), and (mv h 1 , mv h 1 ) is the motion vector of the top-right control point, and (mv h 2 , mv h 2 ) is the motion vector of the bottom-left control point. All these three motion vectors are referred to as control point motion vectors (CPMV). (x, y) represents the coordinates of the representative point relative to the top-left sample within the current block, and (mv h (x, y), mv v (x, y)) is the motion vector derived for the sample located at (x, y). The CP motion vectors can be signaled (such as in the affine AMVP mode) or derived on-the-fly (such as in the affine Merge mode). w and h are the width and height of the current block. In fact, the division is implemented using a rounding operation by right shift. In VTM, the representative point is defined as the center position of the sub-block. For example, when the coordinates of the top-left corner of the sub-block relative to the top-left sample within the current block are (xs, ys), the coordinates of the representative point are defined as (xs + 2, ys + 2). For each sub-block (i.e., 4×4 in VTM), the representative point is used to derive the motion vector of the entire sub-block.

[0145] To further simplify the motion compensation prediction, sub-block-based affine transform prediction is applied. To derive the motion vector for each M×N (both M and N are set to 4 in the current VVC) sub-block, as Figure 14 shown, the motion vector of the center sample of each sub-block can be calculated according to equations (1) and (2) and rounded to 1 / 16 fractional precision. Then a 1 / 16-pixel motion compensation interpolation filter is applied to generate the prediction for each sub-block using the derived motion vector. The affine mode introduces a 1 / 16-pixel interpolation filter.

[0146] After MCP, the high-accuracy motion vectors of each sub-block are rounded and saved with the same accuracy as the regular motion vectors.

[0147] 2.4.3.1 Signaling of Affine Prediction

[0148] Similar to the translational motion model, affine prediction also has two modes for signaling side information. They are the AFFINE_INTER mode and the AFFINE_MERGE mode.

[0149] 2.4.3.2 AF_INTER Mode

[0150] For a CU with both width and height greater than 8, the AF_INTER mode can be applied. In the bitstream, the affine flag at the CU level is signaled to indicate whether the AF_INTER mode is used.

[0151] In this mode, for each reference picture list (list 0 or list 1), an affine AMVP candidate list is constructed with three types of affine motion predictors in the following order, where each candidate includes the estimated CPMV of the current block. The difference of the best CPMV found at the encoder side (such as Figure 17 the mv in 0 mv 1 mv 2 ) and the estimated CPMV are signaled. In addition, the index of the affine AMVP candidate from which the estimated CPMV is derived is further signaled.

[0152] 1) Inherited Affine Motion Predictor

[0153] The checking order is similar to the checking order of the intra-MVP in the HEVC AMVP list construction. First, the left-inherited affine motion predictor is derived from the first block in {A1, A0}, which is affine coded / decoded and has the same reference picture as the current block. Second, the above-mentioned inherited affine motion predictor is derived from the first block in {B1, B0, B2}, which is affine coded / decoded and has the same reference picture as the current block. Figure 16 Five blocks A1, A0, B1, B0, B2 are depicted.

[0154] Once it is found that the neighboring block is affine coded / decoded, the CPMV of the coding unit covering the neighboring block is used to derive the predictor of the CPMV of the current block. For example, if A1 is coded / decoded in a non-affine mode and A0 is coded / decoded in a 4-parameter affine mode, the left-inherited affine MV predictor will be derived from A0. In this case, the CPMV of the CU covering A0 (such as Figure 18B represented as the upper-left CPMV and represented as the upper-right CPMV) is used to derive the estimated CPMV of the current block (the upper-left (coordinates (x0, y0)), upper-right (coordinates (x1, y1)) and lower-right (coordinates (x2, y2)) positions of the current block represented as ).

[0155] 2) Constructed affine motion predictor

[0156] As Figure 17 shown, the constructed affine motion predictor consists of control point motion vectors (CPMVs) derived from neighboring inter - coded blocks with the same reference picture. If the current affine motion model is a 4 - parameter affine, the number of CPMVs is 2; otherwise, if the current affine motion model is a 6 - parameter affine, the number of CPMVs is 3. The upper - left CPMV is derived from the MV at the first block in group {A, B, C}, which is inter - coded and has the same reference picture as the current block. The upper - right CPMV is derived from the MV at the first block in group {D, E}, which is inter - coded and has the same reference picture as the current block. The lower - left is derived from the MV at the first block in group {F, G}, which is inter - coded and has the same reference picture as the current block.

[0157] – If the current affine motion model is a 4 - parameter affine, the constructed affine motion predictor is inserted into the candidate list only when and are both found (that is, and are used as the CPMVs for estimating the upper - left (coordinates (x0, y0)) and upper - right (coordinates (x1, y1)) positions of the current block).

[0158] – If the current affine motion model is a 6 - parameter affine, the constructed affine motion predictor is inserted into the candidate list only when and are both found (that is, and are used as the CPMVs for estimating the upper - left (coordinates (x0, y0)), upper - right (coordinates (x1, y1)) and lower - right (coordinates (x2, y2)) positions of the current block).

[0159] When inserting the constructed affine motion predictor into the candidate list, no pruning process is applied.

[0160] 3) Conventional AMVP motion predictor

[0161] The following applies until the number of affine motion predictors reaches the maximum value.

[0162] 1) If available, derive the affine motion predictor by setting all CPMVs equal to .

[0163] 2) If available, derive the affine motion predictor by setting all CPMVs equal to to derive an affine motion predictor.

[0164] 3) If available, by setting all CPMVs equal to to derive an affine motion predictor.

[0165] 4) If available, by setting all CPMVs equal to HEVC TMVP to derive an affine motion predictor.

[0166] 5) By setting all CPMVs equal to zero MV to derive an affine motion predictor.

[0167] Note that it has been derived in the constructed affine motion predictor

[0168] In the AF_INTER mode, when using the 4 / 6-parameter affine mode, 2 / 3 control points are required, and thus 2 / 3 MVDs need to be coded / decoded for these control points, as Figure 15A shown. In JVET-K0337, the following is proposed to derive MV, i.e., to predict mvd 0 from mvd 1 and mvd 2 .

[0169]

[0170]

[0171]

[0172] where mvd i and mv 1 are the predicted motion vector, motion vector difference, and motion vector of the top-left pixel (i = 0), top-right pixel (i = 1), or bottom-left pixel (i = 2), respectively, as Figure 15B shown. It should be noted that the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the sum of the two components respectively. That is, newMV = mvA + mvB, and the two components of newMV are set to (xA + xB) and (yA + yB) respectively.

[0173] 2.4.3.3 AF_MERGE Mode

[0174] When applying a CU in the AF_MERGE mode, it obtains the first block coded / decoded in the affine mode from valid neighboring reconstructed blocks. And the selection order of candidate blocks is from left, top, top-right, bottom-left to top-left, as Figure 18Aas shown (denoted as A, B, C, D, E in sequence). For example, if the adjacent lower left block is coded in affine mode (as indicated by A0 in Figure 18B ), then the control point (CP) motion vectors mv 0 N of the upper left, upper right, and lower left of the adjacent CU / PU containing block A are extracted 1 N and mv 2 N . And based on mv 0 N , mv 1 N and mv 2 N the motion vectors mv 0 C of the upper left / upper right / lower left of the current CU / PU are calculated 1 C and mv 2 C (only for the 6-parameter affine model). It should be noted that in VTM-2.0, if the current block is affine coded, the sub-block located in the upper left corner (e.g., the 4×4 block in VTM) stores mv0, and the sub-block located in the upper right corner stores mv1. If the current block is coded with the 6-parameter affine model, the sub-block located in the lower left corner stores mv2; otherwise (with the 4-parameter affine model), LB stores mv2'. Other sub-blocks store the MVs for MC.

[0175] After deriving the CPMV mv 0 C , mv 1 C and mv 2 C of the current CU, the MVF of the current CU is generated according to the simplified affine motion model in equations (1) and (2). To identify whether the current CU is coded in AF_MERGE mode, when at least one adjacent block is coded in affine mode, the affine flag is signaled in the bitstream.

[0176] In JVET-L0142 and JVET-L0632, the following steps are used to construct the affine Merge candidates:

[0177] 1) Insert the inherited affine candidates

[0178] An inherited affine candidate means that the candidate is derived from the affine motion model of its valid neighboring affine coding / decoding blocks. The two largest inherited affine candidates are derived from the affine motion models of the neighboring blocks and inserted into the candidate list. For the left predictor, the scanning order is {A0, A1}; for the above predictor, the scanning order is {B0, B1, B2}.

[0179] 2) Insert the constructed affine candidates

[0180] If the number of candidates in the affine Merge candidate list is less than MaxNumAffineCand (e.g., 5), the constructed affine candidates are inserted into the candidate list. A constructed affine candidate means that the candidate is constructed by combining the neighboring motion information of each control point.

[0181] a) First, derive the motion information of the control points from Figure 19 the specified spatial neighbors and temporal neighbors shown. CPk (k = 1, 2, 3, 4) represents the k-th control point. A0, A1, A2, B0, B1, B2, and B3 are used to predict the spatial positions of CPk (k = 1, 2, 3); T is used to predict the temporal position of CP4.

[0182] The coordinates of CP1, CP2, CP3, and CP4 are (0, 0), (W, 0), (H, 0), and (W, H) respectively, where W and H are the width and height of the current block.

[0183] Obtain the motion information of each control point in the following order of precedence:

[0184] For CP1, check the precedence order of B2 -> B3 -> A2. If B2 is available, use B2. Otherwise, if B2 is not available, use B3. If both B2 and B3 are not available, use A2. If all three candidates are not available, the motion information of CP1 cannot be obtained.

[0185] For CP2, check the precedence order of B1 -> B0;

[0186] For CP3, check the precedence order of A1 -> A0;

[0187] For CP4, use T.

[0188] b) Second, use the combination of control points to construct the affine Merge candidates.

[0189] I. The motion information of three control points is required to construct a 6-parameter affine candidate. These three control points can be selected from one of the following four combinations ({CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}). The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4} will be transformed into a 6-parameter motion model represented by the upper-left, upper-right, and lower-left control points.

[0190] II. The motion information of two control points is required to construct a 4-parameter affine candidate. These two control points can be selected from one of the following two combinations ({CP1, CP2}, {CP1, CP3}). These two combinations will be transformed into a 4-parameter motion model represented by the upper-left and upper-right control points.

[0191] III. Insert the combinations for constructing affine candidates into the candidate list in the following order:

[0192] {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4}, {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3}.

[0193] i. For each combination, check the reference indices of each CP in list X of the control points. If they are all the same, then this combination has a valid CPMV for list X. If this combination does not have a valid CPMV for list 0 and list 1, then this combination is marked as invalid. Otherwise, it is valid, and the CPMV is put into the sub-block Merge list.

[0194] 3) Fill in zero motion vectors

[0195] If the number of candidates in the affine Merge candidate list is less than 5, then insert zero motion vectors with zero reference indices into the candidate list until the list is full.

[0196] More specifically, for the sub-block Merge candidate list, the MV of the 4-parameter Merge candidate is set to (0, 0) and the prediction direction is set to unidirectional prediction from list 0 (for P slices) and bidirectional prediction (for B slices).

[0197] 2.4.4 Merge with Motion Vector Difference (MMVD)

[0198] In JVET-L0054, the Unified Motion Vector Expression (UMVE, also known as MMVD) was proposed. UMVE and the proposed motion vector expression method are used for the skip or Merge mode.

[0199] UMVE reuses the same Merge candidates as those included in the regular Merge candidate list in VVC. Among the Merge candidates, a base candidate can be selected and further extended by the proposed motion vector representation method.

[0200] UMVE provides a new method for representing the motion vector difference (MVD), in which the starting point, motion amplitude, and motion direction are used to represent the MVD.

[0201] The proposed technique uses the Merge candidate list as it is. However, only the candidates of the default Merge type (MRG_TYPE_DEFAULT_N) are considered for the extension of UMVE.

[0202] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list as follows.

[0203] Table 1 Base candidate IDX

[0204]

[0205] If the number of base candidates is equal to 1, the base candidate IDX is not signaled.

[0206] The distance index is the motion amplitude information. The distance index indicates a predefined distance from the starting point information. The predefined distances are as follows:

[0207] Table 2 Distance IDX

[0208]

[0209]

[0210] The direction index represents the direction of the MVD relative to the starting point. The direction index can represent the following four directions.

[0211] Table 3 Direction IDX

[0212] Direction IDX 00 01 10 11 x - axis + – N / A N / A y - axis N / A N / A + –

[0213] The UMVE flag is signaled immediately after the skip flag or the Merge flag is sent. If the skip or Merge flag is true, the UMVE flag is parsed. If the UMVE flag is equal to 1, the UMVE syntax is parsed. However, if it is not 1, the affine (AFFINE) flag is parsed. If the affine flag is equal to 1, it is the affine mode, but if it is not 1, the skip / Merge index is parsed for the skip / Merge mode of VTM.

[0214] Additional line buffers due to UMVE candidates are not required. Since the skipped / Merge candidates of the software are directly used as base candidates. Using the input UMVE index, the supplement of the MV is determined before motion compensation. There is no need to reserve a long line buffer for this.

[0215] Under the current normal test conditions, the first or second Merge candidate in the Merge candidate list can be selected as the base candidate.

[0216] UMVE is also referred to as Merge with MV Difference (MMVD).

[0217] 2.4.5 Decoder-side Motion Vector Refinement (DMVR)

[0218] In bidirectional prediction operation, for the prediction of a block region, two prediction blocks formed by the motion vectors (MVs) of list 0 and the MVs of list 1 are respectively combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two motion vectors of bidirectional prediction are further refined.

[0219] In the JEM design, the motion vectors are refined through bilateral template matching processing. Bilateral template matching is applied in the decoder to perform a distortion-based search between the bilateral template and the reconstructed samples in the reference picture, so as to obtain refined MVs without transmitting additional motion information. Figure 22 An example is shown. As Figure 22 shown, from the initial MV0 from list 0 and MV1 from list 1 respectively, a bilateral template is generated as a weighted combination (i.e., average) of the two prediction blocks. The template matching operation includes calculating the cost metric between the generated template and the sample region (around the initial prediction block) in the reference picture. For each of the two reference pictures, the MV that produces the minimum template cost is regarded as the updated MV of that list to replace the original MV. In JEM, nine MV candidates are searched for each list. The nine MV candidates include the original MV and 8 surrounding MVs, and the 8 surrounding MVs have an offset of one luminance sample towards the original MV in the horizontal or vertical direction, or in both the horizontal and vertical directions. Finally, the Figure 22 two new MVs (i.e., MV0' and MV1') shown are used to generate the final bidirectional prediction result. The sum of absolute differences (SAD) is used as the cost metric. It should be noted that when calculating the cost of the prediction block generated by a surrounding MV, the rounded MV (to integer pixels) rather than the true MV is actually used to obtain the prediction block.

[0220] To further simplify the processing of DMVR, JVET-M0147 proposed several modifications to the design in JEM. More specifically, the DMVR design adopted by VTM-4.0 (to be released) has the following main features:

[0221] · Early termination with (0,0) position SAD between list 0 and list 1;

[0222] · Block size W*H >= 64 && H >= 8 of DMVR;

[0223] · For DMVR with CU size > 16*16, divide the CU into multiple 16x16 sub-blocks;

[0224] · Reference block size (W+7)*(H+7) (for luminance);

[0225] · Integer pixel search based on 25-point SAD (i.e., (+-)2 refinement search range, single stage);

[0226] · DMVR based on bilinear interpolation;

[0227] · MVD mirroring between list 0 and list 1 allows bilateral matching;

[0228] · Sub-pixel refinement based on the "parameter error surface equation";

[0229] · Luminance / chrominance MC with reference block filling (if needed);

[0230] · Refined MVs are only used for MC and TMVP.

[0231] 2.4.6 Combined Intra and Inter Prediction

[0232] In JVET-L0100, multi-hypothesis prediction is proposed, where combined intra and inter prediction is one method to generate multi-hypotheses.

[0233] When multi-hypothesis prediction is applied to improve the intra mode, the multi-hypothesis prediction combines an intra prediction and a Merge index prediction. In a Merge CU, when the flag is true, a flag is signaled for the Merge mode to select an intra mode from the intra candidate list. For the luminance component, the intra candidate list is derived from 4 intra prediction modes including DC, planar, horizontal, and vertical modes, and the size of the intra candidate list can be 3 or 4, depending on the shape of the block. When the CU width is greater than twice the CU height, the horizontal mode is not included in the intra mode list, and when the CU height is greater than twice the CU width, the vertical mode is removed from the intra mode list. A weighted average is used to combine an intra prediction mode selected by the intra mode index and a Merge index prediction selected by the Merge index. For the chrominance component, DM is always applied without additional signaling. The weights for the combined prediction are described as follows. When the DC or planar mode is selected, or the CB width or height is less than 4, equal weights are applied. For those CBs with a width and height greater than or equal to 4, when the horizontal / vertical mode is selected, a CB is first divided vertically / horizontally into four equally sized regions. Each set of weights represented as (w_intra i , w_inter i ) is applied to the corresponding region, where i ranges from 1 to 4, and (w_intra 1 , w_inter 1 ) = (6, 2), (w_intra 2 , w_inter 2 ) = (5, 3), (w_intra 3 , w_inter 3 ) = (3, 5), and (w_intra 4 , w_inter 4 ) = (2, 6). (w_intra 1 , w_inter 1 ) is used for the region closest to the reference sample, and (w_intra 4 , w_inter 4 ) is used for the region farthest from the reference sample. Then, the combined prediction can be calculated by adding the two weighted predictions and shifting right by 3 bits. In addition, the intra prediction mode of the intra hypothesis of the predictor can be saved for subsequent neighboring CU reference.

[0234] 2.5 Loop Filtering (ILR) in JVET-M0427

[0235] Loop Filtering (ILR) is also known as Luminance Mapping with Chroma Scaling (LMCS).

[0236] The basic idea of the In-Loop Reshaper (ILR) is to transform the original (in the first domain) signal (predicted / reconstructed signal) into a second domain (reshaping domain).

[0237] The in-loop luminance reshaper is implemented as a pair of Look-Up Tables (LUTs), but only one of the two LUTs needs to be signaled because the other LUT can be computed from the signaled LUT. Each LUT is a one-dimensional, 10-bit, 1024-entry mapping table (1D-LUT). One LUT is the forward LUT FwdLUT, which maps the input luminance code value Y i to the changed value Y r : Y r = FwdLUT[Y i . The other LUT is the inverse LUT InvLUT, which maps the changed code value Y r to ( denoting the reconstructed value of Y i ).

[0238] 2.5.1 PWL Model

[0239] Conceptually, the Piece-Wise Linear (PWL) is implemented as follows:

[0240] Let x1, x2 be two input pivot points, and y1, y2 be the corresponding output pivot points of a segment. The output value y for any input value x between x1 and x2 can be interpolated by the following equation:

[0241] y = ((y2 - y1) / (x2 - x1)) * (x - x1) + y1

[0242] In fixed-point implementation, the equation can be rewritten as:

[0243] y = ((m * x + 2 FP_PREC-1 ) >> FP_PREC) + c

[0244] where m is a scalar, c is an offset, and FP_PREC is a constant value used to specify the precision.

[0245] Note that in the CE-12 software, the PWL model is used to pre-compute the 1024-entry FwdLUT and InvLUT mapping tables; however, the PWL model also allows for computing equivalent mapping values on-the-fly without pre-computing the LUTs.

[0246] 2.5.2 Testing CE12-2

[0247] 2.5.2.1 Luminance Reshaping

[0248] Test 2 of loop luminance shaping (i.e., proposed CE12-2) provides a lower complexity pipeline that also eliminates the decoding latency of block intra prediction in inter slice strip reconstruction. For both inter and intra slices, intra prediction is performed in the shaping domain.

[0249] Regardless of the slice type, intra prediction is always performed in the shaping domain. With such an arrangement, intra prediction can start immediately after the previous TU reconstruction is completed. This arrangement can also provide a unified rather than slice-dependent processing for intra modes. Figure 23 A block diagram of the mode-based CE12-2 decoding process is shown.

[0250] CE12-2 also tests a 16-segment piecewise linear (PWL) model for luminance and chrominance residual scaling instead of the 32-segment PWL model of CE12-1.

[0251] Inter slice strip reconstruction using the loop luminance shaper in CE12-2 (the light green shaded blocks indicate signals in the shaping domain: luminance residuals; intra luminance prediction; and intra luminance reconstruction).

[0252] 2.5.2.2 Luminance-dependent chrominance residual scaling

[0253] Luminance-dependent chrominance residual scaling is a multiplication process implemented using fixed-point integer arithmetic. Chrominance residual scaling compensates for the interaction between the luminance and chrominance signals. Chrominance residual scaling is applied at the TU level. More specifically, the following applies:

[0254] – For intra, average reconstructed luminance.

[0255] – For inter, average predicted luminance.

[0256] The average is used to identify the index in the PWL model. This index identifies the scaling factor cScaleInv. The chrominance residuals are multiplied by this number.

[0257] It should be noted that the chrominance scaling factor is calculated from the forward-mapped predicted luminance value rather than the reconstructed luminance value.

[0258] 2.5.2.3 Signaling of ILR side information

[0259] Parameters (current) are sent in the slice group header (similar to ALF). These require 40 - 100 bits.

[0260] The following table is based on version 9 of JVET-L1001. The syntax to be added is highlighted in underlined bold italic font below.

[0261] In the sequence parameter set of 7.3.2.1, the RBSP syntax can be set as follows:

[0262]

[0263]

[0264]

[0265] In 7.3.3.1, the common slice group header syntax can be modified by inserting the following underlined bold italic text:

[0266]

[0267]

[0268] New syntax table slice group reshaper models can be added as follows:

[0269]

[0270] In the semantics of the regular sequence parameter set RBSP, the following semantics can be added:

[0271] The sps_reshaper_enabled_flag equal to 1 specifies that the reshaper is used in the coded video sequence (CVS). The sps_reformer_enabled_flag equal to 0 specifies that the reshaper is not used in the CVS.

[0272] In the slice group header syntax, the following semantics can be added:

[0273] The tile_group_reshaper_model_present_flag equal to 1 specifies that tile_group_reshaper_model() is present in the slice group header. The tile_group_reshaper_model_present_flag equal to 0 specifies that tile_group_reshaper_model() is not present in the slice group header. When the tile_group_reshaper_model_present_flag is not present, it is inferred to be equal to 0.

[0274] The tile_group_reshaper_enabled_flag equal to 1 specifies that the reshaper is enabled for the current slice group. The tile_group_reshaper_enabled_flag equal to 0 specifies that the reshaper is not enabled for the current slice group. When the tile_group_reshaper_enable_flag is not present, it is inferred to be equal to 0.

[0275] When tile_group_reshaper_chroma_residual_scale_flag equals 1, chroma residual scaling is specified to be enabled for the current tile group. When tile_group_reshaper_chroma_residual_scale_flag equals 0, chroma residual scaling is specified not to be enabled for the current tile group. When tile_group_reshaper_chroma_residual_scale_flag is absent, it is inferred to be equal to 0.

[0276] The syntax of tile_group_reshaper_model() can be added as follows:

[0277] reshape_model_min_bin_idx specifies the index of the minimum bin (or segment) number that will be used in the reshaper construction process. The value of reshape_model_min_bin_idx should be in the range of 0 to MaxBinIdx (including 0 and MaxBinIdx). The value of MaxBinIdx should be equal to 15.

[0278] reshape_model_delta_max_bin_idx specifies the maximum allowed bin (or segment) number index that will be used in the reshaper construction process. Set the value of reshape_model_max_bin_idx to be equal to MaxBinIdx – reshape_model_delta_max_bin_idx.

[0279] reshaper_model_bin_delta_abs_cw_prec_minus1 plus 1 specifies the number of bits used for the representation of the syntax reshape_model_bin_delta_abs_CW[i].

[0280] reshape_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value of the i-th bin.

[0281] reshaper_model_bin_delta_sign_CW_flag[i] specifies the sign of reshape_model_bin_delta_abs_CW[i] as follows:

[0282] – If reshape_model_bin_delta_sign_CW_flag[i] equals 0, the corresponding variable RspDeltaCW[i] is positive.

[0283] – Otherwise (if reshape_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] is negative.

[0284] When reshape_model_bin_delta_sign_CW_flag[i] does not exist, infer that it is equal to 0.

[0285] The variable RspDeltaCW[i] = (1 - 2 * reshape_model_bin_delta_sign_CW[i]) * reshape_model_bin_delta_abs_CW[i].

[0286] Derive the variable RspCW[i] according to the following steps:

[0287] – Set the variable OrgCW to be equal to (1 << BitDepth Y ) / (MaxBinIdx + 1).

[0288] – If reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx, then RspCW[i] = OrgCW + RspDeltaCW[i].

[0289] – Otherwise, RspCW[i] = 0.

[0290] If the value of BitDepth Y is equal to 10, the value of RspCW[i] should be in the range of 32 to 2 * OrgCW - 1.

[0291] The variable InputPivot[i] (where i ranges from 0 to MaxBinIdx + 1, including 0 and MaxBinIdx + 1) is derived as follows:

[0292] InputPivot[i] = i * OrgCW.

[0293] The variables ReshapePivot[i] (where i ranges from 0 to MaxBinIdx + 1, including 0 and MaxBinIdx + 1), ScaleCoef[i], and InvScaleCoeff[i] (where i ranges from 0 to MaxBinIdx, including 0 and MaxBinIdx) are derived as follows:

[0294]

[0295]

[0296] The variable ChromaScaleCoef[i] (where i ranges from 0 to MaxBinIdx, inclusive of 0 and MaxBinIdx) can be derived as follows:

[0297] ChromaResidualScaleLut

[64] = {16384, 16384, 16384, 16384, 16384, 16384, 16384, 8192, 8192, 8192, 8192, 5461, 5461, 5461, 5461, 4096, 4096, 4096, 4096, 3277, 3277, 3277, 3277, 2731, 2731, 2731, 2731, 2341, 2341, 2341, 2048, 2048, 2048, 1820, 1820, 1820, 1638, 1638, 1638, 1638, 1489, 1489, 1489, 1489, 1365, 1365, 1365, 1365, 1260, 1260, 1260, 1260, 1170, 1170, 1170, 1170, 1092, 1092, 1092, 1092, 1024, 1024, 1024, 1024};

[0298]

[0299] The following text can be added in the weighted sample prediction process that combines the Merge and intra prediction. The added part is marked in italic font with an underline.

[0300] 8.4.6.6 Weighted Sample Prediction Process Combining Merge and Intra Prediction

[0301] The inputs to this process are:

[0302] - the width cbWidth of the current coding block,

[0303] - the height cbHeight of the current coding block,

[0304] - two arrays predSamplesInter and predSamplesIntra of (cbWidth) x (cbHeight),

[0305] - the intra prediction mode predModeIntra,

[0306] - the variable cIdx that specifies the color component index.

[0307] The output of this process is an array predSamplesComb of (cbWidth) x (cbHeight) of predicted sample values.

[0308] The variable bitDepth is derived as follows:

[0309] - If cIdx is equal to 0, set bitDepth to be equal to BitDepth Y .

[0310] - Otherwise, set bitDepth to be equal to BitDepth C .

[0311] The predicted samples predSamplesComb[x][y] (where x = 0..cbWidth - 1 and y = 0..cbHeight - 1) are derived as follows:

[0312] - The weight w is derived as follows:

[0313] - If predModeIntra is INTRA_ANGULAR50, specify w in Table 4 with nPos equal to y and nSize equal to cbHeight.

[0314] - Otherwise, if predModeIntra is INTRA_ANGULAR18, specify w in Table 4 with nPos equal to x and nSize equal to cbWidth.

[0315] - Otherwise, set w to be equal to 4.

[0316] – If cIdx equals 0, then predSamplesInter is derived as follows:

[0317] - If the tile group reshaper enabled flag equals 1, then shiftY = 14

[0318] idxY = predSamplesInter[x][y] >> Log2(OrgCW)

[0319] predSamplesInter[x][y] = Clip1 Y (ReshapePivot[idxY]+(ScaleCoeff[idxY]* (predSamplesInter[x][y] - InputPivot[i dxY])+(1 << (shiftY – 1))) >> shiftY)(8 - xxx)

[0320] - Otherwise, (tile_group_reshaper_enabled_flag equals 0)

[0321] predSamplesInter[x][y] = predSamplesInter[x][y]

[0322] – The predicted sample predSamplesComb[x][y] is derived as follows:

[0323]

[0324] Table 4 Specification of w as a function of position nP and size nS

[0325]

[0326] The following text, shown in bold italic with an underline, can be added to the picture reconstruction process:

[0327] 8.5.5 Picture Reconstruction Process

[0328] The input to this process is:

[0329] - Position (xCurr, yCurr), which specifies the top-left sample point of the current block relative to the top-left sample point of the current picture component

[0330] – Variables nCurrSw and nCurrSh, which respectively specify the width and height of the current block

[0331] – Variable cIdx, which specifies the color component of the current block

[0332] – An array predSamples of (nCurrSw) x (nCurrSh), which specifies the predicted samples of the current block

[0333] – An array resSamples of (nCurrSw) x (nCurrSh), which specifies the residual samples of the current block.

[0334] Depending on the value of the color component cIdx, the following assignments are made:

[0335] – If cIdx is equal to 0, recSamples corresponds to the reconstructed picture sample array S L , and the function clipCidx1 corresponds to Clip1 Y .

[0336] – Otherwise, if cIdx is equal to 1, recSamples corresponds to the reconstructed chrominance sample array S Cb , and the function clipCidx1 corresponds to Clip1 C .

[0337] – Otherwise (cIdx is equal to 2), recSamples corresponds to the reconstructed chrominance sample array S Cr , and the function clipCidx1 corresponds to Clip1 C .

[0338] When the value of tile_group_reshaper_enabled_flag equals 1, according to the mapping process specified in Clause 8.5.5.1 derive the block of the reconstructed sample array recSamples of (nCurrSw) x (nCurrSh) at the position (xCurr, yCurr). Otherwise, the block of the reconstructed sample array recSamples of (nCurrSw) x (nCurrSh) at the position (xCurr, yCurr) is derived as follows: xxx

[0339] recSamples[xCurr + i][yCurr + j] = clipCidx1(predSamples[i][j] + resSamples[i][j]) (8 - xxx) where i = 0..nCurrSw - 1, j = 0..nCurrSh–1

[0340] 8.5.5.1 Picture reconstruction using mapping processing

[0341] This clause specifies picture reconstruction using mapping processing. Picture reconstruction using mapping processing for luminance sample values is specified in 8.5.5.1.1. Picture reconstruction using mapping processing for chrominance sample values is specified in 8.5.5.1.2.

[0342] 8.5.5.1.1 Picture reconstruction using mapping processing for luminance sample values

[0343] The inputs to this processing are as follows:

[0344] – An array predSamples of (nCurrSw) x (nCurrSh), which specifies the luminance prediction samples of the current block

[0345] – An array resSamples of (nCurrSw) x (nCurrSh), which specifies the luminance residual samples of the current block.

[0346] The outputs of this processing are as follows:

[0347] – An array predMapSamples of mapped luminance prediction samples of (nCurrSw) x (nCurrSh)

[0348] – An array recSamples of reconstructed luminance samples of (nCurrSw) x (nCurrSh).

[0349] predMapSamples is derived as follows:

[0350] – If (CuPredMode[xCurr][yCurr] == MODE_INTRA) || (CuPredMode[xCurr][yCurr] == MODE_CPR) || (CuPredMode[xCurr][yCurr] == MODE_INTER && mh_intra_flag[xCurr][yCurr])

[0351] predMapSamples[xCurr + i][yCurr + j] = predSamples[i][j] (8 - xxx)

[0352] where i = 0..nCurrSw-1, j = 0..nCurrSh–1

[0353] – Otherwise ((CuPredMode[xCurr][yCurr] == MODE_INTER &&!mh_intra_flag[xCurr][yCurr])), the following applies:

[0354] shiftY = 14

[0355] idxY = predSamples[i][j] >> Log2(OrgCW) predMapSamples[xCurr+i][yCurr+j] = ReshapePivot[idxY] + (ScaleCoeff[idxY] * (predSamples[i][j] - InputPivot[idxY]) + (1 << (shiftY–1))) >> shiftY (8-xxx)

[0356] where i = 0..nCurrSw-1, j = 0..nCurrSh–1

[0357] recSamples is derived as follows:

[0358] recSamples[xCurr+i][yCurr+j] = Clip1 Y (predMapSamples[xCurr+i][yCurr+j] + resSamples[i][j]]) (8- (8 - xxx) )

[0359] where i = 0..nCurrSw-1, j = 0..nCurrSh–1

[0360] 8.5.5.1.2 Picture reconstruction using mapped processing of chrominance sample values

[0361] The input to this processing is:

[0362] – An array mapping predMapSamples of (nCurrSwx2) x (nCurrShx2), which specifies the mapped luminance prediction samples of the current block,

[0363] – An array predSamples of (nCurrSw) x (nCurrSh), which specifies the chrominance prediction samples of the current block,

[0364] – An array resSamples of (nCurrSw) x (nCurrSh), which specifies the chrominance residual samples of the current block.

[0365] The output of this process is the reconstructed chroma sample array recSamples.

[0366] recSamples is exported as follows:

[0367] – If (!tile_group_reshaper_chroma_residual_scale_flag || ((nCurrSw) x (nCurrSh) <= 4))

[0368] recSamples[xCurr + i][yCurr + j] = Clip1 C (predSamples[i][j] + resSamples[i][j]) (8 - xx)

[0369] where i = 0..nCurrSw - 1, j = 0..nCurrSh – 1

[0370] – Otherwise (tile_group_reshaper_chroma_residual_scale_flag && ((nCurrSw) x (nCurrSh) > 4)), the following applies:

[0371] The variable varScale is exported as follows:

[0372] 1. invAvgLuma = Clip1 Y ((Σ i Σ j predMapSamples[(xCurr << 1) + i][(yCurr << 1) + j] + nCurrSw * nCurrSh * 2) / (nCurrSw * nCurrSh * 4))

[0373] 2. The variable idxYInv is an identifier indexed by introducing the piecewise function specified in Clause 8.5.6.2, and is exported with the sample value invAvgLuma as the input.

[0374] 3. varScale = ChromaScaleCoef[idxYInv]

[0375] recSamples is exported as follows:

[0376] – If tu_cbf_cIdx[xCurr][yCurr] is equal to 1, the following applies:

[0377] shiftC = 11

[0378] recSamples[xCurr + i][yCurr + j] = ClipCidx1

[0379] (predSamples[i][j] + Sign(resSamples[i][j]) * ((Abs(resSamples[i][j]) * varScale + (1 << (shiftC – 1))) >> shiftC)) (8 - xxx)

[0380] where i = 0..nCurrSw - 1, j = 0..nCurrSh – 1

[0381] – Otherwise (tu_cbf_cIdx[xCurr][yCurr] equals 0)

[0382] recSamples[xCurr + i][yCurr + j] = ClipCidx1(predSamples[i][j]) (8 - xxx)

[0383] where i = 0..nCurrSw - 1, j = 0..nCurrSh – 1

[0384] 8.5.6 Picture inverse mapping processing

[0385] This clause is invoked when the value of tile_group_reshaper_enabled_flag is equal to 1. The input is the reconstructed picture luminance sample array S L and the output is the modified reconstructed picture luminance sample array S' after inverse mapping processing L .

[0386] The inverse mapping processing of luminance sample values is specified in 8.4.6.1

[0387] 8.5.6.1 Picture inverse mapping processing of luminance sample values

[0388] The input to this processing is the luminance position (xP, yP), which specifies the luminance sample position relative to the top - left luminance sample of the current picture

[0389] The output of this processing is the inverse - mapped luminance sample value invLumaSample

[0390] The value of invLumaSample is derived by applying the following sequential steps

[0391] 1. The variable idxYInv is an identifier that introduces the piece - wise function index specified in clause 8.5.6.2, with the luminance sample value S L [xP][yP] as the input for derivation

[0392] The value of 2.reshapeLumaSample is derived as follows:

[0393] shiftY = 14

[0394] invLumaSample = InputPivot[idxYInv] + (InvScaleCoeff[idxYInv] * (S L [xP][yP] - ReshapePivot[idxYInv]) + (1 << (shiftY – 1))) >> shiftY (8 - xx)

[0395] 3.clipRange = ((reshape_model_min_bin_idx > 0) && (reshape_model_max_bin_idx < MaxBinIdx));

[0396] – If clipRange equals 1, the following applies:

[0397] minVal = 16 << (BitDepth Y – 8)

[0398] maxVal = 235 << (BitDepth Y – 8)

[0399] invLumaSample = Clip3(minVal, maxVal, invLumaSample)

[0400] – Otherwise (clipRange equals 0),

[0401] invLumaSample = ClipCidx1(invLumaSample).

[0402] 8.5.6.2 Identification of the Piecewise Function Index for the Luminance Component

[0403] The input to this process is the luminance sample value S.

[0404] The output of this process is the index idxS, which identifies the segment to which the sample S belongs. The variable idxS is derived as follows:

[0405]

[0406] Note that an alternative implementation for finding the identifying idxS is as follows:

[0407]

[0408]

[0409] 2.5.2.4 Use of ILR

[0410] On the encoder side, each picture (or slice group) is first transformed into the integer domain. And all encoding and decoding processes are performed in the integer domain. For intra prediction, neighboring blocks are in the integer domain; for inter prediction, the reference blocks (generated from the original domain of the decoded picture buffer) are first transformed into the integer domain. Then the residuals are generated and encoded / decoded into the bitstream.

[0411] After encoding / decoding of the entire picture (or slice group) is completed, the samples in the integer domain are transformed back to the original domain, and then the deblocking filter and other filters are applied.

[0412] The forward integer shaping of the prediction signal is disabled in the following cases.

[0413] – The current block is intra-coded;

[0414] – The current block is coded as CPR (Current Picture Reference, also known as Intra Block Copy, IBC);

[0415] – The current block is coded as a Combined Inter / Intra Pattern (CIIP), and forward integer shaping is disabled for the intra prediction block.

[0416] 2.6 Virtual Pipeline Data Unit (VPDU)

[0417] The Virtual Pipeline Data Unit (VPDU) is defined as non-overlapping MxM luminance (L) / NxN chrominance (C) units in a picture. In a hardware decoder, multiple pipeline stages process consecutive VPDUs simultaneously; different stages process different VPDUs. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so it is important to keep the VPDU size small. In the HEVC hardware decoder, the VPDU size is set to the maximum transform block (TB) size. Increasing the maximum TB size from 32x32-L / 16x16-C (as in HEVC) to 64x64-L / 32x32-C (as in the current VVC) can bring encoding / decoding gain, resulting in a 4-fold increase in VPDU size (64x64-L / 32x32-C) compared to HEVC. However, in addition to quadtree (QT) coding tree unit (CU) splitting, the VVC also employs a ternary tree (TT) and a binary tree (BT) to achieve additional encoding / decoding gain, and the TT and BT partitions can be recursively applied to 128x128-L / 64x64-C coding tree blocks (CTUs), resulting in a 16-fold increase in VPDU size (128x128-L / 64x64-C) compared to HEVC.

[0418] In the current design of VVC, the size of the VPDU is defined as 64x64-L / 32x32-C.

[0419] 2.7 APS

[0420] VVC uses Adaptive Parameter Sets (APS) to carry ALF parameters. The slice group header contains an aps_id that is conditionally present when ALF is enabled. The APS contains the aps_id and the ALF parameters. A new NUT (NAL unit type, as in AVC and HEVC) value is assigned for the APS (from JVET-M0132). For the regular test conditions (coming soon) in VTM-4.0, it is recommended to use only aps_id = 0 and send the APS with each picture. Currently, the range of APS ID values is 0…31, and the APS can be shared across pictures (and can be different in different slice groups within a picture). When present, the ID value shall be fixed-length coded. The ID value cannot be reused with different content within the same picture.

[0421] 2.8 Post - reconstruction filter

[0422] 2.8.1 Diffusion Filter (DF)

[0423] In JVET-L0157, a diffusion filter was proposed, where the intra / inter prediction signal of a CU can be further modified with the diffusion filter.

[0424] 2.8.1.1 Uniform diffusion filter

[0425] The uniform diffusion filter is implemented by convolving the prediction signal with a fixed mask (given as h I or h IV ), defined as follows.

[0426] In addition to the prediction signal itself, one row of reconstructed samples to the left and above the block is used as the input to the filtering signal, where these reconstructed samples can be avoided for inter blocks.

[0427] Let pred be the prediction signal on a given block obtained by intra or motion - compensated prediction. To handle the boundary points of the filter, the prediction signal needs to be extended to the prediction signal pred ext . This extended prediction can be formed in two ways: as an intermediate step, add one row of reconstructed samples to the left or above the prediction signal, and then mirror the resulting signal in all directions; or, only mirror the prediction signal itself in all directions. The latter extension is used for inter blocks. In this case, only the prediction signal itself includes the input of the extended prediction signal pred ext .

[0428] If the filter h is to be used I , it is recommended to use the above-mentioned boundary extension and replace the prediction signal pred with h I *pred. Here, the filter mask h I is given as:

[0429]

[0430] If the filter h is to be used IV , it is recommended to replace the prediction signal pred with h IV *pred.

[0431] Here, the filter h IV is given as

[0432] h IV = h I *h I *h I *h I .

[0433] 2.8.1.2 Oriented Diffusion Filter

[0434] Instead of using a signal-adaptive diffusion filter, an oriented filter - a horizontal filter h hor and a vertical filter h ver - is used, which still has a fixed mask. More precisely, the uniform diffusion filtering corresponding to the mask h I in the previous section is limited to being applied along the vertical or horizontal direction. By applying the fixed filter mask

[0435]

[0436] to the prediction signal, the vertical filter is implemented, and the horizontal filter is implemented by using the transposed mask .

[0437] 2.8.2 Bilateral Filter (BF)

[0438] The bilateral filter was proposed in JVET-L0406 and is typically applied to luminance blocks with non-zero transform coefficients and a stripe quantization parameter greater than 17. Therefore, there is no need to signal the use of the bilateral filter. If the bilateral filter is applied, it is performed on the decoded samples immediately after the inverse transform. Additionally, the filter parameters (i.e., weights) are explicitly derived from the codec information.

[0439] The filtering process is defined as:

[0440]

[0441] where P0,0 is the intensity of the current sample, and P′ 0,0 is the corrected intensity of the current sample, P k,0 and W k are the intensity and the weight parameter of the k-th neighboring sample, respectively. Figure 24 Depicts an example of a current sample and its four neighboring samples (i.e., K = 4).

[0442] More specifically, the weight W k associated with the k-th neighboring sample is defined as follows:

[0443] W k (x) = Distance k × Range k (x) (2)

[0444] where

[0445]

[0446] and σ d depends on the coding / decoding mode and the coding / decoding block size. When the TU is further divided, the described filtering process is applied to the intra-coded block and the inter-coded block to enable parallel processing.

[0447] To better capture the statistical characteristics of the video signal and improve the performance of the filter, the weight function obtained by Equation (2) is adjusted with the σ d parameter, as shown in Table 5, which depends on the coding / decoding mode and the parameters (minimum size) of the block segmentation.

[0448] Table 5 Values of σ d for different block sizes and coding / decoding modes

[0449] Minimum (block width, block height) Intra mode Inter mode 4 82 62 8 72 52 Other 52 32

[0450] To further improve the coding / decoding performance, for the inter-coded block, when the TU is not divided, the intensity difference between the current sample and one of its neighboring samples is replaced by the representative intensity difference between two windows covering the current sample and the neighboring samples. Therefore, the equation of the filtering process is modified to:

[0451]

[0452] where P k,m and P 0,m represent the m-th sample in the windows centered on P k,0 and P 0,0 respectively. In this proposal, the window size is set to 3×3. Covering P 2,0 and P 0,0The two windows are as Figure 25 shown.

[0453] 2.8.3 Hadamard Transform Domain Filter (HF)

[0454] In JVET-K0068, a loop filter in the 1D Hadamard transform domain applied at the CU level after reconstruction implements multiplication-free operation. The proposed filter is applied to all CU blocks that meet the predefined conditions, and the filter parameters are derived from the codec information.

[0455] Generally, the proposed filtering is applied to the luma reconstruction blocks with non-zero transform coefficients, except for 4x4 blocks and cases where the slice quantization parameter is greater than 17. The filter parameters are explicitly derived from the codec information. If the proposed filter is applied, it is performed immediately on the decoded samples after the inverse transform.

[0456] For each pixel from the reconstruction block, pixel processing includes the following steps:

[0457] · Scan 4 neighboring pixels (including the current pixel) around the processed pixel according to the scan pattern;

[0458] · Read the 4-point Hadamard transform of the pixel;

[0459] · Spectral filtering based on the following formula:

[0460]

[0461] where (i) is the index of the spectral component in the Hadamard spectrum, R(i) is the spectral component of the reconstructed pixel corresponding to the index, and σ is the filtering parameter derived from the codec quantization parameter QP using the following equation:

[0462] σ = 2 (1+0.126*(QP-27))

[0463] An example of the scan pattern is as Figure 26 shown.

[0464] For the pixels located on the CU boundary, the scan pattern is adjusted to ensure that all required pixels are within the current CU.

[0465] 3. Disadvantages of Existing Implementations

[0466] The current design of ILR may have the following problems:

[0467] 1. The shaping model information may never be signaled in the sequence, but in the current slice (or slice group), the tile_group_reshaper_enable_flag is set to equal 1.

[0468] 2. The stored integer model may come from a strip (or tile group) that cannot be used as a reference.

[0469] 3. An image can be divided into multiple strips (or tile groups), and each strip (or tile group) can signal integer model information.

[0470] 4. Some values and ranges (such as the range of RspCW[i]) are defined only when the bit depth is equal to 10.

[0471] 5. reshape_model_delta_max_bin_idx is not well constrained.

[0472] 6. reshaper_model_bin_delta_abs_cw_prec_minus1 is not well constrained.

[0473] 7. ReshapePivot[i] may be greater than 1<<BitDepth-1.

[0474] 8. Fixed shear parameters are used without considering the use of ILR (i.e., the minimum value is equal to 0 and the maximum value is equal to (1<<BD)-1). Here, BD represents the bit depth.

[0475] 9. The reshaping operation for the chrominance component only considers the 4:2:0 color format.

[0476] 10. Different from the luminance component (where each luminance value can be reshaped differently), only one factor is selected and used for the chrominance component. Such scaling can be incorporated into the quantization / dequantization step to reduce additional complexity.

[0477] 11. The shearing in the picture inverse mapping process can consider the upper and lower boundaries separately.

[0478] 12. For the case where i is not within the range of reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx, ReshapePivot[i] is not set correctly.

[0479] 4. Example embodiments and techniques

[0480] The detailed embodiments described below should be regarded as examples for explaining general concepts. These embodiments should not be interpreted narrowly. In addition, these embodiments can be combined in any way.

[0481] 1. It is proposed to initialize the integer model before decoding a sequence.

[0482] a. Optionally, initialize the shaping model before decoding an I slice (or picture, or slice group).

[0483] b. Optionally, initialize the shaping model before decoding an Instantaneous Decoding Refresh (IDR) slice (or picture, or slice group).

[0484] c. Optionally, initialize the shaping model before decoding a Clean Random Access (CRA) slice (or picture, or slice group).

[0485] d. Optionally, initialize the shaping model before decoding an Intra Random Access Point (I-RAP) slice (or picture, or slice group). The I-RAP slice (or picture, or slice group) may include an IDR slice (or picture, or slice group) and / or a CRA slice (or picture, or slice group) and / or a Broken Link Access (BLA) slice (or picture, or slice group).

[0486] e. In an example of initializing the shaping model, set OrgCW to be equal to (1 << BitDepth Y ) / (MaxBinIdx + 1), and for i = 0, 1, …, MaxBinIdx, ReshapePivot[i] = InputPivot[i] = i * OrgCW.

[0487] f. In an example of initializing the shaping model, for i = 0, 1, …, MaxBinIdx, ScaleCoef[i] = InvScaleCoeff[i] = 1 << shiftY.

[0488] g. In an example of initializing the shaping model, set OrgCW to be equal to (1 << BitDepth Y ) / (MaxBinIdx + 1), and for i = 0, 1, …, MaxBinIdx, RspCW[i] = OrgCW.

[0489] h. Optionally, the default shaping model information can be signaled at the sequence level (such as in the SPS) or at the picture level (such as in the PPS), and the shaping model is initialized to the default model.

[0490] i. Optionally, when the shaping model is not initialized, the constrained ILR should be disabled.

[0491] 2. It is proposed to signal the shaping model information (such as the information in tile_group_reshaper_model()) only in I slices (or pictures, or slice groups).

[0492] a. Optionally, the shaping model information (such as the information in tile_group_reshaper_model()) can only be signaled in an IDR slice (or picture, or slice group).

[0493] b. Optionally, the shaping model information (such as the information in tile_group_reshaper_model()) can only be signaled in a CRA slice (or picture, or slice group).

[0494] c. Optionally, the shaping model information (such as the information in tile_group_reshaper_model()) can only be signaled in an I-RAP slice (or picture, or slice group). The I-RAP slice (or picture, or slice group) can include an IDR slice (or picture, or slice group) and / or a CRA slice (or picture, or slice group) and / or a broken-link access (BLA) slice (or picture, or slice group).

[0495] d. Optionally, the shaping model information (such as the information in tile_group_reshaper_model()) can be signaled at the sequence level (such as in the SPS) or at the picture level (such as in the PPS) or at the APS.

[0496] 3. It is recommended to prohibit the use of shaping information from a picture / slice / slice group on a specific picture type (such as an IRAP picture).

[0497] a. In one example, if an I slice (or picture, or slice group) is sent after the first slice (or picture, or slice group) but before the second slice (or picture, or slice group), or the second slice (or picture, or slice group) itself is an I slice (or picture, or slice group), the shaping model information signaled in the first slice (or picture, or slice group) cannot be used by the second slice (or picture, or slice group).

[0498] b. Optionally, if an IDR slice (or picture, or slice group) is sent after the first slice (or picture, or slice group) but before the second slice (or picture, or slice group), or the second slice (or picture, or slice group) itself is an IDR slice (or picture, or slice group), the shaping model information signaled in the first slice (or picture, or slice group) cannot be used by the second slice (or picture, or slice group).

[0499] c. Optionally, if the CRA strip (or picture, or slice group) is sent after the first strip (or picture, or slice group) but before the second strip (or picture, or slice group), or the second strip (or picture, or slice group) itself is a CRA strip (or picture, or slice group), the shaping model information signaled in the first strip (or picture, or slice group) cannot be used by the second strip (or picture, or slice group).

[0500] d. Optionally, if the I-RAP strip (or picture, or slice group) is sent after the first strip (or picture, or slice group) but before the second strip (or picture, or slice group), or the second strip (or picture, or slice group) itself is an I-RAP strip (or picture, or slice group), the shaping model information signaled in the first strip (or picture, or slice group) cannot be used by the second strip (or picture, or slice group). The I-RAP strip (or picture, or slice group) may include an IDR strip (or picture, or slice group) and / or a CRA strip (or picture, or slice group) and / or a broken-link access (BLA) strip (or picture, or slice group).

[0501] 4. In one example, a flag is signaled in the I strip (or picture, or slice group). If the flag is X, the shaping model information is signaled in this strip (or picture, or slice group); otherwise, the shaping model is initialized before decoding this strip (or picture, or slice group). For example, X = 0 or 1.

[0502] a. Optionally, a flag is signaled in the IDR strip (or picture, or slice group). If the flag is X, the shaping model information is signaled in this strip (or picture, or slice group); otherwise, the shaping model is initialized before decoding this strip (or picture, or slice group). For example, X = 0 or 1.

[0503] b. Optionally, a flag is signaled in the CRA strip (or picture, or slice group). If the flag is X, the shaping model information is signaled in this strip (or picture, or slice group); otherwise, the shaping model is initialized before decoding this strip (or picture, or slice group). For example, X = 0 or 1.

[0504] c. Optionally, a flag is signaled in the I-RAP strip (or picture, or slice group). If the flag is X, the shaping model information is signaled in this strip (or picture, or slice group); otherwise, the shaping model is initialized before decoding this strip (or picture, or slice group). For example, X = 0 or 1. The I-RAP strip (or picture, or slice group) may include an IDR strip (or picture, or slice group) and / or a CRA strip (or picture, or slice group) and / or a broken-link access (BLA) strip (or picture, or slice group).

[0505] 5. In one example, if an image is divided into several stripes (or tile groups), each stripe (or tile group) should share the same shaping model information.

[0506] a. In one example, if an image is divided into several stripes (or tile groups), only the first stripe (or tile group) can signal the shaping model information.

[0507] 6. It should depend on the bit depth to initialize, manipulate, and constrain the variables used in shaping.

[0508] a. In one example, MaxBinIdx = f(BitDepth), where f is a function. For example, MaxBinIdx = 4 * 2 (BitDepth-8) - 1.

[0509] b. In one example, RspCW[i] should be in the range of g(BitDepth) to 2 * OrgCW - 1. For example, RspCW[i] should be in the range of 8 * 2 (BitDepth-8) to 2 * OrgCW - 1.

[0510] 7. In one example, if RspCW[i] equals 0, then for i = 0, 1,..., MaxBinIdx, ScaleCoef[i] = InvScaleCoeff[i] = 1 << shiftY.

[0511] 8. It is proposed that reshape_model_delta_max_bin_idx should be in the range of 0 to MaxBinIdx - reshape_model_min_bin_idx, including 0 and MaxBinIdx - reshape_model_min_bin_idx.

[0512] a. Optionally, reshape_model_delta_max_bin_idx should be in the range of 0 to MaxBinIdx, including 0 and MaxBinIdx.

[0513] b. In one example, clip reshape_model_delta_max_bin_idx to the range from 0 to MaxBinIdx - reshape_model_min_bin_idx, including 0 and MaxBinIdx - reshape_model_min_bin_idx.

[0514] c. In one example, clip reshape_model_delta_max_bin_idx to the range from 0 to MaxBinIdx, inclusive of 0 and MaxBinIdx.

[0515] d. In one example, reshaper_model_min_bin_idx must be less than or equal to reshaper_model_max_bin_idx.

[0516] e. In one example, clip reshaper_model_max_bin_idx to the range from reshaper_model_min_bin_idx to MaxBinIdx, inclusive of reshaper_model_min_bin_idx and MaxBinIdx.

[0517] f. In one example, clip reshaper_model_min_bin_idx to the range from 0 to reshaper_model_max_bin_idx, inclusive of 0 and reshaper_model_max_bin_idx.

[0518] g. A consistent bitstream may require one or some of the above constraints.

[0519] 9. It is proposed to set reshape_model_max_bin_idx equal to reshape_model_min_bin_idx + reshape_model_delta_maxmin_bin_idx, where reshape_model_delta_maxmin_bin_idx, which is an unsigned integer, is a syntax element signaled after reshape_model_min_bin_idx.

[0520] b. In one example, reshape_model_delta_maxmin_bin_idx shall be in the range from 0 to MaxBinIdx - reshape_model_min_bin_idx, inclusive of 0 and MaxBinIdx - reshape_model_min_bin_idx.

[0521] c. In one example, reshape_model_delta_maxmin_bin_idx shall be in the range from 1 to MaxBinIdx - reshape_model_min_bin_idx, inclusive of 1 and MaxBinIdx - reshape_model_min_bin_idx.

[0522] d. In one example, reshape_model_delta_maxmin_bin_idx shall be clipped to the range from 0 to MaxBinIdx - reshape_model_min_bin_idx, inclusive of 0 and MaxBinIdx - reshape_model_min_bin_idx.

[0523] e. In one example, reshape_model_delta_maxmin_bin_idx shall be clipped to the range from 1 to MaxBinIdx - reshape_model_min_bin_idx, inclusive of 1 and MaxBinIdx - reshape_model_min_bin_idx.

[0524] e. A consistent bitstream may require one or some of the above constraints.

[0525] 10. It is proposed that reshaper_model_bin_delta_abs_cw_prec_minus1 be less than a threshold T.

[0526] a. In one example, T can be a fixed number, such as 6 or 7.

[0527] b. In one example, T can depend on the bit depth.

[0528] c. A consistent bitstream may require the above constraints.

[0529] 11. It is proposed that RspCW[i] can be predicted from RspCW[i - 1], i.e., RspCW[i] = RspCW[i - 1] + RspDeltaCW[i].

[0530] a. In one example, when reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx, RspCW[i] can be predicted from RspCW[i - 1].

[0531] b. In one example, RspCW[i] is predicted by OrgCW, i.e., when i equals 0, RspCW[i] = OrgCW + RspDeltaCW[i].

[0532] c. In one example, RspCW[i] is predicted by OrgCW, i.e., when i equals reshaper_model_min_bin_idx, RspCW[i] = OrgCW + RspDeltaCW[i].

[0533] 12. It is proposed to not signal reshape_model_bin_delta_sign_CW[i], and RspDeltaCW[i] = reshape_model_bin_delta_abs_CW[i] is always positive.

[0534] a. In one example, RspCW[i] = MinV + RspDeltaCW[i].

[0535] i. In one example, MinV = 32;

[0536] ii. In one example, MinV = g(BitDepth). For example, MinV = 8 * 2 (BitDepth-8) 。

[0537] 13. The calculation of invAvgLuma can depend on the color format.

[0538] a. In one example, invAvgLuma = Clip1 Y ((∑ i ∑ j predMapSamples[(xCurr << scaleX) + i][(yCurr << scaleY) + j] + ((nCurrSw << scaleX) * (nCurrSh << scaleY) >> 1)) / ((nCurrSw << scaleX) * (nCurrSh << scaleY)))

[0539] i. For the 4:2:0 format, scaleX = scaleY = 1;

[0540] ii. For the 4:4:4 format, scaleX = scaleY = 0;

[0541] iii. For the 4:2:2 format, scaleX = 1 and scaleY = 0.

[0542] 14. It is proposed that the clipping in the image inverse mapping process can consider the upper and lower boundaries separately.

[0543] a. In one example, invLumaSample = Clip3(minVal, maxVal, invLumaSample), where minVal and maxVal are calculated according to different conditions.

[0544] i. For example, if reshape_model_min_bin_idx > 0, then minVal = T1 << (BitDepth – 8); otherwise, minVal = 0; for example, T1 = 16.

[0545] ii. For example, if reshape_model_max_bin_idx < MaxBinIdx, then maxVal = T2 << (BitDepth – 8); otherwise, maxVal = (1 << BitDepth) - 1; for example, T2 = 235.

[0546] In another example, T2 = 40.

[0547] 15. It is proposed to constrain ReshapePivot[i] to ReshapePivot[i] <= T, for example, T = (1 << BitDepth) - 1.

[0548] a. For example, ReshapePivot[i + 1] = min(ReshapePivot[i] + RspCW[i], T).

[0549] 16. The chroma QP offset (denoted as dChromaQp) can be implicitly derived for each block or TU and added to the chroma QP instead of the per-pixel domain residual values of the reshaped chroma components. In this way, the reshaping of the chroma components is incorporated into the quantization / dequantization process.

[0550] a. In one example, dChromaQp can be derived based on a representative luma value denoted as repLumaVal.

[0551] b. In one example, repLumaVal can be derived using some or all of the luma prediction values of the block or TU.

[0552] c. In one example, repLumaVal can be derived using some or all of the luma reconstruction values of the block or TU.

[0553] d. In one example, repLumaVal can be derived as the average of some or all of the luma prediction or reconstruction values of the block or TU.

[0554] e. Assume ReshapePivot[idx] <= repLumaVal < ReshapePivot[idx + 1], then InvScaleCoeff[idx] can be used to derive dChromaQp.

[0555] i. In one example, dQp can be selected as argmin abs(2^(x / 6 + shiftY) – InvScaleCoeff[idx]), where x = -N…, M. For example, N = M = 63.

[0556] ii. In one example, dQp can be selected as argmin abs(1 – (2^(x / 6 + shiftY) / InvScaleCoeff[idx])), where x = -N…, M. For example, N = M = 63.

[0557] iii. In one example, for different values of InvScaleCoeff[idx], dChromaQp can be pre-computed and stored in a lookup table.

[0558] 17. A luma QP offset (denoted as dQp) can be implicitly derived for each block or TU and added to the luma QP instead of shaping the per-pixel domain residual values of the luma component. In this way, the shaping of the luma component is incorporated into the quantization / dequantization process.

[0559] a. In one example, dQp can be derived based on a representative luma value denoted as repLumaVal.

[0560] b. In one example, repLumaVal can be derived using some or all of the luma prediction values of the block or TU.

[0561] c. In one example, repLumaVal can be derived as the average of some or all of the luma prediction values of the block or TU.

[0562] d. Assume idx = repLumaVal / OrgCW, then InvScaleCoeff[idx] can be used to derive dQp.

[0563] i. In one example, dQp can be selected as argmin abs(2^(x / 6 + shiftY) – InvScaleCoeff[idx]), where x = -N…, M. For example, N = M = 63.

[0564] ii. In one example, dQp can be selected as argmin abs(1–(2^(x / 6+shiftY) / InvScaleCoeff[idx])), where x = -N…,M. For example, N = M = 63.

[0565] iii. In one example, for different values of InvScaleCoeff[idx], dQp can be pre-computed and stored in a lookup table.

[0566] In this case, dChromaQp can be set to be equal to dQp.

[0567] 5. Example Implementations of the Disclosed Technology

[0568] Figure 27A is a block diagram of a video processing apparatus 2700. The apparatus 2700 can be used to implement one or more of the methods described herein. The apparatus 2700 can be implemented in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 2700 can include one or more processors 2702, one or more memories 2704, and video processing hardware 2706. The (one or more) processors 2702 can be configured to implement one or more of the methods described herein. The (one or more) memories 2704 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 2706 can be used to implement some of the techniques described herein in hardware circuitry and can be partially or fully part of the processor 2702 (e.g., a graphics processing unit core GPU or other signal processing circuitry).

[0569] Figure 27B is another example of a block diagram of a video processing system that can implement the disclosed technology. Figure 27B is a block diagram showing an example video processing system 4100 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 4100. The system 4100 can include an input 4102 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values) or can be received in a compressed or encoded format. The input 4102 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc. and wireless interfaces such as Wi-Fi or cellular interfaces.

[0570] System 4100 may include an encoding / decoding component 4104, which may implement various encoding / decoding or encoding methods described herein. The encoding / decoding component 4104 may reduce the average bit rate of the video from the input 4102 to the output of the encoding / decoding component 4104 to produce an encoded / decoded representation of the video. Thus, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the encoding / decoding component 4104 may be stored or transmitted via communication represented by component 4106. Component 4108 may use the stored or communicated bitstream (or encoded / decoded) representation of the video received at the input 4102 to generate pixel values or displayable video to be sent to the display interface 4110. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although some video processing operations are referred to as "encoding / decoding" operations or tools, it should be understood that encoding / decoding tools or operations are used at the encoder, and corresponding decoding tools or operations that reverse the encoding / decoding results will be performed by the decoder.

[0571] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described herein may be implemented in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0572] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied in the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block may correspond to bits juxtaposed within the bitstream defined by the syntax or propagated at different positions. For example, a macroblock may be encoded based on the transformed and encoded / decoded error residual values and also using bits in the header and other fields in the bitstream.

[0573] It should be understood that by allowing the use of the techniques disclosed herein, the disclosed methods and techniques will be beneficial for incorporation into video encoder and / or decoder embodiments within video processing devices such as smartphones, laptops, desktops, and similar devices.

[0574] Figure 28A is a flowchart of an example method 2810 of video processing. Method 2810 includes: at step 2812, determining shaping model information shared in common by a plurality of video units for the conversion between a plurality of video units of a video region of a video and an encoded / decoded representation of the plurality of video units. Method 2810 further includes: at step 2814, performing the conversion between the encoded / decoded representation of the video and the video.

[0575] Figure 28B It is a flowchart of an example method 2820 for video processing. The method 2820 includes: at step 2822, for the conversion between the coded representation of a video including one or more video regions and the video, determining the value of a variable in the shaping model information as a function of the bit depth of the video. The method 2820 further includes: at step 2824, performing the conversion based on the determination.

[0576] Figure 28C It is a flowchart of an example method 2830 for video processing. The method 2830 includes: at step 2832, for the conversion between the coded representation of a video including one or more video regions and the video, determining whether to enable or disable in-loop reshaping (ILR). The method 2830 further includes: at step 2834, performing the conversion based on the determination. In some implementations, in the case where the shaping model information is not initialized, the determination determines to disable ILR.

[0577] Figure 28D It is a flowchart of an example method 2840 for video processing. The method 2840 includes: at step 2842, for the conversion between a first video region of a video and the coded representation of the first video region, determining based on a rule whether shaping information from a second video region can be used for the conversion. The method 2840 further includes: at step 2844, performing the conversion according to the determination.

[0578] Figure 28E It is a flowchart of an example method 2850 for video processing. The method 2850 includes: at step 2852, performing the conversion between the coded representation of a video including one or more video regions and the video. In some implementations, the coded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions. In some implementations, the shaping model information provides information for reconstructing video units of a video region and / or scaling chrominance residuals of chrominance video units based on the representation of video units in a first domain and a second domain. In some implementations, the shaping model information is initialized based on an initialization rule. In some implementations, the shaping model information is included in the coded representation only if the video region is coded using a specific coding type. In some implementations, the conversion is performed between the current video region of the video and the coded representation of the current video region such that the current video region is coded using the specific coding type, wherein the coded representation conforms to a format rule that specifies that the shaping model information in the coded representation is conditional based on the value of a flag in the coded representation at the video region level.

[0579] In some implementations, the shaping model information provides information for reconstructing video regions based on representations in a first domain and a second domain, and / or for scaling the chrominance residuals of chrominance video units. In some implementations, the shaping model information includes a parameter set that includes a syntax element that specifies the difference between the maximum allowed binary number index and the maximum binary number index to be used in reconstruction, and wherein the parameter is within a range. In some implementations, the shaping model information includes a parameter set that includes the maximum binary number index to be used in reconstruction, and wherein the maximum binary number index is derived as a first value that is equal to the sum of the minimum binary number index to be used in reconstruction and a syntax element that is an unsigned integer and is signaled after the minimum binary number index.

[0580] In some implementations, the shaping model information includes a parameter set that includes a first syntax element that derives the number of bits used to represent a second syntax element that specifies an absolute delta codeword value from a corresponding binary number, and the value of the first syntax element is less than a threshold. In some implementations, the shaping model information includes a parameter set that includes an i-th parameter that represents the slope of the i-th binary number used in the ILR and that has a value based on the value of the (i - 1)-th parameter, where i is a positive integer. In some implementations, the shaping model information for the ILR includes a parameter set that includes reshape_model_bin_delta_sign_CW[i], where reshape_model_bin_delta_sign_CW[i] is not signaled and RspDeltaCW[i] = reshape_model_bin_delta_abs_CW[i] is always positive. In some implementations, the shaping model information includes a parameter set that includes an invAvgLuma parameter for scaling using a luma value depending on the color format of the video region. In some implementations, the transform includes a picture inverse mapping process that transforms reconstructed picture luma samples to modified reconstructed picture luma samples, and the picture inverse mapping process includes clipping, where the upper and lower boundaries are set to be separated from each other. In some implementations, the shaping model information includes a parameter set that includes a pivot amount that is constrained such that Pivot[i] <= T. In some implementations, the chrominance quantization parameter (QP) has an offset for deriving its value for each block or transform unit. In some implementations, the luma quantization parameter (QP) has an offset for deriving its value for each block or transform unit.

[0581] The following clause-based format can be used to describe various techniques and embodiments.

[0582] A first set of clauses describes some features and aspects of the disclosed techniques listed in the previous section.

[0583] 1. A method for visual media processing, comprising: performing a conversion between a current video block and a bitstream representation of the current video block, wherein, during the conversion, according to side information associated with a loop shaping step, using the loop shaping step to transform a representation of the current video block from a first domain to a second domain, wherein the loop shaping step is partially based on shaping model information, and wherein, during the conversion, the shaping model information is used for one or more of the following: an initialization step, a signaling step, or a decoding step.

[0584] 2. The method according to clause 1, wherein the decoding step relates to an I slice, picture, or slice group.

[0585] 3. The method according to clause 1, wherein the decoding step relates to an Instantaneous Decoding Refresh (IDR) slice, picture, or slice group.

[0586] 4. The method according to clause 1, wherein the decoding step relates to a Clean Random Access (CRA) slice, picture, or slice group.

[0587] 5. The method according to clause 1, wherein the decoding step relates to an Intra Random Access Point (I-RAP) slice, picture, or slice group.

[0588] 6. The method according to clause 1, wherein an I-RAP slice, picture, or slice group may include one or more of the following: an IDR slice, picture, or slice group, a CRA slice, picture, or slice group, or a Broken Link Access (BLA) slice, picture, or slice group.

[0589] 7. The method according to any one or more of clauses 1 to 6, wherein the initialization step occurs before the decoding step.

[0590] 8. The method according to any one or more of clauses 1 to 7, wherein the initialization step includes: setting the OrgCW quantity to (1 << BitDepthY) / (MaxBinIdx + 1), ReshapePivot[i] = InputPivot[i] = i * OrgCW (i = 0, 1,..., MaxBinIdx), where OrgCW, BitDepthY, MaxBinIdx, and ReshapePivot are quantities associated with the shaping model information.

[0591] 9. The method according to any one or more of clauses 1 to 7, wherein the initialization step includes: setting the quantity to ScaleCoef[i] = InvScaleCoeff[i] = 1 << shiftY (i = 0, 1,..., MaxBinIdx), where ScaleCoef, InvScaleCoeff, and shiftY are quantities associated with the shaping model information.

[0592] 10. The method according to any one or more of clauses 1 to 7, wherein the initialization step includes: setting the OrgCW quantity to (1 << BitDepthY) / (MaxBinIdx + 1), RspCW[i] = OrgCW (i = 0, 1,..., MaxBinIdx), and OrgCW is set to be equal to (1 << BitDepthY) / (MaxBinIdx + 1). RspCW[i] = OrgCW (i = 0, 1,..., MaxBinIdx), where OrgCW, BitDepthY, MaxBinIdx, and RspCW are quantities associated with the shaping model information.

[0593] 11. The method according to clause 1, wherein the initialization step includes: setting the shaping model information to a default value, and wherein the signaling step includes: signaling the default value included in the sequence parameter set (SPS) or the picture parameter set (PPS).

[0594] 12. The method according to clause 1, further comprising: disabling the loop shaping step when it is determined that the shaping model information is not initialized during the initialization step.

[0595] 13. The method according to clause 1, wherein during the signaling step, the shaping model information is signaled in any one of the following: I slice, picture, or slice group, IDR slice, picture, or slice group, intra random access point (I-RAP) slice, picture, or slice group.

[0596] 14. The method according to clause 1, wherein the I-RAP slice, picture, or slice group may include one or more of the following: IDR slice, picture, or slice group, CRA slice, picture, or slice group, or broken link access (BLA) slice, picture, or slice group.

[0597] 15. The method according to clause 1, wherein the signaling step includes: signaling the shaping model information included in the sequence parameter set (SPS) or the picture parameter set (PPS).

[0598] 16. A visual media processing method includes: performing a conversion between a current video block and a bitstream representation of the current video block, wherein during the conversion, according to side information associated with a loop filtering step, using the loop filtering step to transform a representation of the current video block from a first domain to a second domain, wherein the loop filtering step is partially based on filtering model information, and wherein during the conversion, the filtering model information is used for one or more of: an initialization step, a signaling step, or a decoding step, and when the filtering model information is signaled in a first picture, disabling the filtering model information in a second picture.

[0599] 17. The method according to clause 16, further comprising: if an intermediate picture is sent after the first picture but before the second picture, when the filtering model information is signaled in the first picture, disabling the utilization of the filtering model information in the second picture.

[0600] 18. The method according to clause 17, wherein the second picture is an I picture, an IDR picture, a CRA picture, or an I-RAP picture.

[0601] 19. The method according to any one or more of clauses 17 to 18, wherein the intermediate picture is an I picture, an IDR picture, a CRA picture, or an I-RAP picture.

[0602] 20. The method according to any one or more of clauses 16 to 19, wherein the picture includes a slice or a group.

[0603] 21. A visual media processing method includes: performing a conversion between a current video block and a bitstream representation of the current video block, wherein during the conversion, according to side information associated with a loop filtering step, using the loop filtering step to transform a representation of the current video block from a first domain to a second domain, wherein the loop filtering step is partially based on filtering model information, and wherein during the conversion, the filtering model information is used for one or more of: an initialization step, a signaling step, or a decoding step, and during the signaling step, signaling a flag in a picture such that the filtering model information is sent in the picture based on the flag, otherwise initializing the filtering model information in the initialization step before decoding the picture in the decoding step.

[0604] 22. The method according to clause 21, wherein the picture is an I picture, an IDR picture, a CRA picture, or an I-RAP picture.

[0605] 23. The method according to any one or more of clauses 21 to 22, wherein the picture includes a slice or a group.

[0606] 24. The method according to clause 21, wherein the value of the flag is 0 or 1.

[0607] 25. A visual media processing method, comprising: performing a conversion between a current video block and a bitstream representation of the current video block, wherein, during the conversion, using a loop filtering step to transform a representation of the current video block from a first domain to a second domain according to side information associated with the loop filtering step, wherein the loop filtering step is partially based on filtering model information, and wherein, during the conversion, using the filtering model information for one or more of: an initialization step, a signaling step, or a decoding step, and after partitioning a picture into a plurality of units, the filtering model information associated with each of the plurality of units is the same.

[0608] 26. A visual media processing method, comprising: performing a conversion between a current video block and a bitstream representation of the current video block, wherein, during the conversion, using a loop filtering step to transform a representation of the current video block from a first domain to a second domain according to side information associated with the loop filtering step, wherein the loop filtering step is partially based on filtering model information, and wherein, during the conversion, using the filtering model information for one or more of: an initialization step, a signaling step, or a decoding step, and after partitioning a picture into a plurality of units, signaling the filtering model information only in a first one of the plurality of units.

[0609] 27. The method according to any one or more of clauses 25 to 26, wherein the unit corresponds to a strip or a slice group.

[0610] 28. A visual media processing method, comprising: performing a conversion between a current video block and a bitstream representation of the current video block, wherein, during the conversion, using a loop filtering step to transform a representation of the current video block from a first domain to a second domain according to side information associated with the loop filtering step, wherein the loop filtering step is partially based on filtering model information, and wherein, during the conversion, using the filtering model information for one or more of: an initialization step, a signaling step, or a decoding step, and wherein, manipulating the filtering model information based on a bit depth value.

[0611] 29. The method according to clause 28, wherein the filtering model information includes a MaxBinIdx variable related to the bit depth value.

[0612] 30. The method according to clause 28, wherein the filtering model information includes a RspCW variable related to the bit depth value.

[0613] 30. The method according to clause 29, wherein the reshaping model information includes a reshape_model_delta_max_bin_idx variable, the value of which ranges from 0 to the MaxBinIdx variable.

[0614] 31. The method according to clause 29, wherein the reshaping model information includes a reshape_model_delta_max_bin_idx variable clipped in the range starting from 0 to a value corresponding to MaxBinIdx - reshape_model_min_bin_idx, where reshape_model_min_bin_idx is another variable in the reshaping model information.

[0615] 32. The method according to clause 29, wherein the reshaping model information includes a reshape_model_delta_max_bin_idx variable clipped in the range starting from 0 to the MaxBinIdx variable.

[0616] 33. The method according to clause 29, wherein the reshaping model information includes a reshaper_model_min_bin_idx variable and a reshaper_model_delta_max_bin_idx variable, where the reshaper_model_min_bin_idx variable is less than the reshaper_model_min_bin_idx variable.

[0617] 34. The method according to clause 29, wherein the reshaping model information includes a reshaper_model_min_bin_idx variable and a reshaper_model_delta_max_bin_idx variable, where the reshaper_model_max_bin_idx is clipped to the range from reshaper_model_min_bin_idx to MaxBinIdx.

[0618] 35. The method according to clause 29, wherein the reshaping model information includes a reshaper_model_min_bin_idx variable and a reshaper_model_delta_max_bin_idx variable, where the reshaper_model_min_bin_idx is clipped to the range from 0 to reshaper_model_max_bin_idx.

[0619] 36. The method according to clause 29, wherein the shaping model information includes a reshaper_model_bin_delta_abs_cw_prec_minus1 variable that is less than a threshold value.

[0620] 37. The method according to clause 36, wherein the threshold value is a fixed number.

[0621] 38. The method according to clause 36, wherein the threshold value is based on the bit depth value.

[0622] 39. The method according to clause 28, wherein the shaping model information includes an invAvgLuma variable calculated as invAvgLuma = Clip1Y((i j predMapSamples[(xCurr << scaleX) + i][(yCurr << scaleY) + j] + ((nCurrSw << scaleX) * (nCurrSh << scaleY) >> 1)) / ((nCurrSw << scaleX) * (nCurrSh << scaleY))).

[0623] 40. The method according to clause 28, wherein for a 4:2:0 format, scaleX = scaleY = 1.

[0624] 41. The method according to clause 28, wherein for a 4:4:4 format, scaleX = scaleY = 1.

[0625] 42. The method according to clause 28, wherein for a 4:2:2 format, scaleX = 1 and scaleY = 0.

[0626] 43. The method according to clause 28, wherein the shaping model information includes an invLumaSample variable calculated as invLumaSample = Clip3(minVal, maxVal, invLumaSample).

[0627] 44. The method according to clause 43, wherein if reshape_model_min_bin_idx > 0, then minVal = T1 << (BitDepth – 8), otherwise minVal = 0.

[0628] 45. The method according to clause 43, wherein if reshape_model_max_bin_idx < MaxBinIdx, then maxVal = T2 << (BitDepth – 8), otherwise maxVal = (1 << BitDepth) - 1.

[0629] 46. The method according to clause 44, wherein T1 = 16.

[0630] 47. The method according to clause 45, wherein T2 takes a value of 235 or 40.

[0631] 48. The method according to clause 28, wherein the shaping model information includes a ReshapePivot amount constrained in such a way that ReshapePivot[i] <= T.

[0632] 49. The method according to clause 48, wherein T can be calculated as T = (1 << BitDepth) - 1, where BitDepth corresponds to the bit depth value.

[0633] 50. The method according to any one or more of clauses 1 to 49, wherein the video processing is implemented on the encoder side.

[0634] 51. The method according to any one or more of clauses 1 to 49, wherein the video processing is implemented on the decoder side.

[0635] 52. An apparatus in a video system, comprising a processor and a non - transitory memory having instructions thereon, wherein, when the processor executes the instructions, the instructions cause the processor to implement the method according to any one or more of clauses 1 to 49.

[0636] 53. A computer program product stored on a non - transitory computer - readable medium, the computer program product comprising program code for implementing the method according to any one or more of clauses 1 to 49.

[0637] The second set of clauses describes some features and aspects of the disclosed techniques listed in the previous sections (e.g., including Examples 1 to 7).

[0638] 1. A video processing method, comprising: determining shaping model information shared in common by a plurality of video units of a video region of a video for conversion between the plurality of video units and an encoded / decoded representation of the plurality of video units; and performing conversion between the encoded / decoded representation of the video and the video, wherein the shaping model information provides information for constructing video samples in a first domain and a second domain and / or scaling chrominance residuals of chrominance video units.

[0639] 2. The method according to clause 1, wherein the plurality of video units correspond to a plurality of stripes or a plurality of slice groups.

[0640] 3. The method according to clause 2, wherein the plurality of video units are associated with the same picture.

[0641] 4. The method according to clause 1, wherein the shaping model information appears only once in the codec representation of the plurality of video units.

[0642] 5. A video processing method, comprising: determining values of variables in shaping model information as a function of the bit depth of a video for a conversion between a codec representation of a video comprising one or more video regions and the video; and performing the conversion based on the determination, wherein the shaping information is applicable to in-loop filtering (ILR) of some of the one or more video regions, and wherein the shaping information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling chrominance residuals of chrominance video units.

[0643] 6. The method according to clause 5, wherein the variable comprises a MaxBinIdx variable related to the bit depth value.

[0644] 7. The method according to clause 5, wherein the variable comprises a RspCW variable related to the bit depth value.

[0645] 8. The method according to clause 7, wherein the variable RspCW is used to derive a variable ReshapePivot, and the variable ReshapePivot is used to derive reconstruction of video units of a video region.

[0646] 9. The method according to clause 5, wherein if RspCW[i] is equal to 0 (where i ranges from 0 to MaxBinIdx), the variable comprises InvScaleCoef[i] that satisfies InvScaleCoeff[i]=1<<shiftY, where shiftY is an integer representing precision.

[0647] 10. The method according to clause 9, wherein shiftY is 11 or 14.

[0648] 11. A video processing method, comprising: performing a conversion between a codec representation of a video comprising one or more video regions and the video, wherein the codec representation comprises shaping model information applicable to in-loop filtering (ILR) of some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations of video units in a first domain and a second domain, and / or scaling chrominance residuals of chrominance video units, and wherein the shaping model information has been initialized based on initialization rules.

[0649] 12. The method according to clause 11, wherein the initialization of the shaping model information occurs before decoding a sequence of the video.

[0650] 13. The method according to clause 11, wherein the initialization of the shaping model information occurs before decoding at least one of the following: i) a video region decoded with an intra (I) codec type, ii) a video region decoded with an instantaneous decoding refresh (IDR) codec type, iii) a video region decoded with a clean random access (CRA) codec type, or iv) a video region decoded with an intra random access point (I-RAP) codec type.

[0651] 14. The method according to clause 13, wherein the video region decoded with the intra random access point (I-RAP) includes at least one of the following: i) the video region decoded with an instantaneous decoding refresh (IDR) codec type, ii) the video region decoded with a clean random access (CRA) codec type, and / or iii) the video region decoded with a broken link access (BLA) codec type.

[0652] 15. The method according to clause 11, wherein the initialization of the shaping model information includes: setting the OrgCW quantity to (1<<BitDepth Y ) / (MaxBinIdx + 1), ReshapePivot[i]=InputPivot[i]=i*OrgCW (i = 0, 1,..., MaxBinIdx), where OrgCW, BitDepth Y , MaxBinIdx, and ReshapePivot are quantities associated with the shaping model information.

[0653] 16. The method according to clause 11, wherein the initialization of the shaping model information includes: setting the quantities to ScaleCoef[i]=InvScaleCoeff[i]=1<<shiftY (i = 0, 1,..., MaxBinIdx), where ScaleCoef, InvScaleCoeff, and shiftY are quantities associated with the shaping model information.

[0654] 17. The method according to clause 11, wherein the initialization of the shaping model information includes: setting the OrgCW quantity to (1<<BitDepth Y ) / (MaxBinIdx + 1), RspCW[i]=OrgCW (i = 0, 1,..., MaxBinIdx), and OrgCW is set to be equal to (1<<BitDepth Y) / (MaxBinIdx + 1), RspCW[i]=OrgCW (i = 0, 1, …, MaxBinIdx), where OrgCW, BitDepthY, MaxBinIdx, and RspCW are quantities associated with the shaping model information.

[0655] 18. The method according to clause 11, wherein the initialization of the shaping model information includes: setting the shaping model information to a default value, and wherein the default value is included in a sequence parameter set (SPS) or a picture parameter set (PPS).

[0656] 19. A video processing method includes: determining to enable or disable in-loop reshaping (ILR) for the conversion between a coded representation of a video including one or more video regions and the video; and performing the conversion based on the determination, and wherein the coded representation includes shaping model information of the ILR applicable to some of the one or more video regions, and wherein the shaping model information provides information for reconstructing a video region based on a first domain and a second domain and / or scaling a chrominance residual of a chrominance video unit, and wherein in the case where the shaping model information is not initialized, the determination determines to disable the ILR.

[0657] 20. A video processing method includes: performing a conversion between a coded representation of a video including one or more video regions and the video, wherein the coded representation includes shaping model information of in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for video units for reconstructing a video region based on a first domain and a second domain and / or scaling a chrominance residual of a chrominance video unit, and wherein the shaping model information is included in the coded representation only when the video region is coded using a specific coding type.

[0658] 21. The method according to clause 20, wherein the video region is a slice or a picture, or a group of slices, and wherein the specific coding type is an intra (I) coding type.

[0659] 22. The method according to clause 20, wherein the video region is a slice or a picture, or a group of slices, and wherein the specific coding type is an instantaneous decoding refresh (IDR) coding type.

[0660] 23. The method according to clause 20, wherein the video region is a slice or a picture, or a group of slices, and wherein the specific coding type is a clean random access (CRA) coding type.

[0661] 24. The method according to clause 20, wherein the video region is a slice or a picture, or a group of pictures, and wherein the specific codec type is an Intra Random Access Point (I-RAP) codec type.

[0662] 25. The method according to clause 24, wherein the intermediate video region encoded and decoded using the Intra Random Access Point (I-RAP) codec type includes at least one of the following: i) the video region encoded and decoded using the Instantaneous Decoding Refresh (IDR) codec type; ii) the video region encoded and decoded using the Clean Random Access (CRA) codec type; and / or iii) the video region encoded and decoded using the Broken Link Access (BLA) codec type.

[0663] 26. The method according to clause 20, wherein the shaping model information is included at the sequence level, the picture level, or in an Adaptive Parameter Set (APS).

[0664] 27. A video processing method, comprising: determining, based on a rule, whether shaping information from a second video region can be used for the conversion between a first video region of a video and the encoded and decoded representation of the first video region; and performing the conversion according to the determination.

[0665] 28. The method according to clause 27, wherein the shaping information is used for video units that reconstruct a video region based on a first domain and a second domain, and / or for chrominance residuals of scaled chrominance video units.

[0666] 29. The method according to clause 27, wherein the rule does not allow the first video region to use the shaping model information in the case where the encoded and decoded representation includes an intermediate video region encoded and decoded using a specific codec type between the first video region and the second video region.

[0667] 30. The method according to clause 29, wherein the intermediate video region is encoded and decoded using an Intra (I) codec type.

[0668] 31. The method according to any one of clauses 29 to 30, wherein the intermediate video region is encoded and decoded using the Instantaneous Decoding Refresh (IDR) codec type.

[0669] 32. The method according to any one of clauses 29 to 31, wherein the intermediate video region is encoded and decoded using the Clean Random Access (CRA) codec type.

[0670] 33. The method according to any one of clauses 29 to 32, wherein the intermediate video region is encoded and decoded using the Intra Random Access Point (I-RAP) codec type.

[0671] 34. The method according to clause 33, wherein the intermediate video region encoded and decoded using the intra random access point (I-RAP) codec type includes at least one of the following: i) the video region encoded and decoded using the instantaneous decoding refresh (IDR) codec type, ii) the video region encoded and decoded using the clean random access (CRA) codec type, and / or iii) the video region encoded and decoded using the broken link access (BLA) codec type.

[0672] 35. The method according to any one of clauses 27 to 34, wherein the first video region, the second video region, and / or the intermediate video region correspond to stripes.

[0673] 36. The method according to any one of clauses 27 to 35, wherein the first video region, the second video region, and / or the intermediate video region correspond to pictures.

[0674] 37. The method according to any one of clauses 27 to 36, wherein the first video region, the second video region, and / or the intermediate video region correspond to slice groups.

[0675] 38. A video processing method, comprising: performing a conversion between a current video region of a video and an encoded and decoded representation of the current video region such that the current video region is encoded and decoded using a specific codec type, wherein the encoded and decoded representation conforms to format rules that specify that the shaping model information in the encoded and decoded representation is conditionally based on a value of a flag in the encoded and decoded representation at the video region level.

[0676] 39. The method according to clause 38, wherein the shaping model information provides information for reconstructing video units of a video region based on a first domain and a second domain and / or chrominance residuals of scaled chrominance video units.

[0677] 40. The method according to clause 38, wherein in a case where the value of the flag has a first value, the shaping model information is signaled in the video region, otherwise the shaping model information is initialized before decoding the video region.

[0678] 41. The method according to clause 38, wherein the first value is 0 or 1.

[0679] 42. The method according to clause 38, wherein the video region is a stripe, or a picture, or a slice group, and the specific codec type is an intra (I) codec type.

[0680] 43. The method according to clause 38, wherein the video region is a slice, or a picture, or a group of pictures, and the specific codec type is an Instantaneous Decoding Refresh (IDR) codec type.

[0681] 44. The method according to clause 38, wherein the video region is a slice, or a picture, or a group of pictures, and the specific codec type is a Clean Random Access (CRA) codec type.

[0682] 45. The method according to clause 38, wherein the video region is a slice, or a picture, or a group of pictures, and the specific codec type is an Intra Random Access Point (I-RAP) codec type.

[0683] 46. The method according to clause 45, wherein the video region encoded using the Intra Random Access Point (I-RAP) codec includes at least one of the following: i) the video region encoded using the Instantaneous Decoding Refresh (IDR) codec type, ii) the video region encoded using the Clean Random Access (CRA) codec type, and / or iii) the video region encoded using the Broken Link Access (BLA) codec type.

[0684] 47. The method according to any one of clauses 1 to 46, wherein the performing the conversion includes generating the video from the encoded representation.

[0685] 48. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein, when the processor executes the instructions, the instructions cause the processor to implement the method according to any one of clauses 1 to 47.

[0686] 49. A computer program product stored on a non-transitory computer-readable medium, the computer program product including program code for implementing the method according to any one of clauses 1 to 47.

[0687] The third set of clauses describes some features and aspects of the disclosed techniques listed in the previous sections (e.g., including Examples 8 and 9).

[0688] 1. A video processing method, comprising: performing a conversion between a coded representation of a video comprising one or more video regions and the video, wherein the coded representation comprises shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region and / or scaling chrominance residuals of chrominance video units based on representations in a first domain and a second domain, and wherein the shaping model information comprises a parameter set, the parameter set comprising a syntax element that specifies a difference between a maximum allowed binary number index and a maximum binary number index to be used in the reconstruction, and wherein the parameter is within a range.

[0689] 2. The method according to clause 1, wherein the syntax element is within a range from 0 to a difference between the maximum allowed binary number index and a minimum binary number index to be used in the reconstruction.

[0690] 3. The method according to clause 1, wherein the syntax element is within a range from 0 to the maximum allowed binary number index.

[0691] 4. The method according to any one of clauses 1 to 3, wherein the maximum allowed binary number index is equal to 15.

[0692] 5. The method according to clause 1 or 4, wherein the syntax element has been clipped to the range.

[0693] 6. The method according to clause 1, wherein the minimum binary number index is equal to or less than the maximum binary number index.

[0694] 7. The method according to clause 1, wherein the maximum binary number index has been clipped to a range from the minimum binary number index to the maximum allowed binary number index.

[0695] 8. The method according to clause 1, wherein the minimum binary number index has been clipped to a range from 0 to the maximum binary number index.

[0696] 9. A video processing method, comprising: performing a conversion between a codec representation of a video comprising one or more video regions and the video, wherein the codec representation comprises shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region and / or scaling chrominance residuals of chrominance video units based on representations in a first domain and a second domain, and wherein the shaping model information comprises a parameter set, the parameter set comprising a maximum binary number index to be used for the reconstruction, and wherein the maximum binary number index is derived as a first value, the first value being equal to the sum of a minimum binary number index to be used in the reconstruction and a syntax element that is an unsigned integer and is signaled after the minimum binary number index.

[0697] 10. The method according to clause 9, wherein the syntax element specifies a difference between the maximum binary number index and an allowed maximum binary number index.

[0698] 11. The method according to clause 10, wherein the syntax element is in a range from 0 to a second value, the second value being equal to a difference between the allowed maximum binary number index and the minimum binary number index.

[0699] 12. The method according to clause 10, wherein the syntax element is in a range from 1 to a second value, the second value being equal to a difference between the minimum binary number index and the allowed maximum binary number.

[0700] 13. The method according to clause 11 or 12, wherein the syntax element has been clipped to the range.

[0701] 14. The method according to clause 9, wherein the syntax element of the parameter set is within a required range in a compliant bitstream.

[0702] 15. The method according to any one of clauses 1 to 14, wherein performing the conversion comprises generating the video from the codec representation.

[0703] 16. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein, when the processor executes the instructions, the instructions cause the processor to implement the method according to any one of clauses 1 to 15.

[0704] 17. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for implementing the method according to any one of clauses 1 to 15.

[0705] The fourth set of clauses describes some features and aspects of the disclosed techniques listed in the previous sections (e.g., including Examples 10 to 17).

[0706] 1. A video processing method, comprising: performing a conversion between a coded representation of a video comprising one or more video regions and the video, wherein the coded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling chrominance residuals of chrominance video units, and wherein the shaping model information includes a parameter set, the parameter set including a first syntax element that derives the number of bits for representing a second syntax element, the second syntax element specifying an absolute incremental codeword value from a corresponding binary number, and wherein the value of the first syntax element is less than a threshold.

[0707] 2. The method according to clause 1, wherein the first syntax element specifies the difference between the number of bits for representing the second syntax element and 1.

[0708] 3. The method according to clause 1, wherein the threshold is a fixed value.

[0709] 4. The method according to clause 1, wherein the threshold is a variable whose value depends on the bit depth.

[0710] 5. The method according to clause 4, wherein the threshold is BitDepth – 1, where BitDepth represents the bit depth.

[0711] 6. The method according to clause 1, wherein the first syntax element is located in a compliant bitstream.

[0712] 7. A video processing method, comprising: performing a conversion between a coded representation of a video comprising one or more video regions and the video, wherein the coded representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region based on representations in a first domain and a second domain, and / or scaling chrominance residuals of chrominance video units, and wherein the shaping model information includes a parameter set, the parameter set including an i-th parameter that represents the slope of an i-th binary number used in the ILR and has a value based on the (i - 1)-th parameter, where i is a positive integer.

[0713] 8. The method according to clause 7, wherein for the case where reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx, the i-th parameter is predicted from the (i - 1)-th parameter, and reshaper_model_min_bin_idx and reshaper_model_max_bin_index indicate the minimum binary number index and the maximum binary number index used in the construction.

[0714] 9. The method according to clause 7, wherein for i equal to 0, the i-th parameter is predicted from OrgCW.

[0715] 10. The method according to clause 7, wherein for the value of i equal to reshaper_model_min_bin_idx indicating the minimum binary number index used in the construction, the i-th parameter is predicted from another parameter.

[0716] 11. A video processing method, comprising: performing a conversion between a coded representation of a video comprising one or more video regions and the video, wherein the coded representation comprises loop filtering (ILR) shaping model information applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region and / or scaling chrominance residuals of chrominance video units based on representations in a first domain and a second domain, and wherein the shaping model information for the ILR comprises a set of parameters, the set of parameters comprising reshape_model_bin_delta_sign_CW[i], not signaling reshape_model_bin_delta_sign_CW[i] and RspDeltaCW[i] = reshape_model_bin_delta_abs_CW[i] is always positive.

[0717] 12. The method according to clause 11, wherein the variable CW[i] is calculated as the sum of MinV and RspDeltaCW[i].

[0718] 13. The method according to clause 12, wherein MinV is 32 or g(BitDepth), and BitDepth corresponds to a bit depth value.

[0719] 14. A video processing method includes: performing a conversion between a codec representation of a video including one or more video regions and the video, where the codec representation includes shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, where the shaping model information provides information for reconstructing video units of a video region and / or scaling the chrominance residuals of chrominance video units based on representations in a first domain and a second domain, and where the shaping model information includes a parameter set that includes a parameter invAvgLuma that uses a luminance value for the scaling depending on the color format of the video region.

[0720] 15. The method according to clause 14, wherein the invAvgLuma is calculated as invAvgLuma = Clip1 Y ((∑ i ∑ j predMapSamples[(xCurr << scaleX) + i][(yCurr << scaleY) + j] + (cnt >> 1)) / cnt, where predMapSamples represents reconstructed luminance samples, (xCurrY, yCurrY) represents the top-left chrominance sample of the current chrominance transform block relative to the top-left chrominance sample of the current picture, (i, j) represents the position of the luminance sample involved in deriving the invAvgLuma relative to the top-left chrominance sample of the current chrominance transform block, and cnt represents the number of luminance samples involved in deriving the invAvgLuma.

[0721] 16. The method according to clause 15, wherein for the 4:2:0 format, scaleX = scaleY = 1.

[0722] 17. The method according to clause 15, wherein for the 4:4:4 format, scaleX = scaleY = 0.

[0723] 18. The method according to clause 15, wherein for the 4:2:2 format, scaleX = 1 and scaleY = 0.

[0724] 19. A video processing method includes: performing a conversion between a current video block of a video and a codec representation of the video, where the conversion includes a picture inverse mapping process of transforming reconstructed picture luminance samples into modified reconstructed picture luminance samples, where the picture inverse mapping process includes clipping, and where the upper boundary and the lower boundary are set to be separated from each other.

[0725] 20. The method according to clause 19, wherein the modified reconstructed picture luma sample comprises values invLumaSample, minVal, maxVal calculated as invLumaSample = Clip3(minVal, maxVal, invLumaSample).

[0726] 21. The method according to clause 20, wherein minVal = T1 << (BitDepth – 8) if the smallest binary number index used in the construction is greater than 0, otherwise minVal = 0.

[0727] 22. The method according to clause 20, wherein maxVal = T2 << (BitDepth – 8) if the largest binary number index used in the construction is less than the maximum allowed binary number index, otherwise maxVal = (1 << BitDepth) - 1, where BitDepth corresponds to the bit depth value.

[0728] 23. A video processing method, comprising: performing a conversion between a coded representation of a video comprising one or more video regions and the video, wherein the coded representation comprises shaping model information for in-loop reshaping (ILR) applicable to some of the one or more video regions, wherein the shaping model information provides information for reconstructing video units of a video region and / or scaling chroma residuals of chroma video units based on representations in a first domain and a second domain, and wherein the shaping model information comprises a parameter set that includes pivot amounts constrained such that Pivot[i] <= T.

[0729] 24. The method according to clause 23, wherein T is calculated as T = (1 << BitDepth) - 1, where BitDepth corresponds to the bit depth value.

[0730] 25. A video processing method, comprising: performing a conversion between a coded representation of a video comprising one or more video regions and the video, wherein the coded representation comprises information applicable to in-loop reshaping (ILR) and provides parameters for reconstructing video units of a video region and / or scaling chroma residuals of chroma video units based on representations in a first domain and a second domain, and wherein a chroma quantization parameter (QP) has an offset for deriving its value for each block or transform unit.

[0731] 26. The method according to clause 25, wherein the offset is derived based on a representative luma value (repLumaVal).

[0732] 27. The method according to clause 26, wherein the representative luminance value is derived using part or all of the luminance prediction value of a block or transform unit.

[0733] 28. The method according to clause 26, wherein the representative luminance value is derived using part or all of the luminance reconstruction value of a block or transform unit.

[0734] 29. The method according to clause 26, wherein the representative luminance value is derived as the average of part or all of the luminance prediction value or luminance reconstruction value of a block or transform unit.

[0735] 30. The method according to clause 26, wherein ReshapePivot[idx] <= repLumaVal < ReshapePivot[idx + 1] and InvScaleCoeff[idx] are used to derive the offset.

[0736] 31. The method according to clause 30, wherein the luminance QP offset is selected as argmin abs(2^(x / 6 + shiftY) – InvScaleCoeff[idx]), where x = -N…,M, and N and M are integers.

[0737] 32. The method according to clause 30, wherein the luminance QP offset is selected as argmin abs(1 – (2^(x / 6 + shiftY) / InvScaleCoeff[idx])), where x = -N…,M, and N and M are integers.

[0738] 33. The method according to clause 30, wherein for different values of InvScaleCoeff[idx], the offset is pre-computed and stored in a lookup table.

[0739] 34. A video processing method, comprising: performing a conversion between a coded representation of a video comprising one or more video regions and the video, wherein the coded representation comprises information applicable to in-loop filtering (ILR), and providing video units for reconstructing the video regions based on representations in a first domain and a second domain, and / or parameters for scaling chrominance residuals of chrominance video units, and wherein a luminance quantization parameter (QP) has an offset for deriving its value for each block or transform unit.

[0740] 35. The method according to clause 34, wherein the offset is derived based on a representative luminance value (repLumaVal).

[0741] 36. The method according to clause 35, wherein the representative luminance value is derived using part or all of the luminance prediction value of a block or transform unit.

[0742] 37. The method according to clause 35, wherein the representative luminance value is derived as an average of some or all of the luminance prediction values of the block or transform unit.

[0743] 38. The method according to clause 35, wherein idx = repLumaVal / OrgCW, and InvScaleCoeff[idx] is used to derive the offset.

[0744] 39. The method according to clause 38, wherein the offset is selected as argmin abs(2^(x / 6 + shiftY) – InvScaleCoeff[idx]), x = -N…, M, and N and M are integers.

[0745] 40. The method according to clause 38, wherein the offset is selected as argmin abs(1 – (2^(x / 6 + shiftY) / InvScaleCoeff[idx])), x = -N…, M, and N and M are integers.

[0746] 41. The method according to clause 38, wherein for different InvScaleCoeff[idx] values, the offset is pre-computed and stored in a lookup table.

[0747] 42. The method according to any one of clauses 1 to 41, wherein performing the conversion includes generating the video from the codec representation.

[0748] 43. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein when the processor executes the instructions, the instructions cause the processor to implement the method according to any one of clauses 1 to 42.

[0749] 44. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for implementing the method according to any one of clauses 1 to 42.

[0750] As used herein, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion process from the pixel representation of a video to the corresponding bitstream representation, and vice versa. For example, the bitstream representation of a current video block may correspond to a juxtaposed position within the bitstream defined by the syntax or bits propagated at different positions. For example, a macroblock may be encoded based on the transformed and coded / decoded error residual values and also using bits in the header and other fields in the bitstream. Additionally, as described in the above solution, during the conversion, the decoder may parse the bitstream based on a determination while knowing that some fields may or may not be present. Similarly, the encoder may determine whether to include some syntax fields and correspondingly generate the coded / decoded representation by including or not including the syntax fields in the coded / decoded representation.

[0751] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein may be implemented in digital electronic circuitry or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in a combination of one or more of them. The disclosed embodiments and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances affecting a machine-readable propagated signal, or one or more of them. The term "data processing apparatus" includes all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiver apparatus.

[0752] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed on one or more computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0753] The processing and logical flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processing and logical flows can also be executed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0754] For example, processors suitable for executing a computer program include general and special-purpose microprocessors, and any one or more of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor that executes the instructions and one or more storage devices that store the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, e.g., magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to one or more mass storage devices to receive data therefrom or transfer data thereto, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable hard disks; magneto-optical disks; and CD ROM and DVD ROM disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.

[0755] Although this patent document contains many details, it should not be construed as limiting any subject matter or the scope of the claims, but rather as a description of the features of particular embodiments of a particular technology. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, the various functions described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. In addition, although the above features may be described as acting in certain combinations, and even initially claimed as such, in some cases, one or more features from a claimed combination may be removed from the combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.

[0756] Likewise, although the operations are described in a particular order in the drawings, this should not be construed as requiring that such operations be performed in the particular order or sequence shown, or that all of the illustrated operations be performed, to achieve the desired result. In addition, the separation of various system components in the embodiments of this patent document should not be construed as required in all embodiments.

[0757] Only some implementations and examples are described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.

Claims

1. A video processing method, comprising: determining shaping model information shared by a plurality of video units in a video region of a video and bitstreams of the plurality of video units for conversion between the plurality of video units and the bitstreams; and performing conversion between the bitstream of the video and the video, wherein the shaping model information provides information for constructing video samples in a first domain and a second domain and / or scaling chrominance residuals of chrominance video units, wherein the shaping model information appears only once in the bitstreams of the plurality of video units.

2. The method according to claim 1, wherein the plurality of video units correspond to a plurality of stripes or a plurality of slice groups.

3. The method according to claim 2, wherein the plurality of video units are associated with the same picture.

4. The video processing method according to claim 1, further comprising: for the conversion, determining values of variables in the shaping model information as a function of the bit depth of the video, wherein the shaping model information is applicable to in-loop reshaping (ILR) of the video region, and wherein the shaping model information further provides information for reconstructing video units in the video region and / or scaling chrominance residuals of chrominance video units based on representations in the first domain and the second domain.

5. The method according to claim 4, wherein the variable includes a MaxBinIdx variable related to the bit depth value.

6. The method according to claim 4, wherein the variable includes an RspCW variable related to the bit depth value.

7. The method according to claim 6, wherein the variable RspCW is used to derive a variable ReshapePivot, and the variable ReshapePivot is used to derive the reconstruction of video units in the video region.

8. The method according to claim 4, wherein if RspCW[i] is equal to 0 and i is in the range from 0 to MaxBinIdx, the variable includes InvScaleCoef[i] that satisfies InvScaleCoeff[i] = 1 << shiftY, where shiftY is an integer representing precision.

9. The method according to claim 8, wherein shiftY is 11 or 14.

10. The video processing method according to claim 1, wherein, the bitstream includes shaping model information applicable to in-loop reshaping (ILR) in the video region, wherein the shaping model information further provides information for reconstructing the video units in the video region and / or scaling chrominance residuals of chrominance video units based on representations of video units in the first domain and the second domain, and wherein the shaping model information has been initialized based on initialization rules.

11. The method according to claim 10, wherein the initialization of the shaping model information occurs before decoding a sequence of the video.

12. The method according to claim 10, wherein the initialization of the shaping model information occurs before decoding at least one of the following: i) a video region decoded with an intra (I) codec type, ii) a video region decoded with an instantaneous decoding refresh (IDR) codec type, iii) a video region decoded with a clean random access (CRA) codec type, or iv) a video region decoded with an intra random access point (I-RAP) codec type.

13. The method according to claim 12, wherein the video region decoded with the intra random access point (I-RAP) includes at least one of the following: i) the video region decoded with the instantaneous decoding refresh (IDR) codec type, ii) the video region decoded with the clean random access (CRA) codec type, and / or iii) the video region decoded with a broken link access (BLA) codec type.

14. The method according to claim 10, wherein the initialization of the shaping model information includes: Set the OrgCW quantity to (1<<BitDepth Y ) / (MaxBinIdx + 1), ReshapePivot[i] = InputPivot[i] = i * OrgCW, for i = 0, 1, …, MaxBinIdx, where OrgCW, BitDepth Y , MaxBinIdx, and ReshapePivot are quantities associated with the shaping model information.

15. The method according to claim 10, wherein the initialization of the shaping model information includes: setting the quantities to ScaleCoef[i]=InvScaleCoeff[i]=1<<shiftY, for i = 0, 1, …, MaxBinIdx, where ScaleCoef, InvScaleCoeff, and shiftY are quantities associated with the shaping model information.

16. The method according to claim 10, wherein the initialization of the shaping model information includes: Set the OrgCW quantity to (1 << BitDepth Y ) / (MaxBinIdx + 1), and RspCW[i] = OrgCW for i = 0, 1, …, MaxBinIdx, where OrgCW, BitDepth Y , MaxBinIdx, and RspCW are quantities associated with the shaping model information.

17. The method according to claim 10, wherein, the initialization of the shaping model information includes setting the shaping model information to a default value, and wherein the default value is included in the sequence parameter set (SPS) or the picture parameter set (PPS).

18. The video processing method according to claim 1, further includes: determining to enable or disable loop shaping ILR for the conversion; wherein the bitstream includes shaping model information for the ILR applicable to the video region, wherein the shaping model information further provides information for reconstructing the video region based on a first domain and a second domain, and / or scaling the chroma residuals of chroma video units, and wherein in the case where the shaping model information is not initialized, the determination determines to disable the ILR.

19. The video processing method according to claim 1, wherein, the bitstream includes shaping model information for loop shaping (ILR) in the video region, wherein the shaping model information further provides information for video units for reconstructing the video region based on a first domain and a second domain, and / or scaling the chroma residuals of chroma video units, wherein the shaping model information is included in the bitstream only if the video region is decoded with a specific codec type.

20. The method according to claim 19, wherein the video region is a slice or a picture, or a group of slices, and wherein, the specific codec type is an intra (I) codec type.

21. The method according to claim 19, wherein the video region is a slice or a picture, or a group of slices, and wherein, the specific codec type is an Instantaneous Decoding Refresh (IDR) codec type.

22. The method according to claim 19, wherein the video region is a slice or a picture, or a group of slices, and wherein, the specific codec type is a Clean Random Access (CRA) codec type.

23. The method according to claim 19, wherein the video region is a slice or a picture, or a group of slices, and wherein, the specific codec type is an Intra Random Access Point (I-RAP) codec type.

24. The method according to claim 23, wherein the intermediate video region encoded with the Intra Random Access Point (I-RAP) codec type includes at least one of the following: i) the video region encoded with the Instantaneous Decoding Refresh (IDR) codec type ; ii) the video region encoded with the Clean Random Access (CRA) codec type; and / or iii) the video region encoded with the Broken Link Access (BLA) codec type.

25. The method according to claim 19, wherein the shaping model information is included at the sequence level, picture level, or Adaptive Parameter Set (APS).

26. The video processing method according to claim 1, further comprising: determining, based on a rule, whether shaping model information from a second video region can be used for the conversion between the first video region of the video and the bitstream of the first video region; and performing the conversion according to the determination.

27. The method according to claim 26, wherein, the shaping model information is used for video units for reconstructing a video region based on a first domain and a second domain, and / or chrominance residuals of scaled chrominance video units.

28. The method according to claim 26, wherein the rule does not allow the first video region to use the shaping model information in the case where the bitstream includes an intermediate video region encoded with a specific codec type between the first video region and the second video region.

29. The method according to claim 28, wherein, the intermediate video region is encoded with an Intra (I) codec type.

30. The method according to claim 28, wherein, the intermediate video region is encoded with an Instantaneous Decoding Refresh (IDR) codec type.

31. The method according to claim 28, wherein, the intermediate video region is encoded with a Clean Random Access (CRA) codec type.

32. The method according to claim 28, wherein, the intermediate video region is encoded with an Intra Random Access Point (I-RAP) codec type.

33. The method according to claim 32, wherein the intermediate video region encoded and decoded using the intra random access point (I-RAP) codec type includes at least one of the following: i) the video region encoded and decoded using the instantaneous decoding refresh (IDR) codec type, ii) the video region encoded and decoded using the clean random access (CRA) codec type, and / or iii) the video region encoded and decoded using the broken link access (BLA) codec type.

34. The method according to claim 29, wherein the first video region, the second video region, and / or the intermediate video region correspond to stripes.

35. The method according to claim 29, wherein the first video region, the second video region, and / or the intermediate video region correspond to pictures.

36. The method according to claim 29, wherein the first video region, the second video region, and / or the intermediate video region correspond to slice groups.

37. The video processing method according to claim 1, wherein, the conversion causes the current video region of the video to be encoded and decoded using a specific codec type, wherein the bitstream conforms to format rules that specify that the shaping model information in the bitstream is conditional based on the value of a flag in the bitstream at the video region level.

38. The method according to claim 37, wherein, the shaping model information further provides information for reconstructing video units of a video region based on a first domain and a second domain, and / or scaling chrominance residuals of chrominance video units.

39. The method according to claim 37, wherein, in a case where the value of the flag has a first value, the shaping model information is signaled in the video region, otherwise the shaping model information is initialized before decoding the video region.

40. The method according to claim 39, wherein the first value is 0 or 1.

41. The method according to claim 37, wherein the video region is a stripe, or a picture, or a slice group, and the specific codec type is the intra (I) codec type.

42. The method according to claim 37, wherein the video region is a stripe, or a picture, or a slice group, and the specific codec type is the instantaneous decoding refresh (IDR) codec type.

43. The method according to claim 37, wherein the video region is a stripe, or a picture, or a slice group, and the specific codec type is the clean random access (CRA) codec type.

44. The method according to claim 37, wherein the video region is a stripe, or a picture, or a slice group, and the specific codec type is the intra random access point (I-RAP) codec type.

45. The method according to claim 44, wherein the video region encoded and decoded using the intra random access point (I-RAP) comprises at least one of the following: i) the video region encoded using the instant decoding refresh (IDR) encoding and decoding type, ii) the video region encoded using the clean random access (CRA) encoding and decoding type, and / or iii) the video region encoded using the broken link access (BLA) encoding and decoding type.

46. The method according to any one of claims 1 to 45, wherein the performing the conversion comprises generating the video from the bitstream.

47. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein, when the processor executes the instructions, the instructions cause the processor to implement the method according to any one of claims 1 to 46.

Citation Information

Patent Citations

  • Integrated image reshaping and video coding

    WO2019006300A1