LOSSLESS ENCODING MODES FOR VIDEO ENCODING
Patent Information
- Application Number
- MX2021015957
- Authority / Receiving Office
- MX · MX
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-06
- Filing Date
- 2021-12-16
- Publication Date
- 2026-02-25
- Estimated Expiration
- 2040-06-29
AI Technical Summary
Current video coding standards like VVC face challenges in efficiently handling lossless coding modes due to unsupported large block sizes and suboptimal selection of remaining coding schemes, which can lead to increased decoder complexity and inefficiencies.
Propose methods to align the maximum remaining coding block size for lossless CUs with the maximum block size supported by transform skip mode, and adaptively select the remaining coding scheme based on predefined conditions or explicit syntax signaling, while disabling decoder-side tools like DMVR and BDOF when not beneficial for lossless encoding.
Improves encoding efficiency and reduces decoder complexity by optimizing block sizes and coding schemes for lossless encoding, ensuring compatibility with existing standards and hardware implementations.
Smart Images

Figure MX431875B0 
Figure MX431875B1
Abstract
Description
LOSSLESS ENCODING MODES FOR VIDEO ENCODING FIELD OF INVENTION This application generally relates to video encoding and compression. More specifically, this disclosure relates to improvements and simplifications of lossless encoding for video encoding. BACKGROUND OF THE INVENTION Several video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include Versatile Video Coding (VVC), Joint Scan Test Model (JEM), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Expert Group (MPEG) coding, and others. Video coding generally uses predictive methods (e.g., inter-prediction, intra-prediction, etc.) that take advantage of the redundancy present in video images or sequences. A major goal of video coding techniques is to compress video data in a way that uses a lower bit rate while avoiding or minimizing degradation to video quality. The first version of the HEVC standard was finalized in October 2013, offering approximately 50% bit rate savings or equivalent perceptual quality compared to the previous-generation H.264 / MPEG AVC video coding standard. While HEVC provides significant coding improvements over its predecessor, there is evidence that superior coding efficiency can be achieved with additional coding tools on top of HEVC. Based on this, both VECG and MPEG began exploring new coding technologies for future video coding standardization. A Joint Video Exploration Team (JVET) was formed in October 2015 by ITU-T VECG and ISO / IEC MPEG to initiate a significant study of advanced technologies that could enable substantial improvements in coding efficiency.A reference software called the joint exploration model was maintained by JVET by integrating several additional coding tools on top of the HEVC test model (HM). In October 2017, the ITU-T and ISO / IEC issued a Call for Proposals (CfP) for video compression capabilities beyond HEVC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, which demonstrated compression efficiency gains over HEVC of approximately 40%. Based on these evaluation results, JVET launched a new project to develop a new, next-generation video coding standard called Versatile Video Coding (VVC). That same month, a reference software codebase, called the VVC Test Model (VTM), was established to demonstrate a reference implementation of the VVC standard. BRIEF DESCRIPTION OF THE INVENTION In general, this disclosure describes examples of techniques related to lossless encoding modes in video coding. According to a first aspect of this disclosure, a method of lossless coding modes for video coding is provided, including: dividing a video image into a plurality of coding units (CUs) comprising a lossless CU; determining a remaining coding block size of the lossless CU; and in response to determining that the remaining coding block size of the lossless CU is greater than a predefined maximum value, splitting the remaining coding block into two or more remaining blocks for further coding. Pursuant to a second aspect of this disclosure, a method of lossless coding modes for video coding is provided, including: dividing a video image into a plurality of coding units (CUs) comprising a lossless CU; and selecting a remaining coding scheme for the lossless CU, wherein the remaining coding scheme selected for the lossless CU is the same as the remaining coding scheme used by untransformed hop-mode CUs. According to a third aspect of this disclosure, a lossless encoding mode apparatus for video encoding is provided, including: one or more processors; and a memory configured to store instructions executable by the one or more processors; wherein the one or more processors, after execution of the instructions, are configured to: divide a video image into a plurality of CUs comprising a lossless CU; determine a remaining encoding block size of the lossless CU; and in response to determining that the remaining encoding block size of the lossless CU is greater than a predefined maximum value, separate the remaining encoding block into two or more remaining blocks for further encoding. According to a fourth aspect of this disclosure, a lossless encoding mode apparatus for video encoding is provided, including: one or more processors; and a memory configured to store instructions executable by the one or more processors, wherein the one or more processors, after instruction execution, are configured to: divide a video image into a plurality of CUs comprising a lossless CU; and select a remaining encoding scheme for the lossless CU, wherein the remaining encoding scheme selected for the lossless CU is the same as the remaining encoding scheme used by the non-transform jump-mode CUs. Pursuant to a fifth aspect of this disclosure, a video coding apparatus is provided, including: one or more processors; and a non-transient storage medium configured to store instructions executable by the one or more processors; wherein the instructions, when executed, cause the one or more processors to perform acts comprising: dividing a video image into a plurality of CUs comprising a lossless CU; determining a remaining coding block size of the lossless CU; and in response to the determination that the remaining coding block size of the lossless CU is greater than a predefined maximum value, splitting the remaining coding block into two or more remaining blocks for further coding. Pursuant to a sixth aspect of this disclosure, a video encoding apparatus is provided, including: one or more processors; and a non-transient storage medium configured to store instructions executable by the one or more processors; wherein the instructions, when executed, cause the one or more processors to perform acts comprising: dividing a video image into a plurality of CUs comprising a lossless CU; and selecting a remaining encoding scheme for the lossless CU, wherein the remaining encoding scheme selected for the lossless CU is the same as the remaining encoding scheme used by the transformless hop-mode CUs. BRIEF DESCRIPTION OF THE FIGURES A more specific description of the examples in this disclosure will be provided by reference to specific examples illustrated in the accompanying drawings. Because these drawings depict only some examples and are therefore not considered limiting in scope, the examples will be described and explained in detail through the use of the accompanying drawings. FIG. 1 is a block diagram illustrating an exemplary video encoder according to some implementations of the present disclosure. FIG. 2A is a schematic diagram illustrating a quaternary block split in the multi-type tree structure according to some implementations of the present disclosure. FIG. 2B is a schematic diagram illustrating a horizontal binary block split in the multi-type tree structure according to some implementations of the present disclosure. FIG. 2C is a schematic diagram illustrating a vertical binary block split in the multi-type tree structure according to some implementations of the present disclosure. FIG. 2D is a schematic diagram illustrating a horizontal ternary block split in the multi-type tree structure according to some implementations of the present disclosure. FIG. 2E is a schematic diagram illustrating a vertical ternary block split in the multi-type tree structure according to some implementations of the present disclosure. FIG. 3 is a block diagram illustrating an exemplary video decoder according to some implementations of the present disclosure. FIG. 3A is a schematic diagram illustrating an example of decoder-side motion vector refinement (DMVR) according to some implementations of the present disclosure. FIG. 4 is a schematic diagram illustrating an example of an image that is divided into CTUs and further divided into tiles and groups of tiles according to some implementations of this disclosure. FIG. 5 is a schematic diagram illustrating another example of an image that is divided into CTUs and further divided into tiles and groups of tiles according to some implementations of this disclosure. FIG. 6A is a schematic diagram illustrating an example of TT and BT division rejected iviA / a / zuzi / uioyor according to some implementations of the present disclosure. FIG. 6B is a schematic diagram illustrating an example of TT and BT splitting rejected according to some implementations of this disclosure. FIG. 6C is a schematic diagram illustrating an example of TT and BT splitting rejected according to some implementations of this disclosure. FIG. 6D is a schematic diagram illustrating an example of TT and BT splitting rejected according to some implementations of this disclosure. FIG. 6E is a schematic diagram illustrating an example of TT and BT splitting rejected according to some implementations of this disclosure. FIG. 6F is a schematic diagram illustrating an example of TT and BT splitting rejected according to some implementations of this disclosure. FIG. 6G is a schematic diagram illustrating an example of TT and BT splitting rejected according to some implementations of this disclosure. FIG. 6H is a schematic diagram illustrating an example of TT and BT splitting rejected according to some implementations of this disclosure. FIG. 7 is a block diagram illustrating an exemplary apparatus of lossless coding modes for video coding according to some implementations of the present disclosure. FIG. 8 is a flowchart illustrating an exemplary process of lossless coding modes for video coding according to some implementations of the present disclosure. FIG. 9 is a flowchart that illustrates another exemplary process of lossless coding modes for video coding according to some implementations of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION Specific implementations will now be discussed in detail, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are raised to aid in understanding the subject matter presented here. However, it will be apparent to someone skilled in the art that several alternatives may be used. For example, it will be apparent to someone skilled in the art that the subject matter presented here can be implemented in many types of electronic devices with digital video capabilities. Reference throughout this specification to “a modality”, “a modality”, “an example”, “some modalities”, “some examples”, or similar language means that a particular feature, structure, or characteristic described is included in at least one modality or example. Features, structures, elements or characteristics described in relation to one or some modalities are also applicable to other modalities, unless expressly specified otherwise. Throughout this disclosure, the terms “first,” “second,” “third,” etc., are all used as nomenclature solely for references to relevant elements, such as devices, components, compositions, steps, etc., without implying any spatial or chronological order, unless expressly stated otherwise. For example, a “first device” and a “second device” may refer to two separately formed devices, or to two parts, components, or operational states of the same device, and may be named arbitrarily. As used herein, the term “if” or “when” may be understood as “after” or “in response to,” depending on the context. These terms, if they appear in a claim, may not indicate that the relevant limitations or features are conditional or optional. The terms “module,” “sub-module,” “circuit,” “sub-circuit,” “circuit assembly,” “circuit sub-assembly,” “unit,” or “sub-unit” may include memory (shared, dedicated, or pooled) that stores code or instructions that can be executed by one or more processors. A module may include one or more circuits with or without stored code or instructions. The module or circuit may include one or more components that are directly or indirectly connected. These components may or may not be attached to, or located adjacent to, another component. A unit or module can be implemented simply through hardware, or through a combination of hardware and software. In a pure software implementation, for example, the unit or module may include functionally related blocks of code or software components that are directly or indirectly linked to each other to perform a particular function. Figure 1 shows a block diagram illustrating an exemplary block-based hybrid video encoder 100, which can be used in conjunction with many video coding standards that utilize block-based processing. VVC is built upon the block-based hybrid video coding framework. In encoder 100, the input video signal is processed block by block, which can be called coding units (CUs). In VTM-1.0, a CU can be up to 128x128 pixels. However, unlike HEVC, which divides blocks solely on the basis of quaternary trees, in VVC, a coding unit tree (CTU) is divided into CUs to accommodate various local features based on quaternary / binary / ternary trees. By definition, a coding block tree (CTB) is an NxN block of samples for some value of N, so dividing a component into CTBs is a split.A CTU includes a luma sample CTB, two corresponding chroma sample CTBs for an image that has three sets of samples, or a sample CTB for a monochrome image or an image that is encoded using three separate color planes and syntax structures used to encode the samples. Additionally, the concept of multiple-split unit types in HEVC is eliminated; that is, the separation of CU, prediction unit (PU), and transform unit (TU) no longer exists in WC. Instead, each CU is always used as the basic unit for both prediction and transform without further divisions. In a multi-tree structure, a CTU is first split using a quaternary tree structure. Then, each leaf node of the quaternary tree can be further split using binary and ternary tree structures. As shown in Figures 2A to 2E, there are five types of splitting: quaternary split (Figure 2A), horizontal binary split (Figure 2B), vertical binary split (Figure 2C), horizontal ternary split (Figure 2D), and vertical ternary split (Figure 2E). For each given video block, a prediction is formed based on either an inter-prediction or an intra-prediction approach. In inter-prediction, one or more predictors are formed through motion estimation and motion compensation, based on pixels from previously reconstructed frames. In intra-prediction, predictors are formed based on pixels reconstructed within the current frame. Through mode decision, a better predictor can be selected to predict a given block. A prediction remainder, representing the difference between a current video block and its predictor, is sent to a Transform circuit set 102. The Transform coefficients are then sent from the Transform circuit set 102 to the Quantization circuit set 104 for entropy reduction. The quantized coefficients are then fed to an Entropy Coding circuit set 106 to generate a compressed video bitstream. As shown in FIG. 1, prediction-related information 110 from an Interprediction circuit set and / or an Intraprediction circuit set 112, such as video block splitting information, motion vectors, reference frame index, and intraprediction mode, is also fed through the Entropy Coding circuit set 106 and stored in a compressed video bitstream 114. In the encoder 100, decoder-related circuit assemblies are also needed in order to reconstruct pixels for prediction purposes. First, a prediction remainder is reconstructed through an Inverse Quantization 116 and an Inverse Transform circuit assembly 118. This reconstructed prediction remainder is combined with a Block Predictor 120 to generate unfiltered reconstructed pixels for a current video block. In spatial prediction (or “intra prediction”), it uses pixel samples from adjacent blocks already encoded (which are called reference samples) in the same video frame as the current video block to predict the current video block. Temporal prediction (also referred to as “inter-prediction”) uses reconstructed pixels from previously encoded video images to predict the current video block. Temporal prediction reduces the inherent temporal redundancy in the video signal. The temporal prediction signal for a given encoding unit (CU) or encoding block is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference images are supported, a reference image index is also sent, which is used to identify which reference image in the reference image store the temporal prediction signal originates from. After spatial and / or temporal prediction is performed, a set of intra / inter-mode decision circuits 121 in the encoder 100 selects the best prediction mode, for example, based on the distortion rate optimization method. The block predictor 120 is then subtracted from the current video block; and the resulting prediction remainder is decorrelated using the transform circuit set 102 and the quantization circuit set 104. The resulting quantized remainder coefficients are inversely quantized by the inverse quantization circuit set 116 and inversely transformed by the transform circuit set 118 to form the reconstructed remainder, which is then added back to the prediction block to form the reconstructed CU signal.In addition, loop filtering 115, such as an unblocking filter, adaptive sample compensation (SAO), and / or an adaptive loop filter (ALF), can be applied to the rebuilt CU before it is placed in the image buffer memory's reference image store 117 and used to encode future video blocks. To form the output video bitstream 114, the encoding mode (inter or intra), prediction mode information, motion information, and remaining quantized coefficients are all sent to the entropy encoding unit 106 to be compressed and packaged to form the bitstream. For example, an unblocking filter is available in AVC, HEVC, and the current version of WC. In HEVC, an additional loop filter called SAO (Sample Adaptive Compensation) is defined to further improve encoding efficiency. In the current version of the VVC standard, yet another loop filter called ALF (Adaptive Loop Filter) is being actively researched and has a good chance of being included in the final standard. These loop filter operations are optional. Performing these operations helps improve coding efficiency and visual quality. They can also be turned off by the coder to save computational complexity. It should be noted that intra prediction is generally based on unfiltered reconstructed pixels, while inter prediction is based on filtered reconstructed pixels if these filter options are turned on by the encoder 100. Figure 3 is a block diagram illustrating an example block-based video decoder 200, which can be used in conjunction with many video coding standards. This decoder 200 is similar to the section related to resident reconstruction in the encoder 100 of Figure 1. In the decoder 200, an incoming video bitstream 201 is first decoded via Entropy Decoding 202 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed via Inverse Quantization 204 and Inverse Transform 206 to obtain a reconstructed prediction remainder. A block prediction mechanism, implemented in an Intra / Inter Mode Selector 212, is configured to perform either Intra Prediction 208 or Motion Compensation 210, based on the decoded prediction information.A series of unfiltered reconstructed pixels are obtained by summarizing the remaining reconstructed prediction from the Inverse Transform 206 and a predictive output generated by the block predictor mechanism, using an adder 214. The reconstructed block can also pass through a loop filter 209 before being stored in an image buffer memory 213, which functions as a reference image store. The video reconstructed in the image buffer memory 213 can be sent to drive a display device, as well as used to predict future video blocks. In situations MA / a / ¿u<'i / uioyof where the Loop Filter 209 is turned on, a filtering operation is performed on these reconstructed pixels to derive a final Reconstructed Video Output 222. In general, the basic intra-prediction scheme applied in the VVC remains the same as that of the HEVC, except that several modules are further extended and / or improved, for example, intra-subdivision coding mode (ISP), extended intra-prediction with large-angle intra-directions, combination intra-position-dependent prediction (PDPC), and intra-4-take interpolation. Image splitting, mosaic groups, mosaics and CTUs in VVC. In VVC, a mosaic is defined as a rectangular region of CTUs within a particular mosaic column and a particular mosaic row in an image. A mosaic group is a group of an integer number of tiles in an image that are exclusively contained within a single NAL unit. Essentially, the concept of a mosaic group is the same as that of a portion as defined in HEVC. For example, images are divided into mosaic groups and individual mosaics. A mosaic is a sequence of CTUs that cover a rectangular region of an image. A mosaic group contains a number of mosaics from an image. Two mosaic group modes are supported: rasterized mosaic group mode and rectangular mosaic group mode. In rasterized mosaic group mode, a mosaic group contains a sequence of rasterized mosaics from an image. In rectangular mosaic group mode, a mosaic group contains a number of mosaics from an image that collectively form a rectangular region of the image. The mosaics within a rectangular mosaic group are in the mosaic raster order of the mosaic group. FIG. 4 shows an example of rasterized mosaic group division of an image, where the image is divided into 12 mosaics and 3 rasterized mosaic groups. FIG. 5 shows an example of rectangular mosaic group division of an image, where the image is divided into 24 mosaics (columns of 6 mosaics and rows of 4 mosaics) and 9 rectangular mosaic groups. Large block size transforms with high frequency VVC zeroing In VTM4, large block-sized transforms, up to 64x64 in size, are enabled, which is primarily useful for higher-resolution video, such as 4K and 1080p streams. High-frequency transform coefficients are zeroed out for transform blocks with a size (width or height, or both width and height) equal to 64, so only the lower-frequency coefficients are retained. For example, for an MxN transform block, where M is the block width and N is the block height, when M equals 64, only the leftmost 32 columns of transform coefficients are retained. Similarly, when N equals 64, only the top 32 rows of transform coefficients are retained. When transform hopping mode is used for a large block, the entire block is used without zeroing out any values. Virtual Conduit Data Units (VPDUs) in VVC Virtual conduit data units (VPDUs) are defined as non-overlapping units in an image. In hardware decoders, successive VPDUs are processed by multiple conduit stages simultaneously. The VPDU size is roughly proportional to the buffer size in most conduit stages, so it is important to keep the VPDU size small. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, the ternary tree (TT) and binary tree (BT) splitting can lead to an increase in the VPDU size. In order to maintain the size of VPDUs as 64x64 luma samples, the following normative splitting restrictions (with syntax signaling modification) are applied in VTM5, as shown in FIGS. 6A to 6H. For convenience, we labeled the examples in FIGS. 6A to 6D from left to right for the upper examples, and FIGS. 6E to 6H from left to right for the lower examples. - TT separation is not permitted for a CU with either width or height, or both width and height equal to 128 (FIGS. 6A, 6B, 6E, 6F, 6G and 6H). - For a CU 128xN with N<128 (i.e., width equal to 128 and height less than or equal to 128), horizontal BT is not permitted. (FIG. 6D). - For a CU Nx128 with N<128 (i.e., height equal to 128 and width less than or equal to 128), vertical BT is not permitted. (FIG. 6C). VVC Transform Coefficient Encoding Transform coefficient coding refers to the process of encoding transform coefficient quantization level values of a TU. In HEVC, the transform coefficients of a coding block are encoded using non-overlapping coefficient groups (or subblocks), and each CG contains the coefficients of a 4x4 block of a coding block. The CGs within a coding block, and the transform coefficients within a CG, are encoded according to predefined scan orders. The encoding of transform coefficient levels of a CG with at least one non-zero transform coefficient can be separated into multiple scan passes. In the first pass, the first bin (denoted by binO, also referred to as the significant_coeff_flag, which indicates that the magnitude of the coefficient is greater than 0) is encoded.Next, two scan passes for context encoding of the second / third bins (denoted by bin1 and bin2, respectively, also referred to as coeff_abs_greater1 flag and coeff_abs_greater2_flag) can be applied. Finally, two more scan passes for encoding the signal information and the remaining coefficient level values (also referred to as coeff_abs_level_remaining) are invoked, if necessary. Note that only bins in the first three scan passes are encoded in a regular mode, and these bins are referred to as regular bins in the following descriptions. In VVC 3, for each subblock, the regular encoded bins and the branch encoded bins are separated in encoding order; first, all the regular encoded bins for a subblock are transmitted, and then the branch encoded bins are transmitted. The transform coefficient levels of a subblock are encoded in four passes over the scan positions as follows: - Pass 1: significance coding (sig_flag), marker greater than 1 (gt1_flag), parity The par_level_flag (IVIA / a / ZUZl / UlOyO / ) and markers greater than 2 (gt2_flag) are processed in encoding order. If sig_flag is equal to 1, the gtl flag is encoded first (which specifies whether the absolute level is greater than 1). If gtlflag is equal to 1, the par_level_flag is additionally encoded (this specifies the absolute level parity minus 2). - Step 2: Encoding of the remaining absolute level (remainder) is processed for all scan positions with gt2_flag equal to 1 or gt1_flag equal to 1. The non-binary syntax element is binarized with Golomb-Rice code and the resulting bins are encoded in the derivation mode of the arithmetic encoding engine. - Pass 3: The absolute level (absLevel) of the coefficients for which sig_flag is not encoded in the first pass (because reaching the limit of regular encoded bins) are fully encoded in the derivation mode of the arithmetic encoding engine using a GolombRice code. - Step 4: Sign encoding (sign_flag) is processed for all scan positions with sig_coeff_flag equal to 1. It is guaranteed that no more than 32 regular encoded bins (sig_flag, parflag, gt1_flag and gt2_flag) are encoded or decoded for a 4x4 subblock. For 2x2 chroma subblocks, the number of regular encoded bins is limited to 8. The Rice parameter (ricePar) for encoding the remainder of the non-binary syntax element (in Pass 3) is derived similarly to HEVC. At the beginning of each subblock, ricePar is set to 0. After encoding a remainder of the syntax element, the Rice parameter is modified according to the predefined equation. To encode the non-binary syntax element absLevel (in Pass 4), the sum of absolute values sumAbs in a local template is determined. The variables ricePar and posZero are determined based on dependent quantization, and sumAbs is determined using a lookup table. The intermediate variable CodeValue is derived as follows: - If absLevel[k] is equal to 0, the codeValue is set equal to posZero; - otherwise, if absLevel[k] is less than or equal to posZero, the codeValue is set equal to absLevel[k]-1: - otherwise (absLevel[k] is greater than posZero), codeValue is set equal to absLevel[k]. The value of codeValue is encoded using a Golomb-Rice code with the Rice parameter ricePar. In the following disclosure description, the transform coefficient coding is also referred to as remaining coding. Decoder-side Motion Vector Refinement (DMVR) in WC Decoder-side Motion Vector Refinement (DMVR) is a technique for blocks encoded in bi-prediction fusion mode and controlled by SPS signaling marker sps_dmvr_enabled_flag. Under this mode, the two movement vectors (MV) of the block can be further refined using bilateral coincidence prediction (BM). Figure 3A is a schematic diagram illustrating an example of decoder-side motion vector refinement (DMVR). As shown in Figure 3A, the bilateral matching method is used to refine the motion information of a current CU 322 in the current image 320 by searching for the closest match between its reference blocks 302 and 312 along the motion path of the current CU in its two associated reference images, i.e., refPic in LO List 300 and refPic in L1 List 310. The rectangular blocks with patterns 322, 302, and 312 indicate the current CU and its two reference blocks based on the initial motion information from the merge mode. The rectangular blocks with patterns 304 and 314 indicate a pair of reference blocks based on a candidate MV used in the motion refinement search process, i.e., the motion vector refinement process. The MV differences between the candidate MV and the initial MV (also called the original MV) are MVdiff and -MVdiff, respectively. The candidate MV and the initial MV are both bidirectional motion vectors. During DMVR, a number of such candidate MVs around the initial MV can be checked. Specifically, for each given candidate MV, its two associated reference blocks can be located from their reference images in List 0 and List 1, respectively, and the difference between them is calculated. Such a block difference is usually measured in SAD (or sum of absolute differences), or row-subsampled SAD (i.e., SAD calculated with each row of the block involved). In the end, the candidate MV with the lowest SAD between its two reference blocks becomes the refined MV and is used to generate the bi-predicted signal as the current prediction for the current CU. In VVC, DMVR is applied to a CU that satisfies the following conditions: • Encoded with CU level fusion mode (not subblock fusion mode) with biprediction MV; • With respect to the current image, one reference image of the CU is in the past (i.e., with a smaller POC than the current image POC) and another reference image is in the future (i.e., with a larger POC than the current image POC); • The POC distances (i.e., absolute POC difference) of both reference images to the current image are the same; • CU has more than 64 luma samples in size and the height of CU is greater than 8 luma samples The refined motion vector (MV) derived through the DMVR process is used to generate interprediction samples and is also used in temporal motion vector prediction for future image encoding. The original MV is used in the unlocking process and also in spatial motion vector prediction for future CU encoding. Some additional features of DMVR are illustrated in the following sub-clauses. Bi-directional Optical Flow (BDOF) in WC The bidirectional optical flow (BDOF) tool is included in VTM5. BDOF, previously referred to as BIO, was included in JEM. Compared to the JEM version, BDOF in VTM5 is a simpler version that requires significantly less computation, especially in terms of the number of multiplications and the multiplier size. BDOF is controlled by the SPS indicator. IVIA / a / ZUZ l / UI 090 / sps_bdof_enabled_flag. BDOF is used to refine the bi-prediction signal of a CU at the 4x4 sub-block level. BDOF is applied to a CU if it meets the following conditions: 1) the CU height is not 4, and the CU is not 4x8 in size; 2) the CU is not encoded using affine mode or ATMVP fusion mode; 3) the CU is encoded using "true" bi-prediction mode, meaning one of the two reference images precedes the current image in the display order and the other follows the current image in the display order. BDOF is only applied to the luma component. As its name suggests, BDOF mode is based on the concept of optical flow, which assumes that an object's movement is smooth. BDOF adjusts the prediction sample by calculating the gradient of the current block to improve encoding efficiency. Decoder-side control for DMVR and BDOF in VVC In current VVC, BDOF / DMVR are always applied if your SPS indicator is enabled and some bi-prediction and size restrictions are met for a regular merge candidate. DMVR applies to a regular fusion mode when all of the following conditions are true: - sps dmvr enabled flag is equal to 1 - general_merge_flag[ xCb ][ yCb ] is equal to 1 - both predFlagLO[ 0 ][ 0 ] and predFlagL1[ 0 ][ 0 ] are equal to 1 mmvd_merge_flag[ xCb ][ yCb ] is equal to 0 DiffPicOrderCnt(currPic, RefPicList[O][refldxLO]) is equal to DiffPicOrderCnt(RefPicList[1][ refldxLI ], currPic) Bcwldx[ xCb ][ yCb ] is equal to 0 Both luma_weight_10_flag[ refldxLO ] and luma_weight_11_flag[ refldxLI ] are equal to 0 cbWidth is greater than or equal to 8 cbHeight is greater than or equal to 8 cbHeight*cbWidth is greater than or equal to 128 BDOF is applied to bi-prediction when all of the following conditions are true: - sps bdof enabled flag is equal to 1. - predFlagLO[ xSbldx ][ ySbldx ] and predFlagL1[ xSbldx ][ ySbldx ] are both equal to 1. - DiffPicOrderCnt(currPic, RefPicList[O][ refldxLO ])*DiffPicOrderCnt(currPic, RefPicList[1][ refldxLI ]) is less than 0. - MotionModelldx[ xCb ][ yCb ] is equal to 0. - merge_subblock_flag[ xCb ][ yCb ] is equal to 0. - sym_mvd_flag[ xCb ][ yCb ] is equal to 0. - Bcwldx[ xCb ][ yCb ] is equal to 0. - luma_weight_10_flag[ refldxLO ] and luma_weight_11_flag[ refldxLI ] are both equal to 0. - cbHeight is greater than or equal to 8 - cldx is equal to 0. iviA / a / zuz ι / ui Remaining encoding for transform jump mode CU in VVC VTM5 enables transform hopping mode, which is used for luma blocks up to and including 32x32 in size. When a CU is coded in transform hopping mode, its prediction remainder is quantized and coded using the transform hopping remainder coding process. This remainder coding process is a modification of the transform coefficient coding process described in the previous section. In transform hopping mode, the remainders of a CU are also coded in non-overlapping 4x4 subblock units. Unlike the regular transform coefficient coding process, in transform hopping mode, the last coefficient position is not signaled; instead, the coded_subblock_flag is signaled for all 4x4 subblocks in the CU in the forward scan order, i.e., from the top left subblock to the last subblock. For each subblock, if the coded subblock flag is equal to 1 (i.e., there is at least one non-zero quantized remainder in the subblock), the coding of the remaining quantized levels is performed in three scan passes: - First scan pass: significance flag (sig_coeff_flag), sign flag (coeff_sign_flag), absolute level greater than 1 flag (abs_level_gtx_flag[O]), and parity flag (par_level_flag) are encoded. For a given scan position, if coeff_sig_flag is equal to 1, then the sign flag is encoded, followed by abs_level_gtx_flag[O] (which specifies whether the absolute level is greater than 1). If abs_level_gtx_flag[O] is equal to 1, then par_level_flag is additionally encoded to specify the parity of the absolute level. - Greater than scan pass x: For each scan position whose absolute level is greater than 1, up to four abs_level_gtx_flag[¡] for i=1.. .4 are coded to indicate whether the absolute level at the given position is greater than 3, 5, 7, or 9, respectively. - Remaining Scan Pass: The remaining absolute level is encoded for all scan positions with abs_level_gtx_flag[4] equal to 1 (i.e., the absolute level is greater than 9). The remaining absolute levels are binarized using the reduced rice parameter derivation template. The bins in scan passes #1 and #2 (the first scan pass and the larger of the two scan passes) are context-coded until the maximum number of context-coded bins in the CU has been exhausted. The maximum number of context-coded bins in a remaining block is limited to 2 * block width * block height, or equivalently, 2 context-coded bins per sample position on average. The bins in the last scan pass (the remaining scan pass) are derivation-coded. Lossless encoding in HEVC Lossless encoding in HEVC is achieved by simply bypassing the transform, quantization, and loop filters (unblocking filter, adaptive sample compensation, and adaptive loop filter). The design is intended to enable lossless encoding with minimal changes required to the standard HEVC encoder and decoder implementation for mainstream applications. In HEVC, lossless encoding mode can be turned on or off at the individual CU level. This is done through a signaled cu_transquant_bypass_flag syntax at the CU level. To reduce signaling overhead where lossless encoding mode is unnecessary, the cu_transquant_bypass_flag syntax is not always signaled. It is signaled only when another syntax called transquant_bypass_enabled_flag has a value of 1. In other words, the transquant_bypass_enabled_flag syntax is used to enable the signaling of the cu_transquant_bypass_flag syntax. In HEVC, the `transquant_bypass_enabled_flag` flag is set in the Picture Settings Parameter (PPS) to indicate whether the `cu_transquant_bypass_flag` flag needs to be set for each CU within a PPS-referenced image. If this flag is set to 1, the `cu_transquant_bypass_flag` flag is sent at the CU level to signal whether the current CU is encoded in lossless mode. If this flag is set to 0 in the PPS, the `cu_transquant_bypass_flag` flag is not sent, and all CUs in the image are encoded with the transform, quantization, and loop filters involved in the process, which will generally result in some level of video quality degradation. To encode an entire image losslessly, you must set the `transquant_bypass_enabled_flag` flag in the PPS to 1 and set the CU-level `cu_transquant_bypass_flag` flag to 1 for each CU in the image.Detailed syntax signaling related to lossless mode in HEVC is illustrated below. - transquant_bypass_enabled_flag equal to 1 specifies that cu_transquant_bypass_flag is present. transquant_bypass_enabled_flag equal to 0 specifies that cu_transquant_bypass_flag is not present. - The `cu_transquant_bypass_flag` set to 1 specifies that the transform and scaling process as specified in clause 8.6 and the loop-filter process as specified in clause 8.7 are derived. When `cu_transquant_bypass_flag` is not present, it is inferred to be set to 0. iviA / a / zuz ι / ui oyor pie parameter set rbsp() { Descriptor ppspicparametersetid uc(v) pps_seq_parameter_set_id uc(v) transquantbypassenabledflag 0(1)} coding unit( xO. yO. log2CbSizc ) { Descriptor if( transquant bypass enabled flag) cu_trans<|uant_bypass_flag ac(v) if( slice type 1=1) cu_skip_flag| x() || and() | ac(v)} transfonn unit( x(). y(). xBasc. s Base. log2TrafoSizc. trafoDcpth. blkldx ) { Descriptor if( cbfChroina && !cu transquant bs pass ílag ) ac(v) chroma_qp_offset() i residual coding( xO. y(>. log2TrafoSiz.c. cldx ) ( Descriptor if( transfonn skip cnablcd flag && !cu transquant bypass flag && ( log2TrafoSizc <= Log2MaxTransfonnSkipSizc )) transform skip flagl xO || yO || cldx | ac(v) ... if( cu transquant bypass flag 11 ( CuPrcdModc| x(> || y() | = = MODE INTRA && implicate rdpcm cnablcd flag&& transfonn skip flag| í iviA / a / zuzi / uioyor In VVC, the maximum CU size is 64x64, and VPDU is also set to 64x64. The maximum block size for coefficient coding in VVC is 32x32 because of the coefficient zeroing mechanism for width / heights greater than 32. Under this constraint, the current transform hop only supports CUs up to 32x32, so the maximum block size for remaining coding can be aligned with the maximum block size for coefficient coding, which is 32x32. However, in VVC, the restriction for the remaining coding block size for a lossless CU is not defined. As a result, it is currently possible in VVC to generate a remaining block in lossless coding mode larger than 32x32, which would require the remaining coding support for blocks larger than 32x32. This is not preferred for codec implementation.In this publication, several methods are proposed to address this problem. Another issue associated with lossless encoding support in VVC is how to choose the remaining encoding scheme (also referred to as the coefficient). In current VVC, two different remaining encoding schemes are available. For a given block (or CU), the selection of the remaining encoding scheme is based on the transform jump flag of that block (or CU). Therefore, if under lossless mode the transform jump flag is assumed to be 1 in VVC, as in HEVC, the remaining encoding scheme used under transform jump mode would always be used for lossless CU mode. However, the current remaining encoding scheme used when the transform jump flag is true is primarily designed for screen content encoding.This may not be optimal for lossless encoding of regular content (i.e., non-screen-related content). In this disclosure, several methods are proposed for selecting the remaining encoding for lossless CUs. In current VVC, two decoder-side tools, namely BDOF and DMVR, refined the decoded pixel by filtering the current block, thereby improving encoding performance. However, in lossless encoding, because the prediction pixels are already perfectly predicted, BDOF and DMVR do not contribute to the encoding gain. Therefore, BDOF and DMVR should not be applied in lossless encoding because these decoder-side tools are not beneficial for VVC. However, in current WC, BDOF and DMVR are always applied if the SPS indicator is enabled and any bi-prediction and size constraints are met for a regular merge candidate. Therefore, for lossless WC encoding, it is beneficial for the performance efficiency of VVC lossless encoding if DMVR and BDOF are controlled at a lower level, i.e., the slice level and the cu level. Remaining block division for lossless CU According to an example in the disclosure, it is proposed to align the maximum remaining encoding block size for a lossless CU with the maximum block size supported by the transform hop mode. In one example, the transform hop mode can only be enabled for a remaining block whose width and height are both less than or equal to 32, meaning that the maximum remaining encoding block size under transform hop mode is 32x32. According to the example, the maximum width and / or height of the remaining block for a lossless CU is also set to 32, with a maximum remaining block size of 32x32. Whenever the width / height of a lossless CU is greater than 32, the remaining CU block is divided into multiple smaller remaining blocks with a size of 32xN and / or Nx32 such that the width or height of the smaller remaining blocks is no greater than 32.For example, a 128x32 lossless CU is divided into four 32x32 remaining blocks for further encoding. In another example, a 64x64 lossless CU is divided into four 32x32 remaining blocks. According to another example in the disclosure, it is proposed to align the maximum block size for remaining encoding for a lossless CU with the VPDU size. In one example, the maximum width / height of the remaining block for a lossless CU is set to the VPDU size (e.g., 64x64 in current VVC). Whenever the width / height of a lossless CU is greater than 64, the remaining CU block is divided into multiple smaller remaining blocks with a size of 64xN and / or Nx64, such that the width or height of the smaller remaining blocks is no greater than the width and / or height of the VPDU. For example, a 128x128 lossless CU is divided into four 64x64 remaining blocks for remaining encoding. In another example, a 128x32 lossless CU is divided into two 64x32 remaining blocks. Selection of remaining encoding scheme for a lossless mode CU In current VVC, different remaining encoding schemes are used per CU depending on whether the CU is encoded with transform hop mode. The current remaining encoding used under transform hop mode is generally more suitable for screen content encoding. According to an example in this disclosure, a lossless CU uses the same remaining encoding scheme as that used by transform-hopping mode CUs. According to another example in this disclosure, a lossless CU uses the same remaining encoding scheme as that used by transformless hop-mode CUs. According to another example in this disclosure, the remaining encoding scheme for lossless CUs is adaptively selected from existing remaining encoding schemes based on certain predefined conditions and / or procedures. These predefined conditions and / or procedures are followed by both the encoder and decoder, so no signaling in the bitstream is needed to indicate the selection. In one example, a simple screen content detection scheme can be specified and used by both the encoder and decoder. Based on the detection scheme, a current video block can be classified as screen content or regular content. If it is screen content, the remaining encoding scheme used under transform hopping mode is selected. Otherwise, the other remaining encoding scheme is selected. According to another example in this disclosure, syntax is signaled in the bitstream to explicitly specify which remaining encoding scheme is used by a lossless CU. Such syntax can be a binary flag, with each binary value indicating the selection of one of the two remaining encoding schemes. Syntax can be signaled at different levels. For example, it can be signaled in the Sequence Parameter Series (SPS), Picture Parameter Series (PPS), Portion Header, Tile Group Header, or tile. It can also be signaled at the CTU or CU level. When such syntax is signaled, all lossless CUs at the same or lower level would use the same remaining encoding scheme indicated by the syntax. For example, when syntax is signaled at the SPS level, all lossless CUs in the sequence would use the same remaining encoding scheme indicated.When syntax is signaled at the portion-header level, all lossless CUs in an image use the same remaining encoding scheme indicated in the associated PPS. If there is syntax at the CU level to indicate whether a CU is encoded in lossless mode, such as `cutransquant_bypassflag`, the syntax indicating the remaining encoding scheme is conditionally signaled based on the CU's lossless mode flag. For example, only when the lossless mode flag `cu_transquant_bypass_flag` indicates that the current CU is encoded in lossless mode is the syntax indicating the remaining encoding scheme signaled for that CU. When syntax is signaled by a flag at the portion header level, all lossless-encoded CUs within that portion use the same remaining encoding scheme identified based on the flagged flag.The remaining encoding scheme for each of the pluralities of CUs is selected based on the first signaled indicator, where the remaining encoding scheme selected for the lossless CU is the remaining encoding scheme used by transform hop mode CUs or transformless hop mode CUs depending on the indicator signaled in the portion header. According to an example in this disclosure, even for a lossless-encoded CU, a transform hop mode indicator is signaled. In this case, regardless of whether a CU is lossless-encoded or not, the selection of the remaining encoding scheme for the CU is based on its transform hop mode indicator. DMVR Deactivation In current WC, DMVR power-on / power-off control is not defined by the lossless encoding mode. In one disclosure example, it is proposed to control DMVR power-on / power-off at the slice level using a 1-bit signaling flag, `slice_disable_dmvr_flag`. In this example, the `slice_disable_dmvr_flag` needs to be signaled if `spsdmvrenabledflag` is set to 1 and the `transquant_bypass_enabled_flag` is set to 0. If the `slice_disable_dmvr_flag` is not signaled, it is inferred to be 1. If `slice_disable_dmvr_flag` is 1, the DMVR is powered off. In this case, the signaling is as follows: if( sps_dm\r_cnablcd_flag && !transquant_bypass_cnablcd_flag) slicedisahlcdmvrflaji u(l) In another example, it is proposed to control the on / off state of DMVR at the cu level using cu_transquant_bypass_flag. In one example, the cu-level control for DMVR is as follows: DMVR is applied to a regular fusion mode when all of the following conditions are true: -sps_dmvr_enabled_flag is equal to 1 - cu_transquant_bypass_flag is set equal to 0 - general_merge_fla[xCb][yCb] is equal to 1 - both predFlagLO[ 0 ][ 0 ] and predFlagLI [ 0 ][ 0 ] are equal to 1 - mmvd_merge_flag[ xCb ][ yCb ] is equal to 0 - DiffP¡cOrderCnt(currPic, RefPicList[ 0 ][refldxL0] is equal to DiffPicOrderCnt(RefP¡cL¡st[1][ refldxLI ], currPic) - Bcwldx[ xCb ][ yCb ] is equal to 0 - Both luma_weight_10_flag[ refldxLO ] and luma_weight_11_flag[ refldxLI ] are equal to 0 - cbWidth is greater than or equal to 8 - cbHeight is greater than or equal to 8 - cbHeight*cbWeight is greater than or equal to 128 BDOF Deactivation In current WC, BDOF on / off control is not defined by the lossless encoding mode. In one example in this disclosure, it is proposed to control BDOF on / off by means of a 1-bit signaling flag, slice_disable_bdof_flag. In one example, the slice_disable_bdof_flag flag is signaled if sps_bdof_enabled_flag is set to 1 or the transquant_bypass_enabled_flag flag is set to 0. If slice_disable_bdof_flag is not signaled, it is inferred to be 1. If the slice_disable_bdof_flag is 1, BDOF is disabled. In this case, the signaling is illustrated as follows: if( sps bdof cnablcd flag && Itransquantbypasscnablcdflag) sliccdisahlebdofflag u(l) In another example from this disclosure, it is proposed to control the on / off state of BDOF at the CU level using cu_transquant_bypass_flag. In one example, the cu-level control for BDOF is as follows: BDOF is applied to a rectangular blending mode when all of the following conditions are true: - sps_bdof_enabled_flag is equal to 1 - cu_transquant_bypass_flag is set equal to 0 - predFlagLO[ xSbldx ][ ySbldx ] and predFlagL1[ xSbldx ][ ySbldx ] are both equal to 1. - DiffPicOrderCnt(currPic, RefPicL¡st[0][refldxL0]*D¡ffPicOrderCnt(currPic, RefPicList[1 ][refldxL1 ] is less than 0. - MotionModelldc [xCb ][ yCb ] is equal to 0. - merge_subblock_flag[ xCb ][ yCb ] is equal to 0. - sym_mvd_flag[ xCb ][ yCb ] is equal to 0. - Bcwldx[ xCb ][ yCb ] is equal to 0 - luma_weight_10_flag[ refldxLO ] and luma_weight_11_flag[ refldxLI ] are both equal to 0. - cbHeight is greater than or equal to 8 - cldx equals 0 Deactivation of BDOF and DMVR In the current VVC, both BDOF and DMVR are always applied for decoder-side refinement to improve decoding efficiency and are controlled by each SPS flag and the condition that certain bi-prediction and size constraints are met for a regular merge candidate. In an example in this disclosure, it is proposed to disable both BDOF and DMVR using a 1-bit slice flag, `slice_disable_bdof_dmvr_flag`. If the `slice_disable_bdof_dmvr_flag` flag is set to 1, both BDOF and DMVR are turned off. If the `slice_disable_bdof_dmvr_flag` flag is not flagged, it is inferred to be 1. In one example, `slice_disable_bdof_dmvr_flag` is flagged if the following condition is met. if( (sps bdof cnablcd flag | sps dmvr cnablcd flag) && llransquant bypass cnablcd flag) slice disable bdof dmvr flag u(l) The methods described above can be implemented using a device that includes one or more circuit assemblies, which may include application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The device may utilize the circuit assemblies in combination with the other hardware or software components to perform the methods described above. Each module, submodule, unit, or subunit described above may be implemented at least partially using one or more of the circuit assemblies. Figure 7 is a block diagram illustrating a video encoding apparatus according to some implementations of this disclosure. The apparatus 700 can be a terminal, such as a mobile phone, tablet computer, digital broadcast terminal, tablet device, or personal digital assistant. As shown in FIG. 7, the apparatus 700 may include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 714, and a communication component 716. Processing component 702 generally controls the overall operations of apparatus 700, such as deployment operations, telephone calls, data communication, camera operation, and recording operations. Processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps described above. Additionally, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702. Memory 704 is configured to store various types of data to support operations of the Apparatus 700. Examples of such data include instructions, contact data, calendar data, messages, images, videos, etc., for any application or method operating on the Apparatus 700. Memory 704 can be implemented using any type of volatile or non-volatile storage device, or a combination thereof. Specifically, memory 704 can be Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or a compact disk. Power supply component 706 supplies power to different components of appliance 700. Power supply component 706 may include a power supply management system, one or more power supplies, and other components associated with the generation, management, and distribution of power to appliance 700. The multimedia component 708 includes a display that provides an output interface between the device 700 and a user. In some examples, the display may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the display includes a touch panel, it may be implemented as a touchscreen that receives an input signal from a user. The touch panel may include one or more touch sensors to detect a touch, a swipe, and a gesture on the touch panel. The touch sensor may detect not only the boundary of a touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some examples, the multimedia component 708 may include a front-facing and / or rear-facing camera. When the device 700 is in an operating mode, such as a shooting mode or a video mode, MA / a / ¿u<'i / uioyof The front camera and / or the rear camera can receive external multimedia data. The audio component 710 is configured to output and / or input an audio signal. For example, the audio component 710 includes a microphone (MIC). When the device 700 is in an operating mode, such as call mode, recording mode, or speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can then be stored in memory 704 or transmitted via communication component 716. In some examples, the audio component 710 also includes a speaker for outputting an audio signal. The I / O interface 712 provides an interface between the processing component 702 and a peripheral interface module. The peripheral interface module can be a keyboard, click wheel, button, or similar device. These buttons may include, but are not limited to, a home button, a volume button, a boot button, and a lock button. Sensor component 714 includes one or more sensors to provide status evaluation of various aspects of device 700. For example, sensor component 714 can detect the on / off status of device 700 and the relative locations of components. For example, the components are a display and a keypad of device 700. Sensor component 714 can also detect a change in the position of device 700 or a component of device 700, the presence or absence of a user touch on device 700, the orientation or acceleration / deceleration of device 700, and a change in the temperature of device 700. Sensor component 714 can include a proximity sensor configured to detect the presence of a nearby object without any physical touch. Sensor component 714 can also include an optical sensor, such as a CMOS or CCD image sensor used in an imaging application.In some examples, the 714 sensor component may also include an acceleration sensor, a gyroscopic sensor, a magnetic sensor, a pressure sensor, or a temperature sensor. The communication component 716 is configured to facilitate wired or wireless communication between the device 700 and other devices. The device 700 can access a wireless network based on a communication standard, such as Wi-Fi, 4G, or a combination thereof. For example, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. Alternatively, the communication component 716 may also include a Near Field Communication (NFC) module to enable short-range communication. For example, the NFC module can be implemented using Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, or other technologies. In one example, the 700 device can be implemented by one or more of the Application-Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field-Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic elements to carry out the above method. A non-transient, computer-readable storage medium can be, for example, a Hard Disk Drive (HDD), a Solid State Drive (SSD), flash memory, a Hybrid Drive or Hybrid Solid State Drive (SSHD), a Read-Only Memory (ROM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, and so on. FIG. 8 is a flowchart that illustrates an exemplary process of techniques related to lossless coding modes in video coding according to some implementations of the present disclosure. In step 801, the 720 processor divides a video image into a plurality of encoding units (CUs), at least one of which is a lossless CU. In step 802, the processor 720 determines a remaining lossless encoding block size for the CU. In step 803, the 720 processor splits the remaining encoding block into two or more remaining blocks for the remaining encoding, in response to determining that the remaining encoding block size of the lossless CU is greater than a predefined maximum value. In some examples, a video encoding apparatus is provided. The apparatus includes a 720 processor; and a 704 memory configured to store instructions executable by the processor; wherein the processor, after executing the instructions, is configured to carry out a method as illustrated in FIG. 8. In some other examples, a non-transient, computer-readable storage medium 704 is provided, which has instructions stored on it. When the instructions are executed by means of a processor 720, the instructions cause the processor to carry out a method as illustrated in FIG. 8. FIG. 9 is a flowchart illustrating an exemplary process of techniques related to lossless coding modes in video coding according to some implementations of the present disclosure. In step 901, the 720 processor divides a video image into a plurality of encoding units (CUs), at least one of which is a lossless CU. In step 902, the 720 processor selects a remaining encoding scheme for the lossless CU, wherein the remaining encoding scheme selected for the lossless CU is the same as the remaining encoding scheme used by means of transformless hop-mode CUs. In some examples, a video encoding apparatus is provided. The apparatus includes a 720 processor and a 704 memory configured to store instructions executable by the processor; where the processor, after executing the instructions, is configured to carry out a method as illustrated in FIG. 9. In some other examples, a non-transient, computer-readable storage medium 704 is provided, which has instructions stored on it. When the instructions are executed by a processor 720, the instructions cause the processor to carry out a method as illustrated in FIG. 9. MA / a / zuz i / ui oao / The description in this disclosure is for illustrative purposes only and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative implementations will be apparent to those skilled in the art who benefit from the teachings presented in the above descriptions and associated drawings. The examples were chosen and described to explain the principles of the disclosure and to enable other practitioners to understand the disclosure for various implementations and to better utilize the underlying principles and various implementations with modifications as appropriate to the intended use. It should therefore be understood that the scope of the disclosure is not limited to the specific examples of the disclosed implementations and that Modifications 10 and other implementations are intended to be included within the scope of this disclosure.
Claims
CLAIMS 1. A method for video decoding, characterized in that it comprises: determining a remaining coding scheme for a block of a video image in a transform hop mode based on a first indication and a second indication derived from a bitstream, wherein the first indication represents the transform hop mode, and the second indication represents either a first remaining coding scheme or a second remaining coding scheme for the block in the transform hop mode.
2. The method for video decoding according to claim 1, further characterized in that the second remaining encoding scheme is configured for blocks in a non-transformed skip mode.
3. The method for video decoding according to claim 1, further characterized in that it additionally comprises: signaling a 1-bit indicator for the block to control the on / off of the decoder-side motion vector refinement (DMVR) at the slice level or picture level.
4. The method for video decoding according to claim 1, further characterized in that it additionally comprises: signaling a 1-bit indicator for the block to control the on / off of the decoder-side motion vector refinement (DMVR) at the CU level.
5. The method for video decoding according to claim 1, further characterized in that it additionally comprises: signaling a 1-bit indicator for the block to control the switching on and off of the bi-directional optical flow (BDOF) at the portion level or picture level.
6. The method for video decoding according to claim 1, further characterized in that it additionally comprises: signaling a 1-bit indicator for the block to control the turning on and off of the bi-directional optical flow (BDOF) at the CU level.
7. The method for video decoding according to claim 1, further characterized in that it additionally comprises: signaling a 1-bit indicator for the block to control on / off decoder-side motion vector refinement (DMVR) and bi-directional optical flow (BDOF) at the slice or picture level.
8. A video decoding apparatus, characterized in that it comprises: one or more processors; and a memory; wherein the one or more processors are configured to: determine a remaining encoding scheme for a block of a video image in a transform hop mode based on a first indication and a second indication derived from a bitstream, wherein the first indication represents the transform hop mode, and the second indication represents a first remaining encoding scheme or a second remaining encoding scheme for the block in the transform hop mode.
9. The video decoding apparatus according to claim 8, further characterized in that the second remaining encoding scheme is configured for blocks in a non-transformed skip mode.
10. The video encoding apparatus according to claim 8, further characterized in that the one or more processors are additionally configured to: signal a 1-bit indicator for the block to control the on / off of the decoder-side motion vector refinement (DMVR) at the slice or picture level.
11. The video encoding apparatus according to claim 8, further characterized in that the one or more processors are configured to: signal a 1-bit indicator for the block to control the on / off of the decoder-side motion vector refinement (DMVR) at the CU level.
12. The video encoding apparatus according to claim 8, further characterized in that the one or more processors are additionally configured to: signal a 1-bit indicator for the block to control the bi-directional optical flow (BDOF) on and off at the portion or picture level.
13. The video encoding apparatus according to claim 8, further characterized in that the one or more processors are additionally configured to: signal a 1-bit indicator for the block to control the switching on and off of the bi-directional optical flow (BDOF) at the CU level.
14. The video encoding apparatus according to claim 8, further characterized in that the one or more processors are additionally configured to: signal a 1-bit indicator for the block to control the power on and off of both, decoder-side motion vector refinement (DMVR), and bi-directional optical flow (BDOF) at the slice or picture level.