Method and apparatus for video encoding
By segmenting video images into lossless encoding/decoding units (CUs) and adaptively dividing residual encoding/decoding blocks or selecting the same encoding/decoding scheme, the problem of insufficient compression efficiency in existing technologies is solved, achieving more efficient video encoding/decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2020-06-29
- Publication Date
- 2026-05-15
AI Technical Summary
Existing video encoding and decoding technologies still have room for improvement in compression efficiency, especially based on the HEVC standard, where more efficient encoding and decoding tools are needed to further reduce bitrate while maintaining video quality.
The lossless encoding and decoding mode is adopted, and the video images are divided into lossless encoding and decoding units (CUs). The residual encoding and decoding blocks are divided into multiple blocks as needed, or the same residual encoding and decoding scheme as the non-transform skip mode CU is selected to adapt to the characteristics of different video blocks.
It improves the compression efficiency of video encoding and decoding, reduces bit rate requirements, and maintains or improves video quality.
Smart Images

Figure CN115567708B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention application with application number 202080043764.2, application date June 29, 2020, entitled "Lossless encoding and decoding mode for video encoding and decoding". Technical Field
[0002] This application generally relates to video encoding / decoding and compression. More specifically, this disclosure relates to improvements and simplifications in lossless encoding / decoding for video encoding / decoding. Background Technology
[0003] Various video codec technologies can be used to compress video data. Video coding and decoding are performed according to one or more video codec standards. Examples of video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (H.265 / HEVC), High-Level Video Codec (H.264 / AVC), and Moving Picture Experts Group (MPEG) codecs. Video codecs typically use prediction methods that utilize redundancy present in video images or sequences (e.g., inter-frame prediction, intra-frame prediction, etc.). A key goal of video codec technologies is to compress video data to a lower bitrate while avoiding or minimizing video quality degradation.
[0004] The first version of the HEVC standard was completed in October 2013. Compared to its predecessor, H.264 / MPEG_AVC, the first version of HEVC offered approximately 50% bitrate savings or equivalent perceived quality. Despite the significant codec improvements offered by HEVC compared to its predecessor, evidence suggests that superior codec efficiency could be achieved using additional codec tools. Based on this, both VCEG and MPEG began exploring new codec technologies for future video codec standardization. In October 2015, ITU-TVECG and ISO / IEC MPEG formed a Joint Video Exploration Group (JVET) to begin important research into advanced technologies that could significantly improve codec efficiency. JVET maintains a reference software called the Joint Exploration Model (JEM) by integrating several additional codec tools on top of the HEVC test model (HM).
[0005] In October 2017, the ITU-T and ISO / IEC issued a joint call for proposals (CfP) on video compression capabilities exceeding HEVC. In April 2018, at the 10th JVET meeting, 23 CfP responses were received and evaluated, demonstrating compression efficiency gains exceeding HEVC by approximately 40%. Based on these evaluation results, JVET launched a new project to develop a next-generation video codec standard named Universal Video Codec (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard. Summary of the Invention
[0006] Generally, this disclosure describes examples of techniques related to lossless encoding / decoding modes in video encoding and decoding.
[0007] According to a first aspect of this disclosure, a method for a lossless encoding / decoding mode for video encoding / decoding is provided, comprising: segmenting a video image into a plurality of CUs including lossless encoding / decoding units (CUs); determining a residual encoding / decoding block size of the lossless CUs; and, in response to determining that the residual encoding / decoding block size of the lossless CUs is greater than a predefined maximum value, dividing the residual encoding / decoding block into two or more residual blocks for residual encoding / decoding.
[0008] According to a second aspect of this disclosure, a method for a lossless encoding / decoding mode for video encoding / decoding is provided, comprising: segmenting a video image into a plurality of CUs including lossless encoding / decoding units (CUs); and selecting a residual encoding / decoding scheme for the lossless CUs, wherein the residual encoding / decoding scheme selected for the lossless CUs is the same as the residual encoding / decoding scheme used by a non-transform skip mode CU.
[0009] According to a third aspect of this disclosure, an apparatus for a lossless encoding / decoding mode for video encoding / decoding is provided, comprising: one or more processors; and a memory configured to store instructions executable by the one or more processors; wherein the one or more processors, when executing the instructions, are configured to: segment a video image into a plurality of CUs including lossless encoding / decoding units (CUs); determine a residual encoding / decoding block size of the lossless CUs; and, in response to determining that the residual encoding / decoding block size of the lossless CUs is greater than a predefined maximum value, divide the residual encoding / decoding block into two or more residual blocks for residual encoding / decoding.
[0010] According to a fourth aspect of this disclosure, an apparatus for a lossless encoding / decoding mode for video encoding / decoding is provided, comprising: one or more processors; and a memory configured to store instructions executable by the one or more processors; wherein the one or more processors, when executing the instructions, are configured to: segment a video frame into a plurality of CUs including lossless encoding / decoding units (CUs); and select a residual encoding / decoding scheme for the lossless CUs, wherein the residual encoding / decoding scheme selected for the lossless CUs is the same as the residual encoding / decoding scheme used by a non-transform skip mode CU.
[0011] According to a fifth aspect of this disclosure, an apparatus for video encoding and decoding is provided, comprising: one or more processors; and a non-transitory storage medium configured to store instructions executable by the one or more processors; wherein the instructions, when executed, cause the one or more processors to perform actions including: segmenting a video image into a plurality of CUs including lossless encoding and decoding units (CUs); determining a residual encoding and decoding block size of the lossless CUs; and, in response to determining that the residual encoding and decoding block size of the lossless CUs is greater than a predefined maximum value, dividing the residual encoding and decoding block into two or more residual blocks for residual encoding and decoding.
[0012] According to a sixth aspect of this disclosure, an apparatus for video encoding and decoding is provided, comprising: one or more processors; and a non-transitory storage medium configured to store instructions executable by the one or more processors; wherein the instructions, when executed, cause the one or more processors to perform actions including: segmenting a video image into a plurality of CUs including lossless encoding and decoding units (CUs); and selecting a residual encoding and decoding scheme for the lossless CUs, wherein the residual encoding and decoding scheme selected for the lossless CUs is the same as the residual encoding and decoding scheme used by a non-transform skip mode CU. Attached Figure Description
[0013] A more detailed description of the examples of this disclosure will be presented with reference to the specific examples shown in the accompanying drawings. Given that these drawings depict only a few examples and are therefore not intended to limit the scope, the examples will be described and explained using additional features and details through the use of the drawings.
[0014] Figure 1 This is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.
[0015] Figure 2A This is a schematic diagram illustrating quad block segmentation in a multi-type tree structure according to some embodiments of the present disclosure.
[0016] Figure 2BThis is a schematic diagram illustrating horizontal binary block segmentation in a multi-type tree structure according to some embodiments of the present disclosure.
[0017] Figure 2C This is a schematic diagram illustrating vertical binary block segmentation in a multi-type tree structure according to some embodiments of the present disclosure.
[0018] Figure 2D This is a schematic diagram illustrating the horizontal ternary block segmentation in a multi-type tree structure according to some embodiments of the present disclosure.
[0019] Figure 2E This is a schematic diagram illustrating vertical ternary block segmentation in a multi-type tree structure according to some embodiments of the present disclosure.
[0020] Figure 3 This is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.
[0021] Figure 3A This is a schematic diagram illustrating an example of decoder-side motion vector refinement (DMVR) according to some embodiments of the present disclosure.
[0022] Figure 4 This is a schematic diagram illustrating an example of a CTU divided into tiles and tile groups according to some embodiments of the present disclosure.
[0023] Figure 5 This is a schematic diagram illustrating another example of images divided into CTUs and further subdivided into tiles and tile groups according to some embodiments of the present disclosure.
[0024] Figure 6A This is a schematic diagram illustrating an example of an implementation of this disclosure where TT and BT splitting is not allowed.
[0025] Figure 6B This is a schematic diagram illustrating an example of an implementation of this disclosure where TT and BT splitting is not allowed.
[0026] Figure 6C This is a schematic diagram illustrating an example of an implementation of this disclosure where TT and BT splitting is not allowed.
[0027] Figure 6D This is a schematic diagram illustrating an example of an implementation of this disclosure where TT and BT splitting is not allowed.
[0028] Figure 6E This is a schematic diagram illustrating an example of an implementation of this disclosure where TT and BT splitting is not allowed.
[0029] Figure 6FThis is a schematic diagram illustrating an example of an implementation of this disclosure where TT and BT splitting is not allowed.
[0030] Figure 6G This is a schematic diagram illustrating an example of an implementation of this disclosure where TT and BT splitting is not allowed.
[0031] Figure 6H This is a schematic diagram illustrating an example of an implementation of this disclosure where TT and BT splitting is not allowed.
[0032] Figure 7 This is a block diagram illustrating an exemplary apparatus for a lossless encoding / decoding mode for video encoding / decoding according to some embodiments of the present disclosure.
[0033] Figure 8 This is a flowchart illustrating an exemplary process for a lossless encoding / decoding mode for video encoding / decoding according to some embodiments of the present disclosure.
[0034] Figure 9 This is a flowchart illustrating another exemplary process of a lossless encoding / decoding mode for video encoding / decoding according to some embodiments of the present disclosure. Detailed Implementation
[0035] Referring now to the detailed description, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be used. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.
[0036] Throughout this specification, references to "an embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language indicate that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Unless otherwise expressly stated, the features, structures, elements, or characteristics described in connection with one or more embodiments also apply to other embodiments.
[0037] Throughout this disclosure, unless otherwise expressly stated, the terms “first,” “second,” “third,” etc., are used only to refer to related elements (e.g., equipment, components, compositions, steps, etc.) and do not indicate any spatial or temporal order. For example, “first equipment” and “second equipment” can refer to two separately formed devices, or two parts, components, or operating states of the same device, and can be named arbitrarily.
[0038] As used herein, depending on the context, the terms "if" or "when" may be understood to mean "at the time of" or "in response to". If these terms appear in the claims, they do not necessarily indicate that the relevant limitation or feature is conditional or optional.
[0039] The terms "module," "submodule," "circuit," "subcircuit," "circuit system," "subcircuit system," "unit," or "subunit" may include memory (shared, dedicated, or combined) storing code or instructions executable by one or more processors. A module may include one or more circuits, with or without stored code or instructions. A module or circuit may include one or more components, directly or indirectly connected. These components may or may not be physically attached to each other or positioned adjacent to each other.
[0040] Units or modules can be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a purely software implementation, for example, a unit or module may include functionally related code blocks or software components that are directly or indirectly linked together to perform a specific function.
[0041] Figure 1 A block diagram is shown illustrating an exemplary block-based hybrid video encoder 100 that can be used in conjunction with many video codec standards that use block-based processing. VVC is built on a block-based hybrid video codec framework. In encoder 100, the input video signal is processed block by block, and a block may be referred to as a codec unit (CU). In VTM-1.0, a CU can be up to 128 × 128 pixels. However, unlike HEVC, which is based solely on quadtree-based block partitioning, in VVC, a codec tree unit (CTU) is divided into multiple CUs based on a quadtree / binary tree / tritree to accommodate varying local characteristics. By definition, a codec tree block (CTB) is an N × N sample block for a certain value of N, such that dividing the components into CTBs is a partitioning operation. A CTU includes the CTB of luminance samples of an image with a three-sample array, two corresponding CTBs of chrominance samples, or the CTB of samples of a monochrome image or an image encoded using three separate color planes and a syntax structure for encoding and decoding samples. In addition, the concept of multiple segmentation unit types in HEVC has been removed; that is, the distinction between CU, prediction unit (PU), and transformation unit (TU) no longer exists in VVC. Instead, each CU is always used as the basic unit for both prediction and transformation without further segmentation.
[0042] In a multi-type tree structure, a CTU is first partitioned using a quadtree structure. Then, each quadtree leaf node can be further partitioned using binary and ternary tree structures. For example... Figures 2A to 2E As shown, there are five types of splitting, including quadruple partitioning ( Figure 2A ), horizontal binary segmentation ( Figure 2B Vertical binary segmentation Figure 2C ), horizontal ternary segmentation ( Figure 2D ) and vertical ternary segmentation ( Figure 2E ).
[0043] For each given video block, a prediction is formed based on either an inter-frame prediction method or an intra-frame prediction method. In inter-frame prediction, one or more predictors are formed based on pixels from previously reconstructed frames, through motion estimation and motion compensation. In intra-frame prediction, predictors are formed based on reconstructed pixels in the current frame. Through mode decision-making, the optimal predictor is selected to predict the current block.
[0044] The prediction residual, representing the difference between the current video block and its prediction factor, is sent to the transform circuit 102. The transform coefficients are then sent from the transform circuit 102 to the quantization circuit 104 for entropy reduction. The quantized coefficients are then fed to the entropy encoding / decoding circuit 106 to generate a compressed video bitstream. Figure 1 As shown, prediction-related information 110 (such as video block segmentation information, motion vectors, reference picture indexes, and intra-prediction modes) from the inter-frame prediction circuit and / or intra-frame prediction circuit 112 is also fed through the entropy encoding / decoding circuit 106 and stored in the compressed video bitstream 114.
[0045] In encoder 100, for prediction purposes, decoder-related circuitry is also required to reconstruct pixels. First, the prediction residual is reconstructed via inverse quantization 116 and inverse transform circuit 118. This reconstructed prediction residual is combined with block prediction factor 120 to generate unfiltered reconstructed pixels for the current video block.
[0046] Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video frame as the current video block to predict the current video block.
[0047] Timing prediction (also known as "inter-frame prediction") uses reconstructed pixels from already encoded and decoded video frames to predict the current video block. Timing prediction reduces the inherent temporal redundancy in the video signal. A timing prediction signal for a given codec unit (CU) or codec block is typically sent by signaling one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its timing reference. Additionally, if multiple reference frames are supported, a reference frame index is sent separately, which is used to identify which reference frame in the reference frame memory the timing prediction signal originates from.
[0048] After performing spatial and / or temporal prediction, the intra / inter-frame mode decision circuit 121 in encoder 100 selects the optimal prediction mode, for example, based on a rate-distortion optimization method. The block prediction factor 120 is then subtracted from the current video block; and the resulting prediction residual is decorrelated using transform circuit 102 and quantization circuit 104. The resulting quantized residual coefficients are dequantized by inverse quantization circuit 116 and inverse transformed by inverse transform circuit 118 to form the reconstruction residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Before the reconstructed CU is placed in the reference picture memory of picture buffer 117 and used for encoding and decoding subsequent video blocks, the reconstructed CU may be further applied with loop filtering 115, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive loop filter (ALF). To form the output video bitstream 114, the encoding / decoding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantized residual coefficients are all sent to entropy encoding / decoding unit 106 for further compression and packing to form the bitstream.
[0049] For example, deblocking filters are available in current versions of AVC, HEVC, and VVC. In HEVC, an additional loop filter called Sample Adaptive Offset (SAO) is defined to further improve encoding and decoding efficiency. In the current version of the VVC standard, another loop filter called Adaptive Loop Filter (ALF) is under active investigation and is very likely to be included in the final standard.
[0050] These loop filter operations are optional. Performing these operations helps improve encoding / decoding efficiency and visual quality. They can also be turned off based on decisions made by encoder 100 to save computational complexity.
[0051] It should be noted that intra-frame prediction is typically based on unfiltered reconstructed pixels, while inter-frame prediction is based on filtered reconstructed pixels (if these filter options are enabled in encoder 100).
[0052] Figure 3 This is a block diagram illustrating an exemplary block-based video decoder 200 that can be used in conjunction with many video codec standards. The decoder 200 is similar to [the one residing in...]. Figure 1The reconstruction-related part is located in the encoder 100. In the decoder 200, the input video bitstream 201 is first decoded by entropy decoding 202 to derive the quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed by inverse quantization 204 and inverse transform 206 to obtain the reconstructed prediction residuals. The block predictor mechanism implemented in the intra / inter-frame mode selector 212 is configured to perform intra-frame prediction 208 or motion compensation 210 based on the decoded prediction information. The reconstructed prediction residuals from the inverse transform 206 and the prediction output generated by the block predictor mechanism are summed using a summer 214 to obtain a set of unfiltered reconstructed pixels.
[0053] Before the reconstructed blocks are stored in the image buffer 213, which serves as a reference image memory, the reconstructed blocks can be further passed through the loop filter 209. The reconstructed video in the image buffer 213 can be sent to drive the display device and to predict subsequent video blocks. With the loop filter 209 open, filtering operations are performed on these reconstructed pixels to produce the final reconstructed video output 222.
[0054] Typically, the basic intra prediction scheme used in VVC is the same as that in HEVC, except that several modules are further extended and / or improved, such as intra-segmentation (ISP) encoding / decoding mode, extended intra prediction with wide-angle intra-direction, position-dependent intra prediction combination (PDPC), and 4-tap frame interpolation.
[0055] Segmentation of images, tile groups, tiles, and CTUs in VVC
[0056] In VVC, a tile is defined as a rectangular region of a CTU within a specific tile column and a specific tile row in an image. A tile group is a combination of an integer number of tiles in an image that are exclusively contained within a single NAL unit. Essentially, the concept of a tile group is the same as that of a stripe defined in HEVC. For example, an image is divided into tile groups and tiles.
[0057] A tile is a CTU sequence that covers a rectangular area of an image. A tile group contains multiple tiles of an image. Two tile group modes are supported: raster scan tile group mode and rectangular tile group mode. In raster scan tile group mode, a tile group contains a sequence of tiles raster scanned into the image. In rectangular tile group mode, a tile group contains multiple tiles that together form a rectangular area of the image. The tiles within a rectangular tile group are arranged in the order of the tile raster scan of the tile group.
[0058] Figure 4 An example of raster scan tile group segmentation of an image is shown, where the image is divided into 12 tiles and 3 raster scan tile groups.
[0059] Figure 5This example demonstrates the segmentation of an image into rectangular tile groups, where the image is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular tile groups.
[0060] Large block size transformation using high-frequency zeroing in VVC
[0061] In VTM4, large block size transforms up to 64×64 are enabled, primarily for higher resolution video such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both) equal to 64, high-frequency transform coefficients are zeroed out, leaving only low-frequency coefficients. For example, for an M×N transform block, where M is the block width and N is the block height, when M equals 64, only the left 32 columns of transform coefficients are retained. Similarly, when N equals 64, only the top 32 rows of transform coefficients are retained. When transform skip mode is applied to large blocks, the entire block is used without zeroing out any values.
[0062] Virtual Pipeline Data Unit (VPDU) in VVC
[0063] Virtual Pipeline Data Units (VPDUs) are defined as non-overlapping units in an image. In a hardware decoder, consecutive VPDUs are processed simultaneously through multiple pipeline stations. The VPDU size is roughly proportional to the buffer size in most pipeline stations, so keeping the VPDU size small is important. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, ternary tree (TT) and binary tree (BT) partitioning can lead to an increase in VPDU size.
[0064] To maintain the VPDU size at 64×64 luminance samples, the following standard segmentation constraint is applied in VTM5 (using syntax signaling modification), such as... Figures 6A to 6H As shown. For convenience, we will label the examples above from left to right as follows. Figures 6A to 6D Examples are shown in the image, and the examples below are labeled from left to right. Figures 6E to 6H Examples are shown in the text.
[0065] TT splitting is not allowed for CUs with a width or height equal to 128, or both width and height equal to 128. Figure 6A , Figure 6B , Figure 6E , Figure 6F , Figure 6G and Figure 6H ).
[0066] For a 128×N CU with N≤128 (i.e., width equal to 128 and height less than or equal to 128), horizontal BT is not allowed. Figure 6D ).
[0067] For an N×128CU with N≤128 (i.e., height equal to 128 and width less than or equal to 128), vertical BT is not allowed. Figure 6C ).
[0068] Transform coefficient encoding and decoding in VVC
[0069] Transform coefficient encoding / decoding refers to the encoding / decoding process of the transform coefficient quantization level values of the TU. In HEVC, the transform coefficients of the codec block are encoded / decoded using non-overlapping coefficient groups (or sub-blocks), and each CG contains coefficients from a 4×4 block of the codec block. CGs within a codec block and transform coefficients within a CG are encoded / decoded according to a predefined scan order. Encoding / decoding of the transform coefficient level of a CG with at least one non-zero transform coefficient can be divided into multiple scan channels. In the first channel, the first binary bit (represented by bin0, also known as significant_coeff_flag, which indicates that the magnitude of the coefficient is greater than 0) is encoded / decoded. Next, two scan channels are applied for context encoding / decoding of the second / third binary bits (represented by bin1 and bin2, also known as coeff_abs_greater1_flag and coeff_abs_greater2_flag, respectively). Finally, if needed, two additional scan channels are invoked for encoding / decoding the symbol information and residual values of the coefficient level (also known as coeff_abs_level_remaining). Note that only the bits in the first three scan channels are encoded and decoded in normal mode, and these bits are referred to as normal bits in the following description.
[0070] In VVC3, for each subblock, the normally encoded and decoded bits and the bypass encoded and decoded bits are separated in the encoding / decoding order; first, all the normally encoded and decoded bits for the subblock are sent, and then the bypass encoded and decoded bits are sent. The transform coefficient levels of the subblock are encoded and decoded in four channels by scanning positions, as follows:
[0071] —Channel 1: Encoding and decoding of validity (sig_flag), greater than 1 flag (gt1_flag), parity (par_level_flag), and greater than 2 flag (gt2_flag) are processed in the order of encoding and decoding. If sig_flag equals 1, then gt1_flag (which specifies whether the absolute level is greater than 1) is encoded and decoded first. If gt1_flag equals 1, then par_level_flag (which specifies the parity of the absolute level minus 2) is encoded and decoded separately.
[0072] —Channel 2: For all scan positions where gt2_flag equals 1 or gt1_flag equals 1, the remaining absolute level (remaining portion) is processed for encoding and decoding. Non-binary syntax elements are binarized using Columbus-Rice code, and the resulting binary bits are encoded and decoded in the bypass mode of the arithmetic encoding / decoding engine.
[0073] —Channel 3: The absolute level (absLevel) of the coefficients that were not encoded or decoded in the first channel (due to reaching the limit of binary bits after normal encoding and decoding) is fully encoded and decoded using Columbus-Rice code in the bypass mode of the arithmetic codec engine.
[0074] —Channel 4: For all scan positions where sig_coeff_flag equals 1, the encoding and decoding of the sign (sign_flag) are processed.
[0075] For 4×4 sub-blocks, no more than 32 normally encoded / decoded bits (sig_flag, par_flag, gt1_flag, and gt2_flag) are guaranteed to be encoded or decoded. For 2×2 chroma sub-blocks, the number of normally encoded / decoded bits is limited to 8.
[0076] Similar to HEVC, the Rice parameter (ricePar) used for encoding and decoding the remainder of non-binary syntax elements (in channel 3) is derived. At the beginning of each sub-block, ricePar is set to 0. After encoding and decoding the remainder of the syntax elements, the Rice parameter is modified according to a predefined equation. To encode and decode the non-binary syntax element absLevel (in channel 4), the sum of the absolute values sumAbs in the local template is determined. The variables ricePar and posZero are determined through a table lookup based on the relevant quantization and sumAbs. The intermediate variable codeValue is derived as follows:
[0077] —If absLevel[k] equals 0, then codeValue is set to equal posZero;
[0078] —Otherwise, if absLevel[k] is less than or equal to posZero, then codeValue is set to equal to absLevel[k]-1;
[0079] —Otherwise (absLevel[k] is greater than posZero), codeValue is set to equal absLevel[k].
[0080] The value of codeValue is encoded or decoded using Columbus-Rice code with the Rice parameter ricePar.
[0081] In the following description of this disclosure, transform coefficient encoding / decoding is also referred to as residual encoding / decoding.
[0082] Decoder-Side Motion Vector Refinement (DMVR) in VVC
[0083] Decoder-side motion vector refinement (DMVR) is a technique for blocks encoded and decoded in bidirectional prediction merging mode and controlled by the SPS signaling flag sps_dmvr_enabled_flag. In this mode, the two motion vectors (MVs) of the block can be further refined using bilateral matching (BM) prediction.
[0084] Figure 3A This is a schematic diagram illustrating an example of decoder-side motion vector refinement (DMVR). Figure 3A As shown, the bilateral matching method is used to refine the motion information of the current CU 322 by searching for the closest match between two reference blocks 302 and 312 of the current CU 322 along the motion trajectory of the current CU 322 in the current image 320 (i.e., refPic in list L0300 and refPic in list L1310). Patterned rectangular blocks 322, 302, and 312 indicate the current CU and its two reference blocks based on the initial motion information from the merging pattern. Patterned rectangular blocks 304 and 314 indicate a pair of reference blocks based on the MV candidates used in the motion refinement search process (i.e., the motion vector refinement process).
[0085] The difference between the candidate MV and the initial MV (also known as the original MV) is MV. diff and-MV diff Both the candidate MV and the initial MV are bidirectional motion vectors. During DMVR, multiple such MV candidates around the initial MV can be examined. Specifically, for each given MV candidate, its two associated reference blocks can be located from its reference images in List 0 and List 1, respectively, and the difference between them is calculated. This block difference is typically measured as SAD (or sum of absolute differences) or row subsampling SAD (i.e., SAD calculated using blocks included in every other row). Finally, the MV candidate with the lowest SAD between its two reference blocks becomes the refined MV and is used to generate the bidirectional prediction signal as the actual prediction for the current CU.
[0086] In VVC, DMVR is applied to CUs that meet the following conditions:
[0087] • Encoded and decoded using a CU-level merging mode (not a sub-block merging mode) with bidirectional predictive MV;
[0088] • Relative to the current image, one reference image of the CU is in the past (i.e., has a POC smaller than the current image's POC) and another reference image is in the future (i.e., has a POC larger than the current image's POC);
[0089] • The POC distance (i.e., absolute POC difference) from the two reference images to the current image is the same;
[0090] • The CU has a size of more than 64 luminance samples and a height of more than 8 luminance samples.
[0091] The refined MV derived from the DMVR process is used to generate inter-frame prediction samples and also for temporal motion vector prediction for future image encoding / decoding. However, the original MV is used in the deblocking process and also for spatial motion vector prediction for future CU encoding / decoding. Some additional features of DMVR are shown in the following sub-clauses.
[0092] Bidirectional optical flow (BDOF) in VVC
[0093] The Bidirectional Optical Flow (BDOF) tool is included in VTM5. The BDOF previously known as BIO is now included in JEM. Compared to the JEM version, the BDOF in VTM5 is a simpler version, requiring fewer computations, particularly in terms of the number of multiplications and the size of the multipliers. BDOF is controlled by the SPS flag `sps_bdof_enabled_flag`.
[0094] BDOF is used to refine the bidirectional prediction signal of the CU at the 4×4 sub-block level. BDOF is applied to the CU if the following conditions are met: 1) the CU height is not 4 and the CU size is not 4×8; 2) the CU is not encoded / decoded using affine mode or ATMVP merging mode; 3) the CU is encoded / decoded using a “true” bidirectional prediction mode, i.e., one of the two reference images is displayed before the current image in the order of display, and the other image is displayed after the current image in the order of display. BDOF is applied only to the luma component.
[0095] As its name suggests, the BDOF mode is based on the concept of optical flow, which assumes that the motion of objects is smooth. BDOF adjusts the prediction samples by calculating the gradient of the current block to improve encoding and decoding efficiency.
[0096] Decoder-side control for DMVR and BDOF in VVC
[0097] In the current VVC, BDOF / DMVR is always applied if the SPS flag of BDOF / DMVR is enabled and some bidirectional prediction and size constraints are met for regular merge candidates.
[0098] DMVR is applied in regular merge mode when all of the following conditions are true:
[0099] —sps_dmvr_enabled_flag equals 1
[0100] —general_merge_flag[xCb][yCb] equals 1
[0101] Both —predFlag10[0][0] and predFlagL1[0][0] are equal to 1.
[0102] —mmvd_merge_flag[xCb][yCb] equals 0
[0103] —DiffPicOrderCnt(currPic,RefPicList[0][RefIdx10]) is equal to DiffPicOrderCnt(RefPicList[1][refIdxL1],currPic)
[0104] —BcwIdx[xCb][yCb] equals 0
[0105] — Both luma_weight_l0_flag[refidx10] and luma_weight_l1_flag[refIdxL1] are equal to 0.
[0106] —cbWidth is greater than or equal to 8
[0107] —cbHeight is greater than or equal to 8
[0108] —cbHeight×cbWidth is greater than or equal to 128
[0109] BDOF is applied to bidirectional forecasting when all of the following conditions are true:
[0110] —sps_bdof_enabled_flag equals 1.
[0111] —predFlag10[xSbIdx][ySbIdx] and predFlag11[xSbIdx][ySbIdx] are both equal to 1.
[0112] —DiffPicOrderCnt(currPic,RefPicList[0][RefIdx10])×DiffPicOrderCnt(currPic,RefPicList[1][RefIdx1]) is less than 0.
[0113] —MotionModelIdc[xCb][yCb] equals 0.
[0114] —merge_subblock_flag[xCb][yCb] equals 0.
[0115] —sym_mvd_flag[xCb][yCb] equals 0.
[0116] —BcwIdx[xCb][yCb] equals 0.
[0117] —luma_weight_l0_flag[refidx10] and luma_weight_l1_flag[refIdxL1] are both equal to 0.
[0118] —cbHeight is greater than or equal to 8
[0119] —cIdx equals 0.
[0120] Residual encoding and decoding for transform skip mode CU in VVC
[0121] VTM5 allows transform skip mode for luma blocks up to 32×32 (including 32×32). When a CU is encoded and decoded in transform skip mode, its prediction residuals are quantized and encoded using a transform skip residual encoding / decoding process. This residual encoding / decoding process is a modification of the transform coefficient encoding / decoding process described in the previous section. In transform skip mode, the CU's residuals are also encoded and decoded in units of 4×4 non-overlapping subblocks. Unlike the regular transform coefficient encoding / decoding process, in transform skip mode, the last coefficient position is not signaled; instead, the coded_subblock_flag is signaled for all 4×4 subblocks in the CU in a forward scan order (i.e., from the top-left subblock to the last subblock).
[0122] For each subblock, if coded_subblock_flag equals 1 (i.e., there is at least one non-zero quantization residual in the subblock), then encoding and decoding of the quantization residual level are performed in the three scan channels:
[0123] —First scan channel: Validity flag (sig_coeff_flag), sign flag (coeff_sign_flag), absolute level greater than 1 flag (abs_level_gtx_flag[0]), and parity flag (par_level_flag) are encoded and decoded. For a given scan position, if coeff_sig_flag equals 1, then coeff_sign_flag is encoded and decoded, followed by abs_level_gtx_flag[0] (which specifies whether the absolute level is greater than 1) being encoded and decoded. If abs_level_gtx_flag[0] equals 1, then par_level_flag is additionally encoded and decoded to specify the parity of the absolute level.
[0124] —Greater than x scan channel: For each scan position where the absolute level is greater than 1, up to four abs_level_gtx_flag[i] (where i = 1...4) are encoded to indicate whether the absolute level at the given position is greater than 3, 5, 7 or 9 respectively.
[0125] — Remaining scan channels: For all scan positions where abs_level_gtx_flag[4] equals 1 (i.e., the absolute level is greater than 9), the remaining absolute level is encoded and decoded. The remaining absolute level is binarized using a simplified Rice parameter-derived template.
[0126] The bits in scan channels #1 and #2 (the first scan channel and scan channels greater than x) are context-coded until the maximum number of context-coded bits in the CU is exhausted. The maximum number of context-coded bits in the residual block is limited to 2 × block_width × block_height, or equivalently, an average of 2 context-coded bits per sample location. The bits in the last scan channel (the remaining scan channels) are bypassed and decoded.
[0127] Lossless encoding and decoding in HEVC
[0128] Lossless encoding / decoding modes in HEVC are achieved through simple bypass transformation, quantization, and loop filters (deblocking filter, sample adaptive offset, and adaptive loop filter). This design aims to achieve lossless encoding / decoding with minimal changes required for implementing conventional HEVC encoders and decoders for mainstream applications.
[0129] In HEVC, lossless codec mode can be enabled or disabled at individual CU levels. This is done via the cu_transquant_bypass_flag syntax, which is signaled at the CU level. To reduce signaling overhead in cases where lossless codec mode is unnecessary, the cu_transquant_bypass_flag syntax is not always signaled. It is only signaled when another syntax, called transquant_bypass_enabled_flag, has a value of 1. In other words, the transquant_bypass_enabled_flag syntax is used to enable the cu_transquant_bypass_flag syntax signaling.
[0130] In HEVC, the syntax `transquant_bypass_enabled_flag` is signaled in the Picture Parameter Set (PPS) to indicate whether the syntax `cu_transquant_bypass_flag` needs to be signaled for each CU within the picture referencing that PPS. If this flag is set to 1, the syntax `cu_transquant_bypass_flag` is sent at the CU level to signal whether the current CU is being encoded / decoded in lossless mode. If this flag is set to 0 in the PPS, `cu_transquant_bypass_flag` is not sent, and all CUs in the picture are encoded / decoded using the transforms, quantizations, and loop filters involved in the process, which typically results in some degree of video quality degradation. To losslessly encode / decode the entire picture, for each CU in the picture, the flag `transquant_bypass_enabled_flag` in the PPS must be set to 1, and the CU-level flag `cu_transquant_bypass_flag` must be set to 1. The detailed syntax signaling related to lossless mode in HEVC is shown below.
[0131] A `transquant_bypass_enabled_flag` value of 1 indicates that `cu_transquant_bypass_flag` exists. A `transquant_bypass_enabled_flag` value of 0 indicates that `cu_transquant_bypass_flag` does not exist.
[0132] A cu_transquant_bypass_flag value of 1 indicates that scaling and transformation procedures as specified in Clause 8.6 and loop filter procedures as specified in Clause 8.7 are bypassed. When cu_transquant_bypass_flag does not exist, it is inferred to be equal to 0.
[0133]
[0134]
[0135]
[0136]
[0137]
[0138] In VVC, the maximum CU size is 64×64, and the VPDU is also set to 64×64. Due to the coefficient zeroing mechanism for widths / heights greater than 32, the maximum block size for coefficient encoding / decoding in VVC is 32×32. Under this constraint, the current transform skips supporting only CUs up to 32×32, allowing the maximum block size for residual encoding / decoding to align with the maximum block size of 32×32 for coefficient encoding / decoding. However, in VVC, the constraint on the block size for residual encoding / decoding of lossless CUs is not defined. Therefore, currently in VVC, residual blocks larger than 32×32 can be generated in lossless encoding / decoding mode, which would require support for residual encoding / decoding of blocks larger than 32×32. This is not preferred for codec implementation. Several methods are proposed in this disclosure to address this problem.
[0139] Another issue associated with lossless codec support in VVC is how to select the residual (or coefficient) codec scheme. In current VVC, two different residual codec schemes are available. For a given block (or CU), the selection of the residual codec scheme is based on the transform skip flag of that given block (or CU). Therefore, if in lossless mode, assuming the transform skip flag is 1 in VVC, as in HEVC, the residual codec scheme used in transform skip mode will always be used for the lossless mode CU. However, the current residual codec scheme used when the transform skip flag is true is primarily designed for screen content encoding and decoding. Using the current residual codec scheme for lossless encoding and decoding of regular content (i.e., non-screen content) may not be optimal. In this disclosure, several methods are proposed for selecting the residual codec for lossless CUs.
[0140] In current VVC, two decoder-side tools (BDOF and DMVR) refine the decoded pixels by filtering the current block, thereby improving encoding / decoding performance. However, in lossless encoding / decoding, since the predicted pixels are already perfectly predicted, BDOF and DMVR do not contribute to the encoding / decoding gain. Therefore, BDOF and DMVR should not be applied to lossless encoding / decoding because these decoder-side tools offer no benefit to VVC. However, in current VVC, BDOF and DMVR are always applied if the SPS flag of BDOF and DMVR is enabled and some bidirectional prediction and size constraints are met for regular merge candidates. Therefore, for lossless VVC encoding / decoding, controlling DMVR and BDOF at lower levels (i.e., stripe level and CU level) is beneficial to the performance efficiency of lossless VVC encoding / decoding.
[0141] Residual block partitioning for lossless CU
[0142] According to examples in this disclosure, it is proposed to align the maximum residual codec block size for a lossless CU with the maximum block size supported by a transform skip mode. In one example, the transform skip mode may be enabled only for residual blocks whose width and height are both less than or equal to 32, indicating that the maximum residual codec block size under transform skip mode is 32×32. According to the example, the maximum width and / or height of the residual block for the lossless CU is also set to 32, resulting in a maximum residual block size of 32×32. Whenever the width / height of the lossless CU is greater than 32, the CU residual block is divided into multiple smaller residual blocks of size 32×N and / or N×32, such that the width or height of the smaller residual blocks is no greater than 32. For example, a 128×32 lossless CU is divided into four 32×32 residual blocks for residual codec. In another example, a 64×64 lossless CU is divided into four 32×32 residual blocks.
[0143] According to another example of this disclosure, it is proposed to align the maximum block size for residual encoding / decoding of a lossless CU with the size of the VPDU. In one example, the width / height of the maximum residual block for the lossless CU is set to the VPDU size (e.g., 64×64 in a current VVC). Whenever the width / height of the lossless CU is greater than 64, the CU residual block is divided into multiple smaller residual blocks of size 64×N and / or N×64, such that the width or height of the smaller residual blocks is no greater than the width and / or height of the VPDU. For example, a 128×128 lossless CU is divided into four 64×64 residual blocks for residual encoding / decoding. In another example, a 128×32 lossless CU is divided into two 64×32 residual blocks.
[0144] Selection of residual encoding / decoding scheme for lossless CU
[0145] In the current VVC, depending on whether the CU is encoded and decoded in transform skip mode, the CU utilizes different residual encoding and decoding schemes. The current residual encoding and decoding used in transform skip mode is generally more suitable for screen content encoding and decoding.
[0146] According to the examples in this disclosure, the lossless CU uses the same residual encoding / decoding scheme as the residual encoding / decoding scheme used by the transform skip mode CU.
[0147] According to another example of this disclosure, the lossless CU uses the same residual encoding / decoding scheme as the non-transform skip mode CU.
[0148] According to another example of this disclosure, a residual codec scheme for the lossless CU is adaptively selected from existing residual codec schemes based on certain conditions and / or predefined procedures. Both the encoder and decoder follow such conditions and / or predefined procedures, so that no signaling is required in the bitstream to indicate the selection. In one example, a simple screen content detection scheme can be specified and utilized in both the encoder and decoder. Based on the detection scheme, the current video block can be classified as screen content or regular content. If it is screen content, the residual codec scheme used in transform skip mode is selected. Otherwise, another residual codec scheme is selected.
[0149] According to another example of this disclosure, a syntax is signaled in the bitstream to explicitly specify which residual codec scheme is used by the lossless CU. Such a syntax can be binary flags, each binary value indicating the selection of one of two residual codec schemes. The syntax can be signaled at different levels. For example, it can be signaled at the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), strip header, tile group header, or tile level. It can also be signaled at the CTU or CU level. When this syntax is signaled, all lossless CUs at the same or lower levels will use the same residual codec scheme indicated by that syntax. For example, when the syntax is signaled at the SPS level, all lossless CUs in the sequence will use the same indicated residual codec scheme. When the syntax is signaled at the PPS level, all lossless CUs in the picture will use the same residual codec scheme indicated in the associated PPS. If a syntax (e.g., cu_transquant_bypass_flag) exists at the CU level to indicate whether a CU is encoded or decoded in lossless mode, the syntax for indicating a residual codec scheme is conditionally signaled based on the CU's lossless mode flag. For example, the syntax for indicating a residual codec scheme is signaled for a CU only if the lossless mode flag cu_transquant_bypass_flag indicates that the current CU is encoded or decoded in lossless mode. When the syntax is signaled via a flag at the stripe header level, all CUs encoded or decoded in lossless mode within this stripe will use the same residual codec scheme identified by the signaled flag. The residual codec scheme for each of the multiple CUs is selected based on a first signaled flag, wherein, according to the flag signaled at the stripe header, the residual codec scheme selected for the lossless CU is the residual codec scheme used by either a transform-skip mode CU or a non-transform-skip mode CU.
[0150] According to the examples in this disclosure, even for CUs encoded in lossless mode, a transform skip mode flag is signaled. In this case, regardless of whether the CU is encoded in lossless mode, the choice of the residual encoding / decoding scheme for the CU is based on its transform_skip_mode_flag.
[0151] Disable DMVR
[0152] In current VVC, there is no defined control for DMVR on / off for lossless codec modes. In the examples of this disclosure, it is proposed to control DMVR on / off at the slice level using a 1-bit signaling flag, `slice_disable_dmvr_flag`. In one example, if `sps_dmvr_enabled_flag` is set to 1 and `transquant_bypass_enabled_flag` is set to 0, then the `slice_disable_dmvr_flag` flag needs to be signaled. If the `slice_disable_dmvr_flag` flag is not signaled, it is inferred to be 1. If `slice_disable_dmvr_flag` equals 1, then the DMVR is off. In this case, the signaling is as follows:
[0153] if(sps_dmvr_enabled_flag&&!transquant_bypass_enabled_flag) slice_disable_dmvr_flag u(1)
[0154] In another example, the CU-level control of the DMVR's on / off state is proposed using the `cu_transquant_bypass_flag`. In one example, CU-level control of the DMVR follows these steps:
[0155] DMVR is applied in regular merge mode when all of the following conditions are true:
[0156] —sps_dmvr_enabled_flag equals 1
[0157] —cu_transquant_bypass_flag is set to 0
[0158] —general_merge_flag[xCb][yCb] equals 1
[0159] Both predFlagL0[0][0] and predFlagL1[0][0] are equal to 1.
[0160] —mmvd_merge_flag[xCb][yCb] equals 0
[0161] —DiffPicOrderCnt(currPic,RefPicList[0][RefIdxL0]) is equal to DiffPicOrderCnt(RefPicList[1][refIdxL1],currPic)
[0162] —BcwIdx[xCb][yCb] equals 0
[0163] —luma_weight_l0_flag[refidxL0] and luma_weight_l1_flag[refIdxL1] are both equal to 0.
[0164] —cbWidth is greater than or equal to 8
[0165] —cbHeight is greater than or equal to 8
[0166] —cbHeight×cbWidth is greater than or equal to 128
[0167] Disable BDOF
[0168] In the current VVC, there is no defined control for enabling / disabling BDOF for lossless codec modes. In the examples of this disclosure, it is proposed to control BDOF enabling / disabling via a 1-bit signaling flag, `slice_disable_bdof_flag`. In one example, if `sps_bdof_enabled_flag` is set to 1 or `transquant_bypass_enabled_flag` is set to 0, the `slice_disable_bdof_flag` flag is signaled. If the `slice_disable_bdof_flag` flag is not signaled, it is inferred to be 1. If the `slice_disable_bdof_flag` flag is equal to 1, BDOF is disabled. In this case, the signaling is shown as follows:
[0169] if(sps_bdof_enabled_flag&&!transquant_bypass_enabled_flag) slice_disable_bdof_flag u(1)
[0170] In another example of this disclosure, it is proposed to control the on / off state of BDOF at the CU level via cu_transquant_bypass_flag. In one example, the CU-level control of BDOF follows the following:
[0171] BDOF is applied to the regular merge pattern when all of the following conditions are true:
[0172] —sps_bdof_enabled_flag equals 1.
[0173] —cu_transquant_bypass_flag is set to 0
[0174] —predFlagL0[xSbIdx][ySbIdx] and predFlagL1[xSbIdx][ySbIdx] are both equal to 1.
[0175] —DiffPicOrderCnt(currPic,RefPicList[0][RefIdxL0])*DiffPicOrderCnt(currPic,RefPicList[1][RefIdx1]) is less than 0.
[0176] —MotionModelIdc[xCb][yCb] equals 0.
[0177] —merge_subblock_flag[xCb][yCb] equals 0.
[0178] —sym_mvd_flag[xCb][yCb] equals 0.
[0179] —BcwIdx[xCb][yCb] equals 0.
[0180] —luma_weight_l0_flag[refidxL0] and luma_weight_l1_flag[refIdxL1] are both equal to 0.
[0181] —cbHeight is greater than or equal to 8
[0182] —cIdx equals 0.
[0183] Disable BDOF and DMVR
[0184] In current VVC, both BDOF and DMVR are consistently used for decoder-side refinement to improve encoding / decoding efficiency and are controlled by each SPS flag, satisfying certain bidirectional prediction and size constraints for regular merging candidates. In the examples of this disclosure, it is proposed to disable both BDOF and DMVR via a 1-bit slice_disable_bdof_dmvr_flag flag. If the slice_disable_bdof_dmvr_flag flag is set to 1, both BDOF and DMVR are disabled. If the slice_disable_bdof_dmvr_flag flag is not signaled, it is inferred to be 1. In one example, slice_disable_bdof_dmvr_flag is signaled if the following condition is met.
[0185] if((sps_bdof_enabled_flag||sps_dmvr_enabled_flag)&& ! transquant_bypass_enabled_flag) slice_disable_bdof_dmvr_flag u(1)
[0186] The methods described above can be implemented using an apparatus comprising one or more circuits, including application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The apparatus can also perform the methods using circuits combined with other hardware or software components. Each module, submodule, unit, or subunit disclosed above can be implemented at least partially using one or more circuits.
[0187] Figure 7 This is a block diagram illustrating an apparatus for video encoding and decoding according to some embodiments of the present disclosure. Apparatus 700 may be a terminal, such as a mobile phone, tablet computer, digital broadcasting terminal, tablet device, or personal digital assistant.
[0188] like Figure 7 As shown, device 700 may include one or more of the following components: processing component 702, memory 704, power supply component 706, multimedia component 708, audio component 710, input / output (I / O) interface 712, sensor component 714, and communication component 716.
[0189] Processing component 702 typically controls the overall operation of device 700, such as operations related to display, telephone calls, data communication, camera operation, and recording. Processing component 702 may include one or more processors 720 for executing instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 702 may include one or more modules for facilitating interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module for facilitating interaction between multimedia component 708 and processing component 702.
[0190] Memory 704 is configured to store different types of data to support the operation of device 700. Examples of such data include instructions for any application or method operating on device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, and memory 704 may be static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0191] Power supply assembly 706 supplies power to various components of device 700. Power supply assembly 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 700.
[0192] Multimedia component 708 includes a screen that provides an output interface between device 700 and the user. In some examples, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen that receives input signals from the user. The touch panel may include one or more touch sensors for sensing touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some examples, multimedia component 708 may include a front-facing camera and / or a rear-facing camera. When device 700 is in an operating mode such as shooting mode or video mode, the front-facing camera and / or rear-facing camera can receive external multimedia data.
[0193] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC). When device 700 is in an operating mode (such as call mode, recording mode, and voice recognition mode), the microphone is configured to receive external audio signals. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some examples, audio component 710 also includes a speaker for outputting audio signals.
[0194] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules. These peripheral interface modules can be keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0195] Sensor assembly 714 includes one or more sensors for providing state assessment in various aspects of device 700. For example, sensor assembly 714 may detect the on / off state of device 700 and the relative position of components. Components, for example, are the display and keyboard of device 700. Sensor assembly 714 may also detect changes in position of device 700 or its components, the presence or absence of user contact on device 700, the orientation or acceleration / deceleration of device 700, and temperature changes of device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical touch. Sensor assembly 714 may also include optical sensors, such as CMOS or CCD image sensors used in imaging applications. In some examples, sensor assembly 714 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0196] Communication component 716 is configured to facilitate wired or wireless communication between device 700 and other devices. Device 700 may access a wireless network based on communication standards such as WiFi, 4G, or combinations thereof. In the example, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In the example, communication component 716 may also include a near-field communication (NFC) module for facilitating short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0197] In the example, device 700 may be implemented by one or more of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field-programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components to perform the above-described method.
[0198] Non-transitory computer-readable storage media can be, for example, hard disk drives (HDDs), solid-state drives (SSDs), flash memory, hybrid drives or solid-state hybrid drives (SSHDs), read-only memory (ROMs), optical disc read-only memory (CD-ROMs), magnetic tapes, floppy disks, etc.
[0199] Figure 8 This is a flowchart illustrating an exemplary process of a technique related to lossless encoding / decoding modes in video encoding / decoding according to some embodiments of the present disclosure.
[0200] In step 801, the processor 720 divides the video image into multiple codec units (CUs), at least one of the CUs being a lossless CU.
[0201] In step 802, the processor 720 determines the residual codec block size of the lossless CU.
[0202] In step 803, in response to determining that the residual codec block size of the lossless CU is greater than a predefined maximum value, the processor 720 divides the residual codec block into two or more residual blocks for residual codec.
[0203] In some examples, an apparatus for video encoding and decoding is provided. The apparatus includes a processor 720; and a memory 704 configured to store instructions executable by the processor; the processor, when executing the instructions, is configured to perform actions such as... Figure 8 The method shown.
[0204] In some other examples, a non-transitory computer-readable storage medium 704 is provided, having instructions stored therein. When the instructions are executed by processor 720, the instructions cause the processor to perform actions such as Figure 8 The method shown.
[0205] Figure 9 This is a flowchart illustrating an exemplary process of a technique related to lossless encoding / decoding modes in video encoding / decoding according to some embodiments of the present disclosure.
[0206] In step 901, the processor 720 divides the video image into multiple codec units (CUs), at least one of the CUs being a lossless CU.
[0207] In step 902, the processor 720 selects a residual encoding / decoding scheme for the lossless CU, wherein the residual encoding / decoding scheme selected for the lossless CU is the same as the residual encoding / decoding scheme used by the non-transform skip mode CU.
[0208] In some examples, an apparatus for video encoding and decoding is provided. The apparatus includes a processor 720; and a memory 704 configured to store instructions executable by the processor; the processor, when executing the instructions, is configured to perform actions such as... Figure 9 The method shown.
[0209] In some other examples, a non-transitory computer-readable storage medium 704 is provided, having instructions stored therein. When the instructions are executed by processor 720, the instructions cause the processor to perform actions such as Figure 9 The method shown.
[0210] The description in this disclosure has been presented for illustrative purposes and is not intended to be exhaustive or limited thereto. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.
[0211] The examples were chosen and described to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure and, best of all, to utilize the basic principles and the various embodiments with modifications suitable for the intended particular purpose. Therefore, it will be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of this disclosure.
Claims
1. A method for video encoding, comprising: Determine if the current block is in lossless mode; In response to determining that the current block is in the lossless mode, it is determined whether to apply a first residual codec scheme or a second residual codec scheme to the current block, wherein the first residual codec scheme is a residual codec scheme predefined for non-transform skip blocks, and the second residual codec scheme is a residual codec scheme predefined for transform skip blocks; A first indication and a second indication are sent by signaling, wherein the first indication indicates whether the current block is in the lossless mode, and the second indication indicates whether the first residual codec scheme or the second residual codec scheme is applied to the current block in the lossless mode.
2. The method for video encoding according to claim 1, further comprising: A 1-bit flag is sent for the current block using a signal. This 1-bit flag is used to control the opening and closing of the decoder-side motion vector refinement (DMVR) at the strip or picture level.
3. The method for video encoding according to claim 1, further comprising: A 1-bit flag is sent for the current block using a signal. This 1-bit flag is used to control the opening and closing of the decoder-side motion vector refinement DMVR at the CU level.
4. The method for video encoding according to claim 1, further comprising: A 1-bit flag is sent for the current block signal, and the 1-bit flag is used to control the opening and closing of the bidirectional optical flow (BDOF) at the strip level or picture level.
5. The method for video encoding according to claim 1, further comprising: A 1-bit flag is sent for the current block, and the 1-bit flag is used to control the opening and closing of the bidirectional optical flow (BDOF) at the CU level.
6. The method for video encoding according to claim 1, further comprising: A 1-bit flag is sent for the current block using a signal. This 1-bit flag is used to control the opening and closing of both the decoder-side motion vector refinement (DMVR) and bidirectional optical flow (BDOF) at the strip or picture level.
7. An apparatus for video encoding, comprising: One or more processors; as well as The memory is configured to store instructions that can be executed by the one or more processors; The one or more processors wherein, when executing the instructions, are configured to perform the method for video encoding as described in any one of claims 1 to 6.
8. A non-transitory computer-readable storage medium storing a plurality of programs executed by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform the video encoding method of any one of claims 1 to 6 and store a bitstream generated by the video encoding method of any one of claims 1 to 6.