Method and system for inter prediction compensation
Patent Information
- Application Number
- CN202280007448.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-02-18
- Filing Date
- 2022-02-21
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-02-21
Smart Images

Figure CN116569552B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application is based on and claims priority to U.S. Provisional Patent Application No. 63 / 151,786, entitled "SYSTEMS AND METHODS FOR INTERPREDICTION COMPENSATION", filed February 21, 2021, and U.S. Provisional Patent Application No. 63 / 160,774, entitled "SYSTEMS AND METHODS FOR INTERPREDICTION COMPENSATION", filed March 13, 2021, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of video processing technology, specifically to inter-frame prediction compensation methods and systems. Background Technology
[0004] Video is a set of still images (or "still frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. Currently, there are various video coding formats that use standardized video coding techniques, the most common being those based on prediction, transform, quantization, entropy coding, and loop filtering. Video coding standards that specify particular video coding formats are developed by standardization organizations; for example, the High Efficiency Video Coding (HEVC / H.265) standard, the Universal Video Coding (VVC / H.266) standard, and the Audio and Video Standard (AVS). As more and more advanced video coding techniques are adopted by video standards, the coding efficiency of new video coding standards is also increasing. Summary of the Invention
[0005] Embodiments of this disclosure provide methods and apparatus for video processing. In some exemplary embodiments, a computer-based method for video processing is provided. The video processing method includes: determining whether to enable inter-frame predictor correction for a coded block; and when inter-frame predictor correction is enabled for a coded block, performing inter-frame predictor correction by: acquiring a plurality of prediction samples from the top and left boundaries of a prediction block corresponding to the coded block; acquiring a plurality of reconstruction samples from top-adjacent reconstruction samples and left-adjacent reconstruction samples of the coded block; and obtaining a corrected prediction block based on the plurality of prediction samples, the plurality of reconstruction samples, and the prediction block.
[0006] In some exemplary embodiments, an apparatus is provided. The apparatus includes: a memory configured to store instructions; and at least one processor configured to execute instructions to cause the apparatus to perform the following steps: determining whether to enable inter-frame predictor correction for a coded block; and when inter-frame predictor correction is enabled for a coded block, performing inter-frame predictor correction by: acquiring a plurality of prediction samples from the top boundary and left boundary of a prediction block corresponding to the coded block; acquiring a plurality of reconstruction samples from top adjacent reconstruction samples and left adjacent reconstruction samples of the coded block; and obtaining the corrected prediction block based on the plurality of prediction samples, the plurality of reconstruction samples, and the prediction block.
[0007] In some exemplary embodiments, a non-volatile computer-readable storage medium is provided. In some embodiments, the non-volatile computer-readable storage medium stores an instruction set that can be executed by at least one processor of the device to cause the device to perform the video processing method described above. Attached Figure Description
[0008] Embodiments and corresponding aspects of this disclosure are illustrated in the following detailed description and accompanying drawings, in which various features are not drawn to scale.
[0009] Figure 1 A structural diagram of a video sequence according to some embodiments of the present disclosure is shown;
[0010] Figure 2 A flowchart of an encoder for a video encoding system according to some embodiments of the present disclosure is shown;
[0011] Figure 3 A flowchart of a decoder for a video encoding system according to some embodiments of the present disclosure is shown;
[0012] Figure 4 A structural block diagram of an apparatus for encoding or decoding video according to some embodiments of the present disclosure is shown;
[0013] Figure 5 A schematic diagram illustrating the derivation of a sub-block time motion vector predictor (TMVP) according to some embodiments of the present disclosure is shown;
[0014] Figure 6 A schematic diagram of adjacent blocks for spatial motion vector predictor (SMVP) derivation according to some embodiments of the present disclosure is shown;
[0015] Figure 7 A schematic diagram of a sub-block in a coding unit (CU) associated with a video frame in motion vector angle predictor (MVAP) processing, according to some embodiments of the present disclosure, is shown.
[0016] Figure 8A schematic diagram of a 4×4 reference block adjacent to the MVAP is shown according to some embodiments of the present disclosure;
[0017] Figure 9 A schematic diagram illustrating motion derivation in the final motion vector expression (UMVE) according to some embodiments of the present disclosure is shown;
[0018] Figure 10 A schematic diagram of an exemplary angle-weighted prediction (AWP) according to some embodiments of the present disclosure is shown;
[0019] Figure 11 A schematic diagram is shown of eight different prediction directions supported in AWP mode according to some embodiments of the present disclosure;
[0020] Figure 12 A schematic diagram of seven different weight arrays in an AWP mode according to some embodiments of the present disclosure is shown;
[0021] Figure 13 A schematic diagram of juxtaposed blocks and sub-blocks for candidate pruning is shown according to some embodiments of the present disclosure;
[0022] Figure 14A and Figure 14B A schematic diagram of two control point-based affine models according to some embodiments of the present disclosure is shown;
[0023] Figure 15 A schematic diagram showing the motion vector of the center sample of each sub-block according to some embodiments of the present disclosure is provided;
[0024] Figure 16 A schematic diagram of integer search points in decoder-side motion vector refinement (DMVR) according to some embodiments of the present disclosure is shown;
[0025] Figure 17 A schematic diagram is shown illustrating the estimation of local illumination compensation (LIC) model parameters using a reference image and a current image of adjacent blocks according to some embodiments of the present disclosure;
[0026] Figure 18 A schematic diagram illustrating LIC model parameter estimation according to some embodiments of the present disclosure is shown;
[0027] Figures 19A-19D A schematic diagram of LIC model parameter estimation using four pairs of samples according to some embodiments of the present disclosure is shown;
[0028] Figure 20 A schematic diagram of sub-block-level inter-frame prediction according to some embodiments of the present disclosure is shown;
[0029] Figures 21A-21CA schematic diagram illustrating a sample of derived LIC model parameters according to some embodiments of the present disclosure is shown;
[0030] Figures 22A-22C A schematic diagram illustrating a sample of LIC model parameters derived at the sub-block level according to some embodiments of the present disclosure is shown;
[0031] Figures 23A-23C A schematic diagram is shown for deriving LIC model parameter samples according to some embodiments of the present disclosure;
[0032] Figure 24 A flowchart of a video processing method according to some embodiments of the present disclosure is shown;
[0033] Figure 25 A flowchart of a video processing method according to some embodiments of the present disclosure is shown. Detailed Implementation
[0034] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings, wherein, unless otherwise specified, the same numbers in different drawings represent the same or similar elements. The embodiments presented in the following description of the embodiments of this disclosure do not represent all implementations consistent with this disclosure. Rather, they are merely examples of apparatuses and methods related to the content of this disclosure as described in the appended claims. An alternative embodiment of this disclosure will be described in more detail below. If the terms and definitions provided in this disclosure conflict with those introduced, the terms and definitions provided in this disclosure shall prevail.
[0035] The Audio and Video Standards (AVS) Working Group is the standards development organization for the AVS series of video standards. The AVS Working Group is currently developing the AVS3 video standard, the third generation of the AVS series. AVS3's predecessors, AVS1 and AVS2, were released in 2006 and 2016, respectively. The AVS3 standard is based on the same hybrid video coding system as modern video compression standards such as AVS1, AVS2, H.264 / AVC, and H.265 / HEVC.
[0036] The High Performance Model (HPM) was selected by the AVS working group as the new reference software platform for the development of the AVS3 standard. The initial technologies in HPM inherited from the AVS2 standard, and were then modified and enhanced using new advanced video coding techniques to improve its compression performance. Compared to its predecessor, AVS2, the final completed first phase of AVS3 achieved a coding performance improvement of over 20%. AVS continues to improve the compression performance of its coding techniques, and is developing the second phase of the AVS3 standard based on the first phase to further enhance coding efficiency.
[0037] Video is a collection of still images (or "frames") arranged chronologically to store visual information. Video capture devices (e.g., cameras) can be used to capture and store these images in chronological order, and video playback devices (e.g., televisions, computers, smartphones, tablets, video players, or any end-user terminal with a display capability) can be used to display these images in chronological order. Furthermore, in some applications, video capture devices can transmit captured video in real time to video playback devices (e.g., computers with monitors), for example, for surveillance, conferencing, or live streaming.
[0038] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed. For example, video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software on a processor (e.g., a processor in a general-purpose computer) or specialized hardware. The module or circuit used for compression is generally called an "encoder," while the module or circuit used for decompression is generally called a "decoder." Encoders and decoders can be collectively referred to as "codecs." Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuits, such as at least one microprocessor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), discrete logic, or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process embedded in a computer-readable medium. Video compression and decompression can be implemented using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, the AVS standard, etc. In some applications, a codec can decompress a video using a first encoding standard and then recompress the decompressed video using a second encoding standard. In this case, the codec can be called a "transcoder".
[0039] The video encoding process identifies and retains useful information that can be used to reconstruct images. If information that was ignored during the video encoding process cannot be fully reconstructed, the encoding process can be called "lossy." Otherwise, it can be called "lossless." Most encoding processes are lossy, which is a trade-off to reduce the storage space and transmission bandwidth required during video encoding.
[0040] In many cases, useful information about the image being encoded (referred to as the "current image") can be the changes relative to a reference image (e.g., a previously encoded or reconstructed image). These changes can include variations in pixel position, brightness, or color. For example, changes in the position of a set of pixels representing an object can reflect the object's movement between the reference image and the current image.
[0041] A coded image that does not reference another image (i.e., the image is its own reference image) is called an "I-frame image". An image is called a "P-frame image" if some or all of its blocks (e.g., portions of a video image) are predicted using intra-frame prediction or inter-frame prediction with a reference image (e.g., one-way prediction). An image is called a "B-frame image" if at least one block in the image is predicted using two reference images (e.g., two-way prediction).
[0042] In this disclosure, a simplified Local Luminance Compensation (LIC) procedure can be applied to the encoder and decoder during video encoding or decoding. In simplified LIC, the samples used for LIC model parameter derivation are limited based on their location to reduce the memory required to store unrefined prediction samples and to reduce pipeline latency in LIC operations. In some embodiments, a Local Chroma Compensation (LCC) procedure can be applied to extend the compensation to the chroma components of the coded block to compensate for chroma variations between the current block and the prediction block during inter-frame prediction. Since both LIC and LCC are applied to inter-frame prediction blocks to further correct predicted sample values by compensating for luminance or chroma variations between the current block and the prediction block, LIC and LCC are also referred to as inter-frame prediction correction in this disclosure.
[0043] Figure 1 A structural diagram of a video sequence according to some embodiments of the present disclosure is shown. Video sequence 100 may be live video or video that has already been captured and archived. Video sequence 100 may be real-life video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real-life video with augmented reality effects). Video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) to receive video from a video content provider. Figure 1 As shown, video sequence 100 may include a series of images arranged chronologically along a timeline, including images 102, 104, 106, and 108. Images 102-106 are consecutive, and there are more images between images 106 and 108.
[0044] When video is compressed or decompressed, useful information about the image being encoded (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). These changes can include variations in pixel position, brightness, or color. For example, changes in the position of a set of pixels can reflect the movement of an object represented by those pixels between two images (e.g., a reference image and the current image).
[0045] For example, such as Figure 1 As shown, image 102 is an I-frame image, which serves as its reference image. Image 104 is a P-frame image, using image 102 as its reference image, as indicated by the arrow. Image 106 is a B-frame image, using images 104 and 108 as its reference images, as indicated by the arrow. In some embodiments, the reference image of an image may immediately precede or follow the image, or it may not immediately precede or follow the image. For example, the reference image of image 104 may be an image preceding image 102, that is, an image not immediately preceding image 104. Figure 1 The above-mentioned reference images 102-106 shown are merely examples and are not intended to limit the scope of this disclosure.
[0046] Due to computational complexity, in some embodiments, the video codec can segment an image into multiple basic segments and encode or decode the image segment by segment. That is, the video codec does not need to encode or decode the entire image at once. In this disclosure, such basic segments are referred to as basic processing units (“BPUs”). For example, Figure 1 An exemplary structure 110 for images (e.g., any one of images 102-108) of video sequence 100 is also shown. For example, structure 110 can be used to segment image 108. Figure 1 As shown, image 108 is divided into 4×4 basic processing units. In some embodiments, the basic processing unit may be referred to as a “coding tree unit” (“CTU”) in some video coding standards (e.g., AVS3, H.265 / HEVC, or H.266 / VVC), or as a “macroblock” in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC). In AVS3, the CTU can be the largest block unit, up to 128×128 luminance samples (plus corresponding chrominance samples, depending on the chrominance format).
[0047] Figure 1The basic processing units shown are for illustrative purposes only. Basic processing units in an image can have variable sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of any shape and size. The size and shape of the basic processing units used for an image can be chosen based on a balance between coding efficiency and the level of detail to be preserved within the basic processing unit.
[0048] A basic processing unit can be a logical unit that may include a set of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color image may include a luminance component (Y) representing achromatic luminance information, at least one chrominance component (e.g., Cb and Cr) representing color information, and associated syntax elements, wherein the luminance and chrominance components may have the same size as the basic processing unit. In some video coding standards, the luminance and chrominance components may be referred to as “code tree blocks” (“CTBs”). Operations performed on a basic processing unit may be repeated on its luminance and chrominance components.
[0049] During the various operational stages of video encoding, the size of a basic processing unit may still be too large for the processing operations, and therefore it can be further divided into segments referred to herein as "basic processing subunits." For example, in the mode determination stage, the encoder can divide the basic processing unit into multiple basic processing subunits and determine the prediction type for each individual basic processing subunit. Figure 1 As shown, the basic processing unit 112 in structure 110 is further divided into 4×4 basic processing subunits. For example, in AVS3, the CTU can be further divided into coding units (CUs) using a quadtree, binary tree, or extended binary tree. Figure 1The basic processing subunits in the text are for illustrative purposes only. Different basic processing units for the same image can be divided into basic processing subunits in different schemes. In some video coding standards (e.g., AVS3, H.265 / HEVC, or H.266 / VVC), a basic processing subunit may be called a “coding unit” (“CU”), or in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC) a “block”. The size of a basic processing subunit can be equal to or smaller than the basic processing unit size. Similar to the basic processing unit, a basic processing subunit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., in a video frame buffer). The operations performed on a basic processing subunit can be repeated on its luminance and chrominance components. This division can be performed at further levels as needed, and different schemes can be used to divide the basic processing units at different stages. At the leaf nodes of the segmentation structure, encoded information such as the coding mode (e.g., intra-frame prediction mode or inter-frame prediction mode), the motion information required for the corresponding coding mode (e.g., reference index, motion vector (MV) etc.), and the quantization residual coefficients are sent.
[0050] In some cases, the basic processing unit (BJU) may still be too large to be processed in certain operational stages of video coding, such as the prediction or transform stages. Therefore, the encoder can further divide the BJU into smaller segments (e.g., called "prediction blocks" or "PBs") at which prediction operations can be performed. Similarly, the encoder can further divide the BJU into smaller segments (e.g., called "transform blocks" or "TBs") at which transform operations can be performed. The partitioning scheme of the same BJU can differ between the prediction and transform stages. For example, prediction blocks (PBs) and transform blocks (TBs) of the same CU can have different sizes and numbers. The operations in the mode determination, prediction, and transform stages will be discussed in later paragraphs. Figure 2 and Figure 3 The examples provided illustrate this in detail.
[0051] Figure 2 A schematic diagram of an encoder 200 for a video encoding system (e.g., AVS3 or H.26x series) according to some embodiments of this disclosure is shown. The input video is processed block by block. As described above, in the AVS3 standard, the CTU is the largest block unit, with a maximum of 128 × 128 luma samples (plus corresponding chroma samples, depending on the chroma format). A CTU can be further divided into CUs using quadtrees, binary trees, or ternary trees. Figure 2As shown, encoder 200 can receive video sequence 202 generated by a video capture device (e.g., a camera). The term "receive" as used above can refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or otherwise inputting data. Encoder 200 can encode video sequence 202 into video bitstream 228. Similar to... Figure 1 Video sequence 100 and video sequence 202 may include a set of images (referred to as "original images") arranged in chronological order. Similar to... Figure 1 In structure 110, any original image of video sequence 202 can be divided by encoder 200 into basic processing units, basic processing sub-units, or processing regions. In some embodiments, encoder 200 can perform basic processing unit-level processing on the original images of video sequence 202. For example, encoder 200 can perform iterative processing. Figure 2 In the processing, encoder 200 can encode basic processing units in one iteration of the processing. In some embodiments, encoder 200 can encode regions of the original image of video sequence 202 (e.g., Figure 1 The slices 114-118 in the middle are processed in parallel.
[0052] Components 202, 2042, 2044, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the "forward path." In Figure 2 In this process, encoder 200 can feed the basic processing unit (referred to as "raw BPU") of the original images of video sequence 202 to two prediction stages, namely intra-frame prediction (also known as "intra-frame image prediction" or "spatial prediction") stage 2042 and inter-frame prediction (also known as "inter-frame image prediction," "motion compensation," "motion compensation prediction," or "temporal prediction") stage 2044, to perform prediction operations and generate corresponding prediction data 206 and prediction BPU 208. Specifically, encoder 200 can receive the raw BPU and prediction reference 224, which can be generated from the reconstruction path of previous iterations of the process.
[0053] The purpose of the intra-frame prediction phase 2042 and the inter-frame prediction phase 2044 is to reduce information redundancy by extracting prediction data 206 from prediction data 206 and prediction reference 224. Prediction data 206 can be used to reconstruct the original BPU into a predicted BPU 228. In some embodiments, intra-frame prediction can use pixels from at least one encoded neighboring BPU in the same image to predict the current BPU. That is, the prediction reference 224 in intra-frame prediction can include neighboring BPUs, allowing spatially adjacent samples to be used to predict the current block. Intra-frame prediction can reduce the inherent spatial redundancy of the image.
[0054] In some embodiments, inter-frame prediction can use regions of one or more encoded images (“reference images”) to predict the current BPU. That is, prediction reference 224 in inter-frame prediction can include encoded images. Inter-frame prediction can reduce the inherent temporal redundancy of images.
[0055] In the forward path, encoder 200 performs prediction operations in the intra-frame prediction phase 2042 and the inter-frame prediction phase 2044. For example, in the intra-frame prediction phase 2042, encoder 200 may perform intra-frame prediction. For the original BPU of the image being encoded, prediction reference 224 may include one or more adjacent BPUs in the same image that have been encoded (in the forward path) and reconstructed (in the reconstruction path). Encoder 200 can generate the predicted BPU 208 by extrapolating adjacent BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, encoder 200 may perform extrapolation at the pixel level, for example, by extrapolating the value of the corresponding pixel for each pixel of BPU 208. The adjacent BPU used for extrapolation can be positioned relative to the original BPU from various directions, such as vertical (e.g., at the top of the original BPU), horizontal (e.g., to the left of the original BPU), diagonal (e.g., to the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra-frame prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the adjacent BPU used, the size of the adjacent BPU used, the extrapolation parameters, the orientation of the adjacent BPU used relative to the original BPU, etc.
[0056] For example, in the inter-frame prediction stage 2042, encoder 200 can perform inter-frame prediction. For the original BPU of the current image, prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference image can be encoded and reconstructed using the BPU. For example, encoder 200 can add the reconstructed residual BPU 222 to prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same image are generated, encoder 200 can generate a reconstructed image as a reference image. Encoder 200 can perform a "motion estimation" operation to search for matching regions within the range of the reference image (referred to as a "search window"). The position of the search window in the reference image can be determined based on the position of the original BPU in the current image. For example, the search window can be centered at a position in the reference image that has the same coordinates as the original BPU in the current image, and can extend outward by a preset distance. When encoder 200 identifies (e.g., using a pixel recursive algorithm, block matching algorithm, etc.) a region similar to the original BPU in the search window, encoder 200 can determine that the region is a matching region. The matching region may have a different size than the original BPU (e.g., less than, equal to, greater than, or with a different shape). Because the reference image and the current image are temporally separated in the timeline (e.g., as...), Figure 1 As shown in the image, the matching region can be considered to "move" to the original BPU's position over time. The encoder 200 records the direction and distance of this movement as a "motion vector (MV)". In other words, MV is the positional difference between the reference block in the reference image and the current block in the current image. In inter-frame prediction, the reference block is used as the predictor for the current block; therefore, the reference block is also called the prediction block. When using multiple reference images (e.g., such as...),... Figure 1 In the image 106, encoder 200 can search for matching regions and determine the associated MV for each reference image. In some embodiments, encoder 200 can assign weights to the pixel values of the matching regions of each matching reference image.
[0057] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, prediction data 206 may include, for example, a reference index, the location (e.g., coordinates) of the matching region, the MV associated with the matching region, the number of reference images, the weights associated with the reference images, or other motion information.
[0058] To generate the predicted BPU 208, encoder 200 can perform a "motion compensation" operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on the predicted data 206 (e.g., MV) and the predicted reference 224. For example, encoder 200 can move the matching region of the reference image according to the MV, where encoder 200 can predict the original BPU of the current image. When using multiple reference images (e.g., such as...), Figure 1 In the image 106, encoder 200 can move the matching region of the reference image based on the corresponding MVs and the average pixel value of the matching region. In some embodiments, if encoder 200 has already assigned weights to the pixel values of the matching regions of each matching reference image, encoder 200 can add the weighted sum of the pixel values of the moved matching regions.
[0059] In some embodiments, inter-frame prediction can use unidirectional or bidirectional prediction, and can be either unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 Image 104 in the diagram is a one-way inter-frame prediction image, where the reference image (i.e., image 102) precedes image 104. In one-way prediction, only one MV pointing to a reference image is used to generate the prediction signal for the current block.
[0060] On the other hand, bidirectional inter-frame prediction can use one or more reference images in two temporal directions relative to the current image. For example, Figure 1 Image 106 in the image is a bidirectional inter-frame prediction image, where reference images (e.g., images 104 and 108) are in the opposite time direction relative to image 104. In bidirectional prediction, two MVs are used to generate the prediction signal for the current block, where each MV points to a corresponding reference image. After generating video stream 228, the MVs and reference indices can be sent to the decoder in video stream 228 to identify where the prediction signal for the current block originates.
[0061] For inter-frame prediction CUs, motion parameters may include the video difference (MV), reference picture index, and reference picture list, or other additional information required using the index or coding features to be used. Motion parameters can be signaled explicitly or implicitly. In AVS3, under certain inter-frame coding modes (e.g., skip mode or direct mode), motion parameters (e.g., MV difference and reference picture index) are not encoded and signaled in the video bitstream 228. Instead, motion parameters can be derived on the decoder side using preset rules, which are the same as those defined in encoder 200. Details of skip mode and direct mode will be described in detail in the following paragraphs.
[0062] Following the intra-frame prediction phase 2042 and the inter-frame prediction phase 2044, in the mode determination phase 230, the encoder 200 can select a prediction mode (e.g., one of intra-frame prediction or inter-frame prediction) for the current iteration of the processing. For example, the encoder 200 can execute a rate-distortion optimization method, wherein the encoder 200 selects the prediction mode based on the bit rate of the candidate prediction mode and the distortion of the reference image reconstructed under the candidate prediction mode to minimize the cost function value. Based on the selected prediction mode, the encoder 200 can generate a corresponding prediction BPU 208 (e.g., prediction block) and prediction data 206.
[0063] In some embodiments, the predicted BPU208 may be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU208 is typically slightly different from the original BPU. To record this difference, after generating the predicted BPU208, the encoder 200 may subtract it from the original BPU to generate a residual BPU210, also referred to as the prediction residual.
[0064] For example, encoder 200 can subtract the pixel values (e.g., grayscale or RGB values) of the predicted BPU 208 from the corresponding pixel values of the original BPU. Each pixel of the residual BPU 210 can have a residual value as the result of the subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 can have fewer bits, but can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU can be compressed.
[0065] After generating the residual BPU210, encoder 200 can feed the residual BPU210 to transform stage 212 and quantization stage 214 to generate quantized residual coefficients 216. To further compress the residual BPU210, in transform stage 212, encoder 200 can reduce its spatial redundancy by decomposing the residual BPU210 into a set of two-dimensional “base patterns,” each associated with a “transform coefficient.” The base patterns can have the same size (e.g., the size of the residual BPU210). Each base pattern can represent a frequency component of the residual BPU210 (e.g., the frequency of brightness variation). No base pattern can be copied from any combination of any other base patterns (e.g., a linear combination). In other words, the decomposition decomposes the variation of the residual BPU210 into the frequency domain. This decomposition is analogous to the discrete Fourier transform of a function, where the base patterns are analogous to the base functions of the discrete Fourier transform (e.g., trigonometric functions), and the transform coefficients are analogous to the coefficients associated with the base functions.
[0066] Different transformation algorithms can use different coding patterns. Various transformation algorithms can be used in the transformation stage 212, such as discrete cosine transform, discrete sine transform, etc. The transformation in the transformation stage 212 is reversible. That is, the encoder 200 can recover the residual BPU210 through the inverse operation of the transformation (called the "inverse transform"). For example, to recover the pixels of the residual BPU210, the inverse transform can be to multiply the values of the corresponding pixels in the coding pattern by the corresponding correlation coefficients and sum the products to produce a weighted sum. For video coding standards, the encoder 200 and the corresponding decoder (e.g., Figure 3 The decoder 300 can use the same transform algorithm (the same encoding pattern). Therefore, the encoder 200 can record only the transform coefficients, and the decoder 300 can reconstruct the residual BPU210 based on these transform coefficients without receiving the encoding pattern from the encoder 200. Compared to the residual BPU210, the transform coefficients have fewer bits, but can be used to reconstruct the residual BPU210 without significant quality degradation. Therefore, the residual BPU210 can be further compressed.
[0067] Encoder 200 can further compress the transform coefficients in quantization stage 214. During the transform process, different coding patterns can be represented by different change frequencies (e.g., brightness change frequencies). Because the human eye is generally better at recognizing low-frequency changes, encoder 200 can ignore information about high-frequency changes without causing a significant degradation in decoding quality. For example, in quantization stage 214, encoder 200 can generate quantization residual coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to its nearest integer. After the above operation, some transform coefficients of high-frequency coding patterns can be converted to zero, and transform coefficients of low-frequency coding patterns can be converted to smaller integers. Encoder 200 can ignore the zero-value quantization residual coefficients 216, further compressing the transform coefficients. The quantization process is also reversible, whereby the quantization residual coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (called "inverse quantization").
[0068] Because encoder 200 does not consider the remainder of such divisions during rounding operations, quantization stage 214 may be lossy. Generally, quantization stage 214 can cause the greatest information loss during the encoding process. The greater the information loss, the fewer bits are needed for the quantization residual coefficients 216. To obtain different levels of information loss, encoder 200 can use different values of the quantization parameter or any other parameter of the quantization process.
[0069] Encoder 200 can feed the prediction data 206 and quantization residual coefficients 216 to the binary encoding stage 226 to generate the video bitstream 228, thus completing the forward path. In the binary encoding stage 226, encoder 200 can use binary encoding techniques to encode the prediction data 206 and quantization residual coefficients 216. For example, the binary encoding techniques mentioned above can be entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding (CABAC), or any other lossless or lossy compression algorithm.
[0070] For example, the CABAC encoding process in binary encoding stage 226 may include a binaryization step, a context modeling step, and a binary arithmetic encoding step. If the syntax element is not binary, encoder 200 first maps the syntax element to a binary sequence. Encoder 200 may choose to encode in either a context encoding mode or a bypass encoding mode. In some embodiments, for context encoding mode, the probability model of the binary to be encoded is selected by the “context,” which refers to previously encoded syntax elements. The binary and the selected context model are then passed to an arithmetic encoding engine, which can encode the binary and update the corresponding probability distribution of the context model. In some embodiments, for bypass encoding mode, the binary is encoded with a fixed probability (e.g., probability equal to 0.5) without selecting a probability model via the “context.” In some embodiments, bypass encoding mode is selected for a particular binary, thereby accelerating the entropy encoding process by ignoring the loss of encoding efficiency.
[0071] In some embodiments, in addition to the prediction data 206 and the quantization residual coefficients 216, the encoder 200 may also encode other information in the binary encoding stage 226, such as the prediction mode selected in the prediction stage (e.g., intra-frame prediction stage 2042 or inter-frame prediction stage 2044), parameters of the prediction operation (e.g., intra-frame prediction mode, motion information, etc.), the transform type of the transform stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. That is, the encoded information can be sent to the binary encoding stage 226, thereby further reducing the bitrate before being packaged into the video stream 228. The encoder 200 can use the output data of the binary encoding stage 226 to generate the video stream 228. In some embodiments, the video stream 228 can be further packaged for network transmission.
[0072] Components 218, 220, 222, 224, 232, and 234 can be referred to as "reconstruction paths." Reconstruction paths can be used to ensure that encoder 200 and its corresponding decoder (e.g., Figure 3 Both decoders (300) use the same reference data for prediction.
[0073] During this process, after quantization phase 214, encoder 200 can feed quantization residual coefficients 216 to inverse quantization phase 218 and inverse transform phase 220 to generate reconstruction residual BPU 222. In inverse quantization phase 218, encoder 200 can perform inverse quantization on quantization residual coefficients 216 to generate reconstruction transform coefficients. In inverse transform phase 220, encoder 200 can generate reconstruction residual BPU 222 based on the reconstruction transform coefficients. Encoder 200 can add reconstruction residual BPU 222 to prediction BPU 208 to generate prediction reference 224, which will be used in prediction phases 2042 and 2044 in the next iteration.
[0074] In the reconstruction path, if intra-prediction mode has already been selected in the forward path, encoder 200 can directly feed prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image) to intra-prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current image) after generating prediction reference 224. If inter-prediction mode has already been selected in the forward path, encoder 200 can feed prediction reference 224 (e.g., the current image where all BPUs have been encoded and reconstructed) to loop filter stage 232 after generating prediction reference 224. At the loop filter stage, encoder 200 can apply loop filters to prediction reference 224 to reduce or eliminate distortions introduced by inter-prediction (e.g., block artifacts). Encoder 200 can apply various loop filter techniques in loop filter stage 232, such as deblocking, sample adaptive offset (SAO), adaptive loop filter (ALF), etc. In SAO, a nonlinear amplitude mapping is introduced in the inter-frame prediction loop after the deblocking filter to reconstruct the original signal amplitude using a lookup table, which is described by several additional parameters determined by histogram analysis on the encoder side.
[0075] The reference image after loop filtering can be stored in buffer 234 (or "decoded image buffer") for later use (e.g., as an inter-frame prediction reference image for future images of video sequence 202). Encoder 200 can store one or more reference images in buffer 234 for use in inter-frame prediction stage 2044. In some embodiments, encoder 200 can encode the parameters of the loop filter (e.g., loop filter strength), as well as the quantization residual coefficients 216, prediction data 206, and other information in binary encoding stage 226.
[0076] Encoder 200 can iteratively execute the above processing steps to encode each original BPU of the original image (in the forward path) and generate a prediction reference 224, which is used to encode the next original BPU of the original image (in the reconstruction path). After encoding all the original BPUs of the original image, encoder 200 can continue to encode the next image in video sequence 202.
[0077] It should be noted that other variations of the encoding process can also be used to encode the video sequence 202. In some embodiments, the encoder 200 may execute the various processing stages in different orders. In some embodiments, at least one stage of the encoding process may be combined into a single stage. In some embodiments, a single stage of the encoding process may be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, the encoding process may include... Figure 2 Additional stages not shown. In some embodiments, the encoding process may be omitted. Figure 2 One or more stages in the process.
[0078] For example, in some embodiments, encoder 200 can operate in a transform skip mode. In transform skip mode, transform stage 212 is bypassed, and a TB transform skip flag can be signaled. This can improve compression of certain types of video content, such as computer-generated images or images or graphics mixed with camera view content (e.g., scrolling text). Furthermore, encoder 200 can also operate in a lossless mode. In lossless mode, transform stage 212, quantization stage 214, and other processing stages affecting the decoded image (e.g., SAO and deblocking filters) are bypassed. The residual signal from intra-frame prediction stage 2042 or inter-frame prediction stage 2044 is fed to binary encoding stage 226 using the same neighborhood “context” applied to the quantization transform coefficients. This process can employ mathematically lossless reconstruction. Therefore, both transform and transform skip residual coefficients are encoded within non-overlapping CGs. That is, each CG may include at least one transform residual coefficient or at least one transform skip residual coefficient.
[0079] Figure 3 A flowchart of a decoder 300 for a video encoding system (e.g., AVS3 or H.26x series) according to some embodiments of the present disclosure is shown. The decoder 300 can perform actions corresponding to... Figure 2 The compression and decompression processes within the process. Figure 2 and Figure 3 In the text, the corresponding stages in compression and decompression are labeled with the same numbers.
[0080] In some embodiments, the decompression process can be similar to Figure 2The reconstruction path in the code. Decoder 300 can correspondingly decode video stream 228 into video stream 304. Video stream 304 can be similar to... Figure 2 The video sequence 202 in the video. However, due to the compression and decompression processes (e.g., Figure 2 Information loss during the quantization stage 214) means that video stream 304 may differ from video sequence 202. Similar to... Figure 2 The encoder 200 and decoder 300 can perform decoding processing on each image encoded in the video stream 228 at the level of a basic processing unit (BPU). For example, the decoder 300 can perform the processing in an iterative manner, wherein the decoder 300 can decode the basic processing unit in one iteration. In some embodiments, the decoder 300 can perform the decoding process in parallel on a region (e.g., slices 114-118) of each image encoded in the video stream 228.
[0081] exist Figure 3 In the binary decoding stage 302, decoder 300 can feed a portion of the video bitstream 228 associated with the basic processing unit (referred to as the "encoding BPU") of the encoded image to the binary decoding stage 302. In the binary decoding stage 302, decoder 300 can unpack the video bitstream and decode it into prediction data 206 and quantization residual coefficients 216. Decoder 300 can use the prediction data 206 and quantization residual coefficients to reconstruct the video stream 304 corresponding to the video bitstream 228.
[0082] Decoder 300 can perform the inverse operation of the binary encoding technique used by encoder 200 (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm) in binary decoding stage 302. In some embodiments, in addition to prediction data 206 and quantization residual coefficients 216, decoder 300 can decode other information in binary decoding stage 302, such as prediction mode, parameters of prediction operation, transform type, quantization parameters (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. In some embodiments, if video stream 228 is transmitted over the network in packet form, decoder 300 can unpack it before feeding video stream 228 to binary decoding stage 302.
[0083] Decoder 300 can feed the quantized residual coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate the reconstructed residual BPU 222. Decoder 300 can feed the prediction data 206 to the intra-frame prediction stage 2042 and the inter-frame prediction stage 2044 to generate the prediction BPU 208. Specifically, for the coding basic processing unit (referred to as the "current BPU") of the coded picture being decoded (referred to as the "current picture"), the prediction data 206 decoded by decoder 300 from binary decoding stage 302 can include various types of data depending on the prediction mode used by encoder 200 to encode the current BPU. For example, if encoder 200 uses intra-frame prediction to encode the current BPU, the prediction data 206 can include encoding information, such as a prediction mode indicator (e.g., a flag value) indicating intra-frame prediction, parameters of the intra-frame prediction operation, etc. Parameters of the intra-frame prediction operation can include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, the sizes of neighboring BPUs, extrapolation parameters, the orientation of neighboring BPUs relative to the original BPU, etc. For example, if encoder 200 uses inter-frame prediction to encode the current BPU, the prediction data 206 may include encoding information, such as prediction mode indicators (e.g., flag values) indicating inter-frame prediction, parameters of the inter-frame prediction operation, etc. Parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights associated with each reference image, the positions (e.g., coordinates) of one or more matching regions in the corresponding reference images, and one or more MVs associated with each matching region, etc.
[0084] Therefore, the prediction mode indicator can be used to select whether to invoke the inter-frame or intra-frame prediction module. Then, the parameters for the corresponding prediction operation can be sent to the corresponding prediction module to generate the prediction signal. Specifically, based on the prediction mode indicator, the decoder 300 can decide whether to perform intra-frame prediction in the intra-frame prediction stage 2042 or inter-frame prediction in the inter-frame prediction stage 2044. The specific steps for performing such intra-frame or inter-frame prediction are described in [the following text is incomplete and requires further context]. Figure 2 This has been discussed previously and will not be repeated below. After performing intra-frame prediction or inter-frame prediction, decoder 300 can generate prediction BPU208.
[0085] After generating the prediction BPU 208, the decoder 300 can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded image buffer in computer memory). The decoder 300 can feed the prediction reference 224 to the intra-frame prediction stage 2042 and the inter-frame prediction stage 2044 for performing prediction operations in the next iteration.
[0086] For example, if intra-frame prediction is used to decode the current BPU in intra-frame prediction stage 2042, then after generating prediction reference 224 (e.g., the decoded current BPU), decoder 300 can directly feed prediction reference 224 to intra-frame prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current image). If inter-frame prediction is used to decode the current BPU in inter-frame prediction stage 2044, then after generating prediction reference 224 (e.g., a reference image where all BPUs have been decoded), decoder 300 can feed prediction reference 224 to loop filter stage 232 to reduce or eliminate distortion (e.g., block artifacts). Furthermore, prediction data 206 may also include parameters of the loop filter (e.g., loop filter strength). Therefore, decoder 300 can... Figure 2 The described method applies a loop filter to prediction reference 224. For example, a loop filter such as deblocking, SAO, or ALF can be applied to form a loop-filtered reference image, which is stored in buffer 234 (e.g., a decoded picture buffer (DPB) in computer memory) for later use (e.g., in the inter-frame prediction stage 2044 for predicting future encoded images of video stream 228). In some embodiments, the reconstructed image from buffer 234 can also be sent to a display, such as a television, personal computer, smartphone, or tablet, for viewing by an end user.
[0087] Decoder 300 iteratively performs the decoding process to decode each encoded BPU of the encoded image and generate a prediction reference 224 for encoding the next encoded BPU of that encoded image. After decoding all encoded BPUs of the encoded image, decoder 300 can output the image to video stream 304 for display and continue decoding the next encoded image in video stream 228.
[0088] Figure 4 A structural block diagram of an apparatus 400 for encoding or decoding video according to some embodiments of the present disclosure is shown. Figure 4As shown, device 400 may include processor 402. When processor 402 executes the instructions described above, device 400 may become a dedicated device for video encoding or video decoding. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any combination of any number of central processing units (or “CPU”), graphics processing units (or “GPU”), neural processing units (“NPU”), microcontroller units (“MCU”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), system-on-a-chip (SoCs), application-specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may also be grouped into a set of processors comprising a single logic component. For example, such as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b and processor 402n.
[0089] The device 400 may also include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, etc.). For example, such as Figure 4 As shown, the data it stores may include program instructions (e.g., for implementing...). Figure 2 and Figure 3 The processor 402 can access the program instructions and data for processing (e.g., video sequence 202, video stream 228, or video stream 304) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAM), read-only memories (ROM), optical discs, magnetic disks, hard disks, solid-state drives, flash drives, secure digital cards (SD cards), memory sticks, small flash memory (CF cards), etc. The memory 404 may also be grouped into a set of memories as a single logical component. Figure 4 (Not shown in the image).
[0090] Bus 410 may be a communication device for transmitting data between components within device 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Fast Port), etc.
[0091] For ease of explanation and to avoid ambiguity, the processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry" in this disclosure. Data processing circuitry can be implemented in hardware or as a combination of software, hardware, or firmware. Furthermore, data processing circuitry can be a single, independent module or can be wholly or partially integrated into other components of the device 400.
[0092] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (“NFC”) adapters, cellular network chips, etc.
[0093] In some embodiments, the device 400 may optionally further include a peripheral interface 408 to provide connectivity to at least one peripheral device. Figure 4 As shown, peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to video archives), etc.
[0094] It should be noted that the video codec (e.g., the codec that performs the process of encoder 200 or decoder 300) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of process encoder 200 or decoder 300 can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. Similarly, some or all stages of process encoder 200 or decoder 300 can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).
[0095] exist Figure 2 and Figure 3 In the inter-frame prediction phase 2044, the reference index is used to indicate which previously encoded image the reference block comes from. The motion vector (MV), which is the positional difference between the reference block in the reference image and the current block in the current image, is used to indicate the position of the reference block in the reference image. For bidirectional prediction (e.g., ...), Figure 1In the example, image 106), two reference blocks are used to generate a combined prediction block: one from a reference image in reference image list 0 (e.g., image 104) and the other from a reference image in reference image list 1 (e.g., image 108). Therefore, bidirectional prediction requires two reference indices (e.g., reference indices for list 0 and list 1) and two motion vectors (e.g., motion vectors for list 0 and list 1). The motion vectors are determined by the encoder and communicated to the decoder as signals. In some embodiments, to save on transmission costs, the motion vector difference (MVD) is communicated as a signal in the bitstream instead. For the decoder, a motion vector predictor (MVP) can be derived based on spatial and temporal neighboring block motion information, and the MV can be obtained by adding the MVD parsed from the bitstream to the MVP.
[0096] As described above, different modes can be used to implement the video encoding or decoding process. In some conventional inter-frame coding modes, encoder 200 can explicitly signal MVs (multiple MVs), the corresponding reference picture index for each reference picture list, and reference picture list usage flags or other information based on each CU. On the other hand, when CUs encode in skip mode or direct mode, motion information, including reference indices and motion vectors, is not signaled to decoder 300 in video bitstream 228. Instead, motion information can be derived in decoder 300 using the same rules as encoder 200. Since skip mode and direct mode share the same motion information derivation rules, they have the same motion information. The difference between the two modes is that in skip mode, signaling of the prediction residual is skipped by setting the residual value to zero. In direct mode, the prediction residual is still signaled in the bitstream.
[0097] For example, when encoding a CU using skip mode, the CU is associated with a PU and has no significant residual coefficients, no encoded MV difference, or a reference image index. In skip mode, signaling for residual data can be skipped by setting the residual value to zero. In direct mode, the remaining data is transmitted while exporting motion information and partitions.
[0098] On the other hand, in inter-frame mode, when the motion vector difference and reference index are signaled to decoder 300, encoder 200 can select any allowed value for the motion vector and reference index. Compared to inter-frame mode for signaling motion information, the bit values dedicated to motion information can therefore be stored in skip mode or direct mode. However, encoder 200 and decoder 300 need to follow the same rules to derive the motion vector and reference index to perform inter-frame prediction 2044. In some embodiments, the derivation of motion information can be based on spatially or temporally adjacent blocks. Therefore, skip mode and direct mode are suitable when the motion information of the current block is close to the motion information of its spatially or temporally adjacent blocks.
[0099] For example, in AVS3, skip mode or direct mode allows motion information (e.g., reference indexes, MVs, etc.) to be inherited from spatial or temporal (co-localization) neighbors, from which a candidate list of motion candidates can be generated. In some embodiments, to derive motion information for inter-frame prediction 2044 in skip mode or direct mode, encoder 200 may first derive the candidate list of motion candidates and select a motion candidate to perform inter-frame prediction 2044. When signaling the video stream 228, encoder 200 may signal the index of the selected candidate. On the decoder side, decoder 300 may obtain the index parsed from video stream 228, derive the same candidate list, and use the same motion candidate (including motion vectors and reference image indices) to perform inter-frame prediction 2044.
[0100] The AVS3 specification includes different skip and direct modes, such as normal skip and direct modes, final motion vector representation mode, angle-weighted prediction mode, enhanced temporal motion vector prediction mode, and affine motion compensation skip / direct mode. The motion candidate list can include multiple candidate objects obtained based on different methods. For example, for the normal skip and direct model, the motion candidate list can have 12 candidates, including temporal motion vector predictor (TMVP) candidates (i.e., temporal candidates), one or more spatial motion vector predictor (SMVP) candidates (i.e., spatial candidates), one or more motion vector angle prediction (MVAP) candidates (i.e., sub-block-based spatial candidates), and one or more history-based motion vector predictor (HMVP) candidates (i.e., history-based candidates). In some embodiments, the encoder or decoder can first export and add TMVP and SMVP candidates to the candidate list. After adding TMVP and SMVP candidates, the encoder or decoder exports and adds MVAP and HMVP candidates. In some embodiments, the number of MVAP candidates added to the candidate list can vary depending on the number of available directions during the MVAP process. For example, the number of MVAP candidates can be between 0 and a maximum number (e.g., 5). After adding MVAP candidates, one or more HMVP candidates can be added to the candidate list until the total number of candidates reaches the target number, and the maximum number can also be signaled in the bitstream.
[0101] In some embodiments, the first candidate is the TMVP derived from the MV of the juxtaposed block in a predefined reference frame. The predefined reference frame is defined as a reference frame with a reference index value of 0 in List 1 of B-frames or a reference frame in List 0 of P-frames. When the juxtaposed block MV is unavailable, the MV predictor (MVP) derived based on the MV of the spatially adjacent block is used as the block-level TMVP.
[0102] In some other embodiments, a sub-block level TMVP may be used. Figure 5 A schematic diagram illustrating the derivation of a sub-block temporal motion vector predictor (TMVP) according to some embodiments of this disclosure is shown. Specifically, when sub-block level TMVP is enabled, the current block 500 is cross-divided into four sub-blocks 510, 520, 530, and 540. Motion vectors can be derived for each sub-block 510, 520, 530, or 540. Figure 5As shown, for each sub-block, corner samples (e.g., samples 512, 522, 532, or 542) are used to find juxtaposed blocks in the reference image. Motion vectors stored in the temporal motion information buffer are extracted and scaled, covering samples with the same coordinates in the reference image as the corner samples. The scaled motion vectors are used as the TMVP of the sub-block. If list 0 motion vectors are available in the temporal motion information buffer (i.e., the juxtaposed block has list 0 motion vectors), list 0 motion vectors are extracted and scaled, and used as the MV of list 0 of the TMVP of the sub-block. If list 1 motion vectors are available in the temporal motion information buffer (i.e., the juxtaposed block has list 1 motion vectors), list 1 motion vectors are extracted and scaled, and used as the MV of list 1 of the TMVP of the sub-block. If neither list 0 nor list 1 motion vectors are available in the temporal motion information buffer, a block-level TMVP can be derived and used as the TMVP of the sub-block. In some embodiments, the sub-block TMVP can only be used for blocks with a width and height greater than or equal to 16, such that the width and height of each sub-block are not less than 8.
[0103] Figure 6 A schematic diagram of adjacent blocks 600 for spatial motion vector predictor (SMVP) derivation according to some embodiments of this disclosure is shown. Figure 6As shown, the second, third, and fourth candidates are SMVPs derived from six adjacent blocks 610, 620, 630, 640, 650, and 660. In some embodiments, the second candidate is a bidirectional prediction candidate. The third candidate is a unidirectional prediction candidate with reference frames from list 0. The fourth candidate is a unidirectional prediction candidate with reference frames from list 1. For these three candidates, encoder 200 or decoder 300 may examine the motion information of the six adjacent blocks in the order of blocks 610, 620, 630, 640, 650, and 660, and borrow the MV and reference index of a first available block with the same prediction type as the current candidate. For example, for the second candidate, the motion information of the first block using bidirectional prediction is borrowed in the order of blocks 610, 620, 630, 640, 650, and 660. If encoder 200 or decoder 300 determines that such a block does not exist, and two or more adjacent blocks use unidirectional prediction with reference image list 0, and two or more adjacent blocks use unidirectional prediction with reference image list 1, then for the second candidate, encoder 200 or decoder 300 can borrow and combine the motion information of the first block using unidirectional prediction with reference image list 0 and the motion information of the first block using reference image list 1, in the order of blocks 610, 620, 630, 640, 650, and 660, to obtain the motion vector and reference index of the second candidate. Otherwise, encoder 200 or decoder 300 can set both the motion vector and reference index to zero. For the third candidate, encoder 200 or decoder 300 can borrow the motion information of the first block using unidirectional prediction with reference image list 0. If no such block exists among the six neighboring blocks, and two or more neighboring blocks use bidirectional prediction, then for the third candidate, encoder 200 or decoder 300 can borrow motion information from list 0 of the first bidirectional prediction block in the order of blocks 660, 650, 640, 630, 620, and 610. Alternatively, encoder 200 or decoder 300 can set both the motion vector and reference index to zero. Similarly, for the fourth candidate, encoder 200 or decoder 300 can borrow motion information from the first block using unidirectional prediction of reference picture list 1. If no such block exists among the six neighboring blocks, and two or more neighboring blocks use bidirectional prediction, then for the fourth candidate, encoder 200 or decoder 300 can borrow motion information from list 1 of the first bidirectional prediction block in the order of blocks 660, 650, 640, 630, 620, and 610. Alternatively, encoder 200 or decoder 300 can set both the motion vector value and reference index value to zero.
[0104] As mentioned above, MVAP candidates are ranked after SMVP candidates. In some embodiments, there can be a maximum of five MVAP candidates. Therefore, candidates five through ninth can all be MVAP candidates. Figure 7 As shown, it illustrates sub-blocks S1-S8 in a coding unit (CU) 710 associated with video frame 700 during an MVAP process consistent with some embodiments of this disclosure. During the MVAP process, CU 710 is divided into sub-blocks S1-S8. In some embodiments, each sub-block S1-S8 is 8×8 in size. For each 8×8 sub-block, motion information including a reference index and the MV is predicted based on reference motion information.
[0105] like Figure 7 As shown, the reference motion information for sub-block S3 in the current CU710 is the motion information (e.g., reference MV) of the horizontally adjacent block 720 and the vertically adjacent block 730 of the current CU710 in five different directions (D0-D4). Specifically, direction D0 is called the horizontal direction, direction D1 is called the vertical direction, direction D2 is called the horizontal upward direction, direction D3 is called the horizontal downward direction, and direction D4 is called the vertical downward direction. In other words, as... Figure 7 As shown, MVAP candidates are sub-block level candidates and can be derived by predicting motion information angles from five directions based on reference motion information, which consists of the MVs and reference indices of neighboring blocks 720 and 730. Neighbor motion information will be preferentially checked at the 4×4 block level.
[0106] Figure 8 A 4×4 reference block A0-A adjacent to MVAP is shown according to some embodiments of this disclosure. 2m+2n The diagram is shown in Figure 800. If the motion information in the aforementioned 4×4 adjacent blocks is unavailable, it is filled with the adjacent available MV and reference index. The filled adjacent motion information can be used as reference motion information for angle prediction.
[0107] Refer again Figure 7 The availability of the five directions (D0-D4) can be checked by comparing with reference motion information. Only the available directions are used for prediction. Figure 7 The MV of each 8×8 sub-block S1-S8 within the current block 700. Therefore, the number of MVAP candidates can be from 0 to 5 depending on the availability of the prediction directions (D0-D4). For the first MVAP candidate, when A m-1+H / 8 and A m+n-1 The candidate is available when there is different motion information. For the second MVAP candidate, when A m+n+1+W / 8 and A m+n+1 The candidate is available when there is different motion information. For the third MVAP candidate, when A m+n-1 and A m+n The candidate is available when there is different motion information. For the fourth MVAP candidate, when A W / 8-1 and A m-1The candidate is available when there is different motion information. For the fifth MVAP candidate, when A m+n+1+W / 8 and A 2m+n+1 The candidate is available when there is different motion information.
[0108] Since MV prediction is applied to each 8×8 sub-block S1-S8 within the current coding block 700, MVAP candidates are sub-block level candidates. In other words, different sub-blocks S1-S8 within the current coding block 700 can have different MVs and reference indices.
[0109] The HMVP candidate follows the MVAP candidate and is derived from the motion information of previously encoded or decoded blocks. For example, after encoding (or decoding) an inter-coded block, Figure 2 Encoder 200 (or Figure 3 The decoder 300 can add motion information associated with the encoded / decoded block to the last entry of the HMVP table. In some embodiments, the size of the HMVP table can be set to 8, but this disclosure is not specifically limited thereto. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule can be used. If there are already 8 candidates in the table, the first candidate is removed when the current motion information is inserted into the table to keep the number of candidates in the table no greater than 8. In some embodiments, redundancy checks can be applied first to determine whether the same motion candidate already exists in the table. If the same candidate motion is found in the table, the candidate motion can be moved to the last entry of the table instead of inserting a new identical entry. The candidates in the HMVP table are used as HMVP candidates for skip mode and direct mode.
[0110] In some embodiments, the encoder may first check whether the HMVP candidate stored in the HMVP table is the same as any motion candidate in the candidate list. In response to an HMVP candidate being different from a motion candidate in the candidate list, the encoder adds the HMVP candidate to the candidate list. Otherwise, the encoder does not add the HMVP candidate to the candidate list. This process can be referred to as a "pruning" process.
[0111] For example, you can check from the last entry in the HMVP table to the first entry. If a candidate in the HMVP table is not the same as any candidate in the candidate list (e.g., a TMVP or SMVP candidate), the candidate in the HMVP table is added to the candidate list for normal skip and direct mode as an HMVP candidate. If a candidate in the HMVP table is the same as either a TMVP or an SMVP candidate, that candidate is not added to the candidate list for normal skip and direct mode to avoid candidate redundancy. Check the candidates in the HMVP table and insert them one by one into the candidate list for normal skip and direct mode until the candidate list for normal skip and direct mode is full, or check all candidates in the HMVP table. If the candidate list for normal skip and direct mode is not full after inserting an HMVP candidate, you can repeat the process for the last candidate until the candidate list is full.
[0112] Figure 9 A schematic diagram illustrating motion derivation in the final motion vector expression (UMVE) according to some embodiments of this disclosure is shown. In the AVS3 standard, UMVE is used as another normal skip and direct mode, in addition to the conventional normal skip and direct modes (where implicitly derived motion information is used to find reference blocks for inter-frame prediction). In UMVE, a basic motion candidate is selected from a UMVE candidate list containing only two candidates derived from motion information of spatially adjacent blocks, based on the basic candidate index of the signaling in the bitstream. The basic motion candidate is then further refined based on motion vector offset information from the signaling. The motion vector offset information includes an index specifying the offset distance and an index indicating the offset direction. The basic motion vector can be set as the starting point for refinement. The direction index indicates the direction of the motion vector offset relative to the starting point. Figure 9 As shown, the direction index can represent one of four directions, while the distance index specifies the magnitude of the motion vector offset. Tables 1 and 2 define the mapping from the distance index to the offset value. In some embodiments, a signal notification flag is used in the image header to indicate and use either the five MVD offsets in Table 1 or the eight MVD offsets in Table 2.
[0113] Table 1: Five MVD Offsets in UMVE Mode
[0114] MVD offset (pixels) 1 / 4 1 / 2 1 2 4
[0115] Table 2: 8 MVD Offsets in UMVE Mode
[0116]
[0117] Figure 10 A schematic diagram of an exemplary angle-weighted prediction (AWP) according to some embodiments of the present disclosure is shown.
[0118] In the AVS3 standard, AWP mode is used as another skip and direct mode. AWP mode is indicated by a flag in the signaling of the bitstream. In AWP mode, a motion vector candidate list is first constructed, containing five different unidirectional predicted motion vectors derived from spatially adjacent blocks and temporal motion vector predictors. To construct the unidirectional prediction candidate list, temporally juxtaposed blocks (denoted as T) and spatially adjacent blocks (e.g., ...) are checked sequentially. Figure 6 The motion information for the coded blocks 610, 620, 630, 640, 650, and 660 is shown. If the spatially adjacent blocks are unidirectional prediction blocks, the motion information can be directly inserted into the candidate list. If the adjacent blocks are bidirectional prediction blocks, the motion information from list 0 or list 1 is inserted as a unidirectional prediction candidate based on the parity of the index of the current candidate to be inserted. If the candidate list is not full after inserting all adjacent motion information, additional candidates can be derived based on the existing candidates in the list until the candidate list is full.
[0119] After constructing the unidirectional prediction candidate list, two unidirectional predicted motion vectors are selected from the motion vector candidate list based on two candidate indices in the bitstream to obtain two reference blocks. Unlike the bidirectional prediction inter-frame mode, where the two reference blocks are averaged with equal weights to obtain the final predicted block, in AWP mode, different samples can have different weights during the averaging process. Figure 10 As shown, the weights for each sample are predicted based on a reference weight array, and the values of the weights are set to 0 to 8. In some embodiments, the weight prediction is similar to intra-sample prediction. For each sample, the reference weights referenced by the current prediction direction are used as the weights for the current sample, depending on the prediction direction.
[0120] Figure 11 A schematic diagram is shown illustrating eight different prediction directions (1110-1180) supported in AWP mode according to some embodiments of this disclosure. Figure 11 As shown, prediction direction 1160 is horizontal, while prediction direction 1120 is vertical. Figure 12 A schematic diagram of seven different weight arrays (1210-1270) in an AWP mode according to some embodiments of the present disclosure is shown.
[0121] For a size w×h equal to 2 m ×2 n The encoded blocks, where m,n∈{3…6}, are supported in AWP mode. Figure 11 The eight predicted directions shown (1110-1180) and Figure 12The diagram shows seven different reference weight arrays (1210-1270). Therefore, 56 predictions, or 56 different weight distributions, can be obtained within the coded block. After determining the weight for each sample, Figure 2 Encoder 200 (or Figure 3 The decoder 300 in the code can derive the final predicted block by taking a weighted average of two reference blocks in a sample-based manner. In some embodiments, the formula for calculating the final predicted block P is as follows:
[0122] P = (P0 * W0 + P1 * W1) >> 3
[0123] Where "*" represents the dot product, P0 and P1 represent two reference blocks respectively, and W0 and W1 represent the derived weight matrices respectively, where W0+W1 is a matrix in which all elements are equal to 8.
[0124] In the AVS3 standard, an Enhanced Temporal Motion Vector Predictor (ETMVP) can be applied to derive motion information. In ETMVP, the current coded block can be divided into 8×8 sub-blocks, and each sub-block derives a motion vector based on the motion information of its corresponding temporally adjacent blocks. A motion candidate list is constructed when an ETMVP flag, signaled in the stream, indicates that ETMVP is enabled. Each candidate in the list contains a set of motion vectors, with one motion vector per 8×8 sub-block.
[0125] For the first candidate, the motion vector of each 8×8 sub-block is derived from the motion vector of the corresponding juxtaposed block in the reference image with a reference index value of 0 in reference image list 0. For the second candidate, the current block is first shifted down by 8 samples, and then the motion vector of each sub-block is derived from the motion vector of the corresponding juxtaposed block of the shifted block in the reference image with a reference index value of 0 in reference image list 0. For the third candidate, the current block is first shifted to the right by 8 samples, and then the motion vector of each sub-block is derived from the motion vector of the corresponding juxtaposed block of the shifted block in the reference image with a reference index value of 0 in reference image list 0. For the fourth candidate, the current block is first shifted up by 8 samples, and then the motion vector of each sub-block is derived from the motion vector of the corresponding juxtaposed block of the shifted block in the reference image with a reference index value of 0 in reference image list 0. For the fifth candidate, the current block is first shifted to the lower left by 8 samples, and then the motion vector of each sub-block is derived from the motion vector of the corresponding juxtaposed block of the shifted block in the reference image with a reference index value of 0 in reference image list 0. When inserting a candidate into the candidate list, the encoder 200 or decoder 300 can reduce the second candidate to the fifth candidate by comparing the motion information of two predefined sub-blocks in the reference image.
[0126] Figure 13A schematic diagram of a juxtaposed block 1300 and sub-blocks (A1-A4, B1-B4, and C1-C4) for candidate pruning according to some embodiments of the present disclosure is shown. Figure 13 As shown, during the pruning process applied to the corresponding co-occurring block 1300, for the second candidate, the motion information of sub-block A2 and sub-block C4 needs to be compared. For the third candidate, the motion information of sub-block A3 and sub-block B4 needs to be compared. For the fourth candidate, the motion information of sub-block A4 and sub-block C2 needs to be compared. For the fifth candidate, the motion information of sub-block A4 and sub-block B3 needs to be compared. If the motion information of the two compared sub-blocks is different, it is a valid candidate and is inserted into the candidate list. Otherwise, the current candidate is invalid and will not be inserted into the candidate list. If the number of valid candidates is less than 5 after checking all five candidates, the last valid candidate is repeated until the number of candidates equals 5 to complete the candidate list. Therefore, after constructing the candidate list, the decoder 300 can also select candidates by the candidate index of the signaling in the bitstream.
[0127] In some embodiments, when deriving the motion vector for each sub-block of the ETMVP, the list 0 motion vector of the juxtaposed block is used to derive the list 0 motion vector for the corresponding sub-block. The list 1 motion vector of the juxtaposed block is used to derive the list 1 motion vector for the corresponding sub-block. The reference index values for both list 0 and list 1 of the current sub-block are set to 0. Therefore, if the juxtaposed block is a list 0 unidirectional prediction block, the corresponding sub-block also uses list 0 unidirectional prediction. If the juxtaposed block is a list 1 unidirectional prediction block, the corresponding sub-block also uses list 1 unidirectional prediction. If the juxtaposed block is a bidirectional prediction block, the corresponding sub-block also uses bidirectional prediction. If the juxtaposed block is an intra-block, the motion vector of the corresponding sub-block can be set to the default motion vector derived from the spatially adjacent block.
[0128] MV represents the physical motion between two images at different times. However, the motion described above only represents translation, as all samples in the coded block have the same positional offset. To compensate for other motions, such as zooming in, zooming out, or rotation, the AVS3 standard employs affine motion compensation. In affine motion compensation, different samples in the coded block can have different motion vectors. The motion vector for each sample is derived from the motion vectors of the control points (CPs) according to an affine model. In some embodiments, affine motion compensation can only be applied to blocks with a size greater than or equal to 16×16.
[0129] Figure 14A and Figure 14B A schematic diagram of two control-point-based affine models for encoding blocks 1400a and 1400b according to some embodiments of the present disclosure is shown. In some embodiments, Figure 14A Control points 1410a and 1420a in the middle and Figure 14BControl points 1410b, 1420b, and 1430b are set at the corners of coding blocks 1400a and 1400b, respectively. Figure 14A As shown, for a four-parameter affine model, two control points, 1410a and 1420a, are required. Figure 14B As shown, for a six-parameter affine model, three control points 1410b, 1420b, and 1430b are required. To reduce the computational complexity of the model and the bandwidth of motion compensation, the granularity of affine motion compensation is changed from the sample level to the sub-block level. In the AVS3 standard, affine motion compensation is performed using 4×4 or 8×8 lumen sub-blocks, where each 4×4 or 8×8 sub-block has a motion vector for motion compensation. To derive the motion vector of each 8×8 or 4×4 lumen sub-block, the motion vector at the center position of each sub-block can be calculated based on two or three control points (CPs) and rounded to 1 / 16 fractional precision.
[0130] Figure 15 A schematic diagram showing the motion vector of the center sample of each sub-block of the coded block 1500 according to some embodiments of the present disclosure is illustrated. Specifically, Figure 15 An example of a four-parameter affine model is given, where the motion vector of each sub-block can be derived from the motion vectors MV1 and MV2 of two control points 1510 and 1520. After deriving the sub-block motion vectors, motion compensation is performed to generate a predicted block of the sub-block with the derived motion vectors.
[0131] Affine motion compensation can be performed using two different modes. In affine inter-frame mode, the motion vector difference between control points 1510 and 1520 (i.e., the difference between the CPMV and the CPMV predictor) and the reference index are signaled in the bitstream. On the other hand, in affine skip / direct mode, the motion vector difference and the reference index are not signaled, but are derived by the decoder 300. Specifically, for affine skip / direct mode, the motion vector of the control point (CPMV) of the current block is generated based on the motion information of spatially adjacent blocks. In some embodiments, there are five candidates in the candidate list for affine skip / direct mode. The index is signaled to indicate the candidate to be used for the current block. For example, the candidate list for affine skip / direct mode may sequentially include three types of candidates: received affine candidates, constructed affine candidates, and zero motion vectors.
[0132] For a received affine candidate, the CPMVs of the current block can be extrapolated from the CPMVs of spatially neighboring blocks. At most two received affine candidates can be derived from the affine motion models of neighboring blocks: one from the left neighboring block and one from the top neighboring block. When a neighboring affine block is identified, its CPMV is used to derive the CPMV of the current block. For the constructed affine candidates, the CPMVs of the current block are derived by combining motion information (e.g., MVs) from different neighboring blocks. If the candidate list for the affine skip / direct mode is not full after inserting inherited and constructed affine candidates, zero MVs are inserted until the candidate list is full.
[0133] On the other hand, for affine inter-frame mode, the difference between the CPMV and CPMVP (CPMV predictor) of the current coded block, as well as the index of the CPMV predictor, can be signaled in the bitstream. Encoder 200 is configured to signal an affine flag in the bitstream to indicate whether affine inter-frame mode is used. Further, if affine inter-frame mode is used, another flag is signaled to indicate whether a four-parameter affine model or a six-parameter affine model is used. Encoder 200 and decoder 300 can respectively construct affine CPMVP candidate lists on the encoder and decoder sides. In some embodiments, the affine CPMVP candidate list includes multiple candidates and is constructed by sequentially using the following four types of CPMVP candidates: received affine candidates, constructed affine candidates, translational motion vectors from neighboring blocks, and zero motion vectors. For received affine candidates, CPMVPs are extrapolated from the CPMVs of neighboring blocks. For constructed affine candidates, CPMVPs are derived by combining motion vectors from different neighboring blocks. The encoder 200 can use the index of the CPMVP in the signaling bitstream 228 to indicate which candidate is used as the CPMVP of the current block, and then the decoder 300 can add the MVD of the signaling in the bitstream 228 to the CPMVP to obtain the CPMV of the current block.
[0134] Figure 16 A schematic diagram of integer search points in decoder-side motion vector refinement (DMVR) according to some embodiments of the present disclosure is shown. In some embodiments, decoder 300 may perform DMVR on the decoder side to refine motion vectors according to a symmetric mechanism, such that encoder 200 does not need to explicitly perform MVD in signaling bitstream 228. Specifically, DMVR can only be applied to code blocks of bidirectional predictive coding. After deriving list 0 motion vector MV0 and list 1 motion vector MV1, decoder 300 may further perform a refinement process to refine these two motion vectors MV0 and MV1. In some embodiments, decoder 300 performs DMVR at the 16×16 sub-block level. Before refining motion vectors MV0 and MV1, motion vectors MV0 and MV1 are adjusted to integer precision and set as initial MVs.
[0135] exist Figure 16 In the illustrated embodiment, DMVR is performed based on a search process, where the samples used in the search process are located within a window of size (sub-block width + 7) × (sub-block height + 7), the center point of which is generated by the initial MV reference. Because an 8-tap interpolation filter is used in normal motion compensation, setting the data window to (sub-block width + 7) × (sub-block height + 7) does not increase memory bandwidth. The integer reference samples in the window are taken from reference images in lists 0 and 1. The optimal integer position is set at the location where the sum of the differences between the reference blocks in list 0 and list 1 is minimized.
[0136] like Figure 16 As shown, the triangle filled with a dashed pattern is the initial position 1610 of the initial MVs reference. For each sub-block, 21 integer positions 1620 displayed as triangles are examined to calculate the sum of absolute differences (SAD) between the two reference blocks. Among the 21 positions, the position with the smallest SAD between the two reference blocks is identified as the optimal integer position. After the integer position search, if the optimal integer position falls within... Figure 16 Within the central region 1630 (i.e., the optimal integer position is one of nine center positions), the decoder 300 can further perform sub-pixel estimation based on a mathematical model. In sub-pixel estimation, an error surface is derived based on the SAD values of surrounding integer positions, and the SAD value of the sub-pixel position is calculated based on the error surface. The sub-pixel position with the minimum SAD value can be obtained as a reference position for refinement. Therefore, a reference block at the refinement reference position is obtained as the refined prediction block for the current sub-block. When the refinement reference position is a sub-pixel position, the decoder 300 performs interpolation filtering to derive the sample value at the sub-pixel. Specifically, during the interpolation filtering process, when integer samples outside the search window are needed, samples at the search window boundaries are filled to avoid actually extracting samples outside the search window.
[0137] DMVR can be performed in skip and direct modes to refine motion vectors to improve prediction without enabling flag signaling. In some embodiments, decoder 300 may perform DMVR when the current block meets the following conditions: (1) the current block is a bidirectional prediction block; (2) the current block is encoded in skip or direct mode; (3) the current encoded block does not use affine mode; (4) the current frame is located between two reference frames in display order; (5) the distance between the current frame and the two reference frames is the same; and (6) the width and height of the current block are greater than or equal to 8.
[0138] In some embodiments, bidirectional optical flow (BIO) can be applied to refine the predicted sample values of bidirectional prediction blocks in skip and direct modes. Bidirectional prediction can take a weighted average of two reference blocks to obtain a combined prediction block. In BIO, the combined prediction block can be further refined based on optical flow theory. Specifically, BIO can be applied only to coding blocks using bidirectional prediction coding, calculating gradient values in the horizontal and vertical directions for each sample in the List 0 and List 1 reference blocks. The current coding block can be divided into 16×16 sub-blocks for gradient calculation, just as in BIO. The integer reference samples used for gradient calculation are within a window of (sub-block width + 7) × (sub-block height + 7), which is the same as the DMVR sub-block search window. After calculating the gradient, a refined value for each sample is calculated based on the optical flow equation. To reduce complexity, in the AVS3 standard, the refined value is calculated for clusters with 4×4 samples, rather than at the sample level. The calculated refined value is added to the combined prediction block to obtain the refined prediction block in BIO. In BIO, gradient computation can use an 8-tap filter with integer reference samples as input. Table 3 shows the filter coefficients of the 8-tap gradient filter.
[0139] Table 3: Coefficients of the 8-tap gradient filter
[0140] 0 –4,11,–39,–1,41,–14,8,–2 1 / 4 –2,6,–19,–31,53,–12,7,–2 1 / 2 0,–1,0,–50,50,0,1,0 3 / 4 2,–7,12,–53,31,19,–6,2
[0141] In some embodiments, encoder 200 does not need to indicate the use of BIO by using a signaling enable flag in the bitstream. BIO can be applied to a coding block when the following conditions are met: (1) the coding block is a luminance coding block; (2) the coding block is a bidirectional prediction block; (3) the list 0 reference frame and the list 1 reference frame are located on either side of the current frame in the order of display; and (4) the current motion vector precision is one-quarter of a pixel.
[0142] In some embodiments, bidirectional gradient correction (BGC) can be applied to refine the predicted sample values of bidirectional prediction blocks in inter-frame mode. BGC computes the difference between two reference blocks as a temporal gradient, where one reference block is from a reference image in list 0 and the other is from a reference image in list 1. The computed temporal gradient is then scaled and added to the combined prediction block generated from the two reference blocks to further correct the prediction block. Specifically, the prediction block (denoted as Pred) BI The corrected prediction block Pred can be generated by weighted averaging of reference block Pred0 from reference image 0 and reference block Pred1 from reference image 1. The formula for calculating the corrected prediction block Pred using BGC is as follows:
[0143]
[0144] Here, k is the correction intensity factor, which can be set to 3 in the AVS3 standard. For coded blocks encoded using bidirectional prediction inter-frame mode and meeting the BGC application conditions, the enable flag BgcFlag can be signaled to indicate whether BGC is enabled. When BGC is enabled, the index BgcIdx is further signaled to indicate that a temporal gradient can be used to correct the prediction block. In some embodiments, both the enable flag BgcFlag and the index BgcId can be signaled using context-coded binary. BGC is only applicable to bidirectional prediction mode. For skip and direct modes, the enable flag BgcFlag and the index BgcIdx, along with other motion information, can be received from adjacent blocks.
[0145] In some embodiments, inter-frame prediction filtering (InterPF) is another process provided in the AVS3 standard for refining prediction blocks and can be applied to the final stage of inter-frame prediction. The filtered block obtained after InterPF is the final prediction block. InterPF can only be applied to prediction blocks encoded in normal direct mode. When encoding the current block in normal direct mode, encoder 200 can signal an enable flag to indicate whether InterPF is enabled. If InterPF is enabled, encoder 200 can further signal an index to indicate which filter to apply. Specifically, encoder 200 can select the filter to apply from two filter candidates. Decoder 300 performs the same filtering operation as encoder 200 based on the filter index signaled in the bitstream.
[0146] For example, when applying InterPF, if the InterPF index value is 0, the current predicted sample is filtered by a weighted average of the adjacent reconstructed samples to the left and above. The formula for performing the filtering process is as follows:
[0147] Pred(x,y)=(Pred_inter(x,y)×5+Pred_Q(x,y)×3)>>3
[0148] Pred_Q(x,y)=(Pred_V(x,y)+Pred_H(x,y)+1)>>2
[0149] Pred_Q(x,y)=(Pred_V(x,y)+Pred_H(x,y)+1)>>2
[0150] Pred_H(x,y)=((w-1-x)×Rec(-1,y)+(x+1)×Rec(w,-1)+(w>>1))>>log2(w)
[0151] Here, Pred_inter(x,y) represents the predicted sample to be filtered at position (x,y), and Pred(x,y) represents the predicted sample filtered at position (x,y). Rec(i,j) represents the neighboring pixel reconstructed at position (i,j). The width and height of the current coding block are represented by w and h, respectively.
[0152] If the InterPF index value is equal to 1, the filtering process can be performed according to the following formula:
[0153] Pred(x,y)=(f(x)×Rec(-1,y)+f(y)×Rec(x,-1)+(64-f(x)-f(y))×Pred_inter(x,y)+32)>>6
[0154] Where Pred_inter(x,y) represents the predicted sample to be filtered at position (x,y), and Pred(x,y) represents the predicted sample to be filtered at position (x,y). Rec(i,j) represents the neighboring pixels reconstructed at position (i,j). f(x) and f(y) represent the position-related weights that can be obtained through a lookup table. Table 4 shows the lookup correspondence between the position-related weights f(x) and f(y) of InterPF.
[0155] Table 4: Lookup tables for f(x) and f(x) in InterPF
[0156]
[0157] Figure 17 A schematic diagram illustrating the estimation of local illumination compensation (LIC) model parameters using a reference image and adjacent blocks in a current image, according to some embodiments of this disclosure, is shown. Figure 17 As shown, in inter-frame prediction, if reference block 1712 has the same or similar content as the current block 1722, reference block 1712 in the previously encoded / decoded reference picture 1710 can be found to predict the current block 1722 in the current picture 1720.
[0158] However, illumination variations frequently occur between different images due to changes in lighting conditions, camera position, or object motion. In cases of illumination variation, the values of samples in reference block 1712 and current block 1722 may not be close to each other, even if they contain the same content. This is because reference block 1712 and current block 1722 are located in different images 1710 and 1720 at different times, and therefore can have different illuminance levels. Therefore, to compensate for illumination variations between images in inter-frame prediction, illumination compensation (IC) based on a linear model can be applied to generate compensated prediction blocks with illumination level values closer to the current block. Encoder 200 and decoder 300 can derive two parameters of the linear model, namely the scaling factor parameter a and the offset parameter b, and apply them to the prediction blocks as follows:
[0159] y = a × x + b
[0160] Where x represents the predicted sample from the prediction block of reference image 1710, and y represents the predicted sample after illumination compensation.
[0161] In some embodiments, the illumination compensation model can be derived at the image level and applied to all coded blocks within an image. In some other embodiments, the illumination compensation model can be derived at the coded block level and applied only to specific coded blocks. Coded block-level illumination compensation is called Local Illumination Compensation (LIC). When LIC is performed, encoder 200 and decoder 300 derive model parameters for the current coded block 1722 in the same manner. Therefore, encoder 200 does not need the model parameters in signaling bitstream 228. Figure 17 In this embodiment, encoder 200 and decoder 300 can derive model parameters using reconstructed samples from neighboring block 1724 and predicted samples from neighboring block 1714 (shown in shaded areas). Specifically, linear model parameters are first estimated based on the relationship between the predicted sample values and reconstructed sample values of neighboring blocks 1714 and 1724. Then, the estimated linear model is applied to the predicted samples to generate illumination-compensated predicted samples 1722 with values closer to the original sample values of the current coded block. Figure 17 In this embodiment, the predicted samples of adjacent block 1714 can be used for parameter derivation. Therefore, the decoder 300 extracts a coded block larger than the current block 1722 from the reference image buffer, and this operation also increases bandwidth.
[0162] Figure 18 A schematic diagram illustrating LIC model parameter estimation according to some embodiments of the present disclosure is shown. Figure 18 In one embodiment, to reduce bandwidth, local illumination compensation based on the current prediction block (CPB) is proposed. For example... Figure 18As shown, instead of adjacent block samples, predicted samples on the left and top boundaries within predicted block 1812 are used to estimate model parameters. Figure 18 As shown in the shaded area 1814 in the image, the predicted samples within the left and upper boundaries of the predicted block 1812 and the reconstructed samples of the adjacent block 1824 of the current encoded block 1822 in the current image 1820 are used to derive the model parameters.
[0163] In some embodiments, the encoder 200 or decoder 300 may apply the least squares method to estimate the model parameters, but its computational complexity is very high. Figures 19A-19D A schematic diagram illustrating LIC model parameter estimation using four pairs of samples according to some embodiments of this disclosure is shown. Figures 19A-19D In this embodiment, to simplify parameter estimation, a four-point estimation method is used, that is, parameter estimation is performed using four pairs of samples. For example... Figure 19A As shown, four predicted samples 1912, 1914, 1916, and 1918 on the top boundary within prediction block 1910 and four reconstructed samples 1932, 1934, 1936, and 1938 on the top adjacent block of the current coding block 1920 were used. Figure 19B As shown, four prediction samples 1912, 1914, 1916, and 1918 on the left boundary within prediction block 1910 and four reconstructed samples 1932, 1934, 1936, and 1938 on the left adjacent block of the current block 1920 were used. Figure 19C and Figure 19D In the estimation process, two predicted samples 1912 and 1914 on the top boundary and two predicted samples 1916 and 1918 on the left boundary of the current coding block 1910 are used, along with two reconstructed samples 1932 and 1934 on the top adjacent block of the current coding block 1920 and two reconstructed samples 1936 and 1938 on the left adjacent block of the current coding block 1920. First, the predicted samples 1912-1918 can be sorted according to their values. The average of the two larger values among the predicted samples 1912-1918 is calculated and denoted as x_max. Then, the average of the two corresponding reconstructed sample values is calculated and denoted as y_max. Specifically, the reconstructed samples 1932, 1934, 1936, and 1938 correspond to the predicted samples 1912, 1914, 1916, and 1918, respectively. Calculate the average of the two smaller values in the predicted samples 1912-1918, denoted as x_min. Then calculate the average of the two corresponding reconstructed sample values, denoted as y_min. The model parameters a and b of the linear model are derived using the following equation based on the values of x_max, y_max, x_min, and y_min:
[0164] a=((y_max–y_min)×(1<<shift) / (x_max-x_min))>>shift
[0165] b = y_min - a × x_min
[0166] The parameter "shift" is the number of bits to shift, and the operator " / " represents integer division.
[0167] In the above Figure 19C-19D In the CPB-based LIC design, the prediction sample 1914 in the upper right corner and the prediction sample 1918 in the lower left corner of prediction block 1910 are needed for parameter derivation. However, in actual hardware implementations, the coded block is usually divided into sub-blocks, and inter-frame prediction is performed at the sub-block level.
[0168] Figure 20 A schematic diagram of sub-block-level inter-frame prediction according to some embodiments of the present disclosure is shown. For example... Figure 20 As shown, a 64×64 coded block 2000 is divided into 16 16×16 sub-blocks 2011-2044. When the decoder 300 performs inter-frame prediction on sub-block 2011, the LIC model parameter estimates can be used for the prediction samples of sub-blocks 2012-2014 and sub-blocks 2021, 2031, and 2041. In other words, the LIC cannot refine the prediction samples of sub-block 2011 between the completion of inter-frame prediction for the other sub-blocks 2012-2014, 2021, 2031, and 2041. Therefore, the storage space required to store the unrefined prediction samples increases, and pipeline latency also increases.
[0169] To address the aforementioned issues, in this disclosure, encoder 200 and decoder 300 can perform a simplified LIC. In the simplified LIC, the samples used for LIC model parameter derivation are limited based on the sample location, thereby reducing the required memory and also reducing pipeline latency. Figures 21A-21C A schematic diagram illustrating a sample of derived LIC model parameters according to some embodiments of this disclosure is shown. Figures 21A-21C In some embodiments, the location of samples used for parameter derivation is restricted to reduce memory usage for storing predicted samples and processing wait time for CPB-based LIC. For example, as Figures 21A-21C As shown, only samples around the top left corner of the encoding block are used to derive the model parameters for the current encoding block. Figure 21A In the embodiment, for the predicted samples, only the first K1 samples 2110a outside the top boundary of the prediction block 2100a and the first K2 samples 2120a outside the left boundary of the prediction block 2100a are used (in Figure 21A (Using shaded samples as an example) to derive model parameters. Optionally, in Figure 21BIn the embodiment, for the predicted samples, only the first K1 samples 2110b within the top boundary of the prediction block 2100b and the first K2 samples 2120b within the left boundary of the prediction block 2100b are used (in Figure 21B (Using shaded samples as an example) to derive model parameters. Figure 21C In the embodiment, for the reconstructed samples, only the first K1 samples 2110c in the top adjacent block adjacent to the top boundary of the current coding block 2100c and the first K2 samples 2120c in the left adjacent block adjacent to the left boundary of the current coding block 2100c are used (in Figure 21C (Shown as shaded samples) to derive model parameters. Limiting values K1 and K2 are two integers used to restrict the positions of samples used in parameter derivation. For example, limiting values K1 and K2 can be 16 or 8, without specific limitation in this disclosure. In various embodiments, encoder 200 and decoder 300 may apply least squares estimation or four-point estimation or other existing parameter estimation methods to derive LIC model parameters.
[0170] In some embodiments, the constraint values K1 and K2 are variable, depending on the type of the selected inter-frame prediction mode or motion vector predictor candidate. For example, for UMVE and AWP modes, inter-frame prediction (or "motion compensation") is performed at the coded block level. For ETMVP and affine modes, inter-frame prediction is performed at the sub-block level. For normal skip and direct modes, inter-frame prediction is performed at the coded block level for block-level TMVP, SMVP, and HMVP as coded block-level candidates, and at the sub-block level for sub-block-level TMVP and MVAP as sub-block-level candidates. For inter-frame prediction modes that perform inter-frame prediction at the sub-block level (e.g., ETMVP and affine modes) or MVP candidates for normal skip and direct modes (e.g., sub-block-level TMVP and MVAP), the values of K1 and K2 can be set to the size of the sub-block. On the other hand, for inter-frame prediction modes that perform inter-frame prediction at the coding block level (e.g., UMVE mode and AWP mode) or MVP candidates for normal skip and direct modes (e.g., block-level TMVP, SMVP, and HMVP), the values of K1 and K2 can be set to preset values, which may be different from the values for coding blocks that perform inter-frame prediction at the sub-block level.
[0171] For example, inter-frame prediction can be performed at the 8×8 sub-block level for sub-block-level TMVP candidates in normal skip and direct modes, MVAP candidates in normal skip and direct modes, and ETMVP mode. Therefore, the values of K1 and K2 can be set to 8. For affine mode, whether inter-frame prediction is performed at the 4×4 or 8×8 sub-block level depends on the image level flag. Therefore, the values of K1 and K2 can be set to 4 or 8 according to the image level flag of the signaling. For UMVE mode, AWP mode, block-level TMVP candidates in normal skip and direct modes, and SMVP and HMVP candidates in normal skip and direct modes, inter-frame prediction can be performed on the coded blocks, and the values of K1 and K2 can both be set to 16 or other preset values. In some embodiments, if the values of K1 and K2 are smaller than the size of the coded block, the position constraint can be changed to the size of the coded block. In some embodiments, to improve coding efficiency, the predefined values need to be larger than the size of the sub-block. Therefore, in normal skip and direct modes, the restrictions on the sample locations used to derive model parameters for sub-block level TMVP candidates and MVAP candidates are more stringent (e.g., smaller) than the restrictions on the sample locations used to derive model parameters for block level TMVP candidates, SMVP candidates, or HMVP candidates.
[0172] In some embodiments, to further reduce implementation costs, encoder 200 and decoder 300 may perform inter-frame prediction only at the coding block level and apply LIC to coding units. LIC is not applied to MVP candidates in prediction modes or normal skip and direct modes that perform inter-frame prediction at the sub-block level. For example, in some embodiments, for normal skip and direct modes, if a sub-block level TMVP or MVAP candidate is selected for a coding unit, LIC is not applied to the coding unit. LIC is applied to coding units using other candidates in normal skip and direct modes. As another example, in some embodiments, LIC is not applied in affine and ETMVP modes because inter-frame prediction can be performed on sub-blocks in affine and ETMVP modes. Since MVP candidates in normal skip and direct modes are indicated by candidate indices, in some embodiments, to simplify candidate type determination, encoder 200 or decoder 300 may directly check the candidate indices to determine whether LIC is enabled. For example, in the motion candidate list for normal skip and direct modes, if sub-block TMVP is enabled, the candidate with index 0 is a sub-block TMVP candidate. Therefore, if the skip or direct index is signaled as 0, LIC can be disabled, and there are no signaling LIC-related syntax elements for the coded block. For example, in the motion candidate list for normal skip and direct modes, candidates with indices equal to 3, 4, 5, 6, or 7 can be MVAP candidates or HMVP candidates, depending on the number of MVAP candidates. Applying a conventional approach, LIC can be directly disabled for candidates with indices from 3 to 7, regardless of the candidate type. Therefore, if the skip or direct index is signaled as 3, 4, 5, 6, or 7 for the coded block, LIC can be directly disabled, and there are no signaling LIC syntax elements for that coded block.
[0173] like Figures 21A-21C As shown, in some embodiments, model parameters are derived only based on samples around the top left corner of the current coding block (e.g., the first K1 samples on the top boundary and the first K2 samples on the left boundary) and applied to all samples within the current coding block. In some embodiments, to further improve the accuracy of model parameters for samples in the bottom or right portion of the coding block, encoder 200 or decoder 300 may perform sub-block level parameter derivation.
[0174] Figures 22A-22C A schematic diagram illustrating a sample of LIC model parameters derived at the sub-block level according to some embodiments of this disclosure is shown. Figures 22A-22C In this embodiment, the model parameters for each sub-block can be derived based on samples at the corresponding top and left boundaries of the current sub-block. Therefore, dependencies between different sub-blocks 2011-2044 are removed. For example, as... Figure 22AAs shown, for sub-block 2022 of prediction block 2200a, prediction sample 2210a outside the top boundary of sub-block 2012 of prediction block 2200a and prediction sample 2220a outside the left boundary of sub-block 2021 of prediction block 2200a are used to derive model parameters. Optionally, as Figure 22B As shown, within prediction block 2200b, prediction samples 2210b on the top boundary of sub-block 2012 and prediction samples 2220b on the left boundary of sub-block 2021 are used to derive model parameters. Figure 22C As shown, the model parameters are derived using the reconstruction sample 2210c of the adjacent block that is adjacent to the top boundary of the sub-block 2012 of the current block 2200c and the reconstruction sample 2220c of the adjacent block that is adjacent to the left boundary of the sub-block 2021 of the current coding block 2200c.
[0175] Similarly, for sub-block 2044, in Figure 22A In the embodiment, prediction samples 2230a outside the top boundary of sub-block 2014 of prediction block 2200a and prediction samples 2240a outside the left boundary of sub-block 2041 of prediction block 2200a are used to derive model parameters. Figure 22B In the embodiment, prediction samples 2230b on the top boundary of sub-block 2014 within prediction block 2200b and prediction samples 2240b on the left boundary of sub-block 2041 within prediction block 2200b are both used to derive model parameters. Figure 22C In this embodiment, reconstruction sample 2230c of the adjacent block adjacent to the top boundary of sub-block 2014 of the current block 2200c and reconstruction sample 2240c of the adjacent block adjacent to the left boundary of sub-block 2041 of the current block 2200c are used to derive model parameters. It should be noted that... Figure 17 and Figure 18 The parameter derivation method described in [the document] can be applied, and for the sake of brevity, it will not be elaborated further in this paper. By using samples on the boundaries of the corresponding coding blocks to derive the parameters of each sub-block, the accuracy of the parameters can be improved.
[0176] Figures 23A-23C A schematic diagram illustrating a method for deriving LIC model parameter samples according to some embodiments of the present disclosure is shown. Figures 23A-23C In this embodiment, consistency between parameters of different sub-blocks can be improved, and potential performance degradation due to inconsistency can be avoided. In other words, the model parameters of the current sub-block can be derived based on samples from the boundary of the first sub-block to the corresponding boundary associated with the current sub-block. Figures 23A-23C As shown, when driving the current sub-block (e.g., Figures 23A-23CWhen the LIC model parameters for sub-block 2023 are sub-block level, the corresponding top boundary is the top boundary of sub-block 2013, and the corresponding left boundary is the left boundary of sub-block 2021.
[0177] Therefore, as Figure 23A As shown, prediction samples 2310a outside the top boundary of sub-blocks 2011, 2012, and 2013 of prediction block 2300a and prediction samples 2320a outside the left boundary of sub-blocks 2011 and 2021 of prediction block 2300a are used to derive model parameters. Alternatively, as... Figure 23B As shown, prediction samples 2310b on the top boundaries of sub-blocks 2011, 2012, and 2013 within prediction block 2300b and prediction samples 2320b on the left boundaries of sub-blocks 2011 and 2021 within prediction block 2300b are used to derive model parameters. Figure 23C As shown, the reconstruction sample 2310c of the adjacent block that is adjacent to the top boundary of the sub-blocks 2011, 2012 and 2013 of the current block 2300c and the reconstruction sample 2320c of the adjacent block that is adjacent to the left boundary of the sub-blocks 2011 and 2021 of the current block 2300c are used to derive the model parameters.
[0178] In some embodiments, LIC may be applied only to the luminance component to compensate for illuminance variations, but this disclosure is not specifically limited thereto. For example, in some embodiments, the chrominance component of an image may also vary in response to changes in illuminance. Therefore, if LIC is applied only to the luminance component, chrominance differences may exist. Consistent with some embodiments of this disclosure, LIC may be applied to the chrominance component to compensate for chrominance variations between the current coding block and the prediction block. Optionally, in embodiments of this disclosure, illuminance compensation may be extended to the chrominance component, and chrominance component illuminance compensation may be referred to as Local Chromaticity Compensation (LCC).
[0179] In LCC, adjacent predicted chroma samples and adjacent reconstructed chroma samples can be used to derive a linear model, which is then applied to the predicted chroma samples of the current coding block to produce compensated predicted chroma samples. For LCC applied to the chroma component, various LIC methods described in the paragraphs above can be employed. Considering that chroma textures are simpler than luma textures, the linear model of LCC can be simplified by removing the scaling factor 'a' and keeping only the offset 'b'. That is, when LCC is applied to the predicted chroma block, the values of the predicted chroma samples in the current coding block are directly set to 'b', without the need for multiplication and addition to calculate the compensated predicted sample values.
[0180] In some embodiments, enabled LIC (i.e., Luminance Component IC) is a prerequisite for enabling LCC (i.e., Chroma Component IC). In other words, in one example, when LIC is enabled and applied to the same coding block, LCC can only be enabled and applied to that coding block. In some embodiments, LIC compensation applied to the luminance coding block and LCC compensation applied to the chroma coding block can be enabled or disabled separately in the coding unit. Specifically, in some embodiments, when LCC is enabled for the current coding block, encoder 200 and decoder 300 further examine adjacent reconstructed chroma samples and predicted chroma samples used to derive model parameters to determine whether simplified LCC (i.e., LCC with only offset b) should be applied.
[0181] In some embodiments, within the coding unit, for the luma component, the positional constraints on the samples used for model parameter derivation depend on the type of inter-frame prediction mode or MVP candidate. On the other hand, for the chroma component, if the chroma coded block is encoded using sub-block-level inter-frame prediction, LCC can be disabled.
[0182] For example, if sub-block-level TMVP or MVAP candidates from normal skip mode, direct mode, affine mode, or ETMVP mode are used for the coding unit, encoder 200 or decoder 300 can enable LIC and set the sampling position limits K1 and K2 of the luma coding block of the coding unit to 8. In other words, the left 8 adjacent reconstructed samples on the top boundary of the luma coding block and the top 8 adjacent reconstructed samples on the left boundary of the luma coding block, as well as the left 8 predicted samples within the top boundary of the luma coding block and the top 8 predicted samples within the left boundary of the luma coding block, are all used to derive the LIC model parameters. Therefore, LIC can be enabled and applied to the luma coding block of the coding unit. Simultaneously, LCC can be disabled and not applied to the chroma coding block of the coding unit.
[0183] Similar to the operation in LIC, when LCC is applied, in some embodiments, encoder 200 or decoder 300 can use predicted chroma samples within the current block instead of predicted chroma samples from adjacent blocks to derive model parameters to reduce bandwidth.
[0184] Figure 24 A flowchart of a video processing method 2400 according to some embodiments of the present disclosure is shown. In some embodiments, the video processing method 2400 may be generated by an encoder (e.g., Figure 2 The encoder 200 or decoder (e.g., in the encoder 200) or decoder (e.g., Figure 3 Decoder 300 in the middle) performs this to perform the inter-frame prediction phase (e.g., Figure 2-3 Inter-frame prediction is performed on the luma-coded blocks in the inter-frame prediction stage 2044 of the encoder. For example, the encoder is a device (e.g., Figure 4 The device 400 in the middle can be implemented as at least one software or hardware component for processing video sequences (e.g., Figure 2 The video sequence 202 in the video is encoded or transcoded to encode or transcode the bitstream of a video frame or video sequence including at least one CU (e.g., Figure 2 The video bitstream (228) in the video stream is encoded or decoded. Similarly, the decoder is a device (e.g., Figure 4 The device 400 in the middle can be implemented as at least one software or hardware component for decoding the bitstream (e.g., Figure 3 The video bitstream 228 in the video stream is used to reconstruct the video frames or video sequence of the bitstream (e.g., Figure 3 The video stream in the video stream returned a 304 error. The processor (e.g., ...) Figure 4 The processor 402 in the middle can execute video processing method 2400.
[0185] Referring to video processing method 2400, in step 2410, the device determines whether to enable inter-frame predictor correction (e.g., local luminance compensation) for the coded blocks of the inter-frame prediction process of the encoded or decoded bitstream. Specifically, according to some embodiments of this disclosure, when LIC / LCC is applied in the AVS3 standard, the interaction between LIC / LCC technology in the AVS3 standard and other inter-frame prediction refinement techniques is considered. In some embodiments, in step 2410, if any of affine motion compensation, ETMVP, AWP, inter-frame prediction filter, bidirectional gradient correction, bidirectional optical flow, or decoder-side motion vector refinement is enabled for the coded block, the device may disable local luminance compensation.
[0186] For example, in some embodiments, LIC / LCC is not applied together with InterPF to avoid increasing pipeline stages. Therefore, for a single code block, at most one of InterPF and LIC / LCC is applied. When InterPF is enabled for a code block, encoder 200 can skip the signaling for the LIC / LCC enable flag because decoder 300 can infer that LIC / LCC is disabled. Alternatively, if LIC / LCC is enabled for a code block, encoder 200 can skip the signaling for the InterPF enable flag and the InterPF index because InterPF cannot be enabled for that code block at this time.
[0187] When LIC / LCC is applied to a coding block along with DMVR, BIO, and BGC, the motion vectors are first refined by DMVR. Then, BIO and BGC are applied to the prediction block of the current coding block to obtain a refined prediction block. Finally, LIC / LCC is used to compensate the refined prediction block to obtain the final prediction block.
[0188] In some embodiments, since BGC uses temporal gradients to refine prediction samples, LIC / LCC is not applied together with BGC. In some embodiments, for skip mode and direct mode, if the BGC enable flag is explicitly signaled, the signaling of the LIC / LCC flag is skipped and its signaling is inferred to be false when the signaled BGC enable flag is true. Optionally, when the LIC / LCC enable flag is signaled true, encoder 200 can skip the BGC enable flag, and decoder 300 can infer that the BGC enable flag is false. If the BGC enable flag is not received via signaling but from skip candidates and direct candidates, the BGC flag is set to false when LIC / LCC is enabled, regardless of the BGC flags of skip candidates and direct candidates. In the above embodiments, when LIC / LCC is applied to an encoded block together with DMVR and BIO, the motion vector is first refined by DMVR. Then, BIO is applied to the prediction block of the current block to obtain a refined prediction block. Then, LIC / LCC is used to compensate the refined prediction block to obtain the final prediction block.
[0189] In some embodiments, to further reduce computational complexity, LIC / LCC is not applied with BIO and is only applied with DMVR when DMVR refines motion vectors without directly refining prediction values. Therefore, encoder 200 can skip the signaling of the LIC / LCC enable flag if the BGC enable flag or InterPF enable flag is signaled true. Optionally, when LIC / LCC is enabled, encoder 200 can skip the signaling of the BGC enable flag or InterPF enable flag, and decoder 300 can infer that the BGC enable flag or InterPF enable flag is false when the LIC / LCC enable flag is signaled equal to true. Therefore, when LIC / LCC is applied to an encoded block with DMVR, the motion vectors are first refined by DMVR. Then, LIC / LCC is applied to the prediction block of the current encoded block to obtain the final prediction block.
[0190] In some embodiments, LIC is not applied in conjunction with any of DMVR, BIO, BGC, and InterPF. Similarly, encoder 200 may skip the signaling of the LIC / LCC enable flag if the BGC enable flag or the InterPF enable flag is signaled as true. Optionally, encoder 200 may skip the signaling of the BGC enable flag or the InterPF enable flag, and decoder 300 may infer that the BGC enable flag and the InterPF enable flag are false when the LIC enable flag is signaled as true. In other words, if local luma compensation or local chroma compensation is enabled for the encoded block, BGC or DMVR may be disabled.
[0191] In some embodiments, encoder 200 or decoder 300 may apply LIC / LCC only to code blocks encoded in normal skip and direct modes. In other embodiments, encoder 200 or decoder 300 may apply LIC / LCC to code blocks encoded in normal skip and direct modes, or to code blocks encoded in other skip and direct modes, such as the AWP mode, ETMVP mode, and affine mode described above. Optionally, LIC may be performed as a final stage in generating the final predicted sample.
[0192] In some embodiments, encoder 200 or decoder 300 applies LIC / LCC to code blocks encoded in normal skip mode and direct mode, AWP mode and ETMVP mode, but does not apply LIC / LCC to code blocks encoded in affine skip mode and direct mode, because the prediction samples for affine mode have been refined by the sub-block MV-based technique in the AVS3 standard.
[0193] In some embodiments, in step 2410, if the coded block is encoded in affine mode or ETMVP mode, the apparatus may disable inter-frame predictor correction (e.g., local luma compensation or local chroma compensation) to reduce encoder complexity. That is, encoder 200 or decoder 300 applies LIC / LCC to coded blocks encoded in normal skip and direct modes, as well as AWP mode. For coded blocks encoded in affine skip and direct modes, as well as ETMVP mode, LIC / LCC is not applied. Therefore, encoder 200 does not need to determine whether to enable LIC on coded blocks in affine skip and direct modes, as well as ETMVP mode, while encoder complexity is also reduced.
[0194] In some embodiments, in addition to skip mode and direct mode, LIC / LCC can be further extended to inter-frame mode, where motion vector difference and prediction residual can be signaled in the bitstream. LIC / LCC can be applied to prediction blocks in inter-frame mode, and the compensated prediction samples generated by LIC / LCC can be used as the final predictor. Therefore, reconstructed samples can be generated by adding the reconstruction residual to the final predictor. In some embodiments, the reconstructed samples can be further processed by other loop filters, such as deblocking filters, sample adaptive offsets, or adaptive loop filters.
[0195] In skip mode, encoder 200 can skip signaling for prediction residuals, thus exhibiting less signaling overhead compared to direct mode. Since LIC / LCC may increase signaling costs, in some embodiments, LIC / LCC is applied only to code blocks encoded in direct mode and not to code blocks encoded in skip mode to reduce signaling costs. In some embodiments, direct mode may include normal direct mode, AWP direct mode, ETMVP direct mode, and affine direct mode. Skip mode may include normal skip mode, AWP skip mode, ETMVP skip mode, and affine skip mode.
[0196] In some embodiments, in step 2410, the apparatus may enable an inter-frame predictor for correction (e.g., LIC / LCC) in response to both the sequence header enable flag and the picture header enable flag associated with the coding block being true for coding blocks encoded in skip mode, and enable LIC in response to the sequence header enable flag associated with the coding block being true for coding blocks encoded in direct mode.
[0197] Specifically, if there is a significant difference in illumination between the current image and the reference image, LIC / LCC is enabled for skip mode. Therefore, encoder 200 can signal a flag in the image header to indicate whether LIC / LCC can be applied to the encoded blocks of the current image encoded in skip mode. If encoder 200 signals the LIC / LCC enable flag to be true in the sequence header, encoder 200 can also signal the image header LIC / LCC enable flag in the image header. Optionally, the image header LIC / LCC enable flag is not signaled in the image header, but is inferred to be false. Therefore, if the image header LIC / LCC enable flag is true, LIC / LCC can be applied to the encoded blocks of the current image encoded in skip mode. Otherwise, LIC / LCC cannot be applied to the encoded blocks of the current image encoded in skip mode. In some embodiments, for direct mode, encoder 200 does not signal the flag in the image header, and if the LIC / LCC sequence header flag is signaled to be true, LIC / LCC can be applied to the encoded blocks encoded in direct mode.
[0198] When inter-frame predictor correction is enabled for a coded block (step 2410 - Yes), the apparatus performs inter-frame predictor correction in steps 2422-2428. Specifically, in step 2422, the apparatus determines sample position constraint values K1 and K2, where K1 is a horizontal position constraint value and K2 is a vertical-horizontal position constraint value. In some embodiments, the value of K1 may be equal to the value of K2. The sample position constraint values specify the range of predicted samples and reconstructed samples used for the derivation of local brightness compensation model parameters. Therefore, only the brightness prediction block corresponding to the coded block (e.g., Figure 21A or Figure 21B The first part of the first boundary (e.g., the top boundary) of the prediction block 2100a or 2100b in the prediction block (e.g., Figure 21A or Figure 21B The predicted brightness samples adjacent to the first K1 samples (2110a or 2110b) or the second part of the second boundary (e.g., the left boundary). Figure 21A or Figure 21B The predicted brightness samples adjacent to the first K2 samples (2120a or 2120b) in the code can be used for model parameter derivation. It is worth noting that only the predicted brightness samples adjacent to the encoded block (e.g., Figure 21C The first part of the first boundary of block 2100c) (e.g., Figure 21C The reconstructed brightness samples adjacent to the first K1 samples (2110c) or the second part of the second boundary (e.g., Figure 21C The reconstructed brightness samples adjacent to the first K2 samples (2120c) can be used in the derivation of model parameters.
[0199] In some embodiments, the apparatus determines limit values K1 and K2 based on the inter-frame prediction mode or motion vector prediction candidate type to be applied in the inter-frame prediction process, and determines the size of the first portion and the second portion based on the first limit value and the second limit value, respectively.
[0200] In some embodiments, the values of the horizontal position parameter K1 and the vertical position parameter K2 can be equal to 8 or 16. For example, if the coding mode of the coded block indicates that the lumen prediction block is obtained at the sub-block level, the apparatus can determine that the values of the horizontal position parameter K1 and the vertical position parameter K2 are equal to 8. Otherwise, the apparatus can determine that the values of the horizontal position parameter K1 and the vertical position parameter K2 are both equal to 16.
[0201] In step 2424, the device acquires predicted brightness samples and reconstructed brightness samples. Specifically, as described above... Figure 21A and Figure 21B As described, the device can select a portion of samples as predicted luminance samples from the first m luminance samples adjacent to and along the top boundary of the luminance prediction block and from the first n luminance samples adjacent to and along the left boundary of the luminance prediction block, where m is less than the width of the luminance prediction block and n is less than the height of the luminance prediction block. Figure 21A In the image, the predicted brightness sample is outside the brightness prediction block, but adjacent to it. Figure 21B In the image, the predicted brightness sample is located inside the brightness prediction block.
[0202] Specifically, such as Figure 21CAs shown, the device can select a portion of the samples as reconstructed luminance samples from the first m luminance samples adjacent to and along the top boundary of the coding block and from the first n luminance samples adjacent to and along the left boundary of the coding block, where m is less than the width of the coding block and n is less than the height of the luminance component coding block. Reconstructed luminance samples 2110c and 2120c are outside of, but adjacent to, coding block 2100c.
[0203] In some embodiments, if the sampling position constraint value K1 or K2 is greater than the width or height of the coding block, then the constraint value K1 or K2 is limited to the width or height of the coding block. Optionally, the device derives m as the smaller of the constraint value K1 and the width of the coding block, and derives n as the smaller of the constraint value K2 and the height of the coding block. Therefore, if the width of the coding block is greater than the horizontal position parameter value, the device can determine the limiting value of the horizontal position parameter as the value of the horizontal position parameter. Otherwise, the device determines the limiting value of the horizontal position parameter as the width value of the coding block. Similarly, if the height of the coding block is greater than the vertical position parameter value, the device determines the limiting value of the vertical position parameter as the value of the vertical position parameter. Otherwise, the device determines the limiting value of the vertical position parameter as the height value of the coding block. Therefore, predicted luminance samples and reconstructed luminance samples are obtained based on the limiting values of the horizontal and vertical position parameters.
[0204] In step 2426, the device derives model parameters for inter-frame predictor correction (e.g., local brightness compensation) based on the obtained predicted brightness samples and reconstructed brightness samples. In step 2428, the device applies inter-frame predictor correction to obtain the corrected brightness prediction block based on at least one model parameter and the brightness prediction block used for the inter-frame prediction process.
[0205] Figure 25 A flowchart illustrating a video processing method according to some embodiments of the present disclosure is shown. Similar to... Figure 24 The video processing method 2400 and video processing method 2500 shown can be generated by an encoder (e.g., Figure 2 The encoder 200 or decoder (e.g., in the encoder 200) or decoder (e.g., Figure 3 Decoder 300 in the middle) performs this to perform the inter-frame prediction phase (e.g., Figure 2-3 Inter-frame prediction is performed on chroma-coded blocks in the inter-frame prediction phase 2044. For example, the processor (e.g., Figure 4 The processor 402 in the middle can execute video processing method 2500.
[0206] Referring to video processing method 2500, in step 2510, the apparatus determines whether to enable inter-frame predictor correction (e.g., local chroma compensation (LCC)) for the coded block of the inter-frame prediction process. When inter-frame predictor correction is enabled for the coded block (step 2510 - Yes), the apparatus performs inter-frame predictor correction in steps 2522-2528. Similar to the luminance compensation in steps 2422-2428 of video processing method 2400, in the chroma compensation in steps 2522-2528, in step 2522, the apparatus may determine position constraint values K1 and K2, where K1 is a horizontal position constraint value and K2 is a vertical-horizontal position constraint value. In some embodiments, the value of K1 may be equal to the value of K2. The sample position constraint values specify the range of predicted samples and reconstructed samples used in local chroma compensation. Therefore, only the chroma prediction block corresponding to the coded block (e.g., Figure 21A or Figure 21B The third part of the first boundary (e.g., the top boundary) of the prediction block 2100a or 2100b in the prediction block (e.g., Figure 21A or Figure 21B The predicted chromaticity samples adjacent to the first K1 samples (2110a or 2110b) or the fourth part (e.g., the second boundary, the left boundary) of the second boundary (e.g., the left boundary). Figure 21A or Figure 21B The predicted chromaticity samples adjacent to the first K2 samples (2120a or 2120b) in the code can be used in the model parameter derivation. It is worth noting that only those adjacent to the encoded block (e.g., Figure 21C The third part of the first boundary of block 2100c) (e.g., Figure 21C The reconstructed chroma samples adjacent to the first K1 samples (2110c) or the fourth part of the second boundary of the coded block (e.g., Figure 21C The reconstructed chromaticity samples adjacent to the first K2 samples (2120c) can be used in the derivation of model parameters.
[0207] In some embodiments, the values of the horizontal position parameter K1 and the vertical position parameter K2 can be equal to 4 or 8. For example, if the coding mode of the coded block indicates that the chroma prediction block is obtained at the sub-block level, the apparatus can determine that the values of the horizontal position parameter and the vertical position parameter are equal to 4. Otherwise, the apparatus can determine that the values of both the horizontal position parameter K1 and the vertical position parameter K2 are equal to 8.
[0208] In some embodiments, the sizes of the first and second portions used for luminance compensation, and the sizes of the third and fourth portions used for chrominance components, may be different. In other words, the horizontal position constraint value K1 of the chrominance component may be different from the horizontal position constraint value K1 of the luminance component, and the vertical position constraint value K2 of the chrominance component may be different from the vertical position constraint value K1 of the luminance component. For example, the constraint values K1 and K2 of the chrominance component may be half of the constraint values K1 and K2 of the luminance component.
[0209] In step 2524, the device can acquire predicted chromaticity samples and reconstructed chromaticity samples. Specifically, as described above, the device can select a subset of samples as predicted chromaticity samples from the first k chromaticity samples adjacent to and along the top boundary of the chromaticity prediction block and from the first p chromaticity samples adjacent to and along the left boundary of the chromaticity prediction block, where k is less than the width of the chromaticity prediction block and p is less than the height of the chromaticity prediction block. Figure 21A In the image, the predicted chromaticity sample is outside the chromaticity prediction block, but adjacent to it. Figure 21B In this process, the predicted chromaticity samples can be located inside the chromaticity prediction block.
[0210] Specifically, such as Figure 21C As shown, the device can select a subset of samples as reconstructed chroma samples from the first k chroma samples adjacent to and along the top boundary of the coding block, and from the first p chroma samples adjacent to and along the left boundary of the coding block, where k is less than the width of the coding block and p is less than the height of the coding block. Reconstructed chroma samples 2110c and 2120c are outside of, but adjacent to, coding block 2100c.
[0211] In some embodiments, if the sample position constraint value K1 or K2 is greater than the width or height of the coded block, then the constraint value K1 or K2 is limited to the width or height of the coded block. Alternatively, the device derives k as the smaller of the constraint value K1 and the width of the coded block, and derives p as the smaller of the constraint value K2 and the height of the coded block. Therefore, if the width value of the coded block is greater than the value of the horizontal position parameter, the device determines the limiting value of the horizontal position parameter as the value of the horizontal position parameter. Otherwise, the device determines the limiting value of the horizontal position parameter as the width value of the coded block. Similarly, if the height value of the coded block is greater than the value of the vertical position parameter, the device determines the limiting value of the vertical position parameter as the parameter value of the vertical position. Otherwise, the device determines the limiting value of the vertical position parameter as the height value of the coded block. Therefore, predicted chroma samples and reconstructed chroma samples can be obtained based on the limiting values of the horizontal and vertical position parameters.
[0212] Specifically, in some embodiments, to reduce memory and latency in chroma compensation, methods such as... Figures 21A-21CThe sample location constraints applied to the Cb and Cr components are shown, or as... Figures 22A-22C and Figures 23A-23C The illustrated sub-block-based parameter derivation method derives model parameters for the Cb and Cr components. In one example, the constraints on the sample locations used to derive the parameters depend on the color format of the video sequence. For example, in the YUV 4:4:4 format, where the three components (i.e., the luma component Y and the two chroma components U and V) have the same sampling rate, the location constraints K1 and K2 are the same for all three components.
[0213] On the other hand, in the YUV 4:2:2 format, the sampling rates of the two chrominance components U and V are half the sampling rate of the luma component Y in the horizontal dimension, and the horizontal position constraint K1 of the chrominance components is also half the horizontal position constraint K1 of the luma component. For example, if the predicted luma sample and reconstructed luma sample from the 16 samples at the top boundary and the 16 samples at the left boundary are used to derive the parameters for the luma coding block, then the predicted chrominance sample and reconstructed chrominance sample from the 8 samples at the top boundary and the 16 samples at the left boundary will be used to derive the parameters for the chrominance coding block.
[0214] Similarly, in the YUV 4:2:0 format, the sampling rates of the two chrominance components U and V are half that of the two-dimensional luma component Y, and the position constraints K1 and K2 of the chrominance components are both half that of the luma components. For example, if the parameters of the luma coding block are derived using the predicted luma samples and reconstructed luma samples from the 16 samples at the top boundary and the 16 samples at the left boundary, then the predicted chrominance samples and reconstructed chrominance samples from the 8 samples at the top boundary and the 8 samples at the left boundary will be used to derive the parameters of the chrominance coding block.
[0215] Similarly, when applying sub-block level model parameters for derivation, the sub-block size of the chroma component can also depend on the color format of the video sequence. In the YUV 4:4:4 format with three components Y, U, and V having the same sampling rate, the sub-block size of the chroma component is the same as that of the luma component. In the YUV 4:2:2 format, the sampling rates of the chroma components U and V are half that of the luma component Y in the horizontal dimension, and the sub-block size of the chroma component is also half that of the luma component in the horizontal dimension. For example, if the luma-coded block is divided into 16×16 sub-blocks, the chroma-coded block can be divided into 8×16 sub-blocks. In the YUV 4:2:0 format, the sampling rates of the chroma components U and V are half that of the two-dimensional luma component Y, and the sub-block size of the chroma component is also half that of the two-dimensional luma component. For example, if the luma-coded block is divided into 16×16 sub-blocks, the chroma-coded block is divided into 8×8 sub-blocks.
[0216] In some embodiments, the sample position constraint depends on the inter-frame prediction mode and the MVP candidate used for inter-frame prediction. For example, when encoder 200 or decoder 300 performs sub-block level inter-frame prediction on a coding block, the sample position constraint value is reduced to the size of the sub-block. Furthermore, when LCC is enabled, the sample position constraint value for the chroma coding block is also reduced to the size of the chroma sub-block. For example, in the YUV 4:2:0 format, if encoder 200 or decoder 300 selects a sub-block level TMVP candidate or MVAP candidate for the coding unit, the horizontal and vertical position constraints for the luma coding block in the coding unit can be set to 8, and the horizontal and vertical position constraints for the chroma coding block in the coding unit can be set to 4. If encoder 200 or decoder 300 selects a block level TMVP candidate, SMVP candidate, or HMVP candidate for the coding unit, the horizontal and vertical position constraints for the luma coding block of the coding unit will be set to 16, while the horizontal and vertical position constraints for the chroma coding block of the coding unit will be set to 8.
[0217] After obtaining the predicted chromaticity sample and the reconstructed chromaticity sample, in step 2526, the device obtains model parameters for local chromaticity compensation based on the obtained predicted chromaticity sample and the reconstructed chromaticity sample, and in step 2528, the corrected chromaticity prediction block is obtained based on the predicted chromaticity sample, the reconstructed chromaticity sample and the chromaticity prediction block.
[0218] In some embodiments, given that the chroma component has less texture information than the luma component, the chroma component compensation model can be further simplified using a single-parameter model. During the simplification of the single-parameter model, the scaling factor parameter 'a' of the linear model is removed, while the offset parameter 'b' is retained. In other words, in chroma compensation, the chroma prediction sample value can be directly set to a specific value 'b', which is independent of the predicted chroma sample value before compensation. Therefore, in this embodiment, the chroma prediction sample value in the current coded block becomes the value obtained from the reconstructed chroma samples of the top and left adjacent blocks.
[0219] Furthermore, the device can determine whether to perform chroma compensation based on the derived linear model parameters a and b. For example, a value of zero for factor parameter a can be set as a prerequisite for performing chroma compensation. In step 2446, the device first derives the model parameters a and b. If the value of factor parameter a is equal to zero, then in step 2528, the device can apply LCC based on a single-parameter model to generate a compensated prediction block for the inter-frame prediction process. Otherwise, the device does not apply local chroma compensation to the current chroma-coded block.
[0220] In some other embodiments, the condition for performing LCC can be determined based on the predicted chromaticity sample and the reconstructed chromaticity sample obtained in step 2524. Optionally, step 2526 can be replaced by a step of checking the situation based on the predicted chromaticity sample and the reconstructed chromaticity sample. If the above condition is met, LCC is applied to the chromaticity prediction block in step 2528; otherwise, step 2528 is skipped.
[0221] In one example, such as Figure 19C and Figure 19D As shown, encoder 200 or decoder 300 may first select four pairs of samples (e.g., the first pair of samples 1912, 1932, the second pair of samples 1914, 1934, the third pair of samples 1916, 1936, and the fourth pair of samples 1918, 1938, wherein the prediction samples 1912-1918 are obtained from the left and / or top boundaries within the prediction block 1910, and the reconstruction samples 1932-1938 are obtained from the top and / or left adjacent blocks of the current block 1920).
[0222] Then, encoder 200 or decoder 300 classifies the four sample pairs based on the values of the predicted samples 1912-1918. For ease of understanding, assume that four pairs of samples are being classified, where the values of the predicted samples 1912-1918 are arranged in non-decreasing order (i.e., the value of the predicted sample in each pair is greater than or equal to the value of the predicted sample in the previous pair). In other words, encoder 200 or decoder 300 can classify the predicted chroma samples and obtain two smaller predicted chroma samples and two larger predicted chroma samples. Since the predicted chroma samples and reconstructed chroma samples are paired, encoder 200 or decoder 300 can also obtain two reconstructed chroma samples corresponding to the two smaller predicted chroma samples and two additional reconstructed chroma samples corresponding to the two larger predicted chroma samples.
[0223] Then, encoder 200 or decoder 300 calculates the first two pairs to obtain (x_min, y_min) and the last two pairs to obtain (x_max, y_max), where x_min represents the average of predicted samples 1912 and 1914 (i.e., the two smaller predicted chroma samples), y_min represents the average of reconstructed samples 1932 and 1934 (i.e., the two predicted chroma samples corresponding to the two smaller predicted chroma samples), x_max represents the average of predicted samples 1916 and 1918 (i.e., the two larger predicted chroma samples), and y_max represents the average of reconstructed samples 1936 and 1938 (i.e., the two reconstructed chroma samples corresponding to the two larger predicted chroma samples). In other words, encoder 200 or decoder 300 can derive a first minimum x_min as the average of the two smaller predicted chroma samples and a second minimum y_min as the average of the two reconstructed chroma samples respectively corresponding to the two smaller predicted chroma samples. Furthermore, encoder 200 or decoder 300 can derive a first maximum value x_max as the average of two larger predicted chromaticity samples and a second maximum value y_max as the average of two reconstructed chromaticity samples corresponding to the two larger predicted chromaticity samples, respectively.
[0224] Then, encoder 200 or decoder 300 calculates the difference between y_max and y_min (denoted as y_diff) and the difference between x_max and x_min (denoted as x_diff). If the value of y_diff and / or the value of x_diff are zero or less than a threshold, then in step 2528, encoder 200 or decoder 300 can set the predicted sample value of the current chroma coding block to (y_min+y_max) / 2. Otherwise, local chroma compensation does not change the predicted chroma sample of the current chroma coding block. In other words, if the value of y_diff and / or the value of x_diff do not meet the condition, enabling LCC is skipped.
[0225] That is, encoder 200 or decoder 300 can determine a first difference x_diff based on the predicted chroma samples, determine a second difference y_diff based on the reconstructed chroma samples, and derive a corrected chroma prediction block based on the first difference x_diff and the second difference y_diff. In some embodiments, encoder 200 or decoder 300 can determine the first difference x_diff as the difference between a first maximum value x_max and a first minimum value x_min, and determine the second difference y_diff as the difference between a second maximum value y_max and a second minimum value y_min. If the first difference x_diff or the second difference y_diff is less than a threshold, encoder 200 or decoder 300 can determine that each corrected predicted chroma sample in the corrected chroma prediction block is the average of the second minimum value y_min and the second maximum value y_max. Otherwise, encoder 200 or decoder 300 can skip the chroma compensation process and derive the corrected chroma prediction block as the chroma prediction block (i.e., local chroma compensation cannot change the predicted chroma samples of the current chroma coding block).
[0226] In other words, the device can determine whether to skip enabling LCC based on a threshold and sample pairs. In response to determining whether to perform local chromaticity compensation based on the threshold and samples, the device can derive individual model parameters for LCC based on the reconstructed chromaticity samples.
[0227] In some embodiments, the threshold can be a fixed value (e.g., 0, 1, 2, or 4), or it can depend on the bit depth of the sample value, for example, the bit depth value of the chroma sample (e.g., 1 << (10 bits deep)). In some embodiments, the threshold can be 1 << (bit depth - 8), where the parameter "bit depth" indicates the bit depth of the sample value, and "<<" represents a left shift operation. In some other embodiments, when chroma compensation is enabled, the threshold can be signaled in the bitstream (e.g., in the sequence header or picture header).
[0228] In some embodiments, the LCC may share the same enable flag and mode index with the LIC used for signaling, which is not specifically limited herein. In some embodiments, the LCC may also have separate enable flags and mode indexes. In the case of shared signaling, if the encoder 200 signals the chroma compensation threshold of the LCC in the bitstream, the encoder 200 may signal the chroma compensation threshold only when the LIC is enabled.
[0229] In summary, according to the various embodiments proposed in this disclosure, by applying simplified Local Luminance Compensation (LIC) and Local Chroma Compensation (LCC) procedures and restricting sample locations to derive model parameters, the inter-frame prediction process can reduce the latency of LIC / LCC operations by storing fewer unrefined prediction samples. Furthermore, LCC also compensates for the chromaticity difference between the current image and the reference image, thereby improving the accuracy of inter-frame prediction used for encoding and decoding video.
[0230] The various embodiments described herein are performed in method or process steps that can be implemented, on the one hand, by a computer program product embodied in a computer-readable medium, including computer-executable instructions, such as program code, executed by a computer in a networked environment. Typically, a program module may include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. The computer-executable instructions, associated data structures, and program modules represent examples of program code for performing steps of the methods disclosed herein. A particular sequence of such executable instructions or associated data structures represents examples of corresponding actions for implementing the functionality described in such steps or processes.
[0231] In some embodiments, a non-volatile computer-readable storage medium including instructions is also provided. In some embodiments, the medium may store all or part of the video bitstream having an enable flag and an index. The enable flag is associated with the video data and may indicate whether local brightness compensation is enabled for coded blocks of the inter-frame prediction process. The index is associated with local brightness compensation and is used to indicate the selected local brightness compensation mode.
[0232] In some embodiments, the aforementioned medium may store instructions that can be executed by a device (e.g., the disclosed encoder and decoder) to perform the methods described above. Common forms of non-volatile media include, for example, floppy disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a hole pattern, RAM, PROMs and EPROMs, FLASH-EPROMs or any other flash memory, NVRAM, caches, registers, any other memory chips or cassette memories and their network versions. The device may include at least one processor (CPU), input / output interface, network interface, or memory.
[0233] It should be noted that the relational terms used herein (e.g., “first” and “second”) are used only to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between these entities or operations. Furthermore, the words “including,” “having,” “containing,” and “comprising,” as well as other similar forms, are intended to be semantically equivalent and open-ended, as at least one item following any of these terms does not imply an exhaustive list of that at least one item, nor does it imply limitation to the at least one item listed.
[0234] As used herein, unless otherwise expressly stated, the term "or" includes all possible combinations unless impractical. For example, if it is specified that a database may include A or B, then unless otherwise specified or impractical, the database may include A, B, or A and B. As a second example, if it is specified that a database may include A, B, or C, then unless otherwise specified or impractical, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0235] It should be understood that the above embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. When executed by a processor, the software can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0236] In the foregoing description, embodiments have been described with reference to numerous specific details, which may vary depending on the implementation. Certain adjustments and modifications may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art in light of the detailed description and practice of this disclosure herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims. The sequence of steps shown in the figures is also intended for illustrative purposes only and is not intended to limit one to any particular order of steps. Therefore, those skilled in the art will understand that these steps may be performed in a different order when implementing the same method.
[0237] The embodiments may be further described using the following terms:
[0238] 1. A video processing method, comprising:
[0239] Determine whether to enable inter-frame predictor correction for the coded block;
[0240] When inter-frame predictor correction is enabled for a coded block, inter-frame predictor correction is performed using the following steps:
[0241] Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block;
[0242] Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block;
[0243] Based on multiple predicted samples and multiple reconstructed samples, at least one parameter for inter-frame predictor correction is obtained; and
[0244] Based on at least one parameter and a prediction block, a corrected prediction block is obtained.
[0245] 2. The video processing method according to Clause 1, wherein acquiring multiple prediction samples and multiple reconstructed samples includes:
[0246] The horizontal and vertical position parameters are determined based on the coding pattern of the coding block;
[0247] Multiple horizontal positions can be determined based on horizontal position parameters, or multiple vertical positions can be determined based on vertical position parameters;
[0248] Multiple prediction samples can be obtained from the top boundary of the prediction block based on multiple horizontal positions, or from the left boundary of the prediction block based on multiple vertical positions; and
[0249] Multiple reconstructed samples can be obtained from the top adjacent reconstructed sample based on multiple vertical positions, or multiple reconstructed samples can be obtained from the left adjacent reconstructed sample based on multiple vertical positions.
[0250] 3. The video processing method according to Clause 2, wherein the value of the horizontal position parameter is equal to the value of the vertical position parameter.
[0251] 4. The video processing method according to Clause 2 or Clause 3, wherein the values of the horizontal position parameter and the vertical position parameter are equal to 8 or 16.
[0252] 5. The video processing method according to Clause 4 further includes:
[0253] If the coding mode of the coded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal and vertical position parameters are determined to be 8; otherwise, the values of the horizontal and vertical position parameters are determined to be 16.
[0254] 6. The video processing method according to any one of clauses 2-5 further includes:
[0255] If the width of the coded block is greater than the value of the horizontal position parameter, then the limiting value of the horizontal position parameter is determined to be the value of the horizontal position parameter; otherwise, the limiting value of the horizontal position parameter is determined to be the width of the coded block.
[0256] If the height of the coded block is greater than the value of the vertical position parameter, then the limiting value of the vertical position parameter is determined to be the value of the vertical position parameter; otherwise, the limiting value of the vertical position parameter is determined to be the height of the coded block; and
[0257] Based on the amplitude limits of the horizontal and vertical position parameters, multiple predicted samples and multiple reconstructed samples are obtained.
[0258] 7. A video processing method, comprising:
[0259] Determine whether to enable inter-frame predictor correction for the coded block;
[0260] When inter-frame predictor correction is enabled for a coded block, inter-frame predictor correction is performed using the following steps:
[0261] Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block;
[0262] Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block; and
[0263] Based on multiple predicted samples, multiple reconstructed samples, and predicted blocks, a corrected predicted block is obtained.
[0264] 8. The video processing method according to Clause 7, wherein obtaining the corrected prediction block further includes:
[0265] Based on multiple predicted samples and multiple reconstructed samples, at least one parameter for inter-frame predictor correction is obtained; and
[0266] The corrected prediction block is obtained based on at least one parameter and the prediction block.
[0267] 9. The video processing method according to Clause 7 or Clause 8, wherein acquiring multiple prediction samples and multiple reconstructed samples includes:
[0268] The horizontal and vertical position parameters are determined based on the coding pattern of the coding block;
[0269] Multiple horizontal positions can be determined based on horizontal position parameters, or multiple vertical positions can be determined based on vertical position parameters;
[0270] Multiple prediction samples can be obtained from the top boundary of the prediction block based on multiple horizontal positions, or from the left boundary of the prediction block based on multiple vertical positions; and
[0271] Multiple reconstructed samples can be obtained from the top adjacent reconstructed sample based on multiple vertical positions, or multiple reconstructed samples can be obtained from the left adjacent reconstructed sample based on multiple vertical positions.
[0272] 10. The video processing method according to Clause 9, wherein the value of the horizontal position parameter is equal to the value of the vertical position parameter.
[0273] 11. The video processing method according to Clause 9 or Clause 10, wherein the values of the horizontal position parameter and the vertical position parameter are equal to 4 or 8.
[0274] 12. The video processing method according to Clause 11 further includes:
[0275] If the coding mode of the coded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal and vertical position parameters are determined to be 4; otherwise, the values of the horizontal and vertical position parameters are determined to be 8.
[0276] 13. The video processing method according to any one of clauses 9-12 further includes:
[0277] If the width of the coded block is greater than the value of the horizontal position parameter, then the limiting value of the horizontal position parameter is determined to be the value of the horizontal position parameter; otherwise, the limiting value of the horizontal position parameter is determined to be the width of the coded block.
[0278] If the height of the coded block is greater than the value of the vertical position parameter, then the limiting value of the vertical position parameter is determined to be the value of the vertical position parameter; otherwise, the limiting value of the vertical position parameter is determined to be the height of the coded block; and
[0279] Based on the amplitude limits of the horizontal and vertical position parameters, multiple predicted samples and multiple reconstructed samples are obtained.
[0280] 14. The video processing method according to any one of clauses 7-13 further includes:
[0281] A first difference is determined based on multiple predicted samples, and a second difference is determined based on multiple reconstructed samples; and
[0282] The corrected prediction block is obtained based on the first difference and the second difference.
[0283] 15. The video processing method according to Clause 14 further includes:
[0284] If either the first or second difference is less than the threshold, then each corrected prediction sample in the corrected prediction block is determined to be equal to the first value; otherwise, the corrected prediction block is used as the prediction block.
[0285] 16. The video processing method according to Clause 15, wherein the threshold depends on the bit depth of the sample value.
[0286] 17. The video processing method according to Clause 16, wherein the threshold is 1 << (bit depth - 8), where "bit depth" is the bit depth of the sample value, and "<<" indicates a left shift operation.
[0287] 18. The video processing method according to any one of clauses 7-17 further comprises:
[0288] The multiple predicted samples are sorted, and two smaller predicted samples and two larger predicted samples are obtained.
[0289] The first minimum value is used as the average of the two smaller predicted samples, and the second minimum value is used as the average of the two reconstructed samples corresponding to the two smaller predicted samples respectively;
[0290] The first maximum value is used as the average of the two larger predicted samples, and the second maximum value is used as the average of the two reconstructed samples corresponding to the two larger predicted samples; and
[0291] The first difference is defined as the difference between the first maximum value and the first minimum value, and the second difference is defined as the difference between the second maximum value and the second minimum value.
[0292] 19. The video processing method according to Clause 18, wherein the first value is the average of the second minimum value and the second maximum value.
[0293] 20. An apparatus comprising:
[0294] Memory, configured to store instructions; and
[0295] At least one processor is configured to execute instructions to cause the device to perform the following steps:
[0296] Determine whether to enable inter-frame predictor correction for the coded block;
[0297] When inter-frame predictor correction is enabled for a coded block, inter-frame predictor correction is performed using the following steps:
[0298] Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block;
[0299] Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block;
[0300] Based on multiple predicted samples and multiple reconstructed samples, at least one parameter for inter-frame predictor correction is obtained; and
[0301] Based on at least one parameter and a prediction block, a corrected prediction block is obtained.
[0302] 21. The apparatus according to claim 20, wherein at least one processor is configured to execute instructions to cause the apparatus to obtain a plurality of predicted samples and a plurality of reconstructed samples by means of the following steps:
[0303] The horizontal and vertical position parameters are determined based on the coding pattern of the coding block;
[0304] Multiple horizontal positions can be determined based on horizontal position parameters, or multiple vertical positions can be determined based on vertical position parameters;
[0305] Multiple prediction samples can be obtained from the top boundary of the prediction block based on multiple horizontal positions, or from the left boundary of the prediction block based on multiple vertical positions; and
[0306] Multiple reconstructed samples can be obtained from the top adjacent reconstructed sample based on multiple vertical positions, or multiple reconstructed samples can be obtained from the left adjacent reconstructed sample based on multiple vertical positions.
[0307] 22. The apparatus according to Clause 21, wherein the value of the horizontal position parameter is equal to the value of the vertical position parameter.
[0308] 23. The apparatus according to clause 21 or 22, wherein the values of the horizontal position parameter and the vertical position parameter are equal to 8 or 16.
[0309] 24. The apparatus according to clause 23, wherein at least one processor is configured to execute instructions to further cause the apparatus to perform the following steps:
[0310] If the coding mode of the coded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal and vertical position parameters are determined to be 8; otherwise, the values of the horizontal and vertical position parameters are determined to be 16.
[0311] 25. The apparatus according to any one of clauses 21-24, wherein at least one processor is configured to execute instructions to further cause the apparatus to perform the following steps:
[0312] If the width of the coded block is greater than the value of the horizontal position parameter, then the limiting value of the horizontal position parameter is determined to be the value of the horizontal position parameter; otherwise, the limiting value of the horizontal position parameter is determined to be the width of the coded block.
[0313] If the height of the coded block is greater than the value of the vertical position parameter, then the limiting value of the vertical position parameter is determined to be the value of the vertical position parameter; otherwise, the limiting value of the vertical position parameter is determined to be the height of the coded block; and
[0314] Based on the amplitude limits of the horizontal and vertical position parameters, multiple predicted samples and multiple reconstructed samples are obtained.
[0315] 26. An apparatus comprising:
[0316] Memory, configured to store instructions; and
[0317] At least one processor is configured to execute instructions to cause the device to perform the following steps:
[0318] Determine whether to enable inter-frame predictor correction for the coded block;
[0319] When inter-frame predictor correction is enabled for a coded block, inter-frame predictor correction is performed using the following steps:
[0320] Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block;
[0321] Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block; and
[0322] Based on multiple predicted samples, multiple reconstructed samples, and predicted blocks, a corrected predicted block is obtained.
[0323] 27. The apparatus according to Clause 26, wherein at least one processor is configured to execute instructions to further cause the apparatus to obtain the corrected prediction block by further steps:
[0324] Based on multiple predicted samples and multiple reconstructed samples, at least one parameter for inter-frame predictor correction is obtained; and
[0325] The corrected prediction block is obtained based on at least one parameter and the prediction block.
[0326] 28. The apparatus according to clause 26 or 27, wherein at least one processor is configured to execute instructions to further cause the apparatus to obtain a plurality of prediction samples and a plurality of reconstruction samples by means of the following steps:
[0327] The horizontal and vertical position parameters are determined based on the coding pattern of the coding block;
[0328] Multiple horizontal positions can be determined based on horizontal position parameters, or multiple vertical positions can be determined based on vertical position parameters;
[0329] Multiple prediction samples can be obtained from the top boundary of the prediction block based on multiple horizontal positions, or from the left boundary of the prediction block based on multiple vertical positions; and
[0330] Multiple reconstructed samples can be obtained from the top adjacent reconstructed sample based on multiple vertical positions, or multiple reconstructed samples can be obtained from the left adjacent reconstructed sample based on multiple vertical positions.
[0331] 29. The apparatus according to Clause 28, wherein the value of the horizontal position parameter is equal to the value of the vertical position parameter.
[0332] 30. The apparatus according to Clause 28 or Clause 29, wherein the values of the horizontal position parameter and the vertical position parameter are equal to 4 or 8.
[0333] 31. The apparatus according to claim 30, wherein at least one processor is configured to execute instructions to further cause the apparatus to perform the following steps:
[0334] If the coding mode of the coded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal and vertical position parameters are determined to be 4; otherwise, the values of the horizontal and vertical position parameters are determined to be 8.
[0335] 32. The apparatus according to any one of clauses 28-31, wherein at least one processor is configured to execute instructions to further cause the apparatus to perform the following steps:
[0336] If the width of the coded block is greater than the value of the horizontal position parameter, then the limiting value of the horizontal position parameter is determined to be the value of the horizontal position parameter; otherwise, the limiting value of the horizontal position parameter is determined to be the width of the coded block.
[0337] If the height of the coded block is greater than the value of the vertical position parameter, then the limiting value of the vertical position parameter is determined to be the value of the vertical position parameter; otherwise, the limiting value of the vertical position parameter is determined to be the height of the coded block; and
[0338] Based on the amplitude limits of the horizontal and vertical position parameters, multiple predicted samples and multiple reconstructed samples are obtained.
[0339] 33. The apparatus according to any one of clauses 26-32, wherein at least one processor is configured to execute instructions to further cause the apparatus to perform the following steps:
[0340] A first difference is determined based on multiple predicted samples, and a second difference is determined based on multiple reconstructed samples; and
[0341] The corrected prediction block is obtained based on the first difference and the second difference.
[0342] 34. The apparatus according to clause 33, wherein at least one processor is configured to execute instructions to further cause the apparatus to perform the following steps:
[0343] If either the first or second difference is less than the threshold, then each corrected prediction sample in the corrected prediction block is determined to be equal to the first value; otherwise, the corrected prediction block is used as the prediction block.
[0344] 35. The apparatus according to clause 34, wherein the threshold depends on the bit depth of the sample value.
[0345] 36. The apparatus according to Clause 35, wherein the threshold is 1 << (bit depth - 8), where "bit depth" is the bit depth of the sample value, and "<<" indicates a left shift operation.
[0346] 37. The apparatus according to any one of clauses 26-36, wherein at least one processor is configured to execute instructions to further cause the apparatus to perform the following steps:
[0347] The multiple predicted samples are sorted, and two smaller predicted samples and two larger predicted samples are obtained.
[0348] The first minimum value is used as the average of the two smaller predicted samples, and the second minimum value is used as the average of the two reconstructed samples corresponding to the two smaller predicted samples respectively;
[0349] The first maximum value is used as the average of the two larger predicted samples, and the second maximum value is used as the average of the two reconstructed samples corresponding to the two larger predicted samples; and
[0350] The first difference is defined as the difference between the first maximum value and the first minimum value, and the second difference is defined as the difference between the second maximum value and the second minimum value.
[0351] 38. The apparatus according to Clause 37, wherein the first value is the average of the second minimum value and the second maximum value.
[0352] 39. A non-volatile computer-readable storage medium storing an instruction set executable by at least one processor of a device to cause the device to perform a video processing method, the video processing method comprising:
[0353] Determine whether to enable inter-frame predictor correction for the coded block;
[0354] When inter-frame predictor correction is enabled for a coded block, inter-frame predictor correction is performed using the following steps:
[0355] Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block;
[0356] Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block;
[0357] Based on multiple predicted samples and multiple reconstructed samples, at least one parameter for inter-frame predictor correction is obtained; and
[0358] Based on at least one parameter and a prediction block, a corrected prediction block is obtained.
[0359] 40. The non-volatile computer-readable storage medium according to Clause 39, wherein acquiring a plurality of prediction samples and a plurality of reconstruction samples includes:
[0360] The horizontal and vertical position parameters are determined based on the coding pattern of the coding block;
[0361] Multiple horizontal positions can be determined based on horizontal position parameters, or multiple vertical positions can be determined based on vertical position parameters;
[0362] Multiple prediction samples can be obtained from the top boundary of the prediction block based on multiple horizontal positions, or from the left boundary of the prediction block based on multiple vertical positions; and
[0363] Multiple reconstructed samples can be obtained from the top adjacent reconstructed sample based on multiple vertical positions, or multiple reconstructed samples can be obtained from the left adjacent reconstructed sample based on multiple vertical positions.
[0364] 41. The non-volatile computer-readable storage medium as described in Clause 40, wherein the value of the horizontal position parameter is equal to the value of the vertical position parameter.
[0365] 42. The non-volatile computer-readable storage medium as described in Clause 40 or Clause 41, wherein the values of the horizontal position parameter and the vertical position parameter are equal to 8 or 16.
[0366] 43. The non-volatile computer-readable storage medium according to Clause 42, wherein the video processing method further comprises:
[0367] If the coding mode of the coded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal and vertical position parameters are determined to be 8; otherwise, the values of the horizontal and vertical position parameters are determined to be 16.
[0368] 44. The non-volatile computer-readable storage medium according to any one of clauses 40-43, wherein the video processing method further comprises:
[0369] If the width of the coded block is greater than the value of the horizontal position parameter, then the limiting value of the horizontal position parameter is determined to be the value of the horizontal position parameter; otherwise, the limiting value of the horizontal position parameter is determined to be the width of the coded block.
[0370] If the height of the coded block is greater than the value of the vertical position parameter, then the limiting value of the vertical position parameter is determined to be the value of the vertical position parameter; otherwise, the limiting value of the vertical position parameter is determined to be the height of the coded block; and
[0371] Based on the amplitude limits of the horizontal and vertical position parameters, multiple predicted samples and multiple reconstructed samples are obtained.
[0372] 45. A non-volatile computer-readable storage medium storing an instruction set executable by at least one processor of a device to cause the device to perform a video processing method, the video processing method comprising:
[0373] Determine whether to enable inter-frame predictor correction for the coded block;
[0374] When inter-frame predictor correction is enabled for a coded block, inter-frame predictor correction is performed using the following steps:
[0375] Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block;
[0376] Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block; and
[0377] Based on multiple predicted samples, multiple reconstructed samples, and predicted blocks, a corrected predicted block is obtained.
[0378] 46. The non-volatile computer-readable storage medium according to clause 45, wherein the corrected prediction block further comprises:
[0379] Based on multiple predicted samples and multiple reconstructed samples, at least one parameter for inter-frame predictor correction is obtained; and
[0380] The corrected prediction block is obtained based on at least one parameter and the prediction block.
[0381] 47. A non-volatile computer-readable storage medium as described in Clause 45 or Clause 46, wherein acquiring a plurality of prediction samples and a plurality of reconstruction samples comprises:
[0382] The horizontal and vertical position parameters are determined based on the coding pattern of the coding block;
[0383] Multiple horizontal positions can be determined based on horizontal position parameters, or multiple vertical positions can be determined based on vertical position parameters;
[0384] Multiple prediction samples can be obtained from the top boundary of the prediction block based on multiple horizontal positions, or from the left boundary of the prediction block based on multiple vertical positions; and
[0385] Multiple reconstructed samples can be obtained from the top adjacent reconstructed sample based on multiple vertical positions, or multiple reconstructed samples can be obtained from the left adjacent reconstructed sample based on multiple vertical positions.
[0386] 48. The non-volatile computer-readable storage medium as described in Clause 47, wherein the value of the horizontal position parameter is equal to the value of the vertical position parameter.
[0387] 49. The non-volatile computer-readable storage medium as described in Clause 47 or Clause 48, wherein the values of the horizontal position parameter and the vertical position parameter are equal to 4 or 8.
[0388] 50. The non-volatile computer-readable storage medium as described in Clause 49, wherein the video processing method further comprises:
[0389] If the coding mode of the coded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal and vertical position parameters are determined to be 4; otherwise, the values of the horizontal and vertical position parameters are determined to be 8.
[0390] 51. The video processing method further comprises the non-volatile computer-readable storage medium according to any one of clauses 47-50:
[0391] If the width of the coded block is greater than the value of the horizontal position parameter, then the limiting value of the horizontal position parameter is determined to be the value of the horizontal position parameter; otherwise, the limiting value of the horizontal position parameter is determined to be the width of the coded block.
[0392] If the height of the coded block is greater than the value of the vertical position parameter, then the limiting value of the vertical position parameter is determined to be the value of the vertical position parameter; otherwise, the limiting value of the vertical position parameter is determined to be the height of the coded block; and
[0393] Based on the amplitude limits of the horizontal and vertical position parameters, multiple predicted samples and multiple reconstructed samples are obtained.
[0394] 52. The video processing method further comprises the non-volatile computer-readable storage medium according to any one of clauses 45-51:
[0395] A first difference is determined based on multiple predicted samples, and a second difference is determined based on multiple reconstructed samples; and
[0396] The corrected prediction block is obtained based on the first difference and the second difference.
[0397] 53. The video processing method further includes, according to the non-volatile computer-readable storage medium described in Clause 52:
[0398] If either the first or second difference is less than the threshold, then each corrected prediction sample in the corrected prediction block is determined to be equal to the first value; otherwise, the corrected prediction block is used as the prediction block.
[0399] 54. The non-volatile computer-readable storage medium as described in Clause 53, wherein the threshold depends on the bit depth of the sample value.
[0400] 55. The non-volatile computer-readable storage medium as described in Clause 54, wherein the threshold is 1 << (bit depth - 8), where "bit depth" is the bit depth of the sample value, and "<<" indicates a left shift operation.
[0401] 56. The video processing method further comprises the non-volatile computer-readable storage medium according to any one of clauses 45-55, wherein:
[0402] The multiple predicted samples are sorted, and two smaller predicted samples and two larger predicted samples are obtained.
[0403] The first minimum value is used as the average of the two smaller predicted samples, and the second minimum value is used as the average of the two reconstructed samples corresponding to the two smaller predicted samples respectively;
[0404] The first maximum value is used as the average of the two larger predicted samples, and the second maximum value is used as the average of the two reconstructed samples corresponding to the two larger predicted samples; and
[0405] The first difference is defined as the difference between the first maximum value and the first minimum value, and the second difference is defined as the difference between the second maximum value and the second minimum value.
[0406] 57. The non-volatile computer-readable storage medium as described in Clause 56, wherein the first value is the average of the second minimum value and the second maximum value.
[0407] 58. A non-volatile computer-readable medium for storing a bit stream, wherein the bit stream comprises:
[0408] An enable flag associated with the video data, indicating whether inter-frame predictor correction is enabled for the coded blocks of the inter-frame prediction process; and
[0409] An index associated with inter-frame predictor correction, which indicates the selected mode of inter-frame predictor correction;
[0410] The enabled inter-frame predictor correction is performed through the following steps:
[0411] Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block;
[0412] Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block;
[0413] Based on multiple predicted samples and multiple reconstructed samples, at least one parameter for inter-frame predictor correction is obtained; and
[0414] The corrected prediction block is obtained based on at least one parameter and the prediction block.
[0415] 59. The non-volatile computer-readable storage medium as described in Clause 58, wherein acquiring a plurality of prediction samples and a plurality of reconstruction samples comprises:
[0416] The horizontal and vertical position parameters are determined based on the coding pattern of the coding block;
[0417] Multiple horizontal positions can be determined based on horizontal position parameters, or multiple vertical positions can be determined based on vertical position parameters;
[0418] Multiple prediction samples can be obtained from the top boundary of the prediction block based on multiple horizontal positions, or from the left boundary of the prediction block based on multiple vertical positions; and
[0419] Multiple reconstructed samples can be obtained from the top adjacent reconstructed sample based on multiple vertical positions, or multiple reconstructed samples can be obtained from the left adjacent reconstructed sample based on multiple vertical positions.
[0420] 60. The non-volatile computer-readable storage medium as described in Clause 59, wherein the value of the horizontal position parameter is equal to the value of the vertical position parameter.
[0421] 61. The non-volatile computer-readable storage medium as described in Clause 59 or Clause 60, wherein the values of the horizontal position parameter and the vertical position parameter are equal to 8 or 16.
[0422] 62. The non-volatile computer-readable storage medium according to Clause 61, wherein the video processing method further comprises:
[0423] If the coding mode of the coded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal and vertical position parameters are determined to be 8; otherwise, the values of the horizontal and vertical position parameters are determined to be 16.
[0424] 63. The non-volatile computer-readable storage medium according to any one of clauses 59-62, wherein the video processing method further comprises:
[0425] If the width of the coded block is greater than the value of the horizontal position parameter, then the limiting value of the horizontal position parameter is determined to be the value of the horizontal position parameter; otherwise, the limiting value of the horizontal position parameter is determined to be the width of the coded block.
[0426] If the height of the coded block is greater than the value of the vertical position parameter, then the limiting value of the vertical position parameter is determined to be the value of the vertical position parameter; otherwise, the limiting value of the vertical position parameter is determined to be the height of the coded block; and
[0427] Based on the amplitude limits of the horizontal and vertical position parameters, multiple predicted samples and multiple reconstructed samples are obtained.
[0428] 64. A non-volatile computer-readable medium for storing a bit stream, wherein the bit stream comprises:
[0429] An enable flag associated with the video data, indicating whether inter-frame predictor correction is enabled for the coded blocks of the inter-frame prediction process; and
[0430] An index associated with inter-frame predictor correction, which indicates the selected mode of inter-frame predictor correction;
[0431] The enabled inter-frame predictor correction is performed through the following steps:
[0432] Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block;
[0433] Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block; and
[0434] The corrected prediction block is obtained based on multiple prediction samples, multiple reconstructed samples, and prediction blocks.
[0435] 65. The non-volatile computer-readable storage medium according to Clause 64, wherein the corrected prediction block further comprises:
[0436] Based on multiple predicted samples and multiple reconstructed samples, at least one parameter for inter-frame predictor correction is obtained; and
[0437] The corrected prediction block is obtained based on at least one parameter and the prediction block.
[0438] 66. The non-volatile computer-readable storage medium according to Clause 64 or Clause 65, wherein obtaining a plurality of prediction samples and a plurality of reconstruction samples includes:
[0439] The horizontal and vertical position parameters are determined based on the coding pattern of the coding block;
[0440] Multiple horizontal positions can be determined based on horizontal position parameters, or multiple vertical positions can be determined based on vertical position parameters;
[0441] Multiple prediction samples can be obtained from the top boundary of the prediction block based on multiple horizontal positions, or from the left boundary of the prediction block based on multiple vertical positions; and
[0442] Multiple reconstructed samples can be obtained from the top adjacent reconstructed sample based on multiple vertical positions, or multiple reconstructed samples can be obtained from the left adjacent reconstructed sample based on multiple vertical positions.
[0443] 67. The non-volatile computer-readable storage medium as described in Clause 66, wherein the value of the horizontal position parameter is equal to the value of the vertical position parameter.
[0444] 68. The non-volatile computer-readable storage medium as described in Clause 66 or Clause 67, wherein the values of the horizontal position parameter and the vertical position parameter are equal to 4 or 8.
[0445] 69. The non-volatile computer-readable storage medium as described in Clause 68, wherein the video processing method further comprises:
[0446] If the coding mode of the coded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal and vertical position parameters are determined to be 4; otherwise, the values of the horizontal and vertical position parameters are determined to be 8.
[0447] 70. The video processing method further comprises: a non-volatile computer-readable storage medium according to any one of clauses 66-69.
[0448] If the width of the coded block is greater than the value of the horizontal position parameter, then the limiting value of the horizontal position parameter is determined to be the value of the horizontal position parameter; otherwise, the limiting value of the horizontal position parameter is determined to be the width of the coded block.
[0449] If the height of the coded block is greater than the value of the vertical position parameter, then the limiting value of the vertical position parameter is determined to be the value of the vertical position parameter; otherwise, the limiting value of the vertical position parameter is determined to be the height of the coded block; and
[0450] Based on the amplitude limits of the horizontal and vertical position parameters, multiple predicted samples and multiple reconstructed samples are obtained.
[0451] 71. The video processing method further comprises: a non-volatile computer-readable storage medium according to any one of clauses 64-70.
[0452] A first difference is determined based on multiple predicted samples, and a second difference is determined based on multiple reconstructed samples; and
[0453] The corrected prediction block is obtained based on the first difference and the second difference.
[0454] 72. The video processing method further includes, according to the non-volatile computer-readable storage medium described in Clause 71:
[0455] If either the first or second difference is less than the threshold, then each corrected prediction sample in the corrected prediction block is determined to be equal to the first value; otherwise, the corrected prediction block is used as the prediction block.
[0456] 73. The non-volatile computer-readable storage medium as described in Clause 72, wherein the threshold depends on the bit depth of the sample value.
[0457] 74. The non-volatile computer-readable storage medium as described in Clause 73, wherein the threshold is 1 << (bit depth - 8), where "bit depth" is the bit depth of the sample value, and "<<" indicates a left shift operation.
[0458] 75. The video processing method further comprises the non-volatile computer-readable storage medium according to any one of clauses 64-74:
[0459] The multiple predicted samples are sorted, and two smaller predicted samples and two larger predicted samples are obtained.
[0460] The first minimum value is used as the average of the two smaller predicted samples, and the second minimum value is used as the average of the two reconstructed samples corresponding to the two smaller predicted samples respectively;
[0461] The first maximum value is used as the average of the two larger predicted samples, and the second maximum value is used as the average of the two reconstructed samples corresponding to the two larger predicted samples; and
[0462] The first difference is defined as the difference between the first maximum value and the first minimum value, and the second difference is defined as the difference between the second maximum value and the second minimum value.
[0463] 76. The non-volatile computer-readable storage medium as described in Clause 75, wherein the first value is the average of the second minimum value and the second maximum value.
[0464] Exemplary embodiments are disclosed in the accompanying drawings and description. However, many variations and modifications can be made to these embodiments. Therefore, although specific terms are used, they are used in a general and descriptive sense only and not for limiting purposes.
Claims
1. A video processing method, comprising: Determine whether to enable inter-frame predictor correction for the coded block; as well as When the inter-frame predictor correction is enabled for the coded block, the inter-frame predictor correction is performed through the following steps: Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block, wherein the prediction samples are located within the prediction block; Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block; as well as Based on the multiple predicted samples, the multiple reconstructed samples, and the predicted block, the corrected prediction block is obtained; Obtaining multiple predicted samples and multiple reconstructed samples includes: The first parameter is determined based on the encoding mode of the encoding block, and the first parameter includes a horizontal position parameter and a vertical position parameter; The second parameter is determined based on the width of the coded block and the horizontal position parameter in the first parameter, wherein the second parameter includes the limiting value of the horizontal position parameter; The third parameter is determined based on the height of the coded block and the vertical position parameter in the first parameter, and the third parameter includes the amplitude limit value of the vertical position parameter; Based on the second parameter and the third parameter, multiple predicted samples and multiple reconstructed samples are obtained.
2. The video processing method according to claim 1, wherein, The corrected prediction block also includes: Based on the plurality of predicted samples and the plurality of reconstructed samples, at least one parameter for the inter-frame predictor correction is obtained; and The corrected prediction block is obtained based on the at least one parameter and the prediction block.
3. The video processing method according to claim 1, wherein, Obtaining the plurality of predicted samples and the plurality of reconstructed samples includes: Multiple horizontal positions can be determined based on the horizontal position parameters, or multiple vertical positions can be determined based on the vertical position parameters; The plurality of prediction samples are obtained from the top boundary of the prediction block based on the plurality of horizontal positions, or from the left boundary of the prediction block based on the plurality of vertical positions; and The multiple reconstructed samples are obtained from the top adjacent reconstructed sample based on the multiple vertical positions, or the multiple reconstructed samples are obtained from the left adjacent reconstructed sample based on the multiple vertical positions.
4. The video processing method according to claim 3, wherein, The value of the horizontal position parameter is equal to the value of the vertical position parameter.
5. The video processing method according to claim 3, wherein, The value of the horizontal position parameter or the value of the vertical position parameter is equal to 4, 8 or 16.
6. The video processing method according to claim 5, further comprising: If the encoding mode of the encoded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal position parameter and the vertical position parameter are set to equal 8; otherwise, the values of the horizontal position parameter and the vertical position parameter are set to equal 16; or If the encoding mode of the encoded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal position parameter and the vertical position parameter are set to equal 4; otherwise, the values of the horizontal position parameter and the vertical position parameter are set to equal 8.
7. The video processing method according to claim 1, wherein determining the second parameter based on the width of the coded block and the horizontal position parameter in the first parameter comprises: If the width of the coded block is greater than the value of the horizontal position parameter, then the limiting value of the horizontal position parameter is set to be equal to the value of the horizontal position parameter; otherwise, the limiting value of the horizontal position parameter is set to be equal to the width of the coded block. The step of determining the third parameter based on the height of the coded block and the vertical position parameter in the first parameter includes: if the height of the coded block is greater than the value of the vertical position parameter, then setting the limiting value of the vertical position parameter to be equal to the value of the vertical position parameter; otherwise, setting the limiting value of the vertical position parameter to be equal to the height of the coded block. The step of obtaining multiple predicted samples and multiple reconstructed samples based on the second parameter and the third parameter includes: obtaining the multiple predicted samples and the multiple reconstructed samples based on the limiting value of the horizontal position parameter or the limiting value of the vertical position parameter.
8. The video processing method according to claim 1, further comprising: A first difference is determined based on the plurality of predicted samples, and a second difference is determined based on the plurality of reconstructed samples; as well as The corrected prediction block is obtained based on the first difference and the second difference.
9. The video processing method according to claim 8, further comprising: If either the first difference or the second difference is less than a threshold, then each corrected prediction sample in the corrected prediction block is set to equal to the first value; otherwise, the corrected prediction block is used as the prediction block.
10. The video processing method according to claim 9, wherein, The threshold depends on the bit depth of the sample value.
11. The video processing method according to claim 10, wherein, The threshold is 1 << (bit depth - 8), where "bit depth" is the bit depth of the sample value, and "<<" indicates a left shift operation.
12. The video processing method according to claim 8, further comprising: The multiple predicted samples are sorted to obtain two smaller predicted samples and two larger predicted samples; The first minimum value is used as the average of the two smaller predicted samples, and the second minimum value is used as the average of the two reconstructed samples corresponding to the two smaller predicted samples respectively; The first maximum value is used as the average of the two larger predicted samples, and the second maximum value is used as the average of the two reconstructed samples corresponding to the two larger predicted samples respectively. as well as The first difference is determined as the difference between the first maximum value and the first minimum value, and the second difference is determined as the difference between the second maximum value and the second minimum value.
13. The video processing method according to claim 12, further comprising: If either the first difference or the second difference is less than a threshold, then each corrected prediction sample in the corrected prediction block is set to the average of the second minimum and the second maximum; otherwise, the corrected prediction block is used as the prediction block.
14. A video processing apparatus, comprising: The memory is configured to store instructions; as well as At least one processor is configured to execute the instructions to cause the device to perform the following steps: Determine whether to enable inter-frame predictor correction for the coded block; and When the inter-frame predictor correction is enabled for the coded block, the inter-frame predictor correction is performed through the following steps: Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block, wherein the prediction samples are located within the prediction block; Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block; as well as Based on the multiple predicted samples, the multiple reconstructed samples, and the predicted block, the corrected prediction block is obtained; Obtaining multiple predicted samples and multiple reconstructed samples includes: The first parameter is determined based on the encoding mode of the encoding block, and the first parameter includes a horizontal position parameter and a vertical position parameter; The second parameter is determined based on the width of the coded block and the horizontal position parameter in the first parameter, wherein the second parameter includes the limiting value of the horizontal position parameter; The third parameter is determined based on the height of the coded block and the vertical position parameter in the first parameter, and the third parameter includes the amplitude limit value of the vertical position parameter; Based on the second parameter and the third parameter, multiple predicted samples and multiple reconstructed samples are obtained.
15. The apparatus according to claim 14, wherein, The at least one processor is configured to execute the instructions to further enable the device to acquire the plurality of predicted samples and the plurality of reconstructed samples through the following steps: Multiple horizontal positions can be determined based on the horizontal position parameters, or multiple vertical positions can be determined based on the vertical position parameters; The multiple prediction samples are obtained from the top boundary of the prediction block based on the multiple horizontal positions, or from the left boundary of the prediction block based on the multiple vertical positions; as well as The multiple reconstructed samples are obtained from the top adjacent reconstructed sample based on the multiple vertical positions, or the multiple reconstructed samples are obtained from the left adjacent reconstructed sample based on the multiple vertical positions.
16. The apparatus according to claim 15, wherein, The at least one processor is configured to execute the instructions to further cause the device to perform the following steps: If the encoding mode of the encoded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal position parameter and the vertical position parameter are set to equal 8; otherwise, the values of the horizontal position parameter and the vertical position parameter are set to equal 16; or If the encoding mode of the encoded block indicates that the prediction block is obtained at the sub-block level, then the values of the horizontal position parameter and the vertical position parameter are set to equal 4; otherwise, the values of the horizontal position parameter and the vertical position parameter are set to equal 8.
17. The apparatus according to claim 15, wherein, The at least one processor is configured to execute the instructions to further cause the device to perform the following steps: If the width of the coded block is greater than the value of the horizontal position parameter, then the limiting value of the horizontal position parameter is set to be equal to the value of the horizontal position parameter; otherwise, the limiting value of the horizontal position parameter is set to be equal to the width of the coded block. If the height of the coded block is greater than the value of the vertical position parameter, then the limiting value of the vertical position parameter is set to be equal to the value of the vertical position parameter; otherwise, the limiting value of the vertical position parameter is set to be equal to the height of the coded block. as well as Based on the amplitude limit value of the horizontal position parameter or the amplitude limit value of the vertical position parameter, the plurality of predicted samples and the plurality of reconstructed samples are obtained.
18. The apparatus according to claim 14, wherein, The at least one processor is configured to execute the instructions to further cause the device to perform the following steps: A first difference is determined based on the plurality of predicted samples, and a second difference is determined based on the plurality of reconstructed samples; as well as The corrected prediction block is obtained based on the first difference and the second difference.
19. The apparatus according to claim 18, wherein, The at least one processor is configured to execute the instructions to further cause the device to perform the following steps: The multiple predicted samples are sorted to obtain two smaller predicted samples and two larger predicted samples; The first minimum value is used as the average of the two smaller predicted samples, and the second minimum value is used as the average of the two reconstructed samples corresponding to the two smaller predicted samples respectively; The first maximum value is used as the average of the two larger predicted samples, and the second maximum value is used as the average of the two reconstructed samples corresponding to the two larger predicted samples respectively. as well as The first difference is determined as the difference between the first maximum value and the first minimum value, and the second difference is determined as the difference between the second maximum value and the second minimum value; as well as If either the first difference or the second difference is less than a threshold, then each corrected prediction sample in the corrected prediction block is set to the average of the second minimum and the second maximum; otherwise, the corrected prediction block is used as the prediction block.
20. A computer-readable storage medium storing an instruction set and a video bitstream, the instruction set being executable by one or more processors of a video processing method to process the video bitstream, the video processing method comprising: Determine whether to enable inter-frame predictor correction for the coded block; as well as When the inter-frame predictor correction is enabled for the coded block, the inter-frame predictor correction is performed through the following steps: Multiple prediction samples are obtained from the top and left boundaries of the prediction block corresponding to the coding block, wherein the prediction samples are located within the prediction block; Multiple reconstructed samples are obtained from the top adjacent reconstructed sample and the left adjacent reconstructed sample of the coded block; as well as Based on the multiple predicted samples, the multiple reconstructed samples, and the predicted block, the corrected prediction block is obtained; Obtaining multiple predicted samples and multiple reconstructed samples includes: The first parameter is determined based on the encoding mode of the encoding block, and the first parameter includes a horizontal position parameter and a vertical position parameter; The second parameter is determined based on the width of the coded block and the horizontal position parameter in the first parameter, wherein the second parameter includes the limiting value of the horizontal position parameter; The third parameter is determined based on the height of the coded block and the vertical position parameter in the first parameter, and the third parameter includes the amplitude limit value of the vertical position parameter; Based on the second parameter and the third parameter, multiple predicted samples and multiple reconstructed samples are obtained.
Citation Information
Patent Citations
Simplified local illumination compensation
US20190260996A1
Method and device for picture encoding and decoding
WO2020106564A2
Parameter derivation for inter prediction
WO2020192717A1
Method and apparatus of local illumination compensation for inter prediction
WO2020236039A1