Video coding method and apparatus, device, system, and storage medium
By identifying candidate weight derivation modes and using template-based prediction modes, the method addresses the inaccuracy in current video coding methods, resulting in improved prediction accuracy and coding performance.
Patent Information
- Application Number
- JP2025521191
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-10-13
- Publication Date
- 2025-10-09
AI Technical Summary
Current video coding methods suffer from inaccurate construction of candidate prediction mode lists, leading to decreased prediction accuracy and coding performance.
The method involves identifying N candidate weight derivation modes and a candidate prediction mode list, which includes prediction modes derived from dividing a template of a current block, to enhance the accuracy of the prediction mode list and improve coding performance.
This approach improves the accuracy of predicting current blocks, thereby enhancing the overall coding performance by accurately deriving prediction modes.
Smart Images

Figure 2025534003000001_ABST
Abstract
Description
[Technical Field]
[0001] The present application relates to the technical field of video coding, and in particular to a video coding method and apparatus, a device, a system, and a storage medium. [Background technology]
[0002] Digital video technology can be incorporated into various video devices, such as digital televisions, smartphones, computers, e-readers, or video players. With the development of video technology, the amount of video data increases. To facilitate the transmission of video data, video compression techniques are used in video devices so that the video data can be transmitted or stored more efficiently.
[0003] Since video has temporal or spatial redundancy, prediction can eliminate or reduce the redundancy in the video and improve compression efficiency. Currently, to improve prediction efficiency, a current block can be predicted using multiple prediction modes. For example, a candidate prediction mode list is constructed, and multiple prediction modes are selected from the candidate prediction mode list to predict the current block. However, the currently constructed candidate prediction mode list is not accurate enough, resulting in a decrease in prediction accuracy of the current block. Summary of the Invention
[0004] In the embodiments of the present application, a video coding method, apparatus, device, system, and storage medium are provided, which improve the accuracy of constructing a candidate prediction mode list, improve the prediction accuracy of a current block, and thereby improve coding performance.
[0005] In a first aspect, the present application provides a video decoding method applied to a decoder. The method includes: identifying N candidate weight derivation modes, where N is a positive integer; identifying a candidate prediction mode list, where the candidate prediction mode list includes at least one candidate prediction mode, the at least one candidate prediction mode including a prediction mode identified based on dividing a template of a current block; identifying a first weight derivation mode and K first prediction modes based on the N candidate weight derivation modes and the candidate prediction mode list, where K is a positive integer greater than 1; predicting the current block based on the first weight derivation mode and the K first prediction modes to obtain a predicted value of the current block.
[0006] In a second aspect, an embodiment of the present application provides a video encoding method, the method including: identifying N candidate weight derivation modes, where N is a positive integer; identifying a candidate prediction mode list, where the candidate prediction mode list includes at least one candidate prediction mode, the at least one candidate prediction mode including a prediction mode identified based on dividing a template of a current block; identifying a first weight derivation mode and K first prediction modes based on the N candidate weight derivation modes and the candidate prediction mode list, where K is a positive integer greater than 1; predicting the current block based on the first weight derivation mode and the K first prediction modes to obtain a predicted value of the current block.
[0007] In a third aspect, the present application provides a video decoding device configured to perform the method of the first aspect or any of its embodiments, specifically the device comprises functional units configured to perform the method of the first aspect or any of its embodiments.
[0008] In a fourth aspect, the present application provides a video encoding apparatus configured to perform the method of the second aspect or any of its embodiments, specifically the apparatus comprises functional units configured to perform the method of the second aspect or any of its embodiments.
[0009] In a fifth aspect, there is provided a video decoder, the video decoder comprising a processor and a memory, the memory configured to store a computer program, the processor configured to perform the method of the first aspect or any embodiment thereof by calling and executing the computer program stored in the memory.
[0010] In a sixth aspect, there is provided a video encoder, the video encoder comprising a processor and a memory, the memory configured to store a computer program, the processor configured to perform the method of the second aspect or any embodiment thereof by calling and executing the computer program stored in the memory.
[0011] In a seventh aspect, there is provided a video coding system, comprising a video encoder and a video decoder, the video decoder configured to perform the method of the first aspect or any of its embodiments, and the video encoder configured to perform the method of the second aspect or any of its embodiments.
[0012] In an eighth aspect, there is provided a chip configured to perform the method of the first or second aspect or each embodiment of the first or second aspect. Specifically, the chip includes a processor configured to call up and execute a computer program from a memory, thereby causing a device equipped with the chip to perform the method of the first or second aspect or each embodiment of the first or second aspect.
[0013] In a ninth aspect, there is provided a computer-readable storage medium configured to store a computer program that causes a computer to perform the method of the first or second aspect or each embodiment of the first or second aspect.
[0014] In a tenth aspect, there is provided a computer program product, the computer program product comprising computer program instructions configured to cause a computer to perform the method of the first or second aspect, or an embodiment of the first or second aspect.
[0015] In an eleventh aspect, there is provided a computer program which, when executed on a computer, is configured to cause the computer to carry out the method of the first or second aspect or a respective embodiment of the first or second aspect.
[0016] In a twelfth aspect, a bitstream is provided, the bitstream being generated according to the method of the second aspect. Optionally, the bitstream includes a first index, the first index being used to indicate a first combination of one weight derivation mode and K prediction modes, where K is a positive integer greater than 1.
[0017] Based on the above technical proposal, N candidate weight derivation modes and a candidate prediction mode list are identified during coding (including encoding and decoding) of the current block. The candidate prediction mode list includes at least one candidate prediction mode, and the at least one candidate prediction mode includes a prediction mode identified based on dividing a template of the current block. That is, in the embodiment of the present application, when identifying the candidate prediction mode, the prediction mode is derived from the template obtained by dividing. This allows for accurate derivation of the prediction mode and further improves the accuracy of identifying the candidate prediction mode list. Performing prediction based on the accurately identified candidate prediction mode list improves prediction accuracy and coding performance. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 1 is a block diagram illustrating a video coding system according to an embodiment of the present application. [Figure 2] FIG. 2 is a block diagram illustrating a video encoder according to an embodiment of the present application. [Figure 3] FIG. 3 is a block diagram illustrating a video decoder according to an embodiment of the present application. [Figure 4] FIG. 4 is a schematic diagram illustrating the weight assignment. [Figure 5] FIG. 5 is a schematic diagram illustrating weight assignment. [Figure 6A] FIG. 6A is a schematic diagram illustrating inter prediction. [Figure 6B] FIG. 6B is a schematic diagram illustrating weighted inter prediction. [Figure 7A] FIG. 7A is a schematic diagram illustrating intra prediction. [Figure 7B] FIG. 7B is a schematic diagram illustrating intra prediction. [Figure 8A] FIG. 8A is a schematic diagram illustrating intra prediction. [Figure 8B] FIG. 8B is a schematic diagram illustrating intra prediction. [Figure 8C] FIG. 8C is a schematic diagram illustrating intra prediction. [Figure 8D] FIG. 8D is a schematic diagram illustrating intra prediction. [Figure 8E] FIG. 8E is a schematic diagram illustrating intra prediction. [Figure 8F] FIG. 8F is a schematic diagram illustrating intra prediction. [Figure 8G] FIG. 8G is a schematic diagram illustrating intra prediction. [Figure 8H] FIG. 8H is a schematic diagram illustrating intra prediction. [Figure 8I] FIG. 8I is a schematic diagram illustrating intra prediction. [Figure 9] FIG. 9 is a schematic diagram illustrating intra prediction modes. [Figure 10] FIG. 10 is a schematic diagram illustrating intra prediction modes. [Figure 11] FIG. 11 is a schematic diagram illustrating intra prediction modes. [Figure 12] FIG. 12 is a schematic diagram illustrating matrix-based intra prediction (MIP). [Figure 13] FIG. 13 is a schematic diagram illustrating template-based intra-mode derivation (TIMD) prediction. [Figure 14A] FIG. 14A is a histogram corresponding to decoder-side intra-mode derivation (DIMD). [Figure 14B] FIG. 14B is a schematic diagram showing DIMD prediction. [Figure 15] FIG. 15 is a schematic diagram illustrating combined prediction. [Figure 16] FIG. 16 is a schematic diagram showing a template. [Figure 17] FIG. 17 is a flowchart illustrating a video decoding method according to an embodiment of the present application. [Figure 18] FIG. 18 is a schematic diagram illustrating template division. [Figure 19] FIG. 19 is a schematic diagram showing adjacent blocks. [Figure 20]FIG. 20 is a schematic diagram illustrating the division of the reconstruction sample area. [Figure 21A] FIG. 21A is a schematic diagram illustrating the identification of a third candidate prediction mode. [Figure 21B] FIG. 21B is another schematic diagram illustrating the identification of a third candidate prediction mode. [Figure 22A] FIG. 22A is a schematic diagram showing a template. [Figure 22B] FIG. 22B is a schematic diagram illustrating the derivation of template weights. [Figure 23A] FIG. 23A is a schematic diagram showing a blending region. [Figure 23B] FIG. 23B is another schematic diagram showing the blending region. [Figure 24] FIG. 24 is a flowchart illustrating a video encoding method according to an embodiment of the present application. [Figure 25] FIG. 25 is a block diagram illustrating a video decoding device according to an embodiment of the present application. [Figure 26] FIG. 26 is a block diagram illustrating a video encoding device according to an embodiment of the present application. [Figure 27] FIG. 27 is a block diagram illustrating an electronic device according to an embodiment of the present application. [Figure 28] FIG. 28 is a block diagram illustrating a video coding system according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION
[0019] The present application is applicable to the fields of image coding, video coding, hardware video coding, dedicated circuit video coding, real-time video coding, etc. For example, the technical solution of the present application can be combined with an audio video coding standard (AVS), and examples of AVS include the H.264 / audio video coding (AVC) standard, the H.265 / high efficiency video coding (HEVC) standard, and the H.266 / versatile video coding (VVC) standard. Alternatively, the technical solution of the present application may be combined with other proprietary or industry standards, including ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, and ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), including scalable video coding (SVC) and multi-view video coding (MVC) extensions. Note that the technology of the present application is not limited to any particular coding standard or technology.
[0020] For ease of understanding, a video coding system according to an embodiment of the present application will first be introduced with reference to FIG.
[0021] FIG. 1 is a block diagram illustrating a video coding system according to an embodiment of the present application. FIG. 1 is merely an example, and video coding systems according to an embodiment of the present application include, but are not limited to, those illustrated in FIG. 1. As illustrated in FIG. 1, the video coding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is configured to encode (which may also be understood as compressing) video data to generate a bitstream and transmit the bitstream to the decoding device. The decoding device is configured to decode the bitstream generated by the encoding device to obtain decoded video data.
[0022] The encoding device 110 according to an embodiment of the present application may be understood as a device having a video encoding function, and the decoding device 120 may be understood as a device having a video decoding function, that is, the encoding device 110 and the decoding device 120 according to an embodiment of the present application may include a wider range of devices, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes (STBs), televisions, cameras, display devices, digital media players, video game consoles, and in-vehicle computers.
[0023] In some embodiments, encoding device 110 may transmit encoded video data (e.g., a bitstream) to decoding device 120 over channel 130. Channel 130 may include one or more media and / or devices that may transmit encoded video data from encoding device 110 to decoding device 120.
[0024] In one example, channel 130 includes one or more communication media that enable encoding device 110 to transmit encoded video data directly to decoding device 120 in real time. In this example, encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to decoding device 120. The communication media can include wireless communication media, such as a radio frequency spectrum. Alternatively, the communication media can also include wired communication media, such as one or more physical transmission lines.
[0025] In another example, channel 130 includes a storage medium that can store the video data encoded by encoding device 110. The storage medium can include various locally accessible data storage media, such as optical disks, digital versatile discs (DVDs), flash memory, etc. In this example, decoding device 120 can retrieve the encoded video data from the storage medium.
[0026] In another example, channel 130 may include a storage server, which may store video data encoded by encoding device 110. In this example, decoding device 120 may download the encoded video data stored on the storage server from the storage server. Alternatively, the storage server may store the encoded video data and transmit the encoded video data to decoding device 120; examples of storage servers include a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.
[0027] In some embodiments, encoding device 110 includes a video encoder 112 and an output interface 113. Output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.
[0028] In some embodiments, encoding device 110 may further include a video source 111 in addition to video encoder 112 and output interface 113 .
[0029] Video source 111 may include at least one of a video acquisition device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system, where the video input interface is configured to receive video data from a video content provider, and the computer graphics system is configured to generate video data.
[0030] The video encoder 112 generates a bitstream by encoding video data from the video source 111. The video data may include one or more pictures or a sequence of pictures. The bitstream includes encoding information for the pictures or the sequence of pictures. The encoding information may include encoding image data and related data. The related data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. An SPS may include parameters that apply to one or more sequences, and a PPS may include parameters that apply to one or more pictures. A syntax structure is a collection of zero or more syntax elements arranged in a specified order in the bitstream.
[0031] The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data may also be stored on a storage medium or a storage server for subsequent reading by the decoding device 120.
[0032] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122 .
[0033] In some embodiments, the decoding device 120 may further include a display device 123 in addition to the input interface 121 and the video decoder 122 .
[0034] The input interface 121 includes a receiver and / or a modem and is capable of receiving encoded video data via a channel 130.
[0035] The video decoder 122 is configured to decode the encoded video data to obtain decoded video data, and transmit the decoded video data to a display device 123 .
[0036] The display device 123 displays the decoded video data and may be built into the decoding device 120 or may be external to the decoding device 120. The display device 123 may include various types of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0037] Furthermore, FIG. 1 is merely an example, and the technical solutions of the embodiments of the present application are not limited to FIG. 1. For example, the technology of the present application can also be applied to one-sided video encoding or one-sided video decoding.
[0038] A video encoding framework according to an embodiment of the present application will now be described.
[0039] 2 is a block diagram illustrating a video encoder according to an embodiment of the present application. The video encoder 200 can be configured to perform lossy or lossless compression on images. The lossless compression can be visually lossless or mathematically lossless compression.
[0040] The video encoder 200 is applicable to image data in a luminance-chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents luminance (luma), Cb (U) represents blue chrominance, and Cr (V) represents red chrominance, and U and V represent chrominance (chroma) components for describing color and saturation. For example, in the chrominance format, 4:2:0 indicates four luminance components and two chrominance components (YYYYCbCr) for every four pixels, 4:2:2 indicates four luminance components and four chrominance components (YYYYCbCrCbCr) for every four pixels, and 4:4:4 indicates full pixel display (YYYYCbCrCbCrCbCrCbCr).
[0041] For example, the video encoder 200 reads video data and, for each image in the video data, divides each image into several coding tree units (CTUs). In some examples, a CTU may be referred to as a "tree block," "largest coding unit (LCU)," or "coding tree block (CTB)." Each CTU may be associated with a pixel block having the same size as the CTU in the image, where each pixel may correspond to one luminance (luma) sample and two chrominance (chroma) samples. Thus, each CTU may be associated with one luminance sample block and two chrominance sample blocks. The size of the CTU may be, for example, 128×128, 64×64, 32×32, etc. The CTU may then be divided into several coding units (CUs) to be encoded. A CU may be a rectangular or square block. The CU can be further divided into a prediction unit (PU) and a transform unit (TU), so that encoding, prediction, and transformation are separated and processing can be flexible. In one example, the CTU is divided into CUs in a quadtree manner, and the CU is divided into TUs and PUs in a quadtree manner.
[0042] Video encoders and video decoders can support various PU sizes. Assuming that a specific CU size is 2N×2N, the video encoder and video decoder can support a PU size of 2N×2N or N×N for intra prediction, and can support symmetric PUs having sizes of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter prediction. The video encoder and video decoder can also support asymmetric PUs having sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter prediction.
[0043] 2, the video encoder 200 may include a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filtering unit 260, a decoded image buffer 270, and an entropy encoding unit 280. Note that the video encoder 200 may include more, fewer, or different functional units.
[0044] Alternatively, in this application, a current block may be referred to as a current CU or a current PU, a prediction block may be referred to as a predicted image block or an image prediction block, and a reconstructed image block may be referred to as a reconstruction block or an image reconstruction block.
[0045] In some embodiments, the prediction unit 210 includes an inter prediction unit 211 and an intra prediction unit 212. Because there is a strong correlation between adjacent samples in a video picture, the video coding technology uses an intra prediction method to eliminate spatial redundancy between adjacent samples. Because there is a strong similarity between adjacent images in a video, the video coding technology uses an inter prediction method to eliminate temporal redundancy between adjacent images, thereby improving coding efficiency.
[0046] The inter prediction unit 211 can be used for inter prediction. Inter prediction may include motion estimation and motion compensation. In inter prediction, image information of a different image can be referenced. The motion information is used to find a reference block from the reference image, and a prediction block is generated based on the reference block, thereby eliminating temporal redundancy. The image used for inter prediction can be a P frame and / or a B frame. A P frame refers to a forward predicted image, and a B frame refers to a bidirectionally predicted image. In inter prediction, the motion information is used to find a reference block from the reference image, and a prediction block is generated based on the reference block. The motion information includes a reference image list including the reference images, a reference image index, and a motion vector. The motion vector can be an integer-sample motion vector or a fractional-sample motion vector. If the motion vector is a fractional-sample motion vector, an interpolation filter must be applied to the reference image to generate the required fractional sample block. The integer sample block or fractional sample block in the reference image found based on the motion vector is called a reference block. There are techniques for using the reference block as a prediction block as it is, and there are techniques for generating a prediction block by processing based on the reference block. Generating a prediction block by processing based on the reference block can also be understood as using the reference block as a prediction block and then processing based on the prediction block to generate a new prediction block.
[0047] The intra prediction unit 212 predicts sample information in the current image block by only referring to information in the same image to eliminate spatial redundancy. The image used for intra prediction may be an I-frame.
[0048] Intra prediction includes various prediction modes. Taking the H series of international digital video coding standards as an example, the H.264 / AVC standard includes eight angular prediction modes and one non-angular prediction mode, while H.265 / HEVC has been expanded to include 33 angular prediction modes and two non-angular prediction modes. The intra prediction modes used in HEVC include planar mode, direct current (DC), and 33 angular modes, for a total of 35 prediction modes. The intra modes used in VVC include planar mode, DC, and 65 angular modes, for a total of 67 prediction modes.
[0049] Furthermore, with an increase in the number of angle modes, intra prediction becomes more accurate and better meets the needs of the development of high-definition and ultra-high-definition digital video.
[0050] The residual unit 220 may generate a residual block of the CU based on the sample block of the CU and the prediction block of the PU of the CU. For example, the residual unit 220 may generate the residual block of the CU such that each sample in the residual block is equal to the difference between a sample in the sample block of the CU and a corresponding sample in the prediction block of the PU of the CU.
[0051] The transform / quantization unit 230 may quantize the transform coefficients. The transform / quantization unit 230 may quantize the transform coefficients associated with the TUs of a CU based on a quantization parameter (QP) value associated with the CU. The video encoder 200 may adjust the degree of quantization applied to the transform coefficients associated with a CU by adjusting the QP value associated with the CU.
[0052] The inverse transform / quantization unit 240 may reconstruct a residual block from the quantized transform coefficients by applying inverse quantization and inverse transform, respectively, to the quantized transform coefficients.
[0053] Reconstruction unit 250 may generate a reconstructed image block associated with a TU by adding samples in the reconstructed residual block to corresponding samples in one or more prediction blocks generated by prediction unit 210. In this manner, by reconstructing the sample blocks of each TU of a CU, video encoder 200 may reconstruct the sample blocks of the CU.
[0054] The loop filtering unit 260 processes the inverse transformed and dequantized samples to compensate for distortion information and provide a better reference for encoding subsequent samples. For example, the loop filtering unit 260 may perform a deblocking filtering process to reduce blocking artifacts of sample blocks associated with a CU.
[0055] In some embodiments, the loop filtering unit 260 includes a deblocking filtering unit and a sample adaptive offset (SAO) / adaptive loop filtering (ALF) unit, where the deblocking filtering unit is used for deblocking and the SAO / ALF unit is used to remove ringing effects.
[0056] The decoding image buffer 270 can store the reconstructed sample blocks. The inter prediction unit 211 can perform inter prediction on a PU of another image using a reference image including the reconstructed sample blocks. The intra prediction unit 212 can perform intra prediction on another PU in the same image as the CU using the reconstructed sample blocks in the decoding image buffer 270.
[0057] The entropy encoding unit 280 may receive the quantized transform coefficients from the transform / quantization unit 230. The entropy encoding unit 280 may generate entropy-coded data by performing one or more entropy coding operations on the quantized transform coefficients.
[0058] FIG. 3 is a block diagram illustrating a video decoder according to an embodiment of the present application.
[0059] 3, the video decoder 300 includes an entropy decoding unit 310, a prediction unit 320, an inverse quantization / transform unit 330, a reconstruction unit 340, a loop filtering unit 350, and a decoded image buffer 360. Note that the video decoder 300 may include more, fewer, or different functional units.
[0060] The video decoder 300 may receive a bitstream. The entropy decoding unit 310 may extract syntax elements from the bitstream by parsing the bitstream. As part of parsing the bitstream, the entropy decoding unit 310 may parse entropy-coded syntax elements in the bitstream. The prediction unit 320, the inverse quantization / transform unit 330, the reconstruction unit 340, and the loop filtering unit 350 may decode the video data based on the syntax elements extracted from the bitstream, i.e., generate decoded video data.
[0061] In some embodiments, the prediction unit 320 includes an intra prediction unit 322 and an inter prediction unit 321 .
[0062] The intra prediction unit 322 may generate a predictive block of the PU by performing intra prediction. The intra prediction unit 322 may generate a predictive block of the PU based on sample blocks of spatially neighboring PUs using an intra prediction mode. The intra prediction unit 322 may further identify the intra prediction mode of the PU based on one or more syntax elements parsed from the bitstream.
[0063] The inter prediction unit 321 may construct a first reference image list (List 0) and a second reference image list (List 1) based on syntax elements parsed from the bitstream. Furthermore, if a PU is encoded using inter prediction, the entropy decoding unit 310 may analyze motion information of the PU. The inter prediction unit 321 may identify one or more reference blocks for the PU based on the motion information of the PU. The inter prediction unit 321 may generate a prediction block for the PU based on the one or more reference blocks for the PU.
[0064] The inverse quantization / transform unit 330 may inverse quantize (i.e., dequantize) the transform coefficients associated with the TU. The inverse quantization / transform unit 330 may use a QP value associated with the CU of the TU to determine the degree of quantization.
[0065] After dequantizing the transform coefficients, the inverse quantization / transform unit 330 may apply one or more inverse transforms to the dequantized transform coefficients to generate a residual block associated with the TU.
[0066] The reconstruction unit 340 reconstructs the sample block of the CU using the residual block associated with the TU of the CU and the prediction block of the PU of the CU. For example, the reconstruction unit 340 can reconstruct the sample block of the CU by adding the samples in the residual block to the corresponding samples in the prediction block to obtain a reconstructed image block.
[0067] The loop filtering unit 350 may perform a deblocking filtering process to reduce block artifacts in the sample blocks associated with the CU.
[0068] The video decoder 300 may store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 may use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or may transmit the reconstructed image to a display device for display.
[0069] The basic flow of video coding is as follows: On the encoding side, an image (frame) is divided into blocks, and a prediction unit 210 performs intra prediction or inter prediction on a current block to generate a predicted block of the current block. A residual unit 220 calculates a residual block based on the difference between the predicted block and the original block of the current block, i.e., the difference between the predicted block and the original block of the current block, and the residual block is also called residual information. The residual block is transformed and quantized by a transform / quantization unit 230, thereby removing information that is not sensitive to the human eye and eliminating visual redundancy. Alternatively, the residual block before being transformed and quantized by the transform / quantization unit 230 may be called a time-domain residual block, and the time-domain residual block after being transformed and quantized by the transform / quantization unit 230 may be called a frequency residual block or a frequency-domain residual block. The entropy encoding unit 280 may receive the quantized transformation coefficients output by the transform / quantization unit 230 and output a bitstream by entropy coding the quantized transformation coefficients. For example, the entropy encoding unit 280 may remove character redundancies based on a target context model and probability information of the binary bitstream.
[0070] On the decoding side, the entropy decoding unit 310 can obtain prediction information, a quantization coefficient matrix, etc., of a current block by analyzing the bitstream. The prediction unit 320 generates a prediction block for the current block by performing intra prediction or inter prediction on the current block based on the prediction information. The inverse quantization / transform unit 330 uses the quantization coefficient matrix obtained from the bitstream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block. The reconstruction unit 340 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed block forms a reconstructed image. The loop filtering unit 350 loop filters the reconstructed image based on an image or block to obtain a decoded image. The encoding side also requires a similar process to that on the decoding side to obtain a decoded image. The decoded image may be called a reconstructed image, and the reconstructed image may be a reference image for inter prediction of a subsequent image.
[0071] In addition, the block division information and mode or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering determined by the encoding side are carried in the bitstream as necessary. The decoding side analyzes the bitstream and determines the same block division information and mode or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering as the encoding side by analyzing the bitstream and analyzing the existing information. This ensures that the decoded image obtained by the encoding side is the same as the decoded image obtained by the decoding side.
[0072] The above is a basic flow of video coding in a block-based mixed coding framework. As technology develops, some modules or steps of the framework or flow may be optimized. This application applies to, but is not limited to, the basic flow of video coding in the block-based mixed coding framework.
[0073] In an embodiment of the present application, the current block may be a current coding unit (CU) or a current prediction unit (PU), etc. According to the need for parallel processing, an image may be divided into slices, etc., and slices within the same image may be processed in parallel, i.e., there is no data dependency between them. The term "frame" is a commonly used term, and one frame may generally be understood to be one image. Furthermore, the term "frame" in the present application may be replaced with "image" or "slice," etc.
[0074] The video coding standard currently under development, Versatile Video Coding (VVC), has an inter-prediction mode called geometric partitioning mode (GPM). The video coding standard currently under development, Audio-Video Coding (AVS), has an inter-prediction mode called angular weighted prediction (AWP). Although these two modes have different names and implementations, they share a common principle.
[0075] While conventional unidirectional prediction uses only one reference block with the same size as the current block, conventional bidirectional prediction uses two reference blocks with the same size as the current block, and the value of each sample in the prediction block is the average of the values of samples at corresponding positions in the two reference blocks, i.e., all samples in each reference block account for 50%. In bidirectional weighted prediction, the proportions of the two reference blocks may be different, for example, all samples in the first reference block account for 75% and all samples in the second reference block account for 25%. However, all samples in the same reference block have the same proportion. However, all samples in the same reference block have the same proportion. Other optimization methods, such as decoder side motion vector refinement (DMVR) technology and bidirectional optical flow (BIO), may cause some changes to the reference samples or prediction samples, but are not related to the above principle. BIO may also be referred to as BDOF. GPM or AWP also uses two reference blocks with the same size as the current block, but at some sample positions, the sample values at the corresponding positions of the first reference block are used 100%, at other sample positions, the sample values at the corresponding positions of the second reference block are used 100%, and in the boundary region (also called the blending region), the sample values at the corresponding positions of these two reference blocks are used at a fixed rate. The weights in the boundary region also gradually transition. How these weights are specifically assigned is determined by the GPM or AWP mode. The weight of each sample position is specified based on the GPM or AWP mode.Of course, in some cases, for example, when the block size is very small, in some GPM or AWP modes, it is not possible to ensure that at some sample positions, the sample values at the corresponding positions of the first reference block are 100% utilized, and at other sample positions, the sample values at the corresponding positions of the second reference block are 100% utilized. In GPM or AWP, two reference blocks having sizes different from the size of the current block are used, that is, a necessary portion of each reference block is considered to be the reference block. That is, the portion with a weight other than 0 is considered to be the reference block, and the portion with a weight of 0 is eliminated. This is a specific implementation and is not the focus of the discussion of the present invention.
[0076] For example, FIG. 4 is a schematic diagram showing weight allocation. As shown in FIG. 4, a schematic diagram of weight allocation for multiple division modes of a GPM in a 64x64 current block according to an embodiment of the present application is shown, where the GPM has 64 division modes. FIG. 5 is a schematic diagram showing weight allocation. As shown in FIG. 5, a schematic diagram of weight allocation for multiple division modes of an AWP in a 64x64 current block according to an embodiment of the present application is shown, where the AWP has 56 division modes. In both FIGS. 4 and 5, for various division modes, a black area indicates a weight value of 0% for the corresponding position of the first reference block, a white area indicates a weight value of 100% for the corresponding position of the first reference block, and a gray area indicates a weight value of greater than 0% and less than 100% with different color shades. The weight value for the corresponding position of the second reference block is 100% minus the weight value for the corresponding position of the first reference block.
[0077] The method of deriving weights in GPM and AWP is different. GPM identifies angles and offsets based on various modes, and then calculates weight matrices for various modes. AWP first generates one-dimensional weight lines, and then tile the entire matrix with the one-dimensional weight lines using a method similar to intra-angle prediction.
[0078] In earlier coding technologies, only rectangular partitioning methods exist, regardless of whether the partitioning is for CUs, PUs, or transform units (TUs). In GPM and AWP, the effect of non-rectangular partitioning of prediction is achieved when no partitioning is performed. In GPM and AWP, a weight mask, i.e., a weight map, of the weights of two reference blocks is used. This mask determines the weights of the two reference blocks used to generate a predicted block. This can be simply understood as follows: Some positions of the predicted block originate from the first reference block, and some positions of the predicted block originate from the second reference block. A blending area is obtained by weighting the corresponding positions of the two reference blocks, resulting in a smoother transition. In GPM and AWP, since the current block is not divided into two CUs or PUs by a partition line, the entire current block is subjected to transform, quantization, inverse transform, inverse quantization, etc. of the predicted residual.
[0079] In GPM, a weight matrix simulates a geometric partition, or more precisely, a prediction partition. To implement GPM, two predictors are required in addition to a weight matrix, and each predictor is specified by one unidirectional motion information. The two unidirectional motion information come from a motion information candidate list, e.g., a merge motion information candidate list (mergeCandList). GPM uses two indices in the bitstream to specify the two unidirectional motion information from mergeCandList.
[0080] Inter-prediction uses motion information to represent "motion." Basic motion information includes reference frame (also called reference picture) information and motion vector (MV) information. In general, bidirectional prediction uses two reference blocks to predict a current block. The two reference blocks may be one forward reference block and one backward reference block. Alternatively, both reference blocks may be forward reference blocks or backward reference blocks. "Forward" means that a time corresponding to a reference picture is located before the current picture, and "backward" means that a time corresponding to a reference picture is located after the current picture. In other words, "forward" means that the position of a reference picture in a video is located before the current picture, and "backward" means that the position of a reference picture in a video is located after the current picture. In other words, "forward" means that the picture order count (POC) of the reference picture is smaller than the POC of the current picture, and "backward" means that the POC of the reference picture is larger than the POC of the current picture. To use bidirectional prediction, two reference blocks must naturally be found. Therefore, two sets of reference image information and motion vector information are required. Each set can be considered as one unidirectional motion information, and combining these two sets results in one bidirectional motion information. In concrete implementation, the same data structure can be used for both unidirectional and bidirectional motion information. In bidirectional motion information, both sets of reference image information and motion vector information are valid, while in unidirectional motion information, one of the two sets of reference image information and motion vector information is invalid.
[0081] In some embodiments, two reference picture lists are supported, denoted as RPL0 and RPL1, where RPL is an abbreviation for Reference Picture List. In some embodiments, P slices can use only RPL0, while B slices can use RPL0 and RPL1. For one slice, each reference picture list contains several reference pictures, and the encoder and decoder use a reference picture index to find a specific reference picture. In some embodiments, motion information is represented by a reference picture index and a motion vector. For example, for the bidirectional motion information, a reference picture index refIdxL0 corresponding to RPL0, a motion vector mvL0 corresponding to RPL0, and a reference picture index refIdxL1 corresponding to RPL1, a motion vector mvL0 corresponding to RPL1 are used. The reference picture index corresponding to RPL0 and the reference picture index corresponding to RPL1 can be understood as the reference picture information. In some embodiments, two flags indicate whether the motion information corresponding to RPL0 is used, and whether the motion information corresponding to RPL0 is used. These flags are denoted as predFlagL0 and predFlagL1, respectively. Furthermore, predFlagL0 and predFlagL1 can be understood to indicate whether the unidirectional motion information is "valid" or "not valid." Although the data structure of motion information is not explicitly stated, the reference image index, motion vector, and "valid" flag corresponding to each reference image list collectively represent motion information. In some standard texts, motion information does not appear, and motion vectors are used instead. The reference image index and the flag indicating whether the corresponding motion information is used can be considered to be appended to the motion vector. For convenience of explanation, the term "motion information" is used in this application, but it should be understood that the term "motion vector" can also be used.
[0082] The motion information used for the current block can be stored. Based on the adjacent positional relationship, the motion information of a previously coded (both encoded and decoded) block (e.g., a neighboring block) can be used for a subsequent block to be coded in the current image. Because it uses spatial correlation, such coded motion information is called spatial motion information. The motion information used for each block of the current image can be stored. Based on the reference relationship, the motion information of a previously coded image can be used for a subsequent image to be coded. Because it uses temporal correlation, such motion information of a previously coded image is called temporal motion information. In a method for storing the motion information used for each block of the current image, a fixed-size matrix, such as a 4x4 matrix, is used as the smallest unit, and one set of motion information is stored independently in each smallest unit. In this way, each time a block is coded, the motion information of that block can be stored in the smallest units corresponding to the block's position. In this way, when spatial motion information or temporal motion information is used, the motion information corresponding to the position can be directly found based on the position. For example, when conventional unidirectional prediction is used for one 16x16 block, all 4x4 minimum units corresponding to the block store the motion information of this unidirectional prediction. When GPM or AWP is used for one block, all minimum units corresponding to the block identify the motion information to be stored in each minimum unit based on the GPM or AWP mode, the first motion information, the second motion information, and the position of each minimum unit. In one method, when all 4x4 samples corresponding to one minimum unit are derived from the first motion information, this minimum unit stores the first motion information. When all 4x4 samples corresponding to one minimum unit are derived from the second motion information, this minimum unit stores the second motion information.If a 4x4 sample corresponding to one minimum unit comes from both the first motion information and the second motion information, AWP selects and stores one of them, while GPM combines and stores the two motion information as bidirectional motion information if they point to different reference image lists; otherwise, it stores only the second motion information.
[0083] Optionally, the mergeCandList is constructed based on spatial motion information, temporal motion information, history-based motion information, and other types of motion information. For example, mergeCandList derives spatial motion information using positions 1 to 5 in FIG. 6A , and derives temporal motion information using positions 6 or 7 in FIG. 6A . For history-based motion information, each time a block is coded, the motion information of the block is added to a first-in first-out (FIFO) list. During the addition, several checks are required, such as whether the information overlaps with existing motion information in the list. In this way, the motion information in the history-based list can be referenced when coding the current block.
[0084] In some embodiments, the syntax description of the GPM is as shown in Table 1.
[0085] [Table 1]
[0086] As shown in Table 1, in merge mode, CIIP (combined inter-intra prediction) or GPM can be used for the current block if regular_merge_flag is not 1. If CIIP is not used for the current block, GPM is used for the current block, that is, as shown in the syntax "if(!ciip_flag[x0][y0])" in Table 1.
[0087] As shown in Table 1 above, GPM requires three pieces of information to be transmitted in the bitstream: merge_gpm_partition_idx, merge_gpm_idx0, and merge_gpm_idx1. x0, y0 are used to specify the coordinates (x0, y0) of the upper left luminance sample of the current block relative to the upper left luminance sample of the image. merge_gpm_partition_idx is used to specify the partition shape of the GPM, which is a "simulated partition" as described above. merge_gpm_partition_idx is an index of the weight matrix derivation mode or weight derivation mode described in this specification. merge_gpm_idx0 is the first merge candidate index, and the first merge candidate index is used to specify the first motion information or the first merge candidate based on the mergeCandList. merge_gpm_idx1 is the second merge candidate index, which is used to identify the second motion information or the second merge candidate based on mergeCandList. merge_gpm_idx1 needs to be decoded only if MaxNumGpmMergeCand > 2, i.e., the length of the candidate list is greater than 2. Otherwise, merge_gpm_idx1 can be identified directly.
[0088] In some embodiments, the GPM decoding process includes the following steps.
[0089] The input information for the decoding process includes the coordinates (xCb, yCb) of the luminance position of the upper left corner of the current block relative to the luminance position of the upper left corner of the image, the width cbWidth of the luminance component of the current block, the height cbHeight of the luminance component of the current block, luminance motion vectors mvA and mvB with 1 / 16 fractional sample accuracy, chrominance motion vectors mvCA and mvCB, reference image indices refIdxA and refIdxB, and prediction list flags predListFlagA and predListFlagB.
[0090] For example, motion information may be represented by a combination of a motion vector, a reference image index, and a prediction list flag. VVC supports two reference image lists, each of which may have multiple reference images. Unidirectional prediction uses only one reference block in one reference image in one reference image list as a reference, while bidirectional prediction uses one reference block in one reference image in one reference image list and one reference block in one reference image in another reference image list as references. GPM in VVC uses two unidirectional predictions. In the above mvA and mvB, mvCA and mvCB, refIdxA and refIdxB, and predListFlagA and predListFlagB, A may be understood as a first prediction mode, and B may be understood as a second prediction mode. X represents A or B, predListFlagX indicates whether the first or second reference image list is used for X, refIdxX indicates the reference image index in the reference image list used for X, mvX indicates the luma motion vector used for X, and mvCX indicates the chroma motion vector used for X. Note that in VVC, the motion information described in this specification can be represented by a combination of the motion vector, reference image index, and prediction list flag.
[0091] The output information of the decoding process includes a (cbWidth)X(cbHeight) matrix of luma prediction samples predSamplesL, an optional (cbWidth / SubWidthC)X(cbHeight / SubHeightC) matrix of Cb chroma prediction samples, and an optional (cbWidth / SubWidthC)X(cbHeight / SubHeightC) matrix of Cr chroma prediction samples.
[0092] For illustrative purposes, the luminance component will be taken as an example, and the processing of the chrominance components is similar to that of the luminance component.
[0093] Assume that the size of predSamplesLAL and predSamplesLBL is (cbWidth)X(cbHeight), and they are matrices of prediction samples obtained based on two prediction modes. predSamplesL is derived based on the following method: predSamplesLAL is determined based on the luma motion vector mvA, the chroma motion vector mvCA, the reference image index refIdxA, and the prediction list flag predListFlagA, and predSamplesLBL is determined based on the luma motion vector mvB, the chroma motion vector mvCB, the reference image index refIdxB, and the prediction list flag predListFlagB. That is, prediction is performed based on the motion information of each of the two prediction modes, and the detailed process is not described in detail. Generally, GPM is a merge mode, and both of the two prediction modes of GPM can be considered to be merge modes.
[0094] Based on merge_gpm_partition_idx[ xCb ][ yCb ], the GPM partition angle index variable angleIdx and distance index variable distanceIdx are identified using Table 2.
[0095] [Table 2]
[0096] Because GPM can be used for any of the three components (e.g., Y, Cb, Cr), some standard texts refer to the process of generating a GPM prediction sample matrix for one component as a subprocess called the weighted sample prediction process for GPM. This subprocess is called for each of the three components, but the parameters used are different. Here, we will use only the luma component as an example. The prediction matrix predSamplesL[xL][yL] (xL = 0..cbWidth - 1, yL = 0..cbHeight - 1) for the current luma block is derived from the weighted sample prediction process for GPM. nCbW is set to cbWidth, nCbH is set to cbHeight, and the prediction sample matrices predSamplesLAL and predSamplesLBL generated by the two prediction modes, angleIdx, and distanceIdx are input.
[0097] In some embodiments, the weighted sample prediction and derivation process for the GPM includes the following steps.
[0098] The inputs of this process include the width nCbW of the current block, the height nCbH of the current block, two (nCbW)X(nCbH) predicted sample matrices predSamplesLA and predSamplesLB, a GPM division angle index variable angleIdx, a GPM distance index variable distanceIdx, and a component index variable cIdx. In this example, luminance is used. A cIdx of 0 indicates the luminance component.
[0099] The output of this process contains the predicted sample values of the GPM's (nCbW)X(nCbH) matrix pbSamples.
[0100] Illustratively, the variables nW, nH, shift1, offset1, displacementX, displacementY, partFlip, and shiftHor are derived based on the following method. nW = ( cIdx = = 0 ) ? nCbW : nCbW * SubWidthC, nH = ( cIdx = = 0 ) ? nCbH : nCbH * SubHeightC, shift1 = Max( 5, 17 - BitDepth ), where BitDepth is the bit depth of the coding. offset1 = 1 << ( shift1 - 1 ), where "<<" means left shift. displacementX = angleIdx, displacementY = (angleIdx + 8) % 32, partFlip = ( angleIdx >= 13 && angleIdx <= 27 ) ? 0 : 1, shiftHor = ( angleIdx % 16 = = 8 | | ( angleIdx % 16 != 0 && nH >= nW ) ) ? 0 : 1.
[0101] The variables offsetX and offsetY are derived based on the following method. If the value of shiftHor is 0, offsetX = ( -nW ) >> 1, offsetY = ( ( -nH ) >> 1 ) + ( angleIdx < 16 ? ( distanceIdx * nH ) >> 3 : -( ( distanceIdx * nH ) >> 3 ) ) . If the value of shiftHor is 1, offsetX = ( ( -nW ) >> 1 ) + ( angleIdx < 16 ? ( distanceIdx * nW ) >> 3 : -( ( distanceIdx * nW ) >> 3 ) , offsetY = ( - nH ) >> 1.
[0102] The variables xL and yL are derived based on the following method. xL = ( cIdx = = 0 ) ? x : x * SubWidthC, yL = ( cIdx = = 0 ) ? y : y * SubHeightC.
[0103] The variable wValue indicating the weight of the prediction sample at the current position is derived based on the following method: wValue is the weight of the value predSamplesLA[ x ][ y ] of the prediction sample in the prediction matrix of the first prediction mode at (x, y), and ( 8 - wValue ) is the weight of the value predSamplesLB[ x ][ y ] of the prediction sample in the prediction matrix of the first prediction mode at (x, y).
[0104] The distance matrix disLut is specified based on Table 3.
[0105] [Table 3]
[0106] weightIdx = ( ( ( xL + offsetX ) << 1 ) + 1 ) * disLut[ displacementX ] + ( ( ( yL + offsetY ) << 1 ) + 1 ) * disLut[ displacementY ] , weightIdxL = partFlip ? 32 + weightIdx : 32 - weightIdx, wValue = Clip3( 0, 8, ( weightIdxL + 4 ) >> 3 ) .
[0107] The predicted sample values pbSamples[ x ][ y ] are derived based on the following method. pbSamples[ x ][ y ] = Clip3( 0, ( 1 << BitDepth ) - 1, ( predSamplesLA[ x ][ y ] * wValue + predSamplesLB[ x ][ y ] * ( 8 - wValue ) + offset1 ) >> shift1 ) .
[0108] Note that one weight value is derived for each position of the current block, and one GPM predicted value pbSamples[x][y] is calculated. In this method, the weight wValue does not need to be written in the form of a matrix; however, if the wValue for each position is stored in a single matrix, the matrix can be understood as a weight matrix. The principle of calculating a GPM predicted value by calculating and weighting each sample is the same as the principle of calculating all weights and then uniformly weighting them to calculate a GPM predicted sample matrix. The term weight matrix is used in many explanations in this application for ease of understanding, and diagrams drawn using weight matrices are more intuitive, but in fact, explanations can also be given using weights for each position. For example, the weight matrix derivation mode can also be expressed as a weight derivation mode.
[0109] In some embodiments, as shown in Figure 6B, the GPM decoding process can be expressed as follows: Analyze the bitstream to determine whether the GPM technique is used for the current block, and if the GPM technique is used for the current block, determine a weight derivation mode (or partition mode, or weight matrix derivation mode), first motion information, and second motion information; Identify a first prediction block based on the first motion information, identify a second prediction block based on the second motion information, identify a weight matrix based on the weight matrix derivation mode, and identify a prediction block for the current block based on the first prediction block, the second prediction block, and the weight matrix.
[0110] In the intra prediction method, a current block is predicted using coded reconstructed samples surrounding the current block as reference samples. FIG. 7A is a schematic diagram illustrating intra prediction. As shown in FIG. 7A, the size of the current block is 4×4, and the samples in the left column and the top row of the current block are reference samples for the current block. In intra prediction, these reference samples are used to predict the current block. All of these reference samples may be available, i.e., may have been coded. Alternatively, some of these reference samples may be unavailable. For example, if the current block is located at the leftmost position of the entire image, the reference sample to the left of the current block may be unavailable. Or, when coding the current block, the reference sample to the bottom left of the current block has not yet been coded, so the reference sample to the bottom left is also unavailable. If a reference sample is unavailable, it may be filled using available reference samples or some other value or method, or it may not be filled.
[0111] 7B is a schematic diagram illustrating intra prediction. As shown in FIG. 7B, a multiple reference line (MRL) intra prediction method can improve coding efficiency by using more reference samples. For example, four reference rows / columns are used as reference samples for a current block.
[0112] Furthermore, there are multiple prediction modes for intra prediction, and FIGS. 8A to 8I are schematic diagrams illustrating intra prediction. As shown in FIGS. 8A to 8I, H.264 can include nine modes when performing intra prediction on a 4×4 block. In mode 0 shown in FIG. 8A, the sample above the current block is copied vertically to the current block as a predicted value. In mode 1 shown in FIG. 8B, the reference sample to the left of the current block is copied horizontally to the current block as a predicted value. In mode 2 (DC) shown in FIG. 8C, the average value of eight samples A to D and I to L is used as the predicted value for all samples. In modes 3 to 8 shown in FIGS. 8D to 8I, the reference sample is copied to a corresponding position on the current block along a certain angle. Because some positions on the current block cannot exactly correspond to the reference sample, it is necessary to use a weighted average value of the reference sample or an interpolated fractional sample of the reference sample.
[0113] Other modes include Planar mode and Planar mode. The number of angular prediction modes has increased with technological development and block size expansion. FIG. 9 is a schematic diagram showing intra prediction modes. As shown in FIG. 9, for example, the intra prediction modes used in HEVC include Planar mode, DC mode, and 33 angular modes, for a total of 35 prediction modes. FIG. 10 is a schematic diagram showing intra prediction modes. As shown in FIG. 10, the intra modes used in VVC include Planar mode, DC mode, and 65 angular modes, for a total of 67 prediction modes. FIG. 11 is a schematic diagram showing intra prediction modes. As shown in FIG. 11, the intra modes used in AVS3 include DC mode, Planar mode, Bilinear mode, PCM, and 62 angular modes, for a total of 66 prediction modes.
[0114] There are also techniques for improving prediction, such as improving fractional sample interpolation of reference samples and filtering predicted samples. For example, the multiple intra prediction filter (MIPF) in AVS3 generates predictions using different filters for blocks of different sizes. For samples at different positions within the same block, one filter is used to generate predictions for samples close to the reference sample, and another filter is used to generate predictions for samples far from the reference sample. Techniques for filtering predicted samples, such as the intra prediction filter (IPF) in AVS3, can filter predictions using a reference sample.
[0115] In intra prediction, coding efficiency can be improved by using an intra mode coding technique using a Most Probable Modes (MPM) list. A mode list is constructed using the intra prediction modes of coded surrounding blocks, intra prediction modes (e.g., neighboring modes) derived based on the intra prediction modes of coded surrounding blocks, and commonly used or highly probable intra prediction modes (e.g., DC, planar, bilinear modes, etc.). Because textures have spatial continuity, referencing the intra prediction modes of coded surrounding blocks utilizes spatial correlation. The MPM can be used to predict the intra prediction mode. That is, the probability that the MPM is used for the current block is considered to be higher than the probability that the MPM is not used for the current block. Therefore, fewer codewords are used for the MPM during binarization, thereby saving overhead and improving coding efficiency.
[0116] In some embodiments, matrix-based intra prediction (MIP) (sometimes referred to as matrix weighted intra prediction) may be used for intra prediction. As shown in FIG. 12, to predict a block having a width of W and a height of H, MIP requires H reconstructed samples in a column to the left of the current block and W reconstructed samples in a row above the current block as input. MIP generates a predicted block based on three steps: averaging of reference samples, matrix vector multiplication, and interpolation. Matrix vector multiplication is the core of MIP. MIP can be considered as a process of generating a predicted block using input samples (reference samples) in a matrix vector multiplication manner. Various matrices are provided for MIP, and differences in prediction methods are reflected in the differences in the matrices. For the same input sample, different results are obtained when different matrices are used. In addition, the processes of averaging and interpolating reference samples are designed with a trade-off between performance and complexity in mind. For blocks with large sizes, the averaging of reference samples can achieve an effect similar to downsampling, allowing the input to be adapted to a relatively small matrix. Interpolation achieves the effect of upsampling. In this way, it is no longer necessary to provide MIP matrices for each block size, but instead only matrices with one or a few specific sizes. As the need for compression performance increases and hardware performance improves, more complex MIPs may appear in next-generation standards.
[0117] MIP mode is a bit similar to Planar, but obviously, MIP mode is more complex and flexible than Planar.
[0118] In some embodiments, an intra prediction technique called template-based intra mode derivation (TIMD) may be used. For example, as shown in FIG. 13, the left region of the current block and the region above the current block are used as templates. Except for boundary cases, when coding the current block, theoretically, reconstructed values can be obtained on the left and upper sides of the current block. This is also the basis of many template matching methods. In TIMD, the left region of the current block and the region above the current block shown in FIG. 13 are used as templates, and adjacent samples on the left and upper sides of the template are used as reference samples for the template. A decoder predicts a template using a certain intra prediction mode and compares the predicted value with the reconstructed value to obtain the cost of the intra prediction mode on the template. Examples of cost calculations include the sum of absolute differences (SAD), sum of absolute transformed difference (SATD), and sum of squared error (SSE). Since the template and the current block are adjacent to each other, they are correlated. Therefore, the behavior of a prediction mode on a template can be used to estimate the behavior of this prediction mode on a current block. In TIMD, several candidate intra-prediction modes are used to predict a template, the costs of the candidate intra-prediction modes on the template are obtained, and the predicted values of one or two intra-prediction modes with the lowest costs are used as the intra-predicted values of the current block.
[0119] Research has shown that if the difference between the two costs corresponding to two intra-prediction modes on a template is not large, a weighted average can be performed on the predicted values of the two intra-prediction modes to improve compression performance. The weights of the predicted values of the two prediction modes are related to the costs. In some embodiments, the weights are inversely proportional to the costs.
[0120] In summary, TIMD utilizes the predictive effect of the intra prediction mode on the template to screen the intra prediction modes and weight the two intra prediction modes according to the cost on the template. The advantage of TIMD is that when the TIMD mode is selected for the current block, it is not necessary to specifically indicate which intra prediction mode is being used, but rather it is derived by the decoder itself through the above process, thereby saving some overhead.
[0121] In some embodiments, an intra prediction technique called decoder-side intra mode derivation (DIMD) may be used. DIMD also derives a prediction mode using reconstructed samples located to the left and above the current block. However, instead of predicting a template, DIMD analyzes the gradient of the reconstructed samples. As shown in FIG. 14A, DIMD analyzes the gradient of the center sample of a window, derives a matching intra prediction mode based on the gradient, and analyzes all samples that need to be checked to obtain the histogram shown in FIG. 14A. Of course, the so-called histogram is merely an aid for understanding, and various simple implementations may be used. In some embodiments, DIMD selects the two intra prediction modes with the highest values in the histogram, adds a planar mode to them, and weights the predicted values of the three intra prediction modes in total. The weights are related to the results of the analysis.
[0122] As an example, the prediction process in DIMD is shown in Figure 14B. Specifically, the intra-prediction modes corresponding to the two highest values in the histogram, i.e., M1 and M2, are selected, and the planar mode is added to them, resulting in a total of three intra-prediction modes. Weights ω1, ω2, and ω3 corresponding to the three intra-prediction modes are identified, and predicted values Pred1, Pred2, and Pred corresponding to the three intra-prediction modes are determined. The predicted values corresponding to the three intra-prediction modes are weighted based on the weights corresponding to the three intra-prediction modes to obtain the final predicted block.
[0123] As can be seen from the above, DIMD can screen intra-prediction modes using gradient analysis of reconstructed samples and weight two intra-prediction modes and planar mode according to the analysis results. The advantage of DIMD is that when DIMD mode is selected for the current block, it is not necessary to specifically indicate which intra-prediction mode is being used; it is determined by the decoder itself through the above process, thereby saving some overhead.
[0124] TIMD and DIMD have many similarities, and in some embodiments, the terms are interchangeable. Both TIMD and DIMD also support weighted predictions for two or more intra-prediction modes.
[0125] In GPM, two inter-prediction blocks are combined using a weight matrix. In fact, the use of a weight matrix can be extended to combine any two prediction blocks, such as two inter-prediction blocks, two intra-prediction blocks, or one inter-prediction block and one intra-prediction block. Furthermore, in screen content coding (SCC), one or two prediction blocks can be intra block copy (IBC) or palette prediction blocks.
[0126] In this application, intra prediction, inter prediction, IBC prediction, and palette prediction are referred to as different prediction methods. For ease of explanation, they are collectively referred to as prediction modes. A prediction mode can be understood as information on which a coder (including both an encoder and a decoder) can generate a prediction block for a current block. For example, in intra prediction, the prediction mode can be an intra prediction mode such as DC, Planar, or various intra angular prediction modes. Of course, some other auxiliary information, such as an optimization method for intra reference samples or an optimization method (e.g., filtering) after generating an initial prediction block, can also be superimposed. For example, in inter prediction, the prediction mode can be skip mode, merge mode, merge with motion vector difference (MMVD) mode, or advanced motion vector predition (AMVP), and can be unidirectional prediction, bidirectional prediction, or multi-hypothesis prediction. When unidirectional prediction is used for the inter prediction mode, one motion information must be identified in one prediction mode, and a prediction block can be identified based on the one motion information. When bidirectional prediction is used for the inter prediction mode, two motion information must be identified in one prediction mode, and a prediction block can be identified based on the two motion information.
[0127] Thus, the information that needs to be specified for the GPM can be expressed as one weight derivation mode and two prediction modes. The weight derivation mode is used to specify a weight matrix or weights, and the two prediction modes are used to specify one prediction block or predicted value, respectively. The weight derivation mode is sometimes called a partitioning mode, but since it is a simulated partitioning, it is referred to as the weight derivation mode in this application.
[0128] Optionally, the two prediction modes can be from the same or different prediction schemes, including but not limited to intra prediction, inter prediction, IBC, and palette.
[0129] A specific example is as follows: GPM is used for the current block. This example is used for an inter-coded block, and merge modes for intra prediction and inter prediction can be used. As shown in Table 4, adding one syntax element, intra_mode_idx, indicates which prediction mode is an intra prediction mode. For example, when intra_mode_idx is 0, it indicates that both prediction modes are inter prediction modes, i.e., mode0IsInter is 1 and mode0IsInter is 1. When intra_mode_idx is 1, it indicates that the first prediction mode is an intra prediction mode and the second prediction mode is an inter prediction mode, i.e., mode0IsInter is 0 and mode0IsInter is 1. When intra_mode_idx is 2, it indicates that the first prediction mode is an inter prediction mode and the second prediction mode is an intra prediction mode, i.e., mode0IsInter is 1 and mode0IsInter is 0. When intra_mode_idx is 3, it indicates that both of the two prediction modes are intra prediction modes, that is, mode0IsInter is 0 and mode0IsInter is 0.
[0130] [Table 4]
[0131] In some embodiments, as shown in Figure 15, the GPM decoding process can be expressed as follows: Analyze the bitstream to determine whether the GPM technique is used for the current block, and if the GPM technique is used for the current block, determine a weight derivation mode (or partition mode, or weight matrix derivation mode), a first prediction mode, and a second prediction mode; Identify a first prediction block based on the first prediction mode, identify a second prediction block based on the second prediction mode, identify a weight matrix based on the weight matrix derivation mode, and identify a prediction block for the current block based on the first prediction block, the second prediction block, and the weight matrix.
[0132] The template matching method is primarily used for inter-prediction. In template matching, the correlation between neighboring samples is used to select a region surrounding the current block as a template. Before coding the current block, the blocks to its left and above it are already coded according to the coding order. Of course, existing hardware decoder implementations do not necessarily ensure that the blocks to its left and above it have already been decoded before decoding the current block. Of course, the reference here is to inter-blocks. For example, in HEVC, when generating a prediction block for an inter-coded block, the prediction process for the inter-block can be performed in parallel without requiring neighboring reconstructed samples. However, for intra-coded blocks, the reconstructed samples to the left and above must be used as reference samples. In theory, the samples to the left and above are available, i.e., this can be achieved by making corresponding adjustments to the hardware design. In contrast, the samples to the right and below are unavailable in the coding order of existing standards such as VVC.
[0133] As shown in FIG. 16, rectangular regions on the left and top of the current block are set as templates. The height of the left template is generally the same as the height of the current block, and the width of the top template is generally the same as the width of the current block. The template may have a different height or width than the current block. The optimal matching position of the template in the reference image is found to identify the motion information or motion vector of the current block. This process can be explained as follows: In a certain reference image, a search is performed within a certain range around a starting position. Search rules such as the search range and search step size can be preset. Each time a position is moved to, the degree of matching between the template corresponding to that position and templates around the current block is calculated. The so-called matching degree can be measured using several distortion costs, such as the sum of absolute difference (SAD) and sum of absolute transformed difference (SATD). Transforms commonly used for SATD include the Hadamard transform and mean-square error (MSE). The smaller the SAD, SATD, MSE, etc. values, the higher the matching degree. The cost is calculated using the template prediction block corresponding to that position and the template reconstruction blocks surrounding the current block. In addition to searching for integer sample positions, fractional sample positions can also be searched. The motion information of the current block can be determined based on the position with the highest degree of matching. By utilizing the relationship between adjacent samples, motion information suitable for the template may also be motion information suitable for the current block. Of course, the template matching method is not necessarily effective for all blocks, so several methods can be used to determine whether the template matching method is used for the current block. For example, a control switch can be used for the current block to indicate whether the template matching method is used.One name for this template matching method is decoder-side motion vector derivation (DMVD). Both the encoder and decoder can use a template to perform a search to derive motion information, or find better motion information based on the original motion information. There is no need to transmit specific motion vectors or motion vector differentials; both the encoder and decoder perform searches according to the same rules, ensuring consistency between encoding and decoding. While template matching can improve compression performance, it also requires the decoder to perform a "search," which introduces a certain degree of decoder complexity.
[0134] Although the above describes a method for applying template matching to inter prediction, the template matching method can also be used for intra prediction. For example, an intra prediction mode is identified using a template. Similarly, for a current block, certain areas above and to the left of the current block can be used as templates, such as the rectangular area on the left and the rectangular area above the current block shown in FIG. 14. When coding the current block, reconstructed samples in the template can be used. This process can be explained as follows: A set of candidate intra prediction modes is identified for the current block, where the candidate intra prediction modes constitute a subset of all available intra prediction modes. Of course, the candidate intra prediction modes may also be a universal set of all available intra prediction modes. They can be identified based on a trade-off between performance and complexity. The set of candidate intra prediction modes can be identified according to MPM or some other rule (e.g., equal-interval screening). For each candidate intra prediction mode, a cost, such as SAD, SATD, or MSE, is calculated using the template. Prediction is performed using the mode to obtain a predicted block, and the cost is calculated using the predicted block and the reconstructed block of the template. A mode with a low cost may be more likely to match the template, and by utilizing the similarity between neighboring samples, an intra-prediction mode that performs well with the template may be an intra-prediction mode that performs well with the current block. One or more modes with a low cost are selected. Of course, the above two steps may be repeated. For example, after selecting one or more modes with a low cost, a set of candidate intra-prediction modes is identified again, costs are calculated again for the newly identified set of candidate intra-prediction modes, and one or more modes with a low cost are selected. This may also be understood as coarse selection and fine selection.The finally selected intra prediction mode is determined as the intra prediction mode of the current block, or the finally selected multiple intra prediction modes are determined as candidate intra prediction modes of the current block. Of course, the set of candidate intra prediction modes can also be sorted using only the template matching method, for example, by sorting the MPM list, i.e., for each mode in the MPM list, a predicted block is obtained against a template to determine a cost, and the modes are sorted in ascending order of cost. Generally, the earlier a mode is located in the MPM list, the less overhead it will have in the bitstream. This can improve compression efficiency.
[0135] A template matching method can be used to identify two prediction modes of the GPM. When the template matching method is used in the GPM, one control switch can be used for the current block to control whether template matching is used for the two prediction modes of the current block, or two control switches can be used to control whether template matching is used for each of the two prediction modes.
[0136] Another aspect is how to use template matching. For example, when a GPM is used in merge mode, for example, in the GPM of VVC, one piece of motion information is identified from mergeCandList using merge_gpm_idxX, where X is 0 or 1. For the Xth motion information, one method is to optimize using a template matching method based on the motion information. That is, when one piece of motion information is identified from mergeCandList based on merge_gpm_idxX and template matching is used for the motion information, optimization is performed based on the motion information using the template matching method. Another method is to identify one piece of motion information by directly searching based on default motion information, rather than identifying one piece of motion information from mergeCandList using merge_gpm_idxX.
[0137] If the Xth prediction mode is an intra prediction mode and a template matching method is used for the Xth prediction mode of the current block, the intra prediction mode can be identified using the template matching method without needing to indicate the index of the intra prediction mode in the bitstream, or it is necessary to identify a candidate set or an MPM list using the template matching method and indicate the index of the corresponding intra prediction mode in the bitstream.
[0138] In the intra prediction and inter prediction methods for GPM, a prediction value in GPM is obtained by weighting one intra prediction value and one inter prediction value using the weight of the GPM mode. The method of deriving prediction mode information (motion information) for inter prediction is similar to that in the VVC standard. For a prediction mode for intra prediction, an intra prediction mode candidate list must be constructed for the corresponding portion in the GPM mode. This list may be referred to as an MPM list. The encoder also signals (i.e., writes) the index of the intra prediction mode selected for the current block to the bitstream. During decoding, the decoder constructs the MPM list for the GPM mode using the same method and identifies the intra prediction mode based on the index of the intra prediction mode obtained through decoding. For example, the corresponding portion in the GPM mode may be understood as a white portion or a black portion in the division diagram of FIG. 4 or FIG. 5. For ease of description, these may be referred to as a first portion and a second portion hereinafter. For example, the first portion is a white portion, and the second portion is a black portion. The first part corresponds to the first prediction mode, and the second part corresponds to the second prediction mode. The first and second parts are more intuitive and easier to understand, but may not actually appear in a concrete algorithm.
[0139] When constructing an MPM list of intra prediction modes for a corresponding portion in the weight derivation mode in the GPM, several types of pre-defined intra prediction modes are sequentially added to the MPM list until the length of the list reaches 3. Optionally, the several types of pre-defined intra prediction modes include DIMD-derived intra prediction modes and TIMD-derived intra prediction modes.
[0140] As can be seen from the above, a GPM has three elements: one weight matrix and two prediction modes. The advantage of a GPM is that the weight matrix allows for more independent combinations. On the other hand, a GPM requires specifying more information, which increases the overhead in the bitstream. Taking a GPM as an example, a GPM is selectively used in merge mode. In the bitstream, merge_gpm_partition_idx specifies the weight matrix, merge_gpm_idx0 specifies the first prediction mode, and merge_gpm_idx1 specifies the second prediction mode. There are multiple possible options for each of the weight matrix and the two prediction modes. For example, there are 64 possible options for the weight matrix in VVC. For merge_gpm_idx0 and merge_gpm_idx1, there are a maximum of six possible options in VVC. VVC specifies that merge_gpm_idx0 and merge_gpm_idx1 are different. Accordingly, there are 65x6x5 possible options for a GPM. Furthermore, when MMVD is used to optimize two pieces of motion information (prediction modes), multiple possible options can be provided for each prediction mode, resulting in a fairly large number. On the other hand, it can be seen that a template matching method can also be used to optimize two pieces of motion information (prediction modes), thereby providing more options. However, even if template matching is used to optimize two pieces of motion information (prediction modes), current technological advances require a block-level switch that indicates whether template matching is used for the current block.
[0141] If two intra-prediction modes are used in GPM, and 67 common intra-prediction modes in VVC are applicable to each prediction mode, and the two intra-prediction modes are different, there are 64 x 67 x 66 possible choices. Of course, to save overhead, it is possible to restrict each prediction mode to only a subset of all common intra-prediction modes, but there are still many possible choices.
[0142] When one intra prediction mode and one inter prediction mode are used in the GPM, the number of possible choices can be inferred based on the above cases of intra prediction mode and inter prediction mode.
[0143] In some embodiments, the indication of one weight derivation mode and two prediction modes in the GPM is signaled to the bitstream using respective syntax elements and parsed from the bitstream. That is, one weight derivation mode has its own syntax element or elements, the first prediction mode has its own syntax element or elements, and the second prediction mode has its own syntax element or elements. Of course, the standard may restrict the second prediction mode and the first prediction mode in some cases, such as that they cannot be the same, or that some optimization methods can be used for both two prediction modes (which can also be understood as being used for the current block). However, the three are relatively independent in terms of signaling (i.e., writing) and parsing the syntax elements. This so-called relative independence can be understood as having a certain degree of relevance, but after the restrictions are removed, other possible options remain independent.
[0144] For events with equal probability, fixed-length coding is relatively appropriate. Also, when the probability is clearly defined, coding efficiency can be improved by using short codes for events with high probability and long codes for events with low probability. In addition, the probability estimation for the two different dimensional modes, weight derivation mode and prediction mode, is separate.
[0145] A weight derivation mode and two prediction modes are used together to generate a predicted block. This predicted block is applied to the current block. Therefore, the weight derivation mode and the two prediction modes are related. For example, the current block contains the edges of two objects that are moving relative to each other. This is an ideal scene for inter-GPM. In theory, this "segmentation" should occur at the edges of the objects. However, in practice, the possible "segmentation" options are limited and cannot cover all edges. A close "segmentation" may be selected. There may be multiple close "segments." The selection depends on which combination of "segmentation" and two prediction modes produces the best results. Similarly, the selection of a prediction mode may depend on which combination produces the best results. This is because, in natural video, it is difficult to perfectly match the current block with the corresponding prediction mode, even in areas where this prediction mode is used. The final selection may lead to the highest coding efficiency. Another common scenario for GPM is when the current block contains areas where objects are moving relative to each other. For example, there may be twisting or distortion caused by arm swinging. In such areas, the "division" becomes more complex and may ultimately depend on which combination produces the best results. Another scenario is intra-prediction. In natural images, some parts have complex textures, some parts gradually change from one texture to another, and some parts cannot be represented in a simple one-way direction. Therefore, intra-GPM can provide more complex prediction blocks. Intra-coded blocks generally have larger residuals than inter-coded blocks under the same quantization level. Therefore, the selection of a prediction mode may ultimately depend on which combination produces the best results.
[0146] The "combination" referred to in various places above means selecting a combination of a weight derivation mode and a prediction mode by considering the weight derivation mode and the prediction mode in combination, rather than selecting the weight derivation mode and the prediction mode separately in two dimensions or three dimensions. When expressed as a syntax element, the "combination" syntax element can be used to specify one weight derivation mode and two prediction modes based on the combination.
[0147] That is, the encoder and decoder can each generate the same N candidate combinations. For example, both the encoder and the decoder build a list containing the N candidate combinations. Each candidate combination can be used to derive a combination of one weight derivation mode and two prediction modes. The encoder only needs to signal which candidate combination it finally selected to the bitstream, and the decoder only needs to analyze which candidate combination it finally selected. In this application, this list is referred to as a GPM combination candidate list or a candidate combination list.
[0148] In one example, the combinations in the GPM combination candidate list are sorted roughly in descending order of the probability that the combination will be selected. In this case, the candidate combinations sorted first are assigned shorter code words than in existing methods, while some combinations that are less likely to be selected are assigned longer code words. This can improve overall coding efficiency. In this way, compared to existing methods that are divided into three parts, the method of the present technical proposal theoretically provides greater flexibility and makes it easier to approximate the most efficient correspondence between probabilities and code words.
[0149] Of course, as mentioned above, GPM can sometimes have a fairly large number of possible combinations. To represent a fairly large number of candidates, longer codewords are required. However, by eliminating some combinations with low probability of occurrence in advance, the cost of combinations with high probability of occurrence can be reduced. While existing methods can partially eliminate cases with low probability of occurrence, adopting a combination-based method allows for more flexible response. For example, in conventional methods, once a "split" is eliminated, all possible options that include that "split" are eliminated.
[0150] Another advantage is that the syntax is simpler and there is no need to judge between different cases when parsing.
[0151] How to encode gpm_cand_idx depends on the probability, as mentioned above. For example, Exponential-Golomb coding is used. If the number of candidates is small, that is, if only a few most probable modes can be selected, fixed-length coding can also be used. For example, if there are only 16 candidates, all 16 candidates are encoded with the same bit length.
[0152] A different number of candidate combinations may be set for blocks having different sizes. For example, for small blocks, prediction results under similar weight derivation modes or prediction modes are slightly different, while for large blocks, prediction results under similar weight derivation modes or prediction modes have a larger difference. Therefore, one method is to set a smaller number of candidate combinations for small blocks and a larger number of candidate combinations for large blocks. The size of a block may be determined based on the width or height of the block or the number of samples in the block. As an example, the number of candidates is set to 8 for a block with a sample number less than (or equal to or less than) 256, and the number of candidates is set to 16 for a block with a sample number greater than (or greater than) 256.
[0153] The process of constructing the GPM combination candidate list is explained below.
[0154] In some embodiments, more relevant information can be used to analyze the probability of occurrence of each combination, such as using mode information of neighboring blocks or reconstructed samples.
[0155] In one approach, the GPM combination candidate list is constructed based on a template.
[0156] Generally, the height of the upper template corresponds to the width of the left template, and the value may be 1, 2, 4, etc. As an example, when constructing a GPM combination candidate list based on templates, it is possible to appropriately reduce the computational complexity by using an upper template with a height of 1 and / or a left template with a width of 1. Note that an upper template with a height of 1 can be understood as the upper template of the current block including one row of decoded or encoded samples adjacent to the upper side of the current block. A left template with a width of 1 can be understood as the left template of the current block including one column of decoded or encoded samples adjacent to the left side of the current block.
[0157] When a template is used, more related information, i.e., reconstructed information around the current block, can be used for the current block, so the relationship between the above three elements can be better utilized. It can also be said that the current block is estimated using reconstructed information around the current block.
[0158] One method is to predict the template for each combination using the GPM method to obtain a predicted block of the template under the combination. Since the reconstruction value for the template is obtained, the predicted block of the template under the combination and the reconstructed block of the template can be used to calculate the predicted distortion cost. For example, SAD, SATD, or SSE can be calculated. Each combination is sorted according to its predicted distortion cost, or a list is constructed in which only the first N combinations with the smallest predicted distortion cost are kept. In this way, a GPM combination candidate list can be constructed.
[0159] In the above-described method, for a certain combination, a first prediction mode is used to generate a first predicted value of the template, a second prediction mode is used to generate a second predicted value of the template, a weight derivation mode is used to derive weights for sample positions within the template, and a predicted value of the template is determined based on the first predicted value, the second predicted value, and the weights.
[0160] To ensure consistency between encoding and decoding, the encoder and decoder must use the same method to construct the GPM combination candidate list. As mentioned above, the number of all possible GPM combinations can be quite large. The above method is exhaustive. In a specific implementation, the GPM combination candidate list can be constructed using a fast algorithm, but the algorithm used by the encoder and the decoder should be the same. For example, hierarchical screening can be performed on various combinations. Alternatively, some combinations that are estimated to have a high probability based on known information can be checked preferentially, and an early termination condition can be set.
[0161] In some embodiments, the embodiments are applied to intra-coded blocks but not to blocks subject to screen content coding. This does not mean that the present technical solution cannot be applied to blocks subject to screen content coding, but merely aims to explain the present technical solution using the simplest example. This is because, for blocks that are intra-coded but not subject to screen content coding, only the intra prediction mode needs to be considered, and there is no need to consider modes for screen content coding such as IBC or palette, or various inter modes. As described above, the present technical solution can be applied in all cases where GPM is applicable.
[0162] Assume that the GPM has 64 possible weight derivation modes and 67 possible intra prediction modes. These can be found in the VVC standard. However, the number of possible weights for the GPM is not limited to only 64, nor is it limited to a specific 64. On the other hand, the selection of 64 GPMs in VVC considers a trade-off between improved prediction performance and reduced bitstream overhead. Furthermore, since the proposed technique does not use fixed logic to encode weight derivation modes, theoretically, a greater variety of weights can be used in the proposed technique, and the technique can be used more flexibly. Similarly, the number of intra prediction modes for the GPM is not limited to only 67, nor is it limited to a specific 67. Theoretically, all possible intra prediction modes can be used for the GPM. For example, by making the intra angle prediction modes more granular and generating more intra angle prediction modes, more intra angle prediction modes can be used for the GPM. For example, the MIP mode in VVC can be used in this technical solution, but considering that MIP has multiple sub-modes that can be selected, MIP is not added to this embodiment for ease of understanding.In addition, there is also a wide-angle mode, which can also be used in this technical solution, but its description is omitted in this embodiment.
[0163] If two intra prediction modes are not allowed to be the same, a total of 64*67*66 possible combinations exist in this embodiment. When an exhaustive method is adopted, all of these possible combinations are used to predict templates and calculate the distortion costs of these combinations. It is not necessary to try several intra prediction modes. This is because an MPM list for the current block can be derived based on the prediction modes for neighboring blocks. For example, in VVC, an MPM list with a length of 6 can be derived for the current block. Furthermore, in subsequent technological advances, a secondary MPM solution exists, which can derive an MPM list with a length of 22. This solution can also be said to derive the sum of the lengths of the first and second MPM lists as 22. In this technical solution, MPM can be used to screen intra prediction modes. Of course, an MPM list applicable to the GPM mode for the current block can also be constructed. For example, prediction modes used in all neighboring blocks of the current block are added to the MPM list. For example, if the MPM list does not include a special prediction mode such as DC, horizontal prediction mode, or vertical prediction mode, one or more special prediction modes are added to the candidate intra-prediction modes of the proposed technique. For example, an intra-prediction mode associated with a weight division line is added to the candidate intra-prediction modes of the proposed technique. In one example, one or more intra-angle prediction modes parallel or nearly parallel to the division line are added. In another example, one or more intra-angle prediction modes perpendicular or nearly perpendicular to the division line are added. Alternatively, the candidate intra-prediction modes of the proposed technique may be determined based on the weight derivation mode. Alternatively, the candidate intra-prediction modes of the proposed technique may be determined for each of two intra-prediction modes. In summary, at least one GPM intra-prediction mode candidate set / list is obtained. Of course, the total number of usable prediction modes may be limited to reduce the complexity of the decoding side. For example, there is a limit of up to six prediction modes. The above methods may be used separately or in any combination.
[0164] As can be seen from the above description, when using intra-prediction modes in a GPM, it is necessary to construct an MPM list or to screen a list or set of candidate prediction modes. This is advantageous for reducing overhead or complexity. As an example of complexity reduction, in the above-mentioned GPM combination-based coding, screening intra-prediction modes reduces the number of possible combinations that need to be tried, thereby reducing the amount of calculation and complexity. However, currently, the constructed candidate prediction mode list is not accurate enough. For example, several preset types of prediction modes are identified as candidate prediction modes, which reduces the prediction accuracy of the current block.
[0165] To solve the above technical problems, in an embodiment of the present application, N candidate weight derivation modes and a candidate prediction mode list are identified during coding of a current block. The candidate prediction mode list includes at least one candidate prediction mode, and the at least one candidate prediction mode includes a prediction mode identified based on dividing a template of the current block. That is, in the embodiment of the present application, when identifying the candidate prediction modes, the prediction mode is derived from the template obtained by the division. This allows for accurate derivation of the prediction mode, and further, by performing prediction based on the accurately derived prediction mode, prediction accuracy can be improved, thereby improving coding performance.
[0166] Hereinafter, referring to FIG. 17, a video decoding method according to an embodiment of the present application will be described taking the decoding side as an example.
[0167] 17 is a flowchart showing a video decoding method according to an embodiment of the present application. The embodiment of the present application is applied to the video decoder shown in FIG. 1 and FIG. 3. As shown in FIG. 17, the method of the embodiment of the present application includes:
[0168] S101: Identify N candidate weight derivation modes.
[0169] N is a positive integer. Optionally, N is a preset value or a default value. Optionally, N is indicated to the decoding side by the encoding side. For example, the encoding side identifies N candidate weight derivation modes and signals N in the bitstream. In this way, the decoding side obtains N by decoding the bitstream. Optionally, N may be identified by other methods on the decoding side, and the embodiments of the present application are not limited thereto.
[0170] As can be seen from the above, in the embodiment of the present application, a predicted block is generated based on one weight derivation mode and K prediction modes, and the predicted block is applied to the current block, that is, weights are determined based on the weight derivation mode, the current block is predicted based on the K prediction modes to obtain K predicted values, and the K predicted values are weighted based on the weights to obtain a predicted value of the current block.
[0171] That is, when decoding a current block, the decoding side needs to identify N candidate weight derivation modes and multiple candidate prediction modes. Then, one weight derivation mode is selected from the N candidate weight derivation modes, and K prediction modes are selected from the multiple candidate prediction modes. Then, the current block is predicted using the selected one weight derivation mode and the K prediction modes to obtain a predicted value of the current block.
[0172] In the embodiment of the present application, the method for identifying the N candidate weight derivation modes by the decoding side is not limited.
[0173] In one possible embodiment, there are 56 weight derivation modes in the AWP and 64 weight derivation modes in the GPM, and the N candidate weight derivation modes include at least one weight derivation mode of the 56 weight derivation modes in the AWP or at least one weight derivation mode of the 64 weight derivation modes in the GPM.
[0174] In one possible embodiment, several weight derivation modes in the AWP or GPM can be screened to obtain N candidate weight derivation modes. That is, the N candidate weight derivation modes in the embodiment of the present application are a subset of all weight derivation modes in the AWP or GPM. For example, in a weight derivation mode, the same "division" angle can correspond to multiple offsets. For example, weight derivation modes 10, 11, 12, and 13 shown in FIG. 4 or FIG. 5 have the same "division" angle but different offsets. In the embodiment of the present application, modes corresponding to several offsets can be eliminated. Of course, modes corresponding to several "division" angles can also be eliminated. In this way, the total number of possible combinations can be reduced, thereby making the differences between each possible combination more apparent. Of course, different screening methods can be set for blocks of different sizes. For example, fewer weight derivation modes are used for small blocks and more weight derivation modes are used for large blocks. Also, different screening methods can be set for blocks of different shapes. The shape of the block can refer to the ratio of width to height.
[0175] In this embodiment, the screening method of the N candidate weight derivation modes on the encoding side is the same as that on the decoding side. In one example, the screening method of the N candidate weight derivation modes is default on both the encoding side and the decoding side. In another example, the encoding side instructs the screening method of the N candidate weight derivation modes to the decoding side, so that the decoding side screens the same N candidate weight derivation modes as the encoding side in the same way.
[0176] In some embodiments, weight derivation modes corresponding to preset division angles and / or preset offsets are eliminated from the preset M weight derivation modes to obtain N weight derivation modes. The same division angle in a weight derivation mode may correspond to multiple offsets; for example, as shown in Figure 4, weight derivation modes 10, 11, 12, and 13 have the same division angle but different offsets. Therefore, weight derivation modes corresponding to some preset offsets and / or weight derivation modes corresponding to some preset division angles may be eliminated.
[0177] In some embodiments, the screening conditions corresponding to different blocks may be different. Thus, when identifying the N weight derivation modes corresponding to the current block, the screening conditions corresponding to the current block are first identified, and the N weight derivation modes are selected from the M preset weight derivation modes based on the screening conditions corresponding to the current block.
[0178] In some embodiments, the screening conditions corresponding to the current block include screening conditions corresponding to the size of the current block and / or screening conditions corresponding to the shape of the current block. During prediction, for relatively small blocks, the difference in the impact of similar weight derivation modes on the prediction result is not significant. For relatively large blocks, the difference in the impact of similar weight derivation modes on the prediction result is more obvious. Based on this, in embodiments of the present application, different N values are set for blocks of different sizes, i.e., a relatively large N value is set for relatively large blocks, and a relatively small N value is set for relatively small blocks.
[0179] In one possible embodiment, N candidate weight derivation modes are presented to the decoding side.
[0180] In some embodiments, the screening condition comprises an array having N elements, the N elements corresponding to the N weight derivation modes in a one-to-one correspondence, and the element corresponding to each weight derivation mode is used to indicate whether the weight derivation mode is available.
[0181] The array may be a one-dimensional array or a two-dimensional array.
[0182] For example, taking GPM as an example, the total number of possible weight derivation modes is 64. On the encoding side, a lookup table containing 64 elements is set up, and the value of each element indicates whether the weight derivation mode corresponding to the element is used.
[0183] In one example, taking a one-dimensional array as an example, a specific example is as follows: An array of g_sgpm_splitDir is set. g_sgpm_splitDir
[64] = { 1,1,1,0,1,0,1,0, 1,0,1,0,1,0,1,0, 1,0,1,1,1,0,1,0, 1,0,1,0,1,0,1,0, 0,0,0,0,1,1,0,1, 0,0,1,0,0,1,0,0, 1,0,1,1,0,1,0,0, 1,0,0,1,0,0,1,0 }.
[0184] If the value of g_sgpm_splitDir[x] is 1, it indicates that the weight derivation mode with index x is available. Otherwise, it indicates that the weight derivation mode with index x is not available. In this example, 26 candidate weight derivation modes are identified on the decoding side based on this array.
[0185] In another example, the N candidate weight derivation modes can be represented by one array. The array contains only the indices of the available weight derivation modes. For example, the 26 candidate weight derivation modes are represented by an array g_sgpm_splitDir
[26] ={0, 1, 6, 8, 10, 12, 14, 16, 18, 19, 20, 22, 24, 26, 28, 30, 36, 37, 42, 45, 48, 50, 51, 53, 56, 59} having a length of 26. On the decoding side, the weight derivation modes corresponding to the indices are identified as candidate weight derivation modes based on the indices of the weight derivation modes contained in the array, thereby obtaining 26 candidate weight derivation modes.
[0186] In some embodiments, if the screening conditions corresponding to the current block include a screening condition corresponding to the size of the current block and a screening condition corresponding to the shape of the current block, and if, for the same weight derivation mode, both the screening condition corresponding to the size of the current block and the screening condition corresponding to the shape of the current block indicate that the weight derivation mode is available, the weight derivation mode is identified as one of the N weight derivation modes. If at least one of the screening condition corresponding to the size of the current block and the screening condition corresponding to the shape of the current block indicates that the weight derivation mode is unavailable, the weight derivation mode does not belong to the N weight derivation modes.
[0187] In some embodiments, screening conditions corresponding to different block sizes and screening conditions corresponding to different block shapes can each be realized using multiple arrays.
[0188] In some embodiments, the screening conditions corresponding to different block sizes and the screening conditions corresponding to different block shapes can be implemented in a two-dimensional array, i.e., the two-dimensional array includes both the screening conditions corresponding to the block sizes and the screening conditions corresponding to the block shapes.
[0189] For example, the screening conditions corresponding to a block having a size of A and a shape of B are shown below: The screening conditions are expressed as a two-dimensional array. g_sgpm_splitDir
[64] = { (1,1),(1,1),(1,1),(1,0),(1,0),(0,0),(1,0),(1,1), (1,1),(0,0),(1,1),(1,0),(1,0),(0,0),(1,0),(1,1), (0,1),(0,0),(1,1),(0,0),(1,0),(0,0),(1,0),(0,0), (1,1),(0,0),(0,1),(1,0),(1,0),(1,0),(1,0),(0,0), (0,0),(0,0),(1,1),(0,0),(1,1),(1,1),(1,0),(0,1), (0,0),(0,0),(1,1),(0,0),(1,0),(0,0),(1,0),(0,0), (1,0),(0,0),(1,1),(1,0),(1,0),(1,0),(0,0),(0,0), (1,1),(0,0),(1,1),(0,0),(0,0),(1,0),(1,1),(0,0) }.
[0190] All values of g_sgpm_splitDir[x] equal to 1 indicate that the weight derivation mode with index x is available, and one of the values of g_sgpm_splitDir[x] equal to 0 indicates that the weight derivation mode with index x is unavailable. For example, g_sgpm_splitDir[4]=(1,0) indicates that weight derivation mode 4 is available for blocks with size A but unavailable for blocks with shape B. Therefore, if the size of a block is A and the shape of the block is B, the weight derivation mode is unavailable.
[0191] Note that the above is an example of 64 weight derivation modes in GPM, but the weight derivation modes of the embodiments of the present application include, but are not limited to, 64 weight derivation modes in GPM and 56 weight derivation modes in AMP.
[0192] In some embodiments, before identifying N candidate weight derivation modes, the decoding side needs to identify whether K different prediction modes are used for weighted prediction for the current block. If the decoding side identifies K different prediction modes to be used for weighted prediction for the current block, the aforementioned S101 is performed to identify N candidate weight derivation modes. If the decoding side identifies K different prediction modes not to be used for weighted prediction for the current block, step S101 is skipped.
[0193] In a possible embodiment, the decoding side can specify whether K different prediction modes are used for the current block for weighted prediction by specifying a prediction mode parameter of the current block.
[0194] Optionally, in an embodiment of the present application, the prediction mode parameter may indicate whether GPM mode or AWP mode can be used for the current block, i.e., whether K different prediction modes can be used for prediction of the current block.
[0195] Note that, in an embodiment of the present application, the prediction mode parameter may be understood as a flag indicating whether the GPM mode or the AWP mode is used. Specifically, the encoder may use one variable as the prediction mode parameter, thereby setting the value of the variable to set the prediction mode parameter. Exemplarily, in the present application, if the GPM mode or the AWP mode is used for the current block, the encoder may set the value of the prediction mode parameter to indicate that the GPM mode or the AWP mode is used for the current block. Specifically, the encoder may set the value of the variable to 1. Exemplarily, in the present application, if the GPM mode or the AWP mode is not used for the current block, the encoder may set the value of the prediction mode parameter to indicate that the GPM mode or the AWP mode is not used for the current block. Specifically, the encoder may set the value of the variable to 0. Furthermore, in an embodiment of the present application, after completing the setting of the prediction mode parameter, the encoder may signal the prediction mode parameter to a bitstream and transmit it to the decoder. As a result, the decoder can obtain the prediction mode parameter after parsing the bitstream.
[0196] Based on the above, the decoding side decodes the bitstream to obtain a prediction mode parameter, and determines whether the GPM mode or the AWP mode is used for the current block based on the prediction mode parameter. If the GPM mode or the AWP mode is used for the current block, that is, if K different prediction modes are used for prediction, N candidate weight derivation modes corresponding to the current block are identified.
[0197] In some embodiments, the embodiment of the present application may also set a condition regarding whether the GPM mode or the AWP mode is used for the current block, that is, if it is determined that the current block satisfies the preset condition, K prediction modes are identified to be used for the current block for weighted prediction, and then N candidate weight derivation modes corresponding to the current block are identified.
[0198] For example, when using GPM mode or AWP mode, the size of the current block can be limited.
[0199] As can be seen, in the prediction method according to an embodiment of the present application, K predicted values need to be generated using K different prediction modes, and then the K predicted values are weighted based on the weights to obtain a predicted value for the current block. To reduce complexity and considering the trade-off between compression performance and complexity, an embodiment of the present application may impose a restriction that the GPM mode or the AWP mode is not used for blocks of a certain size. Therefore, in the present application, the decoder can first determine a size parameter of the current block, and then determine whether the GPM mode or the AWP mode is used for the current block based on the size parameter.
[0200] In an embodiment of the present application, the size parameters of the current block may include the height and width of the current block, so that the decoder can determine whether the GPM mode or the AWP mode is used for the current block based on the height and width of the current block.
[0201] Illustratively, the present application specifies that GPM mode or AWP mode can be used for the current block if the width is greater than threshold 1 and the height is greater than threshold 2. Thus, as can be seen from the above, one possible restriction is to use GPM mode or AWP mode only if the width of the block is greater than (or equal to or greater than) threshold 1 and the height of the block is greater than (or equal to or greater than) threshold 2. The values of threshold 1 and threshold 2 may be 4, 8, 16, 32, 128, 256, etc., and threshold 1 may be equal to threshold 2.
[0202] Illustratively, the present application specifies that GPM mode or AWP mode can be used for the current block if the width is less than threshold 3 and the height is greater than threshold 4. As can be seen from the above, one possible restriction is to use GPM mode or AWP mode only if the width of the block is less than (or equal to or less than) threshold 3 and the height of the block is greater than (or equal to or greater than) threshold 4. The values of threshold 3 and threshold 4 may be 4, 8, 16, 32, 128, 256, etc., and threshold 3 may be equal to threshold 4.
[0203] Furthermore, in embodiments of the present application, restrictions on sample parameters can limit the size of blocks in which GPM mode or AWP mode can be used.
[0204] Illustratively, in this application, the decoder can first determine the sample parameters of the current block, and then determine whether the GPM mode or the AWP mode can be used for the current block based on the sample parameters and threshold 5. As can be seen from the above, one possible restriction is to use the GPM mode or the AWP mode only if the number of samples in the block is greater than (or equal to or greater than) threshold 5. The value of threshold 5 may be 4, 8, 16, 32, 128, 256, 1024, etc.
[0205] That is, in this application, the GPM mode or the AWP mode can be used for the current block only under the condition that the size parameter of the current block meets the size requirement.
[0206] For example, in this application, there may be an image-level flag for specifying whether this application is used for the current image to be decoded. For example, this application may be configured to be used for intraframes (e.g., I frames) but not for interframes (e.g., B frames, P frames). Alternatively, this application may be configured to be not used for intraframes but to be used for interframes. Alternatively, this application may be configured to be used for some interframes but not for other parts of the interframes. Since intraframe prediction can also be used for interframes, this application may also be used for interframes.
[0207] In some embodiments, a flag below the image level can be used to identify whether the present application is used for the current block.
[0208] S102: A candidate prediction mode list is identified.
[0209] The candidate prediction mode list includes at least one candidate prediction mode, the at least one candidate prediction mode including a prediction mode identified based on dividing a template of the current block.
[0210] In some embodiments, when templates are used in TIMD, the entire set of templates, including the left template and the top template, are used together to derive the intra prediction mode for the TIMD. If a template for one side does not exist, for example, when the current block is located at the left or top boundary of the image, only the existing template can be used for the TIMD. However, if templates for both sides exist, these templates are used together. When neighboring reconstructed samples are used for DIMD, the left reconstructed sample and the top reconstructed sample are used together to derive the intra prediction mode for the DIMD. For example, when a reconstructed sample for one side does not exist, such as when the current block is located at the left or top boundary of the image, only the existing reconstructed sample can be used for the DIMD. However, if both the left reconstructed sample and the top reconstructed sample exist, these reconstructed samples are used together. This is not a problem for non-GPM blocks. However, for GPM blocks, the correlation between the first prediction mode and the left and top templates or the left and top reconstructed samples is different from the correlation between the second prediction mode and the left and top templates or the left and top reconstructed samples. If a prediction mode is derived using the entire template, the derived prediction mode will have low accuracy, resulting in an inaccurate candidate prediction mode list being constructed, and if a current block is predicted based on the inaccurate candidate prediction mode list, accurate prediction of the current block cannot be achieved.
[0211] To solve this technical problem, in an embodiment of the present application, a template of a current block is divided. Template division can be understood as dividing the template into a plurality of sub-templates, or dividing a reconstructed sample area where the template is located into a plurality of reconstructed sample sub-areas. In this way, by deriving a prediction mode based on the template obtained by division or the reconstructed sample area obtained by division, the accuracy of deriving the prediction mode can be improved, and the accuracy of constructing a candidate prediction mode list can be improved. By performing prediction based on an accurately constructed candidate prediction mode list, the prediction accuracy can be improved, and decoding performance can be improved.
[0212] In the embodiment of the present application, a specific method for identifying the candidate prediction mode list is not limited.
[0213] In some embodiments, the process of identifying the candidate prediction mode list is independent of the N candidate weight derivation modes. That is, the N candidate weight derivation modes can be understood as corresponding to one candidate prediction mode list. This can reduce the complexity of identifying the candidate prediction mode list and improve decoding efficiency. Note that, in this embodiment, since the candidate prediction mode list is independent of the N candidate weight derivation modes, there is no restriction on the order between S102 and S101. That is, S102 can be performed after S101, before S101, or simultaneously with S101, and this is not limited in the embodiments of the present application.
[0214] In some embodiments, S102 includes the following step S102-A.
[0215] S102-A: For each first candidate weight derivation mode in the N candidate weight derivation modes, identify a candidate prediction mode list corresponding to the first candidate weight derivation mode.
[0216] In one example, the first candidate weight derivation mode is any one of N candidate weight derivation modes. That is, in this example, at least one candidate prediction mode list needs to be identified for each of the N candidate weight derivation modes. As can be seen from the above, one weight derivation mode corresponds to K prediction modes, and the candidate prediction mode list is used to identify the prediction mode. Therefore, in one possible implementation of this example, one candidate prediction mode list is identified for at least one prediction mode among the K prediction modes corresponding to each candidate weight derivation mode in the N candidate weight derivation modes.
[0217] In another example, if the first candidate weight derivation mode belongs to one type of candidate weight derivation mode among N candidate weight derivation modes, in the embodiment of the present application, the N candidate weight derivation modes need to be classified, and at least one candidate prediction mode list is constructed for each type of candidate weight derivation mode.
[0218] Specifically, the decoding side identifies angle indexes corresponding to N candidate weight derivation modes. For example, the decoding side identifies angle indexes corresponding to each of the N candidate weight derivation modes. The method for identifying the angle indexes can be found in the detailed description of the above embodiment, and will not be described again in this specification. Next, the decoding side classifies the N candidate weight derivation modes into M types of candidate weight derivation modes based on the angle index corresponding to each candidate weight derivation mode. Candidate weight derivation modes of the same type have the same angle index. That is, the decoding side classifies candidate weight derivation modes having the same angle index into one type based on the angle index corresponding to each candidate weight derivation mode, thereby obtaining M types of candidate weight derivation modes. Each type of candidate weight derivation mode includes at least one candidate weight derivation mode. Furthermore, the jth type of candidate weight derivation mode among the M types of candidate weight derivation modes is identified as the first weight derivation mode, where j is a positive integer equal to or less than M. In this example, at least one candidate prediction mode list is identified for each type of candidate weight derivation mode among the N candidate weight derivation modes.
[0219] In the embodiment of the present application, the method for identifying the candidate prediction mode list corresponding to each first candidate weight derivation mode among the N candidate weight derivation modes is the same. For ease of explanation, the embodiment of the present application takes the identification of the candidate prediction mode list corresponding to one first candidate weight derivation mode as an example.
[0220] A specific method for identifying the candidate prediction mode list corresponding to the first candidate weight derivation mode in S102-A will be introduced below.
[0221] In some embodiments, the first candidate weight derivation mode corresponds to one candidate prediction mode list.
[0222] In some embodiments, S102-A includes the following step S102-A1:
[0223] S102-A1: Identify a candidate prediction mode list for at least one prediction mode among the K prediction modes corresponding to a first candidate weight derivation mode.
[0224] In this embodiment, a candidate prediction mode list is identified for at least one prediction mode among the K prediction modes corresponding to the first candidate weight derivation mode. Assuming that K=2, selectively, one candidate prediction mode list may be identified for the first prediction mode, but no candidate prediction mode list may be identified for the second candidate prediction mode. Selectively, one candidate prediction mode list may be identified for the second prediction mode, but no candidate prediction mode list may be identified for the first candidate prediction mode. Selectively, one candidate prediction mode list may be identified for the first prediction mode, and one candidate prediction mode list may be identified for the second candidate prediction mode. Selectively, one common candidate prediction mode list may be identified for the first prediction mode and the second prediction mode.
[0225] In an embodiment of the present application, a candidate prediction mode list is identified for at least one prediction mode corresponding to a first candidate weight derivation mode, and at least one prediction mode corresponding to the first candidate weight derivation mode is accurately identified from the constructed candidate prediction mode list.
[0226] In some embodiments, when the at least one prediction mode corresponds to one candidate prediction mode list, the aforementioned S102-A1 includes the following steps S102-A1-11 and S102-A1-12.
[0227] S102-A1-11: Identify a candidate prediction mode list for the i-th prediction mode in the at least one prediction mode, where i is a positive integer.
[0228] S102-A1-12: Identify a candidate prediction mode list for at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode.
[0229] In this embodiment, at least one prediction mode corresponding to the first candidate weight derivation mode corresponds to one candidate prediction mode list. That is, the at least one prediction mode corresponds to the same candidate prediction mode list, that is, corresponds to one candidate prediction mode list. In this way, the complexity of identifying the candidate prediction mode list can be reduced, and decoding efficiency can be improved. In this case, the decoding side identifies one candidate prediction mode list for the at least one prediction mode.
[0230] Specifically, a candidate prediction mode list for an i-th prediction mode among the at least one prediction mode is identified, where the i-th prediction mode is any prediction mode among the at least one prediction mode, and a candidate prediction mode list for the at least one prediction mode is identified based on the candidate prediction mode list for the i-th prediction mode.
[0231] Note that specific methods for identifying a candidate prediction mode list for at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode in S102-A1-12 above include, but are not limited to, the following:
[0232] Method 1: The candidate prediction mode list for the i-th prediction mode is directly identified as the candidate prediction mode list for at least one prediction mode.
[0233] Method 2: Determine whether the candidate prediction mode list for the i-th prediction mode includes a preset prediction mode. If the candidate prediction mode list for the i-th prediction mode includes a preset prediction mode, identify the candidate prediction mode list for the i-th prediction mode as a candidate prediction mode list for at least one prediction mode. If the candidate prediction mode list for the i-th prediction mode does not include a preset prediction mode, add the preset prediction mode to the candidate prediction mode list for the i-th prediction mode to obtain a candidate prediction mode list for at least one prediction mode.
[0234] In the embodiment of the present application, the preset prediction mode in the above Scheme 2 is not limited, and is specifically specified according to actual needs.
[0235] In this embodiment, when the at least one prediction mode corresponds to one candidate prediction mode list, a specific process of identifying a candidate prediction mode list for the at least one prediction mode is introduced.
[0236] In some embodiments, when each prediction mode in the at least one prediction mode corresponds to one candidate prediction mode list, the above S102-A1 includes the following step S102-A1-21.
[0237] S102-A1-21: For the ith prediction mode in the at least one prediction mode, identify a candidate prediction mode list for the ith prediction mode, where i is a positive integer.
[0238] In this embodiment, each prediction mode in the at least one prediction mode corresponds to one candidate prediction mode list. Thus, for a first candidate weight derivation mode, the decoding side identifies one candidate prediction mode list for each prediction mode in the at least one prediction mode corresponding to the first candidate weight derivation mode. For example, the at least one prediction mode includes a first prediction mode and a second prediction mode corresponding to the first candidate weight derivation mode. Accordingly, the decoding side identifies one candidate prediction mode list for the first prediction mode and one candidate prediction mode list for the second prediction mode.
[0239] In this embodiment, the process of identifying one candidate prediction mode list corresponding to each prediction mode in the at least one prediction mode is the same. For ease of explanation, the embodiment of the present application will take the identification of the candidate prediction mode list for the ith prediction mode in the at least one prediction mode as an example.
[0240] The process of identifying the candidate prediction mode list for the i-th prediction mode in S102-A1-11 and S102-A1-21 above will be described below.
[0241] In the embodiment of the present application, the specific types of candidate prediction modes included in the candidate prediction mode list for the i-th prediction mode are not limited.
[0242] In some embodiments, the candidate prediction mode list for the i-th prediction mode includes at least one of a first candidate prediction mode identified based on a template of the current block and a second candidate prediction mode identified based on gradients of reconstructed samples in the template.
[0243] Situation 1: When the candidate prediction mode list for the i-th prediction mode includes the first candidate prediction mode, the embodiment of the present application includes the following steps 11 to 14.
[0244] Step 11: Divide the template of the current block into P sub-templates, where P is a positive integer greater than 1.
[0245] As can be seen from the above, the template of the current block includes the left template of the current block and the upper template of the current block.Currently, the entire template of the current block is used to derive the first candidate prediction mode.For example, in TIMD, the entire template of the current block is used to derive the prediction mode, so that the derived first candidate prediction mode is not accurate enough.
[0246] According to an embodiment of the present application, a template of a current block is divided into P sub-templates, and then a first candidate prediction mode is derived based on the P sub-templates and / or the template of the current block, and the first candidate prediction mode is added to a candidate prediction mode list for the i-th prediction mode, thereby improving the accuracy of the candidate prediction mode list for the i-th prediction mode.
[0247] In the embodiment of the present application, the manner of dividing the template of the current block includes, but is not limited to, some of the following:
[0248] Method 1: At the decoding side, the template of the current block is divided according to a first candidate weight derivation mode, specifically, an angle index corresponding to the first candidate weight derivation mode is determined, and the template of the current block is divided into P sub-templates according to the angle index.
[0249] For example, as shown in Figure 18, taking the weight derivation mode in GPM with index 2 as an example, the white area in the weight matrix for the current block means that the weight corresponding to the predicted value in the first prediction mode is 100%, and the black area means that the weight corresponding to the predicted value in the second prediction mode is 100%. As shown in Figure 18, the first prediction mode is associated with the upper template of the current block, and the second prediction mode is associated with the left template of the current block and a portion of the upper template of the current block. However, currently, the entire template is used to derive the prediction mode, which makes the derivation of the prediction mode inaccurate and results in a large prediction error.
[0250] To solve this technical problem, the present application can realize finer division of the template based on the weight derivation mode. For example, as shown in FIG. 18, the present application can identify an angle index corresponding to a first candidate weight derivation mode and identify a division line of the weight matrix corresponding to the first candidate weight derivation mode based on the angle index. Then, the division line is extended to the template region of the current block to divide the template into two sub-templates. For example, the two sub-templates are referred to as a first sub-template and a second sub-template, where the first sub-template corresponds to a first prediction mode and the second sub-template corresponds to a second prediction mode. That is, the first sub-template is used to derive a first candidate prediction mode corresponding to the first prediction mode, and the second sub-template is used to derive a first candidate prediction mode corresponding to the second prediction mode.
[0251] Method 2: At the decoding side, the template of the current block is divided into P sub-templates based on the size of the current block. For example, if the size of the current block is smaller than a threshold, the template of the current block is divided into fewer sub-templates. If the size of the current block is equal to or larger than the threshold, the template of the current block is divided into more sub-templates.
[0252] The size of the current block includes the width or height of the current block, or the number of samples in the current block.
[0253] Alternatively, if the size of the current block includes the width or height of the current block, the threshold value may be 8, 16, 32, etc.
[0254] Alternatively, if the size of the current block includes the number of samples in the current block, the threshold may be 64, 128, 256, 512, etc.
[0255] In one example, the threshold value is a default value.
[0256] In one example, the threshold value may be derived based on a flag in a higher layer, for example, a flag in a sequence parameter set (SPS) that indicates the value of the threshold value.
[0257] Method 3: Divide the left template and / or the top template of the current block to obtain P sub-templates. For example, divide the left template of the current block equally into halves, quarters, etc., and / or divide the top template of the current block equally into halves, quarters, etc.
[0258] In addition, on the decoding side, in addition to dividing the template of the current block into P sub-templates using the above methods 1 to 3, other methods of division are also possible, and the embodiments of the present application are not limited thereto.
[0259] On the decoding side, after the template of the current block is divided into P sub-templates, the following step 12 is performed.
[0260] Step 12: Select Q prediction templates from the P sub-templates and / or templates of the current block, where Q is a positive integer less than or equal to P+1.
[0261] In an embodiment of the present application, to improve the accuracy of the first candidate prediction mode, the template of the current block is divided into P sub-templates in step 11, and then Q prediction templates are selected from the P sub-templates and / or the template of the current block. The Q prediction templates are then used to derive the first candidate prediction mode, thereby achieving accurate derivation of the first candidate prediction mode. Finally, the derived first candidate prediction mode is added to the candidate prediction mode list corresponding to the i-th prediction mode.
[0262] In the embodiments of the present application, a prediction template may be understood as a template used to derive a prediction mode, and the prediction template may be the above-mentioned sub-template or a template of the current block.
[0263] In the embodiment of the present application, a specific manner for selecting Q prediction templates from P sub-templates and / or templates of the current block is not limited.
[0264] In some embodiments, Q prediction templates are selected from the P sub-templates and / or the template of the current block based on default conditions, e.g., Q prediction templates are selected from the P sub-templates.
[0265] In some embodiments, Q prediction templates are selected from the P sub-templates and / or templates of the current block in steps 12-1 to 12-3 below.
[0266] Step 12-1: Identify the angle index corresponding to the first candidate weight derivation mode.
[0267] Step 12-2: Identify available neighboring blocks corresponding to the i-th prediction mode based on the angle index.
[0268] Step 12-3: Select Q prediction templates from the P sub-templates and / or template of the current block based on the available neighboring blocks corresponding to the i-th prediction mode.
[0269] In this embodiment, it is assumed that five neighboring blocks are available for the current block, and the positions of the five neighboring blocks are illustrated in Figure 19. The coordinates of the upper left corner of the current block are denoted as (x0, y0), the width of the current block is denoted as width, and the height of the current block is denoted as height. The five neighboring blocks are neighboring block AL identified based on coordinates (x0-1, y0-1), neighboring block A identified based on coordinates (x0+width-1, y0-1), neighboring block AR identified based on coordinates (x0+width, y0-1), neighboring block L identified based on coordinates (x0-1, y0+height-1), and neighboring block BL identified based on coordinates (x0-1, y0+height). It can be understood that A is the upper neighboring block of the current block, L is the left neighboring block of the current block, AR is the neighboring block at the upper right corner of the current block, AL is the neighboring block at the upper left corner of the current block, and BL is the neighboring block at the lower left corner of the current block.
[0270] In one example, Table 5 shows the correspondence between the angle index, the first portion (first prediction mode), the second portion (second prediction mode), and the neighboring blocks.
[0271] [Table 5]
[0272] In Table 5 above, it can be understood that A is the upper neighboring block of the current block, L is the left neighboring block of the current block, and L+A are the left neighboring block and the upper neighboring block of the current block.
[0273] In this embodiment, the decoding side determines an angle index corresponding to a first candidate weight derivation mode, determines a range of available neighboring blocks corresponding to the i-th prediction mode based on the angle index, and selects Q prediction templates from the P sub-templates and / or the template of the current block based on the available neighboring blocks corresponding to the i-th prediction mode. In this way, the selected Q prediction templates become associated with the i-th prediction mode. In this way, the first candidate prediction mode corresponding to the i-th prediction mode can be accurately determined based on the Q prediction templates associated with the i-th prediction mode.
[0274] For example, when the i-th prediction mode is the first prediction mode corresponding to the first candidate weight derivation mode, an available neighboring block corresponding to the i-th prediction mode is identified from the neighboring blocks corresponding to the first portion. For example, when the angle index corresponding to the first candidate weight derivation mode is 2, the available neighboring block corresponding to the i-th prediction mode can be identified as A based on Table 5 above.
[0275] For example, when the i-th prediction mode is the second prediction mode corresponding to the first candidate weight derivation mode, the available neighboring blocks corresponding to the i-th prediction mode are identified from the neighboring blocks corresponding to the second portion. For example, when the angle index corresponding to the first candidate weight derivation mode is 2, the available neighboring blocks corresponding to the i-th prediction mode can be identified as L+A based on Table 5 above.
[0276] Then, select Q prediction templates from the P sub-templates and / or templates of the current block based on the available neighboring blocks corresponding to the i-th prediction mode.
[0277] Example 1: If the available neighboring blocks corresponding to the i-th prediction mode include the upper neighboring block of the current block, identify the sub-template located above the current block among the P sub-templates as the prediction template among the Q prediction templates.
[0278] In a possible embodiment, in this example, each of the sub-templates located above the current block in the P sub-templates can be identified as one prediction template in the Q prediction templates.
[0279] In a possible embodiment, in this example, a subtemplate located above the current block among the P subtemplates can be merged with one or more prediction templates among the Q prediction templates. For example, the subtemplates located above the current block among the P subtemplates are subtemplate a, subtemplate b, and subtemplate c, respectively. Subtemplate a, subtemplate b, and subtemplate c are merged into one prediction template. Alternatively, any two of subtemplate a, subtemplate b, and subtemplate c are merged into one prediction template, and the remaining subtemplate is used as another prediction template.
[0280] Example 2: If the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block, the sub-template located to the left of the current block among the P sub-templates is identified as the prediction mode in the Q prediction templates.
[0281] In a possible embodiment, in this example, each of the sub-templates located to the left of the current block in the P sub-templates can be identified as one prediction template in the Q prediction templates.
[0282] In a possible embodiment, in this example, a subtemplate located to the left of the current block among the P subtemplates can be merged with one or more prediction templates among the Q prediction templates. For example, the subtemplates located to the left of the current block among the P subtemplates are subtemplate a, subtemplate b, and subtemplate c, respectively. Subtemplate a, subtemplate b, and subtemplate c are merged into one prediction template. Alternatively, any two of subtemplate a, subtemplate b, and subtemplate c are merged into one prediction template, and the remaining subtemplate is used as another prediction template.
[0283] Example 3: If the available neighboring blocks corresponding to the i-th prediction mode include a left-side neighboring block of the current block and an above-side neighboring block of the current block, at least one of the sub-template located to the left of the current block among the P sub-templates, the sub-template located above the current block among the P sub-templates, and the template of the current block is identified as a prediction template among the Q prediction templates.
[0284] For example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, the sub-template located to the left of the current block among the P sub-templates is identified as the prediction template among the Q prediction templates.
[0285] As another example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, the sub-template located above the current block among the P sub-templates is identified as the prediction template among the Q prediction templates.
[0286] As another example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, the template of the current block is identified as the prediction template among the Q prediction templates.
[0287] As another example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the above neighboring block of the current block, the subtemplate located to the left of the current block among the P subtemplates, the subtemplate located above the current block among the P subtemplates, and the template of the current block are identified as prediction templates among the Q prediction templates.
[0288] As another example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, at least one of the sub-templates located to the left of the current block in the P sub-templates and the template of the current block is identified as a prediction template in the Q prediction templates.
[0289] As another example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, at least one of the sub-templates located above the current block in the P sub-templates and the template of the current block is identified as a prediction template in the Q prediction templates.
[0290] Natural textures have large variations, so when reflected at the sample level, there are many gradual changes of varying degrees in many places. For example, some parts are not adjacent to the left or top, but the texture changes gradually. Therefore, it may be difficult to distinguish high and low correlations within a small range, especially for small blocks. For large blocks, it is easier to distinguish high and low correlations at larger distances.
[0291] Based on the above, the above step 12 includes selecting Q prediction templates from the P sub-templates and / or templates of the current block based on the size of the current block.
[0292] The size of a current block includes, but is not limited to, the width and height of the block, the number of samples in the block, etc.
[0293] In one example, when the size of the current block is equal to or smaller than the second threshold, the P sub-templates and the template of the current block are identified as prediction templates among the Q prediction templates. That is, when the size of the current block is equal to or smaller than the second threshold, the prediction modes derived from the sub-templates and the entire template are not distinguished, and these prediction modes are directly selected as first candidate prediction modes and added to the candidate prediction mode list corresponding to the i-th prediction mode.
[0294] In another example, if the size of the current block is equal to or smaller than a second threshold, the template of the current block is identified as a prediction template among the Q prediction templates. In this example, if the size of the current block is equal to or smaller than the second threshold, it is difficult to distinguish between high and low correlations. Therefore, the prediction mode derived from the template of the current block is directly identified as the first candidate prediction mode and added to the candidate prediction mode list corresponding to the i-th prediction mode.
[0295] In some embodiments, if the size of the current block is greater than the first threshold, steps 12-1 to 12-3 above are performed, i.e., if the size of the current block is greater than the first threshold, step 12-1 of identifying an angle index corresponding to a first candidate weight derivation mode is performed.
[0296] In the embodiment of the present application, the values of the first threshold and the second threshold are not limited.
[0297] For example, the first threshold is equal to the second threshold.
[0298] Alternatively, if the size of the current block includes the width or height of the current block, the first threshold and the second threshold may be 8, 16, 32, and so on.
[0299] Alternatively, if the size of the current block includes the number of samples in the current block, the first threshold and the second threshold may be 64, 128, 256, 512, and so on.
[0300] In one example, the first threshold and the second threshold are default values.
[0301] In one example, the first and second thresholds may be derived based on a flag in a higher layer, for example, by setting a flag in the SPS to indicate the first and second thresholds.
[0302] On the decoding side, after Q prediction templates are selected from the P sub-templates and / or templates of the current block in the above step, the following step 13 is performed.
[0303] Step 13: Identify prediction modes derived from the Q prediction templates.
[0304] On the decoding side, after selecting Q prediction templates from the P sub-templates and / or the template of the current block in the above steps, the Q prediction templates are used to derive a prediction mode.
[0305] In one example, for each of these Q prediction templates, one prediction mode is derived using the prediction template, thereby deriving Q prediction modes.
[0306] In some embodiments, a specific manner for identifying a prediction mode derived from the Q prediction templates may be as follows: R alternative prediction modes are identified for any of the Q prediction templates. The R alternative prediction modes may be all available prediction modes, several preset prediction modes, or several prediction modes corresponding to the prediction template; the embodiments of the present application are not limited thereto. Next, a first cost for predicting the prediction template using the R alternative prediction modes is identified. Since all prediction templates have already been reconstructed, the prediction template is predicted using each of the R alternative prediction modes to obtain a predicted value of the prediction template for each alternative prediction mode. For each alternative prediction mode, a first cost corresponding to the alternative prediction mode is obtained based on the reconstructed value and the predicted value for the alternative prediction mode. Optionally, the first cost may be an approximate cost such as SAD or SATD. Finally, a prediction mode derived from the prediction template is obtained based on the first cost corresponding to each of the R alternative prediction modes. For example, among the R preliminary prediction modes, the preliminary prediction mode corresponding to the smallest first cost is identified as the prediction mode derived from the prediction template.
[0307] On the decoding side, after the prediction modes derived from the Q prediction templates are identified in the above step, the following step 14 is performed.
[0308] Step 14: Identify at least one prediction mode among the prediction modes derived from the Q prediction templates as a first candidate prediction mode.
[0309] As can be seen from the above, the prediction modes derived from the Q prediction templates correspond to a first cost. In some embodiments, at least one prediction mode from the prediction modes derived from the Q prediction templates can be selected and identified as a first candidate prediction mode based on the first cost.
[0310] For example, one or more prediction modes having the smallest first cost among the prediction modes derived from the Q prediction templates are identified as first candidate prediction modes.
[0311] In some embodiments, prediction modes derived from the Q prediction templates are not screened, but rather a prediction mode derived from the Q prediction templates is directly identified as the first candidate prediction mode.
[0312] After the first candidate prediction mode is identified in the above step, the identified first candidate prediction mode is added to the candidate prediction mode list corresponding to the i-th prediction mode.
[0313] As can be seen from the above, in the embodiment of the present application, when identifying the first candidate prediction mode, the template of the current block is divided into P sub-templates, and the first candidate prediction mode is derived based on the P sub-templates and / or the template of the current block, thereby improving the accuracy of deriving the first candidate prediction mode and the quality of the candidate prediction mode list corresponding to the i-th prediction mode.
[0314] The process of identifying the first candidate prediction mode when the candidate prediction mode list for the i-th prediction mode in situation 1 above includes the first candidate prediction mode has been introduced above.
[0315] Below, when the candidate prediction mode list for the i-th prediction mode in situation 2 includes a second candidate prediction mode, a process of identifying the second candidate prediction mode will be introduced.
[0316] Situation 2: If the candidate prediction mode list for the i-th prediction mode includes the second candidate prediction mode, the embodiment of the present application includes the following steps 21 to 24.
[0317] Step 21: Divide the reconstructed sample region where the template of the current block is located into S reconstructed sample sub-regions, where S is a positive integer.
[0318] In an embodiment of the present application, the reconstructed sample area where the template of the current block is located includes the left adjacent reconstructed sample area of the current block and the upper adjacent reconstructed sample area of the current block.Currently, the entire reconstructed sample area where the template of the current block is located is used to derive the second candidate prediction mode.For example, in DIMD, the second candidate prediction mode derived from the entire reconstructed sample area where the template of the current block is located is not accurate enough.
[0319] For ease of explanation, the reconstruction sample area in which the template of the current block is located is called the reconstruction sample area.
[0320] In an embodiment of the present application, a reconstructed sample area where a template of a current block is located is divided into S reconstructed sample sub-areas, a second candidate prediction mode is derived based on the S reconstructed sample sub-areas and / or the reconstructed sample area, and the second candidate prediction mode is added to a candidate prediction mode list for the i-th prediction mode, thereby improving the accuracy of the candidate prediction mode list for the i-th prediction mode.
[0321] In the embodiment of the present application, the manner of dividing the reconstruction sample area where the template of the current block is located includes, but is not limited to, some of the following:
[0322] Method 1: At the decoding side, the template of the current block is divided according to a first candidate weight derivation mode. Specifically, an angle index corresponding to the first candidate weight derivation mode is determined. According to the angle index, the reconstructed sample region where the template of the current block is located is divided into S reconstructed sample sub-regions.
[0323] For example, as shown in Figure 20, taking the weight derivation mode in GPM with index 2 as an example, the white area in the weight matrix for the current block means that the weight corresponding to the predicted value in the first prediction mode is 100%, and the black area means that the weight corresponding to the predicted value in the second prediction mode is 100%. As shown in Figure 20, the first prediction mode is associated with the upper reconstructed sample area of the current block, and the second prediction mode is associated with the left reconstructed sample area of the current block and a portion of the upper reconstructed sample area of the current block. However, currently, the entire reconstructed sample area is used to derive the prediction mode, which makes the derivation of the prediction mode inaccurate and results in a large prediction error.
[0324] To solve this technical problem, the present application may realize a finer division of the reconstructed sample region in which the template is located based on the weight derivation mode. For example, as shown in FIG. 19, the present application may identify an angle index corresponding to a first candidate weight derivation mode and identify a division line of the weight matrix corresponding to the first candidate weight derivation mode based on the angle index. The division line may then be extended to the reconstructed sample region in which the template of the current block is located to divide the reconstructed sample region into two reconstructed sample subregions. For example, the two reconstructed sample subregions may be referred to as a first reconstructed sample subregion and a second reconstructed sample subregion, where the first reconstructed sample subregion corresponds to a first prediction mode and the second reconstructed sample subregion corresponds to a second prediction mode. That is, the first reconstructed sample subregion is used to derive a second candidate prediction mode corresponding to the first prediction mode, and the second reconstructed sample subregion is used to derive a second candidate prediction mode corresponding to the second prediction mode.
[0325] Method 2: At the decoding side, the reconstructed sample area where the template of the current block is located is divided into S reconstructed sample sub-areas based on the size of the current block. For example, if the size of the current block is smaller than a certain threshold, the reconstructed sample area where the template of the current block is located is divided into a relatively small number of reconstructed sample sub-areas. If the size of the current block is equal to or larger than the threshold, the reconstructed sample area where the template of the current block is located is divided into a relatively large number of reconstructed sample sub-areas.
[0326] The size of the current block includes the width or height of the current block, or the number of samples in the current block.
[0327] Alternatively, if the size of the current block includes the width or height of the current block, the threshold value may be 8, 16, 32, etc.
[0328] Alternatively, if the size of the current block includes the number of samples in the current block, the threshold may be 64, 128, 256, 512, etc.
[0329] In one example, the threshold value is a default value.
[0330] In one example, the threshold value may be derived based on a flag in a higher layer, for example, by setting a flag in the SPS to indicate the value of the threshold value.
[0331] Method 3: Divide the left-side adjacent reconstructed sample area and / or the upper-side adjacent reconstructed sample area of the current block to obtain S reconstructed sample sub-areas. For example, the left-side adjacent reconstructed sample area of the current block is equally divided into halves, quarters, etc., and / or the upper-side adjacent reconstructed sample area of the current block is equally divided into halves, quarters, etc.
[0332] In addition, on the decoding side, in addition to dividing the reconstructed sample area where the template of the current block is located into S reconstructed sample sub-areas using the above methods 1 to 3, it is also possible to divide it using other methods, and the embodiments of the present application are not limited to these.
[0333] On the decoding side, the reconstructed sample region where the template of the current block is located is divided into S reconstructed sample sub-regions, and then the following step 12 is performed.
[0334] Step 22: Select G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region in which the template of the current block is located, where G is a positive integer less than or equal to S+1.
[0335] In an embodiment of the present application, to improve the accuracy of the second candidate prediction mode, the reconstructed sample region where the template of the current block is located is divided into S reconstructed sample sub-regions in step 21, and then G reconstructed sample prediction regions are selected from the S reconstructed sample sub-regions and / or the reconstructed sample region where the template of the current block is located. The G reconstructed sample prediction regions are then used to derive the second candidate prediction mode, thereby achieving accurate derivation of the second candidate prediction mode. Finally, the derived second candidate prediction mode is added to the candidate prediction mode list corresponding to the i-th prediction mode.
[0336] In an embodiment of the present application, the reconstructed sample prediction area may be understood as the reconstructed sample area used to derive the prediction mode, and the reconstructed sample prediction area may be the above-mentioned reconstructed sample sub-area or the reconstructed sample area in which the template of the current block is located.
[0337] In the embodiment of the present application, a specific manner for selecting G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region in which the template of the current block is located is not limited.
[0338] In some embodiments, G reconstructed sample prediction regions are selected from the S reconstructed sample sub-regions and / or the reconstructed sample region in which the template of the current block is located based on default conditions, e.g., G reconstructed sample prediction regions are selected from the S reconstructed sample sub-regions.
[0339] In some embodiments, steps 22-1 to 22-3 below select G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or reconstructed sample region in which the template of the current block is located.
[0340] Step 22-1: Identify the angle index corresponding to the first candidate weight derivation mode.
[0341] Step 22-2: Identify available neighboring blocks corresponding to the i-th prediction mode based on the angle index.
[0342] Step 22-3: Based on the available neighboring blocks corresponding to the i-th prediction mode, select G reconstructed sample prediction areas from the S reconstructed sample sub-areas and / or the reconstructed sample area in which the template of the current block is located.
[0343] In this embodiment, it is assumed that five neighboring blocks are available for the current block, and the locations of the five neighboring blocks are illustrated in FIG.
[0344] In one example, Table 5 shows the correspondence between the angle index, the first portion (first prediction mode), the second portion (second prediction mode), and the neighboring blocks.
[0345] In this embodiment, the decoding side determines an angle index corresponding to a first candidate weight derivation mode, determines available neighboring blocks corresponding to the i-th prediction mode based on the angle index, and further selects G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region in which the template of the current block is located based on the available neighboring blocks corresponding to the i-th prediction mode. In this way, the selected G reconstructed sample prediction regions are associated with the i-th prediction mode. In this way, a second candidate prediction mode corresponding to the i-th prediction mode can be accurately identified based on the G reconstructed sample prediction regions associated with the i-th prediction mode.
[0346] For example, when the i-th prediction mode is the first prediction mode corresponding to the first candidate weight derivation mode, an available neighboring block corresponding to the i-th prediction mode is identified from the neighboring blocks corresponding to the first portion. For example, when the angle index corresponding to the first candidate weight derivation mode is 2, the available neighboring block corresponding to the i-th prediction mode can be identified as A based on Table 5 above.
[0347] For example, when the i-th prediction mode is the second prediction mode corresponding to the first candidate weight derivation mode, the available neighboring blocks corresponding to the i-th prediction mode are identified from the neighboring blocks corresponding to the second portion. For example, when the angle index corresponding to the first candidate weight derivation mode is 2, the available neighboring blocks corresponding to the i-th prediction mode can be identified as L+A based on Table 5 above.
[0348] Then, based on the available neighboring blocks corresponding to the i-th prediction mode, G reconstructed sample prediction regions are selected from the S reconstructed sample sub-regions and / or the reconstructed sample region in which the template of the current block is located.
[0349] Example 1: If the available neighboring blocks corresponding to the i-th prediction mode include the upper neighboring block of the current block, identify the reconstructed sample sub-area located above the current block among the S reconstructed sample sub-areas as the reconstructed sample prediction area among the G reconstructed sample prediction areas.
[0350] In a possible embodiment, in this example, each of the reconstructed sample sub-areas located above the current block in the S reconstructed sample sub-areas can be identified as one reconstructed sample prediction area in the G reconstructed sample prediction areas.
[0351] In a possible embodiment, in this example, reconstructed sample sub-areas located above the current block among the S reconstructed sample sub-areas can be merged into one or more reconstructed sample prediction areas among the G reconstructed sample prediction areas. For example, the reconstructed sample sub-areas located above the current block among the S reconstructed sample sub-areas are reconstructed sample sub-area a, reconstructed sample sub-area b, and reconstructed sample sub-area c, respectively. Reconstructed sample sub-area a, reconstructed sample sub-area b, and reconstructed sample sub-area c are merged into one reconstructed sample prediction area. Alternatively, any two of reconstructed sample sub-area a, reconstructed sample sub-area b, and reconstructed sample sub-area c are merged into one reconstructed sample prediction area, and the remaining reconstructed sample sub-area is used as another reconstructed sample prediction area.
[0352] Example 2: If the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block, identify the reconstructed sample sub-area located to the left of the current block among the S reconstructed sample sub-areas as the prediction mode in the G reconstructed sample prediction areas.
[0353] In a possible embodiment, in this example, each of the reconstructed sample sub-areas located to the left of the current block in the S reconstructed sample sub-areas can be identified as one reconstructed sample prediction area in the G reconstructed sample prediction areas.
[0354] In a possible embodiment, in this example, the reconstructed sample sub-areas located to the left of the current block among the S reconstructed sample sub-areas can be merged into one or more reconstructed sample prediction areas among the G reconstructed sample prediction areas. For example, the reconstructed sample sub-areas located to the left of the current block among the S reconstructed sample sub-areas are reconstructed sample sub-area a, reconstructed sample sub-area b, and reconstructed sample sub-area c, respectively. Reconstructed sample sub-area a, reconstructed sample sub-area b, and reconstructed sample sub-area c are merged into one reconstructed sample prediction area. Alternatively, any two of reconstructed sample sub-area a, reconstructed sample sub-area b, and reconstructed sample sub-area c are merged into one reconstructed sample prediction area, and the remaining reconstructed sample sub-area is used as another reconstructed sample prediction area.
[0355] Example 3: If the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, at least one of the reconstructed sample sub-area located to the left of the current block among the S reconstructed sample sub-areas, the reconstructed sample sub-area located above the current block among the S reconstructed sample sub-areas, and the reconstructed sample area in which the template of the current block is located is identified as a reconstructed sample prediction area among the G reconstructed sample prediction areas.
[0356] For example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, the reconstructed sample sub-area located to the left of the current block among the S reconstructed sample sub-areas is identified as the reconstructed sample prediction area among the G reconstructed sample prediction areas.
[0357] As another example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, the reconstructed sample sub-area located above the current block among the S reconstructed sample sub-areas is identified as the reconstructed sample prediction area among the G reconstructed sample prediction areas.
[0358] As another example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, the reconstructed sample area in which the template of the current block is located is identified as the reconstructed sample prediction area among the G reconstructed sample prediction areas.
[0359] As another example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the above neighboring block of the current block, the reconstructed sample sub-area located to the left of the current block among the S reconstructed sample sub-areas, the reconstructed sample sub-area located above the current block among the S reconstructed sample sub-areas, and the reconstructed sample area in which the template of the current block is located are identified as reconstructed sample prediction areas among the G reconstructed sample prediction areas.
[0360] As another example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, at least one of the reconstructed sample sub-areas located to the left of the current block in the S reconstructed sample sub-areas and the reconstructed sample area in which the template of the current block is located is identified as a reconstructed sample prediction area in the G reconstructed sample prediction areas.
[0361] As another example, if the available neighboring blocks corresponding to the i-th prediction mode include the left neighboring block of the current block and the upper neighboring block of the current block, at least one of the reconstructed sample sub-areas located above the current block in the S reconstructed sample sub-areas and the reconstructed sample area in which the template of the current block is located is identified as a reconstructed sample prediction area in the G reconstructed sample prediction areas.
[0362] Natural textures have large variations, so when reflected at the sample level, there are many gradual changes of varying degrees in many places. For example, some parts are not adjacent to the left or top, but the texture changes gradually. Therefore, it may be difficult to distinguish high and low correlations within a small range, especially for small blocks. For large blocks, it is easier to distinguish high and low correlations at larger distances.
[0363] Based on the above, the above step 22 includes selecting G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region in which the template of the current block is located based on the size of the current block.
[0364] The size of a current block includes, but is not limited to, the width and height of the block, the number of samples in the block, etc.
[0365] In one example, when the size of the current block is equal to or smaller than a fourth threshold, the reconstructed sample region in which the S reconstructed sample sub-regions and the template of the current block are located is identified as a reconstructed sample prediction region among the G reconstructed sample prediction regions. That is, when the size of the current block is equal to or smaller than the fourth threshold, some prediction modes derived from the reconstructed sample region in which the reconstructed sample sub-regions and the entire template are located are not distinguished, and these prediction modes are directly set as second candidate prediction modes and added to the candidate prediction mode list corresponding to the i-th prediction mode.
[0366] In another example, if the size of the current block is equal to or smaller than a fourth threshold, the template of the current block is identified as a reconstructed sample prediction region among the G reconstructed sample prediction regions. In this example, if the size of the current block is equal to or smaller than the fourth threshold, it is difficult to distinguish between high and low correlations. Therefore, the prediction mode derived from the reconstructed sample region in which the template of the current block is located is directly identified as the second candidate prediction mode and added to the candidate prediction mode list corresponding to the i-th prediction mode.
[0367] In some embodiments, if the size of the current block is greater than the third threshold, steps 22-1 to 22-3 above are performed, i.e., if the size of the current block is greater than the third threshold, step 22-1 of identifying an angle index corresponding to the first candidate weight derivation mode is performed.
[0368] In the embodiment of the present application, the values of the third threshold and the fourth threshold are not limited.
[0369] For example, the third threshold is equal to the fourth threshold.
[0370] Alternatively, if the size of the current block includes the width or height of the current block, the third threshold and the fourth threshold may be 8, 16, 32, and so on.
[0371] Alternatively, if the size of the current block includes the number of samples in the current block, the third threshold and the fourth threshold may be 64, 128, 256, 512, and so on.
[0372] In one example, the third and fourth thresholds are default values.
[0373] In one example, the third and fourth thresholds may be derived based on a flag in a higher layer, for example, by setting a flag in the SPS to indicate the third and fourth thresholds.
[0374] On the decoding side, after the above step selects G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region where the template of the current block is located, the following step 23 is performed.
[0375] Step 13: Identify a prediction mode derived from the G reconstructed sample prediction regions.
[0376] On the decoding side, in the above step, G reconstructed sample prediction areas are selected from the S reconstructed sample sub-areas and / or the reconstructed sample area where the template of the current block is located, and then these G reconstructed sample prediction areas are used to derive a prediction mode.
[0377] In one example, for each of these G reconstructed sample prediction regions, one prediction mode is derived using the reconstructed sample prediction region, thereby deriving G prediction modes.
[0378] In some embodiments, a specific manner of determining a prediction mode derived from the G reconstructed sample prediction regions may be as follows: For any of the reconstructed sample prediction regions in the G reconstructed sample prediction regions, determine a gradient of each central sample of the reconstructed sample prediction region; and determine a prediction mode derived from the reconstructed sample prediction region based on the gradient of each central sample of the reconstructed sample prediction region.
[0379] On the decoding side, after the prediction mode derived from the G reconstructed sample prediction regions is identified in the above step, the following step 24 is performed.
[0380] Step 24: Identify at least one of the prediction modes derived from the G reconstructed sample prediction regions as a second candidate prediction mode.
[0381] As can be seen from the above, the prediction modes derived from the G reconstructed sample prediction regions correspond to a first cost. In some embodiments, at least one prediction mode from the prediction modes derived from the G reconstructed sample prediction regions can be selected and identified as the second candidate prediction mode based on the first cost.
[0382] For example, one or more prediction modes having the smallest first cost among the prediction modes derived from the G reconstructed sample prediction regions are identified as second candidate prediction modes.
[0383] In some embodiments, prediction modes derived from the G reconstructed sample prediction regions are not screened, but a prediction mode derived from the G reconstructed sample prediction regions is directly identified as the second candidate prediction mode.
[0384] After the second candidate prediction mode is identified in the above step, the identified second candidate prediction mode is added to the candidate prediction mode list corresponding to the i-th prediction mode.
[0385] As can be seen from the above, in the embodiment of the present application, when identifying the second candidate prediction mode, the template of the current block is divided into S reconstructed sample sub-regions, and the second candidate prediction mode is derived based on the S reconstructed sample sub-regions and / or the template of the current block, thereby improving the accuracy of deriving the second candidate prediction mode and the quality of the candidate prediction mode list corresponding to the i-th prediction mode.
[0386] The process of identifying the second candidate prediction mode when the candidate prediction mode list for the i-th prediction mode in Situation 2 above includes the second candidate prediction mode has been introduced above.
[0387] In some embodiments, the candidate prediction mode list for the i-th prediction mode further includes at least one of a third candidate prediction mode corresponding to the first candidate weight derivation mode, a prediction mode for a neighboring block of the current block, and a pre-set prediction mode.
[0388] In the embodiment of the present application, the third candidate prediction mode corresponding to the first candidate weight derivation mode may be understood as a third candidate prediction mode identified based on the first candidate weight derivation mode. In the embodiment of the present application, the specific type of the third candidate prediction mode corresponding to the first candidate weight derivation mode is not limited.
[0389] In one example, the third candidate prediction mode corresponding to the first candidate weight derivation mode includes a prediction mode whose prediction angle is parallel to a dividing line (or weight boundary line) of the first candidate weight derivation mode. For example, assuming that the dividing line of the first candidate weight derivation mode is indicated by a diagonal line as shown in FIG. 21A, at least one prediction mode whose prediction angle is parallel to the dividing line is identified as the third candidate prediction mode corresponding to the first candidate weight derivation mode.
[0390] In another example, the third candidate prediction mode corresponding to the first candidate weight derivation mode includes a prediction mode whose prediction angle is perpendicular to a dividing line (or weight boundary line) of the first candidate weight derivation mode. For example, as shown in FIG. 21B, assuming that the dividing line of the first candidate weight derivation mode is indicated by a diagonal line in the figure, at least one prediction mode whose prediction angle is perpendicular to the dividing line is identified as the third candidate prediction mode corresponding to the first candidate weight derivation mode.
[0391] In some embodiments, to speed up the identification of the third candidate prediction mode, a lookup table is constructed for the angle index angleIdx and the corresponding intra-prediction mode. In this way, the decoding side can calculate the angle index of the first candidate weight derivation mode, and based on the angle index of the first candidate weight derivation mode, can obtain from the lookup table an intra-prediction mode whose prediction angle is parallel to the division line of the first candidate weight derivation mode. Optionally, based on the intra-prediction mode whose prediction angle is parallel to the division line of the first candidate weight derivation mode, can calculate an intra-prediction mode whose prediction angle is perpendicular to the division line of the first candidate weight derivation mode.
[0392] In some embodiments, when using intra prediction modes for neighboring blocks, intra prediction modes for up to five neighboring blocks are available, and the positions of the five neighboring blocks are illustrated in Figure 19. The coordinates of the upper left corner of the current block are denoted as (x0, y0), the width of the current block is denoted as width, and the height of the current block is denoted as height. The five neighboring blocks are neighboring block A L identified based on coordinates (x0-1, y0-1), neighboring block A identified based on coordinates (x0+width-1, y0-1), neighboring block A R identified based on coordinates (x0+width, y0-1), neighboring block L identified based on coordinates (x0-1, y0+height-1), and neighboring block B L identified based on coordinates (x0-1, y0+height). The range of available neighboring blocks is identified from Table 5 above based on whether the i-th prediction mode corresponds to the first part or the second part and the angle index angleIdx corresponding to the first candidate weight derivation mode. It can be understood that A is the upper neighboring block of the current block, and L is the left neighboring block of the current block. If it is determined from Table 5 that the available neighboring block corresponding to the i-th prediction mode is neighboring block A, the intra prediction mode for neighboring block A and the intra prediction mode for neighboring block AR are added to the candidate prediction mode list corresponding to the i-th prediction mode. If it is determined from Table 5 that the available neighboring block corresponding to the i-th prediction mode is neighboring block L, the intra prediction mode for neighboring block L and the intra prediction mode for neighboring block BL are added to the candidate prediction mode list corresponding to the i-th prediction mode. If it is determined from Table 5 that the available neighboring block corresponding to the i-th prediction mode is neighboring block L+A, the intra prediction mode for neighboring block A, the intra prediction mode for neighboring block AR, the intra prediction mode for neighboring block L, and the intra prediction mode for neighboring block BL are added to the candidate prediction mode list corresponding to the i-th prediction mode. As can be seen from the above, the prediction mode for neighboring block AL is always available.Optionally, the order in which the neighboring blocks are checked is L->A->BL->AR->AL.
[0393] In the embodiment of the present application, the specific type of the preset prediction mode is not limited. For example, the preset prediction mode may be at least one of a DC mode, a horizontal mode, a vertical mode, an angular mode, and a planar mode.
[0394] In some embodiments, in order to reduce the complexity of prediction, a limit is imposed on the length of the candidate prediction mode list. For example, the number of candidate prediction modes included in the candidate prediction mode list for the i-th prediction mode is a preset value. In the embodiments of the present application, the specific value of the preset value is not limited, for example, the preset value is 3.
[0395] In some embodiments, a candidate prediction mode list for the i-th prediction mode is constructed by selecting a predetermined number of prediction modes (e.g., three) from the following in a predetermined order: a first candidate prediction mode identified based on a template of the current block; a second candidate prediction mode identified based on gradients of reconstructed samples in the template of the current block; a prediction mode whose prediction angle is parallel to the division line of the first candidate weight derivation mode; a prediction mode whose prediction angle is perpendicular to the division line of the first candidate weight derivation mode; a prediction mode for a neighboring block of the current block; and a planar mode. In embodiments of the present application, the predetermined order is not limited.
[0396] In one example, when constructing a candidate prediction mode list for the i-th prediction mode, the following types of prediction modes are sequentially added to the candidate prediction mode list until the length of the list reaches a preset value (e.g., 3): 1. A prediction mode in which the prediction angle is parallel to the dividing line of the first candidate weight derivation mode. 2. A first candidate prediction mode identified based on the template of the current block. In some embodiments, the first candidate prediction mode is also referred to as a prediction mode derived based on TIMD. 3. A second candidate prediction mode identified based on gradients of reconstructed samples in the template of the current block. In some embodiments, the second candidate prediction mode is also referred to as a prediction mode derived based on DIMD. 4. Prediction mode for the neighboring blocks of the current block. 5. A prediction mode in which the prediction angle is perpendicular to the dividing line of the first candidate weight derivation mode. 6. PLANAR mode.
[0397] In the above embodiment, a specific process for identifying the candidate prediction mode list has been introduced.
[0398] On the decoding side, after the candidate prediction mode list is identified based on the above steps, the following step S103 is executed.
[0399] S103: Identify a first weight derivation mode and K first prediction modes based on the N candidate weight derivation modes and the candidate prediction mode list.
[0400] On the decoding side, N candidate weight derivation modes are identified in step S101, and a candidate prediction mode list is identified in step S102. Then, one candidate weight derivation mode is selected from the N candidate weight derivation modes as a first weight derivation mode, and at least one first prediction mode among K first prediction modes is identified from at least one candidate prediction mode included in the candidate prediction mode list. Finally, a current block is predicted using the identified first weight derivation mode and the K first prediction modes to obtain a predicted value of the current block.
[0401] Note that the first weight derivation mode and the K first prediction modes are used together to determine a predicted value for the current block. In some embodiments, the first weight derivation mode is also referred to as a weight derivation mode for the current block or a weight derivation mode corresponding to the current block. In some embodiments, the K first prediction modes are also referred to as K prediction modes for the current block or K prediction modes corresponding to the current block. In one example, when K=2, the K first prediction modes include a first prediction mode and a second prediction mode corresponding to the current block. In some embodiments, the first prediction mode is referred to as a first prediction mode, and the second prediction mode is referred to as a second prediction mode.
[0402] In the embodiment of the present application, the specific manner in which the decoding side identifies the first weight derivation mode and the K first prediction modes based on the N candidate weight derivation modes and the candidate prediction mode list is not limited.
[0403] In some embodiments, the candidate prediction mode list corresponds to K first prediction modes, i.e., all of the K first prediction modes are selected from the candidate prediction mode list. In this case, the decoding side combines N candidate weight derivation modes with the candidate prediction modes included in the candidate prediction mode list. For example, each of the N candidate weight derivation modes is combined with any K candidate prediction modes included in the candidate prediction mode list to obtain multiple combinations. Each combination includes one candidate weight derivation mode and K candidate prediction modes. Next, the template of the current block is predicted using the candidate weight derivation mode and the K candidate prediction modes included in each combination, and the cost of each combination is determined. Then, one combination is identified from the multiple combinations based on the cost. For example, the combination with the smallest cost is selected from the multiple combinations, and the candidate weight derivation mode included in the combination with the smallest cost is identified as the first weight derivation mode, and the K prediction modes included in the combination with the smallest cost are identified as the K first prediction modes.
[0404] In some embodiments, the candidate prediction mode list is a candidate prediction mode list for a first prediction mode among the K first prediction modes, e.g., K=2, and the candidate prediction mode is a candidate prediction mode list for the first prediction mode. In this case, the decoding side identifies a set of selectable prediction modes corresponding to the second prediction mode. Then, for each of the N candidate weight derivation modes, the decoding side selects one candidate prediction mode from the candidate prediction mode list for the first prediction mode as a possible option for the first prediction mode, and one prediction mode from the set of selectable prediction modes corresponding to the second prediction mode as a possible option for the second prediction mode, thereby obtaining one combination consisting of the candidate weight derivation mode, the possible options for the first prediction mode, and the possible options for the second prediction mode. In this manner, multiple combinations can be obtained. Each combination includes one candidate weight derivation mode and two candidate prediction modes. Next, the candidate weight derivation mode and two candidate prediction modes included in each combination are used to predict a template for the current block, and the cost of each combination is determined. Then, one combination is identified from the multiple combinations based on the cost. For example, the combination with the smallest cost is selected from multiple combinations, a candidate weight derivation mode included in the combination with the smallest cost is identified as a first weight derivation mode, and K prediction modes included in the combination with the smallest cost are identified as K first prediction modes.
[0405] In some embodiments, the candidate prediction mode list includes a candidate prediction mode list corresponding to each of the K first prediction modes, i.e., the decoding side identifies K candidate prediction modes in step S102. For example, assuming K=2, the decoding side identifies a candidate prediction mode list for the first prediction mode and a candidate prediction mode for the second prediction mode. In this case, the decoding side selects one candidate weight derivation mode from N candidate weight derivation modes, selects one candidate prediction mode from the candidate prediction mode list for the first prediction mode, and selects one candidate prediction mode from the candidate prediction mode list for the second prediction mode. In this way, the selected one candidate weight derivation mode and two candidate prediction modes form one combination. According to the above method, multiple combinations can be obtained. Each combination includes one candidate weight derivation mode and two candidate prediction modes. Next, the candidate weight derivation mode and two candidate prediction modes included in each combination are used to predict a template for the current block, and a cost for each combination is determined. Then, one combination is identified from the multiple combinations based on the cost. For example, the combination with the smallest cost is selected from multiple combinations, a candidate weight derivation mode included in the combination with the smallest cost is identified as a first weight derivation mode, and K prediction modes included in the combination with the smallest cost are identified as K first prediction modes.
[0406] Based on the above, one weight derivation mode and K prediction modes can be used together as one combination for the current block. In order to save codewords and reduce coding costs, in some embodiments, the weight derivation mode and K prediction modes corresponding to the current block are combined into one combination, i.e., a first combination, and a first index represents the first combination. Compared with the case where the weight derivation mode and the K prediction modes are represented separately, fewer codewords are used in the embodiments of the present application, thereby reducing coding costs.
[0407] Based on the above, the above S103 includes the following steps S103-A to S103-C.
[0408] S103-A: Decode the bitstream to obtain a first index, which is used to indicate a first combination, where the first combination includes a first weight derivation mode and K first prediction modes.
[0409] S103-B: Identify a candidate combination list based on the N candidate weight derivation modes and the candidate prediction mode list, where the candidate combination list includes at least one candidate combination, and the candidate combination includes one weight derivation mode and K prediction modes.
[0410] S103-C: Identify a first combination from the candidate combination list based on the first index.
[0411] The embodiments of the present application do not limit the specific format of the syntax element of the first index.
[0412] In one possible embodiment, if the GPM technique is used to predict the current block, gpm_cand_idx indicates the first index.
[0413] Since the first index is used to indicate the first combination, in some embodiments, the first index may also be referred to as a first combination index or an index of the first combination.
[0414] In one example, the syntax after adding the first index to the bitstream is shown in Table 6.
[0415] [Table 6]
[0416] gpm_cand_idx is the first index.
[0417] For example, a list of candidate combinations is shown in Table 7.
[0418] [Table 7]
[0419] As shown in Table 7, the candidate combination list includes multiple candidate combinations, and any two of the multiple candidate combinations are not exactly the same; that is, any two candidate combinations have different weight derivation modes and at least one of the K prediction modes. For example, the weight derivation mode in candidate combination 1 is different from the weight derivation mode in candidate combination 2. Alternatively, the weight derivation mode in candidate combination 1 is the same as the weight derivation mode in candidate combination 2, but the K prediction modes in candidate combination 1 are different from the K prediction modes in candidate combination 2 by at least one prediction mode. Alternatively, the weight derivation mode in candidate combination 1 is different from the weight derivation mode in candidate combination 2, and the K prediction modes in candidate combination 1 are different from the K prediction modes in candidate combination 2 by at least one prediction mode.
[0420] For example, in Table 7 above, the order of the candidate combinations in the candidate combination list is the index of the candidate combination. Optionally, the index of the candidate combination in the candidate combination list may be expressed in other ways. This is not a limitation in the embodiment of the present application.
[0421] In this embodiment, the decoding side decodes the bitstream to obtain a first index, identifies the candidate combination list shown in Table 7 above, and queries from the candidate combination list based on the first index to obtain the first weight derivation mode and K prediction modes included in the first combination indicated by the first index.
[0422] For example, the first index is index 1, and in the candidate combination list shown in Table 7, the candidate combination corresponding to index 1 is candidate combination 2, that is, the first combination indicated by the first index is candidate combination 2. In this way, the decoding side identifies the weight derivation mode and K prediction modes included in candidate combination 2 as the first weight derivation mode and K first prediction modes included in the first combination, and predicts the current block using the first weight derivation mode and K first prediction modes to obtain a predicted value of the current block.
[0423] In Scheme 2, the encoding side and the decoding side may each specify the same candidate combination list. For example, both the encoding side and the decoding side specify a list including X candidate combinations. Each candidate combination includes one weight derivation mode and K prediction modes. The encoding side only needs to signal one finally selected candidate combination (e.g., the first combination) to the bitstream. The decoding side analyzes the first combination finally selected by the encoding side. Specifically, the decoding side decodes the bitstream to obtain a first index, and identifies the first combination from the candidate combination list identified by the decoding side based on the first index.
[0424] Hereinafter, a specific process for identifying the candidate combination list based on the N candidate weight derivation modes and the candidate prediction mode list in S103-B above will be described.
[0425] The embodiment of the present application does not limit the specific manner of identifying the candidate combination list based on the N candidate weight derivation modes and the candidate prediction mode list in the above S103-B.
[0426] In some embodiments, the N candidate weight derivation modes are arbitrarily combined with the plurality of candidate prediction modes included in the candidate prediction mode list. Each combination includes one weight derivation mode and two prediction modes. In this way, a plurality of combinations can be obtained. Information related to the current block is used to analyze the occurrence probability of different combinations, and a candidate combination list is constructed according to the occurrence probability of each combination. Optionally, the information related to the current block includes mode information of the neighboring blocks of the current block, reconstructed samples of the current block, etc.
[0427] In some embodiments, the above S103-B includes the following steps S103-B1 and S103-B2.
[0428] S103-B1: Obtain T second combinations based on the N candidate weight derivation modes and the candidate prediction mode list.
[0429] S103-B2: Obtain a candidate combination list based on the T second combinations.
[0430] Any of the second combinations among the T second combinations includes one weight derivation mode and K prediction modes, and the weight derivation mode and K prediction modes included in any one of the second combinations among the T second combinations do not have to be exactly the same as the weight derivation mode and K prediction modes included in any other second combination among the T second combinations, where T is a positive integer greater than 1.
[0431] In this embodiment, the decoding side identifies T second combinations based on the N candidate weight derivation modes and the candidate prediction mode list. In this application, the specific number of the T second combinations is not limited, and may be, for example, 8, 16, 32, etc., and each of the T second combinations includes one weight derivation mode and K prediction modes. The weight derivation mode and the K prediction modes included in any one second combination among the T second combinations are not necessarily exactly the same as the weight derivation mode and the K prediction modes included in any other second combination among the T second combinations.
[0432] The embodiment of the present application does not limit the specific manner of obtaining the T second combinations based on the N candidate weight derivation modes and the candidate prediction mode list in the above S103-B1.
[0433] In some embodiments, the candidate prediction mode list corresponds to K first prediction modes, i.e., all of the K first prediction modes are selected from the candidate prediction mode list. In this case, the decoding side combines N candidate weight derivation modes with the candidate prediction modes included in the candidate prediction mode list. For example, each of the N candidate weight derivation modes is combined with any K candidate prediction modes included in the candidate prediction mode list to obtain T second combinations. Each second combination includes one candidate weight derivation mode and K candidate prediction modes.
[0434] In some embodiments, the candidate prediction mode list is a candidate prediction mode list for a first prediction mode among the K first prediction modes, e.g., K=2, and the candidate prediction mode is a candidate prediction mode list for the first prediction mode. In this case, the decoding side identifies a set of selectable prediction modes corresponding to the second prediction mode. Then, for each of the N candidate weight derivation modes, the decoding side selects one candidate prediction mode from the candidate prediction mode list for the first prediction mode as a possible option for the first prediction mode, and one prediction mode from the set of selectable prediction modes corresponding to the second prediction mode as a possible option for the second prediction mode, thereby obtaining one second combination consisting of the candidate weight derivation mode, the possible options for the first prediction mode, and the possible options for the second prediction mode. In this way, T second combinations can be obtained. Each second combination includes one candidate weight derivation mode and two candidate prediction modes.
[0435] In some embodiments, the candidate prediction mode list includes a candidate prediction mode list corresponding to each of the K first prediction modes, i.e., the decoding side identifies K candidate prediction modes in step S102. For example, assuming K=2, the decoding side identifies a candidate prediction mode list for the first prediction mode and a candidate prediction mode for the second prediction mode. In this case, the decoding side selects one candidate weight derivation mode from N candidate weight derivation modes, selects one candidate prediction mode from the candidate prediction mode list for the first prediction mode, and selects one candidate prediction mode from the candidate prediction mode list for the second prediction mode. In this way, the selected one candidate weight derivation mode and two candidate prediction modes constitute one second combination. According to the above method, T second combinations can be obtained. Each second combination includes one candidate weight derivation mode and two candidate prediction modes.
[0436] The implementation manner of obtaining the candidate combination list based on the T second combinations in the above S103-B2 includes, but is not limited to, the following several manners:
[0437] Method 1: Sort the T second combinations according to a preset rule to obtain a list of candidate combinations.
[0438] Method 2: The above S103-B2 includes the following steps:
[0439] S103-B21: For any of the T second combinations, determine the cost corresponding to the second combination when predicting the template of the current block using the weight derivation mode in the second combination and the K prediction modes.
[0440] S103-B22: Identify a candidate combination list based on the cost corresponding to each second combination in the T second combinations.
[0441] In Method 2, for each second combination among the T second combinations, the weight derivation mode included in the second combination and the K prediction modes are used to predict the template of the current block, thereby obtaining a predicted value of the template corresponding to the second combination.
[0442] Specifically, for each of the T second combinations, the template of the current block is predicted using the K prediction modes in the second combination to obtain K predicted values.
[0443] Next, the template weights corresponding to the second combination are identified based on the weight derivation mode for the second combination.
[0444] In some embodiments, determining the template weights based on the weight derivation mode includes: determining an angle index, a distance index, and a blending parameter based on the weight derivation mode, and determining the template weights based on the angle index, the distance index, the blending parameter, and the size of the template.
[0445] In this application, the template weights can be derived in the same manner as the predictor weights are derived, for example, first, the angle index and the distance index are identified based on the weight derivation mode.
[0446] The manners of identifying the template weights based on the angle index, distance index, and template size include, but are not limited to, the following several manners.
[0447] Method 1: Determine a first parameter of a sample in the template based on an angle index, a distance index, and a size of the template. In some embodiments, the first parameter is also referred to as a weight index (weightIdx). Determine weights of the samples in the template based on the first parameters of the samples in the template. Determine template weights based on the weights of the samples in the template.
[0448] In a possible embodiment, the template weights can be specified as follows:
[0449] The inputs for the template weight derivation process include the width nCbW of the current block, the height nCbH of the current block, the width nVmW of the left template, the height nVmH of the upper template, the "split" angle index variable angleId of the GFM, the distance index variable distanceIdx of the GFM, and the component index variable cIdx. For illustrative purposes, since the present application takes the luminance component as an example, cIdx being 0 indicates the luminance component.
[0450] The variables nW, nH, shift1, offset1, displacementX, displacementY, partFlip, and shiftHor are derived based on the following method. nW = ( cIdx = = 0 ) ? nCbW : nCbW * EubWidthC nH = ( cIdx = = 0 ) ? nCbH : nCbH * EubHeightC shift1 = Max( 5, 17 - BitDepth ), where BitDepth is the bit depth of the coding. offset1 = 1 << ( shift1 - 1 ) displacementX = angleIdx displacementY = (angleIdx + 8) % 32 partFlip = ( angleIdx >= 13 && angleIdx <= 27 ) ? 0 : 1 shiftHor = ( angleIdx % 16 = = 8 | | ( angleIdx % 16 != 0 && nH >= nW ) ) ? 0 : 1
[0451] The offsets offsetX and offsetY are derived based on the following method. - If the value of shiftHor is 0, offsetX = ( -nW ) >> 1 offsetY = ( ( -nH ) >> 1 ) + ( angleIdx < 16 ? ( distanceIdx * nH ) >> 3 : -( ( distanceIdx * nH ) >> 3 ) ) - Otherwise (i.e., shiftHor has the value 1), offsetX = ( ( -nW ) >> 1 ) + ( angleIdx < 16 ? ( distanceIdx * nW ) >> 3 : -( ( distanceIdx * nW ) >> 3 ) offsetY = ( - nH ) >> 1
[0452] The template weight matrix wVemplateValue[x][y] (x = -nVmW..nCbW - 1, y = -nVmH..nCbH - 1, except when both x and y are 0 or greater) is derived based on the following method (note that in this example, the coordinates of the upper left corner of the current block are set to (0,0)): - The variables xL and yL are derived based on the following method: xL = ( cIdx = = 0 ) ? x : x * EubWidthC yL = ( cIdx = = 0 ) ? y : y * EubHeightC
[0453] The disLut is specified according to Table 3 above.
[0454] The first parameter weightIdx is derived based on the following method. weightIdx = ( ( ( xL + offsetX ) << 1 ) + 1 ) * disLut[ displacementX ] + ( ( ( yL + offsetY ) << 1 ) + 1 ) * disLut[ displacementY ]
[0455] In some embodiments, after the first parameter weightIdx is determined based on the above method, the weights of the samples in the template are determined based on the following formula: weightIdxL = partFlip ? 32 + weightIdx : 32 - weightIdx wVemplateValue[x][y] = Clip3( 0, 8, ( weightIdxL + 4 ) >> 3 ) wVemplateValue[x][y] is the weight of sample (x, y) in the template. weightIdxL is the weight index for the first component (e.g., luminance component). wVemplateValue[x][y] is the weight of sample (x, y) in the template. partFlip is an intermediate variable that is determined based on the angle index angleIdx. For example, as described above, partFlip = ( angleIdx >= 13 && angleIdx <= 27 ) ? 0 : 1. That is, the value of partFlip is either 1 or 0. When the value of partFlip is 0, weightIdxL is 32 - weightIdx. When the value of partFlip is 1, weightIdxL is 32 + weightIdx. The value 32 here is merely an example, and the present application is not limited thereto.
[0456] Method 2: Determine the weight of the sample in the template based on the first parameter weightIdx, the first threshold, and the second threshold of the sample in the template.
[0457] In order to reduce the computational complexity of the template weights, Method 2 limits the weights of samples in the template to either the first or second threshold. That is, by limiting the weights of samples in the template to either the first or second threshold, the computational complexity of the template weights is reduced.
[0458] The present application does not limit the specific values of the first threshold and the second threshold.
[0459] Optionally, the first threshold is one.
[0460] Optionally, the second threshold is zero.
[0461] In one example, the weights of the samples in the template can be determined based on the following formula: wVemplateValue[x][y] = (partFlip ? weightIdx: - weightIdx) > 0 ? 1 : 0 wVemplateValue[x][y] is the weight of the sample (x, y) in the template, and 1 in the above "1 : 0" is the first threshold and 0 is the second threshold.
[0462] In the above method 1, the weight of each sample in the template is determined based on the weight derivation mode, and a weight matrix consisting of the weights of each sample in the template is defined as the template weight.
[0463] Method 2: Determine the weights of the current block and the template based on the weight derivation mode. That is, in this method 2, the combined region consisting of the current block and the template is treated as a whole, and the weights of the samples in the combined region are derived based on the weight derivation mode.
[0464] For example, the decoding side determines weights of samples in a combination region consisting of the current block and the template based on the angle index, distance index, template size, and current block size, and determines template weights based on the template size and sample weights in the combination region.
[0465] In the second method, the weights of the samples in the combination region formed by the current block and the template are determined based on the angle index, distance index, template size, and current block size for the current block and the template as a whole. Then, the weights corresponding to the template in the combination region are determined as template weights based on the template size. For example, as shown in Figures 22A and 22B, the weights corresponding to the L-shaped template region in the combination region are determined as template weights.
[0466] In one example, the derivation process of template weights in Scheme 2 is as follows:
[0467] The inputs to this process include the width of the current block nCbW, the height of the current block nCbH, the width of the left template nTmW, the height of the top template nTmH, the GPM "split" angle index variable angleIdx, the GPM distance index variable distanceIdx, and the component index variable cIdx. Since this example only covers luminance, an Idx of 0 indicates the luminance component.
[0468] The output of this process is the template weight matrix wTemplateValue.
[0469] The variables nW, nH, shift1, offset1, displacementX, displacementY, partFlip, and shiftHor are derived based on the following method. nW = ( cIdx = = 0 ) ? nCbW : nCbW * SubWidthC nH = ( cIdx = = 0 ) ? nCbH : nCbH * SubHeightC shift1 = Max( 5, 17 - BitDepth ), where BitDepth is the bit depth of the coding. offset1 = 1 << ( shift1 - 1 ) displacementX = angleIdx displacementY = (angleIdx + 8) % 32 partFlip = ( angleIdx >= 13 && angleIdx <= 27 ) ? 0 : 1 shiftHor = ( angleIdx % 16 = = 8 | | ( angleIdx % 16 != 0 && nH >= nW ) ) ? 0 : 1
[0470] The offsets offsetX and offsetY are derived based on the following method. - If the value of shiftHor is 0, offsetX = ( -nW ) >> 1 offsetY = ( ( -nH ) >> 1 ) + ( angleIdx < 16 ? ( distanceIdx * nH ) >> 3 : -( ( distanceIdx * nH ) >> 3 ) ) - Otherwise (i.e., shiftHor has the value 1), offsetX = ( ( -nW ) >> 1 ) + ( angleIdx < 16 ? ( distanceIdx * nW ) >> 3 : -( ( distanceIdx * nW ) >> 3 ) offsetY = ( - nH ) >> 1
[0471] The template weight matrix wTemplateValue[x][y] (x = -nTmW..nCbW - 1, y = -nTmH..nCbH - 1, except when both x and y are 0 or greater) is derived based on the following method (note that in this example, the coordinates of the upper left corner of the current block are set to (0,0)): - The variables xL and yL are derived based on the following method: xL = ( cIdx = = 0 ) ? x : x * SubWidthC yL = ( cIdx = = 0 ) ? y : y * SubHeightC
[0472] The disLut is specified according to Table 3 above.
[0473] weightIdx = ( ( ( xL + offsetX ) << 1 ) + 1 ) * disLut[ displacementX ] + ( ( ( yL + offsetY ) << 1 ) + 1 ) * disLut[ displacementY ] weightIdxL = partFlip ? 32 + weightIdx : 32 - weightIdx wTemplateValue[x][y] = Clip3( 0, 8, ( weightIdxL + 4 ) >> 3 )
[0474] In some embodiments, for computational simplicity, the template weights may be set to only two possible values, for example, 0 and 1.
[0475] In one example, the weights of the samples in the template can be determined based on the following formula: wVemplateValue[x][y] = (partFlip ? weightIdx: - weightIdx) > 0 ? 1 : 0.
[0476] The above method identifies template weights and predicted values of the K templates corresponding to a second combination, and weights the predicted values of the K templates using the template weights to obtain predicted values of the templates in the second combination.
[0477] Since the template of the current block is a reconstructed region, the reconstruction value of the template can be obtained at the decoding side. In this way, for each second combination among the T second combinations, a cost corresponding to the second combination can be determined based on the predicted value of the template in the second combination and the reconstructed value of the template. Methods for determining the cost corresponding to the second combination include, but are not limited to, SAD, SATD, SSE, etc. Next, a candidate combination list is constructed based on the cost corresponding to each second combination among the T second combinations.
[0478] In an embodiment of the present application, the manner of identifying the predicted value of the template corresponding to the second combination includes at least the following manner.
[0479] In the first method, the predicted value of the template corresponding to the second combination is a single value. That is, at the decoding side, the template is predicted using the K prediction modes included in the second combination to obtain K predicted values. Template weights are determined based on the weight derivation modes included in the second combination. The K predicted values are weighted using the template weights to obtain weighted prediction values. The weighted prediction values are determined as the predicted value of the template corresponding to the second combination.
[0480] In the second method, in some embodiments, a hierarchical screening approach can be used. For example, if a weight derivation mode can obtain a relatively small cost, weight derivation modes similar to this weight derivation mode can be subsequently tried. Conversely, if a weight derivation mode cannot obtain a relatively small cost, weight derivation modes similar to this weight derivation mode are not subsequently tried. For example, if a certain intra prediction mode can obtain a relatively small cost, intra prediction modes similar to this intra prediction mode can be subsequently tried. Conversely, if a certain intra prediction mode cannot obtain a relatively small cost, intra prediction modes similar to this intra prediction mode are not subsequently tried. Of course, these screening methods can also be limited to use cases in which they are combined with the other two factors. For example, if a certain intra prediction mode that is set as the first prediction mode under a certain weight derivation mode cannot obtain a relatively small cost, intra prediction modes similar to this intra prediction mode are not subsequently tried as the first prediction mode under this weight derivation mode.
[0481] In the third method, a cost corresponding to each second combination is determined using a fast cost calculation method. As can be seen from the above, the predicted values of the template corresponding to the second combination include predicted values of the template corresponding to each of the K prediction modes included in the second combination. In this case, the cost corresponding to each of the K prediction modes in the second combination can be determined based on the predicted values of the template and the reconstructed values of the template corresponding to each of the K prediction modes in the second combination. The cost corresponding to the second combination is determined based on the costs corresponding to each of the K prediction modes in the second combination. For example, the sum of the costs corresponding to each of the K prediction modes in the second combination is determined as the cost corresponding to the second combination.
[0482] In an embodiment of the present application, assuming K=2, the weights on the template can be simplified to only two possibilities: 0 and 1. Therefore, for each sample position, the sample value is only from the prediction block of the first prediction mode or the prediction block of the second prediction mode. Therefore, for a certain prediction mode, it is possible to calculate the cost on the template when the prediction mode is the first prediction mode in a weight derivation mode. That is, only the cost on the template of samples with a weight of 1 when the prediction mode is the first prediction mode in the weight derivation mode is calculated. In one example, the cost is written as cost[pred_mode_idx][gpm_idx][0], where pred_mode_idx represents the index of the prediction mode, gpm_idx represents the index of the weight derivation mode, and 0 represents the first prediction mode.
[0483] Also, for a certain prediction mode, it is possible to calculate the cost on the template when the prediction mode is set as the second prediction mode in a certain weight derivation mode. That is, only the cost on the template of a sample with a weight of 1 when the prediction mode is set as the second prediction mode in the weight derivation mode is calculated. In one example, the cost is written as cost[pred_mode_idx][gpm_idx][1], where pred_mode_idx represents the index of the prediction mode, gpm_idx represents the index of the weight derivation mode, and 1 represents that the prediction mode is set as the second prediction mode.
[0484] Then, when calculating the cost corresponding to the combination, the two corresponding costs can be directly added together. For example, when calculating the cost corresponding to prediction modes pred_mode_idx0 and pred_mode_idx1 in weight derivation mode gpm_idx (pred_mode_idx0 is the first prediction mode, pred_mode_idx1 is the second prediction mode, and the cost is written as costTemp), costTemp=cost[pred_mode_idx0][gpm_idx][0]+cost[pred_mode_idx1][gpm_idx][1]. When calculating the cost corresponding to prediction modes pred_mode_idx0 and pred_mode_idx1 in weight derivation mode gpm_idx (pred_mode_idx1 is the first prediction mode, pred_mode_idx0 is the second prediction mode, and the cost is written as costTemp), costTemp = cost[pred_mode_idx1][gpm_idx][0] + cost[pred_mode_idx0][gpm_idx][1].
[0485] One advantage of doing so is that combining weights into a prediction block and then calculating the cost is simplified to directly calculating the costs of the two parts and then adding the costs of the two parts to obtain the cost corresponding to the combination. Because a prediction mode can be combined with multiple other prediction modes and, for the same weight derivation mode, the cost of the part where that prediction mode is the first prediction mode and the cost of the part where that prediction mode is the second prediction mode are fixed, these costs (i.e., cost[pred_mode_idx][gpm_idx][0] and cost[pred_mode_idx][gpm_idx][1] in the above example) can be reserved and reused to reduce the amount of calculation.
[0486] Based on the above method, a cost corresponding to each second combination in the T second combinations can be identified, and then a candidate combination list can be constructed based on the cost corresponding to each second combination in the T second combinations.
[0487] In an embodiment of the present application, the manner of identifying the candidate combination list based on the cost corresponding to each second combination in the T second combinations in S103-B22 includes, but is not limited to, some examples below.
[0488] Example 1: T second combinations are sorted based on the cost corresponding to each second combination in the T second combinations, and the sorted T second combinations are identified as a candidate combination list.
[0489] The candidate combination list generated in this Example 1 includes T first candidate combinations.
[0490] Optionally, the T first candidate combinations in the candidate combination list are sorted in ascending order of cost, that is, the costs corresponding to the T first candidate combinations in the candidate combination list increase sequentially according to the sorting order.
[0491] Sorting the T second combinations based on the cost corresponding to each second combination in the T second combinations may involve sorting the T second combinations according to ascending order of cost.
[0492] Example 2: C second combinations are selected from the T second combinations based on the costs associated with the second combinations, and the list of the C second combinations is identified as a candidate combination list.
[0493] Optionally, the C second combinations are the first C second combinations with the smallest cost among the T second combinations. For example, based on the cost corresponding to each second combination in the T second combinations, the C second combinations with the smallest cost are selected from the T second combinations to form a candidate combination list. In this case, the candidate combination list includes the C candidate combinations.
[0494] Optionally, the C candidate combinations in the candidate combination list are sorted in ascending order of cost, that is, the costs corresponding to the C candidate combinations in the candidate combination list increase sequentially according to the sorting order.
[0495] On the decoding side, based on the above steps, a candidate combination list is identified, a first combination corresponding to a first index is selected from the candidate combination list, a weight derivation mode included in the first combination is identified as a first weight derivation mode, and K prediction modes included in the first combination are identified as K first prediction modes.
[0496] Based on the above steps, the decoding side identifies a first weight derivation mode and K first prediction modes. Then, the following step S104 is executed.
[0497] S104: Predict the current block based on the first weight derivation mode and the K first prediction modes to obtain a predicted value of the current block.
[0498] In an embodiment of the present application, a decoding side identifies N candidate weight derivation modes and a candidate prediction mode list during decoding of a current block. The candidate prediction mode list includes at least one candidate prediction mode, and the at least one candidate prediction mode includes a prediction mode identified based on dividing a template of the current block. That is, in the embodiment of the present application, when identifying a candidate prediction mode, the prediction mode is derived from the template obtained by dividing. This allows for accurate derivation of a prediction mode and improves the accuracy of identifying the candidate prediction mode list. Next, a first weight derivation mode and K first prediction modes are identified based on the N candidate weight derivation modes and the accurately identified candidate prediction modes. This achieves accurate identification of the first weight derivation mode and the K first prediction modes. When predicting the current block based on the accurately identified first weight derivation mode and the K first prediction modes, prediction accuracy can be improved and decoding performance can be improved.
[0499] The embodiments of the present application do not limit the specific process of predicting the current block based on the first weight derivation mode and the K first prediction modes in the above S104 to obtain the predicted value of the current block.
[0500] In Situation 1, when determining the weights of the predictors, a weight gradient parameter (also known as a blending parameter) is not considered. In this case, the weights of the predictors of the current block are determined based on a first weight derivation mode. The current block is predicted based on K first prediction modes to obtain K predicted values of the current block. The weights of the predictors of the current block are weighted to obtain the predicted values of the current block. For the process of deriving the weights of the predictors of the current block based on the first weight derivation mode, please refer to the process of deriving the weights of the predictors of the current block in the above embodiment, and will not be repeated here.
[0501] In the situation 2, the weight gradient parameter is taken into consideration when determining the weight of the predictor value. In this case, the above S104 includes the following steps:
[0502] S104-A1: Identify weight gradient parameters.
[0503] S104-A2: Predict the current block based on the weight gradient parameter, the first weight derivation mode, and the K first prediction modes to obtain a predicted value of the current block.
[0504] In the above situation 1, the weight gradient is fixed. However, in some embodiments, a variable weight gradient can adjust the gradient of the weight change, so that the GPM can obtain blending regions with different widths under the same parting line angle and the same parting line offset.
[0505] For example, as shown in FIGS. 23A and 23B, FIG. 23A is a schematic diagram showing a blending area in GPM in VVC, and FIG. 23B is an example showing a variable weight gradient in GPM.
[0506] The value of blendingCoeff may be 1 / 4, 1 / 2, 1, 2, 4, etc.
[0507] Illustratively, the value of blendingCoeff can be derived from the weight gradient index gpm_blending_idx.
[0508] In some embodiments, the weight gradient index is also referred to as a blending gradient parameter or a blending parameter.
[0509] The embodiment of the present application does not limit the specific implementation process of S104-A2, for example, predicting the current block to obtain a predicted value according to the first weight derivation mode and the K first prediction modes, and then determining the predicted value of the current block according to the weight gradient parameter and the predicted value.
[0510] In some embodiments, the above S104-A2 includes the following steps:
[0511] S104-A21: Determine the weights of the predictors based on the weight gradient parameters and the first weight derivation mode.
[0512] S104-A22: Predict the current block according to the K first prediction modes to obtain K predicted values.
[0513] S104-A23: Weight the K predicted values according to the weights of the predicted values to obtain a predicted value of the current block.
[0514] The execution order of S104-A22 and S104-A21 is not particularly limited, that is, S104-A22 may be executed before S104-A21, after S104-A21, or in parallel with S104-A21.
[0515] In this situation 2, the decoding side determines a weight gradient parameter, determines weights of predictors based on the weight gradient parameter and the first weight derivation mode, predicts the current block based on the K first prediction modes to obtain K predicted values of the current block, and weights the K predicted values of the current block using the weights of the predictors to obtain a predicted value of the current block.
[0516] In an embodiment of the present application, the manner of identifying the weight gradient parameters includes at least some of the following:
[0517] Method 1: The bitstream is decoded to obtain a second index. The second index is used to indicate a weight gradient parameter. The weight gradient parameter is determined based on the second index. Specifically, the encoding side determines the weight gradient parameter, and then signals the second index corresponding to the weight gradient parameter in the bitstream. Next, the decoding side decodes the bitstream to obtain the second index, and determines the weight gradient parameter based on the second index.
[0518] In some embodiments, the second index is also referred to as a weight gradient index.
[0519] In one example, the syntax corresponding to Method 1 is shown in Table 8.
[0520] [Table 8]
[0521] In Table 8, gpm_cand_idx represents the first index and gpm_blending_idx represents the second index.
[0522] In some embodiments, different weight gradient parameters have little effect on the prediction results of the template. When a simplified method is used, i.e., when the weights on the template are only 0 and 1, the weight gradient parameters cannot affect the prediction of the template, i.e., cannot affect the candidate combination list, and in this case, the blending gradient index may not be included in the combination.
[0523] In the embodiment of the present application, a specific manner for identifying the weight gradient parameter based on the second index is not limited.
[0524] In some embodiments, the decoding side identifies a candidate blending parameter list, the candidate blending parameter list including a plurality of candidate blending parameters, and identifies a candidate blending parameter corresponding to a second index in the candidate blending parameter list as a weight gradient parameter.
[0525] The embodiments of the present application do not limit the manner in which the list of candidate blending parameters is identified.
[0526] In one example, the candidate blending parameters in the list of candidate blending parameters are pre-defined.
[0527] In another example, the decoding side selects at least one blending parameter from a plurality of preset blending parameters based on the feature information of the current block to form a list of candidate blending parameters, for example, selects a blending parameter that matches the image information of the current block from a plurality of preset blending parameters based on the image information of the current block to form a list of candidate blending parameters.
[0528] For example, assume that the image information includes image edge sharpness. If the image edge sharpness of the current block is less than a preset value, at least one first type weighted gradient parameter (e.g., 1 / 4, 1 / 2, etc.) from among the plurality of preset weighted gradient parameters is selected to form a list of candidate weighted gradient parameters. If the image edge sharpness of the current block is equal to or greater than the preset value, at least one second type weighted gradient parameter (e.g., 2, 4, etc.) from among the plurality of preset weighted gradient parameters is selected to form a list of candidate weighted gradient parameters.
[0529] Illustratively, a list of candidate weight gradient parameters for an embodiment of the present application is shown in Table 9.
[0530] [Table 9]
[0531] As shown in Table 9, the list of candidate weight gradient parameters includes multiple candidate weight gradient parameters, each of which corresponds to an index.
[0532] For example, in the above Table 9, the order of the candidate weight gradient parameters in the list of candidate weight gradient parameters is the index of the candidate weight gradient parameter. Optionally, the index of the candidate weight gradient parameter in the list of candidate weight gradient parameters may be expressed in other ways. The embodiment of the present application is not limited thereto.
[0533] Based on the above Table 9, the decoding side identifies the candidate weight gradient parameter corresponding to the second index in Table 9 as the weight gradient parameter based on the second index.
[0534] On the decoding side, in addition to decoding the bitstream to obtain the second index based on the above method 1 and then determining the weight gradient parameter based on the second index, the weight gradient parameter can also be determined based on the following method 2.
[0535] In some embodiments, it is also possible to not transmit the weight gradient index in the bitstream, and directly derive the weight gradient index gpm_blending_idx or blendingCoeff based on the block size, etc. On the decoding side, the weight gradient parameter can also be determined based on the following Scheme 2:
[0536] Method 2: The weight gradient parameters are identified by the following steps S104-A11 and S104-A12.
[0537] S104-A11: Identify multiple alternative weight gradient parameters, where G is a positive integer.
[0538] S104-A12: Identify a weight gradient parameter from a plurality of preliminary weight gradient parameters.
[0539] In this method 2, the decoding side determines the weight gradient parameter by itself, which avoids the encoding side from encoding the second index into the bitstream and saves codewords. Specifically, the decoding side first determines multiple preliminary weight gradient parameters, and then determines one preliminary weight gradient parameter from the multiple preliminary weight gradient parameters as the weight gradient parameter.
[0540] The embodiments of the present application do not limit the specific manner in which the decoding side specifies the plurality of preliminary weight gradient parameters.
[0541] In a possible embodiment, the plurality of preliminary weight gradient parameters are preset, i.e., the specification of the plurality of preset weight gradient parameters to the G preliminary weight gradient parameters is agreed upon between the decoding side and the encoding side.
[0542] In another possible embodiment, the plurality of preliminary weight gradient parameters can be specified by an encoding side, for example, the encoding side can specify that the plurality of weight gradient parameters among the plurality of preset weight gradient parameters are the plurality of preliminary weight gradient parameters.
[0543] In another possible embodiment, multiple preliminary weight gradient parameters can be identified based on the size of the current block.
[0544] In another possible embodiment, image information of the current block is identified, and a plurality of preliminary weighting gradient parameters are identified from a plurality of pre-set preliminary weighting gradient parameters based on the image information of the current block.
[0545] On the decoding side, after determining a plurality of preliminary weight gradient parameters, a weight gradient parameter is determined from the plurality of preliminary weight gradient parameters.
[0546] The embodiments of the present application do not limit the specific manner of determining the weight gradient parameter from the plurality of preliminary weight gradient parameters.
[0547] In some embodiments, any one of a plurality of preliminary weight gradient parameters is identified as the weight gradient parameter.
[0548] In some embodiments, a cost corresponding to each preliminary weight gradient parameter in the plurality of preliminary weight gradient parameters is identified. A weight gradient parameter is then identified from the plurality of preliminary weight gradient parameters based on the cost. For example, the weight gradient parameter with the smallest cost may be identified as the gradient parameter corresponding to the current block.
[0549] Method 3: Determine the weight gradient parameters based on the size of the current block.
[0550] As can be seen from the above, there is a certain correlation between the weight gradient parameter and the block size, so in the embodiment of the present application, the weight gradient parameter can also be determined based on the size of the current block.
[0551] In one possible embodiment, a fixed weight gradient parameter is specified as the weight gradient parameter based on the size of the current block.
[0552] For example, if the size of the current block is smaller than a first set threshold, the weight gradient parameter is set to a first value.
[0553] Further, for example, if the size of the current block is equal to or greater than a first set threshold, the weight gradient parameter is set to a second value, the second value being smaller than the first value.
[0554] The embodiments of the present application do not limit the specific values of the first value, the second value, and the first set threshold value.
[0555] Illustratively, the first value is 1 and the second value is 1 / 2.
[0556] Exemplarily, if the size of the current block is expressed as the number of samples (or pixels) in the current block, the first set threshold may be 256, etc.
[0557] In another possible embodiment, a numerical range for the weight gradient parameter is determined based on the size of the current block, and the weight gradient parameter is determined to be a value within that numerical range.
[0558] For example, if the size of the current block is smaller than a first set threshold, the weight gradient parameter is determined to be within the numerical range of the weight gradient parameters. For example, the weight gradient parameter is any one of the weight gradient parameters, such as the minimum weight gradient parameter, the maximum weight gradient parameter, or the intermediate weight gradient parameter, within the numerical range of the weight gradient parameters. As another example, the weight gradient parameter is the weight gradient parameter with the smallest cost within the numerical range of the weight gradient parameters. The method for determining the cost of the weight gradient parameter can be described with reference to other embodiments of the present application, and will not be repeated here.
[0559] Furthermore, for example, if the size of the current block is equal to or greater than the first set threshold, the weight gradient parameter is determined to be within the numerical range of the second weight gradient parameter. For example, the weight gradient parameter may be any one of the minimum weight gradient parameter, the maximum weight gradient parameter, or the intermediate weight gradient parameter within the numerical range of the second weight gradient parameters. As another example, the weight gradient parameter may be the weight gradient parameter with the smallest cost within the numerical range of the second weight gradient parameters. The minimum value of the numerical range of the second weight gradient parameter is smaller than the minimum value of the numerical range of the weight gradient parameter. The numerical range of the weight gradient parameter and the numerical range of the second weight gradient parameter may or may not overlap. This is not limited in the embodiment of the present application.
[0560] In this situation 2, after the weight gradient parameter is determined in the above step, the above step S104-A21 is executed to determine the weight of the predicted value based on the weight gradient parameter and the first weight derivation mode.
[0561] In the embodiment of the present application, the manner of determining the weight of the predictor value based on the weight gradient parameter and the first weight derivation mode includes at least the manners shown in the following several examples.
[0562] In Example 1, when deriving the weights for the predictors using the first weight derivation mode, multiple intermediate variables must be identified, and a weight gradient parameter can be used to adjust one or more of these multiple intermediate variables, and the adjusted variables can be used to derive the weights for the predictors.
[0563] In example 2, a weight index weightIdx corresponding to the current block is determined based on the first weight derivation mode and the current block, the weight index weightIdx is processed using a weight gradient parameter to obtain a processed weight index weightIdx, and a weight for the prediction value wVemplateValue is determined based on the processed weightIdx.
[0564] In one example, the weight gradient parameter can be used to specify the weight of the predictor value, wVemplateValue, as follows: …… weightIdx = ( ( ( xL + offsetX ) << 1 ) + 1 ) * disLut[ displacementX ] + ( ( ( yL + offsetY ) << 1 ) + 1 ) * disLut[ displacementY ] weightIdx = weightIdx * blendingCoeff weightIdxL = partFlip ? 32 + weightIdx : 32 - weightIdx wValue = Clip3( 0, 8, ( weightIdxL + 4 ) >> 3 ) blendingCoeff1 is the weight gradient parameter.
[0565] Then, the current block is predicted based on the K first prediction modes to obtain K predicted values, and the K predicted values are weighted based on the weights of the predicted values to obtain a predicted value of the current block.
[0566] In the above embodiment, the template weight and the predictor weight can be understood as two independent processes that do not interfere with each other. The above method allows the predictor weight to be determined independently.
[0567] In some embodiments, when determining the template weight, a combination region consisting of the template region and the current block is considered. When determining the template weight by determining the weight of the combination region, since the combination region includes the current block, the weight corresponding to the current block in the weight of the combination region is determined as the weight of the prediction value. Note that, when determining the weight of the combination region, the influence of the weight gradient parameter on the weight is also taken into account. For details, please refer to the above embodiments, and the description will not be repeated here.
[0568] In some embodiments, the prediction process is performed sample by sample, and the weights of the predictors correspond to the samples accordingly. In this case, when predicting the current block, a sample A in the current block is predicted using each of K first prediction modes to obtain K predicted values of sample A in the K first prediction modes. The weight of the predicted value of sample A is determined based on the first weight derivation mode and the weight gradient parameter. Then, the weight of the predicted value of sample A is used to weight these K predicted values to obtain a predicted value of sample A. By performing the above steps for each sample in the current block, a predicted value of each sample in the current block can be obtained. The predicted values of each sample in the current block constitute a predicted value of the current block. For example, when K=2, a sample A in the current block is predicted using a first prediction mode to obtain a first predicted value of sample A, and a second prediction mode to predict sample A to obtain a second predicted value of sample A. The first and second predicted values are weighted based on the weight of the predicted value corresponding to sample A to obtain a predicted value of sample A.
[0569] For example, when K=2, if the first prediction mode and the second prediction mode are both intra prediction modes, a first predicted value is obtained by performing prediction using the first intra prediction mode, a second predicted value is obtained by performing prediction using the second intra prediction mode, and the first and second predicted values are weighted based on the weights of the predicted values to obtain a predicted value of the current block. For example, sample A is predicted using the first intra prediction mode to obtain a first predicted value of sample A, and sample A is predicted using the second intra prediction mode to obtain a second predicted value of sample A, and the first and second predicted values are weighted based on the weight of the prediction value corresponding to sample A to obtain a predicted value of sample A.
[0570] In some embodiments, when K is greater than 2, weights of predictors corresponding to two of the K first prediction modes may be determined based on the first weight derivation mode. Weights of predictors corresponding to other of the K first prediction modes may be preset values. For example, when K=3, first weights of predictors corresponding to the first and second prediction modes are derived based on the weight derivation mode, and the weight of predictor corresponding to the third prediction mode is a preset value. In some embodiments, when the total weight of predictors corresponding to the K first prediction modes is fixed, e.g., 8, weights of predictors corresponding to each of the K first prediction modes may be determined based on a preset weight ratio. Assuming that the weight of predictor corresponding to the third prediction mode accounts for 1 / 4 of the total weight of predictors, the weight of predictor corresponding to the third prediction mode may be determined to be 2, and the remaining 3 / 4 of the total weight of predictors is assigned to the first and second prediction modes. For example, if the weight of the predicted value corresponding to the first prediction mode derived based on the first weight derivation mode is 3, the weight of the predicted value corresponding to the first prediction mode is determined to be (3 / 4)*3, and the weight of the predicted value corresponding to the second prediction mode is determined to be (3 / 4)*5.
[0571] Based on the above method, a predicted value of the current block is determined, and the quantized coefficients of the current block are obtained by decoding the bitstream, and the quantized coefficients of the current block are inversely quantized and inversely transformed to obtain a residual value of the current block. The predicted value and the residual value of the current block are added to obtain a reconstructed value of the current block.
[0572] In a video decoding method according to an embodiment of the present application, a decoding side identifies N candidate weight derivation modes and a candidate prediction mode list during decoding of a current block. The candidate prediction mode list includes at least one candidate prediction mode, and the at least one candidate prediction mode includes a prediction mode identified based on dividing a template of the current block. That is, in the embodiment of the present application, when identifying a candidate prediction mode, the prediction mode is derived from the template obtained by dividing. This allows for accurate derivation of a prediction mode and improves the accuracy of identifying the candidate prediction mode list. Next, a first weight derivation mode and K first prediction modes are identified based on the N candidate weight derivation modes and the accurately identified candidate prediction modes. This allows for improved accuracy of identifying the first weight derivation mode and the K first prediction modes. When predicting the current block based on the accurately identified first weight derivation mode and the K first prediction modes, prediction accuracy can be improved and decoding performance can be improved.
[0573] The video decoding method of the present application has been introduced above by taking the decoding side as an example, and the following description will be given by taking the encoding side as an example.
[0574] 24 is a flowchart illustrating a video encoding method according to an embodiment of the present application. The embodiment of the present application is applied to the video encoder shown in FIG. 1 and FIG. 2. As shown in FIG. 24, the method of the embodiment of the present application includes the following contents:
[0575] S201: Identify N candidate weight derivation modes.
[0576] N is a positive integer. Optionally, N is a preset value or a default value. Optionally, N may be specified by other methods on the encoding side, and the embodiment of the present application is not limited thereto.
[0577] As can be seen from the above, in the embodiment of the present application, a predicted block is generated based on one weight derivation mode and K prediction modes, and the predicted block is applied to the current block, that is, weights are determined based on the weight derivation mode, the current block is predicted based on the K prediction modes to obtain K predicted values, and the K predicted values are weighted based on the weights to obtain a predicted value of the current block.
[0578] That is, when encoding a current block, the encoding side needs to identify N candidate weight derivation modes and multiple candidate prediction modes. Then, it selects one weight derivation mode from the N candidate weight derivation modes and selects K prediction modes from the multiple candidate prediction modes. Then, it predicts the current block using the selected one weight derivation mode and the K prediction modes to obtain a predicted value of the current block.
[0579] In the embodiment of the present application, the method for identifying the N candidate weight derivation modes by the decoding side is not limited.
[0580] In one possible embodiment, there are 56 weight derivation modes in the AWP and 64 weight derivation modes in the GPM, and the N candidate weight derivation modes include at least one weight derivation mode of the 56 weight derivation modes in the AWP or at least one weight derivation mode of the 64 weight derivation modes in the GPM.
[0581] In one possible embodiment, several weight derivation modes in the AWP or GPM can be screened to obtain N candidate weight derivation modes. That is, the N candidate weight derivation modes in the embodiment of the present application are a subset of all weight derivation modes in the AWP or GPM. For example, in a weight derivation mode, the same "division" angle can correspond to multiple offsets. For example, weight derivation modes 10, 11, 12, and 13 shown in FIG. 4 or FIG. 5 have the same "division" angle but different offsets. In the embodiment of the present application, modes corresponding to several offsets can be eliminated. Of course, modes corresponding to several "division" angles can also be eliminated. In this way, the total number of possible combinations can be reduced, thereby making the differences between each possible combination more apparent. Of course, different screening methods can be set for blocks of different sizes. For example, fewer weight derivation modes are used for small blocks and more weight derivation modes are used for large blocks. Also, different screening methods can be set for blocks of different shapes. The shape of the block can refer to the ratio of width to height.
[0582] In this embodiment, the screening method of the N candidate weight derivation modes on the encoding side is the same as that on the decoding side. In one example, the screening method of the N candidate weight derivation modes is default on both the encoding side and the decoding side. In another example, the encoding side instructs the screening method of the N candidate weight derivation modes to the encoding side, so that the decoding side screens the same N candidate weight derivation modes as the encoding side in the same way.
[0583] In some embodiments, weight derivation modes corresponding to preset division angles and / or preset offsets are eliminated from the preset M weight derivation modes to obtain N weight derivation modes. The same division angle in a weight derivation mode may correspond to multiple offsets; for example, as shown in Figure 4, weight derivation modes 10, 11, 12, and 13 have the same division angle but different offsets. Therefore, weight derivation modes corresponding to some preset offsets and / or weight derivation modes corresponding to some preset division angles may be eliminated.
[0584] In some embodiments, the screening conditions corresponding to different blocks may be different. Thus, when identifying the N weight derivation modes corresponding to the current block, the screening conditions corresponding to the current block are first identified, and the N weight derivation modes are selected from the M preset weight derivation modes based on the screening conditions corresponding to the current block.
[0585] In some embodiments, the screening conditions corresponding to the current block include screening conditions corresponding to the size of the current block and / or screening conditions corresponding to the shape of the current block. During prediction, for relatively small blocks, the difference in the impact of similar weight derivation modes on the prediction result is not significant. For relatively large blocks, the difference in the impact of similar weight derivation modes on the prediction result is more obvious. Based on this, in embodiments of the present application, different N values are set for blocks of different sizes, i.e., a relatively large N value is set for relatively large blocks, and a relatively small N value is set for relatively small blocks.
[0586] In one possible embodiment, the encoding side indicates N candidate weight derivation modes to the decoding side.
[0587] In some embodiments, the screening condition comprises an array having N elements, the N elements corresponding to the N weight derivation modes in a one-to-one correspondence, and the element corresponding to each weight derivation mode is used to indicate whether the weight derivation mode is available.
[0588] The array may be a one-dimensional array or a two-dimensional array.
[0589] For example, taking GPM as an example, the total number of possible weight derivation modes is 64. On the encoding side, a lookup table containing 64 elements is set up, and the value of each element indicates whether the weight derivation mode corresponding to the element is used.
[0590] In one example, taking a one-dimensional array as an example, a specific example is as follows: An array of g_sgpm_splitDir is set. g_sgpm_splitDir
[64] = { 1,1,1,0,1,0,1,0, 1,0,1,0,1,0,1,0, 1,0,1,1,1,0,1,0, 1,0,1,0,1,0,1,0, 0,0,0,0,1,1,0,1, 0,0,1,0,0,1,0,0, 1,0,1,1,0,1,0,0, 1,0,0,1,0,0,1,0 }.
[0591] If the value of g_sgpm_splitDir[x] is 1, it indicates that the weight derivation mode with index x is available. Otherwise, it indicates that the weight derivation mode with index x is not available. In this example, on the encoding side, 26 candidate weight derivation modes are identified based on this array.
[0592] In another example, the N candidate weight derivation modes can be represented by one array. The array contains only the indices of the available weight derivation modes. For example, the 26 candidate weight derivation modes are represented by an array g_sgpm_splitDir
[26] ={0, 1, 6, 8, 10, 12, 14, 16, 18, 19, 20, 22, 24, 26, 28, 30, 36, 37, 42, 45, 48, 50, 51, 53, 56, 59} having a length of 26. On the encoding side, the weight derivation modes corresponding to the indices are identified as candidate weight derivation modes based on the indices of the weight derivation modes contained in the array, thereby obtaining 26 candidate weight derivation modes.
[0593] In some embodiments, if the screening conditions corresponding to the current block include a screening condition corresponding to the size of the current block and a screening condition corresponding to the shape of the current block, and if, for the same weight derivation mode, both the screening condition corresponding to the size of the current block and the screening condition corresponding to the shape of the current block indicate that the weight derivation mode is available, the weight derivation mode is identified as one of the N weight derivation modes. If at least one of the screening condition corresponding to the size of the current block and the screening condition corresponding to the shape of the current block indicates that the weight derivation mode is unavailable, the weight derivation mode does not belong to the N weight derivation modes.
[0594] In some embodiments, screening conditions corresponding to different block sizes and screening conditions corresponding to different block shapes can each be realized using multiple arrays.
[0595] In some embodiments, the screening conditions corresponding to different block sizes and the screening conditions corresponding to different block shapes can be implemented in a two-dimensional array, i.e., the two-dimensional array includes both the screening conditions corresponding to the block sizes and the screening conditions corresponding to the block shapes.
[0596] For example, the screening conditions corresponding to a block having a size of A and a shape of B are shown below: The screening conditions are expressed as a two-dimensional array. g_sgpm_splitDir
[64] = { (1,1),(1,1),(1,1),(1,0),(1,0),(0,0),(1,0),(1,1), (1,1),(0,0),(1,1),(1,0),(1,0),(0,0),(1,0),(1,1), (0,1),(0,0),(1,1),(0,0),(1,0),(0,0),(1,0),(0,0), (1,1),(0,0),(0,1),(1,0),(1,0),(1,0),(1,0),(0,0), (0,0),(0,0),(1,1),(0,0),(1,1),(1,1),(1,0),(0,1), (0,0),(0,0),(1,1),(0,0),(1,0),(0,0),(1,0),(0,0), (1,0),(0,0),(1,1),(1,0),(1,0),(1,0),(0,0),(0,0), (1,1),(0,0),(1,1),(0,0),(0,0),(1,0),(1,1),(0,0) }.
[0597] All values of g_sgpm_splitDir[x] equal to 1 indicate that the weight derivation mode with index x is available, and one of the values of g_sgpm_splitDir[x] equal to 0 indicates that the weight derivation mode with index x is unavailable. For example, g_sgpm_splitDir[4]=(1,0) indicates that weight derivation mode 4 is available for blocks with size A but unavailable for blocks with shape B. Therefore, if the size of a block is A and the shape of the block is B, the weight derivation mode is unavailable.
[0598] Note that the above is an example of 64 weight derivation modes in GPM, but the weight derivation modes of the embodiments of the present application include, but are not limited to, 64 weight derivation modes in GPM and 56 weight derivation modes in AMP.
[0599] In some embodiments, before identifying the N candidate weight derivation modes, the encoding side needs to identify whether K different prediction modes are used for weighted prediction for the current block. If the encoding side identifies that K different prediction modes are used for weighted prediction for the current block, the encoding side performs the above-mentioned S101 to identify N candidate weight derivation modes. If the encoding side identifies that K different prediction modes are not used for weighted prediction for the current block, step S101 is skipped.
[0600] In a possible embodiment, the encoding side can specify whether K different prediction modes are used for the current block for weighted prediction by specifying a prediction mode parameter for the current block.
[0601] Optionally, in an embodiment of the present application, the prediction mode parameter may indicate whether GPM mode or AWP mode can be used for the current block, i.e., whether K different prediction modes can be used for prediction of the current block.
[0602] Note that, in an embodiment of the present application, the prediction mode parameter may be understood as a flag indicating whether the GPM mode or the AWP mode is used. Specifically, the encoder may use one variable as the prediction mode parameter, thereby setting the value of the variable to set the prediction mode parameter. Exemplarily, in the present application, if the GPM mode or the AWP mode is used for the current block, the encoder may set the value of the prediction mode parameter to indicate that the GPM mode or the AWP mode is used for the current block. Specifically, the encoder may set the value of the variable to 1. Exemplarily, in the present application, if the GPM mode or the AWP mode is not used for the current block, the encoder may set the value of the prediction mode parameter to indicate that the GPM mode or the AWP mode is not used for the current block. Specifically, the encoder may set the value of the variable to 0. Furthermore, in an embodiment of the present application, after completing the setting of the prediction mode parameter, the encoder may signal the prediction mode parameter to a bitstream and transmit it to the decoder. As a result, the decoder can obtain the prediction mode parameter after parsing the bitstream.
[0603] In some embodiments, the embodiment of the present application may also set a condition regarding whether the GPM mode or the AWP mode is used for the current block, that is, if it is determined that the current block satisfies the preset condition, K prediction modes are identified to be used for the current block for weighted prediction, and then N candidate weight derivation modes corresponding to the current block are identified.
[0604] For example, when using GPM mode or AWP mode, the size of the current block can be limited.
[0605] As can be understood, in a video encoding method according to an embodiment of the present application, K predicted values need to be generated using K different prediction modes, and then the K predicted values are weighted based on the weights to obtain a predicted value for a current block. To reduce complexity and considering the trade-off between compression performance and complexity, an embodiment of the present application may impose a restriction that the GPM mode or the AWP mode is not used for blocks of a certain size. Therefore, in the present application, the encoder may first determine a size parameter of the current block, and then determine whether the GPM mode or the AWP mode is used for the current block based on the size parameter.
[0606] In an embodiment of the present application, the size parameters of the current block may include the height and width of the current block, so that the encoder can determine whether the GPM mode or the AWP mode is used for the current block based on the height and width of the current block.
[0607] Illustratively, the present application specifies that GPM mode or AWP mode can be used for the current block if the width is greater than threshold 1 and the height is greater than threshold 2. Thus, as can be seen from the above, one possible restriction is to use GPM mode or AWP mode only if the width of the block is greater than (or equal to or greater than) threshold 1 and the height of the block is greater than (or equal to or greater than) threshold 2. The values of threshold 1 and threshold 2 may be 4, 8, 16, 32, 128, 256, etc., and threshold 1 may be equal to threshold 2.
[0608] Illustratively, the present application specifies that GPM mode or AWP mode can be used for the current block if the width is less than threshold 3 and the height is greater than threshold 4. As can be seen from the above, one possible restriction is to use GPM mode or AWP mode only if the width of the block is less than (or equal to or less than) threshold 3 and the height of the block is greater than (or equal to or greater than) threshold 4. The values of threshold 3 and threshold 4 may be 4, 8, 16, 32, 128, 256, etc., and threshold 3 may be equal to threshold 4.
[0609] Furthermore, in embodiments of the present application, restrictions on sample parameters can limit the size of blocks in which GPM mode or AWP mode can be used.
[0610] Illustratively, in this application, the encoder may first determine the sample parameters of the current block, and then determine whether the GPM mode or the AWP mode can be used for the current block based on the sample parameters and threshold 5. As can be seen from the above, one possible restriction is to use the GPM mode or the AWP mode only if the number of samples in the block is greater than (or equal to or greater than) threshold 5. The value of threshold 5 may be 4, 8, 16, 32, 128, 256, 1024, etc.
[0611] That is, in this application, the GPM mode or the AWP mode can be used for the current block only under the condition that the size parameter of the current block meets the size requirement.
[0612] For example, in this application, there may be an image-level flag for specifying whether this application is used for the current image to be encoded. For example, this application may be configured to be used for intraframes (e.g., I frames) but not for interframes (e.g., B frames, P frames). Alternatively, this application may be configured to be not used for intraframes but to be used for interframes. Alternatively, this application may be configured to be used for some interframes but not for other parts of the interframes. Since intraframe prediction can also be used for interframes, this application may also be used for interframes.
[0613] In some embodiments, a flag below the image level can be used to identify whether the present application is used for the current block.
[0614] S202: A candidate prediction mode list is identified.
[0615] The candidate prediction mode list includes at least one candidate prediction mode, the at least one candidate prediction mode including a prediction mode identified based on dividing a template of the current block.
[0616] In some embodiments, when templates are used in TIMD, the entire set of templates, including the left template and the top template, are used together to derive the intra prediction mode for the TIMD. If a template for one side does not exist, for example, when the current block is located at the left or top boundary of the image, only the existing template can be used for the TIMD. However, if templates for both sides exist, these templates are used together. When neighboring reconstructed samples are used for DIMD, the left reconstructed sample and the top reconstructed sample are used together to derive the intra prediction mode for the DIMD. For example, when a reconstructed sample for one side does not exist, such as when the current block is located at the left or top boundary of the image, only the existing reconstructed sample can be used for the DIMD. However, if both the left reconstructed sample and the top reconstructed sample exist, these reconstructed samples are used together. This is not a problem for non-GPM blocks. However, for GPM blocks, the correlation between the first prediction mode and the left and top templates or the left and top reconstructed samples is different from the correlation between the second prediction mode and the left and top templates or the left and top reconstructed samples. If a prediction mode is derived using the entire template, the derived prediction mode will have low accuracy, resulting in an inaccurate candidate prediction mode list being constructed, and if a current block is predicted based on the inaccurate candidate prediction mode list, accurate prediction of the current block cannot be achieved.
[0617] To solve this technical problem, in an embodiment of the present application, a template of a current block is divided. Template division can be understood as dividing the template into a plurality of sub-templates, or dividing a reconstructed sample area where the template is located into a plurality of reconstructed sample sub-areas. In this way, by deriving a prediction mode based on the template obtained by division or the reconstructed sample area obtained by division, the accuracy of derivation of the prediction mode can be improved, and the accuracy of construction of a candidate prediction mode list can be improved. By performing prediction based on an accurately constructed candidate prediction mode list, prediction accuracy can be improved, and encoding performance can be improved.
[0618] In the embodiment of the present application, a specific method for identifying the candidate prediction mode list is not limited.
[0619] In some embodiments, the process of identifying the candidate prediction mode list is independent of the N candidate weight derivation modes. That is, the N candidate weight derivation modes can be understood as corresponding to one candidate prediction mode list. This can reduce the complexity of identifying the candidate prediction mode list and improve encoding efficiency. Note that, in this embodiment, since the candidate prediction mode list is independent of the N candidate weight derivation modes, there is no restriction on the order between S202 and S201. That is, S202 can be performed after S201, before S201, or simultaneously with S201, and this is not limited in the embodiments of the present application.
[0620] In some embodiments, S202 includes the following step S202-A.
[0621] S202-A: For each first candidate weight derivation mode in the N candidate weight derivation modes, identify a candidate prediction mode list corresponding to the first candidate weight derivation mode.
[0622] In one example, the first candidate weight derivation mode is any one of N candidate weight derivation modes. That is, in this example, at least one candidate prediction mode list needs to be identified for each of the N candidate weight derivation modes. As can be seen from the above, one weight derivation mode corresponds to K prediction modes, and the candidate prediction mode list is used to identify the prediction mode. Therefore, in one possible implementation of this example, one candidate prediction mode list is identified for at least one prediction mode among the K prediction modes corresponding to each candidate weight derivation mode in the N candidate weight derivation modes.
[0623] In another example, if the first candidate weight derivation mode belongs to one type of candidate weight derivation mode among N candidate weight derivation modes, in the embodiment of the present application, the N candidate weight derivation modes need to be classified, and at least one candidate prediction mode list is constructed for each type of candidate weight derivation mode.
[0624] Specifically, the encoding side identifies angle indices corresponding to N candidate weight derivation modes. For example, the encoding side identifies angle indices corresponding to each of the N candidate weight derivation modes. The method for identifying the angle indices can be found in the detailed description of the above embodiment, and will not be described again in this specification. Next, the encoding side classifies the N candidate weight derivation modes into M types of candidate weight derivation modes based on the angle indices corresponding to each candidate weight derivation mode. Candidate weight derivation modes of the same type have the same angle indices. That is, the encoding side classifies candidate weight derivation modes having the same angle indices into one type based on the angle indices corresponding to each candidate weight derivation mode, thereby obtaining M types of candidate weight derivation modes. Each type of candidate weight derivation mode includes at least one candidate weight derivation mode. Furthermore, the jth type of candidate weight derivation mode among the M types of candidate weight derivation modes is identified as the first weight derivation mode, where j is a positive integer equal to or less than M. In this example, at least one candidate prediction mode list is identified for each type of candidate weight derivation mode among the N candidate weight derivation modes.
[0625] In the embodiment of the present application, the method for identifying the candidate prediction mode list corresponding to each first candidate weight derivation mode among the N candidate weight derivation modes is the same. For ease of explanation, the embodiment of the present application takes the identification of the candidate prediction mode list corresponding to one first candidate weight derivation mode as an example.
[0626] A specific method for identifying the candidate prediction mode list corresponding to the first candidate weight derivation mode in S202-A will be introduced below.
[0627] In some embodiments, the first candidate weight derivation mode corresponds to one candidate prediction mode list.
[0628] In some embodiments, S202-A includes the following step S202-A1:
[0629] S202-A1: Identify a candidate prediction mode list for at least one prediction mode among the K prediction modes corresponding to a first candidate weight derivation mode.
[0630] In this embodiment, a candidate prediction mode list is identified for at least one prediction mode among the K prediction modes corresponding to the first candidate weight derivation mode. Assuming that K=2, selectively, one candidate prediction mode list may be identified for the first prediction mode, but no candidate prediction mode list may be identified for the second candidate prediction mode. Selectively, one candidate prediction mode list may be identified for the second prediction mode, but no candidate prediction mode list may be identified for the first candidate prediction mode. Selectively, one candidate prediction mode list may be identified for the first prediction mode, and one candidate prediction mode list may be identified for the second candidate prediction mode. Selectively, one common candidate prediction mode list may be identified for the first prediction mode and the second prediction mode.
[0631] In an embodiment of the present application, a candidate prediction mode list is identified for at least one prediction mode corresponding to a first candidate weight derivation mode, and at least one prediction mode corresponding to the first candidate weight derivation mode is accurately identified from the constructed candidate prediction mode list.
[0632] In some embodiments, when the at least one prediction mode corresponds to one candidate prediction mode list, the aforementioned S202-A1 includes the following steps S202-A1-11 and S202-A1-12.
[0633] S202-A1-11: Identify a candidate prediction mode list for the i-th prediction mode in the at least one prediction mode, where i is a positive integer.
[0634] S202-A1-12: Identify a candidate prediction mode list for at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode.
[0635] In this embodiment, at least one prediction mode corresponding to the first candidate weight derivation mode corresponds to one candidate prediction mode list. That is, the at least one prediction mode corresponds to the same candidate prediction mode list, that is, corresponds to one candidate prediction mode list. In this way, the complexity of identifying the candidate prediction mode list can be reduced, and encoding efficiency can be improved. In this case, the encoding side identifies one candidate prediction mode list for the at least one prediction mode.
[0636] Specifically, a candidate prediction mode list for an i-th prediction mode among the at least one prediction mode is identified, where the i-th prediction mode is any prediction mode among the at least one prediction mode, and a candidate prediction mode list for the at least one prediction mode is identified based on the candidate prediction mode list for the i-th prediction mode.
[0637] Note that specific methods for identifying the candidate prediction mode list for at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode in S202-A1-12 above include, but are not limited to, the following:
[0638] Method 1: The candidate prediction mode list for the i-th prediction mode is directly identified as the candidate prediction mode list for at least one prediction mode.
[0639] Method 2: Determine whether the candidate prediction mode list for the i-th prediction mode includes a preset prediction mode. If the candidate prediction mode list for the i-th prediction mode includes a preset prediction mode, identify the candidate prediction mode list for the i-th prediction mode as a candidate prediction mode list for at least one prediction mode. If the candidate prediction mode list for the i-th prediction mode does not include a preset prediction mode, add the preset prediction mode to the candidate prediction mode list for the i-th prediction mode to obtain a candidate prediction mode list for at least one prediction mode.
[0640] In the embodiment of the present application, the preset prediction mode in the above Scheme 2 is not limited, and is specifically specified according to actual needs.
[0641] In this embodiment, when the at least one prediction mode corresponds to one candidate prediction mode list, a specific process of identifying a candidate prediction mode list for the at least one prediction mode is introduced.
[0642] In some embodiments, when each prediction mode in the at least one prediction mode corresponds to one candidate prediction mode list, the above S202-A1 includes the following step S202-A1-21.
[0643] S202-A1-21: For the ith prediction mode in the at least one prediction mode, identify a candidate prediction mode list for the ith prediction mode, where i is a positive integer.
[0644] In this embodiment, each prediction mode in the at least one prediction mode corresponds to one candidate prediction mode list. Thus, for a first candidate weight derivation mode, the encoding side identifies one candidate prediction mode list for each prediction mode in the at least one prediction mode corresponding to the first candidate weight derivation mode. For example, the at least one prediction mode includes a first prediction mode and a second prediction mode corresponding to the first candidate weight derivation mode. Accordingly, the encoding side identifies one candidate prediction mode list for the first prediction mode and one candidate prediction mode list for the second prediction mode.
[0645] In this embodiment, the process of identifying one candidate prediction mode list corresponding to each prediction mode in the at least one prediction mode is the same. For ease of explanation, the embodiment of the present application will take the identification of the candidate prediction mode list for the ith prediction mode in the at least one prediction mode as an example.
[0646] The process of identifying the candidate prediction mode list for the i-th prediction mode in S202-A1-11 and S202-A1-21 above will be described below.
[0647] In the embodiment of the present application, the specific types of candidate prediction modes included in the candidate prediction mode list for the i-th prediction mode are not limited.
[0648] In some embodiments, the candidate prediction mode list for the i-th prediction mode includes at least one of a first candidate prediction mode identified based on a template of the current block and a second candidate prediction mode identified based on gradients of reconstructed samples in the template.
[0649] Situation 1: When the candidate prediction mode list for the i-th prediction mode includes the first candidate prediction mode, the embodiment of the present application includes the following steps 31 to 34.
[0650] Step 31: Divide the template of the current block into P sub-templates, where P is a positive integer greater than 1.
[0651] As can be seen from the above, the template of the current block includes the left template of the current block and the upper template of the current block.Currently, the entire template of the current block is used to derive the first candidate prediction mode.For example, in TIMD, the entire template of the current block is used to derive the prediction mode, so that the derived first candidate prediction mode is not accurate enough.
[0652] According to an embodiment of the present application, a template of a current block is divided into P sub-templates, and then a first candidate prediction mode is derived based on the P sub-templates and / or the template of the current block, and the first candidate prediction mode is added to a candidate prediction mode list for the i-th prediction mode, thereby improving the accuracy of the candidate prediction mode list for the i-th prediction mode.
[0653] In the embodiment of the present application, the manner of dividing the template of the current block includes, but is not limited to, some of the following:
[0654] Method 1: In the encoding side, the template of the current block is divided according to a first candidate weight derivation mode, specifically, an angle index corresponding to the first candidate weight derivation mode is determined, and the template of the current block is divided into P sub-templates according to the angle index.
[0655] For example, as shown in Figure 18, taking the weight derivation mode in GPM with index 2 as an example, the white area in the weight matrix for the current block means that the weight corresponding to the predicted value in the first prediction mode is 100%, and the black area means that the weight corresponding to the predicted value in the second prediction mode is 100%. As shown in Figure 18, the first prediction mode is associated with the upper template of the current block, and the second prediction mode is associated with the left template of the current block and a portion of the upper template of the current block. However, currently, the entire template is used to derive the prediction mode, which makes the derivation of the prediction mode inaccurate and results in a large prediction error.
[0656] To solve this technical problem, the present application can realize finer division of the template based on the weight derivation mode. For example, as shown in FIG. 18, the present application can identify an angle index corresponding to a first candidate weight derivation mode and identify a division line of the weight matrix corresponding to the first candidate weight derivation mode based on the angle index. Then, the division line is extended to the template region of the current block to divide the template into two sub-templates. For example, the two sub-templates are referred to as a first sub-template and a second sub-template, where the first sub-template corresponds to a first prediction mode and the second sub-template corresponds to a second prediction mode. That is, the first sub-template is used to derive a first candidate prediction mode corresponding to the first prediction mode, and the second sub-template is used to derive a first candidate prediction mode corresponding to the second prediction mode.
[0657] Method 2: In the encoding side, the template of the current block is divided into P sub-templates based on the size of the current block. For example, if the size of the current block is smaller than a certain threshold, the template of the current block is divided into a relatively small number of sub-templates. If the size of the current block is equal to or larger than the threshold, the template of the current block is divided into a relatively large number of sub-templates.
[0658] Method 3: Divide the left template and / or the top template of the current block to obtain P sub-templates. For example, divide the left template of the current block equally into halves, quarters, etc., and / or divide the top template of the current block equally into halves, quarters, etc.
[0659] In addition, on the encoding side, in addition to dividing the template of the current block into P sub-templates using the above methods 1 to 3, it is also possible to divide it using other methods, and the embodiments of the present application are not limited to these.
[0660] On the encoding side, after the template of the current block is divided into P sub-templates, the following step 32 is performed.
[0661] Step 32: Select Q prediction templates from the P sub-templates and / or templates of the current block, where Q is a positive integer less than or equal to P+1.
[0662] In an embodiment of the present application, to improve the accuracy of the first candidate prediction mode, the template of the current block is divided into P sub-templates in step 31, and then Q prediction templates are selected from the P sub-templates and / or the template of the current block. The Q prediction templates are then used to derive the first candidate prediction mode, thereby achieving accurate derivation of the first candidate prediction mode. Finally, the derived first candidate prediction mode is added to the candidate prediction mode list corresponding to the i-th prediction mode.
[0663] In the embodiments of the present application, a prediction template may be understood as a template used to derive a prediction mode, and the prediction template may be the above-mentioned sub-template or a template of the current block.
[0664] In the embodiment of the present application, a specific manner for selecting Q prediction templates from P sub-templates and / or templates of the current block is not limited.
[0665] In some embodiments, Q prediction templates are selected from the P sub-templates and / or the template of the current block based on default conditions, e.g., Q prediction templates are selected from the P sub-templates.
[0666] In some embodiments, steps 32-1 to 32-3 below select Q prediction templates from the P sub-templates and / or templates of the current block.
[0667] Step 32-1: Identify the angle index corresponding to the first candidate weight derivation mode.
[0668] Step 32-2: Identify available neighboring blocks corresponding to the i-th prediction mode based on the angle index.
[0669] Step 32-3: Select Q prediction t...
Claims
1. 1. A video decoding method, comprising: identifying N candidate weight derivation modes, where N is a positive integer; identifying a candidate prediction mode list, the candidate prediction mode list including at least one candidate prediction mode, the at least one candidate prediction mode including a prediction mode identified based on dividing a template of a current block; Identifying a first weight derivation mode and K first prediction modes based on the N candidate weight derivation modes and the candidate prediction mode list, where K is a positive integer greater than 1; predicting the current block based on the first weight derivation mode and the K first prediction modes to obtain a predicted value of the current block; Including, 1. A video decoding method comprising:
2. Identifying the candidate prediction mode list includes: for each first candidate weight derivation mode in the N candidate weight derivation modes, identifying a candidate prediction mode list corresponding to the first candidate weight derivation mode; 2. The method of claim 1 .
3. the first candidate weight derivation mode is any one of the candidate weight derivation modes among the N candidate weight derivation modes; 3. The method of claim 2.
4. When the first candidate weight derivation mode belongs to one type of candidate weight derivation modes among the N candidate weight derivation modes, the method includes: Identifying angle indices corresponding to the N candidate weight derivation modes; Classifying the N candidate weight derivation modes into M types of candidate weight derivation modes based on the angle index, wherein the angle indexes corresponding to the candidate weight derivation modes in the same type of candidate weight derivation modes are the same; Identifying a j-th candidate weight derivation mode among the M candidate weight derivation modes as the first weight derivation mode, where j is a positive integer equal to or less than M; further comprising:
3. The method of claim 2.
5. Identifying a candidate prediction mode list corresponding to the first candidate weight derivation mode includes: identifying a candidate prediction mode list for at least one prediction mode among the K prediction modes corresponding to the first candidate weight derivation mode; 3. The method of claim 2.
6. When the at least one prediction mode corresponds to one candidate prediction mode list, identifying a candidate prediction mode list for at least one prediction mode among the K prediction modes corresponding to the first candidate weight derivation mode includes: identifying a candidate prediction mode list for an i-th prediction mode in the at least one prediction mode, where i is a positive integer; identifying a candidate prediction mode list for the at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode; Including, 6. The method of claim 5.
7. Identifying a candidate prediction mode list for the at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode includes: identifying a candidate prediction mode list for the i-th prediction mode as a candidate prediction mode list for the at least one prediction mode.
7. The method of claim 6.
8. Identifying a candidate prediction mode list for the at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode includes: identifying the candidate prediction mode list for the i-th prediction mode as the candidate prediction mode list for the at least one prediction mode when the candidate prediction mode list for the i-th prediction mode includes a predetermined prediction mode; 7. The method of claim 6.
9. Identifying a candidate prediction mode list for the at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode includes: if the candidate prediction mode list for the i-th prediction mode does not include a preset prediction mode, adding the preset prediction mode to the candidate prediction mode list for the i-th prediction mode to obtain a candidate prediction mode list for the at least one prediction mode.
7. The method of claim 6.
10. When each of the at least one prediction mode corresponds to one candidate prediction mode list, identifying a candidate prediction mode list for at least one prediction mode among the K prediction modes corresponding to the first candidate weight derivation mode includes: for an ith prediction mode among the at least one prediction mode, identifying a candidate prediction mode list for the ith prediction mode, where i is a positive integer.
6. The method of claim 5.
11. the candidate prediction mode list for the i-th prediction mode includes at least one of a first candidate prediction mode identified based on a template of the current block and a second candidate prediction mode identified based on gradients of reconstructed samples in the template.
11. The method according to claim 6 or 10.
12. If the candidate prediction mode list for the i-th prediction mode includes the first candidate prediction mode, the method further comprises: Dividing the template of the current block into P sub-templates, where P is a positive integer greater than 1; selecting Q prediction templates from the P sub-templates and / or the template of the current block, where Q is a positive integer less than or equal to P+1; identifying a prediction mode derived from the Q prediction templates; identifying at least one prediction mode among the prediction modes derived from the Q prediction templates as the first candidate prediction mode; further comprising:
12. The method of claim 11 .
13. Dividing the template of the current block into P sub-templates includes: Identifying an angle index corresponding to the first candidate weight derivation mode; Dividing the template of the current block into P sub-templates based on the angle index; Including, 13. The method of claim 12.
14. Dividing the template of the current block into P sub-templates includes: Dividing the template of the current block into P sub-templates based on a size of the current block.
13. The method of claim 12.
15. Dividing the template of the current block into P sub-templates includes: Dividing a left template and / or a top template of the current block to obtain the P sub-templates.
13. The method of claim 12.
16. Selecting the Q prediction templates from the P sub-templates and / or the template of the current block includes: Identifying an angle index corresponding to the first candidate weight derivation mode; Identifying an available neighboring block corresponding to the i-th prediction mode based on the angle index; selecting the Q prediction templates from the P sub-templates and / or a template of the current block based on available neighboring blocks corresponding to the i-th prediction mode; Including, 13. The method of claim 12.
17. selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on available neighboring blocks corresponding to the i-th prediction mode, When the available neighboring blocks corresponding to the i-th prediction mode include an upper neighboring block of the current block, identifying a sub-template located above the current block among the P sub-templates as a prediction template among the Q prediction templates.
17. The method of claim 16.
18. selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on available neighboring blocks corresponding to the i-th prediction mode, When the available neighboring blocks corresponding to the i-th prediction mode include a left neighboring block of the current block, identifying a sub-template located to the left of the current block among the P sub-templates as a prediction mode in the Q prediction templates.
17. The method of claim 16.
19. selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on available neighboring blocks corresponding to the i-th prediction mode, When the available neighboring blocks corresponding to the i-th prediction mode include a left neighboring block of the current block and an upper neighboring block of the current block, identifying at least one of a subtemplate located to the left of the current block among the P subtemplates, a subtemplate located above the current block among the P subtemplates, and a template of the current block as a prediction template among the Q prediction templates.
17. The method of claim 16.
20. Identifying an angle index corresponding to the first candidate weight derivation mode includes: if a size of the current block is greater than a first threshold, identifying an angle index corresponding to the first candidate weight derivation mode; 17. The method of claim 16.
21. Selecting the Q prediction templates from the P sub-templates and / or the template of the current block includes: selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on a size of the current block; 17. The method of claim 16.
22. selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on a size of the current block, If a size of the current block is equal to or smaller than a second threshold, identifying the P sub-templates and a template of the current block as prediction templates among the Q prediction templates.
22. The method of claim 21 .
23. selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on a size of the current block, If a size of the current block is less than or equal to a second threshold, identifying a template of the current block as a prediction template among the Q prediction templates.
22. The method of claim 21 .
24. Identifying a prediction mode derived from the Q prediction templates includes: identifying R preliminary prediction modes for any of the Q prediction templates; determining a first cost for predicting the prediction template using the R preliminary prediction modes; obtaining a prediction mode derived from the prediction template based on a first cost corresponding to each of the R preliminary prediction modes; Including, 13. The method of claim 12.
25. Identifying at least one prediction mode among prediction modes derived from the Q prediction templates as the first candidate prediction mode includes: identifying a prediction mode derived from the Q prediction templates as the first candidate prediction mode.
13. The method of claim 12.
26. If the candidate prediction mode list for the i-th prediction mode includes a second candidate prediction mode, the method further comprises: Dividing a reconstructed sample region in which the template of the current block is located into S reconstructed sample sub-regions, where S is a positive integer; selecting G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region, where G is a positive integer less than or equal to S+1; identifying a prediction mode derived from the G reconstructed sample prediction regions; identifying at least one prediction mode from the prediction modes derived from the G reconstructed sample prediction regions as the second candidate prediction mode; further comprising:
12. The method of claim 11 .
27. Dividing the reconstructed sample region in which the template of the current block is located into S reconstructed sample sub-regions includes: Identifying an angle index corresponding to the first candidate weight derivation mode; dividing the reconstruction sample area into S reconstruction sample sub-areas based on the angular index; 27. The method of claim 26.
28. Dividing the reconstructed sample region in which the template of the current block is located into S reconstructed sample sub-regions includes: dividing the reconstructed sample region into S reconstructed sample sub-regions based on a size of the current block; 27. The method of claim 26.
29. Dividing the reconstructed sample region in which the template of the current block is located into S reconstructed sample sub-regions includes: Dividing a left-side adjacent reconstructed sample area and / or an upper-side adjacent reconstructed sample area of the current block to obtain the S reconstructed sample sub-areas.
27. The method of claim 26.
30. Selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region comprises: Identifying an angle index corresponding to the first candidate weight derivation mode; Identifying an available neighboring block corresponding to the i-th prediction mode based on the angle index; selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on available neighboring blocks corresponding to the i-th prediction mode; Including, 27. The method of claim 26.
31. selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on available neighboring blocks corresponding to the i-th prediction mode, If the available neighboring blocks corresponding to the i-th prediction mode include an upper neighboring block of the current block, identifying a reconstructed sample sub-area located above the current block in the S reconstructed sample sub-areas as a reconstructed sample prediction area in the G reconstructed sample prediction areas.
31. The method of claim 30.
32. selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on available neighboring blocks corresponding to the i-th prediction mode, If the available neighboring blocks corresponding to the i-th prediction mode include a left neighboring block of the current block, identifying a reconstructed sample sub-area located to the left of the current block in the S reconstructed sample sub-areas as a reconstructed sample prediction area in the G reconstructed sample prediction areas.
31. The method of claim 30.
33. selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on available neighboring blocks corresponding to the i-th prediction mode, When the available neighboring blocks corresponding to the i-th prediction mode include a left neighboring block of the current block and an upper neighboring block of the current block, identifying at least one of a reconstructed sample sub-area located to the left of the current block in the S reconstructed sample sub-areas, a reconstructed sample sub-area located above the current block in the S reconstructed sample sub-areas, and the reconstructed sample area as a reconstructed sample prediction area in the G reconstructed sample prediction areas.
31. The method of claim 30.
34. Identifying an angle index corresponding to the first candidate weight derivation mode includes: if the size of the current block is greater than a third threshold, identifying an angle index corresponding to the first candidate weight derivation mode.
31. The method of claim 30.
35. Selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region comprises: selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on a size of the current block; 27. The method of claim 26.
36. selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on a size of the current block, If the size of the current block is equal to or smaller than a fourth threshold, identifying the S reconstructed sample sub-regions and the reconstructed sample region as a reconstructed sample prediction region in the G reconstructed sample prediction regions.
36. The method of claim 35.
37. selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on a size of the current block, If the size of the current block is equal to or smaller than a fourth threshold, identifying the reconstructed sample region as a reconstructed sample prediction region among the G reconstructed sample prediction regions.
36. The method of claim 35.
38. Identifying a prediction mode derived from the G reconstructed sample prediction regions includes: determining, for one of the G reconstructed sample prediction regions, a gradient of each central sample of the reconstructed sample prediction region; determining a prediction mode derived from the reconstructed sample prediction region based on a gradient of each central sample of the reconstructed sample prediction region; Including, 27. The method of claim 26.
39. Identifying at least one prediction mode from the prediction modes derived from the G reconstructed sample prediction regions as the second candidate prediction mode includes: identifying a prediction mode derived from the G reconstructed sample prediction regions as the second candidate prediction mode.
27. The method of claim 26.
40. The candidate prediction mode list for the i-th prediction mode further includes at least one of a third candidate prediction mode corresponding to the first candidate weight derivation mode, a prediction mode for a neighboring block of the current block, and a preset prediction mode.
12. The method of claim 11 .
41. the number of candidate prediction modes included in the candidate prediction mode list for the i-th prediction mode is a preset value; 41. The method of claim 40.
42. Identifying a candidate prediction mode list for the i-th prediction mode includes: selecting, in a predetermined order, prediction modes whose number is the predetermined value from the first candidate prediction mode, the second candidate prediction mode, a third candidate prediction mode corresponding to the first candidate weight derivation mode, prediction modes for neighboring blocks of the current block, and the predetermined prediction modes, to configure a candidate prediction mode list for the i-th prediction mode.
42. The method of claim 41 .
43. the third candidate prediction mode includes at least one of a prediction mode in which a prediction angle is parallel to a division line of the first candidate weight derivation mode and a prediction mode in which a prediction angle is perpendicular to the division line of the first candidate weight derivation mode, the preset prediction mode includes a planar mode, and the preset order includes a prediction mode in which the prediction angle is parallel to a division line of the first candidate weight derivation mode, the first candidate prediction mode, the second candidate prediction mode, a prediction mode for a neighboring block of the current block, a prediction mode in which the prediction angle is perpendicular to the division line of the first candidate weight derivation mode, and a planar mode.
43. The method of claim 42.
44. Identifying the first weight derivation mode and the K first prediction modes based on the N candidate weight derivation modes and the candidate prediction mode list includes: decoding a bitstream to obtain a first index, the first index being used to indicate a first combination, the first combination including the first weight derivation mode and the K first prediction modes; identifying a candidate combination list based on the N candidate weight derivation modes and the candidate prediction mode list, the candidate combination list including at least one candidate combination, the candidate combination including one weight derivation mode and K prediction modes; Identifying the first combination from the candidate combination list based on the first index; Including, The method according to any one of claims 1 to 10, 12 to 43.
45. Identifying the candidate combination list based on the N candidate weight derivation modes and the candidate prediction mode list includes: obtaining T second combinations based on the N candidate weight derivation modes and the candidate prediction mode list, wherein any one of the T second combinations includes one weight derivation mode and K prediction modes, and the weight derivation mode and the K prediction modes included in any one second combination of the T second combinations are not completely the same as the weight derivation mode and the K prediction modes included in any other second combination of the T second combinations, and the T is a positive integer greater than 1; obtaining the candidate combination list based on the T second combinations; Including, 45. The method of claim 44.
46. Obtaining the candidate combination list based on the T second combinations includes: determining a cost corresponding to any one of the T second combinations when predicting a template of the current block using a weight derivation mode in the second combination and K prediction modes; and identifying the candidate combination list based on a cost corresponding to each second combination in the T second combinations; Including, 46. The method of claim 45.
47. The height of the top template of the current block is 1, and / or the width of the left template of the current block is 1; 47. The method of any one of claims 1 to 10, 12 to 43, and 45 to 46.
48. 1. A video encoding method, comprising: identifying N candidate weight derivation modes, where N is a positive integer; identifying a candidate prediction mode list, the candidate prediction mode list including at least one candidate prediction mode, the at least one candidate prediction mode including a prediction mode identified based on dividing a template of a current block; Identifying a first weight derivation mode and K first prediction modes based on the N candidate weight derivation modes and the candidate prediction mode list, where K is a positive integer greater than 1; predicting the current block based on the first weight derivation mode and the K first prediction modes to obtain a predicted value of the current block; Including, A video encoding method comprising:
49. Identifying the candidate prediction mode list includes: for each first candidate weight derivation mode in the N candidate weight derivation modes, identifying a candidate prediction mode list corresponding to the first candidate weight derivation mode; 49. The method of claim 48.
50. the first candidate weight derivation mode is any one of the candidate weight derivation modes among the N candidate weight derivation modes; 50. The method of claim 49.
51. When the first candidate weight derivation mode belongs to one type of candidate weight derivation modes among the N candidate weight derivation modes, the method includes: Identifying angle indices corresponding to the N candidate weight derivation modes; Classifying the N candidate weight derivation modes into M types of candidate weight derivation modes based on the angle index, wherein the angle indexes corresponding to the candidate weight derivation modes in the same type of candidate weight derivation modes are the same; Identifying a j-th candidate weight derivation mode among the M candidate weight derivation modes as the first weight derivation mode, where j is a positive integer equal to or less than M; further comprising:
50. The method of claim 49.
52. Identifying a candidate prediction mode list corresponding to the first candidate weight derivation mode includes: identifying a candidate prediction mode list for at least one prediction mode among the K prediction modes corresponding to the first candidate weight derivation mode; 50. The method of claim 49.
53. When the at least one prediction mode corresponds to one candidate prediction mode list, identifying a candidate prediction mode list for at least one prediction mode among the K prediction modes corresponding to the first candidate weight derivation mode includes: identifying a candidate prediction mode list for an i-th prediction mode in the at least one prediction mode, where i is a positive integer; identifying a candidate prediction mode list for the at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode; Including, 53. The method of claim 52.
54. Identifying a candidate prediction mode list for the at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode includes: identifying a candidate prediction mode list for the i-th prediction mode as a candidate prediction mode list for the at least one prediction mode.
54. The method of claim 53.
55. Identifying a candidate prediction mode list for the at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode includes: identifying the candidate prediction mode list for the i-th prediction mode as the candidate prediction mode list for the at least one prediction mode when the candidate prediction mode list for the i-th prediction mode includes a predetermined prediction mode; 54. The method of claim 53.
56. Identifying a candidate prediction mode list for the at least one prediction mode based on the candidate prediction mode list for the i-th prediction mode includes: if the candidate prediction mode list for the i-th prediction mode does not include a preset prediction mode, adding the preset prediction mode to the candidate prediction mode list for the i-th prediction mode to obtain a candidate prediction mode list for the at least one prediction mode.
54. The method of claim 53.
57. When each of the at least one prediction mode corresponds to one candidate prediction mode list, identifying a candidate prediction mode list for at least one prediction mode among the K prediction modes corresponding to the first candidate weight derivation mode includes: for an ith prediction mode among the at least one prediction mode, identifying a candidate prediction mode list for the ith prediction mode, where i is a positive integer.
53. The method of claim 52.
58. the candidate prediction mode list for the i-th prediction mode includes at least one of a first candidate prediction mode identified based on a template of the current block and a second candidate prediction mode identified based on gradients of reconstructed samples in the template.
58. The method of claim 53 or 57.
59. If the candidate prediction mode list for the i-th prediction mode includes the first candidate prediction mode, the method further comprises: Dividing the template of the current block into P sub-templates, where P is a positive integer greater than 1; selecting Q prediction templates from the P sub-templates and / or the template of the current block, where Q is a positive integer less than or equal to P+1; identifying a prediction mode derived from the Q prediction templates; identifying at least one prediction mode among the prediction modes derived from the Q prediction templates as the first candidate prediction mode; further comprising:
59. The method of claim 58.
60. Dividing the template of the current block into P sub-templates includes: Identifying an angle index corresponding to the first candidate weight derivation mode; Dividing the template of the current block into P sub-templates based on the angle index; Including, 60. The method of claim 59.
61. Dividing the template of the current block into P sub-templates includes: Dividing the template of the current block into P sub-templates based on a size of the current block.
60. The method of claim 59.
62. Dividing the template of the current block into P sub-templates includes: Dividing a left template and / or a top template of the current block to obtain the P sub-templates.
60. The method of claim 59.
63. Selecting the Q prediction templates from the P sub-templates and / or the template of the current block includes: Identifying an angle index corresponding to the first candidate weight derivation mode; Identifying an available neighboring block corresponding to the i-th prediction mode based on the angle index; selecting the Q prediction templates from the P sub-templates and / or a template of the current block based on available neighboring blocks corresponding to the i-th prediction mode; Including, 60. The method of claim 59.
64. selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on available neighboring blocks corresponding to the i-th prediction mode, When the available neighboring blocks corresponding to the i-th prediction mode include an upper neighboring block of the current block, identifying a sub-template located above the current block among the P sub-templates as a prediction template among the Q prediction templates.
64. The method of claim 63.
65. selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on available neighboring blocks corresponding to the i-th prediction mode, When the available neighboring blocks corresponding to the i-th prediction mode include a left neighboring block of the current block, identifying a sub-template located to the left of the current block among the P sub-templates as a prediction mode in the Q prediction templates.
64. The method of claim 63.
66. selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on available neighboring blocks corresponding to the i-th prediction mode, When the available neighboring blocks corresponding to the i-th prediction mode include a left neighboring block of the current block and an upper neighboring block of the current block, identifying at least one of a subtemplate located to the left of the current block among the P subtemplates, a subtemplate located above the current block among the P subtemplates, and a template of the current block as a prediction template among the Q prediction templates.
64. The method of claim 63.
67. Identifying an angle index corresponding to the first candidate weight derivation mode includes: if a size of the current block is greater than a first threshold, identifying an angle index corresponding to the first candidate weight derivation mode; 64. The method of claim 63.
68. Selecting the Q prediction templates from the P sub-templates and / or the template of the current block includes: selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on a size of the current block; 64. The method of claim 63.
69. selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on a size of the current block, If a size of the current block is equal to or smaller than a second threshold, identifying the P sub-templates and a template of the current block as prediction templates among the Q prediction templates.
69. The method of claim 68.
70. selecting the Q prediction templates from the P sub-templates and / or the template of the current block based on a size of the current block, If a size of the current block is less than or equal to a second threshold, identifying a template of the current block as a prediction template among the Q prediction templates.
69. The method of claim 68.
71. Identifying a prediction mode derived from the Q prediction templates includes: identifying R preliminary prediction modes for any of the Q prediction templates; determining a first cost for predicting the prediction template using the R preliminary prediction modes; obtaining a prediction mode derived from the prediction template based on a first cost corresponding to each of the R preliminary prediction modes; Including, 60. The method of claim 59.
72. Identifying at least one prediction mode among prediction modes derived from the Q prediction templates as the first candidate prediction mode includes: identifying a prediction mode derived from the Q prediction templates as the first candidate prediction mode.
60. The method of claim 59.
73. If the candidate prediction mode list for the i-th prediction mode includes a second candidate prediction mode, the method further comprises: Dividing a reconstructed sample region in which the template of the current block is located into S reconstructed sample sub-regions, where S is a positive integer; selecting G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region, where G is a positive integer less than or equal to S+1; identifying a prediction mode derived from the G reconstructed sample prediction regions; identifying at least one prediction mode from the prediction modes derived from the G reconstructed sample prediction regions as the second candidate prediction mode; further comprising:
59. The method of claim 58.
74. Dividing the reconstructed sample region in which the template of the current block is located into S reconstructed sample sub-regions includes: Identifying an angle index corresponding to the first candidate weight derivation mode; dividing the reconstruction sample area into S reconstruction sample sub-areas based on the angular index; 74. The method of claim 73.
75. Dividing the reconstructed sample region in which the template of the current block is located into S reconstructed sample sub-regions includes: dividing the reconstructed sample region into S reconstructed sample sub-regions based on a size of the current block; 74. The method of claim 73.
76. Dividing the reconstructed sample region in which the template of the current block is located into S reconstructed sample sub-regions includes: Dividing a left-side adjacent reconstructed sample area and / or an upper-side adjacent reconstructed sample area of the current block to obtain the S reconstructed sample sub-areas.
74. The method of claim 73.
77. Selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region comprises: Identifying an angle index corresponding to the first candidate weight derivation mode; Identifying an available neighboring block corresponding to the i-th prediction mode based on the angle index; selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on available neighboring blocks corresponding to the i-th prediction mode; Including, 74. The method of claim 73.
78. selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on available neighboring blocks corresponding to the i-th prediction mode, If the available neighboring blocks corresponding to the i-th prediction mode include an upper neighboring block of the current block, identifying a reconstructed sample sub-area located above the current block in the S reconstructed sample sub-areas as a reconstructed sample prediction area in the G reconstructed sample prediction areas.
78. The method of claim 77.
79. selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on available neighboring blocks corresponding to the i-th prediction mode, If the available neighboring blocks corresponding to the i-th prediction mode include a left neighboring block of the current block, identifying a reconstructed sample sub-area located to the left of the current block in the S reconstructed sample sub-areas as a reconstructed sample prediction area in the G reconstructed sample prediction areas.
78. The method of claim 77.
80. selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on available neighboring blocks corresponding to the i-th prediction mode, When the available neighboring blocks corresponding to the i-th prediction mode include a left neighboring block of the current block and an upper neighboring block of the current block, identifying at least one of a reconstructed sample sub-area located to the left of the current block in the S reconstructed sample sub-areas, a reconstructed sample sub-area located above the current block in the S reconstructed sample sub-areas, and the reconstructed sample area as a reconstructed sample prediction area in the G reconstructed sample prediction areas.
78. The method of claim 77.
81. Identifying an angle index corresponding to the first candidate weight derivation mode includes: if the size of the current block is greater than a third threshold, identifying an angle index corresponding to the first candidate weight derivation mode.
78. The method of claim 77.
82. Selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region comprises: selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on a size of the current block; 74. The method of claim 73.
83. selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on a size of the current block, If the size of the current block is equal to or smaller than a fourth threshold, identifying the S reconstructed sample sub-regions and the reconstructed sample region as a reconstructed sample prediction region in the G reconstructed sample prediction regions.
83. The method of claim 82.
84. selecting the G reconstructed sample prediction regions from the S reconstructed sample sub-regions and / or the reconstructed sample region based on a size of the current block, If the size of the current block is equal to or smaller than a fourth threshold, identifying the reconstructed sample region as a reconstructed sample prediction region among the G reconstructed sample prediction regions.
83. The method of claim 82.
85. Identifying a prediction mode derived from the G reconstructed sample prediction regions includes: determining, for one of the G reconstructed sample prediction regions, a gradient of each central sample of the reconstructed sample prediction region; determining a prediction mode derived from the reconstructed sample prediction region based on a gradient of each central sample of the reconstructed sample prediction region; Including, 74. The method of claim 73.
86. Identifying at least one prediction mode from the prediction modes derived from the G reconstructed sample prediction regions as the second candidate prediction mode includes: identifying a prediction mode derived from the G reconstructed sample prediction regions as the second candidate prediction mode.
74. The method of claim 73.
87. The candidate prediction mode list for the i-th prediction mode further includes at least one of a third candidate prediction mode corresponding to the first candidate weight derivation mode, a prediction mode for a neighboring block of the current block, and a preset prediction mode.
59. The method of claim 58.
88. the number of candidate prediction modes included in the candidate prediction mode list for the i-th prediction mode is a preset value; 88. The method of claim 87.
89. Identifying a candidate prediction mode list for the i-th prediction mode includes: selecting, in a predetermined order, prediction modes whose number is the predetermined value from the first candidate prediction mode, the second candidate prediction mode, a third candidate prediction mode corresponding to the first candidate weight derivation mode, prediction modes for neighboring blocks of the current block, and the predetermined prediction modes, to configure a candidate prediction mode list for the i-th prediction mode.
89. The method of claim 88.
90. the third candidate prediction mode includes at least one of a prediction mode in which a prediction angle is parallel to a division line of the first candidate weight derivation mode and a prediction mode in which a prediction angle is perpendicular to the division line of the first candidate weight derivation mode, the preset prediction mode includes a planar mode, and the preset order includes a prediction mode in which the prediction angle is parallel to a division line of the first candidate weight derivation mode, the first candidate prediction mode, the second candidate prediction mode, a prediction mode for a neighboring block of the current block, a prediction mode in which the prediction angle is perpendicular to the division line of the first candidate weight derivation mode, and a planar mode.
90. The method of claim 89.
91. Identifying the first weight derivation mode and the K first prediction modes based on the N candidate weight derivation modes and the candidate prediction mode list includes: identifying a candidate combination list based on the N candidate weight derivation modes and the candidate prediction mode list, the candidate combination list including at least one candidate combination, the candidate combination including one weight derivation mode and K prediction modes; identifying a first combination from the list of candidate combinations; Including, The method according to any one of claims 48 to 57, 59 to 90.
92. Identifying the candidate combination list based on the N candidate weight derivation modes and the candidate prediction mode list includes: obtaining T second combinations based on the N candidate weight derivation modes and the candidate prediction mode list, wherein any one of the T second combinations includes one weight derivation mode and K prediction modes, and the weight derivation mode and the K prediction modes included in any one second combination of the T second combinations are not completely the same as the weight derivation mode and the K prediction modes included in any other second combination of the T second combinations, and the T is a positive integer greater than 1; obtaining the candidate combination list based on the T second combinations; Including, 92. The method of claim 91 .
93. Obtaining the candidate combination list based on the T second combinations includes: determining a cost corresponding to any one of the T second combinations when predicting a template of the current block using a weight derivation mode in the second combination and K prediction modes; and identifying the candidate combination list based on a cost corresponding to each second combination in the T second combinations; Including, 93. The method of claim 92.
94. The method further includes signaling the first index into a bitstream; the first index is used to indicate a first combination, and the first combination includes the first weight derivation mode and the K first prediction modes; 93. The method of claim 92.
95. The height of the top template of the current block is 1, and / or the width of the left template of the current block is 1; 95. The method of any one of claims 48 to 57, 59 to 90, and 92 to 94.
96. A video decoding device, comprising: The video decoding apparatus includes a weight derivation mode specifying unit, a prediction list specifying unit, a processing unit, and a prediction unit; the weight derivation mode identification unit is configured to identify N candidate weight derivation modes, where N is a positive integer; the prediction list identification unit is configured to identify a candidate prediction mode list, the candidate prediction mode list including at least one candidate prediction mode, the at least one candidate prediction mode including a prediction mode identified based on dividing a template of a current block; the processing unit is configured to identify a first weight derivation mode and K first prediction modes based on the N candidate weight derivation modes and the candidate prediction mode list, where K is a positive integer greater than 1; the prediction unit is configured to predict the current block based on the first weight derivation mode and the K first prediction modes to obtain a predicted value of the current block. A video decoding device comprising:
97. 1. A video encoding device, comprising: The video encoding apparatus includes a weight derivation mode unit, a prediction list identification unit, a processing unit, and a prediction unit; the weight derivation mode unit is configured to identify N candidate weight derivation modes, where N is a positive integer; the prediction list identification unit is configured to identify a candidate prediction mode list, the candidate prediction mode list including at least one candidate prediction mode, the at least one candidate prediction mode including a prediction mode identified based on dividing a template of a current block; the processing unit is configured to identify a first weight derivation mode and K first prediction modes based on the N candidate weight derivation modes and the candidate prediction mode list, where K is a positive integer greater than 1; the prediction unit is configured to predict the current block based on the first weight derivation mode and the K first prediction modes to obtain a predicted value of the current block. A video encoding device comprising:
98. 1. An electronic device comprising: the electronic device includes a processor and a memory; the memory is configured to store a computer program; The processor is configured to execute the method according to any one of claims 1 to 47 or any one of claims 48 to 95 by calling and executing a computer program stored in the memory. An electronic device characterized by:
99. 1. A video coding system, comprising: the video coding system includes a video encoder and a video decoder; said video decoder being configured to perform a method according to any one of claims 1 to 47; The video encoder is configured to perform a method according to any one of claims 48 to 95. A video coding system comprising:
100. 1. A computer-readable storage medium, comprising: the computer-readable storage medium is configured to store a computer program; The computer program causes a computer to carry out the method according to any one of claims 1 to 47 or the method according to any one of claims 48 to 95. A computer-readable storage medium comprising:
Citation Information
Patent Citations
Method, apparatus, and non-transitory computer-readable storage medium for motion vector refinement for geometric partition mode
WO2022218256A1
Template matching refinement in inter-prediction modes
WO2022221013A1
Geometric partition mode in video coding
WO2023200643A2
Spatial geometric partition mode
WO2024008611A1