Video encoding / decoding method and its device, equipment, system, and storage medium

The use of multidirectional prediction modes in video encoding and decoding methods addresses inaccuracies in predicting current blocks, resulting in improved video compression efficiency.

JP2026501854APending Publication Date: 2026-01-16GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025541637
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Current video encoding and decoding methods face inaccuracies in predicting current blocks using multiple prediction modes, leading to inefficient video compression.

Method used

Implementing a video encoding and decoding method that utilizes multidirectional prediction modes, specifically bidirectional prediction, to enhance the accuracy of predicting current blocks.

Benefits of technology

Improves the prediction accuracy of current blocks, thereby enhancing the efficiency of video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026501854000001_ABST
    Figure 2026501854000001_ABST
Patent Text Reader

Abstract

The present application provides a video encoding / decoding method, and an apparatus, device, system, and storage medium thereof, in which, when encoding / decoding a current block, K prediction modes for the current block are determined, and at least one prediction mode among the K prediction modes is a multi-directional prediction mode (e.g., a bidirectional prediction mode). In this way, by predicting the current block using the K prediction modes, the prediction accuracy of the current block can be improved, and the effect of video encoding / decoding can be further improved.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to the technical field of video encoding and decoding, and in particular to a video encoding and decoding method and an apparatus, device, system, and storage medium therefor. [Background technology]

[0002] Digital video technology can be incorporated into various video devices such as digital televisions, smartphones, computers, e-readers, video players, etc. As video technology develops, the amount of data contained in video data increases. To facilitate the transmission of video data, video devices implement video compression techniques to allow the video data to be transmitted or stored more efficiently.

[0003] Since video has temporal or spatial redundancy, prediction can eliminate or reduce the redundancy in the video and improve compression efficiency. Currently, in order to improve prediction efficiency, a current block can be predicted using multiple prediction modes. However, when predicting a current block using multiple prediction modes, there is currently a problem that the prediction is inaccurate. Summary of the Invention [Means for solving the problem]

[0004] An embodiment of the present application provides a video encoding / decoding method, an apparatus, a device, a system, and a storage medium thereof, which can improve the prediction accuracy of a current block by setting at least one prediction mode of the current block to a multidirectional prediction mode (e.g., a bidirectional prediction mode).

[0005] In a first aspect, the present application provides a video decoding method applied to a decoder, the video decoding method comprising: determining K prediction modes for a current block, where at least one prediction mode among the K prediction modes is an N-directional prediction mode, and both K and N are positive integers greater than 1; predicting the current block based on the K prediction modes to obtain a predicted value of the current block.

[0006] In a second aspect, embodiments of the present application provide a video encoding method, the video encoding method comprising: determining K prediction modes for a current block, where at least one prediction mode among the K prediction modes is an N-directional prediction mode, and both K and N are positive integers greater than 1; predicting the current block based on the K prediction modes to obtain a predicted value of the current block.

[0007] In a third aspect, the present application provides a video decoding device for performing the method of the first aspect or each implementation thereof, specifically the device comprising functional units for performing the method of the first aspect or each implementation thereof.

[0008] In a fourth aspect, the present application provides a video encoding device for performing the method of the second aspect or each implementation thereof, specifically the device comprising functional units for performing the method of the second aspect or each implementation thereof.

[0009] In a fifth aspect, there is provided a video decoder comprising a processor and a memory, the memory being adapted to store a computer program, and the processor being adapted to call and execute the computer program stored in the memory to perform the method of the first aspect above or each implementation thereof.

[0010] In a sixth aspect, there is provided a video encoder comprising a processor and a memory, the memory being adapted to store a computer program, the processor being adapted to execute the method of the second aspect or each implementation thereof by calling and executing the computer program stored in the memory.

[0011] In a seventh aspect, there is provided a video encoding / decoding system including a video encoder and a video decoder, wherein the video decoder is used to perform the method of the first aspect or each of its implementations, and the video encoder is used to perform the method of the second aspect or each of its implementations.

[0012] In an eighth aspect, there is provided a chip for implementing the method according to any one of the first and second aspects described above, or each of their respective implementations. Specifically, the chip includes a processor, and the processor is used to cause a device including the chip to execute the method according to any one of the first and second aspects described above, or each of their respective implementations, by calling up and executing a computer program from a memory.

[0013] In a ninth aspect, there is provided a computer-readable storage medium for storing a computer program, the computer program causing a computer to perform a method according to any one of the first and second aspects described above, or each implementation thereof.

[0014] In a tenth aspect, there is provided a computer program product comprising computer program instructions to cause a computer to perform the method of any one of the first and second aspects above, or each implementation thereof.

[0015] In an eleventh aspect, there is provided a computer program which, when run on a computer, causes the computer to carry out the method of any one of the first and second aspects described above, or each implementation thereof.

[0016] In a twelfth aspect, a bitstream is provided, the bitstream being generated based on the method of the second aspect described above, and optionally the bitstream including a first index, the first index being used to indicate a first combination of one weight derivation mode and K prediction modes, where K is a positive integer greater than 1.

[0017] According to the above technical solution, when encoding and decoding a current block, K prediction modes of the current block are determined, and at least one prediction mode among the K prediction modes is a multi-directional prediction mode (e.g., a bidirectional prediction mode). In this way, when predicting the current block using the K prediction modes, the prediction accuracy of the current block can be improved, and thus the video encoding and decoding effect can be improved. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a schematic block diagram of a video encoding / decoding system according to an embodiment of the present application. [Figure 2] 1 is a schematic block diagram of a video encoder according to an embodiment of the present application; [Figure 3] 1 is a schematic block diagram of a video decoder according to an embodiment of the present application; [Figure 4] FIG. 1 is a schematic diagram of weight assignment. [Figure 5] FIG. 1 is a schematic diagram of weight assignment. [Figure 6A] FIG. 1 is a schematic diagram of inter prediction. [Figure 6B] FIG. 1 is a schematic diagram of weighted inter prediction. [Figure 6C] FIG. 2 is a schematic diagram of a transition region. [Figure 6D] FIG. 10 is a schematic diagram of another transition region. [Figure 7A] FIG. 1 is a schematic diagram of intra prediction. [Figure 7B] FIG. 1 is a schematic diagram of intra prediction. [Figure 8A]FIG. 1 is a schematic diagram of intra prediction. [Figure 8B] FIG. 1 is a schematic diagram of intra prediction. [Figure 8C] FIG. 1 is a schematic diagram of intra prediction. [Figure 8D] FIG. 1 is a schematic diagram of intra prediction. [Figure 8E] FIG. 1 is a schematic diagram of intra prediction. [Figure 8F] FIG. 1 is a schematic diagram of intra prediction. [Figure 8G] FIG. 1 is a schematic diagram of intra prediction. [Figure 8H] FIG. 1 is a schematic diagram of intra prediction. [Figure 8I] FIG. 1 is a schematic diagram of intra prediction. [Figure 9] FIG. 1 is a schematic diagram of intra-prediction modes. [Figure 10] FIG. 1 is a schematic diagram of intra-prediction modes. [Figure 11] FIG. 1 is a schematic diagram of intra-prediction modes. [Figure 12] FIG. 1 is a schematic diagram of an MIP. [Figure 13] FIG. 1 is a schematic diagram of TIMD prediction. [Figure 14A] 1 is a bar graph corresponding to DIMD. [Figure 14B] Schematic diagram of DIMD prediction. [Figure 15] FIG. 1 is a schematic diagram of combinatorial prediction. [Figure 16A] FIG. 1 is a schematic diagram of a template. [Figure 16B] FIG. 1 is a schematic diagram of a template. [Figure 16C] FIG. 10 is a schematic diagram of template weight derivation. [Figure 16D] FIG. 2 is a schematic diagram of adjacent blocks. [Figure 16E] FIG. 1 is a schematic diagram of weight assignment. [Figure 16F] FIG. 1 is a schematic diagram of weight assignment. [Figure 16G] FIG. 1 is a schematic diagram of template division. [Figure 16H] FIG. 10 is a schematic diagram of another template division. [Figure 17A] FIG. 1 is a schematic diagram of MMVD. [Figure 17B] FIG. 1 is a schematic diagram of MMVD. [Figure 18] 1 is a schematic flow chart diagram of a video decoding method according to an embodiment of the present application; [Figure 19] FIG. 1 is a schematic diagram of a matching according to an embodiment of the present application. [Figure 20] FIG. 10 is another matching schematic diagram according to an embodiment of the present application. [Figure 21] FIG. 10 is another matching schematic diagram according to an embodiment of the present application. [Figure 22] 1 is a schematic flow chart diagram of a video decoding method according to an embodiment of the present application; [Figure 23] 1 is a schematic block diagram of a video decoding device according to an embodiment of the present application; [Figure 24] 1 is a schematic block diagram of a video encoding device according to an embodiment of the present application; [Figure 25] 1 is a schematic block diagram of an electronic device according to an embodiment of the present application; [Figure 26] 1 is a schematic block diagram of a video encoding / decoding system according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0019] The present application can be applied to image encoding / decoding fields, video encoding / decoding fields, hardware video encoding / decoding fields, dedicated circuit video encoding / decoding fields, real-time video encoding / decoding fields, etc. For example, the present application's solution can be incorporated into audio video coding standards (AVS) such as the H.264 / audio video coding (AVC) standard, the H.265 / high efficiency video coding (HEVC) standard, and the H.266 / versatile video coding (VVC) standard. Alternatively, the present solution may be incorporated into and operate with other proprietary or industry standards, including ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, and ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), including Scalable Video Coding and Decoding (SVC) and Multiview Video Coding and Decoding (MVC) extensions. It should be understood that the present technology is not limited to any particular coding and decoding standard or technology.

[0020] For ease of understanding, a video encoding / decoding system according to an embodiment of the present invention will be described first with reference to FIG.

[0021] FIG. 1 is a schematic block diagram of a video encoding / decoding system according to an embodiment of the present application. Note that FIG. 1 is merely an example, and the video encoding / decoding system according to the embodiment of the present application includes, but is not limited to, the one shown in FIG. 1. As shown in FIG. 1, the video encoding / decoding system 100 includes an encoding device 110 and a decoding device 120. Here, the encoding device is used to encode (which can be understood as compressing) video data to generate a bitstream and transmit the bitstream to a decoding device. The decoding device decodes the bitstream generated by the encoding device to obtain decoded video data.

[0022] The encoding device 110 of the present embodiment can be understood as a device having a video encoding function, and the decoding device 120 can be understood as a device having a video decoding function, that is, the encoding device 110 and the decoding device 120 of the present embodiment include a broader range of devices, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, and in-vehicle computers.

[0023] In some embodiments, encoding device 110 may transmit encoded video data (e.g., a bitstream) to decoding device 120 over channel 130. Channel 130 may include one or more media and / or devices that may transmit encoded video data from encoding device 110 to decoding device 120.

[0024] In one example, channel 130 includes one or more communication media that enable encoding device 110 to transmit encoded video data directly in real time to decoding device 120. In this example, encoding device 110 may modulate the encoded video data in accordance with a communication standard and transmit the modulated video data to decoding device 120. Here, the communication media may include wireless communication media, such as a radio frequency spectrum, and optionally, the communication media may also include wired communication media, such as one or more physical transmission lines.

[0025] In another example, channel 130 includes a storage medium that can store the video data after it has been encoded by encoding device 110. The storage medium can include various locally accessible data storage media, such as optical disks, DVDs, flash memory, etc. In this example, decoding device 120 can obtain the encoded video data from the storage medium.

[0026] In another example, channel 130 may include a storage server that may store video data after it has been encoded by encoding device 110. In this example, decoding device 120 may download the stored encoded video data from the storage server. Optionally, the storage server may store the encoded video data and transmit the encoded video data to decoding device 120 from a web server (e.g., for a website), a File Transfer Protocol (FTP) server, or the like.

[0027] In some embodiments, encoding device 110 includes a video encoder 112 and an output interface 113, where output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.

[0028] In some embodiments, the encoding device 110 includes a video encoder 112 and output In addition to the interface 113, a video source 111 may also be included.

[0029] The video source 111 may include at least one of a video capture device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system, where the video input interface is used to receive video data from a video content provider and the computer graphics system is used to generate the video data.

[0030] The video encoder 112 encodes video data from the video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream includes coding information for a picture or a sequence of pictures in the form of a bitstream. The coding information may include coded image data and associated data. The associated data may include sequence parameter sets (SPSs), picture parameter sets (PPSs), and other syntax structures. An SPS may include parameters that apply to one or more sequences. A PPS may include parameters that apply to one or more pictures. A syntax structure refers to a set of zero or more syntax elements arranged in a specified order within the bitstream.

[0031] Video encoder 112 transmits the encoded video data directly to decoding device 120 via output interface 113. The encoded video data may be stored on a storage medium or storage server for subsequent retrieval by decoding device 120.

[0032] In some embodiments, decoding device 120 includes an input interface 121 and a video decoder 122 .

[0033] In some embodiments, decoding device 120 may further include a display device 123 in addition to input interface 121 and video decoder 122 .

[0034] Here, the input interface 121 includes a receiver and / or a modem, and can receive encoded video data via a channel 130.

[0035] The video decoder 122 is used to decode the encoded video data to obtain decoded video data, and transmit the decoded video data to a display device 123 .

[0036] Display device 123 displays the decoded video data and may be integrated with or external to decoding device 120. Display device 123 may include multiple display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.

[0037] It should be noted that Fig. 1 is merely an example, and the technical solutions of the embodiments of the present application are not limited to Fig. 1. For example, the technology of the present application can also be applied to one-sided video encoding or one-sided video decoding.

[0038] The following describes a video coding framework according to an embodiment of the present application.

[0039] 2 is a schematic block diagram of a video encoder according to an embodiment of the present application. It should be understood that the video encoder 200 may be used to perform lossy compression on images or lossless compression on images. The lossless compression may be visually lossless compression or mathematically lossless compression.

[0040] The video encoder 200 is applicable to image data in a luminance and chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents luminance (Luma), Cb (U) represents blue chrominance, and Cr (V) represents red chrominance, and U and V represent chrominance (Chroma) that describes color and saturation. For example, in color format, 4:2:0 represents four luminance components and two chrominance components (YYYYCbCr) for every four pixels, 4:2:2 represents four luminance components and four chrominance components (YYYYCbCrCbCr) for every four pixels, and 4:4:4 represents full pixel display (YYYYCbCrCbCrCbCrCbCr).

[0041] For example, the video encoder 200 reads video data and divides each frame of the video data into several coding tree units (CTUs). In some examples, a CTU may be referred to as a "tree block," a "largest coding unit" (LCU), or a "coding tree block" (CTB). Each CTU may be associated with a block of pixels of the same size in the image. Each pixel may correspond to one luminance (luma) sample and two chrominance (chroma) samples. Thus, each CTU may be associated with one luminance sample block and two chrominance sample blocks. The size of a CTU may be, for example, 128x128, 64x64, 32x32, etc. Each CTU may be further divided into several coding units (CUs) for encoding. A CU may be a rectangular or square block. A CU can be further divided into a prediction unit (PU) and a transform unit (TU), which separates coding, prediction, and transformation, making processing more flexible. In one example, a CTU is divided into CUs in a quadtree manner, and a CU is divided into TUs and PUs in a quadtree manner.

[0042] Video encoders and video decoders can support various PU sizes. Assuming that a specific CU size is 2Nx2N, the video encoder and video decoder support PU sizes of 2Nx2N or NxN for intra prediction, and support symmetric PUs of 2Nx2N, 2NxN, Nx2N, NxN, or similar sizes for inter prediction. The video encoder and video decoder also support asymmetric PUs of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter prediction.

[0043] 2, the video encoder 200 may include a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / inverse quantization unit 240, a reconstruction unit 250, a loop filtering unit 260, a decoded image buffer 270, and an entropy coding unit 280. Note that the video encoder 200 may include more, fewer, or different functional components.

[0044] Optionally, in this application, the current block may be referred to as a current coding unit (CU) or a current prediction unit (PU), etc. The prediction block may also be referred to as a predicted image block or an image prediction block, and the reconstructed image block may also be referred to as a reconstructed block or a is again composition image It may also be called.

[0045] In some embodiments, the prediction unit 210 includes an inter prediction unit 211 and an intra prediction unit 212. sample Because of the strong correlation between adjacent pixels, video coding and decoding techniques use intra prediction methods to predict the sample Eliminate spatial redundancy between adjacent frames. Because there is a high degree of similarity between adjacent frames in a video, video encoding and decoding techniques use inter-prediction methods to eliminate temporal redundancy between adjacent frames, thereby improving coding efficiency.

[0046] The inter prediction unit 211 may be used to perform inter prediction, which may include motion estimation and motion compensation and may refer to image information from different frames. Inter prediction uses motion information to find a reference block from a reference frame and generates a predicted block based on the reference block to eliminate temporal redundancy. The frames used for inter prediction may be P frames and / or B frames, where P frames refer to forward predicted frames and B frames refer to bidirectionally predicted frames. Inter prediction uses motion information to find a reference block from a reference frame and generates a predicted block based on the reference block. The motion information includes a reference frame list in which the reference frame is located, a reference frame index, and a motion vector. The motion vector may be integer or fractional pixels. If the motion vector is fractional pixels, interpolation filtering must be used within the reference frame to create the required fractional pixel block. Here, the integer or fractional pixel block in the reference frame found based on the motion vector is called the reference block. Some techniques use the reference block directly as the predicted block, while other techniques generate a predicted block by processing based on the reference block. Processing based on a reference block to generate a predictive block may also be understood as treating the reference block as a predictive block and then processing based on that predictive block to generate a new predictive block.

[0047] The intra prediction unit 212 predicts pixel information in the current image block by only referring to information of the same frame image, eliminating spatial redundancy. The frame used for intra prediction may be an I-frame.

[0048] There are multiple prediction modes for intra prediction. Taking the H series of international digital video coding standards as an example, the H.264 / AVC standard has eight angular prediction modes and one non-angular prediction mode, while H.265 / HEVC expands this to 33 angular prediction modes and two non-angular prediction modes. The intra prediction modes used in HEVC include planar mode, DC mode, and 33 angular modes, for a total of 35 prediction modes. The intra modes used in VVC include planar mode, DC mode, and 65 angular modes, for a total of 67 prediction modes.

[0049] Furthermore, with the increase in angle modes, intra prediction becomes more accurate and better meets the needs of the development of high-definition and ultra-high-definition digital video.

[0050] The residual unit 220 may generate a residual block of the CU based on the pixel block of the CU and the prediction block of the PU of the CU. For example, the residual unit 220 may generate the residual block of the CU such that the value of each sample in the residual block is equal to the difference value between the sample in the pixel block of the CU and the corresponding sample in the prediction block of the PU of the CU.

[0051] The transform / quantization unit 230 may quantize the transform coefficients. The transform / quantization unit 230 may quantize the transform coefficients associated with the TUs of a CU based on a quantization parameter (QP) value associated with the CU. The video encoder 200 may adjust the degree of quantization applied to the transform coefficients associated with a CU by adjusting the QP value associated with the CU.

[0052] The inverse transform / inverse quantization unit 240 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficients to reconstruct residual blocks from the quantized transform coefficients.

[0053] Reconstruction unit 250 may add the samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by prediction unit 210 to generate a reconstructed image block associated with the TU. By reconstructing the sample blocks of each TU of a CU in this manner, video encoder 200 can reconstruct the pixel blocks of the CU.

[0054] The loop filtering unit 260 is used to process the pixels after inverse transformation and inverse quantization to correct distortion information and provide a better reference for subsequent coding pixels, and can, for example, perform deblocking filtering operations to reduce blocking artifacts in pixel blocks associated with a CU.

[0055] In some embodiments, the loop filtering unit 260 may include a deblocking filtering unit and a sample adaptive compensation unit. / and an adaptive loop filtering (SAO / ALF) unit, where the deblocking filtering unit is used to remove blocking artifacts and the SAO / ALF unit is used to remove ringing artifacts.

[0056] The decoded image buffer 270 can store the reconstructed pixel blocks. The inter prediction unit 211 can perform inter prediction on PUs of other images using a reference image including the reconstructed pixel blocks. The intra prediction unit 212 can also perform intra prediction on other PUs in the same image as the CU using the reconstructed pixel blocks in the decoded image buffer 270.

[0057] The entropy coding unit 280 may receive the quantized transform coefficients from the transform / quantization unit 230. The entropy coding unit 280 may perform one or more entropy coding operations on the quantized transform coefficients to generate entropy coded data.

[0058] FIG. 3 is a schematic block diagram of a video decoder according to an embodiment of the present application.

[0059] 3, the video decoder 300 includes an entropy decoding unit 310, a prediction unit 320, an inverse quantization / inverse transform unit 330, a reconstruction unit 340, a loop filtering unit 350, and a decoded image buffer 360. Note that the video decoder 300 may include more, fewer, or different functional components.

[0060] The video decoder 300 may receive a bitstream. The entropy decoding unit 310 may parse the bitstream and extract syntax elements from the bitstream. As part of parsing the bitstream, the entropy decoding unit 310 may analyze entropy-encoded syntax elements in the bitstream. The prediction unit 320, the inverse quantization / inverse transform unit 330, the reconstruction unit 340, and the loop filtering unit 350 may decode video data based on the syntax elements extracted from the bitstream, i.e., generate decoded video data.

[0061] In some embodiments, the prediction unit 320 includes an intra prediction unit 322 and an inter prediction unit 321 .

[0062] The intra prediction unit 322 may perform intra prediction to generate a predictive block of the PU. The intra prediction unit 322 may use an intra prediction mode to generate a predictive block of the PU based on pixel blocks of spatially neighboring PUs. The intra prediction unit 322 may further determine the intra prediction mode of the PU based on one or more syntax elements parsed from the bitstream.

[0063] The inter prediction unit 321 may construct a first reference image list (list 0) and a second reference image list (list 1) based on syntax elements parsed from the bitstream. Furthermore, if the PU uses inter predictive coding, the entropy decoding unit 310 may analyze motion information of the PU. The inter prediction unit 321 may determine one or more reference blocks for the PU based on the motion information of the PU. The inter prediction unit 321 may generate a prediction block for the PU based on the one or more reference blocks for the PU.

[0064] The inverse quantization / inverse transform unit 330 may inverse quantize (i.e., dequantize) the transform coefficients associated with the TU. The inverse quantization / inverse transform unit 330 may determine the degree of quantization using a QP value associated with the CU of the TU.

[0065] After dequantizing the transform coefficients, the inverse quantization / inverse transform unit 330 may apply one or more inverse transforms to the dequantized transform coefficients to generate a residual block associated with the TU.

[0066] The reconstruction unit 340 reconstructs the pixel block of the CU using the residual block associated with the TU of the CU and the prediction block of the PU of the CU. For example, the reconstruction unit 340 may add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU and obtain a reconstructed image block.

[0067] The loop filtering unit 350 may perform a deblocking filtering process to reduce blocking artifacts in pixel blocks associated with a CU.

[0068] The video decoder 300 can store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 can use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or can transmit the reconstructed image to a display device for display.

[0069] The basic process of video encoding and decoding is as follows: On the encoding side, an image of one frame is divided into blocks, and for a current block, a prediction unit 210 generates a predicted block of the current block using intra prediction or inter prediction. The residual unit 220 calculates a residual block, i.e., a difference value between the predicted block and the original block of the current block, based on the predicted block and the original block of the current block. The residual block may also be referred to as residual information. The residual block may be subjected to processes such as transformation and quantization by the transform / quantization unit 230 to remove information that is difficult for the human eye to perceive, thereby eliminating visual redundancy. Optionally, the residual block before transformation and quantization by the transform / quantization unit 230 may be referred to as a time-domain residual block, and the time-domain residual block after transformation and quantization by the transform / quantization unit 230 may be referred to as a frequency residual block or a frequency-domain residual block. The entropy coding unit 280 may receive the quantized transform coefficients output from the transform / quantization unit 230, perform entropy coding on the quantized transform coefficients, and output a bitstream. For example, the entropy coding unit 280 may remove character redundancy based on a target context model and probability information of the binary bitstream.

[0070] On the decoding side, the entropy decoding unit 310 analyzes the bitstream to obtain prediction information, a quantization coefficient matrix, etc., for the current block. The prediction unit 320 generates a prediction block for the current block using intra- or inter-prediction based on the prediction information. The inverse quantization / inverse transform unit 330 uses the quantization coefficient matrix obtained from the bitstream to inversely quantize and inversely transform the quantization coefficient matrix to obtain a residual block. The reconstruction unit 340 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed block forms a reconstructed image, and the loop filtering unit 350 performs loop filtering on the reconstructed image based on an image or block to obtain a decoded image. On the encoding side, similar processing to that on the decoding side is required to obtain a decoded image. The decoded image may also be called a reconstructed image, and the reconstructed image may be used as a reference frame for inter-prediction for a subsequent frame.

[0071] The block division information determined by the encoding side, and mode or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering are included in the bitstream if necessary. The decoding side analyzes the bitstream and determines the same block division information, mode or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering as the encoding side by analyzing it based on existing information, thereby ensuring that the decoded image obtained on the encoding side is the same as the decoded image obtained on the decoding side.

[0072] The above is the basic process of video codec in a block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized, and the present application is applicable to, but not limited to, the basic process of video codec in the block-based hybrid coding framework.

[0073] In the embodiments of the present application, the current block may be a current coding unit (CU) or a current prediction unit (PU), etc. According to the requirements of parallel processing, an image may be divided into slices, etc., and slices within the same image can be processed in parallel, i.e., there is no data dependency between them. The term "frame" is a commonly used term, and one frame can generally be understood as one image. In the present application, the frame may also be replaced with an image or a slice, etc.

[0074] Current video codec standards for Versatile Video Coding (VVC) include a Geometric Prediction Mode (GPM). P The Audio Video Coding Standard (AVS) video coding standard, which is currently being developed, includes an inter-prediction mode called Angular Weighted Prediction (AWP). P There is also an inter-prediction mode called inter-frame prediction (IFD). Although these two modes have different names and specific implementations, they share a common principle.

[0075] In addition, while conventional unidirectional prediction finds only one reference block of the same size as the current block, conventional bidirectional prediction uses two reference blocks of the same size as the current block and finds the Sample Values are the corresponding positions of the two reference blocks. Sample Values This is the average value of the motion vectors, i.e., all points in each reference block occupy 50% of the total motion vector. With bidirectional weighted prediction, the proportions of the two reference blocks can be different; if all points in the first reference block occupy 75%, all points in the second reference block occupy 25% of the total motion vector. However, the proportions of all points in the same reference block are the same. Decoder side motion vector correction (DMVR) Other optimization methods such as Motion Vector Refinement (MVRE) technology and Bi-directional Optical Flow (BIO) Reference Sample or Prediction Sample However, this is not related to the above principle. BIO can also be abbreviated as BDOF. On the other hand, GPM and AWP also use two reference blocks of the same size as the current block, but at some pixel positions, the position corresponding to the first reference block is used. Sample Values is used 100%, and at some pixel positions, the position corresponding to the second reference block is used. Sample Values is used 100%, and in the boundary or transition area, the positions corresponding to these two reference blocks are used. Sample Values are used at a fixed ratio. The weights of the boundary areas also transition gradually. How these weights are specifically distributed is determined by the GPM or AWP mode. The weight of each pixel position is determined based on the GPM or AWP mode. Of course, in certain cases, for example when the block size is very small, in some GPM or AWP modes, the weight corresponding to the first reference block may be used at some pixel positions. Sample Values is used 100% and corresponds to the second reference block at some pixel locations. Sample Values It may not be possible to guarantee that 100% of the current block is used. GPM or AWP may use two reference blocks of different sizes than the current block, i.e., each may take a necessary part as a reference block. In other words, the part with a weight other than 0 is used as the reference block, while the part with a weight of 0 is excluded. This is an implementation issue and is not the main point of this application.

[0076] For example, FIG. 4 is a schematic diagram of weight allocation. As shown in FIG. 4, a schematic diagram of weight allocation for multiple partition modes of a GPM on a 64x64 current block according to an embodiment of the present application is shown, where the GPM has 64 partition modes. FIG. 5 is a schematic diagram of weight allocation. As shown in FIG. 5, a schematic diagram of weight allocation for multiple partition modes of an AWP on a 64x64 current block according to an embodiment of the present application is shown, where the AWP has 56 partition modes. In both FIG. 4 and FIG. 5, for each partition mode, a black area indicates that the weight value of the position corresponding to the first reference block is 0%, a white area indicates that the weight value of the position corresponding to the first reference block is 100%, a gray area indicates that the weight value of the position corresponding to the first reference block is greater than 0% and less than 100% according to color depth, and a weight value of the position corresponding to the second reference block is 100% minus the weight value of the position corresponding to the first reference block.

[0077] GPM and AWP derive their weights differently. GPM determines the angle and offset for each mode, then calculates the weight matrix for each mode. AWP first creates a one-dimensional weight line, then uses a method similar to intra-angle prediction to spread the one-dimensional weight line across the matrix.

[0078] It should be understood that early encoding and decoding technologies only used rectangular partitioning, whether for CUs, PUs, or transform units (TUs). However, both GPM and AWP achieve non-rectangular partitioning in prediction without partitioning. GPM and AWP use a weight mask of two reference blocks, i.e., the weight map described above. This mask determines the weights of the two reference blocks when generating a predicted block. Alternatively, it can be understood that some of the positions of the predicted block are obtained from the first reference block and other parts are obtained from the second reference block. The transition area (blending area) is obtained by weighting the corresponding positions of the two reference blocks, thereby achieving a smoother transition. Because GPM and AWP do not divide the current block into two CUs or PUs according to the partitioning line, the current block is processed as a whole, including the transformation, quantization, inverse transformation, and inverse quantization of the predicted residual.

[0079] The GPM uses a weight matrix to simulate a geometrical shape partition, or more precisely, a prediction partition. To implement the GPM, two predictors are required in addition to the weight matrix, and each predictor is determined by one unidirectional motion information. The two unidirectional motion information are obtained from a motion information candidate list, e.g., a merge motion information candidate list (mergeCandList). The GPM determines the two unidirectional motion information from the mergeCandList using two indices in the bitstream.

[0080] In inter-prediction, motion information is used to represent "motion." Basic motion information includes information on a reference frame (or reference picture) and a motion vector (MV). In commonly used bidirectional prediction, two reference blocks are used to predict a current block. The two reference blocks can be one forward reference block and one backward reference block. Optionally, both can be forward or both can be backward. "Forward" refers to a time corresponding to a reference frame occurring before the current frame, while "backward" refers to a time corresponding to a reference frame occurring after the current frame. Alternatively, "forward" refers to a position of a reference frame in a video occurring before the current frame, while "backward" refers to a position of a reference frame in a video occurring after the current frame. Alternatively, "forward" refers to a picture order count (POC) of a reference frame being smaller than that of the current frame, while "backward" refers to a POC of a reference frame being larger than that of the current frame. To enable bidirectional prediction, two reference blocks must be found, which requires two sets of reference frame information and motion vector information. Each of these sets can be understood as one unidirectional motion information, and combining these two sets forms one bidirectional motion information. In a specific implementation, the unidirectional motion information and the bidirectional motion information can use the same data structure, except that the two sets of reference frame information and motion vector information in the bidirectional motion information are both valid, while the one set of reference frame information and motion vector information in the unidirectional motion information is invalid.

[0081] In some embodiments, two reference frame lists, denoted as RPL0 and RPL1, are supported, where RPL is an abbreviation for Reference Picture List. In some embodiments, a P slice can use only RPL0, and a B slice can use RPL0 and RPL1. For one slice, each reference frame list has multiple reference frames, and the codec finds a specific reference frame via a reference frame index. In some embodiments, motion information is represented by a reference frame index and a motion vector. For the above bidirectional motion information, the reference frame index refIdxL0 corresponding to reference frame list 0 and the motion vector mvL0 corresponding to reference frame list 0, and the reference frame index refIdxL1 corresponding to reference frame list 1 and the motion vector mvL1 corresponding to reference frame list 1 are used. 1 Here, the reference frame index corresponding to reference frame list 0 and the reference frame index corresponding to reference frame list 1 can be understood as the above-mentioned reference frame information. In some embodiments, the two flags respectively indicate whether to use the motion information corresponding to reference frame list 0 and the reference frame list 1. 1 The predFlagL0 and predFlagL1 are used to indicate whether or not to use the motion information corresponding to each reference frame list, and are denoted as predFlagL0 and predFlagL1, respectively. predFlagL0 and predFlagL1 can also be understood to indicate whether or not the above-mentioned unidirectional motion information is "valid." Although the data structure of motion information is not explicitly mentioned, a reference frame index, a motion vector, and a "valid or not" flag corresponding to each reference frame list are used to represent the motion information. In some standard texts, motion information does not appear, but a motion vector is used, and the reference frame index and the flag indicating whether or not to use the corresponding motion information can be considered to be associated with the motion vector. For convenience of explanation, the term "motion information" is still used in this application, but it should be understood that the term "motion vector" can also be used to describe it.

[0082] The motion information used by the current block may be saved. Subsequent coded and decoded blocks of the current frame may use motion information from previously coded and decoded blocks, such as neighboring blocks, based on their adjacent positional relationships. Because this utilizes spatial correlation, such coded and decoded motion information is called spatial motion information. The motion information used by each block of the current frame may be saved. Subsequent coded and decoded frames may use motion information from previously coded and decoded frames based on their reference relationships. Because this utilizes temporal correlation, such motion information from coded and decoded frames is called temporal motion information. The motion information used by each block of the current frame is typically stored as a fixed-size matrix, such as a 4x4 matrix, as a minimum unit, with each minimum unit individually storing a set of motion information. In this way, each time a block is coded or decoded, the minimum unit corresponding to that position can store the motion information for that block. In this way, when using spatial or temporal motion information, the motion information corresponding to that position can be directly found based on the position. For example, if a 16x16 block uses conventional unidirectional prediction, all 4x4 minimum units corresponding to this block store the unidirectional prediction motion information. When one block uses GPM or AWP, all minimum units corresponding to this block determine the motion information to be stored in each minimum unit based on the GPM or AWP mode, the first motion information, the second motion information, and the position of each minimum unit. In one method, if all 4x4 pixels corresponding to one minimum unit are obtained from the first motion information, this minimum unit stores the first motion information, and if all 4x4 pixels corresponding to one minimum unit are obtained from the second motion information, this minimum unit stores the second motion information. If 4x4 pixels corresponding to one minimum unit are obtained from both the first motion information and the second motion information, the AWP selects and stores one of the motion information.On the other hand, GPM combines and stores two pieces of motion information as bidirectional motion information if they point to different reference frame lists, otherwise it stores only the second piece of motion information.

[0083] Optionally, the mergeCandList may be constructed based on spatial motion information, temporal motion information, history-based motion information, and other motion information. For example, mergeCandList derives spatial motion information using positions 1 to 5 in FIG. 6A and derives temporal motion information using positions 6 or 7 in FIG. 6A. History-based motion information means that each time a block is encoded or decoded, the motion information of that block is added to a first-in, first-out list. The addition process may require several checks, such as whether the motion information overlaps with existing motion information in the list. In this way, the motion information in the history-based list can be referenced when encoding or decoding the current block.

[0084] In some embodiments, a syntax description for a GPM is as shown in Table 1.

[0085] [Table 1]

[0086] As shown in Table 1, in merge mode, the current block may use CIIP or GPM if regular_merge_flag is not 1. If the current block does not use CIIP, GPM is used, as shown in the syntax "if(!ciip_flag[x0][y0])" in Table 1.

[0087] As can be seen from Table 1 above, the GPM needs to transmit three pieces of information in the bitstream: merge_gpm_partition_idx, merge_gpm_idx0, and merge_gpm_idx1. x0, y0 are used to determine the coordinates (x0, y0) of the upper left luminance pixel of the current block relative to the upper left luminance pixel of the image. merge_gpm_partition_idx determines the partition shape of the GPM and is the "simulation partition" as described above. merge_gpm_partition_idx is the weight matrix derivation mode or weight derivation mode index described in this application. merge_gpm_idx0 is the first merge candidate index, and the first merge candidate index is used to determine the first motion information (or called the first merge candidate) according to the mergeCandList. merge_gpm_idx1 is the second merge candidate index, which is used to determine the second motion information (or called the second merge candidate) according to mergeCandList. If MaxNumGpmMergeCand>2, that is, the length of the candidate list is greater than 2, merge_gpm_idx1 needs to be decoded; otherwise, it can be determined directly.

[0088] The decoding process of GPM is described below.

[0089] The information input to the decoding process includes the coordinates (xCb, yCb) of the luminance location of the top left corner of the current block relative to the top left corner of the image, the width of the luminance component of the current block cbWidth, the height of the luminance component of the current block cbHeight, 1 / 16 pixel precision luminance motion vectors mvA and mvB, chrominance motion vectors mvCA and mvCB, reference frame indices refIdxA and refIdxB, and prediction list flags predListFlagA and predListFlagB.

[0090] For example, motion information may be represented by a combination of a motion vector, a reference frame index, and a prediction list flag. VVC supports two reference frame lists, and each reference frame list may have multiple reference frames. Unidirectional prediction uses only one reference block of one reference frame in one reference frame list as a reference, while bidirectional prediction uses each reference block of each reference frame in each of the two reference frame lists as a reference. GPM in VVC uses two unidirectional predictions. In the above mvA and mvB, mvCA and mvCB, refIdxA and refIdxB, and predListFlagA and predListFlagB, A can be understood as the first prediction mode, and B can be understood as the second prediction mode. X is used to represent A or B, and predListFlagX indicates whether X uses the first or second reference frame list. refIdxX indicates the reference frame index in the reference frame list used by X. mvX indicates the luma motion vector used by X. mvCX indicates the chrominance motion vector used by X. Again, in VVC, the combination of motion vector, reference frame index and prediction list flag can be considered to represent the motion information described in this application.

[0091] The information output by the decoding process includes a luma prediction sample matrix predSamplesL of (cbWidth) by (cbHeight), and optionally a Cb chroma prediction sample matrix of (cbWidth / SubWidthC) by (cbHeight / SubHeightC), and optionally a Cr chroma prediction sample matrix of (cbWidth / SubWidthC) by (cbHeight / SubHeightC).

[0092] For illustrative purposes, the following takes the luminance component as an example, and the processing of the chrominance component is similar to that of the luminance component.

[0093] Assume that the size of predSamplesLAL and predSamplesLBL is (cbWidth) x (cbHeight), and these are prediction sample matrices created according to two prediction modes. predSamplesL is derived as follows, and predSamplesLAL and predSamplesLBL are determined based on luma motion vectors mvA and mvB, chroma motion vectors mvCA and mvCB, reference frame indices refIdxA and refIdxB, and prediction list flags predListFlagA and predListFlagB, respectively. That is, prediction is performed based on motion information of each of the two prediction modes, and a detailed description of the process will not be repeated. Typically, GPM is a merge mode, and both of the two prediction modes of GPM are considered to be merge modes.

[0094] Based on merge_gpm_partition_idx[xCb][yCb], the GPM partition angle index variable angleIdx and distance index variable distanceIdx are determined using Table 2.

[0095] [Table 2]

[0096] Because all three components (e.g., Y, Cb, and Cr) can use GPM, some standard texts encapsulate the process by which one component generates a GPM predicted sample matrix as a single subprocess, called the GPM weighted sample prediction process for geometric partitioning mode. All three components call this process, but the parameters they use are different. Here, we use only the luma component as an example. The current luma block's predicted matrix, predSamplesL[xL][yL] (where xL = 0..cbWidth-1, yL = 0..cbHeight-1), is derived by the GPM weighted prediction process. In this process, nCbW is set to cbWidth, nCbH is set to cbHeight, and the predicted sample matrices predSamplesLAL and predSamplesLBL generated by the two prediction modes, as well as angleIdx and distanceIdx, are used as inputs.

[0097] In some embodiments, the weighted prediction derivation process of the GPM includes the following steps.

[0098] The inputs of this process include the width nCbW of the current block, the height nCbH of the current block, two (nCbW) x (nCbH) predicted sample matrices predSamplesLA and predSamplesLB, a GPM division angle index variable angleIdx, a GPM distance index variable distanceIdx, and a component index variable cIdx. Since this example uses luminance as an example, cIdx is set to 0, representing the luminance component.

[0099] The output of this process is a (nCbW) x (nCbH) GPM predicted sample matrix pbSamples.

[0100] Illustratively, the variables nW, nH, shift1, offset1, displacementX, displacementY, partFlip, and shiftHor are derived as follows:

[0101] nW=(cIdx==0)?nCbW:nCbW*SubWidthC nH=(cIdx==0)?nCbH:nCbH*SubHeightC shift1=Max(5,17-BitDepth), where BitDepth is the bit depth of the encoding / decoding.

[0102] offset1=1<<(shift1-1), where "<<" represents a left shift.

[0103] displacementX=angleIdx displacementY=(angleIdx+8)%32 partFlip=(angleIdx>=13&&angleIdx<=27)?0:1 shiftHor=(angleIdx%16==8||(angleIdx%16!=0&&nH>=nW))?0:1 The variables offsetX and offsetY are derived as follows:

[0104] If the value of shiftHor is 0, offsetX=(-nW)>>1 offsetY=((-nH)>>1)+(angleIdx<16?(distanceIdx*nH)>>3:-((distanceIdx*nH)>>3)) If the value of shiftHor is 1, offsetX=((-nH)>>1)+(angleIdx<16?(distanceIdx*nH)>>3:-((distanceIdx*nH)>>3)) offsetY=(-nW)>>1 The variables xL and yL are derived as follows:

[0105] xL=(cIdx==0)?x:x*SubWidthC yL=(cIdx==0)?y:y*SubHeightC The variable wValue representing the prediction sample weight for the current position is derived as follows: wValue is the weight of the predicted value predSamplesLA[x][y] of the prediction matrix of the first prediction mode for point (x, y), and (8-wValue) is the weight of the predicted value predSamplesLB[x][y] of the prediction matrix of the first prediction mode for point (x, y).

[0106] Here, the distance matrix disLut is determined according to Table 3.

[0107] [Table 3]

[0108] weightIdx=(((xL+offsetX)<<1)+1)*disLut[displacementX]+(((yL+offsetY)<<1)+1)*disLut[displacementY] weightIdxL=partFlip?32+weightIdx:32-weightIdx wValue=Clip3(0,8,(weightIdxL+4)>>3) The predicted sample values ​​pbSamples[x][y] are derived as follows:

[0109] pbSamples[x][y]=Clip3(0,(1<<BitDepth)-1,(predSamplesLA[x][y]*wValue+predSamplesLB[x][y]*(8-wValue)+offset1)> >shift1) In addition, one weight value is derived for each position of the current block, and then one predicted value pbSamples[x][y] of the GPM is calculated. In this manner, the weight wValue does not need to be written in matrix form. However, if all wValues ​​at each position are stored in one matrix, it can be understood that this is one weight matrix. The principle is the same whether the method of calculating and adding weights individually for each point to obtain a predicted value of the GPM, or the method of calculating all weights and adding them all at once to obtain a predicted sample matrix of the GPM. In this application, the term "weight matrix" is used in many descriptions to make the expression easier to understand and to illustrate more intuitively using a weight matrix. However, in practice, the description can also be made using the weight of each position. For example, the weight matrix derivation mode may also be called a weight derivation mode.

[0110] In some embodiments, as shown in FIG. 6B , the GPM decoding process can be described as follows: Analyze the bitstream to determine whether the current block uses the GPM technique. If the current block uses the GPM technique, determine a weight derivation mode (or partition mode or weight matrix derivation mode), and first and second motion information. Determine a first prediction block based on the first motion information, determine a second prediction block based on the second motion information, determine a weight matrix based on the weight matrix derivation mode, and determine a prediction block for the current block based on the first prediction block, the second prediction block, and the weight matrix.

[0111] In some embodiments, the gradient of the GPM weights is fixed. In some embodiments, the gradient of the GPM weights is variable. The GPM variable weight gradient allows the gradient of the weight change to be adjusted, allowing the GPM to obtain transition regions of different widths for the same dividing line angle and dividing line offset. Comparing Figures 6C and 6D, Figure 6C is a schematic diagram of the transition region (blending area) of the GPM in VVC, and Figure 6D is an example of the variable weight gradient of the GPM.

[0112] One way to achieve the GPM variable weight gradient is to transmit a flag called the weight gradient index gpm_blending_idx in the bitstream. When the decoder calculates the weights, it derives the weights based on the analyzed weight gradient index. One possible way is to multiply weightIdx by the weight gradient parameter blendingCoeff, as shown below, taking the VVC weight deriving process as an example. The value of blendingCoeff can be 1 / 4, 1 / 2, 1, 2, 4, etc., and blendingCoeff can derive gpm_blending_idx from the transition gradient index.

[0113] … weightIdx=(((xL+offsetX)<<1)+1)*disLut[displacementX]+(((yL+offsetY)<<1)+1)*disLut[displacementY] weightIdx=weightIdx*blendingCoeff weightIdxL=partFlip?32+weightIdx:32-weightIdx wValue=Clip3(0,8,(weightIdxL+4)>>3) … Of course, the weight gradient index does not have to be transmitted in the bitstream, and the weight gradient index gpm_blending_idx or blendingCoeff may be directly derived based on the block size, etc. In this way, even if there is no such weight gradient index in the bitstream, one weight gradient parameter is considered to be determined during the derivation process.

[0114] Intra prediction will be described below.

[0115] Intra prediction method is used to encode and decode the area around the current block. Reconstituted sample of Reference Sample7A is a schematic diagram of intra prediction. As shown in FIG. 7A, the size of the current block is 4×4, and the left side of the current block is used as a predictor. column and the upper one line of sample is currently in the block Reference Sample and intra prediction is Reference Sample These are used to predict the current block. Reference Sample All of the blocks may be available, i.e., they may all have been coded and decoded, or some may not be available. For example, if the current block is the leftmost block in the entire frame, the blocks to the left of the current block may be Reference Sample Or, when encoding or decoding the current block, the bottom left part of the current block has not been encoded or decoded yet, so the bottom left Reference Sample is also not available. Reference Sample If not available, Reference Sample Or it can be padded with a specific value or some method, or it can be unpadded.

[0116] FIG. 7B is a schematic diagram of intra prediction. As shown in FIG. 7B, the multi-reference line (MRL) intra prediction method can be implemented by using more Reference Sample By using , the efficiency of encoding and decoding can be improved, for example, four reference rows / columns Sample of The current block Reference Sample Use as.

[0117] Furthermore, there are multiple prediction modes for intra prediction, as shown in Figures 8A-8C. 8 8A-8I are schematic diagrams of intra prediction. 8 As shown in I, in H.264, intra prediction for a 4x4 block mainly includes nine modes. Here, mode 0 shown in FIG. 8A is the mode for the upper part of the current block. sample is copied to the current block as a predicted value along the vertical direction. Mode 1 shown in Figure 8B is Reference Sampleis copied to the current block as a predicted value along the horizontal direction. Mode 2 (DC mode) shown in Fig. 8C uses the average value of the eight points A to D and I to L as the predicted value of all points. 8 Modes 3 to 8 shown in I correspond to specific angles. Reference Sample Copy the current block to the corresponding position. Reference Sample may not correspond exactly to Reference Sample Weighted average or interpolated Reference Sample sub-pixels may need to be used.

[0118] In addition, there are modes such as Planar and Planar, and with technological development and block expansion, the number of angular prediction modes is increasing. Figure 9 is a schematic diagram of intra prediction modes. As shown in Figure 9, the intra prediction modes used in HEVC include Planar, DC, and 33 angular modes, for a total of 35 prediction modes. Figure 10 is a schematic diagram of intra prediction modes. As shown in Figure 10, the intra modes used in VVC include Planar, DC, and 65 angular modes, for a total of 67 prediction modes. Figure 11 is a schematic diagram of intra prediction modes. As shown in Figure 11, VS3 uses 66 prediction modes, including DC, Planar, Bilinear, PCM, and 62 angular modes.

[0119] Also, Reference Sample Improvement of sub-pixel interpolation and sample There are also several techniques to improve prediction, such as multiple intra-prediction filtering (MIPF) in AVS3. I ntra P rediction F filter) generates predictions using different filters for different block sizes. sample Regarding Reference Sample close to sample uses one filter to generate a prediction, Reference Sample far from sampleuses a different filter to generate predictions. The intra-prediction filter (IPF) in AVS3 I ntra P rediction F filter) and other prediction targets sample In the technology to filter Reference Sample can be used to filter the predictions.

[0120] Intra prediction can improve encoding and decoding efficiency by using a Most Probable Mode List (MPM) intra-mode coding technique. The mode list is composed of intra prediction modes derived based on the intra prediction modes of surrounding encoded and decoded blocks, such as neighboring modes, and commonly used or highly probable intra prediction modes, such as DC, planar, and bilinear modes. Intra prediction modes that reference surrounding encoded and decoded blocks exploit spatial correlation, since textures have a certain degree of spatial continuity. MPM can be used as a predictor for intra prediction modes. That is, the probability that the current block uses MPM is higher than the probability that it does not. Therefore, fewer codewords are used for MPM during binarization, thereby saving overhead and improving encoding and decoding efficiency.

[0121] In some embodiments, intra prediction can be performed using matrix-based intra prediction (MIP), also referred to as matrix weighted intra prediction. As shown in Figure 12, to predict a block of width W and height H, MIP uses H pixels in one column to the left of the current block. Reconstituted sample and W items in the top row of the current block Reconstituted sample The input is required. MIP is Reference SampleThe prediction block is generated through three steps: averaging, matrix vector multiplication, and interpolation. Here, matrix multiplication is the core of MIP. MIP uses matrix multiplication to sample ( Reference Sample ) to generate a prediction block. MIP provides multiple matrices, and the difference in prediction methods is represented by the difference in the matrices. sample However, using different matrices will give different results. Reference Sample The averaging and interpolation process is designed with a trade-off between performance and complexity. For large blocks, Reference Sample By averaging, an effect similar to downsampling can be achieved, where the input can be fitted to a smaller matrix, and by interpolation, an upsampling effect can be achieved. In this way, it is not necessary to provide a matrix of MIPs for blocks of each size, but only one or more matrices of a specific size. As compression performance demands increase and hardware capabilities improve, it is likely that MIPs of higher complexity will appear in future standards.

[0122] MIP is somewhat similar to Planar, but is clearly more complex and flexible than Planar.

[0123] In some embodiments, an intra prediction technique of template-based intra mode derivation (TIMD) may be used. For example, as shown in FIG. 13, the regions to the left and above a current block are used as templates. When encoding and decoding the current block, except for boundary cases, the left and above sides of the current block can theoretically obtain reconstructed values. This is also the basis of many template adaptation methods. In TIMD, the regions to the left and above the current block shown in FIG. 13 are used as templates, and the regions to the left and above the template are used as templates. sample in the template Reference SampleThe decoder can predict on a template using a certain intra-prediction mode and compare the predicted value with the reconstructed value to obtain the cost of the intra-prediction mode on the template, such as SAD, SATD, or SSE. Because the template and the current block are adjacent and correlated, the representation of the prediction mode on the current block can be estimated based on the representation of the prediction mode on the template. TIMD predicts several candidate intra-prediction modes on the template, obtains their costs on the template, and takes one or two intra-prediction modes with the lowest costs as the intra-predicted values ​​for the current block.

[0124] Research has shown that when the cost difference between two intra-prediction modes on a template is not large, compression performance can be improved by performing a weighted average on the predicted values ​​of the two intra-prediction modes. The weights of the predicted values ​​of the two prediction modes are related to the costs, and in some embodiments, the weights are inversely proportional to the costs.

[0125] In general, TIMD can select an intra-prediction mode using the prediction effect of the intra-prediction mode on the template and weight the two intra-prediction modes based on the cost on the template. The advantage of TIMD is that when the current block selects the TIMD mode, it is not necessary to indicate which intra-prediction mode is specifically used, and the decoder itself derives it through the above process, thereby saving some overhead.

[0126] In some embodiments, the intra prediction technique of decoder-side intra mode derivation (DIMD) may be used. DIMD also uses the left and top edges of the current block. Reconstituted sample The prediction mode is derived using the Reconstituted sampleAs shown in FIG. 14A, DIMD analyzes the gradient of the center point of the window, adapts one intra-prediction mode based on the gradient, and analyzes all points that need to be checked, thereby obtaining a result similar to the bar graph of FIG. 14A. Of course, the so-called bar graph is intended to aid understanding, and can be implemented in various simple forms when specifically implemented. In some embodiments, DIMD selects the two highest intra-prediction modes in the bar graph, adds a planar mode, and weights the predicted values ​​of a total of three intra-prediction modes, with the weights relating to the analysis results.

[0127] In one example, the DIMD prediction process is as shown in Figure 14B, where the intra prediction modes corresponding to the two highest intra prediction modes in the bar graph, i.e., M1 and M2, are selected, and a planar mode is added, resulting in a total of three intra prediction modes. Weights ω1, ω2, and ω3 corresponding to the three intra prediction modes are determined, and predicted values ​​Pred1, Pred2, and Pred3 corresponding to the three intra prediction modes are determined. The predicted values ​​corresponding to the three intra prediction modes are weighted based on the weights corresponding to the three intra prediction modes to obtain the final predicted block.

[0128] As can be seen from the above, DIMD is Reconstituted sample The intra prediction mode is selected using gradient analysis, and Planar is added to the two intra prediction modes and weighted based on the analysis results. The advantage of DIMD is that when the current block selects DIMD mode, it is not necessary to indicate which intra prediction mode is specifically used; the decoder itself derives it through the above process, saving some overhead.

[0129] TIMD and DIMD have many similarities, and in some implementations, their names are reversed. They all support weighting of predictors for two or more intra-prediction modes.

[0130] GPM combines two inter-prediction blocks using a weight matrix. In practice, it can be extended to combine any two prediction blocks, such as two inter-prediction blocks, two intra-prediction blocks, or one inter-prediction block and one intra-prediction block. Furthermore, in screen content coding, IBC (intra block copy) or palette prediction blocks can be used as one or two of the prediction blocks.

[0131] In this application, intra, inter, IBC, and palette are referred to as different prediction methods. For convenience of explanation, the term prediction model is used herein. A prediction mode can be understood as a prediction mode based on which a codec can generate information on a predicted block of a current block. For example, in intra prediction, the prediction mode may be an intra prediction mode such as DC, planar, or various intra angle prediction modes. Of course, intra prediction modes are also possible. Reference Sample It is also possible to overlap one or more auxiliary information, such as an optimization method for generating a preliminary predicted block, or an optimization method (e.g., filtering) after generating a preliminary predicted block. For example, inter prediction may be a skip mode, a merge mode, a merge with motion vector difference (MMVD) mode, or an advanced motion vector prediction (AMVP), and may be unidirectional prediction, bidirectional prediction, or multi-hypothesis prediction. When an inter prediction mode uses unidirectional prediction, the prediction mode must also be able to determine one piece of motion information, and a predicted block can be determined based on the one piece of motion information. When an inter prediction mode uses bidirectional prediction, the prediction mode must also be able to determine two pieces of motion information, and a predicted block can be determined based on the two pieces of motion information.

[0132] In this way, the information that the GPM needs to determine can be expressed as one weight derivation mode and two prediction modes. The weight derivation mode is used to determine a weight matrix or weights, and the two prediction modes each determine one predicted block or predicted value. The weight derivation mode is sometimes called a partitioning mode. However, because it simulates partitioning, it is referred to as the weight derivation mode in this application.

[0133] Optionally, the two prediction modes may be obtained from the same or different prediction schemes, where the prediction schemes include, but are not limited to, intra prediction, inter prediction, IBC, palette, etc.

[0134] A specific example is as follows: If the current block uses GPM, this example is used in an inter-frame coded block, and merge mode is allowed for intra prediction and inter prediction. As shown in Table 4, a syntax element intra_mode_idx is added to indicate which prediction mode is an intra prediction mode. For example, if intra_mode_idx is 0, it indicates that both prediction modes are inter prediction modes, that is, mode0IsInter is 1, and mode 1 IsInter is 1. When intra_mode_idx is 1, it indicates that the first prediction mode is an intra prediction mode and the second prediction mode is an inter prediction mode, that is, mode0IsInter is 0 and mode 1 IsInter is 1. When intra_mode_idx is 2, it indicates that the first prediction mode is an inter prediction mode and the second prediction mode is an intra prediction mode, that is, mode0IsInter is 1 and mode 1 IsInter is 0. If intra_mode_idx is 3, it indicates that both prediction modes are intra prediction modes, that is, mode0IsInter is 0 and mode 1 IsInter is 0.

[0135] [Table 4]

[0136] In some embodiments, as shown in Figure 15, the GPM decoding process is expressed as follows: Analyze the bitstream to determine whether the current block uses the GPM technique. If the current block uses the GPM technique, determine a weight derivation mode (or partition mode or weight matrix derivation mode), and a first prediction mode and a second prediction mode. Determine a first prediction block based on the first prediction mode, a second prediction block based on the second prediction mode, and a weight matrix based on the weight matrix derivation mode. Determine a prediction block for the current block based on the first prediction block, the second prediction block, and the weight matrix.

[0137] Template matching will be described below.

[0138] The template matching method was first used in inter prediction, sample By utilizing the correlation between the current block and the surrounding area, the template is used. When the current block is coded and decoded, the coding and decoding of the left and upper areas of the current block has already been completed according to the coding order. Of course, in the implementation of existing hardware decoders, it is not necessarily guaranteed that the decoding of the left and upper areas of the current block will be completed when the decoding of the current block starts. Of course, this refers to inter-blocks. For example, in HEVC, when an inter-frame coded block generates a predicted block, the surrounding areas are used as templates. Reconstituted sample However, the intra-frame coding blocks are not required to be coded in parallel. Reconstituted sample of Reference Sample Theoretically, the left and top sides are available, which means that it is feasible to make corresponding adjustments to the hardware design. Relatively speaking, the right and bottom sides are not available in the coding order of current standards such as VVC.

[0139] As shown in FIG. 16A, rectangular regions on the left and top sides of the current block are set as templates. The height of the left template is generally the same as the height of the current block, and the width of the top template is generally the same as the width of the current block, although they may differ. The motion information or motion vector of the current block is determined by finding the best-matching position of the template in a reference frame. This process can be roughly described as follows: a search is performed within a certain range from a starting position in a certain reference frame. Search rules such as the search range and search step can be preset. Each time a position is moved to, the matching degree between the template corresponding to that position and templates surrounding the current block is calculated. The matching degree can be measured by distortion costs such as SAD (sum of absolute difference) and SATD (sum of absolute transformed difference). Transforms used in SATD are generally Hadamard transform, MSE (mean-square error), etc., and the smaller the SAD, SATD, MSE, etc., the higher the matching degree. The cost is calculated using the predicted block of the template corresponding to that position and the reconstructed blocks of the template surrounding the current block. In addition to searching for integer pixel positions, sub-pel positions can also be searched, and the motion information of the current block is determined based on the position with the highest matching degree. sampleBy utilizing the correlation between the template and the current block, the motion information suitable for the template may also be suitable for the current block. Of course, the template matching method is not necessarily applicable to all blocks, so several methods can be used to determine whether the current block uses the template matching method. For example, a control switch can be used in the current block to indicate whether the template matching method is used. This template matching method is called decoder-side motion vector derivation (DMVD). Both the encoder and decoder can use the template to perform a search, thereby deriving motion information, or finding better motion information based on the original motion information. There is no need to transmit specific motion vectors or motion vector differentials; both the encoder and decoder perform searches according to the same rules, thereby ensuring consistency between encoding and decoding. While the template matching method can improve compression performance, it also requires a "search" within the decoder, which increases the decoder's complexity to some extent.

[0140] Although the template matching method applied to inter prediction has been described above, this template matching method can also be applied to intra prediction. For example, an intra prediction mode is determined using a template. Similarly, for a current block, regions within a certain range above and to the left of the current block, such as the left rectangular region and the upper rectangular region shown in the above figure, can be used as templates. When encoding or decoding the current block, Reconstituted sampleare available. This process can be roughly described as follows: determine a set of candidate intra prediction modes for the current block, where the candidate intra prediction modes constitute a subset of all available intra prediction modes. Of course, the candidate intra prediction modes may be the entire set of all available intra prediction modes. This may be determined based on a trade-off between performance and complexity. The set of candidate intra prediction modes may be determined based on several rules, such as MPM or equal-interval screening. Calculate the cost of each candidate intra prediction mode on the template, such as SAD, SATD, or MSE. Use the mode to make a prediction on the template to create a predicted block, and calculate the cost using the predicted block and the reconstructed block of the template. Modes with lower costs are more likely to match the template, and neighboring blocks are more likely to be affected. sample By utilizing the similarity between the templates, an intra prediction mode that performs well on the template may also be an intra prediction mode that performs well on the current block. One or more low-cost modes are selected. Of course, the above two steps can be repeated. For example, after selecting one or more low-cost modes, a set of candidate intra prediction modes is determined again, the costs of the newly determined set of candidate intra prediction modes are calculated again, and one or more low-cost modes are selected. This can also be understood as coarse selection and fine selection. The finally selected intra prediction mode is determined as the intra prediction mode for the current block, or several finally selected intra prediction modes are used as candidates for the intra prediction mode for the current block. Of course, the set of candidate intra prediction modes can also be sorted using only the template matching method. For example, the MPM list is sorted, i.e., a prediction block is created on the template for each mode in the MPM list, the costs are determined, and the sets are sorted in ascending order of cost. Generally, the earlier a mode is in the MPM list, the less overhead it will have in the bitstream, thus achieving the goal of improving compression efficiency.

[0141] The template matching method can be used to determine the two prediction modes of the GPM. When the template matching method is used for the GPM, for the current block, one control switch can be used to control whether the two prediction modes of the current block use template matching, or two control switches can be used to control whether each of the two prediction modes uses template matching.

[0142] The other is how to use template matching. For example, when a GPM is used in merge mode, such as a GPM in VVC, merge_gpm_idxX is used to determine one piece of motion information from mergeCandList, where capital X is 0 or 1. For the Xth motion information, one method is to optimize it using a template matching method based on the above motion information. That is, when one piece of motion information is determined from mergeCandList based on merge_gpm_idxX and template matching is used for the motion information, optimization is performed based on the above motion information using the template matching method. Another method is to determine one piece of motion information by directly searching based on one default motion information, rather than determining one piece of motion information from mergeCandList using merge_gpm_idxX.

[0143] If the Xth prediction mode is an intra prediction mode and the Xth prediction mode of the current block uses a template matching method, it is not necessary to indicate the index of the intra prediction mode in the bitstream, and one intra prediction mode can be determined using the template matching method. Alternatively, one candidate set or MPM list can be determined using the template matching method, and in this case, it is necessary to indicate the index of the intra prediction mode in the bitstream.

[0144] In the GPM intra-prediction+inter-prediction scheme, one intra-predicted value and one inter-predicted value are weighted using the weight of the GPM mode to obtain a GPM predicted value. Here, the prediction mode information (motion information) for inter-prediction is similar to the derivation method in the VVC standard, but the prediction mode for intra-prediction requires constructing an intra-prediction mode candidate list (also referred to as an MPM list) for the corresponding part of the GPM mode. The encoder writes the intra-prediction mode index selected by the current block to the bitstream, and the decoder constructs the MPM list for the corresponding GPM mode in the same manner during decoding and determines the intra-prediction mode based on the intra-prediction mode index obtained by decoding. Exemplarily, the corresponding part of the GPM mode can be understood as the white or black part in the division diagram of FIG. 4 or FIG. 5, and for convenience of explanation, they may be referred to as the first part and the second part below. In one example, the first part is the white part, and the second part is the black part. The first part corresponds to the first prediction mode, and the second part corresponds to the second prediction mode. The first and second parts are more intuitive and easier to understand, but may not appear in concrete algorithms.

[0145] When constructing an MPM list of intra prediction modes for the corresponding part of the GPM weight derivation mode, several types of pre-defined intra prediction modes are added to the MPM list in order until the length of the list reaches 3. Optionally, the several types of pre-defined intra prediction modes include an intra prediction mode parallel to the GPM division line, an intra prediction mode derived from the DIMD, an intra prediction mode derived from the TIMD, an intra prediction mode of a neighboring block, an intra prediction mode perpendicular to the GPM division line, and a PLANAR mode.

[0146] Here, the intra prediction mode parallel to the GPM division line is as shown in Figure 16B, and the intra prediction mode perpendicular to the GPM division line is as shown in Figure 16C. A current specific implementation is to determine the GPM division angle index angleIdx based on the GPM weight derivation mode, build a lookup table corresponding to angleIdx and the intra prediction mode, and determine the intra prediction mode parallel to the GPM division line from the lookup table based on angleIdx. The perpendicular intra prediction mode is calculated using the parallel intra prediction mode.

[0147] When the intra prediction mode of a neighboring block is used, the intra prediction modes of up to five neighboring blocks are used, and the positions of the five neighboring blocks are as shown in Figure 16D. The coordinates of the upper left corner of the current block are (x0, y0), the width of the current block is width, and the height of the current block is height. The five neighboring blocks are neighboring block A L determined by coordinates (x0-1, y0-1), neighboring block A determined by coordinates (x0+width-1, y0-1), neighboring block A R determined by coordinates (x0+width, y0-1), neighboring block L determined by coordinates (x0-1, y0+height-1), and neighboring block BL determined by coordinates (x0-1, y0+height).

[0148] The range of available neighboring blocks is determined by referring to Table 5 based on whether the intra prediction mode corresponds to the first part or the second part and the angle index angleIdx corresponding to the GPM weight derivation mode.

[0149] [Table 5] In Table 5, A can be understood as the neighboring block above the current block, and L can be understood as the neighboring block to the left of the current block. If the result obtained by referring to Table 5 is A, the intra prediction mode of neighboring block A and the intra prediction mode of neighboring block AR can be used. If the result obtained by referring to Table 5 is L, the intra prediction mode of neighboring block L and the intra prediction mode of neighboring block BL can be used. If the result obtained by referring to Table 5 is L+A, the intra prediction modes of neighboring blocks A, AR, L, and BL can be used. The prediction mode of neighboring block AL is always available. The checking order of neighboring blocks is L → A → BL → AR → AL.

[0150] As can be seen from the above, a GPM has three elements: one weight matrix and two prediction modes. The advantage of a GPM is that it can achieve more autonomous combinations through the weight matrix. On the other hand, a GPM requires more information to be determined, resulting in greater overhead in the bitstream. Taking the GPM as an example, an optional GPM can be used in merge mode. In the bitstream, merge_gpm_partition_idx, merge_gpm_idx0, and merge_gpm_idx1 are used to determine the weight matrix, first prediction mode, and second prediction mode, respectively. The weight matrix and the two prediction modes each have multiple possible options. For example, the weight matrix in VVC has 64 possible options, while merge_gpm_idx0 and merge_gpm_idx1 each allow up to six possible options in VVC. Of course, VVC requires that merge_gpm_idx0 and merge_gpm_idx1 do not overlap. This means that such a GPM has 65x6x5 possible options. When optimizing two pieces of motion information (prediction modes) using MMVD, multiple possible options can be provided for each prediction mode. This number is quite large. On the other hand, it may be discovered that two pieces of motion information (prediction modes) can be optimized using a template matching method, which also provides more possible options. Even this method of optimizing two pieces of motion information (prediction modes) using template matching requires a block-level switch that indicates whether the block currently uses it, given the current state of technological evolution.

[0151] If a GPM uses two intra prediction modes, each of which can use any of the 67 normal intra prediction modes in VVC, and the two intra prediction modes are different, then there are also 64x67x66 possible choices. Of course, to save overhead, each prediction mode can be restricted to only using a subset of all normal intra prediction modes, but there are still many possible choices.

[0152] If the GPM uses one intra prediction mode and one inter prediction mode, the situation can be inferred based on the above intra prediction mode and inter prediction mode situations.

[0153] In some embodiments, the indication of one weight derivation mode and two prediction modes of the GPM is written into the bitstream using respective syntax elements, and the bitstream is parsed. That is, one weight derivation mode has its own one or more syntax elements, the first prediction mode has its own one or more syntax elements, and the second prediction mode has its own one or more syntax elements. Of course, the standard may restrict in some cases that the second prediction mode must not be the same as the first prediction mode, or that some optimization means can be used for two prediction modes simultaneously (this can also be understood as being used for the current block), but in writing and parsing the syntax elements, these three are relatively independent. The so-called relative independence can also be understood to mean that there is a certain correlation, but other possible options after removing the restrictions are still independent.

[0154] For equally probable events, fixed-length coding is more suitable. When there are distinct probabilities, using short codes for high-probability events and long codes for low-probability events can improve coding efficiency. For the two different dimensional modes, the weight derivation mode and the prediction mode, the probability estimates for them are separated from each other.

[0155] One weight derivation mode and two prediction modes jointly generate one prediction block, which then acts on the current block. There is a correlation between them. For example, if the current block contains the edges of two objects that move relative to each other, this is an ideal scenario for inter-frame GPM. In theory, this "split" should occur at the edge of the objects, but in practice, the possible "splits" are limited, making it impossible to cover any edge. Similar "splits" may be selected, and in such cases, there may not be only one similar "split." The selection of which depends on which "split" and the two prediction modes produce the best combination. Similarly, the selection of which prediction mode may depend on which combination produces the best combination. This is because, even if a portion of the prediction mode is used, it may be difficult for that portion to perfectly match the current block for natural video, and the final selection may result in the highest coding efficiency. In another example, GPM is often used when the current block contains a portion where a single object is in relative motion, such as a twisted or distorted area due to arm swinging. This "splitting" is more ambiguous and may ultimately depend on which combination produces the best results. Another scenario is intra-prediction, where the texture of some parts of a natural image is very complex, with some parts having gradations from one texture to another, and some parts not being representable in a single direction. Intra-frame GPM can provide more complex prediction blocks, but intra-frame coded blocks usually have larger residuals than inter-frame coded blocks under the same quantization, and the choice of prediction mode may ultimately depend on which combination produces the best results.

[0156] The term "combination" is mentioned many times above. That is, instead of selecting a weight derivation mode and a prediction mode separately in two or three dimensions, they can be combined to select a combination of a weight derivation mode and a prediction mode. This is reflected in the syntax element. That is, the "combination" syntax element can be used to determine a weight derivation mode and two prediction modes based on this combination.

[0157] For example, the syntax element gpm_cand_idx can be set as shown in Table 6.

[0158] [Table 6]

[0159] Here, gpm_cand_idx determines the index of the GPM candidate combination. The weight derivation mode and the two prediction modes are determined based on gpm_cand_idx. For example, if the current block is an intra-frame coded block (not applicable to screen content coding), the two prediction modes are all intra-prediction modes. If the current block is an inter-frame coded block, there may be more restrictions on the application scenario. For example, in this application scenario, the two prediction modes are only inter-prediction modes.

[0160] That is, the encoder and decoder can each generate the same N candidate combinations. For example, both the encoder and decoder can create a list of N candidate combinations and derive a combination of one weight derivation mode and two prediction modes from each candidate combination. In the bitstream, the encoder only needs to write which candidate combinations will ultimately be selected, and the decoder analyzes which candidate combinations will ultimately be selected by the encoder. In this application, this list is referred to as a GPM combination candidate list or candidate combination list.

[0161] In one example, the GPM combination candidate list is roughly sorted in descending order according to the probability of the combination being selected, allowing shorter codewords to be used for the top-sorted candidate combinations than in existing methods. Meanwhile, longer codewords are used for combinations with very low selection probabilities. This improves overall coding efficiency. Because the existing methods are divided into three parts, the proposed solution theoretically achieves greater flexibility and makes it easier to approximate the most effective probability-codeword correspondences.

[0162] Of course, as mentioned above, in some cases, the number of possible GPM combinations is quite large. Longer codewords are needed to represent the large number of candidates. However, if we can preempt some combinations with very low probability, we can also reduce the cost of combinations with high probability. While existing methods can preempt situations where the probability of occurrence for each part is too low, using a combinatorial approach offers more flexibility. For example, preempting a "split" in an existing method eliminates all possibilities for that "split."

[0163] Another advantage is that this allows for a simpler syntax and eliminates the need for various decision-making when parsing.

[0164] As mentioned above, how to encode gpm_cand_idx is related to their probability. For example, Exponential-Golomb coding is used. If the number of candidates is relatively small, that is, only a few modes with the highest probability can be selected, a fixed-length code can be used. For example, if there are only 16 candidates, the 16 candidates will use a uniform bit-length encoding.

[0165] A different number of candidate combinations can be set for blocks of different sizes. For example, for small blocks, the difference in the impact of similar weight derivation modes or prediction modes on the prediction result is not significant, while for large blocks, the difference in the impact of similar weight derivation modes or prediction modes on the prediction result is more significant. Therefore, one method is to set a relatively small number of candidate combinations for small blocks and a relatively large number of candidate combinations for large blocks. The size of the block can be determined based on the width and height of the block or the number of pixels of the block. As an example, the number of candidates is set to 8 for blocks with a pixel count of less than (or equal to or less than) 256, and the number of candidates is set to 16 for blocks with a pixel count of 256 or more.

[0166] The process of constructing the GPM combination candidate list is explained below.

[0167] In some embodiments, more relevant information can be used, e.g., mode information of surrounding blocks, Reconstituted sample Using these, the probability of occurrence of various combinations can be analyzed.

[0168] One method is to use templates to build a list of potential GPM combinations.

[0169] Typically, the height of the upper template is the same as the width of the left template, and this value can be 1, 2, 4, etc. As an example, when constructing a GPM combination candidate list using templates, the calculation complexity can be appropriately reduced by using an upper template with a height of 1 and / or a left template with a width of 1. Note that the height of the upper template here being 1 means that the upper template of the current block is located one row above the current block, which has been decoded or coded. sample The width of the left template is 1, which means that the left template of the current block is the decoded or encoded part of the left column of the current block. sample It can be understood to include the following.

[0170] When using a template, the current block can use more relevant information, i.e., reconstructed information around the current block, and therefore can better utilize the correlation between the above three elements. In addition, some situations of the current block can be estimated using reconstructed information around the current block.

[0171] One method is to predict a template for each combination using the GPM method and obtain a predicted block for the template of this combination. Since the template already has a reconstruction value, this combination can be used to calculate a predicted distortion cost for the template's predicted block and the template's reconstructed block, such as SAD, SATD, or SSE. The various combinations can be sorted based on the predicted distortion cost, or a list can be constructed that keeps only the top N combinations with the lowest predicted distortion cost. This allows a GPM combination candidate list to be constructed.

[0172] For a given combination, the method generates a first predicted value of the template using a first prediction mode, generates a second predicted value of the template using a second prediction mode, derives weights for pixel positions on the template using a weight derivation mode, and determines the predicted value of the template based on the first predicted value, the second predicted value, and the weights.

[0173] To ensure consistency between encoding and decoding, both the encoder and decoder must use the same method for constructing the GPM combination candidate list. As mentioned above, the number of all possible GPM combinations can be quite large, and the above method is an exhaustive enumeration method. In a specific implementation, a GPM combination candidate list can be constructed using a fast algorithm, but the algorithm used by the encoder and decoder must be the same. For example, hierarchical screening can be performed on various combinations, or some combinations that are estimated to be more likely based on known information can be prioritized, and some early termination conditions can be set.

[0174] In some embodiments, this embodiment is used for blocks of intra-frame coding and is not applied to blocks of screen content coding. This does not mean that the solution cannot be used in blocks of screen content coding, but is merely used to explain the solution in the simplest example. In blocks that are intra-frame coded and do not use screen content coding, only intra prediction modes need to be considered, and screen content coding modes such as IBC and palette and various inter modes do not need to be considered. As described above, this solution can be used in any situation where GPM is available.

[0175] Here, assume that the GPM has 64 possible weight derivation modes and 67 possible intra prediction modes, which can be found in the VVC standard. However, the GPM is not limited to 64 possible weights, nor is it limited to which of the 64 possible weights it has. It should be noted, however, that the VVC GPM chose 64 weights as a trade-off solution between improved prediction performance and increased bitstream overhead. Because this solution no longer uses fixed logic to encode weight derivation modes, theoretically, this solution can use a wider variety of weights and use them more flexibly. Similarly, the GPM is not limited to 67 possible intra prediction modes, nor is it limited to which of the 67 possible intra prediction modes it has. In theory, all possible intra prediction modes can be used by the GPM. For example, by making the intra angle prediction modes more granular and generating more intra angle prediction modes, the GPM can also use more intra angle prediction modes. For example, the MIP (matrix-based intra prediction) mode of VVC can also be used in this solution, but considering that there are multiple submodes that MIP can select, for ease of understanding, MIP is not added to this embodiment.In addition, some wide-angle modes can also be used in this solution, but this embodiment will not further describe them.

[0176] If two intra prediction modes are not allowed to be the same, a total of 64*67*66 possible combinations are possible in this embodiment. When using an exhaustive search method, a template is predicted using all possible combinations, and the distortion cost of each combination is calculated. Since the MPM list for the current block can be obtained based on the prediction modes of surrounding blocks, it is not necessary to try all intra prediction modes. For example, in VVC, the current block can obtain an MPM list with a length of 6. Furthermore, in some subsequent technological advances, there is a solution called secondary MPM, which may derive an MPM list with a length of 22. That is, the sum of the lengths of the first MPM list and the second MPM list is 22. In this solution, intra prediction mode screening can be performed using MPM. Of course, an MPM list suitable for the GPM mode of the current block can also be constructed. For example, all prediction modes used in all blocks neighboring the current block are added to the MPM list. For example, if the MPM list does not include special prediction modes such as DC, horizontal prediction mode, or vertical prediction mode, one or more of these modes are added to the candidate intra prediction modes of this solution. For example, an intra prediction mode associated with the weight division line may be added to the candidate intra prediction modes of this solution. As an example, one or more intra angle prediction modes whose prediction angles are parallel or nearly parallel to the division line may be added to the candidate intra prediction modes of this solution. As another example, one or more intra angle prediction modes whose prediction angles are perpendicular or nearly perpendicular to the division line may be added to the candidate intra prediction modes of this solution. Alternatively, the candidate intra prediction modes of this solution may be determined based on the weight derivation mode. Alternatively, the candidate intra prediction modes of this solution may be determined for each of the two intra prediction modes. In summary, at least one GPM intra prediction mode candidate set / list may be obtained. Of course, to ensure complexity on the decoding side, the total number of available prediction modes may be limited, for example, to a maximum of six. All of these methods described above may be used alone or in any combination.

[0177] Regarding the weight derivation mode, it is possible to try each weight derivation mode. Of course, it is also possible to screen several weight derivation modes. One method is to not use a template if the weights on the template derived by a weight derivation mode have little effect on a certain prediction mode, since a template is required. For example, in mode 54 (square block) in FIG. 4 above, it is conceivable that the second prediction mode has little effect on the template, or even no effect at all on the template. In this case, the second prediction mode has no effect. One possibility is to not use such a weight derivation mode for such a block. Another possibility is to use only a fixed first or second prediction mode for such a weight derivation mode for such a block. It is important to note that for the same weight derivation matrix for blocks of different shapes, the effects of the two prediction modes may also be different. See Figures 16E and 16F.

[0178] Of course, several weight derivation modes can be screened and tested. In other words, the set of weight derivation modes available in this solution is a subset of all weight derivation modes. For example, the same "division" angle in a weight derivation mode can correspond to multiple offset amounts. For example, modes 10, 11, 12, and 13 in the above figure have the same "division" angle but different offset amounts. In this solution, some modes corresponding to the offset amounts can be eliminated. Of course, some modes corresponding to the "division" angle can also be eliminated. In this way, the total number of possible combinations can be reduced, which makes the differences between each possible combination more clear. Of course, different screening methods can be set for different block sizes. For example, fewer weight derivation modes can be used for small blocks and more weight derivation modes for large blocks. Different screening methods can also be set for different block shapes. One interpretation is that the block shape refers to the ratio of width to height.

[0179] One way to implement this screening method is to use a lookup table. For example, since there are a total of 64 possible weight derivation modes, a table containing 64 elements is set up, and the value of each element indicates whether the corresponding weight derivation mode is to be used. One specific example is as follows: Set an array of g_sgpm_splitDir: g_sgpm_splitDir

[64] ={ 1, 1, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 1, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 0,0,0,0,1,1,0,1, 0,0,1,0,0,1,0,0, 1,0,1,1,0,1,0,0, 1,0,0,1,0,0,1,0 }; A value of 1 in g_sgpm_splitDir[x] indicates that the index can use the weight derivation mode of x. Otherwise, it indicates that the index cannot use the weight derivation mode of x. Using different screening methods for different block sizes or block shapes can be achieved by using multiple arrays or a two-dimensional array.

[0180] For each possible combination, we use them to obtain a predicted value for the template. To predict the template, we use the Reference Sample is the top row and left column of the template. sample Two prediction modes are used to obtain two prediction blocks.

[0181] For weights on the template, the weight of the current block can be obtained based on the weight derivation mode, so that, exemplarily, as shown in FIG. 16G, the weight of the template can be derived in the same way, only the position used is different.

[0182] An example of the template weight derivation process is as follows, and some content of this example can be integrated with some content of the prediction weight derivation.

[0183] The input for this process is -Current block width nCbW, current block height nCbH, - width of the left template nTmW, height of the upper template nTmH, -GPM's "split" angle index variable angleIdx, - GPM distance index variable distanceIdx, and - It has a component index variable cIdx. In this example, since only luminance is used as an example, cIdx is 0 in this example, which represents the luminance component.

[0184] The output of this process is the template weight matrix wTemplateValue.

[0185] The variables nW, nH, shift1, offset1, displacementX, displacementY, partFlip, and shiftHor are derived as follows:

[0186] nW=(cIdx==0)?nCbW:nCbW*SubWidthC nH=(cIdx==0)?nCbH:nCbH*SubHeightC shift1=Max(5,17-BitDepth), where BitDepth is the bit depth for encoding and decoding.

[0187] offset1=1<<(shift1-1) displacementX=angleIdx displacementY=(angleIdx+8)%32 partFlip=(angleIdx>=13&&angleIdx<=27)?0:1 shiftHor=(angleIdx%16==8||(angleIdx%16!=0&&nH>=nW))?0:1 The variables offsetX and offsetY are derived as follows:

[0188] If the value of shiftHor is 0: offsetX=(-nW)>>1 offsetY=((-nH)>>1)+(angleIdx<16?(distanceIdx*nH)>>3:-((distanceIdx*nH)>>3)) Otherwise (i.e., if the value of shiftHor is 1), offsetX=((-nW)>>1)+(angleIdx<16?(distanceIdx*nW)>>3:-((distanceIdx*nW)>>3)) offsetY=(-nH)>>1 The template weight matrix wTemplateValue[x][y] is derived as follows (where x=-nTmW..nCbW-1, y=-nTmH..nCbH-1, except when both x and y are greater than or equal to 0): Note that in this example, the coordinates of the top-left corner of the current block are (0, 0).

[0189] The variables xL and yL are derived as follows:

[0190] xL=(cIdx==0)?x:x*SubWidthC yL=(cIdx==0)?y:y*SubHeightC Here, disLut is determined based on Table 3.

[0191] weightIdx=(((xL+offsetX)<<1)+1)*disLut[displacementX]+(((yL+offsetY)<<1)+1)*disLut[displacementY] weightIdxL=partFlip?32+weightIdx:32-weightIdx wTemplateValue[x][y]=Clip3(0,8,(weightIdxL+4)>>3) In some embodiments, for ease of calculation, the weights of the templates may now be set to only two possible values: 0 and 1.

[0192] As an example, An example of the template weight derivation process is shown below, which is similar to the above, but simply weightIdxL=partFlip?32+weightIdx:32-weightIdx, and Simply replace wTemplateValue[x][y]=Clip3(0,8,(weightIdxL+4)>>3) with the following.

[0193] wTemplateValue[x][y]=(partFlip?weightIdx:-weightIdx)>0?1:0 Of course, a hierarchical screening approach can also be used. For example, if a weight derivation mode can obtain a relatively small cost, similar weight derivation modes are subsequently tried. Conversely, if a weight derivation mode cannot obtain a relatively small cost, similar weight derivation modes are not further tried. For example, if a relatively small cost can be obtained with one intra prediction mode, similar intra modes are subsequently tried. Conversely, if a relatively small cost cannot be obtained with one intra prediction mode, similar intra prediction modes are not further tried. Of course, these screening methods are limited to use in combination with the other two factors. For example, if a certain intra prediction mode cannot obtain a relatively small cost as the first prediction mode in a certain weight derivation mode, cases where an intra prediction mode similar to the certain intra prediction mode is the first prediction mode in the weight derivation mode are not tried.

[0194] Obtain the cost of a certain combination, sort each combination based on this cost in ascending order, and finally select N candidate combinations. Or just maintain a list of N candidate combinations. One possible case is N is 8 or 16 or 32, etc.

[0195] As described above, for each combination, a template is predicted using the GPM method to obtain a predicted block for this combination relative to the template. Since the template already has a reconstructed value, this combination can be used to calculate a prediction distortion cost for the template's predicted block and the template's reconstructed block, such as SAD, SATD, or SSE. All that is needed is the cost of the combination on the template. A fast algorithm can be used to calculate this cost, but this fast algorithm does not first combine the predicted blocks obtained using the GPM method with the two prediction modes to obtain a GPM weighted predicted block, and then calculate the prediction distortion cost using this GPM weighted predicted block and the template's reconstructed block.

[0196] As mentioned before, the weights on the template can be simplified to two possibilities: 0 and 1. For each pixel location, Sample Values is obtained only from the prediction block of the first prediction mode or the prediction block of the second prediction mode. Therefore, for one prediction mode, the cost on the template when it is the first prediction mode of a certain weight derivation mode can be calculated. That is, for the weight derivation mode, when the prediction mode is the first prediction mode, some of the weights with 1 are calculated. sample Only the cost generated in the template by

[0000] is calculated. As an example, the cost is represented as cost[pred_mode_idx][gpm_idx][0], where pred_mode_idx represents the index of the prediction mode, gpm_idx represents the index of the weight derivation mode, and 0 represents the first prediction mode.

[0197] Also, the cost on the template when it is the second prediction mode of a certain weight derivation mode is calculated, that is, in the case of the weight derivation mode, when the prediction mode is the second prediction mode, some of the weights having 1 are calculated. sampleOnly the cost generated in the template by

[0000] is calculated. As an example, the cost is represented as cost[pred_mode_idx][gpm_idx][1], where pred_mode_idx represents the index of the prediction mode, gpm_idx represents the index of the weight derivation mode, and 1 represents the second prediction mode.

[0198] Then, when calculating the cost of one combination, the above two corresponding costs can be directly added. For example, the costs of prediction modes pred_mode_idx0 and pred_mode_idx1 in weight derivation mode gpm_idx are calculated, where pred_mode_idx0 is used as the first prediction mode and pred_mode_idx1 is used as the second prediction mode. When the cost is denoted as costTemp, costTemp=cost[pred_mode_idx0][gpm_idx][0]+cost[pred_mode_idx1][gpm_idx][1]. When the cost of prediction modes pred_mode_idx0 and pred_mode_idx1 in weight derivation mode gpm_idx is calculated, where pred_mode_idx1 is used as the first prediction mode and pred_mode_idx0 is used as the second prediction mode, and the cost is denoted as costTemp, then costTemp = cost[pred_mode_idx1][gpm_idx][0] + cost[pred_mode_idx0][gpm_idx][1].

[0199] One advantage of doing this is that it simplifies the process of first weighting and combining into one prediction block and then calculating the cost to directly calculating the costs of the two parts and then adding the costs to get the combined cost. Because one prediction mode can be combined with multiple other prediction modes, for the same weight derivation mode, the prediction mode has fixed costs as parts of the first and second prediction modes, so these costs, i.e., cost[pred_mode_idx][gpm_idx][0] and cost[pred_mode_idx][gpm_idx][1] in the above example, can be kept and reused, thereby reducing the amount of computation.

[0200] The following describes a merging technique using a motion vector difference (Merge mode with MVD: abbreviated as MMVD).

[0201] MMVD is a special merge mode. In a typical merge technique, there is no need to encode or decode the motion vector difference (MVD). In a typical inter mode, the MVD must be encoded and decoded. MMVD encodes the MVD in a special way. It takes advantage of the fact that MVDs tend to be distributed in a single horizontal or vertical direction, i.e., there are many MVDs with small values ​​and few MVDs with large values, as shown in Figures 17A and 17B.

[0202] MMVD can only represent MVDs of specific values ​​in some specific directions, and cannot represent arbitrary MVDs. mmvd_direction_idx (MMVD direction index) is used to indicate the direction of the MVD. Of course, it can also be understood as whether x and y of the MVD are non-zero and the positive or negative sign, and mmvd_distance_idx is used to indicate the magnitude of the absolute value Mmvddistance of the non-zero one of x and y of the MVD.

[0203] In one example, the relationship between mmvd_distance_idx[x0][y0] and MmvdDistance[x0][y0] is as shown in Table 7.

[0204] [Table 7] Here, ph_mmvd_fullpel_only_flag is an image header flag that can be set to two different combinations of MMVD.

[0205] In one example, the relationship between mmvd_direction_idx[x0][y0] and MmvdSign[x0][y0] is as shown in Table 8.

[0206] [Table 8]

[0207] Illustratively, the MVD of the MMVD is obtained as follows:

[0208] MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][0] MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][1] Currently, due to bandwidth constraints, GPM can only predict using two unidirectional motion information, which limits the prediction effect of GPM and further affects the video compression effect.

[0209] To solve the above technical problem, in an embodiment of the present application, when encoding or decoding a current block, K prediction modes of the current block are determined, and at least one prediction mode of the K prediction modes is a multi-directional prediction mode (e.g., a bidirectional prediction mode). In this way, when predicting the current block using the K prediction modes, the prediction accuracy of the current block can be improved, and the video encoding and decoding effect can be further improved.

[0210] Hereinafter, with reference to FIG. 18, the video decoding method provided in the embodiment of the present application will be described by taking the decoding side as an example.

[0211] 18 is a schematic flowchart of a video decoding method according to an embodiment of the present application, which is applied to the video decoder shown in FIG. 1 and FIG. 3. As shown in FIG. 18, the method according to the embodiment of the present application includes steps S101 and S102.

[0212] In step S101, K prediction modes for the current block are determined.

[0213] Here, at least one prediction mode among the K prediction modes is an N-directional prediction mode, and both K and N are positive integers greater than 1.

[0214] Optionally, the above K is a predetermined value or a default value. Optionally, the encoding side indicates the above K to the decoding side, for example, the encoding side determines K prediction modes and writes K into the bitstream, so that the decoding side decodes the bitstream to obtain K. Optionally, K may be determined by the decoding side in other manners, and the embodiments of the present application are not limited thereto.

[0215] As can be seen from the above, in the embodiment of the present application, K prediction modes jointly generate one prediction block, and this prediction block acts on the current block, that is, the current block is predicted based on the K prediction modes to obtain K prediction values, and weighted processing is performed on the K prediction values ​​to obtain the prediction value of the current block.

[0216] In other words, when decoding the current block, the decoding side needs to determine multiple candidate prediction modes, select K prediction modes from the multiple candidate prediction modes, and predict the current block using the K prediction modes to obtain a predicted value of the current block.

[0217] In some embodiments, before determining the K prediction modes of the current block, the decoding side must first determine whether the current block performs weighted prediction processing using K different prediction modes. If the decoding side determines that the current block performs weighted prediction processing using K different prediction modes, it performs the above-mentioned step S101 of determining the K prediction modes of the current block. If the decoding side determines that the current block does not perform weighted prediction processing using K different prediction modes, it skips the above-mentioned step S101.

[0218] In one possible implementation, the decoding side may determine whether the current block performs weighted prediction processing using K different prediction modes by determining a prediction mode parameter of the current block.

[0219] Optionally, in an embodiment of the present application, the prediction mode parameter may indicate whether the current block can use GPM mode or AWP mode, i.e., whether the current block can perform prediction processing using K different prediction modes.

[0220] As can be understood, in the present embodiment, the prediction mode parameter can be understood as a flag indicating whether the GPM mode or the AWP mode is used. Specifically, the encoder can use one variable as the prediction mode parameter and set the value of the variable to achieve the setting of the prediction mode parameter. Exemplarily, in the present application, if the current block uses the GPM mode or the AWP mode, the encoder can set the value of the prediction mode parameter to indicate that the current block uses the GPM mode or the AWP mode. Specifically, the encoder can set the value of the variable to 1. Exemplarily, in the present application, if the current block does not use the GPM mode or the AWP mode, the encoder can set the value of the prediction mode parameter to indicate that the current block does not use the GPM mode or the AWP mode. Specifically, the encoder can set the value of the variable to 0. Furthermore, in the present embodiment, after the encoder completes setting the prediction mode parameter, the encoder can write the prediction mode parameter into a bitstream and transmit it to a decoder, so that the decoder can obtain the prediction mode parameter after parsing the bitstream.

[0221] Based on this, the decoding side decodes the bitstream to obtain the prediction mode parameter, and further determines whether the current block uses GPM mode or AWP mode based on the prediction mode parameter. If the current block uses GPM mode or AWP mode, i.e., if the prediction process is performed using K different prediction modes, it determines K prediction modes for the current block.

[0222] In some embodiments, the embodiments of the present application may also impose conditional restrictions on the current block using GPM mode or AWP mode, i.e., if it is determined that the current block meets a preset condition, it may determine that the current block performs weighted prediction using K prediction modes, and further determine the K prediction modes for the current block.

[0223] For example, when the GPM mode or the AWP mode is applied, the size of the current block can be limited.

[0224] It can be understood that the prediction method proposed in the embodiment of the present application requires generating K predicted values ​​using K different prediction modes, and then weighting the K predicted values ​​to obtain a predicted value of the current block. Therefore, in order to reduce complexity and consider the trade-off between compression performance and complexity, the embodiment of the present application may restrict the use of the GPM mode or AWP mode for blocks of some sizes. Therefore, in the present application, the decoder may first determine the size parameter of the current block, and then determine whether the current block uses the GPM mode or the AWP mode based on the size parameter.

[0225] In an embodiment of the present application, the size parameters of the current block may include the height and width of the current block, and thus the decoder can determine whether the current block uses GPM mode or AWP mode based on the height and width of the current block.

[0226] For example, in this application, it is determined that the current block can use the GPM mode or the AWP mode if the width is greater than threshold 1 and the height is greater than threshold 2. One possible restriction can be understood to be to use the GPM mode or the AWP mode only if the width of the block is greater than (or equal to) threshold 1 and the height of the block is greater than (or equal to) threshold 2. Here, the values ​​of threshold 1 and threshold 2 may be 4, 8, 16, 32, 128, 256, etc., and threshold 1 may be equal to threshold 2.

[0227] For example, in this application, it is determined that the current block can use the GPM mode or the AWP mode if the width is less than threshold 3 and the height is greater than threshold 4. One possible restriction can be understood to be to use the GPM mode or the AWP mode only if the width of the block is less than (or equal to) threshold 3 and the height of the block is greater than (or equal to) threshold 4. Here, the values ​​of threshold 3 and threshold 4 may be 4, 8, 16, 32, 128, 256, etc., and threshold 3 may be equal to threshold 4.

[0228] Furthermore, in the examples of the present application, sample The parameter restrictions can also be used to restrict the size of blocks that can use GPM mode or AWP mode.

[0229] Illustratively, in this application, the decoder first sample Determine the parameters, then sample It can further be determined whether the current block can use the GPM mode or the AWP mode based on the parameters and the threshold value 5. It can be seen that one possible restriction is to use the GPM mode or the AWP mode only if the number of pixels in the block is greater than (or equal to) the threshold value 5. Here, the value of the threshold value 5 may be 4, 8, 16, 32, 128, 256, 1024, etc.

[0230] That is, in this application, the current block can use the GPM mode or the AWP mode only if the size parameter of the current block meets the size requirement.

[0231] For example, in the present application, a frame-level flag may be used to determine whether the present application is applied to a current frame to be decoded. For example, it may be set to apply the present application to intraframes (e.g., I frames) and not to apply the present application to interframes (e.g., B frames, P frames). Alternatively, it may be set to not apply the present application to intraframes and to apply the present application to interframes. Alternatively, it may be set to apply the present application to some interframes and not to apply the present application to some interframes. Since intraframe prediction can also be used in interframes, the present application may also be applied to interframes.

[0232] In some embodiments, a flag at the frame level or below may also be used to determine whether the present application applies to the current block.

[0233] If the decoding side determines that the current block is predicted using K prediction modes based on the above method, it determines K prediction modes for the current block.

[0234] In an embodiment of the present application, in order to improve the prediction effect of the K prediction modes for the current block, at least one prediction mode among the K prediction modes is an N-directional prediction mode, where N is a positive integer greater than 1. Therefore, the N-directional prediction mode can also be understood as a multi-directional prediction mode.

[0235] It should be noted that the N-directional prediction mode in the embodiments of the present application can also be understood as an N-reference image prediction mode, i.e., a mode of prediction based on N reference images. For example, if the i-th prediction mode of the K prediction modes of the current block is an N-directional prediction mode and N=2, it will include two reference images. In this way, one predicted value of the current block is obtained based on the first reference image, and another predicted value of the current block is obtained based on the second reference image, and these two predicted values ​​are processed (for example, added or weighted added) to obtain a predicted value of the current block in the i-th prediction mode. The N reference images may be N decoded images ahead of the current image or N decoded images behind the current image, or the N reference images may include at least one decoded image ahead of the current image and at least one decoded image behind the current image.

[0236] For example, the K prediction modes of the current block include a first prediction mode and a second prediction mode.

[0237] In one example, the first prediction mode is an N-directional prediction mode, for example, the first prediction mode is a bidirectional prediction mode (i.e., a mode that predicts based on two reference images), a three-directional prediction mode (i.e., a mode that predicts based on three reference images), or a four-directional prediction mode (i.e., a mode that predicts based on four reference images), and the second prediction mode is a unidirectional prediction mode.

[0238] In another example, the second prediction mode is an N-directional prediction mode, e.g., 2 The prediction modes include bidirectional prediction mode (i.e., a mode that predicts based on two reference images), three-directional prediction mode (i.e., a mode that predicts based on three reference images), or four-directional prediction mode (i.e., a mode that predicts based on four reference images), and the first prediction mode is a unidirectional prediction mode.

[0239] In another example, the first prediction mode and the second prediction mode are both N-directional prediction modes, for example, both the first prediction mode and the second prediction mode are bidirectional prediction modes, tridirectional prediction modes, or quaddirectional prediction modes. Note that when the first prediction mode and the second prediction mode are both N-directional prediction modes, the number of reference images and / or selection methods corresponding to the first prediction mode and the second prediction mode may be the same or different, and this is not limited in the embodiments of the present application. For example, the first prediction mode is a bidirectional prediction mode, i.e., corresponding to two reference images, and the second prediction mode is a tridirectional prediction mode, i.e., corresponding to three reference images. As another example, the first prediction mode and the second prediction mode are both bidirectional prediction modes, but the reference image selection methods corresponding to the first prediction mode and the second prediction mode may be the same or different. For example, the reference images corresponding to the first prediction mode are two decoded images ahead of the current image, and the reference images corresponding to the second prediction mode are two decoded images behind the current image, or the reference images corresponding to the first prediction mode and the second prediction mode are both two decoded images ahead of the current image, etc.

[0240] A specific method for determining the K prediction modes of the current block on the decoding side will be described below.

[0241] In case 1, when the decoding side presets the current block, the predicted value of the current block is obtained based on K prediction modes without considering the weight derivation mode. At this time, the above S101 includes the following steps S101-A1 and S101-A2:

[0242] In step S101-A1, a candidate prediction mode list is determined, and the candidate prediction mode list includes a plurality of candidate prediction modes.

[0243] In step S101-A2, K prediction modes are selected from the candidate prediction mode list.

[0244] In Case 1, the decoding side first determines a candidate prediction mode list and then selects K prediction modes from the constructed candidate prediction mode list. For example, the decoding side combines each of the K candidate prediction modes from the plurality of candidate prediction modes into a single combination to obtain a plurality of combinations. For each combination, the decoding side predicts a template for the current block using the K candidate prediction modes included in the combination to obtain a template prediction cost corresponding to the combination. In this way, the combination with the smallest template prediction cost is selected from the plurality of combinations based on the template prediction cost, and the K candidate prediction modes included in the combination with the smallest template prediction cost are set as the K prediction modes for the current block.

[0245] For example, when K=2, a first candidate prediction mode list corresponding to the first prediction mode is constructed, and a second candidate prediction mode list corresponding to the second prediction mode is constructed. list When the first prediction mode is an N-directional prediction mode, each candidate prediction mode in the first candidate prediction mode list is also an N-directional prediction mode. When the second prediction mode is an N-directional prediction mode, candidateEach candidate prediction mode in the prediction mode list is also an N-directional prediction mode. Next, candidate prediction mode 11 is selected from the first candidate prediction mode list, candidate prediction mode 21 is selected from the second candidate prediction mode list, and a template for the current block is predicted using the candidate prediction modes 11 and 21 to obtain a template predicted value. A template prediction cost corresponding to the combination of candidate prediction mode 11 and candidate prediction mode 21 is obtained based on the template predicted value and the template reconstruction value. Similarly, candidate prediction mode 12 is selected from the first candidate prediction mode list, candidate prediction mode 22 is selected from the second candidate prediction mode list, and a template for the current block is predicted using the candidate prediction modes 12 and 22 to obtain a template predicted value. A template prediction cost corresponding to the combination of candidate prediction mode 12 and candidate prediction mode 22 is obtained based on the template predicted value and the template reconstruction value. By this analogy, it is possible to determine the template prediction costs corresponding to the combination of each candidate prediction mode in the first candidate prediction mode list and each candidate prediction mode in the second candidate prediction mode list. Furthermore, based on the template prediction cost, two candidate prediction modes with the smallest template prediction costs are obtained, and the two candidate prediction modes are determined as the first prediction mode and the second prediction mode of the current block.

[0246] In case 2, when the decoding side presets the current block, the predicted value of the current block is obtained according to the weight derivation mode and K prediction modes. At this time, the above S101 includes the following steps:

[0247] In step S101-B1, M candidate weight derivation modes are determined.

[0248] In step S101-B2, a candidate prediction mode list is determined.

[0249] In step S101-B3, K prediction modes for the current block are determined based on the M candidate weight derivation modes and the candidate prediction mode list.

[0250] A specific method for determining the M candidate weight derivation modes will be described below.

[0251] In one possible implementation, the AWP has 56 weight derivation modes and the GPM has 64 weight derivation modes. The M candidate weight derivation modes mentioned above include at least one weight derivation mode of the 56 weight derivation modes in the AWP or at least one weight derivation mode of the 64 weight derivation modes in the GPM.

[0252] In one possible implementation, several weight derivation modes in the AWP or GPM can be screened as M candidate weight derivation modes. In other words, the M candidate weight derivation modes in the present embodiment are a subset of all weight derivation modes in the AWP or GPM. For example, the same "division" angle in a weight derivation mode can correspond to multiple offset amounts. As shown in modes 10, 11, 12, and 13 in FIG. 4 or FIG. 5, these "division" angles are the same but the offset amounts are different. In the present embodiment, some modes corresponding to the offset amounts can be deleted. Of course, some modes corresponding to the "division" angles can also be deleted. This reduces the total number of possible combinations and makes the differences between each possible combination more clear. Of course, different screening methods can be set for different block sizes. For example, fewer weight derivation modes are used for small blocks and more weight derivation modes are used for large blocks. Different screening methods can also be set for different block shapes. One interpretation is that the block shape refers to the ratio of width to height.

[0253] In this implementation, the encoding side and the decoding side use the same method for screening M candidate weight derivation modes. In one example, the method for screening M candidate weight derivation modes is set as a default by both the encoding side and the decoding side. In another example, the encoding side instructs the decoding side to use the method for screening M candidate weight derivation modes, and the decoding side uses the same method to screen and obtain the same M candidate weight derivation modes as the encoding side.

[0254] In some embodiments, weight derivation modes corresponding to a preset division angle and / or a preset offset amount are removed from the plurality of preset weight derivation modes to obtain M candidate weight derivation modes. In the weight derivation modes, the same division angle can correspond to multiple offset amounts. Therefore, as shown in FIG. 4, weight derivation modes 10, 11, 12, and 13 have the same division angle but different offset amounts. In this way, some weight derivation modes corresponding to the preset offset amount can be removed, and / or some weight derivation modes corresponding to the preset division angle can also be removed. In some embodiments, the screening conditions corresponding to different blocks may be different. In this way, when determining the M candidate weight derivation modes corresponding to the current block, first determine the screening conditions corresponding to the current block, and then select the M candidate weight derivation modes from a plurality of pre-set weight derivation modes based on the screening conditions corresponding to the current block.

[0255] In some embodiments, the screening conditions corresponding to the current block include screening conditions corresponding to the size of the current block and / or screening conditions corresponding to the shape of the current block. During prediction, for small blocks, the difference in the impact of similar weight derivation modes on the prediction result is not significant, but for large blocks, the difference in the impact of similar weight derivation modes on the prediction result is more significant. Based on this, in embodiments of the present application, different M values ​​are set for blocks of different sizes, i.e., a large M value is set for large blocks and a small M value is set for small blocks.

[0256] In one possible implementation, the encoding side indicates M candidate weight derivation modes to the decoding side.

[0257] In some embodiments, the above-mentioned screening condition comprises an array, the array comprising M elements, the M elements corresponding one-to-one to the M weight derivation modes, and the element corresponding to each weight derivation mode is used to indicate whether the weight derivation mode is available.

[0258] The above array may be a 1-bit numerical value or a 2-bit numerical value.

[0259] For example, taking GPM as an example, there are a total of 64 possible weight derivation modes, and the encoding side sets up a lookup table containing 64 elements, and the value of each element indicates whether the corresponding weight derivation mode is to be used.

[0260] In one example, taking a 1-bit numeric value as an example, a specific example is as follows, setting the g_sgpm_splitDir array:

[0261] g_sgpm_splitDir

[64] ={ 1, 1, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 1, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 0,0,0,0,1,1,0,1, 0,0,1,0,0,1,0,0, 1,0,1,1,0,1,0,0, 1,0,0,1,0,0,1,0 } Here, a value of 1 in g_sgpm_splitDir[x] indicates that the weight derivation mode of index x can be used. Otherwise, it indicates that the weight derivation mode of index x cannot be used. In this example, the decoding side determines 26 candidate weight derivation modes using this array.

[0262] In another example, one array can be used to indicate M candidate weight derivation modes, and the array contains only the indices of the available weight derivation modes. For example, an array g_sgpm_splitDir

[26] ={0,1,6,8,10,12,14,16,18,19,20,22,24,26,28,30,36,37,42,45,48,50,51,53,56,59} with a length of 26 can be used to indicate 26 candidate weight derivation modes. The decoding side determines the weight derivation mode corresponding to the index as a candidate weight derivation mode based on the weight derivation mode index contained in the numeric value, thereby obtaining 26 candidate weight derivation modes.

[0263] In some embodiments, when the screening conditions corresponding to the current block include screening conditions corresponding to the size of the current block and screening conditions corresponding to the shape of the current block, if, for the same weight derivation mode, both the screening conditions corresponding to the size of the current block and the screening conditions corresponding to the shape of the current block indicate that the weight derivation mode is available, the weight derivation mode is determined to be one of the M candidate weight derivation modes; on the other hand, if at least one of the screening conditions corresponding to the size of the current block and the screening conditions corresponding to the shape of the current block indicates that the weight derivation mode is not available, the weight derivation mode is determined not to constitute one of the M candidate weight derivation modes.

[0264] In some embodiments, screening conditions corresponding to different block sizes and screening conditions corresponding to different block shapes can be realized using multiple arrays, respectively.

[0265] In some embodiments, the screening conditions corresponding to different block sizes and the screening conditions corresponding to different block shapes can be realized using a two-bit array, i.e., the two-bit array includes both the screening conditions corresponding to the block size and the screening conditions corresponding to the block shape.

[0266] For example, the screening conditions corresponding to a block of size A and shape B are as follows, and the screening conditions are expressed as a 2-bit array:

[0267] g_sgpm_splitDir

[64] ={ (1,1),(1,1),(1,1),(1,0),(1,0),(0,0),(1,0),(1,1), (1,1),(0,0),(1,1),(1,0),(1,0),(0,0),(1,0),(1,1), (0,1),(0,0),(1,1),(0,0),(1,0),(0,0),(1,0),(0,0), (1,1),(0,0),(0,1),(1,0),(1,0),(1,0),(1,0),(0,0), (0,0),(0,0),(1,1),(0,0),(1,1),(1,1),(1,0),(0,1), (0,0),(0,0),(1,1),(0,0),(1,0),(0,0),(1,0),(0,0), (1,0),(0,0),(1,1),(1,0),(1,0),(1,0),(0,0),(0,0), (1,1),(0,0),(1,1),(0,0),(0,0),(1,0),(1,1),(0,0) } Here, all 1's in g_sgpm_splitDir[x] indicate that the weight derivation mode with index x is available, and at least one 0 in g_sgpm_splitDir[x] indicates that the weight derivation mode with index x is unavailable. For example, g_sgpm_splitDir[4]=(1,0) indicates that weight derivation mode 4 is available for blocks of size A but not for blocks of shape B. Therefore, if the block size is A and the shape is B, the weight derivation mode is unavailable.

[0268] In the above description, the GPM includes 64 weight derivation modes as an example, but the weight derivation modes in the embodiments of the present application include, but are not limited to, the 64 weight derivation modes included in the GPM and the 56 weight derivation modes included in the AWP.

[0269] The process of determining the candidate prediction mode list is described below.

[0270] In addition, the candidate prediction mode list in the embodiments of the present application includes N-directional prediction modes, such as a bidirectional prediction mode (i.e., a mode that predicts based on two reference images), a three-directional prediction mode (i.e., a mode that predicts based on three reference images), or a four-directional prediction mode (i.e., a mode that predicts based on four reference images), thereby ensuring that at least one prediction mode of the K prediction modes of the current block determined based on the candidate prediction mode list is an N-directional prediction mode.

[0271] In some embodiments, the candidate prediction mode list determination process is independent of the M candidate weight derivation modes. That is, it can be understood that the M candidate weight derivation modes correspond to one candidate prediction mode list, thereby reducing the complexity of determining the candidate prediction mode list and further improving decoding efficiency. Note that in this embodiment, since the candidate prediction mode list is independent of the M candidate weight derivation modes, there is no strict distinction in order between S101-B1 and S101-B2 when they are executed. That is, S101-B1 may be executed after S101-B2, before S101-B2, or simultaneously with S101-B2, but the embodiments of the present application are not limited thereto.

[0272] In some embodiments, for each first candidate weight derivation mode of the M candidate weight derivation modes, a candidate prediction mode list corresponding to the first candidate weight derivation mode is determined.

[0273] In one example, the first candidate weight derivation mode is one of the M candidate weight derivation modes. That is, in this example, at least one candidate prediction mode list needs to be determined for each of the M candidate weight derivation modes. As can be seen from the above, one weight derivation mode corresponds to K prediction modes, and the candidate prediction mode list is used to determine a prediction mode. Therefore, in one possible implementation of this example, one candidate prediction mode list is determined for at least one prediction mode of the K prediction modes corresponding to each candidate weight derivation mode of the M candidate weight derivation modes.

[0274] In another example, if the above-mentioned first candidate weight derivation mode is one type of candidate weight derivation mode among M candidate weight derivation modes, the embodiment of the present application needs to classify the M candidate weight derivation modes and construct at least one candidate prediction mode list for each type of candidate weight derivation mode.

[0275] In the embodiment of the present application, the method for determining the candidate prediction mode list corresponding to each first candidate weight derivation mode among the M candidate weight derivation modes is the same. For ease of explanation, the embodiment of the present application will be described taking as an example the determination of the candidate prediction mode list corresponding to one first candidate weight derivation mode.

[0276] A specific method for determining the candidate prediction mode list corresponding to the first candidate weight derivation mode will be described below.

[0277] In some embodiments, the first candidate weight derivation mode corresponds to one candidate prediction mode list.

[0278] In some embodiments, when each prediction mode among the K prediction modes corresponds to one candidate prediction mode list, for the i-th prediction mode among the K prediction modes, the decoding side determines the candidate prediction mode list for the i-th prediction mode, where i is a positive integer less than or equal to K.

[0279] In this embodiment, each prediction mode among the K prediction modes corresponds to one candidate prediction mode list. Therefore, for a first candidate weight derivation mode, the decoding side determines one candidate prediction mode list for each prediction mode among the K prediction modes corresponding to the first candidate weight derivation mode. For example, the K prediction modes include a first prediction mode and a second prediction mode corresponding to the first candidate weight derivation mode, and the decoding side determines one candidate prediction mode list for the first prediction mode and one candidate prediction mode list for the second prediction mode.

[0280] In this embodiment, the process of determining the candidate prediction mode list corresponding to each prediction mode among the above K prediction modes is the same. For convenience of explanation, in this embodiment, the determination of the candidate prediction mode list for the i-th prediction mode among the above K prediction modes is taken as an example.

[0281] The embodiments of the present application do not limit the specific types of candidate prediction modes included in the candidate prediction mode list of the ith prediction mode.

[0282] In some embodiments, when constructing a candidate prediction mode list for the i-th prediction mode, the following types of prediction modes are added to the candidate prediction mode list in order until the length of the list reaches a preset value (e.g., 3):

[0283] 1. A prediction mode in which the prediction angle is parallel to the dividing line of the first candidate weight derivation mode.

[0284] 2. A first candidate prediction mode determined based on the template of the current block, in some embodiments the first candidate prediction mode is also referred to as a TIMD derived prediction mode.

[0285] 3. Reorganization of the current block template sample The second candidate prediction mode determined based on the gradient of , in some embodiments, the second candidate prediction mode may also be referred to as a DIMD-derived prediction mode.

[0286] 4. Prediction modes of neighboring blocks of the current block.

[0287] 5. Prediction mode where the prediction angle is perpendicular to the dividing line of the first candidate weight derivation mode.

[0288] 6.PLANAR mode.

[0289] 7.N-way prediction mode.

[0290] In the above embodiment, a specific process for determining the candidate prediction mode list has been described.

[0291] After determining the candidate prediction mode list based on the above steps, the decoding side executes step S101-B3.

[0292] The process of determining K prediction modes based on the M candidate weight derivation modes and the candidate prediction mode list in S101-B3 above will be described below.

[0293] In an embodiment of the present application, the decoding side selects one candidate weight derivation mode from the M candidate weight derivation modes as the weight derivation mode of the current block, and determines K prediction modes of the current block from at least one candidate prediction mode included in the candidate prediction mode list. Finally, the current block is predicted using the weight derivation mode of the current block and the K prediction modes of the current block to obtain a predicted value of the current block.

[0294] The weight derivation mode of the current block and the K prediction modes of the current block are jointly used to determine the predicted value of the current block.

[0295] The embodiment of the present application does not limit the specific manner in which the decoding side determines the weight derivation mode of the current block and the K prediction modes of the current block based on the M candidate weight derivation modes and the candidate prediction mode list.

[0296] In some embodiments, when the candidate prediction mode list corresponds to the K prediction modes of the current block, that is, when all of the K prediction modes of the current block are selected from the candidate prediction mode list, the decoding side combines M candidate weight derivation modes with the candidate prediction modes included in the candidate prediction mode list. For example, each candidate weight derivation mode among the M candidate weight derivation modes is combined with any K candidate prediction modes in the candidate prediction mode list to obtain multiple combinations, each combination including one candidate weight derivation mode and K candidate prediction modes. Next, the candidate weight derivation mode and the K candidate prediction modes included in each combination are used to predict a template of the current block (for example, an upper template of the current block, a left template of the current block, or both upper and left templates of the current block), determine a cost for each combination, and further determine one combination from the multiple combinations based on the cost. For example, the combination with the lowest cost is selected from multiple combinations, the candidate weight derivation mode included in the combination with the lowest cost is determined as the weight derivation mode for the current block, and the K prediction modes included in the combination with the lowest cost are determined as the K prediction modes for the current block.

[0297] In some embodiments, when the candidate prediction mode list is a candidate prediction mode list for a prediction mode among the K prediction modes of the current block (e.g., K=2), the candidate prediction mode list is a candidate prediction mode list for the first prediction mode. At this time, the decoding side determines a selectable prediction mode set corresponding to the second prediction mode. Next, for each of the M candidate weight derivation modes, the decoding side selects one candidate prediction mode from the candidate prediction mode list for the first prediction mode as a possible first prediction mode, and selects one prediction mode from the selectable prediction mode set corresponding to the second prediction mode as a possible second prediction mode, thereby obtaining a combination consisting of the candidate weight derivation mode, one possible first prediction mode, and one possible second prediction mode, and thus obtaining multiple combinations. Each combination includes one candidate weight derivation mode and two candidate prediction modes. Next, the candidate weight derivation mode and two candidate prediction modes included in each combination are used to predict a template of the current block, determine a cost for each combination, and further determine one combination from the multiple combinations based on the cost. For example, the combination with the lowest cost is selected from multiple combinations, the candidate weight derivation mode included in the combination with the lowest cost is determined as the weight derivation mode for the current block, and the K prediction modes included in the combination with the lowest cost are determined as the K prediction modes for the current block.

[0298] In some embodiments, when the candidate prediction mode list includes a candidate prediction mode list corresponding to each prediction mode among the K prediction modes of the current block, for example, assuming K=2, the decoding side selects a candidate prediction mode list corresponding to the first prediction mode and a candidate prediction mode list corresponding to the second prediction mode. listIn this way, the decoding side selects one candidate weight derivation mode from M candidate weight derivation modes, selects one candidate prediction mode from the candidate prediction mode list for the first prediction mode, and selects one candidate prediction mode from the candidate prediction mode list for the second prediction mode. At this time, the selected one candidate weight derivation mode and the two candidate prediction modes form one combination. Multiple combinations can be obtained by referring to the above method. Each combination includes one candidate weight derivation mode and two candidate prediction modes. Next, the template of the current block is predicted using the candidate weight derivation mode and the two candidate prediction modes included in each combination, the cost of each combination is determined, and one combination is determined from the multiple combinations based on the cost. For example, the combination with the smallest cost is selected from the multiple combinations, the candidate weight derivation mode included in the combination with the smallest cost is determined as the weight derivation mode of the current block, and the K prediction modes included in the combination with the smallest cost are determined as the K prediction modes of the current block.

[0299] Based on the above description, one weight derivation mode and K prediction modes can act on the current block as one combination. In order to save codewords and reduce encoding costs, in some embodiments, the weight derivation mode and K prediction modes corresponding to the current block are combined into one combination, i.e., a first combination, and a first index is used to indicate the first combination. Compared with separately indicating the weight derivation mode and the K prediction modes, the embodiments of the present application use fewer codewords and further reduce encoding costs.

[0300] Based on this, the above S101-B3 includes the following steps S101-B31 to S101-B33.

[0301] In step S101-B31, the bitstream is decoded to obtain a first index, and the first index is used to indicate a first combination, where the first combination includes a weight derivation mode of the current block and K prediction modes of the current block.

[0302] In step S101-B32, a candidate combination list is determined based on the M candidate weight derivation modes and the candidate prediction mode list, and the candidate combination list includes at least one candidate combination, and the candidate combination includes one weight derivation mode and K prediction modes.

[0303] In step S101-B33, a first combination is determined from the candidate combination list based on the first index.

[0304] The embodiment of the present application does not limit the specific format of the syntax element of the first index.

[0305] In one possible implementation, if the current block is predicted using GPM techniques, gpm_cand_idx is used to represent the first index.

[0306] Since the above-mentioned first index is used to indicate the first combination, in some embodiments, the first index may also be referred to as a first combination index or an index of the first combination.

[0307] In one example, after adding the first index to the bitstream, the syntax is as shown in Table 9.

[0308] [Table 9]

[0309] where gpm_cand_idx is the first index.

[0310] For example, the candidate combination list is as shown in Table 10.

[0311] [Table 10]

[0312] As shown in Table 10, the candidate combination list includes multiple candidate combinations, and any two of the multiple candidate combinations are not completely the same, that is, the weight derivation modes and at least one mode among the K prediction modes included in any two candidate combinations are different. For example, the weight derivation modes of candidate combination 1 and candidate combination 2 are different, or the weight derivation modes of candidate combination 1 and candidate combination 2 are the same and at least one prediction mode among the K prediction modes is different, or the weight derivation modes of candidate combination 1 and candidate combination 2 are different and at least one prediction mode among the K prediction modes is different.

[0313] For example, in the above Table 10, the order of the candidate combinations in the candidate combination list is used as an index. Optionally, the index of the candidate combinations in the candidate combination list may be reflected in other ways, but the embodiment of the present application is not limited thereto.

[0314] In this embodiment, the decoding side decodes the bitstream, obtains a first index, determines the candidate combination list shown in Table 10 above, searches in the candidate combination list based on the first index, and obtains the weight derivation mode and K prediction modes of the current block included in the first combination indicated by the first index.

[0315] For example, the first index is index 1, and in the candidate combination list shown in Table 10, the candidate combination corresponding to index 1 is candidate combination 2. That is, the first combination indicated by the first index is candidate combination 2. In this way, the decoding side determines the weight derivation mode and K prediction modes included in candidate combination 2 as the weight derivation mode and K prediction modes of the current block included in the first combination, predicts the current block using the weight derivation mode and K prediction modes of the current block, and obtains a predicted value of the current block.

[0316] In Scheme 2, the encoding side and the decoding side can each determine the same candidate combination list. For example, the encoding side and the decoding side each determine a list including X candidate combinations, each of which includes one weight derivation mode and K prediction modes. In the bitstream, the encoding side only needs to write the final selected candidate combination, such as the first combination. The decoding side analyzes the first combination finally selected by the encoding side. Specifically, the decoding side decodes the bitstream to obtain a first index, and determines the first combination from the candidate combination list determined by the decoding side based on the first index.

[0317] The specific process of determining the candidate combination list based on the M candidate weight derivation modes and the candidate prediction mode list in the above S101-B32 will be described below.

[0318] The embodiment of the present application does not limit the specific manner of determining the candidate combination list based on the M candidate weight derivation modes and the candidate prediction mode list in the above S101-B32.

[0319] In some embodiments, the M candidate weight derivation modes are arbitrarily combined with multiple candidate prediction modes included in the candidate prediction mode list, with each combination including one weight derivation mode and two prediction modes. In this way, multiple combinations can be obtained. Information related to the current block is used to analyze the probability of occurrence of different combinations, and a candidate combination list is constructed based on the probability of occurrence of each combination. Optionally, the information related to the current block includes mode information of blocks surrounding the current block, mode information of the current block, and the like. Reconstituted sample Includes:

[0320] In some embodiments, the above S101-B32 includes the following steps S101-B321 and S101-B322.

[0321] In step S101-B321, T second combinations are obtained based on the M candidate weight derivation modes and the candidate prediction mode list.

[0322] In step S101-B322, a candidate combination list is obtained based on the T second combinations.

[0323] Here, any of the T second combinations includes one weight derivation mode and K prediction modes, and the weight derivation mode and the K prediction modes included in any two of the T second combinations are not exactly the same, and T is a positive integer greater than 1.

[0324] The embodiment of the present application does not limit the specific manner of obtaining the T second combinations based on the M candidate weight derivation modes and the candidate prediction mode list in the above S101-B321.

[0325] In some embodiments, the candidate prediction mode list may include a candidate prediction mode list corresponding to each of the K prediction modes of the current block. For example, assuming K=2, the decoding side determines a candidate prediction mode list for the first prediction mode and a candidate prediction mode list for the second prediction mode. In this manner, the decoding side selects one candidate weight derivation mode from M candidate weight derivation modes, selects one candidate prediction mode from the candidate prediction mode list for the first prediction mode, and selects one candidate prediction mode from the candidate prediction mode list for the second prediction mode. At this time, the selected one candidate weight derivation mode and the two candidate prediction modes form one second combination. Using the above method, T second combinations can be obtained. Each second combination includes one candidate weight derivation mode and two candidate prediction modes.

[0326] The implementation of obtaining the candidate combination list based on the T second combinations in the above S101-B322 includes, but is not limited to, the following several ways.

[0327] In Method 1, the T second combinations are rearranged according to a preset rule to obtain a candidate combination list.

[0328] In Method 2, for any of the T second combinations, a cost corresponding to the second combination when predicting a template of a current block using the weight derivation mode in that second combination and the K prediction modes is determined, and a candidate combination list is determined based on the cost corresponding to each second combination in the T second combinations.

[0329] In the second method, for each of the T second combinations, a template for the current block is predicted using the weight derivation mode and K prediction modes included in the second combination to obtain a predicted value of the template corresponding to the second combination. Specifically, for each of the T second combinations, a template for the current block is predicted using the K prediction modes included in the second combination to obtain K predicted values. Next, a template weight corresponding to the second combination is determined based on the weight derivation mode for the second combination, and the K predicted values ​​of the template are weighted based on the template weight to obtain a predicted value of the template corresponding to the second combination.

[0330] Since the template of the current block is a reconstructed region, the decoding side can obtain the reconstructed value of the template. Thus, for each second combination in the T second combinations, the cost corresponding to the second combination can be determined based on the template predicted value and the template reconstructed value corresponding to the second combination. Here, methods for determining the cost corresponding to the second combination include, but are not limited to, SAD, SATD, SEE, etc. Next, a candidate combination list is constructed based on the cost corresponding to each second combination in the T second combinations.

[0331] In some embodiments, a cost corresponding to each second combination can be determined using a fast cost calculation method. As can be seen from the above, the template prediction values ​​corresponding to the second combination include template prediction values ​​corresponding to each of the K prediction modes included in the second combination. In this case, the costs corresponding to each of the K prediction modes in the second combination can be determined based on the template prediction values ​​and template reconstruction values ​​corresponding to each of the K prediction modes in the second combination, and the cost corresponding to the second combination can be determined based on the costs corresponding to each of the K prediction modes in the second combination. For example, the sum of the costs corresponding to each of the K prediction modes in the second combination can be determined as the cost corresponding to the second combination.

[0332] In the present embodiment, taking K=2 as an example, the weights on the template can be simplified to only two possibilities, 0 and 1, so that for each pixel position, Sample Values is obtained only from the prediction block of the first prediction mode or the prediction block of the second prediction mode. Therefore, for one prediction mode, the cost on the template when it is the first prediction mode of a certain weight derivation mode can be calculated. That is, when the prediction mode is the first prediction mode under the weight derivation mode, some of the weights with 1 are calculated. sample Only the cost generated in the template by

[0000] is calculated. As an example, the cost is written as cost[pred_mode_idx][gpm_idx][0], where pred_mode_idx represents the index of the prediction mode, gpm_idx represents the index of the weight derivation mode, and 0 represents the first prediction mode.

[0333] In addition, the cost on the template when the prediction mode is the second prediction mode of a certain weight derivation mode can be calculated. That is, when the prediction mode is the second prediction mode under the weight derivation mode, some of the weights with 1 can be calculated. sampleOnly the cost generated in the template by

[0000] is calculated. As an example, the cost is written as cost[pred_mode_idx][gpm_idx][1], where pred_mode_idx represents the index of the prediction mode, gpm_idx represents the index of the weight derivation mode, and 1 represents the second prediction mode.

[0334] When calculating the cost of one combination, the above two corresponding costs can be directly added. For example, when the costs of prediction modes pred_mode_idx0 and pred_mode_idx1 in weight derivation mode gpm_idx are calculated, where pred_mode_idx0 is used as the first prediction mode and pred_mode_idx1 is used as the second prediction mode, and the cost is denoted as costTemp, costTemp=cost[pred_mode_idx0][gpm_idx][0]+cost[pred_mode_idx1][gpm_idx][1]. When the costs of prediction modes pred_mode_idx0 and pred_mode_idx1 in weight derivation mode gpm_idx are calculated, where pred_mode_idx1 is used as the first prediction mode and pred_mode_idx0 is used as the second prediction mode, and the cost is written as costTemp, then costTemp = cost[pred_mode_idx1][gpm_idx][0] + cost[pred_mode_idx0][gpm_idx][1].

[0335] One advantage of doing this is that it simplifies the process of first weighting and combining into one prediction block and then calculating the cost to directly calculating the costs of the two parts and then adding the costs to get the combined cost. Because one prediction mode can be combined with multiple other prediction modes, for the same weight derivation mode, the prediction mode has fixed costs as parts of the first prediction mode and the second prediction mode, so these costs, i.e., cost[pred_mode_idx][gpm_idx][0] and cost[pred_mode_idx][gpm_idx][1] in the above example, can be kept and reused, thereby reducing the amount of computation.

[0336] According to the above method, a cost corresponding to each second combination in the T second combinations can be determined, and then a candidate combination list can be constructed based on the cost corresponding to each second combination in the T second combinations.

[0337] In Example 1, the T second combinations are sorted based on the costs corresponding to each of the T second combinations, and the sorted T second combinations are determined as a candidate combination list. In Example 1, the generated candidate combination list includes the T first candidate combinations.

[0338] In Example 2, C second combinations are selected from T second combinations based on the costs associated with the second combinations, and a list consisting of the C second combinations is determined as a candidate combination list. Optionally, the C second combinations are the C second combinations with the lowest costs among the T second combinations. For example, based on the costs associated with each second combination among the T second combinations, the C second combinations with the lowest costs are selected from the T second combinations to form the candidate combination list. In this case, the candidate combination list includes C candidate combinations. Optionally, the C candidate combinations in the candidate combination list are sorted in ascending order according to the magnitude of their costs, i.e., the costs associated with the C candidate combinations in the candidate combination list increase in order according to the sorting.

[0339] Based on the above steps, the decoding side determines a candidate combination list, selects a first combination corresponding to a first index from the candidate combination list, determines a weight derivation mode included in the first combination as a weight derivation mode for the current block, and determines K prediction modes included in the first combination as K prediction modes for the current block.

[0340] Based on the above steps, the decoding side determines K prediction modes of the current block, and then performs the following step S102.

[0341] In step S102, the current block is predicted based on the K prediction modes of the current block to obtain a predicted value of the current block.

[0342] In an embodiment of the present application, at least one of the K prediction modes of the current block determined by the decoding side is an N-directional prediction mode, and in this manner, when predicting the current block based on the K prediction modes, the prediction accuracy can be improved.

[0343] The embodiment of the present application does not limit the specific types of the K prediction modes of the current block.

[0344] In some embodiments, all of the K prediction modes are inter prediction modes.

[0345] In some embodiments, some of the K prediction modes are intra prediction modes and some are inter prediction modes.

[0346] In some embodiments, the N-way prediction mode of the embodiments of the present application is an inter prediction mode, such as bidirectional motion information or multidirectional motion information.

[0347] In this embodiment, N-directional prediction modes are applied to the current block, and one predicted value of the current block is finally obtained. For example, the K prediction modes include a first prediction mode and a second prediction mode, where the first prediction mode is the N-directional prediction mode. The decoding side predicts the current block using the N-directional prediction mode to obtain a predicted value 1 of the current block, and predicts the current block using the second prediction mode to obtain a predicted value 2 of the current block. Then, the decoding side processes the predicted value 1 and the predicted value 2 to obtain a predicted value of the current block.

[0348] The embodiment of the present application does not limit the specific manner in which the decoding side predicts the current block based on the K prediction modes and obtains the predicted value of the current block.

[0349] In some embodiments, if the above N-way prediction mode is N-way motion information, the above S102 includes the following steps:

[0350] In step S102-A, for the i-th prediction mode among the K prediction modes, if the i-th prediction mode is an N-directional prediction mode, N pieces of first motion information corresponding to the i-th prediction mode are determined (i is a positive integer less than or equal to K).

[0351] In step S102-B, the i-th predicted value of the current block is obtained based on the N first motion information.

[0352] In step S102-C, the current block is predicted using prediction modes other than the i-th prediction mode among the K prediction modes to obtain other predicted values ​​of the current block.

[0353] In step S102-D, a prediction value for the current block is obtained based on the i-th prediction value and other prediction values.

[0354] In this embodiment, when the i-th prediction mode of the K prediction modes of the current block is an N-directional prediction mode, the decoding side first determines N pieces of first motion information corresponding to the i-th prediction mode, where the first motion information can be understood as the initial value of each motion information in the N-directional motion information, or referred to as initial motion information. The i-th predicted value of the current block is obtained based on the N pieces of first motion information. Next, a prediction mode other than the i-th prediction mode of the K prediction modes is used to predict the current block to obtain another predicted value of the current block. Finally, a predicted value of the current block is obtained based on the i-th predicted value and the other predicted values.

[0355] For example, assume that the K prediction modes include a first prediction mode and a second prediction mode. In one example, when the first prediction mode is an N-directional prediction mode and the second prediction mode is a non-N-directional prediction mode, N pieces of first motion information corresponding to the first prediction mode are determined, and a prediction value 1 of the current block is obtained based on the N pieces of first motion information. Next, the current block is predicted using the second prediction mode to obtain a prediction value 2 of the current block, and a prediction value of the current block is further obtained based on the prediction value 1 and the prediction value 2. In another example, when the first prediction mode is a non-N-directional prediction mode and the second prediction mode is an N-directional prediction mode, the current block is predicted using the first prediction mode to obtain a prediction value 1 of the current block, and N pieces of first motion information corresponding to the second prediction mode are determined, and a prediction value 2 of the current block is obtained based on the N pieces of first motion information, and a prediction value of the current block is further obtained based on the prediction value 1 and the prediction value 2. In another example, when the first prediction mode is an N-directional prediction mode and the second prediction mode is also an N-directional prediction mode, N pieces of first motion information corresponding to the first prediction mode are determined, and a predicted value 1 of the current block is obtained based on the N pieces of first motion information. N pieces of first motion information corresponding to the second prediction mode are determined, and a predicted value 2 of the current block is obtained based on the N pieces of first motion information. A predicted value of the current block is further obtained based on the predicted value 1 and the predicted value 2.

[0356] A specific process by which the decoding side determines the N-th piece of first motion information corresponding to the i-th prediction mode will be described below.

[0357] In an embodiment of the present application, each motion information in the N-way motion information includes one reference image and one motion vector, and a reference block of the current block can be obtained in the reference image based on the motion vector. In this way, N reference blocks of the current block can be determined based on the N-way motion information, and a predicted value of the current block in the i-th prediction mode can be obtained based on the N reference blocks.

[0358] In some embodiments, encoding sidewrites the reference image index and the motion vector index corresponding to each piece of motion information in the N-directional motion information into the bitstream. In this way, the decoding side obtains the reference image index and the motion vector index corresponding to the motion information by decoding the bitstream, and further obtains the reference image corresponding to the motion information based on the reference image index. The motion vector corresponding to the motion information is obtained based on the motion vector index. Specifically, the decoding side constructs a reference image list, obtains the reference image corresponding to the motion information from the constructed reference image list based on the reference image index, and constructs a motion information candidate list, and obtains the reference image corresponding to the motion information from the constructed reference image list based on the motion vector index. Movement information candidate list A motion vector corresponding to the motion information is obtained from the

[0359] As can be seen from the above, before determining the N pieces of first motion information, the decoding side first checks the reference image list and Movement information candidate list needs to be constructed.

[0360] In some embodiments, the decoding side builds one reference image list for N-way motion information.

[0361] In some embodiments, the decoding side constructs one reference picture list for each motion information in the N-directional motion information. For example, because the N-directional motion information includes motion information 1 and motion information 2, the decoding side constructs a reference picture list RPL0 for motion information 1 and a reference picture list RPL1 for motion information 2. In one example, the reference pictures in at least one of the reference picture lists RPL0 and RPL1 are preset or default. In another example, the encoding side transmits identification information (e.g., image sequence numbers) of the reference pictures included in at least one of the reference picture lists RPL0 and RPL1 to the decoding side, and in this way, the decoding side can construct at least one of the reference picture lists RPL0 and RPL1 based on the identification information of the reference pictures.

[0362] below, Movement information candidate list Explain the construction process.

[0363] In some embodiments, the decoding side generates one motion vector for N-way motion information. Movement information candidate list Build.

[0364] In some embodiments, the decoding side generates one N-way motion information for each motion information in the N-way motion information. Movement information candidate list Build.

[0365] In some embodiments, the above Movement information candidate list may be a Merge candidate list. Illustratively, the Merge candidate list includes a plurality of candidate motion vectors.

[0366] For example, assume that the K prediction modes include a first prediction mode and a second prediction mode, and both the first prediction mode and the second prediction mode are inter-prediction modes. In this case, the K prediction modes can be understood as K pieces of motion information, such as motion information 1 and motion information 2, where at least one of the motion information 1 and the motion information 2 is N-directional motion information. Assuming that the motion information 1 is bidirectional motion information, the decoding side decodes the bitstream to obtain reference image indexes and motion vector indexes corresponding to each piece of motion information in the bidirectional motion information. The reference image index corresponding to the first motion information in the bidirectional motion information is refIdxL0, the motion vector index corresponding to the first motion information is mvIdxL0, and the reference image index corresponding to the second motion information is refIdxL1. 2 Assume that the motion vector index corresponding to the th motion information is mvIdxL1. In this way, the decoding side determines the reference picture list RPL0 corresponding to the first motion information based on the above method, and further obtains the reference picture refL0 corresponding to the first motion information in the reference picture list RPL0 based on the reference picture index refIdxL0. Next, the decoding side performs the following based on the above method: Movement information candidate list(e.g., Merge candidate list), and further based on the motion vector index mvIdxL0, Movement information candidate list In step S100, the decoding side obtains a motion vector mvL0 corresponding to the first motion information, and determines the motion vector mvL0 and reference image refL0 corresponding to the first motion information as the first motion information or initial motion information corresponding to the first motion information. Similarly, the decoding side determines a reference image list RPL1 corresponding to the second motion information based on the above method, and further obtains a reference image refL1 corresponding to the second motion information in the reference image list RPL1 based on the reference image index refIdxL1. Next, the decoding side determines the reference image list RPL1 corresponding to the second motion information based on the motion vector index mvIdxL1. Movement information candidate list In step S100, a motion vector mvL1 corresponding to the second motion information is obtained, and the motion vector mvL1 and reference image refL1 corresponding to the second motion information are determined as first motion information or initial motion information corresponding to the second motion information. In this way, two pieces of first motion information corresponding to the first prediction mode can be obtained. Similarly, when the second prediction mode is also a bidirectional prediction mode, two pieces of first motion information corresponding to the second prediction mode can be obtained based on the above method.

[0367] As can be seen from the above, each of the N pieces of first motion information includes reference image information corresponding to the motion information and motion vector information, where the reference image information may be a reference image index or a POC of a reference image, etc. The motion vector information may be a first motion vector, also called an initial motion vector or an initial value of a motion vector.

[0368] In the above embodiment, the decoding side performs the following on the basis of the index of the motion vector: Movement information candidate list In some embodiments, the encoding side directly encodes the first motion vector into the bitstream, so that the decoding side can directly decode the first motion vector from the bitstream.

[0369] In an embodiment of the present application, the decoding side determines N pieces of first motion information corresponding to the i-th prediction mode based on the above steps, and then obtains the i-th predicted value of the current block based on the N pieces of first motion information.

[0370] The embodiment of the present application does not limit the specific manner of obtaining the i-th predicted value of the current block based on the N first motion information in S102-B.

[0371] In some embodiments, the decoding side directly predicts the current block using N pieces of first motion information to obtain the i-th predicted value of the current block. Specifically, for each piece of first motion information in the N pieces of first motion information, the decoding side obtains a reference image corresponding to the first motion information based on the reference image information included in the first motion information, and obtains one reference block for the current block in the reference image based on the first motion vector included in the first motion information. In this way, the decoding side obtains one reference block for the current block for each piece of first motion information in the N pieces of first motion information, and further obtains N reference blocks for the current block. Based on these N reference blocks, the i-th predicted value (or i-th predicted block) of the current block can be obtained.

[0372] In some embodiments, as can be seen from the above, each of the N pieces of first motion information includes one first motion vector, and thus the N pieces of first motion information correspond to the N first motion vectors. At this time, the above S102-B includes the following steps S102-B1 and S102-B2.

[0373] In step S102-B1, at least one first motion vector among the N first motion vectors is refined to obtain at least one second motion vector.

[0374] In step S102-B2, obtain the i-th predicted value of the current block based on at least one second motion vector.

[0375] In this embodiment, to further improve the prediction accuracy of the N-directional prediction mode for the current block, at least one first motion vector among the N first motion vectors corresponding to the N-directional prediction mode is refined to obtain at least one second motion vector, for example, the first motion vector mvL0 and / or the first motion vector mvL1 is refined. Here, the second motion vector can be understood as the refined first motion vector. In this way, the i-th predicted value of the current block can be obtained based on at least one second motion vector obtained by refinement of at least one first motion vector.

[0376] In one example, when the decoding side improves all N first motion vectors, N second motion vectors are obtained, and N reference blocks are further obtained based on the N second motion vectors, and the i-th predicted value of the current block is further obtained based on the N reference blocks.

[0377] In another example, if the decoding side improves some of the N first motion vectors and does not improve the other first motion vectors, some reference blocks for the current block are obtained based on the second motion vectors obtained after improvement, and some reference blocks for the current block are obtained based on the unimproved first motion vectors, thereby obtaining a total of N reference blocks for the current block, and further obtaining the i-th predicted value for the current block based on these N reference blocks.

[0378] In the embodiment of the present application, the specific manner of improving the first motion vector to obtain the second motion vector includes, but is not limited to, the following several manners.

[0379] In Method 1, the motion vector differential is used to refine the first motion vector, and in this case, the above S102-B1 includes the following steps S102-B1-A1 to S102-B1-A3.

[0380] In step S102-B1-A1, motion vector difference information is determined.

[0381] In step S102-B1-A2, a first motion vector difference is obtained based on the motion vector difference information.

[0382] In step S102-B1-A3, at least one first motion vector among the N first motion vectors is improved based on the first motion vector differential to obtain at least one second motion vector.

[0383] In the method 1, when the decoding side decides to use a motion vector differential to improve at least one first motion vector among the N first motion vectors, the decoding side first determines motion vector differential information, then uses the motion vector differential information to obtain a first motion vector differential, and then uses the first motion vector differential to improve at least one first motion vector to obtain at least one second motion vector.

[0384] The embodiment of the present application does not limit the specific content of the motion vector difference information, and the information may be any information for deriving the first motion vector difference.

[0385] In some embodiments, the motion vector differential information includes a direction index mmvd_direction_idx and a distance index mmvd_distance_idx. In one example, the direction index mmvd_direction_idx and the distance index mmvd_distance_idx are predetermined values ​​(or default values). In another example, the encoding side writes the determined direction index mmvd_direction_idx and distance index mmvd_distance_idx into a bitstream, and the decoding side decodes the bitstream to obtain the direction index mmvd_direction_idx and the distance index mmvd_distance_idx. Correspondingly, in the above S102-B1-A2, obtaining a first motion vector differential based on the motion vector differential information includes obtaining the first motion vector differential based on the direction index mmvd_direction_idx and the distance index mmvd_distance_idx.

[0386] Illustratively, directions of motion vector differences in embodiments of the present application include a single horizontal direction, a single vertical direction, directions such as 45 degrees up and left, 45 degrees down and left, 45 degrees up and right, and 45 degrees down and right.

[0387] In some embodiments, before decoding the bitstream to obtain a direction index (mmvd_direction_idx) and a distance index (mmvd_distance_idx), the decoding side first determines whether the i-th prediction mode can be improved using a motion vector differential. Specifically, the bitstream is decoded to obtain first information, which is used to indicate whether the i-th prediction mode can be improved using a motion vector differential. If the first information indicates that the i-th prediction mode can be improved using a motion vector differential, the bitstream is decoded to obtain a direction index and a distance index.

[0388] In one example, a specific process for obtaining the first motion vector differential based on the direction index mmvd_direction_idx and the distance index mmvd_distance_idx may be as follows: First, the decoding side obtains the distance information MmvdDistance based on the direction index mmvd_direction_idx in the above Table 7, obtains the direction information MmvdSign based on the direction index mmvd_direction_idx in the above Table 8, and further obtains the first motion vector differential based on MmvdDistance and MmvdSign.

[0389] In some embodiments, refining the first motion vector includes refining two components x and y, respectively. Illustratively, the decoding side obtains the first motion vector differential based on the following formula (1):

[0390] MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][0] MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][1] (1) Here, MmvdOffset[x0][y0][0] represents the first motion vector differential in the x direction, MmvdOffset[x0][y0][1] represents the first motion vector differential in the y direction, MmvdDistance[x0][y0] represents the distance in the x direction, MmvdDistance[x0][y0] represents the distance in the y direction, MmvdSign[x0][y0][0] is the Mmvd sign in the x direction, and MmvdSign[x0][y0][1] is the Mmvd sign in the y direction.

[0391] After determining the first motion vector differential based on the above steps, the decoding side improves at least one of the N first motion vectors based on the first motion vector differential to obtain at least one second motion vector.

[0392] In the embodiment of the present application, the implementation of the above S102-B1-A3 includes, but is not limited to, the following several ways:

[0393] In Method 1, the decoding side improves all of the N first motion vectors based on the first motion vector differentials to obtain N second motion vectors.

[0394] For example, assuming that the N first motion vectors include a first first motion vector and a second first motion vector, the first motion vector differential is used to improve the first first motion vector to obtain a first second motion vector, and the first motion vector differential is used to improve the second first motion vector to obtain a second second motion vector.

[0395] The embodiments of the present application do not limit the specific manner of using the first motion vector differential to improve the first motion vector. For example, the first motion vector differential can be directly used to compensate the first motion vector to obtain a second motion vector. As another example, a predetermined process can be performed on the first motion vector differential, and the processed first motion vector differential can be used to compensate the first motion vector to obtain a second motion vector.

[0396] In one example, the first motion vector differential and the first motion vector are added to obtain the second motion vector.

[0397] For example, the N first motion vectors include a first first motion vector mvL0 and a second first motion vector mvL1, and the decoding side obtains the second motion vector using the following equation (2).

[0398] mvL0'=mvL0+MmvdOffset mvL1'=mvL1+MmvdOffset (2) Here, MmvdOffset is the first motion vector differential, mvL0' is the first second motion vector obtained by improving the first first motion vector mvL0, and mvL1' is the second second motion vector obtained by improving the second first motion vector mvL1.

[0399] In some embodiments, the refinement for mvL0 and mvL1 is performed on the x and y components, respectively, and then the second motion vector can be obtained according to the following equation (3):

[0400] mvL0'[0]=mvL0[0]+MmvdOffset[0] mvL0'[1]=mvL0[1]+MmvdOffset[1] mvL1'[0]=mvL1[0]+MmvdOffset[0] mvL1'[1]=mvL1[1]+MmvdOffset[1] (3) where 0 represents the x direction, 1 represents the y direction, MmvdOffset[0] is the first motion vector differential in the x direction, and MmvdOffset[1] is the first motion vector differential in the y direction. As shown in the above equation (3), the first motion vector difference in the x direction MmvdOffset[0] is used to refine the first first motion vector mvL0[0] in the x direction to obtain the first second motion vector mvL0'[0]; the first motion vector difference in the y direction MmvdOffset[1] is used to refine the first first motion vector mvL0[1] in the y direction to obtain the first second motion vector mvL0'[1]; the first motion vector difference in the x direction MmvdOffset[0] is used to refine the second first motion vector mvL1[0] in the x direction to obtain the second second motion vector mvL1'[0]; and the first motion vector difference in the y direction MmvdOffset[1] is used to refine the second first motion vector mvL1[1] in the y direction to obtain the second second motion vector mvL1'[1].

[0401] In Method 2, the decoding side uses the first motion vector differential to improve some of the N first motion vectors, and uses the second motion vector differential to improve other of the N first motion vectors. At this time, the above S102-B1-A3 includes the following steps S102-B1-A3-11 to S102-B1-A3-13.

[0402] In step S102-B1-A3-11, the first first motion vector is improved based on the first motion vector difference to obtain the first second motion vector.

[0403] In step S102-B1-A3-12, a second motion vector differential is determined based on the first motion vector differential.

[0404] In step S102-B1-A3-13, the second first motion vector is improved based on the second motion vector difference to obtain a second second motion vector.

[0405] In the method 2, for ease of explanation, it is assumed that the N first motion vectors include the first first motion vector and the second first motion vector. In this way, the decoding side calculates the first motion vector difference as follows: Among the N first motion vectors The first motion vector is directly compensated to obtain the first second motion vector. At the same time, a second motion vector differential is obtained based on the first motion vector differential, and the second first motion vector is then compensated using the second motion vector differential to obtain the second second motion vector. Note that in Method 2, the N first motion vectors include two first motion vectors, but the specific number of N first motion vectors is not limited in the embodiments of the present application. That is, when N is a positive integer greater than 2, Method 2 can also be used to improve the N first motion vectors.

[0406] In Method 2, the specific method of improving the first first motion vector based on the first motion vector differential to obtain the first second motion vector can refer to the relevant description of Method 1 above. For example, the first motion vector differential and the first first motion vector are added to obtain the first second motion vector. Specifically, the specific description of Equation (2) and Equation (3) above can be referred to, and will not be repeated here.

[0407] The specific process of determining the second motion vector differential based on the first motion vector differential in S102-B1-A3-12 will be described below.

[0408] In one possible implementation, the first motion vector differential is processed based on the deviation between the first first motion vector and the second first motion vector to obtain a second motion vector differential, for example, the deviation between the first first motion vector and the second first motion vector and the first motion vector differential are added to obtain a second motion vector differential.

[0409] In one possible implementation, the above S102-B1-A3-12 includes the following steps:

[0410] In step S102-B1-A3-121, the image sequence number of the first reference image corresponding to the first first motion vector, the image sequence number of the current image, and the image sequence number of the second reference image corresponding to the second first motion vector are determined.

[0411] In step S102-B1-A3-122, a second motion vector differential is determined based on the image sequence number of the first reference image, the image sequence number of the current image, the image sequence number of the second reference image, and the first motion vector differential.

[0412] As can be seen from the above, each of the N pieces of first motion information corresponds to one reference image and one motion vector. Therefore, for each piece of motion information, the reference image corresponding to the motion information is the reference image corresponding to the first motion vector corresponding to the motion information. For convenience of explanation, in the embodiments of the present application, the reference image corresponding to the first first motion vector is referred to as the first reference image, and the reference image corresponding to the second first motion vector is referred to as the second reference image.

[0413] In this implementation, the decoding side determines the image sequence number of the first reference image corresponding to the first first motion vector, the image sequence number of the current image, and the image sequence number of the second reference image corresponding to the second first motion vector, and further determines the second motion vector differential based on the image sequence number of the first reference image, the image sequence number of the current image, the image sequence number of the second reference image, and the first motion vector differential.

[0414] The embodiments of the present application do not limit the specific manner of determining the second motion vector differential based on the image sequence number of the first reference image, the image sequence number of the current image, the image sequence number of the second reference image, and the first motion vector differential.

[0415] In one example, a mapping is performed on the first motion vector differential based on the correlation between the image sequence number of the first reference image, the image sequence number of the current image, and the image sequence number of the second reference image to obtain a second motion vector differential.

[0416] In another example, the decoding side determines a first difference value between the image sequence number of the second reference image and the image sequence number of the current image, determines a second difference value between the image sequence number of the first reference image and the image sequence number of the current image, and obtains a second motion vector differential based on the first difference value, the second difference value, and the first motion vector differential.

[0417] For example, the difference between the first difference value and the second difference value is determined, and the difference value is added to the first motion vector differential to obtain the second motion vector differential.

[0418] As another example, the ratio between the first difference value and the second difference value is determined, and the product of the ratio and the first motion vector difference is determined as the second motion vector difference. Illustratively, the second motion vector difference is obtained based on the following equation (4):

[0419] MmvdOffset'=(poc1-pocC) / (poc0-pocC)*MmvdOffset (4) where MmvdOffset' is the second motion vector differential, poc1 is the image sequence number of the second reference image, pocC is the image sequence number of the current image, poc0 is the image sequence number of the first reference image, poc1-pocC is the first difference value, poc0-pocC is the second difference value, (poc1-pocC) / (poc0-pocC) is the ratio between the first difference value and the second difference value, and MmvdOffset is the first motion vector differential.

[0420] After the second motion vector differential is obtained through the above steps, the second first motion vector is improved based on the second motion vector differential to obtain a second second motion vector. Here, the method by which the decoding side improves the second first motion vector using the second motion vector differential is almost the same as the method by which the decoding side improves the first first motion vector using the first motion vector differential.

[0421] For example, the second motion vector difference and the second first motion vector are added to obtain the second second motion vector.

[0422] In this method, in addition to improving the N first motion vectors using the above method 1 or method 2, the decoding side can also improve the N first motion vectors using the following method 3.

[0423] In method 3, the decoding side improves some of the N first motion vectors using the first motion vector differentials. ,Ma Tuning searchThe first motion vectors of the other part of the N first motion vectors are improved using the above method. At this time, the above S102-B1-A3 includes the following steps S102-B1-A3-21 to S102-B1-A3-22.

[0424] In step S102-B1-A3-21, the first first motion vector is improved based on the first motion vector difference to obtain the first second motion vector.

[0425] In step S102-B1-A3-22, a matching search is performed within a preset search range of the reference image corresponding to the second first motion vector based on the second first motion vector, to obtain a second second motion vector.

[0426] In the method 3, for ease of explanation, it is assumed that the N first motion vectors include a first first motion vector and a second first motion vector. In this way, the decoding side calculates the first motion vector difference. Among the N first motion vectors The first motion vector is directly compensated to obtain the first second motion vector. At the same time, a matching search is performed within a preset search range of the reference image corresponding to the second first motion vector based on the second first motion vector to obtain the second second motion vector. For example, by searching a small range, the motion vector corresponding to the reference block that has the highest matching degree with the reference block on the first reference image corresponding to the first first motion vector is searched as the second second motion vector. 3 Although the N first motion vectors include two first motion vectors, the embodiment of the present application does not limit the specific number of the N first motion vectors. That is, when N is a positive integer greater than 2, the N first motion vectors can also be improved using Method 3.

[0427] In Method 3, the specific method of improving the first first motion vector based on the first motion vector differential to obtain the first second motion vector can refer to the relevant description of Method 1 above. For example, the first motion vector differential and the first first motion vector are added to obtain the first second motion vector. Specifically, the specific description of Equation (2) and Equation (3) above can be referred to, and will not be repeated here.

[0428] Below, we will describe the specific process of performing a matching search within the preset search range of the reference image corresponding to the second first motion vector based on the second first motion vector in S102-B1-A3-22, and obtaining the second second motion vector.

[0429] As can be seen from the above, the reference image corresponding to the second first motion vector is marked as the second reference image, and in this way, the decoding side performs a search within a predetermined search range of the second reference image based on the second first motion vector, obtains multiple reference blocks, selects one reference block from these reference blocks, and determines the motion vector corresponding to the reference block as the second second motion vector.

[0430] In the embodiment of the present application, a specific method for performing a matching search within a predetermined search range of a reference image corresponding to the second first motion vector based on the second first motion vector and obtaining the second second motion vector is not limited.

[0431] In some embodiments, the decoding side performs a full block search, that is, searches for multiple reference blocks corresponding to the current block in the second reference image, selects one reference block from these reference blocks, and determines the motion vector corresponding to the reference block as the second second motion vector. For example, based on the first second motion vector, a reference block for the current block is determined in the first reference image (i.e., the reference image corresponding to the first first motion vector), which is denoted as reference block 1, and matches the multiple reference blocks searched in the second reference image with reference block 1 respectively to obtain the reference block with the highest matching degree (or the lowest matching cost), and determines the motion vector corresponding to the reference block as the second second motion vector.

[0432] In some embodiments, the decoding side can perform sub-block search, that is, can independently improve a sub-block within the current block, in which case the above S102-B1-A3-22 includes the following steps:

[0433] In step S102-B1-A3-221, the current block is divided into at least one first sub-block.

[0434] In step S102-B1-A3-222, for any one of the at least one first sub-blocks, a matching search is performed within a preset search range of the reference image corresponding to the second first motion vector, and a second second motion vector corresponding to the first sub-block is obtained.

[0435] In this embodiment, when the at least one first sub-block includes one sub-block, and the size of the first sub-block is the same as the size of the current block, the decoding side performs a full block search using the current block as a unit. In this embodiment, when the size of the first sub-block is smaller than the size of the current block, the decoding side searches a partial area within the current block to obtain an improved second motion vector for the partial area, and then performs prediction optimization on the image of the partial area based on the improved second motion vector to improve the image decoding effect of the partial area.

[0436] The embodiments of the present application do not limit the size and shape of the at least one first sub-block divided from the current block. For example, the size of the first sub-block may be, for example, 4x4, 8x8, or 16x16. In some embodiments, the decoding side may divide the entire current block into one or more first sub-blocks. That is, there is no undivided region in the current block. In some embodiments, the decoding side divides a partial region of the current block into at least one first sub-block, and does not divide other partial regions.

[0437] For each first sub-block in at least one first sub-block obtained by the above division, the decoding side Tema Tuning search and execute the second sub-block corresponding to the first sub-block. No. 2 The method of obtaining the motion vector is the same. For convenience of explanation, one first sub-block will be taken as an example.

[0438] The embodiments of the present application do not limit the specific manner of performing a matching search within the preset search range of the reference image corresponding to the second first motion vector in the above S102-B1-A3-222, and obtaining the second second motion vector corresponding to the first sub-block.

[0439] In some embodiments, the decoding side uses the second first motion vector as a search starting point and performs a matching search within a predetermined search range of the second reference image to obtain a plurality of first reference sub-blocks corresponding to the first sub-block. Furthermore, the decoding side selects one first reference sub-block from the plurality of first reference sub-blocks and determines the motion vector corresponding to the first reference sub-block as the second second motion vector corresponding to the first sub-block. For example, the method of selecting one first reference sub-block from the plurality of first reference sub-blocks may be to match the plurality of first reference sub-blocks to obtain a first reference sub-block with the smallest matching cost, and determine the motion vector corresponding to the first reference sub-block as the second second motion vector corresponding to the first sub-block.

[0440] In some embodiments, the above S102-B1-A3-222 includes the following steps S102-B1-A3-2221 to S102-B1-A3-2224.

[0441] In step S102-B1-A3-2221, a search starting point corresponding to the second first motion vector is determined, and using the search starting point as a starting point, a search is performed within a predetermined search range of the reference image corresponding to the second first motion vector to obtain a plurality of first reference sub-blocks corresponding to the first sub-blocks.

[0442] In step S102-B1-A3-2222, a second reference sub-block corresponding to the first sub-block is determined in the reference image corresponding to the first first motion vector based on the first second motion vector.

[0443] In step S102-B1-A3-2223, a matching cost between the plurality of first reference sub-blocks and the second reference sub-block is determined, and a first reference sub-block with the smallest matching cost is selected from the plurality of first reference sub-blocks.

[0444] In step S102-B1-A3-2224, a second motion vector corresponding to the first sub-block is obtained based on the motion vector corresponding to the first reference sub-block with the smallest matching cost.

[0445] As can be seen from the above, the reference image corresponding to the first first motion vector is referred to as the first reference image, and the reference image corresponding to the second first motion vector is referred to as the second reference image. Thus, as shown in FIG. 19, for each first sub-block in at least one first sub-block in the current block, the decoding side determines a plurality of first reference sub-blocks corresponding to the first sub-block in the second reference image based on the second first motion vector mvL1, and determines a second reference sub-block corresponding to the first sub-block in the first reference image based on the first second motion vector mvL0'. For each first reference sub-block in the plurality of first reference sub-blocks, the decoding side determines a matching cost between the first reference sub-block and the second reference sub-block, such as SAD, SATD, or SSE between the first reference sub-block and the second reference sub-block. In this way, the matching cost between different first reference sub-blocks and second reference sub-blocks is calculated, and based on the matching cost, the first reference sub-block with the smallest matching cost is selected from the multiple first reference sub-blocks, and based on the motion vector corresponding to the first reference sub-block with the smallest matching cost, a second second motion vector corresponding to the first sub-block can be obtained.

[0446] A specific process for determining the search starting point corresponding to the second first motion vector will be described below.

[0447] In Example 1, the second first motion vector is directly set as the search starting point. That is, the decoding side uses the second first motion vector as the search starting point to perform a search within a preset search range of the second reference image, and obtains multiple first reference sub-blocks corresponding to the first sub-blocks.

[0448] In Example 2, the second first motion vector is improved based on the first motion vector differential to obtain an improved second first motion vector, and the improved second first motion vector is used as a search starting point. Here, the method for improving the second first motion vector based on the first motion vector differential is substantially the same as the method for improving the first first motion vector based on the first motion vector differential. For example, the first motion vector differential and the second first motion vector are added to obtain the improved second first motion vector. In this way, the decoding side uses the improved second first motion vector as a search starting point to perform a search within a predetermined search range of the second reference image to obtain multiple first reference sub-blocks corresponding to the first sub-blocks.

[0449] In Example 3, a second motion vector differential is determined based on the first motion vector differential, and the second first motion vector is improved based on the second motion vector differential to obtain an improved second first motion vector, and the improved second first motion vector is used as a search starting point. The specific process of determining the second motion vector differential based on the first motion vector differential can refer to the description of the above embodiment, and will not be repeated here. The process of improving the second first motion vector based on the second motion vector differential can refer to the method of the above embodiment. For example, the second motion vector differential and the second first motion vector are added to obtain an improved second first motion vector. In this way, the decoding side uses the improved second first motion vector as a search starting point to perform a search within a predetermined search range of the second reference image to obtain multiple first reference sub-blocks corresponding to the first sub-blocks.

[0450] In the above S102-B1-A3-2222, determining the second reference sub-block corresponding to the first sub-block in the reference image corresponding to the first first motion vector based on the first second motion vector includes at least some of the following methods:

[0451] One method is to directly determine the second reference sub-block corresponding to the first sub-block in the first reference image based on the first second motion vector for each first sub-block, i.e., for first sub-block 1, search the first reference image for the second reference sub-block corresponding to first sub-block 1, and for first sub-block 2, search the first reference image for the second reference sub-block corresponding to first sub-block 2.

[0452] In another method, a reference block corresponding to the current block is first determined, and then second reference sub-blocks corresponding to each first sub-block in the reference block are determined. Specifically, the decoding side determines a reference block corresponding to the current block in the first reference image based on the first second motion vector. Next, for each first sub-block in the current block, second reference sub-blocks corresponding to each first sub-block in the reference block corresponding to the current block are determined.

[0453] In some embodiments, the decoding side can search for one first reference subblock, match the first reference subblock with a second reference subblock, and obtain a matching cost between the first reference subblock and the second reference subblock. In some embodiments, the decoding side can first search for multiple first reference subblocks, and then match each first reference subblock in the multiple first reference subblocks with a second reference subblock one by one, and obtain a matching cost between each first reference subblock and the second reference subblock.

[0454] The process of determining the matching cost between the first reference sub-block and the second reference sub-block will now be described.

[0455] The embodiment of the present application does not limit the specific process for determining the matching cost between the first reference sub-block and the second reference sub-block.

[0456] In some embodiments, the first reference sub-block and the second reference sub-block are image blocks in a decoded reference image, so that each of the first reference sub-block and the second reference sub-block sample of Sample Values Based on this, the decoding side calculates the sample of Sample Values and the corresponding sample of Sample Values , and the matching cost between the first reference sub-block and the second reference sub-block can be obtained. For example, sample 1 of Sample Values and in the second reference sub-block sample 1 of Sample Values and a difference value between the first reference sub-block and the second reference sub-block is determined. sample 2 Sample Values and in the second reference sub-block sample 2 Sample Values By analogy, the difference value between the first and second reference sub-blocks is determined. sample Determine the difference between sample The matching cost between the first reference sub-block and the second reference sub-block is obtained based on the difference value between the first and second reference sub-blocks. For example, sample The sum of the difference values ​​between the first and second reference sub-blocks is determined as the matching cost between the first and second reference sub-blocks.

[0457] In some embodiments, when the embodiments of the present application are applied to GPM, the matching cost between the first reference sub-block and the second reference sub-block can be determined by the following steps:

[0458] In step 1, the weight derivation mode of the current block is determined.

[0459] In step 2, the weight corresponding to the i-th prediction mode is determined based on the weight derivation mode.

[0460] In step 3, for any one of the multiple first reference subblocks, a matching cost between the first reference subblock and the second reference subblock is obtained based on the weight corresponding to the i-th prediction mode and the first reference subblock and the second reference subblock.

[0461] In this embodiment, the i-th prediction mode is an N-directional prediction mode, and a predicted block predicted by the N-directional prediction mode is used in GPM. That is, for a prediction block of the same size as the current block, a portion (i.e., a portion with a weight other than 0) is valid, and another portion (i.e., a portion with a weight of 0) is invalid. Therefore, when calculating a matching cost for N-directional prediction, a matching cost calculation method using a mask similar to GPM prediction can be used. Specifically, a weight derivation mode for the current block is determined. Here, for a specific method of determining the weight derivation mode for the current block, refer to the description of the above embodiment. For example, a combination of the weight derivation mode for the current block and K prediction modes is determined. Next, a weight corresponding to the i-th prediction mode can be determined based on the weight derivation mode for the current block. For example, if the weight derivation mode for the current block is mode 45 in FIG. 4 and the i-th prediction mode is the first prediction mode of the current block, a region with a weight other than 0 corresponding to the i-th prediction mode is a white and gray region. If the i-th prediction mode is the second prediction mode of the current block, the area corresponding to the i-th prediction mode where the weight is not 0 is a black and gray area. In this way, the matching cost between the first reference sub-block and the second reference sub-block can be obtained based on the weight corresponding to the i-th prediction mode.

[0462] Here, in the above step 3, specific methods for obtaining the matching cost between the first reference sub-block and the second reference sub-block based on the weight corresponding to the i-th prediction mode and the first reference sub-block and the second reference sub-block include, but are not limited to, the following methods:

[0463] In one possible implementation, based on the weight corresponding to the i-th prediction mode, only the areas in the first reference sub-block and the second reference sub-block where the weight corresponding to the i-th prediction mode is not 0 are processed to obtain a matching cost between the first reference sub-block and the second reference sub-block.

[0464] In another possible implementation, the decoding side obtains a first difference value based on the first reference sub-block and the second reference sub-block, and multiplies the first difference value by a weight corresponding to the i-th prediction mode to obtain a matching cost between the first reference sub-block and the second reference sub-block. Here, the method based on the first difference value between the first reference sub-block and the second reference sub-block may be to determine the difference value between the first reference sub-block and the second reference sub-block as the first difference value, or to determine the absolute value of the difference value between the first reference sub-block and the second reference sub-block as the first difference value, or may further determine the difference value between the first reference sub-block and the second reference sub-block and perform a preset process on the difference value to obtain the first difference value.

[0465] In one example, the matching cost SADwithMask1 between the first reference sub-block and the second reference sub-block may be obtained based on the following equation (10):

[0466] SADwithMask1=Σ(wValue[x][y]*abs(RefValue0[x][y]-RefValue1[x][y])) (10) where wValue is the weight corresponding to the i-th prediction mode at position (x,y), RefValue0[x][y] is the value corresponding to the second reference sub-block at position (x,y), and RefValue1[x][y] is the value corresponding to the first reference sub-block at position (x,y).

[0467] In some embodiments, the method for calculating wValue may be simplified, for example, to simplify the calculation, the wValue for calculating SADwithMask1 may be limited to 0 and 1.

[0468] Based on the above steps, the decoding side determines a matching cost between each first reference sub-block and a second reference sub-block in the multiple first reference sub-blocks, selects a first reference sub-block with the smallest matching cost from the multiple first reference sub-blocks based on the matching cost, and further obtains a second second motion vector corresponding to the first sub-block based on the motion vector corresponding to the first reference sub-block with the smallest matching cost.

[0469] In one possible implementation, the decoding side can directly determine the motion vector corresponding to the first reference sub-block with the smallest matching cost as the second motion vector corresponding to the first sub-block.

[0470] In another possible implementation, a sub-pixel search can be performed for higher accuracy, such as 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, 1 / 16 pixel accuracy, etc. Specifically, a sub-pixel search is performed around a first reference sub-block with the smallest matching cost to obtain a first sub-pixel motion vector, and a second motion vector corresponding to the first sub-block is obtained based on the first sub-pixel motion vector and the motion vector corresponding to the first reference sub-block with the smallest matching cost.

[0471] In the above, the first sub-block second The specific process of determining the second motion vector has been described. For each first sub-block in the current block, the above method is adopted to determine the second second motion vector corresponding to each first sub-block. Furthermore, based on the first second motion vector and the second second motion vector corresponding to each first sub-block, the predicted value of each first sub-block in the i-th prediction mode can be obtained.

[0472] The above describes a specific process of improving at least one first motion vector among the first motion vectors in N directions based on a motion vector differential in Method 1. As can be seen from the above, improving at least one first motion vector among the first motion vectors in N directions based on a motion vector differential can be performed using any one of Methods 1, 2, and 3 described above. In some embodiments, when a current block is divided into multiple first sub-blocks, some of the first sub-blocks can be improved using Method 1 described above, some of the first sub-blocks can be improved using Method 2 described above, and some of the first sub-blocks can be improved using Method 3 described above. That is, the first sub-blocks in the current block can be improved using at least two of Methods 1, 2, and 3 described above.

[0473] The decoding side may improve at least one of the N first motion vectors using the method shown in Method 1 above, or may improve the first motion vector using the following Method 2.

[0474] In Method 2, the above S102-B1 includes the following step S102-B1-B.

[0475] In step S102-B1-B, based on the at least one first motion vector, a matching search is performed within a preset search range of a reference image corresponding to each of the at least one first motion vector to obtain at least one second motion vector.

[0476] In the method 2, a search method is used for each of at least one of the N first motion vectors to obtain a second motion vector of the first motion vector to be improved.

[0477] For example, if the first and second primary motion vectors among N primary motion vectors are to be improved, the first and second primary motion vectors can be improved by bilateral matching. For example, the decoding side simultaneously searches the first and second primary motion vectors, obtains one reference block each time one MV is searched, and calculates the matching cost (e.g., SAD, SATD, or SSE) of the two reference blocks. When performing a bilateral matching search, the two MVs can be moved simultaneously, or one can be fixed and the other moved, or the latter can be fixed and the former moved. According to a preset search rule, the two MVs with the smallest matching cost are detected as the optimized MVs, i.e., the first and second secondary motion vectors.

[0478] The embodiments of the present application do not limit the specific manner of performing a matching search within a predetermined search range of a reference image corresponding to each of the at least one first motion vector based on the at least one first motion vector, and obtaining at least one second motion vector.

[0479] In some embodiments, the decoding side performs a full block search, that is, searches for multiple reference blocks corresponding to the current block in the reference image, selects one reference block from these reference blocks, and determines the motion vector corresponding to the reference block as a second motion vector.For example, based on the first first motion vector, search for reference block 1 in the first reference image, and based on the second first motion vector, search for reference block 2 in the second reference image, calculate the matching cost between reference block 1 and reference block 2, and then obtain two reference blocks with the highest matching degree (i.e., the smallest matching cost), and determine the motion vectors corresponding to these two reference blocks as the first second motion vector and the second second motion vector.

[0480] In some embodiments, the decoding side can perform sub-block search, that is, can independently improve a sub-block within the current block, in which case the above S102-B1-B includes the following steps:

[0481] In step S102-B1-B1, the current block is divided into at least one second sub-block.

[0482] In step S102-B1-B2, for any second sub-block among at least one second sub-block, a matching search is performed within a preset search range of a reference image corresponding to the first first motion vector to obtain a plurality of third reference sub-blocks corresponding to the second sub-block, and a matching search is performed within a preset search range of a reference image corresponding to the second first motion vector to obtain a plurality of fourth reference sub-blocks corresponding to the second sub-block.

[0483] In step S102-B1-B3, a matching cost between the plurality of third reference sub-blocks and the plurality of fourth reference sub-blocks is determined, and the third reference sub-block and the fourth reference sub-block with the smallest matching cost are selected from the plurality of third reference sub-blocks and the plurality of fourth reference sub-blocks.

[0484] In steps S102-B1-B4, the first second motion vector corresponding to the second sub-block is obtained based on the motion vector corresponding to the third reference sub-block with the smallest matching cost, and the second second motion vector corresponding to the second sub-block is obtained based on the motion vector corresponding to the fourth reference sub-block with the smallest matching cost.

[0485] In this embodiment, for convenience of explanation, it is assumed that at least one first motion vector to be improved among the N first motion vectors includes a first first motion vector and a second first motion vector. In this way, the decoding side improves the first first motion vector and the second first motion vector through a search method to obtain two second motion vectors. Note that in this embodiment, the first motion vector to be improved is described as two first motion vectors, but the embodiment of the present application does not limit the specific number of first motion vectors to be improved among the N first motion vectors. In other words, if the number of first motion vectors to be improved is more than two, multiple first motion vectors can also be improved using the search method of the embodiment of the present application.

[0486] In this embodiment, the current block is divided into at least one second sub-block, and each second sub-block can be improved independently.

[0487] In this embodiment, when the at least one second sub-block includes one second sub-block, and the size of the second sub-block is the same as the size of the current block, the decoding side performs a full block search using the current block as a unit. In this embodiment, when the size of the second sub-block is smaller than the size of the current block, the decoding side searches a partial area within the current block to obtain an improved second motion vector for the partial area, and then performs prediction optimization on the image of the partial area based on the improved second motion vector to improve the image decoding effect of the partial area.

[0488] The embodiments of the present application do not limit the size and shape of the at least one second sub-block divided from the current block. Exemplarily, the size of the second sub-block may be, for example, 4x4, 8x8, or 16x16. In some embodiments, the decoding side may divide the entire current block into one or more second sub-blocks. That is, there is no undivided region in the current block. In some embodiments, the decoding side divides a partial region of the current block into at least one second sub-block, and does not divide other partial regions.

[0489] For each second sub-block in at least one second sub-block obtained by the above division, the decoding side Tema Tuning search and obtain the second motion vector corresponding to the second sub-block in the same manner. For convenience of explanation, one second sub-block will be taken as an example.

[0490] Specifically, as shown in FIG. 20 , the decoding side uses the first first motion vector as a search starting point to search the first reference image and obtain multiple third reference sub-blocks corresponding to the second sub-block. Then, using the second first motion vector as a search starting point, the decoding side uses the second reference image to obtain multiple fourth reference sub-blocks corresponding to the second sub-block. Next, for these multiple third reference sub-blocks and multiple fourth reference sub-blocks, the matching cost between each pair of the third reference sub-block and the fourth reference sub-block is calculated, and the pair of the third reference sub-block and the fourth reference sub-block with the smallest matching cost is obtained. In this way, the first second motion vector corresponding to the second sub-block is obtained based on the motion vector corresponding to the third reference sub-block with the smallest matching cost, and the second second motion vector corresponding to the second sub-block is obtained based on the motion vector corresponding to the fourth reference sub-block with the smallest matching cost.

[0491] In some embodiments, the decoding side can search for one third reference subblock and one fourth reference subblock, and then match the third reference subblock with the fourth reference subblock to obtain a matching cost between the third reference subblock and the fourth reference subblock.In some embodiments, the decoding side can first search for a plurality of third reference subblocks and a plurality of fourth reference subblocks, and then match each third reference subblock in the plurality of third reference subblocks with each fourth reference subblock in the plurality of fourth reference subblocks one by one to obtain a matching cost between each third reference subblock and each fourth reference subblock.

[0492] The process of determining the matching cost between the third reference sub-block and the fourth reference sub-block will now be described.

[0493] The embodiment of the present application does not limit the specific process for determining the matching cost between the third reference sub-block and the fourth reference sub-block.

[0494] In some embodiments, the above S102-B1-B3 includes the following steps:

[0495] In steps S102-B1-B31, for any third reference subblock among the plurality of third reference subblocks and any fourth reference subblock among the plurality of fourth reference blocks, a first matching cost between the third reference subblock and the fourth reference subblock is determined.

[0496] In step S102-B1-B32, the matching cost between the third reference sub-block and the fourth reference sub-block is obtained based on the first matching cost.

[0497] In this embodiment, for any one of the third reference sub-blocks among the plurality of third reference sub-blocks and any one of the fourth reference sub-blocks among the plurality of fourth reference blocks, a first matching cost between the third reference sub-block and the fourth reference sub-block is determined, and a matching cost between the third reference sub-block and the fourth reference sub-block is determined based on the first matching cost. For example, the first matching cost is determined as the matching cost between the third reference sub-block and the fourth reference sub-block.

[0498] Here, the method for determining the first matching cost between the third reference sub-block and the fourth reference sub-block in S102-B1-B31 above may be the same as the method for determining the matching cost between the first reference sub-block and the second reference sub-block above.

[0499] In some embodiments, the third and fourth reference sub-blocks are image blocks in a decoded reference image, so that each of the third and fourth reference sub-blocks sample of Sample Values Based on this, the decoding side calculates each of the third reference sub-blocks. sample of Sample Values And, 4 The corresponding sample of Sample Values , and the matching cost between the third reference sub-block and the fourth reference sub-block can be obtained. For example, sample 1 of Sample Values and in the fourth reference subblock sample 1 of Sample Values and the difference value between the third reference sub-block and the sample 2 Sample Values and in the fourth reference subblock sample 2 Sample Values and by analogy, the difference value between the third and fourth reference sub-blocks is determined. sample Determine the difference between sampleThe matching cost between the third reference sub-block and the fourth reference sub-block is obtained based on the difference value between the third reference sub-block and the fourth reference sub-block. For example, sample The sum of the difference values ​​between the third and fourth reference sub-blocks is determined as the matching cost between the third and fourth reference sub-blocks.

[0500] In some embodiments, when the embodiments of the present application are applied to GPM, the matching cost between the third reference sub-block and the fourth reference sub-block can be determined by the following steps, that is, determining the first matching cost between the third reference sub-block and the fourth reference sub-block in the above S102-B1-B31 includes the following steps:

[0501] In step S102-B1-B311, the weight derivation mode of the current block is determined.

[0502] In step S102-B1-B312, a weight corresponding to the i-th prediction mode is determined based on the weight derivation mode.

[0503] In step S102-B1-B313, obtain a first matching cost between the third reference sub-block and the fourth reference sub-block according to the weight corresponding to the i-th prediction mode and the third reference sub-block and the fourth reference sub-block.

[0504] In this embodiment, the i-th prediction mode is an N-directional prediction mode, and a predicted block predicted by the N-directional prediction mode is used in GPM, i.e., for a predicted block of the same size as the current block, a portion of it (i.e., a portion with a weight other than 0) is valid, and another portion (i.e., a portion with a weight of 0) is invalid. Therefore, when calculating a matching cost for N-directional prediction, a matching cost calculation method using a mask similar to GPM prediction can be used.

[0505] For specific methods of determining the weight derivation mode of the current block and determining the weight corresponding to the i-th prediction mode based on the weight derivation mode, please refer to the relevant descriptions of steps 1 and 2 above, and they will not be repeated here.

[0506] Specific methods for obtaining the first matching cost between the third reference sub-block and the fourth reference sub-block based on the weight corresponding to the i-th prediction mode in S102-B1-B313 above and the third reference sub-block and the fourth reference sub-block include, but are not limited to, the following methods:

[0507] In one possible implementation, based on the weight corresponding to the i-th prediction mode, only the areas in the third reference sub-block and the fourth reference sub-block where the weight corresponding to the i-th prediction mode is not 0 are processed to obtain a matching cost between the third reference sub-block and the fourth reference sub-block.

[0508] In another possible implementation, the decoding side obtains a second difference value based on the third reference sub-block and the fourth reference sub-block, and multiplies the second difference value by a weight corresponding to the i-th prediction mode to obtain a matching cost between the third reference sub-block and the fourth reference sub-block. Here, the method based on the second difference value between the third reference sub-block and the fourth reference sub-block may be to determine the difference value between the third reference sub-block and the fourth reference sub-block as the second difference value, or to determine the absolute value of the difference value between the third reference sub-block and the fourth reference sub-block as the second difference value, or may further be to determine the difference value between the third reference sub-block and the fourth reference sub-block and perform a preset process on the difference value to obtain the second difference value.

[0509] In one example, the first matching cost SADwithMask2 between the third reference sub-block and the fourth reference sub-block may be obtained based on the following equation (11):

[0510] SADwithMask2=Σ(wValue[x][y]*abs(RefValue0[x][y]-RefValue1[x][y])) ( 11 ) Here, wValue is the weight corresponding to the i-th prediction mode at position (x,y), RefValue0[x][y] is the value corresponding to the third reference sub-block at position (x,y), and RefValue1[x][y] is the value corresponding to the fourth reference sub-block at position (x,y).

[0511] In some embodiments, the method for calculating wValue may be simplified, for example, to simplify the calculation, the wValue for calculating SADwithMask1 may be limited to 0 and 1.

[0512] Based on the above steps, the decoding side determines a first matching cost between each third reference subblock in the plurality of third reference subblocks and each fourth reference subblock in the plurality of fourth reference subblocks, and obtains a matching cost between the third subblock and the fourth reference subblock based on the first matching cost.

[0513] In one example, the above S102-B1-B32 includes that the decoding side can directly determine the first matching cost as the matching cost between the third reference sub-block and the fourth reference sub-block.

[0514] In another example, the above S102-B1-B32 includes the following steps:

[0515] In step S102-B1-B321, a second matching cost between the template of the third reference sub-block and the template of the second sub-block is determined.

[0516] In step S102-B1-B322, a third matching cost between the template of the fourth reference sub-block and the template of the second sub-block is determined.

[0517] In step S102-B1-B323, obtain the matching cost between the third reference sub-block and the fourth reference sub-block according to the first matching cost, the second matching cost, and the third matching cost.

[0518] In this example, template matching is used to improve the above motion vector. For example, the decoding side uses template matching and the above-mentioned Noma Tuning search By using the method described above, the motion vector of the current block is improved to improve the accuracy of the motion vector of the current block. Furthermore, when the current block is predicted based on the high-accuracy motion vector, the prediction accuracy of the current block can be improved, and the decoding effect can be further improved.

[0519] The embodiments of the present application do not limit the size and position of the template used in the calculation, and in some embodiments, the template used in the calculation may be an upper template, a left template, or both an upper template and a left template.

[0520] As shown in FIG. 21, in template matching, the template region of the reference block is compared with the template region of the current block. For example, the template of the third reference sub-block is compared with the template of the second sub-block to obtain a second matching cost CostTm0 between the template of the third reference sub-block and the template of the second sub-block; the template of the fourth reference sub-block is compared with the template of the second sub-block to obtain a second matching cost CostTm0; 4 A third matching cost CostTm1 between the template of the reference sub-block and the template of the second sub-block is obtained.

[0521] Here, the method for determining the second matching cost between the template of the third reference sub-block and the template of the second sub-block may be to determine a third difference value between the template of the third reference sub-block and the template of the second sub-block. For example, the absolute value of the difference value between the template of the third reference sub-block and the template of the second sub-block may be determined as the third difference value between the template of the third reference sub-block and the template of the second sub-block. Next, a weight for the template may be determined, and the product of the weight for the template and the third difference value may be determined as the second matching cost CostTm0 between the template of the third reference sub-block and the template of the second sub-block. Similarly, the third matching cost between the template of the fourth reference sub-block and the template of the second sub-block may be determined by determining a fourth difference between the template of the fourth reference sub-block and the template of the second sub-block, for example, determining the absolute value of the difference between the template of the fourth reference sub-block and the template of the second sub-block as the fourth difference between the template of the fourth reference sub-block and the template of the second sub-block, then determining a weight for the template, and determining the product of the weight for the template and the fourth difference as the third matching cost CostTm1 between the template of the fourth reference sub-block and the template of the second sub-block. Here, the method for determining the weight of the template may refer to the description of the above embodiment, and will not be repeated here.

[0522] The decoding side determines a second matching cost between the template of the third reference sub-block and the template of the second sub-block, and a third matching cost between the template of the fourth reference sub-block and the template of the second sub-block, and then obtains a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, the second matching cost, and the third matching cost.

[0523] For example, the first matching cost, the second matching cost, and the third matching cost are added together to obtain the matching cost between the third reference sub-block and the fourth reference sub-block.

[0524] Exemplarily, the matching cost between the third reference sub-block and the fourth reference sub-block may be determined based on the following Equation (11):

[0525] CostAll=CostBi+CostTm0+CostTm1 (11) Here, CostAll is the matching cost between the third reference subblock and the fourth reference subblock, CostBi is the first matching cost between the third reference subblock and the fourth reference subblock, CostTm0 is the second matching cost between the third reference subblock and the second subblock, and CostTm1 is the third matching cost between the fourth reference subblock and the second subblock.

[0526] In another example, a weighted sum is performed on the first matching cost, the second matching cost, and the third matching cost to obtain a matching cost between the third reference sub-block and the fourth reference sub-block. In this example, the first matching cost, the second matching cost, and the third matching cost each correspond to a weight. Optionally, the weights corresponding to the first matching cost, the second matching cost, and the third matching cost may be preset values.

[0527] As another example, a first weight corresponding to the first matching cost is determined, a second weight corresponding to the second matching cost and the third matching cost is determined, and the first matching cost, the second matching cost, and the third matching cost are weighted based on the first weight and the second weight to obtain the matching cost between the third reference sub-block and the fourth reference sub-block.

[0528] In this example, the first matching cost corresponds to a weight noted as a first weight, and the second matching cost and the third matching cost jointly correspond to a weight noted as a second weight. Optionally, the first weight and the second weight may be preset values.

[0529] In one example, the matching cost between the third reference sub-block and the fourth reference sub-block may be determined based on the following equation (11):

[0530] CostAll=CostBi*WeightBi+(CostTm0+CostTm1)*WeightTm (11) Here, WeightBi is the first weight and WeightTm is the second weight.

[0531] The decoding side determines matching points between each of the third and fourth reference subblocks in the plurality of third and fourth reference subblocks based on the above steps, and selects the third and fourth reference subblocks with the smallest matching cost from the plurality of third and fourth reference subblocks, and obtains a first second motion vector based on the motion vector of the third reference subblock with the smallest matching cost, and obtains a second second motion vector based on the motion vector of the fourth reference subblock with the smallest matching cost.

[0532] In one possible implementation, the decoding side can directly determine the motion vector corresponding to the third reference sub-block with the smallest matching cost as the first second motion vector corresponding to the second sub-block, and determine the motion vector corresponding to the fourth reference sub-block with the smallest matching cost as the second second motion vector corresponding to the second sub-block.

[0533] In another possible implementation, a fractional pixel search can be performed for higher accuracy, such as half-pixel accuracy, quarter-pixel accuracy, or sixteenth-pixel accuracy. Specifically, a fractional pixel search is performed around the third reference subblock with the smallest matching cost to obtain a second fractional pixel motion vector, and a first second motion vector corresponding to the second subblock is obtained based on the second fractional pixel motion vector and the motion vector corresponding to the third reference subblock with the smallest matching cost. Similarly, a fractional pixel search is performed around the fourth reference subblock with the smallest matching cost to obtain a third fractional pixel motion vector, and a second second motion vector corresponding to the second subblock is obtained based on the third fractional pixel motion vector and the motion vector corresponding to the fourth reference subblock with the smallest matching cost.

[0534] The above describes a specific process for determining two second motion vectors corresponding to one second sub-block. For each second sub-block in the current block, the above method is used to determine two second motion vectors corresponding to each second sub-block. Then, the predicted value of each second sub-block in the i-th prediction mode can be obtained based on the first second motion vector and the second second motion vector corresponding to each second sub-block.

[0535] The above describes a process of determining a predicted value of a current block in the i-th prediction mode when the i-th prediction mode among the K prediction modes is an N-directional prediction mode.

[0536] If the prediction modes other than the i-th prediction mode among the K prediction modes are also N-directional prediction modes, the predicted value of the current block in the other prediction modes can be determined by referring to the above steps.

[0537] If the K prediction modes further include a prediction mode other than the N-directional prediction mode, for example, a unidirectional prediction mode, the current block is predicted using the unidirectional prediction mode to obtain a predicted value of the current block in the unidirectional prediction mode.

[0538] In the embodiment of the present application, the manner of predicting the current block based on the K prediction modes and obtaining the predicted value of the current block includes, but is not limited to, the following cases:

[0539] In Case 1, when the decoding side predicts the current block based on K prediction modes, the weight derivation mode is not taken into consideration. For example, the decoding side predicts the current block using the K prediction modes of the current block, respectively, to obtain K predicted values ​​of the current block, and processes these K predicted values ​​to obtain the predicted value of the current block. Specifically, the predicted value of the current block is the average value, sum, or weighted sum of the i-th predicted value of the current block determined above and other predicted values.

[0540] In case 2, the decoding side predicts the current block according to the K prediction modes of the current block and the weight derivation mode of the current block, and at this time, the above S102-D includes the following steps:

[0541] In step S102-D1, the weight of the predicted value is determined based on the weight derivation mode.

[0542] Step S102-D2: based on the weight of the predicted value, weight the i-th predicted value and other predicted values ​​to obtain a predicted value of the current block.

[0543] In Case 2, when determining the predicted value of the current block, the weight derivation mode is taken into consideration, and the weight of the predicted value is determined based on the weight derivation mode. In this way, the decoding side can weight the i-th predicted value and other predicted values ​​based on the weight of the predicted value to obtain the predicted value of the current block.

[0544] The embodiments of the present application do not limit the specific process for determining the weight of the predicted value based on the weight derivation mode.

[0545] In some embodiments, when determining the weight of the predictor, the weight gradient parameter (also called the transition parameter) is not considered, and the weight of the predictor of the current block is directly determined based on the weight derivation mode of the current block.

[0546] In some embodiments, when determining the weight of the predictor, a weight gradient parameter is taken into consideration, and then the weight gradient parameter can be determined, and the weight of the predictor is further obtained based on the weight gradient parameter and the weight derivation mode of the current block.

[0547] Illustratively, the value of the weight gradient parameter blendingCoeff may be derived from the weight gradient index gpm_blending_idx.

[0548] In this way, the K predictors (e.g., the i-th predictor and other predictors) can be weighted based on the weights of the predictors determined above to obtain a predictor for the current block.

[0549] In some embodiments, the prediction process sample The corresponding weights of the above predicted values ​​are also sample In this case, when predicting the current block, the weights of the K prediction modes of the current block are used to predict the sample Predict A, sample Obtain K prediction values ​​of K prediction modes of the current block for A, and based on the weight derivation mode and weight gradient parameter of the current block: sample Determine the weight of the predicted value of A. Next, sample Weight these K predictors using the weights of the predictors in A, sample Get the predicted value of A. sample By performing the above steps for each sample We can obtain the predicted value of sample The predicted value of the current block is used to predict the current block. For example, if K=2, the first prediction mode is used to predict the current block. samplePredict A and sample Obtain the first predicted value of A and use the second prediction mode to sample Predict A and sample Obtain a second predicted value of A. sample Weighting the first predicted value and the second predicted value based on the weight of the predicted value corresponding to A, sample Obtain the predicted value of A.

[0550] In some embodiments, when K is greater than 2, weights of predictors corresponding to two of the K prediction modes of the current block may be determined based on the weight derivation mode of the current block, and weights of predictors corresponding to the other of the K prediction modes of the current block may be preset values. For example, when K=3, first weights of predictors corresponding to the first and second prediction modes are derived based on the weight derivation mode, and the weight of predictor corresponding to the third prediction mode is a preset value. In some embodiments, when the total weight of predictors corresponding to the K prediction modes of the current block is constant, e.g., 8, weights of predictors corresponding to each of the K prediction modes of the current block may be determined based on a preset weight ratio. Assuming that the weight of the predictor corresponding to the third prediction mode accounts for 1 / 4 of the total weight of the entire predictor, it may be determined that the weight of the predictor for the third prediction mode is 2, and the remaining 3 / 4 of the total weight of the predictor is assigned to the first and second prediction modes. For example, if the weight of the prediction value corresponding to the first prediction mode derived based on the weight derivation mode of the current block is 3, the weight of the prediction value corresponding to the first prediction mode is determined to be (3 / 4)*3, and the weight of the prediction value corresponding to the second prediction mode is determined to be (3 / 4)*5.

[0551] According to the above method, the predicted value of the current block is determined, and at the same time, the bitstream is decoded to obtain the quantized coefficients of the current block, and the quantized coefficients of the current block are inverse quantized and inverse transformed to obtain the residual value of the current block, and the predicted value and the residual value of the current block are added to obtain the reconstructed value of the current block.

[0552] In the video decoding method provided in the embodiment of the present application, when a decoding side decodes a current block, K prediction modes of the current block are determined, and at least one prediction mode among the K prediction modes is a multi-directional prediction mode (e.g., a bidirectional prediction mode). In this way, when the current block is predicted using the K prediction modes, the prediction accuracy of the current block can be improved, and the video decoding effect can be further improved.

[0553] While the video decoding method of the present invention has been described above using the decoding side as an example, the following description will be given using the encoding side as an example.

[0554] 22 is a schematic flowchart of a video encoding method according to an embodiment of the present application. The embodiment of the present application is applied to the video encoder shown in Figures 1 and 2. As shown in Figure 22, the method of the embodiment of the present application includes S201 to S202.

[0555] In step S201, K prediction modes for the current block are determined.

[0556] Here, at least one prediction mode among the K prediction modes is an N-directional prediction mode, and both K and N are positive integers greater than 1.

[0557] As can be seen from the above, in the embodiment of the present application, K prediction modes jointly generate one prediction block, and this prediction block acts on the current block, that is, the current block is predicted based on the K prediction modes to obtain K prediction values, and weighted processing is performed on the K prediction values ​​to obtain the prediction value of the current block.

[0558] In other words, when encoding the current block, the encoding side needs to determine multiple candidate prediction modes, select K prediction modes from the multiple candidate prediction modes, and predict the current block using the K prediction modes to obtain a predicted value of the current block.

[0559] In some embodiments, before determining the K prediction modes of the current block, the encoding side first needs to determine whether the current block performs weighted prediction processing using K different prediction modes. If the encoding side determines that the current block performs weighted prediction processing using K different prediction modes, it performs the above-mentioned step S201 of determining the K prediction modes of the current block. If the encoding side determines that the current block does not perform weighted prediction processing using K different prediction modes, it skips the above-mentioned step S201.

[0560] In one possible implementation, the encoding side may determine the prediction mode parameter of the current block to determine whether the current block performs weighted prediction processing using K different prediction modes. Specifically, please refer to the above description of the decoding side, and it will not be repeated here.

[0561] In some embodiments, the embodiments of the present application may also impose conditional restrictions on the current block using GPM mode or AWP mode, i.e., if it is determined that the current block meets a preset condition, it may determine that the current block performs weighted prediction using K prediction modes, and further determine the K prediction modes for the current block.

[0562] For example, when the GPM mode or the AWP mode is applied, the size of the current block can be limited.

[0563] In an embodiment of the present application, the size parameters of the current block may include the height and width of the current block, and therefore the encoder can determine whether the current block uses GPM mode or AWP mode based on the height and width of the current block.

[0564] Furthermore, in the examples of the present application, sample The parameter restrictions can also be used to restrict the size of blocks that can use GPM mode or AWP mode.

[0565] That is, in this application, the current block can use the GPM mode or the AWP mode only if the size parameter of the current block meets the size requirement.

[0566] For example, in the present application, a frame-level flag may be used to determine whether the present application is applied to a current frame to be encoded. For example, it may be set to apply the present application to intraframes (e.g., I frames) and not to apply the present application to interframes (e.g., B frames, P frames). Alternatively, it may be set not to apply the present application to intraframes and to use the present application to interframes. Alternatively, it may be set to apply the present application to some interframes and not to apply the present application to some interframes. Since intraframes can also use intraprediction, the present application may also be applied to interframes.

[0567] In some embodiments, a flag at the frame level or below may also be used to determine whether the present application applies to the current block.

[0568] If the encoding side determines that the current block is to be predicted using K prediction modes based on the above method, it determines K prediction modes for the current block.

[0569] In an embodiment of the present application, in order to improve the prediction effect of the K prediction modes for the current block, at least one prediction mode among the K prediction modes is an N-directional prediction mode, where N is a positive integer greater than 1. Therefore, the N-directional prediction mode can also be understood as a multi-directional prediction mode.

[0570] It should be noted that the N-directional prediction mode in the embodiments of the present application can also be understood as an N-reference image prediction mode, i.e., a mode of prediction based on N reference images. For example, if the i-th prediction mode of the K prediction modes of the current block is an N-directional prediction mode and N=2, it will include two reference images. In this way, one predicted value of the current block is obtained based on the first reference image, and another predicted value of the current block is obtained based on the second reference image, and these two predicted values ​​are processed (for example, added or weighted added) to obtain a predicted value of the current block in the i-th prediction mode. The N reference images may be N coded images ahead of the current image or N coded images behind the current image, or the N reference images may include at least one coded image ahead of the current image and at least one coded image behind the current image.

[0571] For example, the K prediction modes of the current block include a first prediction mode and a second prediction mode.

[0572] In one example, the first prediction mode is an N-directional prediction mode, for example, the first prediction mode is a bidirectional prediction mode (i.e., a mode that predicts based on two reference images), a three-directional prediction mode (i.e., a mode that predicts based on three reference images), or a four-directional prediction mode (i.e., a mode that predicts based on four reference images), and the second prediction mode is a unidirectional prediction mode.

[0573] In another example, the second prediction mode is an N-directional prediction mode, for example, the first prediction mode is a bidirectional prediction mode (i.e., a mode that predicts based on two reference images), a three-directional prediction mode (i.e., a mode that predicts based on three reference images), or a four-directional prediction mode (i.e., a mode that predicts based on four reference images), and the first prediction mode is a unidirectional prediction mode.

[0574] In another example, the first prediction mode and the second prediction mode are both N-directional prediction modes, for example, both the first prediction mode and the second prediction mode are bidirectional prediction modes, tridirectional prediction modes, or quaddirectional prediction modes. Note that when the first prediction mode and the second prediction mode are both N-directional prediction modes, the number of reference images and / or selection methods corresponding to the first prediction mode and the second prediction mode may be the same or different, and this is not limited in the embodiments of the present application. For example, the first prediction mode is a bidirectional prediction mode, i.e., corresponding to two reference images, and the second prediction mode is a tridirectional prediction mode, i.e., corresponding to three reference images. As another example, the first prediction mode and the second prediction mode are both bidirectional prediction modes, but the reference image selection methods corresponding to the first prediction mode and the second prediction mode may be the same or different. For example, the reference images corresponding to the first prediction mode are two coded images ahead of the current image, and the reference images corresponding to the second prediction mode are two coded images behind the current image, or the reference images corresponding to the first prediction mode and the second prediction mode are both two coded images ahead of the current image, etc.

[0575] A specific method for determining the K prediction modes of the current block on the encoding side will now be described.

[0576] In case 1, when the encoding side presets the current block, the predicted value of the current block is obtained based on K prediction modes without considering the weight derivation mode. At this time, the above S201 includes the following steps S201-A1 and S201-A2:

[0577] In step S201-A1, a candidate prediction mode list is determined, and the candidate prediction mode list includes a plurality of candidate prediction modes.

[0578] In step S201-A2, K prediction modes are selected from the candidate prediction mode list.

[0579] In Case 1, the encoding side first determines a candidate prediction mode list and then selects K prediction modes from the constructed candidate prediction mode list. For example, the encoding side combines each of the K candidate prediction modes from the plurality of candidate prediction modes to obtain a plurality of combinations. For each combination, the encoding side predicts a template for the current block using the K candidate prediction modes included in the combination to obtain a template prediction cost corresponding to the combination. In this way, the combination with the smallest template prediction cost is selected from the plurality of combinations based on the template prediction cost, and the K candidate prediction modes included in the combination with the smallest template prediction cost are set as the K prediction modes for the current block.

[0580] In case 2, when the encoding side presets the current block, the predicted value of the current block is obtained according to the weight derivation mode and K prediction modes. At this time, the above S201 includes the following steps:

[0581] In step S201-B1, M candidate weight derivation modes are determined.

[0582] In step S201-B2, a candidate prediction mode list is determined.

[0583] In step S201-B3, K prediction modes for the current block are determined based on the M candidate weight derivation modes and the candidate prediction mode list.

[0584] Here, for a specific manner of determining the M candidate weight derivation modes, please refer to the description of the above embodiment, and it will not be repeated here. In one possible implementation form, several weight derivation modes in AWP or GPM can be screened as M candidate weight derivation modes.

[0585] The process of determining the candidate prediction mode list is described below.

[0586] In addition, the candidate prediction mode list in the embodiment of the present application includes N-directional prediction modes, such as bidirectional prediction modes, three-directional prediction modes, or four-directional prediction modes, thereby ensuring that at least one prediction mode among the K prediction modes of the current block determined based on the candidate prediction mode list is an N-directional prediction mode.

[0587] In some embodiments, the candidate prediction mode list determination process is independent of the M candidate weight derivation modes. That is, it can be understood that the M candidate weight derivation modes correspond to one candidate prediction mode list, thereby reducing the complexity of determining the candidate prediction mode list and further improving coding efficiency. Note that in this embodiment, since the candidate prediction mode list is independent of the M candidate weight derivation modes, there is no strict distinction in order during execution between S201-B1 and S201-B2. That is, S201-B1 may be executed after S201-B2, before S201-B2, or simultaneously with S201-B2, but the embodiments of the present application are not limited thereto.

[0588] In some embodiments, for each first candidate weight derivation mode of the M candidate weight derivation modes, a candidate prediction mode list corresponding to the first candidate weight derivation mode is determined.

[0589] In one example, the first candidate weight derivation mode is one of the M candidate weight derivation modes. That is, in this example, at least one candidate prediction mode list needs to be determined for each of the M candidate weight derivation modes. As can be seen from the above, one weight derivation mode corresponds to K prediction modes, and the candidate prediction mode list is used to determine a prediction mode. Therefore, in one possible implementation of this example, one candidate prediction mode list is determined for at least one prediction mode of the K prediction modes corresponding to each candidate weight derivation mode of the M candidate weight derivation modes.

[0590] In another example, if the above-mentioned first candidate weight derivation mode is one type of candidate weight derivation mode among M candidate weight derivation modes, the embodiment of the present application needs to classify the M candidate weight derivation modes and construct at least one candidate prediction mode list for each type of candidate weight derivation mode.

[0591] In the embodiment of the present application, the method for determining the candidate prediction mode list corresponding to each first candidate weight derivation mode among the M candidate weight derivation modes is the same. For ease of explanation, the embodiment of the present application will be described taking as an example the determination of the candidate prediction mode list corresponding to one first candidate weight derivation mode.

[0592] A specific method for determining the candidate prediction mode list corresponding to the first candidate weight derivation mode will be described below.

[0593] In some embodiments, the first candidate weight derivation mode corresponds to one candidate prediction mode list.

[0594] In some embodiments, when each prediction mode among the K prediction modes corresponds to one candidate prediction mode list, for the i-th prediction mode among the K prediction modes, the encoding side determines the candidate prediction mode list for the i-th prediction mode, where i is a positive integer less than or equal to K.

[0595] The embodiments of the present application do not limit the specific types of candidate prediction modes included in the candidate prediction mode list of the ith prediction mode.

[0596] In some embodiments, when constructing a candidate prediction mode list for the i-th prediction mode, the following types of prediction modes are added to the candidate prediction mode list in order until the length of the list reaches a preset value (e.g., 3):

[0597] 1. A prediction mode in which the prediction angle is parallel to the dividing line of the first candidate weight derivation mode.

[0598] 2. A first candidate prediction mode determined based on the template of the current block, in some embodiments the first candidate prediction mode is also referred to as a TIMD derived prediction mode.

[0599] 3. Reorganization of the current block template sample The second candidate prediction mode determined based on the gradient of , in some embodiments, the second candidate prediction mode may also be referred to as a DIMD-derived prediction mode.

[0600] 4. Prediction modes of neighboring blocks of the current block.

[0601] 5. Prediction mode where the prediction angle is perpendicular to the dividing line of the first candidate weight derivation mode.

[0602] 6.PLANAR mode.

[0603] 7.N-way prediction mode.

[0604] After determining the candidate prediction mode list based on the above steps, the encoding side executes step S201-B3.

[0605] The process of determining K prediction modes based on the M candidate weight derivation modes and the candidate prediction mode list in S201-B3 above will be described below.

[0606] In an embodiment of the present application, the encoding side selects one candidate weight derivation mode from the M candidate weight derivation modes as the weight derivation mode of the current block, determines K prediction modes of the current block from at least one candidate prediction mode included in the candidate prediction mode list, and finally predicts the current block using the weight derivation mode of the current block and the K prediction modes of the current block to obtain a predicted value of the current block.

[0607] The weight derivation mode of the current block and the K prediction modes of the current block are jointly used to determine the predicted value of the current block.

[0608] The embodiment of the present application does not limit the specific manner in which the encoding side determines the weight derivation mode of the current block and the K prediction modes of the current block based on the M candidate weight derivation modes and the candidate prediction mode list.

[0609] In some embodiments, when the candidate prediction mode list corresponds to the K prediction modes of the current block, that is, when all of the K prediction modes of the current block are selected from the candidate prediction mode list, the encoding side combines M candidate weight derivation modes with the candidate prediction modes included in the candidate prediction mode list. For example, each of the M candidate weight derivation modes is combined with any K candidate prediction modes in the candidate prediction mode list to obtain multiple combinations, each combination including one candidate weight derivation mode and K candidate prediction modes. Next, the candidate weight derivation mode and the K candidate prediction modes included in each combination are used to predict a template of the current block (for example, an upper template of the current block, a left template of the current block, or both upper and left templates of the current block), determine a cost for each combination, and further determine one combination from the multiple combinations based on the cost. For example, the combination with the lowest cost is selected from multiple combinations, the candidate weight derivation mode included in the combination with the lowest cost is determined as the weight derivation mode for the current block, and the K prediction modes included in the combination with the lowest cost are determined as the K prediction modes for the current block.

[0610] In some embodiments, when the candidate prediction mode list is a candidate prediction mode list for a prediction mode among the K prediction modes of the current block (e.g., K=2), the candidate prediction mode list is a candidate prediction mode list for the first prediction mode. At this time, the encoding side determines a selectable prediction mode set corresponding to the second prediction mode. Next, for each of the M candidate weight derivation modes, the encoding side selects one candidate prediction mode from the candidate prediction mode list for the first prediction mode as a possible first prediction mode, and selects one prediction mode from the selectable prediction mode set corresponding to the second prediction mode as a possible second prediction mode, thereby obtaining a combination consisting of the candidate weight derivation mode, one possible first prediction mode, and one possible second prediction mode, and thus obtaining a plurality of combinations. Each combination includes one candidate weight derivation mode and two candidate prediction modes. Next, the candidate weight derivation mode and two candidate prediction modes included in each combination are used to predict a template of the current block, determine a cost for each combination, and further determine one combination from the plurality of combinations based on the cost. For example, the combination with the lowest cost is selected from multiple combinations, the candidate weight derivation mode included in the combination with the lowest cost is determined as the weight derivation mode for the current block, and the K prediction modes included in the combination with the lowest cost are determined as the K prediction modes for the current block.

[0611] Based on the above description, one weight derivation mode and K prediction modes can act on the current block as one combination. In order to save codewords and reduce encoding costs, in some embodiments, the weight derivation mode and K prediction modes corresponding to the current block are combined into one combination, i.e., a first combination, and a first index is used to indicate the first combination. Compared with separately indicating the weight derivation mode and the K prediction modes, the embodiments of the present application use fewer codewords and further reduce encoding costs.

[0612] In Scheme 2, the encoding side determines a candidate combination list, for example, the encoding side determines a list including X candidate combinations, each candidate combination including one weight derivation mode and K prediction modes, the encoding side finally selects a candidate combination, such as the first combination, and the encoding side writes a first index of the first combination into the bitstream.

[0613] The specific process of determining the candidate combination list based on the M candidate weight derivation modes and the candidate prediction mode list in the above S201-B32 will be described below.

[0614] The embodiment of the present application does not limit the specific manner of determining the candidate combination list based on the M candidate weight derivation modes and the candidate prediction mode list in the above S201-B32.

[0615] In some embodiments, the above S201-B32 includes the following steps S201-B321 and S201-B322.

[0616] In step S201-B321, T second combinations are obtained based on the M candidate weight derivation modes and the candidate prediction mode list.

[0617] In step S201-B322, a candidate combination list is obtained based on the T second combinations.

[0618] Here, any of the T second combinations includes one weight derivation mode and K prediction modes, and the weight derivation mode and the K prediction modes included in any two of the T second combinations are not exactly the same, and T is a positive integer greater than 1.

[0619] The implementation of obtaining the candidate combination list based on the T second combinations in the above S201-B322 includes, but is not limited to, the following several ways.

[0620] In Method 1, the T second combinations are rearranged according to a preset rule to obtain a candidate combination list.

[0621] In Method 2, for any of the T second combinations, a cost corresponding to the second combination when predicting a template of a current block using the weight derivation mode in that second combination and the K prediction modes is determined, and a candidate combination list is determined based on the cost corresponding to each second combination in the T second combinations.

[0622] In the second method, for each of the T second combinations, a template for the current block is predicted using the weight derivation mode and K prediction modes included in the second combination to obtain a predicted value of the template corresponding to the second combination. Specifically, for each of the T second combinations, a template for the current block is predicted using the K prediction modes included in the second combination to obtain K predicted values. Next, a template weight corresponding to the second combination is determined based on the weight derivation mode for the second combination, and the K predicted values ​​of the template are weighted based on the template weight to obtain a predicted value of the template corresponding to the second combination.

[0623] Since the template of the current block is a reconstructed region, the encoding side can obtain the reconstructed value of the template, and thus, for each second combination in the T second combinations, the cost corresponding to the second combination can be determined based on the template predicted value and the template reconstructed value corresponding to the second combination. Here, methods for determining the cost corresponding to the second combination include, but are not limited to, SAD, SATD, SEE, etc. Next, a candidate combination list is constructed based on the cost corresponding to each second combination in the T second combinations.

[0624] In some embodiments, a cost corresponding to each second combination can be determined using a fast cost calculation method. As can be seen from the above, the template prediction values ​​corresponding to the second combination include template prediction values ​​corresponding to each of the K prediction modes included in the second combination. In this case, the costs corresponding to each of the K prediction modes in the second combination can be determined based on the template prediction values ​​and template reconstruction values ​​corresponding to each of the K prediction modes in the second combination, and the cost corresponding to the second combination can be determined based on the costs corresponding to each of the K prediction modes in the second combination. For example, the sum of the costs corresponding to each of the K prediction modes in the second combination can be determined as the cost corresponding to the second combination.

[0625] According to the above method, a cost corresponding to each second combination in the T second combinations can be determined, and then a candidate combination list can be constructed based on the cost corresponding to each second combination in the T second combinations.

[0626] In Example 1, the T second combinations are sorted based on the costs corresponding to each of the T second combinations, and the sorted T second combinations are determined as a candidate combination list. In Example 1, the generated candidate combination list includes the T first candidate combinations.

[0627] In Example 2, C second combinations are selected from T second combinations based on the costs associated with the second combinations, and a list consisting of the C second combinations is determined as a candidate combination list. Optionally, the C second combinations are the C second combinations with the lowest costs among the T second combinations. For example, based on the costs associated with each second combination among the T second combinations, the C second combinations with the lowest costs are selected from the T second combinations to form the candidate combination list. In this case, the candidate combination list includes C candidate combinations. Optionally, the C candidate combinations in the candidate combination list are sorted in ascending order according to the magnitude of their costs, i.e., the costs associated with the C candidate combinations in the candidate combination list increase in order according to the sorting.

[0628] Based on the above steps, the encoding side determines a candidate combination list, selects a first combination corresponding to a first index from the candidate combination list, determines the weight derivation mode included in the first combination as the weight derivation mode of the current block, and determines the K prediction modes included in the first combination as the K prediction modes of the current block.

[0629] Based on the above steps, the encoding side determines K prediction modes for the current block, and then performs the following step S202.

[0630] In step S202, the current block is predicted based on the K prediction modes of the current block to obtain a predicted value of the current block.

[0631] In an embodiment of the present application, at least one of the K prediction modes of the current block determined by the encoding side is an N-directional prediction mode, and in this manner, when predicting the current block based on the K prediction modes, prediction accuracy can be improved.

[0632] The embodiment of the present application does not limit the specific types of the K prediction modes of the current block.

[0633] In some embodiments, all of the K prediction modes are inter prediction modes.

[0634] In some embodiments, some of the K prediction modes are intra prediction modes and some are inter prediction modes.

[0635] In some embodiments, the N-way prediction mode of the embodiments of the present application is an inter prediction mode, such as bidirectional motion information or multidirectional motion information.

[0636] In this embodiment, N-directional prediction modes are applied to the current block, and one predicted value of the current block is finally obtained. For example, the K prediction modes include a first prediction mode and a second prediction mode, where the first prediction mode is the N-directional prediction mode. The encoding side predicts the current block using the N-directional prediction mode to obtain a predicted value 1 of the current block, and predicts the current block using the second prediction mode to obtain a predicted value 2 of the current block, and then processes the predicted value 1 and the predicted value 2 to obtain a predicted value of the current block.

[0637] The embodiment of the present application does not limit the specific manner in which the encoding side predicts the current block based on the K prediction modes and obtains the predicted value of the current block.

[0638] In some embodiments, if the above N-way prediction mode is N-way motion information, the above S202 includes the following steps:

[0639] In step S202-A, for the i-th prediction mode among the K prediction modes, if the i-th prediction mode is an N-directional prediction mode, N pieces of first motion information corresponding to the i-th prediction mode are determined (i is a positive integer less than or equal to K).

[0640] In step S202-B, the i-th predicted value of the current block is obtained based on the N first motion information.

[0641] In step S202-C, the current block is predicted using prediction modes other than the i-th prediction mode among the K prediction modes to obtain other predicted values ​​of the current block.

[0642] In step S202-D, a prediction value for the current block is obtained based on the i-th prediction value and other prediction values.

[0643] In this embodiment, when the i-th prediction mode of the K prediction modes of the current block is an N-directional prediction mode, the encoding side first determines N pieces of first motion information corresponding to the i-th prediction mode, where the first mot...

Claims

1. 1. A video decoding method comprising: determining K prediction modes for a current block, where at least one prediction mode among the K prediction modes is an N-directional prediction mode, and both K and N are positive integers greater than 1; predicting the current block based on the K prediction modes to obtain a predicted value of the current block.

2. Predicting the current block based on the K prediction modes and obtaining a predicted value of the current block includes: For an i-th prediction mode among the K prediction modes, when the i-th prediction mode is the N-directional prediction mode, determining N pieces of first motion information corresponding to the i-th prediction mode, where i is a positive integer equal to or less than K; obtaining an i-th predicted value of the current block based on the N first motion information; predicting the current block using a prediction mode other than the i-th prediction mode among the K prediction modes to obtain another predicted value of the current block; obtaining a predicted value of the current block based on the i-th predicted value and the other predicted values. The video decoding method of claim 1 .

3. The first motion information includes first motion vectors, and obtaining an i-th predicted value of the current block based on the N first motion vectors includes: improving at least one first motion vector among the N first motion vectors to obtain at least one second motion vector; obtaining an i-th predicted value of the current block based on the at least one second motion vector; The video decoding method of claim 2.

4. Refining at least one first motion vector among the N first motion vectors to obtain at least one second motion vector includes: determining motion vector differential information; obtaining a first motion vector differential based on the motion vector differential information; and refining at least one first motion vector among the N first motion vectors based on the first motion vector differential to obtain the at least one second motion vector.

4. The video decoding method of claim 3.

5. The motion vector difference information includes a direction index and a distance index, and determining the motion vector difference information includes: decoding the bitstream to obtain a direction index and a distance index; Obtaining a first motion vector differential based on the motion vector differential information includes: obtaining the first motion vector differential based on the direction index and the distance index.

5. The video decoding method of claim 4.

6. Before decoding the bitstream to obtain a direction index and a distance index, the video decoding method further comprises: decoding a bitstream to obtain first information, the first information being used to indicate whether the i-th prediction mode is improved using a motion vector differential; Determining the motion vector differential information includes: if the i-th prediction mode is improved using the motion vector differential, decoding the bitstream to obtain the direction index and the distance index. The video decoding method of claim 5.

7. Refining at least one first motion vector among the N first motion vectors based on the first motion vector differential to obtain the at least one second motion vector includes: and refining all of the N first motion vectors based on the first motion vector differentials to obtain N second motion vectors.

5. The video decoding method of claim 4.

8. The N is 2, and the N first motion vectors include a first first motion vector and a second first motion vector, and improving at least one first motion vector among the N first motion vectors based on the first motion vector difference to obtain the at least one second motion vector includes: improving the first first motion vector based on the first motion vector differential to obtain a first second motion vector; determining a second motion vector differential based on the first motion vector differential; and refining the second first motion vector based on the second motion vector differential to obtain a second second motion vector.

5. The video decoding method of claim 4.

9. Refining the second first motion vector based on the second motion vector differential to obtain a second second motion vector includes: adding the second motion vector differential and the second first motion vector to obtain the second second motion vector. The video decoding method of claim 8.

10. The N is 2, and the N first motion vectors include a first first motion vector and a second first motion vector, and improving at least one first motion vector among the N first motion vectors based on the first motion vector difference to obtain the at least one second motion vector includes: improving the first first motion vector based on the first motion vector differential to obtain a first second motion vector; and performing a matching search within a preset search range of a reference image corresponding to the second first motion vector based on the second first motion vector to obtain a second second motion vector.

5. The video decoding method of claim 4.

11. performing a matching search within a preset search range of a reference image corresponding to the second first motion vector based on the second first motion vector to obtain a second second motion vector; Dividing the current block into at least one first sub-block; performing a matching search within a preset search range of a reference image corresponding to the second first motion vector for any one of the at least one first sub-blocks, to obtain a second second motion vector corresponding to the first sub-block; The video decoding method of claim 10.

12. performing a matching search within a preset search range of a reference image corresponding to the second first motion vector to obtain a second second motion vector corresponding to the first sub-block; determining a search start point corresponding to the second first motion vector, and performing a search within a preset search range of a reference image corresponding to the second first motion vector using the search start point as a starting point, to obtain a plurality of first reference sub-blocks corresponding to the first sub-blocks; determining a second reference sub-block corresponding to the first sub-block in a reference image corresponding to the first motion vector based on the first second motion vector; determining a matching cost between the plurality of first reference sub-blocks and the second reference sub-block; and selecting a first reference sub-block from the plurality of first reference sub-blocks that has the smallest matching cost; obtaining a second motion vector corresponding to the first sub-block based on the motion vector corresponding to the first reference sub-block with the smallest matching cost; The video decoding method of claim 11.

13. Determining a search start point corresponding to the second first motion vector includes: and setting the second first motion vector as the search starting point.

13. The video decoding method of claim 12.

14. Determining a search start point corresponding to the second first motion vector includes: improving the second first motion vector based on the first motion vector differential to obtain an improved second first motion vector; and setting the improved second first motion vector as the search starting point.

13. The video decoding method of claim 12.

15. Improving the second first motion vector based on the first motion vector differential to obtain an improved second first motion vector includes: adding the first motion vector differential and the second first motion vector to obtain the improved second first motion vector.

15. The video decoding method of claim 14.

16. Determining a search start point corresponding to the second first motion vector includes: determining a second motion vector differential based on the first motion vector differential; improving the second first motion vector based on the second motion vector differential to obtain an improved second first motion vector; and setting the improved second first motion vector as the search starting point.

13. The video decoding method of claim 12.

17. improving the second first motion vector based on the second motion vector differential to obtain an improved second first motion vector; adding the second motion vector differential and the second first motion vector to obtain the improved second first motion vector.

17. The video decoding method of claim 16.

18. Determining a matching cost between the plurality of first reference sub-blocks and the second reference sub-block includes: determining a weight derivation mode for the current block; determining a weight corresponding to the i-th prediction mode based on the weight derivation mode; For any one of the plurality of first reference sub-blocks, obtaining a matching cost between the first reference sub-block and the second reference sub-block based on a weight corresponding to the i-th prediction mode and the first reference sub-block and the second reference sub-block; 13. The video decoding method of claim 12.

19. obtaining a matching cost between the first reference sub-block and the second reference sub-block based on a weight corresponding to the i-th prediction mode and the first reference sub-block and the second reference sub-block, obtaining a first difference value based on the first reference sub-block and the second reference sub-block; multiplying the first matching difference value by a weight corresponding to the i-th prediction mode to obtain a matching cost between the first reference sub-block and the second reference sub-block; 20. The video decoding method of claim 18.

20. Obtaining a first difference value based on the first reference sub-block and the second reference sub-block includes: determining an absolute value of a difference between the first reference sub-block and the second reference sub-block as the first difference value; 20. The video decoding method of claim 19.

21. obtaining a second motion vector corresponding to the first sub-block based on the motion vector corresponding to the first reference sub-block with the smallest matching cost; performing a sub-pel search around the first reference sub-block with the smallest matching cost to obtain a first sub-pel motion vector; obtaining a second motion vector corresponding to the first sub-block based on the first sub-pixel motion vector and a motion vector corresponding to the first reference sub-block with the smallest matching cost; 13. The video decoding method of claim 12.

22. Determining a second motion vector differential based on the first motion vector differential includes: determining an image sequence number of a first reference image corresponding to the first first motion vector, an image sequence number of a current image, and an image sequence number of a second reference image corresponding to the second first motion vector; determining the second motion vector differential based on an image sequence number of the first reference image, an image sequence number of the current image, an image sequence number of the second reference image, and the first motion vector differential.

17. A video decoding method according to claim 8 or 16.

23. determining the second motion vector differential based on an image sequence number of the first reference image, an image sequence number of the current image, an image sequence number of the second reference image, and the first motion vector differential; determining a first difference value between the image sequence number of the second reference image and the image sequence number of the current image; determining a second difference value between the image sequence number of the first reference image and the image sequence number of the current image; obtaining the second motion vector differential based on the first difference value, the second difference value, and the first motion vector differential.

23. The video decoding method of claim 22.

24. Obtaining the second motion vector differential based on the first difference value, the second difference value, and the first motion vector differential includes: determining a ratio between the first difference value and the second difference value; determining the product of the ratio and the first motion vector differential as the second motion vector differential.

24. The video decoding method of claim 23.

25. Refining the first motion vector based on the first motion vector differential to obtain the second motion vector includes: adding the first motion vector differential and the first motion vector to obtain the second motion vector. A video decoding method according to claim 7, 8 or 10.

26. Refining at least one first motion vector among the N first motion vectors to obtain at least one second motion vector includes: performing a matching search within a preset search range of a reference image corresponding to each of the at least one first motion vector based on the at least one first motion vector, to obtain at least one second motion vector; 4. The video decoding method of claim 3.

27. The at least one first motion vector includes a first first motion vector and a second first motion vector, and performing a matching search within a preset search range of a reference image corresponding to each of the at least one first motion vector based on the at least one first motion vector to obtain at least one second motion vector; Dividing the current block into at least one second sub-block; For any second sub-block among the at least one second sub-block, perform a matching search within a preset search range of a reference image corresponding to the first first motion vector to obtain a plurality of third reference sub-blocks corresponding to the second sub-block; perform a matching search within a preset search range of a reference image corresponding to the second first motion vector to obtain a plurality of fourth reference sub-blocks corresponding to the second sub-block; determining a matching cost between the plurality of third reference sub-blocks and the plurality of fourth reference sub-blocks, and selecting a third reference sub-block and a fourth reference sub-block with the smallest matching cost from the plurality of third reference sub-blocks and the plurality of fourth reference sub-blocks; obtaining a first second motion vector corresponding to the second sub-block based on the motion vector corresponding to the third reference sub-block having the smallest matching cost, and obtaining a second second motion vector corresponding to the second sub-block based on the motion vector corresponding to the fourth reference sub-block having the smallest matching cost.

27. The video decoding method of claim 26.

28. Determining a matching cost between the plurality of third reference sub-blocks and the plurality of fourth reference sub-blocks includes: determining a first matching cost between a third reference subblock of the plurality of third reference subblocks and a fourth reference subblock of the plurality of fourth reference blocks; obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost; 28. The video decoding method of claim 27.

29. Determining a first matching cost between the third reference sub-block and the fourth reference sub-block includes: determining a weight derivation mode for the current block; determining a weight corresponding to the i-th prediction mode based on the weight derivation mode; obtaining a first matching cost between the third reference sub-block and the fourth reference sub-block based on a weight corresponding to the i-th prediction mode and the third reference sub-block and the fourth reference sub-block; 29. The video decoding method of claim 28.

30. obtaining a first matching cost between the third reference sub-block and the fourth reference sub-block based on a weight corresponding to the i-th prediction mode and the third reference sub-block and the fourth reference sub-block; obtaining a second difference value based on the third reference sub-block and the fourth reference sub-block; multiplying the second difference value by a weight corresponding to the i-th prediction mode to obtain the first matching cost; 30. The video decoding method of claim 29.

31. Obtaining a second difference value based on the third reference sub-block and the fourth reference sub-block includes: determining an absolute value of a difference between the third reference sub-block and the fourth reference sub-block as the second difference value; 31. The video decoding method of claim 30.

32. obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, determining the first matching cost as a matching cost between the third reference sub-block and the fourth reference sub-block; 29. The video decoding method of claim 28.

33. obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, determining a second matching cost between the template of the third reference sub-block and the template of the second sub-block; determining a third matching cost between the fourth reference sub-block template and the second sub-block template; obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, the second matching cost, and the third matching cost; 29. The video decoding method of claim 28.

34. obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, the second matching cost, and the third matching cost; adding the first matching cost, the second matching cost, and the third matching cost to obtain a matching cost between the third reference sub-block and the fourth reference sub-block.

34. The video decoding method of claim 33.

35. obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, the second matching cost, and the third matching cost; performing a weighted addition on the first matching cost, the second matching cost, and the third matching cost to obtain a matching cost between the third reference sub-block and the fourth reference sub-block; 34. The video decoding method of claim 33.

36. performing a weighted addition on the first matching cost, the second matching cost, and the third matching cost to obtain a matching cost between the third reference sub-block and the fourth reference sub-block; determining a first weight corresponding to the first matching cost; determining second weights corresponding to the second matching cost and the third matching cost; weighting the first matching cost, the second matching cost, and the third matching cost based on the first weight and the second weight to obtain a matching cost between the third reference sub-block and the fourth reference sub-block; 36. The video decoding method of claim 35.

37. obtaining a first second motion vector corresponding to the second sub-block based on the motion vector corresponding to the third reference sub-block with the smallest matching cost, and obtaining a second second motion vector corresponding to the second sub-block based on the motion vector corresponding to the fourth reference sub-block with the smallest matching cost; performing a sub-pel search around the third reference sub-block with the smallest matching cost to obtain a second sub-pel motion vector; obtaining a first second motion vector corresponding to the second sub-block based on the second sub-pixel motion vector and the motion vector corresponding to the third reference sub-block with the smallest matching cost; performing a sub-pel search around the fourth reference sub-block with the smallest matching cost to obtain a third sub-pel motion vector; obtaining a second motion vector corresponding to the second sub-block based on the third sub-pixel motion vector and the motion vector corresponding to the fourth reference sub-block having the smallest matching cost; 28. The video decoding method of claim 27.

38. Obtaining a predicted value of the current block based on the i-th predicted value and the other predicted values ​​includes: determining weights for the predicted values ​​based on the weight derivation mode; weighting the i-th predicted value and the other predicted values ​​based on the weights of the predicted values ​​to obtain a predicted value of the current block; 30. A video decoding method according to claim 18 or 29.

39. 1. A video encoding method comprising: determining K prediction modes for a current block, where at least one prediction mode among the K prediction modes is an N-directional prediction mode, and both K and N are positive integers greater than 1; predicting the current block based on the K prediction modes to obtain a predicted value of the current block.

40. Predicting the current block based on the K prediction modes and obtaining a predicted value of the current block includes: For an i-th prediction mode among the K prediction modes, when the i-th prediction mode is the N-directional prediction mode, determining N pieces of first motion information corresponding to the i-th prediction mode, where i is a positive integer equal to or less than K; obtaining an i-th predicted value of the current block based on the N first motion information; predicting the current block using a prediction mode other than the i-th prediction mode among the K prediction modes to obtain another predicted value of the current block; obtaining a predicted value of the current block based on the i-th predicted value and the other predicted values; 40. The video encoding method of claim 39.

41. The first motion information includes first motion vectors, and obtaining an i-th predicted value of the current block based on the N first motion vectors includes: improving at least one first motion vector among the N first motion vectors to obtain at least one second motion vector; obtaining an i-th predicted value of the current block based on the at least one second motion vector; 41. A video encoding method according to claim 40.

42. Refining at least one first motion vector among the N first motion vectors to obtain at least one second motion vector includes: determining motion vector differential information; obtaining a first motion vector differential based on the motion vector differential information; and refining at least one first motion vector among the N first motion vectors based on the first motion vector differential to obtain the at least one second motion vector.

42. The video encoding method of claim 41.

43. The motion vector difference information includes a direction index and a distance index, and determining the motion vector difference information includes: determining a direction index and a distance index; Obtaining a first motion vector differential based on the motion vector differential information includes: obtaining the first motion vector differential based on the direction index and the distance index.

43. The video encoding method of claim 42.

44. Before determining the direction index and the distance index, the video encoding method further comprises: determining first information, the first information being used to indicate whether the i-th prediction mode is improved using a motion vector differential; Determining the motion vector differential information includes: determining the direction index and the distance index if the motion vector differential is used to improve the i-th prediction mode; 44. A video encoding method according to claim 43.

45. Refining at least one first motion vector among the N first motion vectors based on the first motion vector differential to obtain the at least one second motion vector includes: and refining all of the N first motion vectors based on the first motion vector differentials to obtain N second motion vectors.

43. The video encoding method of claim 42.

46. The N is 2, and the N first motion vectors include a first first motion vector and a second first motion vector, and improving at least one first motion vector among the N first motion vectors based on the first motion vector difference to obtain the at least one second motion vector includes: improving the first first motion vector based on the first motion vector differential to obtain a first second motion vector; determining a second motion vector differential based on the first motion vector differential; and refining the second first motion vector based on the second motion vector differential to obtain a second second motion vector.

43. The video encoding method of claim 42.

47. Refining the second first motion vector based on the second motion vector differential to obtain a second second motion vector includes: adding the second motion vector differential and the second first motion vector to obtain the second second motion vector.

47. A video encoding method according to claim 46.

48. The N is 2, and the N first motion vectors include a first first motion vector and a second first motion vector, and improving at least one first motion vector among the N first motion vectors based on the first motion vector difference to obtain the at least one second motion vector includes: improving the first first motion vector based on the first motion vector differential to obtain a first second motion vector; and performing a matching search within a preset search range of a reference image corresponding to the second first motion vector based on the second first motion vector to obtain a second second motion vector.

43. The video encoding method of claim 42.

49. performing a matching search within a preset search range of a reference image corresponding to the second first motion vector based on the second first motion vector to obtain a second second motion vector; Dividing the current block into at least one first sub-block; performing a matching search within a preset search range of a reference image corresponding to the second first motion vector for any one of the at least one first sub-blocks, to obtain a second second motion vector corresponding to the first sub-block; 49. A video encoding method according to claim 48.

50. performing a matching search within a preset search range of a reference image corresponding to the second first motion vector to obtain a second second motion vector corresponding to the first sub-block; determining a search start point corresponding to the second first motion vector, and performing a search within a preset search range of a reference image corresponding to the second first motion vector using the search start point as a starting point, to obtain a plurality of first reference sub-blocks corresponding to the first sub-blocks; determining a second reference sub-block corresponding to the first sub-block in a reference image corresponding to the first motion vector based on the first second motion vector; determining a matching cost between the plurality of first reference sub-blocks and the second reference sub-block; and selecting a first reference sub-block from the plurality of first reference sub-blocks that has the smallest matching cost; obtaining a second motion vector corresponding to the first sub-block based on the motion vector corresponding to the first reference sub-block with the smallest matching cost; 50. A video encoding method according to claim 49.

51. Determining a search start point corresponding to the second first motion vector includes: and setting the second first motion vector as the search starting point.

51. A video encoding method according to claim 50.

52. Determining a search start point corresponding to the second first motion vector includes: improving the second first motion vector based on the first motion vector differential to obtain an improved second first motion vector; and setting the improved second first motion vector as the search starting point.

51. A video encoding method according to claim 50.

53. Improving the second first motion vector based on the first motion vector differential to obtain an improved second first motion vector includes: adding the first motion vector differential and the second first motion vector to obtain the improved second first motion vector.

53. A video encoding method according to claim 52.

54. Determining a search start point corresponding to the second first motion vector includes: determining a second motion vector differential based on the first motion vector differential; improving the second first motion vector based on the second motion vector differential to obtain an improved second first motion vector; and setting the improved second first motion vector as the search starting point.

51. A video encoding method according to claim 50.

55. improving the second first motion vector based on the second motion vector differential to obtain an improved second first motion vector; adding the second motion vector differential and the second first motion vector to obtain the improved second first motion vector.

55. A video encoding method according to claim 54.

56. Determining a matching cost between the plurality of first reference sub-blocks and the second reference sub-block includes: determining a weight derivation mode for the current block; determining a weight corresponding to the i-th prediction mode based on the weight derivation mode; For any one of the plurality of first reference sub-blocks, obtaining a matching cost between the first reference sub-block and the second reference sub-block based on a weight corresponding to the i-th prediction mode and the first reference sub-block and the second reference sub-block; 51. A video encoding method according to claim 50.

57. obtaining a matching cost between the first reference sub-block and the second reference sub-block based on a weight corresponding to the i-th prediction mode and the first reference sub-block and the second reference sub-block, obtaining a first difference value based on the first reference sub-block and the second reference sub-block; multiplying the first matching difference value by a weight corresponding to the i-th prediction mode to obtain a matching cost between the first reference sub-block and the second reference sub-block; 57. A video encoding method according to claim 56.

58. Obtaining a first difference value based on the first reference sub-block and the second reference sub-block includes: determining an absolute value of a difference between the first reference sub-block and the second reference sub-block as the first difference value; 58. A video encoding method according to claim 57.

59. obtaining a second motion vector corresponding to the first sub-block based on the motion vector corresponding to the first reference sub-block with the smallest matching cost; performing a sub-pel search around the first reference sub-block with the smallest matching cost to obtain a first sub-pel motion vector; obtaining a second motion vector corresponding to the first sub-block based on the first sub-pixel motion vector and a motion vector corresponding to the first reference sub-block with the smallest matching cost; 51. A video encoding method according to claim 50.

60. Determining a second motion vector differential based on the first motion vector differential includes: determining an image sequence number of a first reference image corresponding to the first first motion vector, an image sequence number of a current image, and an image sequence number of a second reference image corresponding to the second first motion vector; determining the second motion vector differential based on an image sequence number of the first reference image, an image sequence number of the current image, an image sequence number of the second reference image, and the first motion vector differential.

55. A video encoding method according to claim 46 or 54.

61. determining the second motion vector differential based on an image sequence number of the first reference image, an image sequence number of the current image, an image sequence number of the second reference image, and the first motion vector differential; determining a first difference value between the image sequence number of the second reference image and the image sequence number of the current image; determining a second difference value between the image sequence number of the first reference image and the image sequence number of the current image; obtaining the second motion vector differential based on the first difference value, the second difference value, and the first motion vector differential.

61. The video encoding method of claim 60.

62. Obtaining the second motion vector differential based on the first difference value, the second difference value, and the first motion vector differential includes: determining a ratio between the first difference value and the second difference value; determining the product of the ratio and the first motion vector differential as the second motion vector differential.

62. The video encoding method of claim 61.

63. Refining the first motion vector based on the first motion vector differential to obtain the second motion vector includes: adding the first motion vector differential and the first motion vector to obtain the second motion vector.

49. A video encoding method according to claim 45, 46 or 48.

64. Refining at least one first motion vector among the N first motion vectors to obtain at least one second motion vector includes: performing a matching search within a preset search range of a reference image corresponding to each of the at least one first motion vector based on the at least one first motion vector, to obtain at least one second motion vector; 42. The video encoding method of claim 41.

65. The at least one first motion vector includes a first first motion vector and a second first motion vector, and performing a matching search within a preset search range of a reference image corresponding to each of the at least one first motion vector based on the at least one first motion vector to obtain at least one second motion vector; Dividing the current block into at least one second sub-block; For any second sub-block among the at least one second sub-block, perform a matching search within a preset search range of a reference image corresponding to the first first motion vector to obtain a plurality of third reference sub-blocks corresponding to the second sub-block; perform a matching search within a preset search range of a reference image corresponding to the second first motion vector to obtain a plurality of fourth reference sub-blocks corresponding to the second sub-block; determining a matching cost between the plurality of third reference sub-blocks and the plurality of fourth reference sub-blocks, and selecting a third reference sub-block and a fourth reference sub-block with the smallest matching cost from the plurality of third reference sub-blocks and the plurality of fourth reference sub-blocks; obtaining a first second motion vector corresponding to the second sub-block based on the motion vector corresponding to the third reference sub-block having the smallest matching cost, and obtaining a second second motion vector corresponding to the second sub-block based on the motion vector corresponding to the fourth reference sub-block having the smallest matching cost.

65. A video encoding method according to claim 64.

66. Determining a matching cost between the plurality of third reference sub-blocks and the plurality of fourth reference sub-blocks includes: determining a first matching cost between a third reference subblock of the plurality of third reference subblocks and a fourth reference subblock of the plurality of fourth reference blocks; obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost; 66. A video encoding method according to claim 65.

67. Determining a first matching cost between the third reference sub-block and the fourth reference sub-block includes: determining a weight derivation mode for the current block; determining a weight corresponding to the i-th prediction mode based on the weight derivation mode; obtaining a first matching cost between the third reference sub-block and the fourth reference sub-block based on a weight corresponding to the i-th prediction mode and the third reference sub-block and the fourth reference sub-block; 67. A video encoding method according to claim 66.

68. obtaining a first matching cost between the third reference sub-block and the fourth reference sub-block based on a weight corresponding to the i-th prediction mode and the third reference sub-block and the fourth reference sub-block; obtaining a second difference value based on the third reference sub-block and the fourth reference sub-block; multiplying the second difference value by a weight corresponding to the i-th prediction mode to obtain the first matching cost; 68. A video encoding method according to claim 67.

69. Obtaining a second difference value based on the third reference sub-block and the fourth reference sub-block includes: determining an absolute value of a difference between the third reference sub-block and the fourth reference sub-block as the second difference value; 69. The video encoding method of claim 68.

70. obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, determining the first matching cost as a matching cost between the third reference sub-block and the fourth reference sub-block; 67. A video encoding method according to claim 66.

71. obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, determining a second matching cost between the template of the third reference sub-block and the template of the second sub-block; determining a third matching cost between the fourth reference sub-block template and the second sub-block template; obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, the second matching cost, and the third matching cost; 67. A video encoding method according to claim 66.

72. obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, the second matching cost, and the third matching cost; adding the first matching cost, the second matching cost, and the third matching cost to obtain a matching cost between the third reference sub-block and the fourth reference sub-block.

72. A video encoding method according to claim 71.

73. obtaining a matching cost between the third reference sub-block and the fourth reference sub-block based on the first matching cost, the second matching cost, and the third matching cost; performing a weighted addition on the first matching cost, the second matching cost, and the third matching cost to obtain a matching cost between the third reference sub-block and the fourth reference sub-block; 72. A video encoding method according to claim 71.

74. performing a weighted addition on the first matching cost, the second matching cost, and the third matching cost to obtain a matching cost between the third reference sub-block and the fourth reference sub-block; determining a first weight corresponding to the first matching cost; determining second weights corresponding to the second matching cost and the third matching cost; weighting the first matching cost, the second matching cost, and the third matching cost based on the first weight and the second weight to obtain a matching cost between the third reference sub-block and the fourth reference sub-block; 74. A video encoding method according to claim 73.

75. obtaining a first second motion vector corresponding to the second sub-block based on the motion vector corresponding to the third reference sub-block with the smallest matching cost, and obtaining a second second motion vector corresponding to the second sub-block based on the motion vector corresponding to the fourth reference sub-block with the smallest matching cost; performing a sub-pel search around the third reference sub-block with the smallest matching cost to obtain a second sub-pel motion vector; obtaining a first second motion vector corresponding to the second sub-block based on the second sub-pixel motion vector and the motion vector corresponding to the third reference sub-block with the smallest matching cost; performing a sub-pel search around the fourth reference sub-block with the smallest matching cost to obtain a third sub-pel motion vector; obtaining a second motion vector corresponding to the second sub-block based on the third sub-pixel motion vector and the motion vector corresponding to the fourth reference sub-block having the smallest matching cost; 66. A video encoding method according to claim 65.

76. Obtaining a predicted value of the current block based on the i-th predicted value and the other predicted values ​​includes: determining weights for the predicted values ​​based on the weight derivation mode; weighting the i-th predicted value and the other predicted values ​​based on the weights of the predicted values ​​to obtain a predicted value of the current block; 68. A video encoding method according to claim 56 or 67.

77. 1. A video decoding device, comprising: a determining unit configured to determine K prediction modes of a current block, at least one prediction mode among the K prediction modes being an N-directional prediction mode, where K and N are both positive integers greater than 1; a prediction unit configured to predict the current block based on the K prediction modes to obtain a predicted value of the current block.

78. 1. A video encoding device, comprising: a determining unit configured to determine K prediction modes of a current block, at least one prediction mode among the K prediction modes being an N-directional prediction mode, where K and N are both positive integers greater than 1; a prediction unit configured to predict the current block based on the K prediction modes to obtain a predicted value of the current block.

79. An electronic device, a memory for storing a computer program; a processor for implementing the method according to any one of claims 1 to 38 or claims 39 to 76 by calling and executing the computer program stored in the memory.

80. a video decoder for implementing the method according to any one of claims 1 to 38; A video encoder for implementing the method of any one of claims 39 to 76.

81. A computer-readable storage medium storing a computer program for causing a computer to execute the method according to any one of claims 1 to 38 or claims 39 to 76.