Video encoding / decoding method, device, equipment, system and storage medium

By determining a reference region and interpolation filter for video blocks, the method improves transformation kernel accuracy and decoding efficiency in video encoding and decoding, addressing suboptimal prediction in existing technologies.

JP2026508708A5Pending Publication Date: 2026-03-30GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies face challenges in improving prediction efficiency and accuracy, particularly in determining transformation kernels for video blocks, leading to suboptimal encoding and decoding performance.

Method used

The proposed method involves determining a reference region and interpolation filter for a current block, using these to predict a block and determine a transformation kernel based on the prediction mode, followed by inverse transformation to obtain a residual block, thereby enhancing the accuracy of transformation kernel determination and improving decoding accuracy.

Benefits of technology

This approach improves the accuracy of transformation kernel determination, reduces the need for specifying individual transformation kernels, and enhances the overall video coding and decoding efficiency by saving codewords.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present application provides a video encoding / decoding method, device, apparatus, system, and storage medium. When predicting a current block, a reference region and an interpolation filter for the current block are determined, a prediction block for the current block is determined based on the reference region and the interpolation filter, a prediction mode corresponding to the prediction block is determined, a transform kernel corresponding to the current block is determined based on the prediction mode, and the transform kernel is used to perform an inverse transform on the transform coefficients of the current block to obtain a residual block for the current block. A reconstructed value of the current block is obtained based on the residual block for the current block and the prediction block. That is, in the present application, when the current block is predicted using an interpolation filter prediction method, a transform kernel corresponding to the current block is determined by determining a conventional prediction mode corresponding to the prediction block, thereby improving the accuracy of determining the transform kernel and improving the video encoding / decoding effect of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video encoding and decoding technologies, and particularly to video encoding and decoding methods, apparatuses, devices, systems, and storage media.

Background Art

[0002] Digital video technology can be incorporated into various video devices such as digital TVs, smartphones, computers, e-readers, or video players. With the development of video technology, the amount of data contained in video data is increasing. To facilitate the transmission of video data, video devices can execute video compression technology to more effectively transmit or store video data.

[0003] Since there is temporal or spatial redundancy in video, redundancy in the video can be removed or reduced by prediction, and the compression efficiency can be improved. To improve the prediction effect, the present application proposes an interpolation filtering prediction method.

Summary of the Invention

Means for Solving the Problems

[0004] Embodiments of the present application provide a video encoding and decoding method, apparatus, device, system, and storage media, which can improve the prediction effect of the current block and enhance the encoding and decoding performance.

[0005] In a first aspect, the present application provides a video decoding method, which is applied to a decoder, and the method includes: determining a reference region and an interpolation filter of a current block, and determining a prediction block of the current block based on the reference region and the interpolation filter; determining an intra prediction mode corresponding to the prediction block, and determining a transform kernel corresponding to the current block based on the intra prediction mode corresponding to the prediction block; The process includes the steps of: performing an inverse transformation on the transformation coefficients of the current block based on the transformation kernel corresponding to the current block to obtain the residual block of the current block; and obtaining the reconstructed block of the current block based on the predicted block and the residual block of the current block.

[0006] In a second aspect, embodiments of the present application provide a video encoding method, which is applied to an encoder. The steps include determining the reference region and interpolation filter of the current block, and determining the predicted block of the current block based on the reference region and the interpolation filter, The steps include determining an intra-prediction mode corresponding to the prediction block, and determining a conversion kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block, The process includes the steps of performing a transformation on the residual block of the current block based on the transformation kernel corresponding to the current block, obtaining the transformation coefficients of the current block, performing encoding based on the transformation coefficients of the current block, and obtaining a bitstream.

[0007] In a third aspect, the present application provides a video decoding device used to carry out the methods of the first aspect or each embodiment thereof. Specifically, the device includes a functional unit for carrying out the methods of the first aspect or each embodiment thereof.

[0008] In a fourth aspect, the present application provides a video encoding apparatus used to carry out the methods of the second aspect or each embodiment thereof. Specifically, the apparatus includes a functional unit for carrying out the methods of the second aspect or each embodiment thereof.

[0009] A fifth aspect provides a video decoder, including a processor and memory. The memory is used to store computer programs, and the processor is used to perform the methods of the first aspect or each embodiment thereof by calling and executing the computer programs stored in the memory.

[0010] The sixth aspect provides a video encoder, including a processor and memory. The memory is used to store computer programs, and the processor is used to perform the methods of the second aspect or each embodiment thereof by calling and executing the computer programs stored in the memory.

[0011] The seventh aspect provides a video encoding and decoding system, which includes a video encoder and a video decoder. The video decoder is used to carry out the method in the first aspect or each embodiment thereof, and the video encoder is used to carry out the method in the second aspect or each embodiment thereof.

[0012] In the eighth aspect, a chip is provided which is used to implement the method in any aspect of the first to second aspects or each embodiment thereof. Specifically, the chip includes a processor for calling and executing a computer program from memory, thereby causing the device on which the chip is installed to execute the method in any aspect of the first to second aspects or each embodiment thereof.

[0013] In the ninth aspect, a computer-readable storage medium is provided and used for storing a computer program, the computer program causing the computer to execute any aspect of the first to second aspects or the methods in each embodiment thereof.

[0014] In the tenth aspect, a computer program product is provided, which includes computer program instructions, the computer program instructions causing a computer to execute any aspect of the first to second aspects or the methods in each embodiment thereof.

[0015] In the eleventh aspect, a computer program is provided, and when the computer program is executed on a computer, the computer is made to execute any aspect of the first to second aspects or the methods in each embodiment thereof.

[0016] Based on the above technical solution, this application provides an interpolation filter prediction method. When making a prediction for a current block, first, the reference region and interpolation filter of the current block are determined, and based on the reference region and interpolation filter, the prediction block of the current block is determined. Next, the prediction mode corresponding to the prediction block is determined, and based on the prediction mode, the transformation kernel corresponding to the current block is determined. An inverse transformation is performed on the transformation coefficients of the current block using the transformation kernel to obtain the residual block of the current block, and based on the residual block and prediction block of the current block, the reconstruction value of the current block is obtained. In other words, in the embodiment of this application, when the current block is predicted using the interpolation filter prediction method, the conventional prediction mode corresponding to the prediction block is determined, which further determines the transformation kernel corresponding to the current block. This results in the determined transformation kernel being better suited to the characteristics of the current block, improving the accuracy of the transformation kernel determination. When the reconstruction value of the current block is determined using this accurately determined transformation kernel, the accuracy of the reconstruction value determination is improved, and the decoding accuracy of the current block can be enhanced. Furthermore, since the embodiment of this application determines the transformation kernel of the current block using a conventional prediction mode corresponding to the prediction block, there is no need to individually specify the transformation kernel, saving codewords and further improving the video coding and decoding effect. [Brief explanation of the drawing]

[0017] [Figure 1] This is a schematic block diagram of a video encoding and decoding system according to an embodiment of the present invention. [Figure 2] This is a schematic block diagram of a video encoder according to an embodiment of the present invention. [Figure 3] This is a schematic block diagram of a video decoder according to an embodiment of the present invention. [Figure 4A]It is a schematic diagram of intra prediction. [Figure 4B] It is a schematic diagram of intra prediction. [Figure 5A] It is a schematic diagram of intra prediction. [Figure 5B] It is a schematic diagram of intra prediction. [Figure 5C] It is a schematic diagram of intra prediction. [Figure 5D] It is a schematic diagram of intra prediction. [Figure 5E] It is a schematic diagram of intra prediction. [Figure 5F] It is a schematic diagram of intra prediction. [Figure 5G] It is a schematic diagram of intra prediction. [Figure 5H] It is a schematic diagram of intra prediction. [Figure 5I] It is a schematic diagram of intra prediction. <0,000,095>It is a schematic diagram of the intra prediction mode. [Figure 7] It is a schematic diagram of the intra prediction mode. [Figure 8] It is a schematic diagram of the intra prediction mode. [Figure 9] It is a schematic diagram of the CCCM principle. [Figure 10] It is a flow schematic diagram of the video decoding method according to an embodiment of the present application. [Figure 11] It is a schematic diagram of the position in the current image of the current block. [Figure 12] It is a schematic diagram of the reconstruction area. [Figure 13A] It is a schematic diagram of one type of reference area. [Figure 13B] It is a schematic diagram of another type of reference area. [Figure 13C] It is a schematic diagram of yet another type of reference area. [Figure 14A] It is a schematic diagram of one type of interpolation filter shape. [Figure 14B] It is a schematic diagram of another type of interpolation filter shape. [Figure 14C]This is a schematic diagram of yet another type of interpolation filter shape. [Figure 14D] This is a schematic diagram of yet another type of interpolation filter shape. [Figure 14E] This is a schematic diagram of yet another type of interpolation filter shape. [Figure 14F] This is a schematic diagram of yet another type of interpolation filter shape. [Figure 14G] This is a schematic diagram of yet another type of interpolation filter shape. [Figure 15] This is a schematic diagram of multiple types of interpolation filter shapes according to embodiments of the present application. [Figure 16] This is a schematic diagram of multiple types of interpolation filter shapes according to embodiments of the present application. [Figure 17] This is a schematic diagram of multiple types of interpolation filter shapes according to embodiments of the present application. [Figure 18A] This is a schematic diagram of multiple types of interpolation filter shapes according to embodiments of the present application. [Figure 18B] This is a schematic diagram of multiple types of interpolation filter shapes according to embodiments of the present application. [Figure 19] This is a schematic diagram of multiple types of interpolation filter shapes according to embodiments of the present application. [Figure 20] This is a schematic diagram of the first reconstruction area. [Figure 21] This is a schematic diagram showing interpolation filters of different shapes moving within different types of reference regions. [Figure 22] This is a schematic diagram showing how interpolation predictions are performed on the current block using an interpolation filter. [Figure 23] This is a schematic diagram of the intra-prediction mode. [Figure 24] This is a schematic diagram for determining the horizontal and vertical slopes. [Figure 25] This is a histogram of gradient amplitude values. [Figure 26] This is a schematic flowchart of a prediction method according to one embodiment of the present invention. [Figure 27] This is a schematic flowchart for determining the prediction mode according to the embodiment of the present invention. [Figure 28]This is a schematic block diagram of a video decoding device according to one embodiment of the present invention. [Figure 29] This is a schematic block diagram of a video encoding device according to one embodiment of the present invention. [Figure 30] This is a schematic block diagram of an electronic device according to an embodiment of the present application. [Figure 31] This is a schematic block diagram of a video encoding and decoding system according to an embodiment of the present invention. [Modes for carrying out the invention]

[0018] This invention can be applied to the fields of image coding and decoding, video coding and decoding, hardware video coding and decoding, dedicated circuit video coding and decoding, and real-time video coding and decoding. For example, the solution of this invention can be combined with audio video coding standards (AVS), such as the H.264 / Advanced video coding (AVC) standard, the H.265 / High efficiency video coding (HEVC) standard, and the H.266 / versatile video coding (VVC) standard. Alternatively, the present invention may be used in conjunction with other proprietary or industry standards, including ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, and ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), and including Scalable Video Coding (SVC) and Multiview Video Coding (MVC) extensions. It is understood that the present invention is not limited to any particular coding or decoding standard or technique.

[0019] To facilitate understanding, the video coding and decoding system according to the embodiment of this application will first be described with reference to Figure 1.

[0020] Figure 1 is a schematic block diagram of a video encoding / decoding system according to an embodiment of the present application. Note that Figure 1 is merely an example, and the video encoding / decoding system of the embodiment of the present application is not limited to that shown in Figure 1. As shown in Figure 1, the video encoding / decoding system 100 includes an encoding device 110 and a decoding device 120. Here, the encoding device is used to encode (understand as compressing) video data to generate a bitstream and transmit the bitstream to the decoding device. The decoding device decodes the bitstream generated by the encoding device and obtains the decoded video data.

[0021] In the embodiments of the present application, the encoding device 110 is understood to be a device having video encoding capabilities, and the decoding device 120 is understood to be a device having video decoding capabilities. That is, the embodiments of the present application include a broader range of devices than the encoding device 110 and the decoding device 120, including, for example, smartphones, desktop computers, mobile computing devices, notebook computers (e.g., laptops), tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, and the like.

[0022] In some embodiments, the encoding device 110 may transmit the encoded video data (e.g., a bitstream) to the decoding device 120 via channel 130. Channel 130 may include one or more media and / or devices capable of transmitting the encoded video data from the encoding device 110 to the decoding device 120.

[0023] In one example, channel 130 includes one or more communication media on which the encoding device 110 can directly transmit encoded video data to the decoding device 120 in real time. In this example, the encoding device 110 may modulate the encoded video data according to a communication standard and transmit the modulated video data to the decoding device 120. Here, the communication medium includes a wireless communication medium (e.g., a radio frequency spectrum), and optionally, the communication medium may further include a wired communication medium (e.g., one or more physical transmission lines).

[0024] In another example, channel 130 may include a storage medium that stores video data encoded by the encoding device 110. The storage medium may include a variety of local-access data storage media such as optical discs, DVDs, and flash memory. In this example, the decoding device 120 may obtain the encoded video data from the storage medium.

[0025] In another example, channel 130 may include a storage server that stores the video data encoded by the encoding device 110. In this example, the decoding device 120 may download the encoded video data stored from the storage server. Selectively, the storage server may store the encoded video data and transmit the encoded video data to the decoding device 120, for example, a web server (e.g., for a website), a File Transfer Protocol (FTP) server, etc.

[0026] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113, where the output interface 113 may include a modulator / demodulator (modem) and / or transmitter.

[0027] In some embodiments, the encoding device 110 may further include a video source 111 in addition to the video encoder 112 and the output interface 113.

[0028] The video source 111 may include at least one of a video acquisition device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system, where the video input interface is used to receive video data from a video content provider and the computer graphics system is used to generate video data.

[0029] The video encoder 112 encodes video data from the video source 111 and generates a bitstream. The video data may include one or more pictures or sequences of pictures. The bitstream contains encoding information for the pictures or sequences of pictures in bitstream format. The encoding information may include encoded image data and related data. The related data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntactic structures. The SPS may include parameters that apply to one or more sequences. The PPS may include parameters that apply to one or more pictures. A syntactic structure refers to a set of zero or more syntactic elements arranged in a specified order within the bitstream.

[0030] The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data may be further stored in a storage medium or storage server for later reading by the decoding device 120.

[0031] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122.

[0032] In some embodiments, the decoding device 120 may further include a display device 123 in addition to the input interface 121 and the video decoder 122.

[0033] Here, the input interface 121 may include a receiver and / or a modem. The input interface 121 may receive encoded video data via channel 130.

[0034] The video decoder 122 is used to decode the encoded video data, obtain the decoded video data, and transmit the decoded video data to the display device 123.

[0035] The display device 123 is used to display the decoded video data. The display device 123 may be integrated with the decoding device 120 or installed outside the decoding device 120. The display device 123 may include a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.

[0036] Furthermore, Figure 1 is merely an example, and the technical solutions of the embodiments of this application are not limited to Figure 1. For example, the technology of this application may also be applied to one-sided video encoding or one-sided video decoding.

[0037] The following describes the video encoding framework according to the embodiment of the present application.

[0038] Figure 2 is a schematic block diagram of a video encoder according to an embodiment of the present invention. It is understood that the video encoder 200 may be used for lossy compression or lossless compression of an image. This lossless compression may be visually lossless compression or mathematically lossless compression.

[0039] The video encoder 200 may be applied to image data in luminance-chromaticity (YCbCr, YUV) format. For example, the YUV ratio may be 4:2:0, 4:2:2, or 4:4:4, where Y represents luminance (Luma), Cb (U) represents blue chromaticity, Cr (V) represents red chromaticity, and U and V represent chromaticity (Chroma) to describe color and saturation. For example, in the color format, 4:2:0 means that there are four luminance components and two chromaticity components (YYYYCbCr) for every four pixels, 4:2:2 means that there are four luminance components and four chromaticity components (YYYYCbCrCbCr) for every four pixels, and 4:4:4 means full pixel display (YYYYCbCrCbCrCbCrCbCr).

[0040] For example, the video encoder 200 reads video data and divides each frame image in the video data into multiple coding tree units (CTUs). In some examples, CTUs are called "tree blocks," "largest coding units" (LCUs), or "coding tree blocks" (CTBs). Each CTU may be associated with a pixel block of the same size in the image. Each pixel may correspond to one luminance (or luma) sample and two chrominance (or chroma) samples. Therefore, each CTU may be associated with one luminance sample block and two chrominance sample blocks. The size of a single CTU may be, for example, 128×128, 64×64, or 32×32. A single CTU can be further divided into multiple coding units (CUs) for coding, and the CUs may be rectangular blocks or square blocks. The CU is further divided into a Prediction Unit (PU) and a Transform Unit (TU), separating encoding, prediction, and transformation to increase processing flexibility. In one example, the CTU is divided into CUs using a quadtree scheme, and the CU is then divided into TUs and PUs using a quadtree scheme.

[0041] Video encoders and video decoders can support a variety of PU sizes. Assuming a specific CU size of 2N×2N, video encoders and video decoders can support 2N×2N or N×N PU sizes for intra-prediction, and symmetric PU sizes of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-prediction. Video encoders and video decoders can further support asymmetric PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-prediction.

[0042] In some embodiments, as shown in Figure 2, the video encoder 200 may include a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / inverse quantization unit 240, a reconstruction unit 250, a loop filter unit 260, a decoded image buffer 270, and an entropy coding unit 280. The video encoder 200 may also include more, fewer, or different functional components.

[0043] Selectively, in this application, the current block may be called the current coding unit (CU) or the current prediction unit (PU), etc. The prediction block may also be called the prediction image block or image prediction block, and the reconstructed image block may also be called the reconstruction block or image reconstruction block.

[0044] In some embodiments, the prediction unit 210 includes an inter-prediction unit 211 and an intra-prediction unit 212. Because there is a strong correlation between adjacent pixels within a single frame of video, the intra-prediction method is used in video coding and decoding techniques to remove spatial redundancy between adjacent pixels. Because there is a strong similarity between adjacent frames of video, the inter-prediction method is used in video coding and decoding techniques to remove temporal redundancy between adjacent frames, thereby improving coding efficiency.

[0045] The interprediction unit 211 may be used for interprediction, which may include motion estimation and motion compensation, and may refer to image information from different frames. Interprediction is used to find reference blocks from reference frames using motion information, generate prediction blocks based on the reference blocks, and remove temporal redundancy. The frames used for interprediction can be P frames and / or B frames, where P frames refer to forward prediction frames and B frames refer to bidirectional prediction frames. Interprediction finds reference blocks from reference frames using motion information and generates prediction blocks based on the reference blocks. Motion information includes a list of reference frames in which the reference frames are located, a reference frame index, and a motion vector. The motion vector may be an integer pixel or a fractional pixel. If the motion vector is a fractional pixel, it is necessary to generate the required fractional pixel blocks using an interpolation filter in the reference frames, and the integer or fractional pixel blocks between reference frames found based on the motion vector are called reference blocks. In some techniques, the reference blocks are directly used as prediction blocks, and in some techniques, prediction blocks are generated by further processing based on the reference blocks. Generating a prediction block by further processing a reference block can also be understood as generating a new prediction block by first using the reference block as a prediction block and then further processing it based on that prediction block.

[0046] The intra-prediction unit 212 refers only to information from the same frame image, predicts pixel information within the current encoded image block, and is used to eliminate spatial redundancy. The frame used for intra-prediction may be an I-frame.

[0047] Intra-prediction has a variety of prediction modes. Taking the international digital video coding standard H-series as an example, the H.264 / AVC standard has 8 types of angle prediction modes and 1 type of non-angle prediction mode, while H.265 / HEVC is extended to 33 types of angle prediction modes and 2 types of non-angle prediction modes. The intra-prediction modes used in HEVC include Planar mode, DC, and 33 types of angle modes, for a total of 35 types of prediction modes. The intra-prediction modes used in VVC include Planar, DC, and 65 types of angle modes, for a total of 67 types of prediction modes.

[0048] Furthermore, with the increase in angle modes, intra-prediction becomes more accurate, better meeting the evolving needs of high-resolution and ultra-high-resolution digital video.

[0049] The residual unit 220 can generate residual blocks of the CU based on the pixel blocks of the CU and the prediction blocks of the CU's PU. For example, by generating residual blocks of the CU, the residual unit 220 ensures that each sample in the residual block has a value equal to the difference between the sample in the pixel block of the CU and the corresponding sample in the prediction block of the CU's PU.

[0050] The conversion / quantization unit 230 can quantize the conversion coefficients. The conversion / quantization unit 230 can quantize the conversion coefficients associated with the TU of the CU based on the quantization parameter (QP) value associated with the CU. The video encoder 200 can adjust the degree of quantization applied to the conversion coefficients associated with the CU by adjusting the QP value associated with the CU.

[0051] The inverse transform / inverse quantization unit 240 can apply inverse quantization and inverse transform to the quantized transformation coefficients, respectively, and reconstruct the residual block from the quantized transformation coefficients.

[0052] The reconstruction unit 250 can add the samples of the reconstructed residual blocks to the corresponding samples of one or more prediction blocks generated by the prediction unit 210 to generate a reconstructed image block associated with the TU. By reconstructing the sample blocks of each TU of the CU in this manner, the video encoder 200 can reconstruct the pixel blocks of the CU.

[0053] The loop filter unit 260 processes the pixels after inverse transformation and inverse quantization to compensate for distortion information and provide a good reference for subsequent encoded pixels. For example, deblocking filtering can be performed to reduce the blocking effect of pixel blocks associated with the CU.

[0054] In some embodiments, the loop filter unit 260 includes a deblocking filter unit and a sample adaptive compensation / adaptive loop filter (SAO / ALF) unit, where the deblocking filter unit is used to eliminate blocking effects and the SAO / ALF unit is used to eliminate ringing effects.

[0055] The decoded image buffer 270 can store the reconstructed pixel blocks. The inter-prediction unit 211 can perform inter-prediction on other PUs of other images using the reference image containing the reconstructed pixel blocks. In addition, the intra-prediction unit 212 can perform intra-prediction on other PUs in the same image as the CU using the reconstructed pixel blocks in the decoded image buffer 270.

[0056] The entropy coding unit 280 can receive quantized transformation coefficients from the transformation / quantization unit 230. The entropy coding unit 280 can generate entropy-coded data by performing one or more entropy coding operations on the quantized transformation coefficients.

[0057] Figure 3 is a schematic block diagram of a video decoder according to an embodiment of the present invention.

[0058] As shown in Figure 3, the video decoder 300 includes an entropy decoding unit 310, a prediction unit 320, an inverse quantization / inverse transformation unit 330, a reconstruction unit 340, a loop filter unit 350, and a decoded image buffer 360. The video decoder 300 may also include more, fewer, or different functional components.

[0059] The video decoder 300 can receive a bitstream. The entropy decoding unit 310 can extract syntactic elements from the bitstream by analyzing it. As part of the bitstream analysis, the entropy decoding unit 310 can analyze the syntactic elements in the bitstream after entropy coding. The prediction unit 320, the inverse quantization / inverse transformation unit 330, the reconstruction unit 340, and the loop filter unit 350 can decode the video data based on the syntactic elements extracted from the bitstream, i.e., generate the decoded video data.

[0060] In some embodiments, the prediction unit 320 includes an intra-prediction unit 322 and an inter-prediction unit 321.

[0061] The intra-prediction unit 322 can generate prediction blocks for the PU by performing intra-prediction. The intra-prediction unit 322 can generate prediction blocks for the PU based on spatially adjacent pixel blocks of the PU using the intra-prediction mode. The intra-prediction unit 322 can further determine the intra-prediction mode for the PU based on one or more syntactic elements analyzed from the bitstream.

[0062] The interprediction unit 321 can construct a first reference image list (list 0) and a second reference image list (list 1) based on syntactic elements analyzed from the bitstream. Furthermore, if the PU uses interprediction coding, the entropy decoding unit 310 can analyze the motion information of the PU. Based on the motion information of the PU, the interprediction unit 321 can determine one or more reference blocks of the PU. Based on one or more reference blocks of the PU, the interprediction unit 321 can generate prediction blocks for the PU.

[0063] The inverse quantization / inverse transformation unit 330 can inverse quantize (i.e., perform inverse quantization) the transformation coefficients associated with the TU. The inverse quantization / inverse transformation unit 330 can determine the degree of quantization using the QP value associated with the CU of the TU.

[0064] After the inverse quantization of the transformation coefficients, the inverse quantization / inverse transformation unit 330 can generate residual blocks associated with the TU by applying one or more inverse transformations to the inversely quantized transformation coefficients.

[0065] The reconstruction unit 340 can reconstruct the pixel blocks of the CU using the residual blocks associated with the TU of the CU and the predicted blocks of the PU of the CU. For example, the reconstruction unit 340 can reconstruct the pixel blocks of the CU by adding the samples of the residual blocks to the corresponding samples of the predicted blocks and obtain a reconstructed image block.

[0066] The loop filter unit 350 can perform deblocking filtering to reduce the blocking effect of pixel blocks associated with the CU.

[0067] The video decoder 300 can store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 may use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent predictions, or it may transmit the reconstructed image to a display device for display.

[0068] The basic flow of video encoding and decoding is as follows: On the encoding side, a single frame image is divided into blocks, and for the current block, the prediction unit 210 generates a predicted block of the current block using intra-prediction or inter-prediction. The residual unit 220 can calculate a residual block, i.e., the difference between the predicted block and the original block of the current block, based on the predicted block and the original block of the current block, and this residual block is also called residual information. This residual block undergoes processes such as transformation and quantization by the transformation / quantization unit 230 to remove information insensitive to the human eye and eliminate visual redundancy. Selectively, the residual block before transformation and quantization by the transformation / quantization unit 230 is called a time-domain residual block, and the time-domain residual block after transformation and quantization by the transformation / quantization unit 230 is called a frequency-domain residual block or frequency-range residual block. The entropy encoding unit 280 receives the quantized transformation coefficients output from the transformation / quantization unit 230, performs entropy encoding on these quantized transformation coefficients, and outputs a bitstream. For example, the entropy coding unit 280 can remove character redundancy based on the target context model and the probabilistic information of the binary bitstream.

[0069] On the decoding side, the entropy decoding unit 310 analyzes the bitstream to obtain prediction information and quantization coefficient matrix for the current block. The prediction unit 320 generates a predicted block for the current block using intra-prediction or inter-prediction based on the prediction information. The inverse quantization / inverse transform unit 330 uses the quantization coefficient matrix obtained from the bitstream to perform inverse quantization and inverse transform on the quantization coefficient matrix to obtain the residual block. The reconstruction unit 340 adds the predicted block and the residual block to obtain the reconstructed block. The reconstructed block constitutes a reconstructed image, and the loop filter unit 350 performs loop filtering on the reconstructed image based on the image or the block to obtain the decoded image. The encoding side also needs to perform the same operations as the decoding side to obtain the decoded image. This decoded image is also called the reconstructed image, and the reconstructed image can be used as a reference frame for inter-prediction of subsequent frames.

[0070] Furthermore, the block partitioning information determined by the encoding side, as well as mode information or parameter information such as prediction, transformation, quantization, entropy coding, and loop filtering, are included in the bitstream as needed. The decoding side analyzes the bitstream and, based on the existing information, determines the same block partitioning information as the encoding side, as well as mode information or parameter information such as prediction, transformation, quantization, entropy coding, and loop filtering, thereby ensuring that the image obtained by the encoding side and the decoded image obtained by the decoding side match.

[0071] The above is the basic flow of video coding and decoding under a block-based mixed coding framework. As technology advances, some modules or steps of this framework or flow may be optimized. This application applies to, but is not limited to, the basic flow of video coding and decoding under the said block-based mixed coding framework.

[0072] In the embodiments of this application, the current block may be the current coding unit (CU) or the current prediction unit (PU), etc. Due to the need for parallel processing, the image may be divided into slices, etc., and slices within the same image can be processed in parallel, that is, there is no data dependency between them. "Frame" is a general term and can usually be understood as one image. The frames described in this application may be replaced with images or slices, etc.

[0073] Intra prediction typically uses each angular mode and non-angular mode to make predictions on the current encoded block and obtain a predicted block. Based on rate distortion information calculated from the predicted block and the original block, the optimal prediction mode for the current encoded unit is selected, and this prediction mode is transmitted to the decoding side via a bitstream. The decoding side analyzes the prediction mode, predicts and obtains a predicted image of the current decoded block, and can obtain a reconstructed image by superimposing it with the residual pixels transmitted via the bitstream. The intra prediction method uses already encoded and decoded reconstructed pixels surrounding the current block as reference pixels to make predictions on the current block. Figure 4A is a schematic diagram of intra prediction. As shown in Figure 4A, the size of the current block is 4x4, and the pixels in the leftmost row and topmost column of the current block are the reference pixels of the current block. Intra prediction uses these reference pixels to make predictions on the current block. These reference pixels may all be available, i.e., all have been encoded and decoded. Some may be unavailable; for example, if the current block is the leftmost of all frames, the reference pixels to the left of the current block are unavailable. Alternatively, if a portion of the lower left side of the current block has not yet been encoded or decoded during the encoding or decoding of the current block, the reference pixels in the lower left side are also unavailable. If reference pixels are unavailable, they may be filled in using available reference pixels or specific values, or in a specific manner, or not filled in at all.

[0074] Figure 4B is a schematic diagram of intraprediction. As shown in Figure 4B, the Multiple Reference Line (MRL) intraprediction method can improve encoding and decoding efficiency by using more reference pixels, for example, by using four reference rows / columns as reference pixels in the current block.

[0075] Furthermore, intraprediction has various prediction modes, and Figures 5A to 5I are schematic diagrams of intraprediction. As shown in Figures 5A to 5I, intraprediction for a 4x4 block in H.264 mainly includes nine types of modes. Here, Mode 0, shown in Figure 5A, duplicates the upper pixels of the current block vertically as predicted values ​​in the current block. Mode 1, shown in Figure 5B, duplicates the left reference pixels horizontally as predicted values ​​in the current block. Mode 2, DC, shown in Figure 5C, uses the average of eight points A to D and I to L as the predicted value for all points. Modes 3 to 8, shown in Figures 5D to 5I, each duplicate reference pixels at specific angles to their corresponding positions in the current block. However, if some positions in the current block cannot accurately correspond to the reference pixels, it is necessary to use a weighted average of the reference pixels or a fractional pixel of the interpolated reference pixels.

[0076] Other modes include Plane and Planar, and with technological advancements and increased block size, the number of angle prediction modes is also increasing. Figure 6 is a schematic diagram of intra-prediction modes. As shown in Figure 6, the intra-prediction modes used in HEVC include Planar, DC, and 33 types of angle modes, for a total of 35 types of prediction modes. Figure 7 is a schematic diagram of intra-prediction modes. As shown in Figure 7, the intra-modes used in VVC include Planar, DC, and 65 types of angle modes, for a total of 67 types of prediction modes. Figure 8 is a schematic diagram of intra-prediction modes. As shown in Figure 8, the modes used in AVS3 include DC, Plane, Bilinear, PCM, and 62 types of angle modes, for a total of 66 types of prediction modes.

[0077] Furthermore, there are several techniques to improve prediction, such as improvements to fractional pixel interpolation of reference pixels and filtering of predicted pixels. For example, the Multiple Intra Prediction Filter (MIPF) in AVS3 generates predicted values ​​using different filters for different block sizes. For pixels at different positions within the same block, pixels close to the reference pixel generate predicted values ​​using one type of filter, while pixels further away from the reference pixel generate predicted values ​​using another type of filter. As for filtering techniques for predicted pixels, for example, there is the Intra Prediction Filter (IPF) in AVS3, which can filter the predicted values ​​using the reference pixel.

[0078] In some embodiments, the current video encoding and decoding employs Adaptive Loop Filter (ALF) technology in the loop filter unit. For example, ALF technology is used to filter the reconstructed image and obtain the final decoded image.

[0079] The following describes the Adaptive Loop Filter (ALF) technology.

[0080] ALF is a filter within a loop filter, designed based on the Wiener filter principle, and is a filter that minimizes the error between the target sample and the input sample. In a loop filter, the target sample is the original image, and the input is the reconstructed image.

[0081] Before performing filtering using ALF, first determine the filter coefficients.

[0082] For example, by constructing the Wiener-Hop fielder equation shown in equation (1) and solving it, the filter coefficients of the interpolation filter can be obtained.

[0083] JPEG2024192733000001.jpg110169

[0084] JPEG2024192733000002.jpg28169

[0085] For example, the Wiener-Hopf equation can be solved by decomposing the autocorrelation matrix using Cholesky decomposition, thereby obtaining the filter coefficients of the filter.

[0086] After determining the filter coefficients based on equation (1) above, the samples awaiting filtering are filtered using equation (2) below, and the filtered samples are obtained.

[0087] JPEG2024192733000003.jpg58169

[0088] The Convolutional Cross Component Model (CCCM) is a process that predicts chromaticity pixels using reconstructed luminance component pixels. Its advantage is that the decoder can obtain the CCCM filter coefficients using already reconstructed pixels, thus eliminating the overhead of storing filter coefficients in the bitstream, as is the case with ALF. As shown in Figure 9, the CCCM coefficients are obtained by calculating the already reconstructed pixels around the current chroma block to be predicted and the already reconstructed pixels around the luminance block at the corresponding position in that chroma block.

[0089] This embodiment proposes an interpolation filter prediction method. When making a prediction for the current block, first the reference region and interpolation filter of the current block are determined, and the predicted block of the current block is determined based on the reference region and interpolation filter. For example, the reference region is filtered using the interpolation filter, the filter coefficients of the filter are calculated and obtained, and an interpolation filter prediction is performed on the current block using the interpolation filter whose filter coefficients have already been determined, thereby obtaining the predicted block of the current block. Next, the prediction mode corresponding to the predicted block is determined, and a transformation kernel corresponding to the current block is determined based on the prediction mode. An inverse transformation is performed on the transformation coefficients of the current block using the transformation kernel to obtain the residual block of the current block, and the reconstructed value of the current block is obtained based on the residual block and the predicted block of the current block. In other words, in this embodiment, when the current block is predicted using the interpolation filter prediction method, the transformation kernel corresponding to the current block is determined by determining the conventional prediction mode corresponding to the predicted block, and the transformation kernel determined in this way fits the characteristics of the current block, improving the accuracy of the transformation kernel determination. When determining the reconstructed value of the current block using this accurately determined transformation kernel, the accuracy of determining the reconstructed value can be improved, and the decoding accuracy of the current block can be increased. Furthermore, since the embodiment of the present invention determines the transformation kernel of the current block by a conventional prediction mode corresponding to the prediction block, there is no need to individually specify the transformation kernel, thus saving codewords and further improving the video coding and decoding effect.

[0090] The video decoding method provided by the embodiment of this application will be described below, with reference to Figure 10, using the decoding side as an example.

[0091] Figure 10 is a schematic flowchart of a video decoding method according to one embodiment of the present invention, and the embodiment of the present invention is applied to the video decoder shown in Figures 1 and 3. As shown in Figure 10, the method of the embodiment of the present invention includes the following steps.

[0092] S101, determine the reference region and interpolation filter of the current block, and determine the predicted block of the current block based on the reference region and interpolation filter.

[0093] The decoding side, when decoding the current block, decodes the bitstream to obtain the quantization coefficients of the current block, performs inverse quantization on the quantization coefficients to obtain the transformation coefficients of the current block, and performs inverse transformation on the transformation coefficients to obtain the residual value of the current block. Next, the prediction mode of the current block is determined, the predicted value of the current block is determined based on the prediction mode, and the reconstructed value of the current block is obtained based on the predicted value and residual value of the current block.

[0094] In some embodiments, the current block is also called the predicted block.

[0095] In the embodiment of the present invention, the decoding side first determines the prediction mode of the current block.

[0096] In some embodiments, the method by which the decryption side determines the prediction mode of the current block includes at least the following:

[0097] Method 1: The encoding side determines the prediction mode for the current block. For example, from among the candidate prediction modes consisting of the conventional prediction mode and the interpolation filter prediction mode shown in Figure 6 or Figure 7, the candidate prediction mode with the lowest cost is determined as the prediction mode for the current block. Next, the encoding side adds the indication information for the prediction mode of the current block to the bitstream. As a result, the decoding side decodes the bitstream to obtain the indication information for the prediction mode of the current block, further determines the prediction mode for the current block based on this indication information, makes a prediction for the current block using that intra prediction mode, and obtains the predicted value for the current block.

[0098] For example, if the terminal device determines that the prediction mode of the current block is a conventional prediction mode, it writes the index of the current block's prediction mode to the bitstream as instruction information for that prediction mode. The decoding side decodes the bitstream to obtain the prediction mode index, and then determines the prediction mode of the current block from among the conventional prediction modes shown in Figure 6 or Figure 7 based on that index.

[0099] Method 2: The encoding side constructs an intra-prediction mode candidate list and selects the intra-prediction mode for the current block from this list. Note that the intra-prediction mode candidate list includes interpolation filter prediction modes. Next, the encoding side writes the sequence number (or index number) of the current block's intra-prediction mode in the intra-prediction mode candidate list to the bitstream. The decoding side then decodes the bitstream to determine the sequence number of the current block's intra-prediction mode in the intra-prediction mode candidate list, and at the same time constructs an intra-prediction mode candidate list in the same way as the encoding side (note that the constructed intra-prediction mode candidate list includes interpolation filter prediction modes). Furthermore, based on the sequence number of the current block's intra-prediction mode in the intra-prediction mode candidate list, the decoding side determines the current block's intra-prediction mode from the constructed intra-prediction mode candidate list. Finally, the determined intra-prediction mode for the current block is used to make a prediction for the current block and obtain the predicted value for the current block.

[0100] Method 3: The encoding side constructs an intra-prediction mode candidate list, which includes interpolation filter prediction modes. Next, it selects an intra-prediction mode for the current block from this intra-prediction mode candidate list. For example, it determines the cost of each candidate prediction mode in the intra-prediction mode candidate list in the current block template, and then determines the intra-prediction mode for the current block based on this cost. Correspondingly, the decoding side constructs an intra-prediction mode candidate list using the same method as the encoding side, and this constructed intra-prediction mode candidate list also includes interpolation filter prediction modes. Next, it determines the cost of each candidate prediction mode in the intra-prediction mode candidate list in the current block template, and then determines the intra-prediction mode for the current block based on this cost. Finally, it makes a prediction for the current block using the determined intra-prediction mode for the current block and obtains the predicted value for the current block.

[0101] Method 4: The encoding and decoding sides use the interpolation filter prediction mode by default to make predictions for the current block.

[0102] In addition to determining whether the current block adopts the interpolation filter prediction mode and performs prediction using methods 1 to 4 described above, the decoding side may also determine whether the current block adopts the interpolation filter prediction mode and performs prediction using the following method 5.

[0103] Method 5: The decoding side decodes the bitstream and obtains third information, which is used to indicate whether the current block should adopt the interpolation filter prediction mode to perform predictions. If the decoding side decides, based on the third information, that the current block should adopt the interpolation filter prediction mode to perform predictions, it determines the reference region and interpolation filter of the current block.

[0104] In method 5, if the encoding side decides that the current block adopts the interpolation filter prediction mode, it writes third information to the bitstream. The decoder then decodes the bitstream to obtain the third information and decides whether or not the current block adopts the interpolation filter prediction mode to perform predictions based on this third information. If the third information instructs the current block to adopt the interpolation filter prediction mode to perform predictions, the decoding side uses this interpolation filter prediction mode to perform predictions on the current block and obtains the predicted block for the current block. If the third information instructs the current block to perform predictions without adopting the interpolation filter prediction mode, the decoding side skips the step of performing predictions on the current block using this interpolation filter prediction mode, instead determines the prediction mode for the current block, uses the determined prediction mode to perform predictions on the current block, and obtains the predicted block for the current block.

[0105] The embodiments of this application may, but are not limited to, any instruction information that can indicate whether or not the current block employs an interpolation filter prediction mode to perform predictions, regarding the specific representation format of the third information described above.

[0106] In one example, the third piece of information may be represented as intra_eip_flag, which allows the current block to determine whether or not to use interpolation filter prediction mode for prediction based on different assigned values ​​of intra_eip_flag. For example, if intra_eip_flag=0, it indicates that the current block will not use interpolation filter prediction mode for prediction, and if intra_eip_flag=1, it indicates that the current block will use interpolation filter prediction mode for prediction. The encoding side then writes the predetermined flag intra_eip_flag to the bitstream, and the decoding side determines the prediction mode of the current block based on the value of the decoded predetermined flag intra_eip_flag. For example, if the predetermined flag intra_eip_flag=1, it indicates that the prediction mode of the current block is interpolation filter prediction mode, and the decoding side then uses interpolation filter prediction mode to make predictions for the current block.

[0107] In some embodiments, to improve the prediction accuracy of the interpolation filter prediction mode, the interpolation filter prediction mode is used for some blocks that meet the requirements, and not for some blocks that do not meet the requirements. Based on this, before decoding the bitstream and obtaining the third information, the decoding side needs to determine whether the position of the current block in the current image satisfies a predetermined position requirement, and whether the size of the current block satisfies a predetermined block size. If it is determined that the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the bitstream is decoded and the third information is obtained.

[0108] The embodiments of this application do not limit the predetermined positional requirements and predicted block size, which are specifically determined based on actual needs.

[0109] In one example, as shown in Figure 11, assuming the position of the upper left corner of the current image is (0,0) and the position of the upper left corner of the current block is (x,y), the predetermined position requirement is that the x of the current block is greater than or equal to the first predetermined value XX, and the y of the current block is greater than or equal to the second predetermined value YY.

[0110] The embodiments of this application do not limit the specific numerical values ​​of the first and second predetermined values ​​described above.

[0111] For example, the first predetermined value and the second predetermined value are the same.

[0112] For example, both the first predetermined value and the second predetermined value are 13. In other words, if the distance from the top edge of the current block to the top edge of the current image is 13 rows of pixels or more, and the distance from the left edge of the current block to the left edge of the current image is 13 columns of pixels or more, then the position of the current block in the current image satisfies the predetermined position requirement.

[0113] In one example, referring to Figure 11, assuming that the current block width is W and height is H, the predetermined block size requirement is that the current block width W is less than or equal to the third predetermined value A, and the current block height H is less than or equal to the fourth predetermined value B.

[0114] The embodiments of this application do not limit the specific numerical values ​​of the third and fourth predetermined values ​​described above.

[0115] For example, the third predetermined value and the fourth predetermined value are the same.

[0116] For example, both the third and fourth predetermined values ​​are 32. In other words, if the current block's width and height are both 32 or less, it indicates that the current block meets the predetermined block size requirement.

[0117] In the embodiment of the present invention, before determining whether the current block will perform predictions using the interpolation filter prediction mode, the decoding side first determines whether the position of the current block in the current image satisfies a predetermined position requirement, and whether the size of the current block satisfies a predetermined block size requirement. If the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the bitstream is decoded to obtain third information, and based on this third information, it is determined whether the current block will perform predictions using the interpolation filter prediction mode. For example, as shown in Figure 11, if the distance from the top edge of the current block to the top edge of the current image is 13 rows of pixels or more, the distance from the left edge of the current block to the left edge of the current image is 13 columns of pixels or more, and the width and height of the current block are both 32 or less, the decoding side decodes the bitstream to obtain third information.

[0118] In some embodiments, the first predetermined value, second predetermined value, third predetermined value, and fourth predetermined value are default values.

[0119] In some embodiments, the first predetermined value, second predetermined value, third predetermined value, and fourth predetermined value are values ​​obtained by the decoding side after decoding from the bitstream.

[0120] In some embodiments, if the position of the current block in the current image does not meet a predetermined position requirement, and / or the size of the current block does not meet a predetermined block size requirement, it is determined that the current block will perform prediction without employing an interpolation filter prediction mode.

[0121] In some embodiments, the decoding side decodes a bitstream to obtain fourth information before determining whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement, the fourth information being used to indicate whether the current sequence is permitted to perform predictions using an interpolation filter prediction mode; and if the fourth information indicates that the current sequence is permitted to perform predictions using an interpolation filter prediction mode, the decoding side further includes determining whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement.

[0122] In the embodiment of the present invention, a higher-level syntactic element, such as a fourth piece of information at the sequence level, indicates whether the current sequence is permitted to perform predictions using an interpolation filter prediction mode. If the fourth piece of information indicates that the current sequence is permitted to perform predictions using an interpolation filter prediction mode, the decoding side determines whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement. Furthermore, if it is determined that the position of the current block in the current image satisfies the predetermined position requirement and that the size of the current block satisfies the predetermined block size requirement, the third piece of information is decoded to determine whether the current block is permitted to perform predictions using an interpolation filter prediction mode.

[0123] In some embodiments, if the fourth information indicates that the current sequence does not allow prediction using an interpolation filter prediction mode, the decoding side skips the steps of determining whether the position of the current block in the current image satisfies a predetermined position requirement, whether the size of the current block satisfies a predetermined block size requirement, and decoding the third information.

[0124] The embodiments of this application do not limit the specific representation format of the fourth information, and may be any instruction information that can indicate whether or not the current sequence is permitted to perform predictions using an interpolation filter prediction mode.

[0125] For example, the fourth piece of information may be represented as sps_eip_enabled_flag, where different values ​​of sps_eip_enabled_flag determine whether the current sequence is allowed to perform predictions using interpolation filter prediction mode. For example, sps_eip_enabled_flag=0 indicates that the current sequence is not allowed to perform predictions using interpolation filter prediction mode, while sps_eip_enabled_flag=1 indicates that the current sequence is allowed to perform predictions using interpolation filter prediction mode.

[0126] For example, the fourth piece of information is transported to the sequence-level parameter set (SPS), as shown in Table 1.

[0127] JPEG2024192733000004.jpg57139

[0128] However, sps_eip_enabled_flag represents the fourth piece of information, and this sps_eip_enabled_flag is passed to seq_parameter_set_rbsp(). For example, sps_eip_enabled_flag=0 indicates that the current sequence is not allowed to perform predictions using the interpolation filter prediction mode, and sps_eip_enabled_flag=1 indicates that the current sequence is allowed to perform predictions using the interpolation filter prediction mode.

[0129] In some embodiments, embodiments of the present invention may further include a general constraints information (GCI) identification bit that indicates whether or not to use interpolation filter prediction techniques. For example, gci_no_eip_constraint_flag indicates whether or not the current video enables interpolation filter prediction techniques. For example, as shown in Table 2, the gci_no_eip_constraint_flag is carried to general constraints information general_constraints_info().

[0130] JPEG2024192733000005.jpg70142

[0131] As shown in Table 2, if gci_no_eip_constraint_flag=1, it means that the current video does not enable interpolation filter prediction. In other words, it means that a constraint is imposed on the sequence-level interpolation filter intra-prediction technique, requiring it to be 0 for all images. In other words, it means that the use of interpolation filter intra-prediction is not permitted for any sequence contained in the current video. If gci_no_eip_constraint_flag=0, it means that the current video enables interpolation filter prediction. In other words, it means that the constraint that the sequence-level interpolation filter intra-prediction technique must be 0 for all images is not imposed.

[0132] As can be seen from the above, if the syntactic elements in the embodiment of this application include the high-level syntactic elements gci_no_eip_constraint_flag and sps_eip_enabled_flag, as well as the block-level intra_eip_flag, the decoder first decodes the high-level syntactic elements. Specifically, it first decodes gci_no_eip_constraint_flag, and if gci_no_eip_constraint_flag=0, it then decodes sps_eip_enabled_flag. If sps_eip_enabled_flag=1, it parses the block syntactic elements.

[0133] For example, block-level syntactic elements are shown in Table 3.

[0134] JPEG2024192733000006.jpg42155

[0135] In Table 3, cbWidth and cbHeight currently represent the width and height of the block, SIZE_A may be understood as the third predetermined value described above, SIZE_B as the fourth predetermined value, XX as the first predetermined value, and YY as the second predetermined value. x0 and y0 represent the coordinate difference between the top-left corner of the block and the top-left corner of the image.

[0136] As is clear from Table 3 above, when the fourth information at the sequence level, sps_eip_enabled_flag=1, indicates that the use of interpolation filter prediction mode is permitted in the current sequence, it is determined whether the position of the current block in the current image satisfies a predetermined position requirement, and whether the size of the current block satisfies a predetermined block size requirement. If it is determined that the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the third information intra_eip_flag is decoded, and based on the decoded third information intra_eip_flag, it is determined whether the current block employs interpolation filter prediction mode to perform prediction.

[0137] This concludes the explanation of the specific process by which the current block determines whether or not to use the interpolation filter prediction mode for prediction.

[0138] In the embodiment of the present invention, when the decoding side decides that the current block will be predicted using the interpolation filter prediction mode, it performs a prediction of the current block using the interpolation filter prediction mode and obtains the predicted value of the current block.

[0139] The following describes the process by which the decoding side makes predictions for the current block using the interpolation filter prediction mode.

[0140] If the decoding side determines that the current block will be predicted using the interpolation filter prediction mode, it first determines the reference region and interpolation filter of the current block.

[0141] The following describes the specific process by which the decryption side determines the reference region of the current block.

[0142] In the embodiments of the present application, the reference region of the current block is part or all of the already reconfigured region surrounding the current block.

[0143] As an example, as shown in Figure 12, the reconstruction region surrounding the current block may include the reconstruction region above the current block, the reconstruction region to the left of the current block, the reconstruction region to the upper right of the current block, the reconstruction region to the lower left of the current block, and the reconstruction region to the upper left of the current block. Note that the block to be predicted in Figure 12 is the current block.

[0144] The embodiments of this application do not currently limit the specific shape and size of the reference region of the block.

[0145] For example, the reference region of the current block includes any one of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, or the upper left reconstruction region of the current block. For example, the reference region of the current block may be the upper reconstruction region of the current block, or the reference region of the current block may be the left reconstruction region of the current block.

[0146] For example, the reference region of the current block includes any two of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region and the left reconstruction region of the current block. Also, for example, the reference region of the current block includes the upper reconstruction region and the lower left reconstruction region of the current block.

[0147] For example, the reference region of the current block includes any three of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region of the current block, the upper right reconstruction region of the current block, and the upper left reconstruction region of the current block. Also, for example, the reference region of the current block includes the left reconstruction region of the current block, the upper left reconstruction region of the current block, and the lower left reconstruction region of the current block.

[0148] For example, the reference region of the current block includes any four of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region of the current block, the upper right reconstruction region of the current block, the upper left reconstruction region of the current block, and the left reconstruction region of the current block. Also, for example, the reference region of the current block includes the left reconstruction region of the current block, the upper left reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper reconstruction region of the current block.

[0149] In one example, the reference region of the current block includes all five reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block.

[0150] In the embodiments of this application, the specific methods by which the decoding side determines the reference region of the current block include, but are not limited to, those described below.

[0151] Method 1: The current block's reference region is used as the default region. For example, the encoding and decoding sides are set to default to the current block's reference region including at least one of the following: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block.

[0152] Method 2: The decryption side decrypts the bitstream to obtain first information, which is used to indicate the type of reference region of the current block. Based on the type of reference region, the reference region of the current block is determined from among a predetermined P reference regions, where P is a positive integer greater than 1.

[0153] In this implementation, the encoding side determines the reference region of the current block from among P pre-configured reference regions. For example, the encoding side calculates the encoding cost corresponding to each of these P reference regions and determines the reference region with the minimum encoding cost as the reference region of the current block. Next, the type of the reference region with the minimum encoding cost is instructed to the decoding side via the first information. As a result, the decoding side decodes the bitstream to obtain the first information, and then determines the reference region of the current block from among the P pre-configured reference regions based on the type of reference region indicated by the first information.

[0154] Furthermore, these pre-configured P reference regions will each be of a different type or shape.

[0155] The embodiments of this application do not specifically limit the number and shape of the P reference regions.

[0156] In one example, the P reference regions include at least one of the first, second, and third reference regions.

[0157] Here, the first reference region shown in Figure 13A includes the reconstruction regions above, to the upper right, to the left, to the lower left, and to the upper left of the current block. The second reference region shown in Figure 13B includes the reconstruction regions above, to the upper right, and to the upper left of the current block. The third reference region shown in Figure 13C includes the reconstruction regions to the left, to the lower left, and to the upper left of the current block. The block to be predicted in Figures 13A to 13C is the current block.

[0158] The embodiments of this application do not limit the specific representation format of the first information. Any directive information that can indicate the type of reference region of the current block is acceptable.

[0159] In one example, the first piece of information is represented as eip_ref_type. For instance, different types of reference regions are indicated depending on the value of eip_ref_type.

[0160] As an example, as shown in Table 4, the correspondence between the three reference regions shown in Figures 13A to 13B and the value of eip_ref_type is as follows:

[0161] JPEG2024192733000007.jpg52137

[0162] Based on Table 4 above, the decryption side decrypts the bitstream to obtain the first information eip_ref_type, and then determines the reference region of the current block based on the value of the first information eip_ref_type. For example, if eip_ref_type=0, the reference region of the current block is determined as the first reference region. As shown in Figure 13A, the first reference region includes the reconstruction regions above, to the upper right, left, lower left, and upper left of the current block. If eip_ref_type=1, the reference region of the current block is determined as the second reference region. As shown in Figure 13B, the second reference region includes the reconstruction regions above, to the upper right, and upper left of the current block. If eip_ref_type=2, the reference region of the current block is determined as the third reference region. As shown in Figure 13C, the third reference region includes the reconstruction regions to the left, to the upper right, and lower left of the current block.

[0163] In the above explanation, we have used the example of the P reference regions being the three reference regions shown in Figures 13A to 13C. However, the P reference regions in the embodiment of this application may include other reference regions besides the three reference regions mentioned above, but the embodiment of this application is not limited to this. The correspondence between the reference regions shown in Table 4 above and the value of eip_ref_type can be adjusted as appropriate depending on the number of reference regions.

[0164] In some embodiments, the decryption side may employ a truncated binary code decryption method and decrypt and obtain the first information from the bitstream.

[0165] For example, the correspondence between truncated binary code, the value of eip_ref_type, and the type of reference area is shown in Table 5.

[0166] JPEG2024192733000008.jpg56149

[0167] In the embodiments of the present invention, the decoding side may employ an equal-probability decoding method, a context model decoding method, or decode the codeword of the truncated binary code.

[0168] In addition to determining the reference region of the current block using method 1 or method 2 described above, the decryption side may also determine the reference region of the current block using method 3 shown below.

[0169] Method 3: Based on the shape of the current block, the reference region of the current block is determined from among the P pre-set reference regions.

[0170] Method 3 improves the accuracy of predictions by using different reference regions for each current block with a different shape.

[0171] For example, if the current block shape is a square, use the first type of reference region.

[0172] Additionally, for example, if the current block shape is a rectangle where the width is greater than the height, a second type of reference area is used.

[0173] Furthermore, for example, if the current block shape is a rectangle where the width is less than the height, a third type of reference area is used.

[0174] In other words, in the embodiment of the present application, the correspondence between the P reference regions and the shape of the current block is predetermined. As a result, the decoding side may determine the reference region of the current block from among the P reference regions based on the correspondence between the P reference regions and the shape of the current block.

[0175] The following describes the process by which the decoding side determines the interpolation filter for the current block.

[0176] In the embodiments of this application, the specific shape of the interpolation filter is not limited.

[0177] Exemplary interpolation filters provided by embodiments of the present application include, but are not limited to, square interpolation filters and interpolation filters whose height is less than their width.

[0178] For example, a square interpolation filter includes, but is not limited to, the 4x4 interpolation filter shown in Figure 14A.

[0179] Furthermore, interpolation filters where the height is greater than the width include, but are not limited to, the 5×3 interpolation filter shown in Figure 14B, the 6×2 interpolation filter shown in Figure 14D, and the 7×1 interpolation filter shown in Figure 14G.

[0180] Furthermore, interpolation filters where the height is smaller than the width include, but are not limited to, the 3×5 interpolation filter shown in Figure 14C, the 2×6 interpolation filter shown in Figure 14E, and the 1×7 interpolation filter shown in Figure 14F.

[0181] JPEG2024192733000009.jpg20168

[0182] In the embodiments of this application, specific methods for determining the interpolation filter of the current block by the decoding side include, but are not limited to, those described below.

[0183] Method 1: The interpolation filter for the current block is set as the default interpolation filter. For example, the default setting for the encoding side and the decoding side is to set the interpolation filter for the current block to any one of the interpolation filters shown in Figures 14A to 14G. For example, the default interpolation filter is a 4x4 interpolation filter.

[0184] Method 2: The decoding side decodes the bitstream to obtain second information, which is used to indicate the shape of the interpolation filter for the current block. Based on the shape of the interpolation filter for the current block, the interpolation filter for the current block is determined from among Q pre-set interpolation filters, where Q is a positive integer greater than 1.

[0185] In this implementation, the encoding side determines the interpolation filter for the current block from among Q pre-configured interpolation filters. For example, the encoding side determines the encoding cost corresponding to each of these Q interpolation filters and determines the interpolation filter with the minimum encoding cost as the interpolation filter for the current block. Next, the shape of the interpolation filter with the minimum encoding cost is instructed to the decoding side via second information. As a result, the decoding side decodes the bitstream to obtain the second information and then determines the interpolation filter for the current block from among the Q pre-configured interpolation filters based on the shape of the interpolation filter indicated by the second information.

[0186] Furthermore, these pre-configured Q interpolation filters will each have a different shape.

[0187] The embodiments of this application do not specifically limit the number and shape of the Q interpolation filters. For example, the Q interpolation filters include at least one of a first interpolation filter, a second interpolation filter, and a third interpolation filter, wherein the first interpolation filter is a square interpolation filter, the second interpolation filter is a rectangular interpolation filter whose width is greater than its height, and the third interpolation filter is a rectangular interpolation filter whose height is greater than its width.

[0188] In one example, Q interpolation filters include multiple interpolation filters as shown in Figures 14A to 14G.

[0189] The embodiments of this application do not limit the specific representation format of the second information. Any instruction information that can specify the shape of the interpolation filter of the current block is acceptable.

[0190] In one example, the second piece of information is represented as `eip_filter_type`. For instance, the value of `eip_filter_type` specifies an interpolation filter of a different shape.

[0191] For example, if the Q interpolation filters are the five interpolation filters shown in Figure 15, the correspondence between the five interpolation filters and the value of eip_filter_type is as shown in Table 6.

[0192] JPEG2024192733000010.jpg47135

[0193] Based on Table 5 above, the decoding side decodes the bitstream to obtain the second information eip_filter_type, and then determines the interpolation filter for the current block based on the value of the second information eip_filter_type. For example, if eip_filter_type=0, the shape of the interpolation filter for the current block is determined to be 4x4. If eip_filter_type=1, the shape of the interpolation filter for the current block is determined to be 3x5. If eip_filter_type=2, the shape of the interpolation filter for the current block is determined to be 5x3. If eip_filter_type=3, the shape of the interpolation filter for the current block is determined to be 2x6. If eip_filter_type=4, the shape of the interpolation filter for the current block is determined to be 6x2.

[0194] In some embodiments, the decryption side may employ a truncated binary code decryption method and decrypt and obtain second information from the bitstream.

[0195] For example, if the pre-configured Q interpolation filters include the five interpolation filters shown in Figure 15, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 7.

[0196] JPEG2024192733000011.jpg52139

[0197] In this case, combining the five types of interpolation filter shapes shown in Table 7 above with the three types of reconstruction region types shown in Table 5 above results in a total of 15 possible combinations of interpolation filters and reconstruction regions.

[0198] In some embodiments, the decoder may decode the bitstream to obtain the second information eip_filter_type, and then determine the interpolation filter for the current block from Table 7 above based on the shape of the interpolation filter indicated by the second information eip_filter_type. Similarly, the bitstream is decoded to obtain the first information eip_ref_type, and then determine the reference region for the current block from Table 5 above based on the value of the first information eip_ref_type.

[0199] For example, the syntactic elements of the embodiment of this application are as shown in Table 8.

[0200] JPEG2024192733000012.jpg157162

[0201] As shown in Table 8, the decoding side decodes the bitstream and first obtains the sequence-level fourth information sps_eip_enabled_flag. This fourth information sps_eip_enabled_flag indicates whether or not the interpolation filter prediction mode is permitted to be used for prediction in the current sequence. Next, it determines whether the position of the current block in the current image satisfies a predetermined position requirement and whether or not the size of the current block satisfies a predetermined block size requirement. If it is determined that the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the third information intra_eip_flag is decoded. This third information intra_eip_flag indicates whether or not the current block will use the interpolation filter prediction mode for prediction. If the third information intra_eip_flag=1 indicates that the current block will use the interpolation filter prediction mode, the bitstream is decoded to obtain the first information eip_ref_type and the second information eip_filter_type. Here, the first piece of information, eip_ref_type, indicates the type of reference region of the current block, allowing the decoding side to obtain the reference region of the current block by table lookup based on the value of the first piece of information, eip_ref_type. Furthermore, the second piece of information, eip_filter_type, indicates the shape of the interpolation filter of the current block, and based on the value of the second piece of information, eip_filter_type, the interpolation filter of the current block can be obtained by table lookup.

[0202] In some embodiments, when the embodiment of the present application includes the seven interpolation filters shown in Figure 16, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 9.

[0203] JPEG2024192733000013.jpg73129

[0204] In this case, combining the seven types of interpolation filter shapes shown in Table 9 above with the three types of reconstruction region types shown in Table 5 above results in a total of 21 possible combinations of interpolation filters and reconstruction regions.

[0205] Similarly, the decryption side can decrypt the syntax shown in Table 8 and, by referring to Table 5 and Table 8 mentioned above, obtain the reference region and interpolation filter of the current block.

[0206] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in Figure 17, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 10.

[0207] JPEG2024192733000014.jpg46133

[0208] In this case, by combining the three types of interpolation filter shapes shown in Table 10 above with the three types of reconstruction region types shown in Table 5 above, there are a total of nine possible combinations of interpolation filters and reconstruction regions.

[0209] Similarly, the decryption side can decrypt the syntax shown in Table 8 and look up Table 5 and Table 10 mentioned above to obtain the reference region and interpolation filter of the current block.

[0210] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in Figure 18A, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 11.

[0211] JPEG2024192733000015.jpg46133

[0212] In this case, combining the three types of interpolation filter shapes shown in Table 11 above with the three types of reconstruction region types shown in Table 5 above results in a total of nine possible combinations of interpolation filters and reconstruction regions.

[0213] Similarly, the decryption side can decrypt the syntax shown in Table 8 and look up Table 5 and Table 11 mentioned above to obtain the reference region and interpolation filter of the current block.

[0214] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in Figure 18B, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 12.

[0215] JPEG2024192733000016.jpg46133

[0216] In this case, combining the three types of interpolation filter shapes shown in Table 12 above with the three types of reconstruction region types shown in Table 5 above results in a total of nine possible combinations of interpolation filters and reconstruction regions.

[0217] Similarly, the decryption side can decrypt the syntax shown in Table 8 and look up Table 5 and Table 11 mentioned above to obtain the reference region and interpolation filter of the current block.

[0218] Generally, using a filter with more taps for the same number of samples yields better interpolation results. Compared to the 2x6 and 6x2 interpolation filters shown in Figure 18A, the interpolation filters shown in Figure 18B have an increased number of taps. For example, they are expanded to 2x8 and 8x2 interpolation filters. In practice, the 2x8 and 8x2 filters, like the 4x4 filter, use 15 samples as input and produce one result as output. From a complexity standpoint, the complexity of these filters is approximate. Therefore, the interpolation filters shown in Figure 18B can improve the interpolation effect without increasing complexity.

[0219] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in Figure 19, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 13.

[0220] JPEG2024192733000017.jpg46133

[0221] In this case, combining the three types of interpolation filter shapes shown in Table 13 above with the three types of reconstruction region types shown in Table 5 above results in a total of nine possible combinations of interpolation filters and reconstruction regions.

[0222] Similarly, the decryption side can decrypt the syntax shown in Table 8 and look up Table 5 and Table 13 mentioned above to obtain the reference region and interpolation filter of the current block.

[0223] In addition to determining the interpolation filter for the current block using method 1 or method 2 described above, the decoding side may also determine the interpolation filter for the current block using method 3 shown below.

[0224] Method 3: Based on the shape of the current block, the interpolation filter for the current block is determined from among Q pre-set interpolation filters.

[0225] In method 3, prediction accuracy is improved by using different interpolation filters for each current block with a different shape.

[0226] For example, if the current block shape is a square, use the interpolation filter for the first type of shape.

[0227] Additionally, for example, if the current block shape is a rectangle where the width is greater than the height, a second type of shape interpolation filter is used.

[0228] Furthermore, for example, if the current block shape is a rectangle where the width is less than the height, a third type of shape interpolation filter is used.

[0229] In other words, in the embodiment of the present invention, the correspondence between the Q interpolation filters and the shape of the current block is predetermined. As a result, the decoding side can determine the interpolation filter for the current block from among the Q interpolation filters based on the correspondence between the Q interpolation filters and the shape of the current block.

[0230] In the embodiment of the present invention, the decoding side determines the reference region and interpolation filter of the current block based on the steps described above, and then determines the predicted block of the current block based on the reference region and interpolation filter.

[0231] The following describes the process by which the decoding side determines the predicted block for the current block based on the reference region of the current block and the interpolation filter.

[0232] In the embodiment of the present invention, the decoding side determines the reference region and interpolation filter of the current block, then determines the filter coefficients of the interpolation filter by filtering the reference region using the interpolation filter, and further performs interpolation filtering on the current block based on the determined filter coefficients to obtain a predicted block of the current block.

[0233] The embodiments of this application do not limit the specific method by which the decoding side determines the predicted block of the current block based on the reference region of the current block and the interpolation filter.

[0234] In some embodiments, determining the predicted block of the current block based on the reference region and interpolation filter of the current block in S101 described above includes the following steps.

[0235] S101-A1, the filter coefficients of the interpolation filter are determined based on the reference region.

[0236] S101-A2, based on the filter coefficients, interpolation filtering prediction is performed on the current block using an interpolation filter, and the predicted block for the current block is obtained.

[0237] The method for determining the filter coefficients of the interpolation filter in S101-A1 described above includes at least the method shown below.

[0238] Method 1: The interpolation filter determined above is slid across the reference region of the current block to construct the Wiener-Hoff equations. Next, the filter coefficients of the interpolation filter are obtained by solving the Wiener-Hoff equations.

[0239] JPEG2024192733000018.jpg69168

[0240] In one example, the Wiener-Hoff equation constructed by sliding the interpolation filter within the reference region of the current block is as shown in equation (3).

[0241] JPEG2024192733000019.jpg57168

[0242] JPEG2024192733000020.jpg39168

[0243] In one example, the decoding side employs a method of decomposing the autocorrelation matrix using Cholesky decomposition, solving the Wiener-Hoff equation shown in equation (3) above to obtain the filter coefficients of the filter.

[0244] The decoding side determines the filter coefficients of the interpolation filter based on equation (3) described above, and then uses the interpolation filter to perform interpolation filtering prediction on the current block based on these filter coefficients, and obtains the predicted block for the current block.

[0245] For example, the decryption side obtains the predicted block of the current block based on the following equation (4).

[0246] JPEG2024192733000021.jpg100168

[0247] Method 2: The above-mentioned S101-A1 includes the following steps S101-A11 to S101-A14.

[0248] S101-A11, currently, the first reconstruction region around the block is determined.

[0249] S101-A12: Based on the reconstruction values ​​of the first reconstruction region, the pixel-average reconstruction value is determined.

[0250] S101-A13: Mean subtraction is performed on the reconstructed values ​​of the pixel points in the reference region based on the pixel-average reconstructed values.

[0251] S101-A14: The pixel values ​​of the pixel points after averaging removal in the reference region are used as input to the interpolation filter, and the interpolation filter is slid within the reference region to obtain the filter coefficients of the interpolation filter.

[0252] In method two, averaging removal is performed on the reference region, and the filter coefficients of the interpolation filter are determined based on the reference region after averaging removal. Since the amount of data is reduced by performing averaging removal on the reference region, the efficiency of determining the filter coefficients can be improved when the filter coefficients are determined based on the reference region after averaging removal.

[0253] Specifically, the decoding side first determines a first reconstruction region, which may be any part of the reconstruction region within the reconstruction region surrounding the current block.

[0254] In embodiments of the present invention, the method by which the decoding side determines the first reconstruction region around the current block includes at least the method shown below.

[0255] Method 1: By default, the decoding side determines one reconstruction area around the current block as the first reconstruction area.

[0256] For example, as shown in FIG. 20, by default, the decoding side determines the area composed of one row above the current block, one column on the left side, and one pixel point in the upper left corner as the first reconstruction area.

[0257] Method 2: Determine the first reconstruction area based on the shape of the current block.

[0258] For example, when the shape of the current block is a square, determine the reconstruction pixel area of one row above the current block and one column on the left side as the first reconstruction area.

[0259] Also, for example, when the shape of the current block is a rectangle with a width larger than the height, determine the reconstruction pixel area of one row above the current block as the first reconstruction area.

[0260] Also, for example, when the shape of the current block is a rectangle with a height larger than the width, determine the reconstruction pixel area of one column on the left side of the current block as the first reconstruction area.

[0261] Note that the method of determining the first reconstruction area based on the shape of the current block includes the above examples, but is not limited thereto.

[0262] After the decoding side determines the first reconstruction area, it determines the pixel average reconstruction value m based on the reconstruction value of the first reconstruction area.

[0263] In one implementation method, the average value of the reconstruction values of the first reconstruction area is determined as the pixel average reconstruction value m.

[0264] In one example, when the first reconstruction area is as shown in FIG. 20, the pixel average reconstruction value m may be calculated by the method shown in Table 14.

[0265] JPEG2024192733000022.jpg113150

[0266] In one example, if the first reconstruction region is the top row and / or leftmost column of the current block, the average of the reconstruction values ​​of the top row and / or leftmost column may be determined as the pixel-average reconstruction value m. In this case, the pixel-average reconstruction value m can be calculated using the method shown in Table 15.

[0267] JPEG2024192733000023.jpg160150

[0268] As shown in Table 15 above, if the first reconstruction region is the top row and / or leftmost column of the current block, the pixel average reconstruction value m can be calculated quickly by using a shift operation instead of division.

[0269] In addition to determining the pixel-average reconstruction value m as the average of the reconstruction values ​​in the first reconstruction region, the decoding side may also determine the pixel-average reconstruction value m using the method shown below.

[0270] In another method, the weighted average of the reconstruction values ​​in the first reconstruction region is determined as the pixel-average reconstruction value m.

[0271] The decoding side may also determine the pixel-averaged reconstruction value m using other methods.

[0272] The decoding side determines the pixel-average reconstruction value, and then performs average removal on the reconstruction values ​​of the pixel points in the reference region based on that pixel-average reconstruction value.

[0273] For example, for each pixel point in the reference region, the pixel value of the pixel point after averaging removal is obtained by dividing the reconstructed value of that pixel point by the above-mentioned average reconstructed value and then rounding it.

[0274] Furthermore, for example, the decoding side obtains the pixel value of a pixel point in the reference region after averaging removal by subtracting the pixel average reconstruction value from the reconstruction value of the pixel point in the reference region. For example, for each pixel point in the reference region, the pixel value of the pixel point in the reference region after averaging removal is obtained by subtracting the aforementioned pixel average reconstruction value from the reconstruction value of that pixel point.

[0275] The embodiments of this application are not limited to a specific method by which the decoding side performs average removal on the reconstructed values ​​of pixel points in the reference region based on the pixel-averaged reconstructed values.

[0276] The decoding side performs average removal on the reconstructed values ​​of the pixel points in the reference region based on the method described above, obtains the pixel values ​​of the pixel points after average removal in the reference region, and then executes the steps S101-A14 described above. The pixel values ​​of the pixel points after average removal in the reference region are used as input to the interpolation filter, and the interpolation filter is slid within the reference region to obtain the filter coefficients of the interpolation filter.

[0277] As an example, Figure 21 shows the process of obtaining the filter coefficients of the interpolation filter by sliding the interpolation filter of the current block on the reference region of the current block after averaging and removal, when the interpolation filter of the current block has five different shapes and the reference region of the current block has three different types. The interpolation filter may be slid horizontally one row at a time or vertically one column at a time on the reference region after averaging and removal. Note that the block to be predicted in Figure 21 is the current block.

[0278] JPEG2024192733000024.jpg68168

[0279] In one example, the Wiener-Hoff equation constructed by sliding the interpolation filter within the reference region of the current block is as shown in equation (5).

[0280] JPEG2024192733000025.jpg79169

[0281] JPEG2024192733000026.jpg38169

[0282] In one example, the decoding side adopts a method of decomposing the autocorrelation coefficient matrix by Cholesky decomposition, solves the Wiener-Hopf equation shown in the above formula (5), and can obtain the filter coefficients of the filter.

[0283] After the decoding side determines the filter coefficients of the interpolation filter based on the above formula (5), it executes the steps of S101-A2 described above, and performs interpolation filtering prediction on the current block using the interpolation filter based on the filter coefficients to obtain the predicted block of the current block.

[0284] In the above formula (5), since the filter coefficients are determined using the reference region after averaging removal, it is necessary to consider the influence of the pixel average reconstruction value m when determining the predicted value of the current block based on the filter coefficients.

[0285] In one possible implementation, substitute the interpolation filter coefficients determined by the above formula (5) into the above formula (4), obtain the predicted values of each point in the current block, and then add the pixel average reconstruction value m to the predicted values of each point to obtain the final predicted values of each point in the current block, and thus obtain the predicted block of the current block.

[0286] In another possible implementation, the above S101-A2 includes the following steps.

[0287] S101-A21. For the r-th point in the current block, determine the pixel values at N positions corresponding to the r-th point based on the shape of the interpolation filter, where r is a positive integer.

[0288] S101-A22. Perform averaging removal on the pixel values at N positions based on the pixel average reconstruction value to obtain the pixel values after averaging removal at N positions.

[0289] S101-A23 obtains the predicted value of the r-th point based on the pixel values ​​and filter coefficients after averaging and removing N positions.

[0290] S101-A24: Based on the predicted values ​​of each point within the current block, the predicted block for the current block is obtained.

[0291] As shown in Figure 22, assuming that the shape of the interpolation filter for the current block is 4x4, the decoding side sequentially performs interpolation predictions for each position within the current block using an interpolation filter with known filter coefficients. Specifically, for the r-th point within the current block, the pixel values ​​of the N positions corresponding to the r-th point are first determined based on the shape of the interpolation filter for the current block. For example, as shown in Figure 22, in a 4x4 interpolation filter, the dark positions are the positions of the r-th point to be processed, and the 15 light positions are the N positions corresponding to the r-th point. Here, the block to be predicted in Figure 22 is the current block.

[0292] Next, the pixel values ​​of the N positions corresponding to the r-th point are determined. For example, for any of the N positions, if that position is in the reconstruction area surrounding the current block, the reconstruction value of that position is determined as the pixel value of that position. If that position is within the current block, the predicted value of that position is determined as the pixel value of that position.

[0293] Since the filter coefficients described above are determined based on the reference region after averaging and removal, the decoding side performs averaging and removal on the pixel values ​​at N positions of the r-th point based on the pixel-averaged reconstruction value, and obtains the pixel values ​​at N positions of the r-th point after averaging and removal. For example, the pixel values ​​at N positions of the r-th point after averaging and removal are obtained by subtracting the pixel-averaged reconstruction value from the pixel values ​​at N positions of the r-th point.

[0294] Next, the predicted value of the r-th point is obtained based on the pixel values ​​and filter coefficients after averaging and removing the values ​​of N positions.

[0295] The embodiments of this application are not limited to a specific method for obtaining a predicted value of the r-th point based on the averaged and removed pixel values ​​and filter coefficients of N positions.

[0296] JPEG2024192733000027.jpg49169

[0297] In another implementation method, the above-mentioned S101-A23 includes the following steps.

[0298] S101-A231, currently, a second reconstruction region around the block is determined, and the maximum and minimum reconstruction values ​​of the second reconstruction region are determined.

[0299] S101-A232 obtains a first predicted value based on the pixel values ​​after averaging and removing N positions, the filter coefficients, and the pixel-averaged reconstruction value.

[0300] Based on S101-A233, the first predicted value, the maximum reconstruction value, and the minimum reconstruction value, the predicted value of the r-th point is determined.

[0301] JPEG2024192733000028.jpg27169

[0302] The embodiments of this application do not limit the specific method for determining the second reconstruction region around the current block.

[0303] In one example, the second reconfiguration region of the current block coincides with the reference region of the current block.

[0304] In one example, the second reconfiguration region of the current block coincides with the first reconfiguration region of the current block.

[0305] In one example, the reconstruction areas above, to the left, to the right, to the left, and to the left of the current block are determined as the second reconstruction area. For example, the reconstruction areas of the top 13 rows, left 13 columns, top 13 rows, top 13 rows and 13 columns of the top left, and bottom 13 columns of the current block are determined as the second reconstruction area.

[0306] Furthermore, the execution order is not limited in the specific implementation process of S101-A231 and S101-A232 described above. For example, S101-A231 may be executed before S101-A232, after S101-A232, or simultaneously with S101-A232.

[0307] The embodiments of this application are not limited to a specific method by which the decoding side obtains a first predicted value based on the pixel values ​​after averaging removal of N positions, the filter coefficients, and the pixel-averaged reconstruction value.

[0308] For example, the second predicted value for the r-th point is obtained by multiplying the pixel values ​​after averaging and removing the N positions of the r-th point by the filter coefficient, and the first predicted value for the r-th point is obtained by adding the second predicted value and the pixel-averaged reconstruction value.

[0309] For example, the decoding side obtains the first predicted value of the r-th point based on the following equation (6).

[0310] JPEG2024192733000029.jpg50167

[0311] Furthermore, for example, the decoding side obtains one predicted value for the r-th point based on equation (6) described above, then performs a prediction process on that predicted value to obtain the first predicted value for the r-th point.

[0312] Based on the steps described above, the decoding side determines the first predicted value for the r-th point in the current block, and then determines the predicted value for the r-th point based on the first predicted value, the maximum reconstruction value, and the minimum reconstruction value.

[0313] For example, if the first predicted value is greater than the minimum reconstruction value and less than the maximum reconstruction value, the first predicted value is determined to be the predicted value for the r-th point.

[0314] Furthermore, for example, if the first predicted value is less than or equal to the minimum reconstruction value, the minimum reconstruction value is determined as the predicted value for the r-th point.

[0315] Furthermore, for example, if the first predicted value is greater than or equal to the maximum reconstruction value, the maximum reconstruction value is determined as the predicted value for the r-th point.

[0316] In one example, the decoding side determines the predicted value of the r-th point based on the following equation (7).

[0317] JPEG2024192733000030.jpg38169

[0318] The above example illustrates the process of determining the predicted value of the r-th point in the current block. The decoding side refers to the above method to determine the predicted value of each point in the current block, and each predicted value in the current block constitutes the predicted block of the current block.

[0319] Based on the steps above, the decoding side performs interpolation filtering prediction on the current block, obtains the predicted block for the current block, and then performs the following steps.

[0320] S102, the intra-prediction mode corresponding to the prediction block is determined, and based on the intra-prediction mode corresponding to the prediction block, the transformation kernel corresponding to the current block is determined.

[0321] As is clear from the above explanation, when decoding the current block, the decoding side decodes the bitstream to obtain the quantization coefficients of the current block, then performs inverse quantization on the quantization coefficients to obtain the transformation coefficients of the current block, and then performs inverse transformation on the transformation coefficients of the current block to obtain the residual block (or residual value) of the current block. At the same time, the prediction mode of the current block is determined, a prediction is made on the current block using that prediction mode to obtain the predicted block of the current block, and the reconstructed block of the current block is obtained by adding the predicted block and the residual block.

[0322] When performing an inverse transformation on the transformation coefficients of the current block, it is necessary to determine the transformation kernel, and the residual value of the current block is obtained by performing an inverse transformation on the transformation coefficients of the current block based on the transformation kernel. Currently, the decoding side employs a conventional intra-prediction mode to make predictions on the current block, and the decoding side can determine the transformation kernel to be used for the current block based on the correspondence between the conventional intra-prediction mode and the transformation kernel. However, in the embodiment of the present invention, when making predictions on the current block, an interpolation filtering prediction mode is used instead of the conventional intra-prediction mode. Therefore, it is not possible to directly determine the transformation kernel corresponding to the current block based on the interpolation filtering prediction mode.

[0323] To solve this technical problem, in the embodiment of the present invention, after determining the predicted block of the current block using an interpolation filtering prediction mode, a conventional intra prediction mode corresponding to the predicted block is determined, and then a transformation kernel corresponding to the current block is determined based on the conventional intra prediction mode.

[0324] The following describes the specific process by which the decoding side determines the intra-prediction mode corresponding to the predicted block.

[0325] As an example, as shown in Figure 7, the conventional intra-prediction modes included in the current VVC are as follows: PLANAR mode: Intra prediction mode index is 0, DC mode: Intra prediction mode index is 1, Angle mode: Intra prediction mode index is 2-66.

[0326] In one example, as shown in Figure 23, the direction of the arrows in the figure indicates the direction of the angular mode prediction present in the VVC, where the prediction mode index used during decoding is 2 to 66. If the current block is a non-square block, some angular directions are replaced with wider angles, for example, -1 to -14 and 67 to 80 in Figure 23.

[0327] In some embodiments, the intra-prediction mode corresponding to the prediction block is the default intra-prediction mode. That is, if the current block employs the interpolation filtering prediction mode to perform prediction and obtains a prediction block, one of the conventional intra-prediction modes is determined by default as the intra-prediction mode corresponding to that prediction block.

[0328] In some embodiments, the decoding side determines the intra-prediction mode corresponding to the prediction block by following these steps:

[0329] S102-A1, determine the angle values ​​of M points in the prediction block, where M is a positive integer.

[0330] S102-A2 determines the intra-prediction mode corresponding to the prediction block based on the angle values ​​of M points.

[0331] In the embodiment of this invention, the intra-prediction mode corresponding to a prediction block is determined by statistically analyzing the intra-prediction modes corresponding to the angle values ​​of M points in the prediction block.

[0332] The embodiments of this application do not limit the specific location and number of M points for determining the angle value in the prediction block. For example, the M points may be one point in the prediction block or multiple points in the prediction block.

[0333] For example, if the above M points constitute a single point, the decoding side determines the angle value of one point within the prediction block (for example, the center point of the prediction block), determines the intra-prediction mode corresponding to that point based on the angle value of that point, and further determines that intra-prediction mode as the intra-prediction mode corresponding to the prediction block.

[0334] Furthermore, for example, if the M points mentioned above are multiple points, the decoding side determines the angle values ​​of these multiple points, determines the intra-prediction mode corresponding to each of these multiple points based on the angle values ​​of these multiple points, and then determines the intra-prediction mode that has the largest number of identical intra-prediction modes among these multiple points as the intra-prediction mode corresponding to the prediction block.

[0335] In some embodiments, when the angular values ​​of M points in a prediction block are determined by a sliding window method, the selection of these M points is related to the shape and size of the sliding window. For example, each of the M points is the center point of the sliding window as it slides within the prediction block.

[0336] In the embodiments of this application, the method for determining the angle value of each of the M points is the same. For simplicity of explanation, the case of determining the angle value of the i-th point among the M points will be explained as an example.

[0337] The embodiments of this application do not limit the specific method for determining the angle value of a point.

[0338] In some embodiments, step S102-A1 includes steps S102-A11 and S102-A12.

[0339] S102-A11, for the i-th point out of M points, determine the horizontal and vertical slopes of the i-th point, where i is a positive integer less than or equal to M.

[0340] S102-A12, the angle value of the i-th point is determined based on the horizontal and vertical slopes of the i-th point.

[0341] In this embodiment, the decoding side first determines the horizontal and vertical slopes of each of the M points (for example, the i-th point), and then determines the angle value of the i-th point based on the horizontal and vertical slopes.

[0342] The embodiments of this application do not limit the specific method for determining the horizontal and vertical slopes of the i-th point.

[0343] In one example, the horizontal gradient value of the i-th point is determined based on the horizontal change between the predicted values ​​of the points surrounding the i-th point within the prediction block and the predicted value of the i-th point itself, and the vertical gradient value of the i-th point is determined based on the vertical change between the predicted values ​​of the points surrounding the i-th point within the prediction block and the predicted value of the i-th point itself.

[0344] In another example, the decoding side determines the predicted value of a point within a sliding window centered on the i-th point within the prediction block, and obtains the horizontal and vertical gradients of the i-th point based on the predicted value of the point within the sliding window, the horizontal gradient operator, and the vertical gradient operator.

[0345] In this example, first, the sliding window is determined, for example, a 3x3 sliding window as shown in Figure 24. The sliding window is slid within the prediction block, and each time it slides, the horizontal and vertical slopes of the center point of the sliding window are determined. To illustrate with an example where the current center point of the sliding window is the i-th point, first, the predicted values ​​for each point in the current sliding window are obtained, for example, 3x3 = 9 predicted values ​​can be obtained. Next, the horizontal and vertical slopes of the i-th point are determined based on the predicted values ​​of these 9 points and the pre-set horizontal and vertical slope operators.

[0346] JPEG2024192733000031.jpg28169

[0347] JPEG2024192733000032.jpg38169

[0348] The embodiments of this application do not limit the specific values ​​of the horizontal gradient operator and the vertical gradient operator.

[0349] JPEG2024192733000033.jpg29135

[0350] The decoding side may determine the horizontal and vertical slopes of the i-th point based on the above steps, and then determine the angle value of the i-th point based on the horizontal and vertical slopes of the i-th point.

[0351] For example, the arctangent value of the ratio of the vertical slope to the horizontal slope at the i-th point is determined as the angle value of the i-th point. This is illustrated in equation (8).

[0352] JPEG2024192733000034.jpg49168

[0353] The decoding side may determine the angle value of the i-th point using a method other than equation (8) above. For example, the decoding side may obtain the angle value of the i-th point by adjusting the angle value determined by equation (8) above.

[0354] The decoding side employs the above method for each of the M points, determines the angle value for each of the M points, and then executes S102-A2 to determine the intra-prediction mode corresponding to the prediction block based on the angle values ​​of the M points.

[0355] The embodiments of this application are not limited to a specific method for determining an intra-prediction mode corresponding to a prediction block based on the angular values ​​of M points.

[0356] In some embodiments, the decoding side selects the angle value 1 that appears most frequently from among the angle values ​​of M points, matches this angle value 1 with the predicted angle of the conventional intra prediction mode to obtain the intra prediction mode corresponding to this angle value 1, and further determines the intra prediction mode corresponding to this angle value 1 as the intra prediction mode corresponding to the prediction block.

[0357] In some embodiments, the above S102-A2 includes the following steps S102-A21 and S102-A22.

[0358] S102-A21 determines the intra-prediction mode corresponding to M points based on the angle values ​​of M points.

[0359] S102-A22 determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction modes corresponding to M points.

[0360] In this implementation method, the decoding side determines the intra-prediction mode corresponding to each of the M points based on the angle value of each point. For example, for each of the M points, the intra-prediction mode corresponding to the angle value of that point is obtained by matching the angle value of that point with the predicted angle of a conventional intra-prediction mode. This makes it possible to obtain the intra-prediction mode corresponding to each of the M points.

[0361] Next, based on the intra-prediction modes corresponding to each of these M points, the intra-prediction mode corresponding to the prediction block is determined.

[0362] In one possible implementation, among the intra-prediction modes that each of the M points corresponds to, the intra-prediction mode that is repeated most frequently is determined as the intra-prediction mode corresponding to the prediction block.

[0363] In another possible implementation, S102-A22 includes the following steps:

[0364] S102-A221, based on the horizontal and vertical gradients of M points, the gradient amplitude values ​​corresponding to M points are determined.

[0365] S102-A222 determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction mode and gradient amplitude value corresponding to M points.

[0366] In this implementation method, the decoding side determines the gradient amplitude value corresponding to each of the M points based on the horizontal and vertical gradients of each of the M points determined above.

[0367] In the embodiments of this application, the specific method by which the decoding side determines the gradient amplitude value corresponding to each of the M points is the same. For simplicity of explanation, the case in which the gradient amplitude value corresponding to the i-th point among the M points is determined will be described as an example.

[0368] The embodiments of this application do not limit the specific method by which the decoding side determines the gradient amplitude value corresponding to the i-th point based on the horizontal and vertical gradients of the i-th point.

[0369] For example, the decoding side multiplies the horizontal gradient and vertical gradient of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.

[0370] Furthermore, for example, the decoding side adds the absolute values ​​of the horizontal and vertical gradients of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.

[0371] For example, the decoding side determines the gradient amplitude value corresponding to the i-th point based on the following equation (9).

[0372] JPEG2024192733000035.jpg49168

[0373] Based on the above steps, the decoding side can determine the gradient amplitude value corresponding to each of the M points. Next, the decoding side performs S102-A222 and determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction mode and gradient amplitude value corresponding to the M points.

[0374] In one example, the intra-prediction mode corresponding to the point with the maximum gradient amplitude among the M points is determined as the intra-prediction mode corresponding to the prediction block.

[0375] In another example, for any of the M points, the gradient amplitude value corresponding to that point is accumulated in the intra-prediction mode corresponding to that point, and the cumulative gradient amplitude value of the intra-prediction modes corresponding to the M points is obtained. Among the intra-prediction modes corresponding to the M points, the intra-prediction mode with the largest cumulative gradient amplitude value is determined as the intra-prediction mode corresponding to the prediction block.

[0376] As an example, as shown in Figure 25, the gradient amplitude values ​​corresponding to each of the M points are accumulated in the corresponding intra-prediction mode. For example, if the intra-prediction modes corresponding to points 1 and 2 among the M points are both intra-prediction mode 1, the gradient amplitude values ​​corresponding to points 1 and 2 are accumulated and added to the gradient amplitude value corresponding to intra-prediction mode 1. By this analogy, the gradient amplitude value histogram shown in Figure 25 can be obtained. This allows the intra-prediction mode with the largest accumulated gradient amplitude value in the gradient amplitude value histogram to be determined as the intra-prediction mode corresponding to the prediction block. For example, the intra-prediction mode corresponding to the accumulated gradient amplitude value shown in dark color in Figure 25 is determined as the intra-prediction mode corresponding to the prediction block.

[0377] In some embodiments, if the gradient amplitude values ​​corresponding to M points are all 0, the first intra-prediction mode is determined as the intra-prediction mode corresponding to the prediction block. That is, if the gradient amplitude values ​​corresponding to all M points are all 0, it means that the horizontal and vertical gradients of each of the M points are both 0. In this case, a pre-set first intra-prediction mode may be determined as the intra-prediction mode corresponding to the prediction block.

[0378] The embodiments of this application are not limited to the first intra-prediction mode described above.

[0379] For example, the first intra-prediction mode described above is the PLANAR mode.

[0380] The decoding side determines the intra-prediction mode corresponding to the prediction block based on the above steps, and then determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to that prediction block.

[0381] The embodiments of this application do not limit the specific method by which the decoding side determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block.

[0382] In some embodiments, the decoding side searches for an image block whose intra-prediction mode is the same as the intra-prediction mode corresponding to the prediction block, based on the intra-prediction mode corresponding to the prediction block, among the already decoded image blocks surrounding the prediction block, and then determines the transformation kernel corresponding to that image block as the transformation kernel corresponding to the current block.

[0383] In some embodiments, the step of determining the translation kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block in S102 above includes the following steps.

[0384] S102-B1, the correspondence between the intra prediction mode and the transformation kernel group is obtained, where one transformation kernel group includes at least one type of transformation kernel.

[0385] S102-B2, within the above correspondence, looks up the first transformation kernel group corresponding to the intra-prediction mode of the prediction block.

[0386] S102-B3: From the first group of translation kernels, determine the translation kernel corresponding to the current block.

[0387] In the embodiment of the present invention, a correspondence exists between the intra-prediction mode and the conversion kernel group. Based on this, the decoding side determines the intra-prediction mode corresponding to the prediction block, and then obtains the pre-configured correspondence between the intra-prediction mode and the conversion kernel group.

[0388] For example, the correspondence between the intra-prediction mode and the conversion kernel group is shown in Table 16.

[0389] JPEG2024192733000036.jpg177160

[0390] Table 16 above merely shows the correspondence between the intra-prediction mode and the conversion kernel group in the embodiment of the present application, and the correspondence between the intra-prediction mode and the conversion kernel group in the embodiment of the present application is not limited to those shown in Table 16.

[0391] Here, each translation kernel group includes at least one type of translation kernel.

[0392] The decoding side obtains the correspondence between intra-prediction modes and transformation kernel groups shown in Table 16. Based on the intra-prediction mode corresponding to the prediction block, it looks up the transformation kernel group corresponding to the intra-prediction mode from the correspondence between intra-prediction modes and transformation kernel groups, and denotes this transformation kernel group as the first transformation kernel group. For example, if the intra-prediction mode corresponding to the prediction block is an angle prediction mode in 64 angular directions, looking up Table 16 above will reveal that the transformation kernel group corresponding to that angle prediction mode in 64 angular directions is 4. As a result, the decoding side determines the transformation kernel corresponding to the current block from at least one type of transformation kernel included in transformation kernel group 4.

[0393] For example, if the first translation kernel group contains one translation kernel, that translation kernel is determined to be the translation kernel corresponding to the current block.

[0394] Furthermore, for example, if the first transformation kernel group contains multiple types of transformation kernels, the decryption side determines the type of transformation kernel corresponding to the current block, and then determines the transformation kernel of that type in the first transformation kernel group as the transformation kernel corresponding to the current block.

[0395] Here, the method by which the decryption side determines the type of transformation kernel corresponding to the current block is not limited to those shown below.

[0396] In one example, the type of translation kernel corresponding to the current block is the default type. In this case, the decryption side determines that this default type is the type of translation kernel corresponding to the current block.

[0397] In another example, the encoding side writes the type of transformation kernel corresponding to the current block into the bitstream. The decoding side then decodes the bitstream to obtain the type of transformation kernel corresponding to the current block.

[0398] As is clear from the above description, in the embodiment of the present invention, the decoding side uses an interpolation filtering prediction mode to determine the predicted block of the current block, then determines the conventional intra-prediction mode corresponding to that predicted block, and then determines the transform kernel corresponding to the current block based on the conventional intra-prediction mode corresponding to that predicted block. That is, in the embodiment of the present invention, the conventional intra-prediction mode derived from the interpolation filtering prediction is used to select the transform kernel group of Non-separable primary transform (NSPT) and Low Frequency non-separable secondary transform (LFNST). This allows the determined transform kernel to better suit the characteristics of the current block, improving the accuracy of the transform kernel determination. By determining the reconstruction value of the current block using a transform kernel with this accuracy, the accuracy of the reconstruction value can be improved, and the decoding accuracy of the current block can be increased. Furthermore, in the embodiment of the present invention, when the transform kernel of the current block is determined via the conventional prediction mode corresponding to the predicted block, it is not necessary to specify the transform kernel individually, saving codewords and further improving the video coding and decoding effect.

[0399] After determining the conversion kernel corresponding to the current block based on the above steps, the decryption side executes the following step S103.

[0400] S103, based on the transformation kernel corresponding to the current block, an inverse transformation is performed on the transformation coefficients of the current block to obtain the residual block of the current block, and further, based on the predicted block and residual block of the current block, the reconstructed block of the current block is obtained.

[0401] In the embodiment of the present invention, the decoding side determines the predicted block of the current block and the transformation kernel corresponding to the current block based on the above steps. As a result, the decoding side decodes the bitstream to obtain the quantization coefficients of the current block, then performs inverse quantization on the quantization coefficients to obtain the transformation coefficients of the current block, and performs inverse transformation on the transformation coefficients using the transformation kernel corresponding to the current block determined above to obtain the residual block (or residual value) of the current block. Finally, the decoding side adds the predicted block and the residual block of the current block to obtain the reconstructed block of the current block.

[0402] In some embodiments, the current block is either a luminance block or a chroma block. That is, in embodiments of the present application, prediction can be performed for either a luminance block or a chroma block using the interpolation filtering prediction mode provided by embodiments of the present application.

[0403] In some embodiments, when the current block is a luminance block, the prediction mode of the current block is an interpolation filtering prediction mode, and the chroma block corresponding to the current block employs a direct derivation mode DM, the PLANAR mode or the intra-prediction mode corresponding to the prediction block is determined as the prediction mode of the chroma block.

[0404] In this embodiment, prediction may be performed on the luminance block (or luminance component) using the interpolation filtering prediction mode provided in the embodiment of the present application. On the other hand, prediction may be performed on the chroma block (or chroma component) using other intra prediction modes.

[0405] Specifically, the decoding side performs prediction and decoding on the current block (i.e., the luminance block) using the interpolation filtering prediction mode described above, and then starts predicting and decoding the chroma block corresponding to that current block (i.e., the luminance block). When performing prediction and decoding on a chroma block, the decoding side first determines the prediction mode used by the chroma block, for example, by decoding the bitstream to obtain the prediction mode of the chroma block. In one example, if it is determined that the chroma block adopts the direct derivation DM mode, the decoding side derives the intra-prediction mode of the chroma block based on the intra-prediction mode of the luminance block.

[0406] For example, if the current block (i.e., the luminance block) is using interpolation filtering prediction mode and the chroma block is using DM mode, then PLANAR mode is determined as the prediction mode for the chroma block, prediction is performed on the chroma block, and the predicted value for the chroma block is obtained.

[0407] In another example, if the current block (i.e., the luminance block) employs interpolation filtering prediction mode and the chroma block employs DM mode, the intra prediction mode corresponding to the prediction block of the current block determined above is determined as the prediction mode for the chroma block, prediction is performed on the chroma block, and the predicted value of the chroma block is obtained.

[0408] In the video decoding method provided by the embodiment of the present application, when making a prediction for the current block, first, the reference region and interpolation filter of the current block are determined, and the predicted block of the current block is determined based on the reference region and interpolation filter. For example, filtering is applied to the reference region using the interpolation filter, the filter coefficients of the filter are calculated, and interpolation filtering prediction of the current block is performed using the interpolation filter with the determined filter coefficients to obtain the predicted block of the current block. Next, the prediction mode corresponding to the predicted block is determined, and a transformation kernel corresponding to the current block is determined based on the prediction mode. An inverse transformation is performed on the transformation coefficients of the current block using the transformation kernel to obtain the residual block of the current block, and the reconstructed value of the current block is obtained based on the residual block and the predicted block of the current block. That is, in the embodiment of the present application, when making a prediction by applying the interpolation filtering prediction method to the current block, the transformation kernel corresponding to the current block can be derived by determining the conventional prediction mode corresponding to the predicted block. As a result, the determined transformation kernel is better suited to the characteristics of the current block, and the accuracy of the transformation kernel determination is improved. When the reconstructed value of the current block is determined using the transformation kernel determined with this accuracy, the accuracy of the reconstructed value can be improved, and the decoding accuracy of the current block can be increased. Furthermore, in the embodiments of this application, since the transformation kernel of the current block is determined via a conventional prediction mode corresponding to the prediction block, there is no need to individually specify the transformation kernel, which saves code and further improves the video coding and decoding effect.

[0409] The prediction method of this application was explained above using the decoding side as an example, but below it will be explained using the encoding side as an example.

[0410] Figure 26 is a flowchart of a prediction method according to one embodiment of the present invention, and the embodiment of the present invention is applied to the video encoder shown in Figures 1 and 2. As shown in Figure 26, the method of the embodiment of the present invention includes the following steps.

[0411] S201 determines the reference region and interpolation filter of the current block, and determines the predicted block of the current block based on the reference region and interpolation filter.

[0412] When encoding the current block, the encoding process first determines the prediction mode for the current block, uses that prediction mode to make a prediction for the current block, and obtains the predicted block (or predicted value) of the current block. Subtracting the predicted block of the current block from the current block obtains the residual block (or residual value) of the current block. Next, a transformation is performed on the residual block of the current block to obtain the transformation coefficients, quantization is performed on the transformation coefficients to obtain the quantization coefficients, and the quantization coefficients are encoded to obtain the bitstream.

[0413] In the embodiment of the present invention, the encoding side first determines the prediction mode of the current block.

[0414] In some embodiments, the method by which the encoding side determines the prediction mode of the current block includes at least the following:

[0415] Method 1: The encoding side selects the candidate prediction mode with the lowest cost from among multiple candidate prediction modes, which consist of the conventional prediction mode and interpolation filtering prediction mode shown in Figure 6 or Figure 7, as the prediction mode for the current block. Next, the encoding side adds the instruction information for the prediction mode of the current block to the bitstream. The decoding side then decodes the bitstream to obtain the instruction information for the prediction mode of the current block and further determines the prediction mode for the current block based on this instruction information.

[0416] Method 2: The encoding side constructs an intra-prediction mode candidate list and selects the intra-prediction mode for the current block from this intra-prediction mode candidate list. Note that this intra-prediction mode candidate list includes interpolation filtering prediction modes. Next, the encoding side writes the sequence number (or index number) of the current block's intra-prediction mode in the intra-prediction mode candidate list to the bitstream.

[0417] Method 3: The encoding side constructs an intra-prediction mode candidate list, which includes interpolation filtering prediction modes. Next, it selects an intra-prediction mode for the current block from this intra-prediction mode candidate list. For example, it determines the cost of each candidate prediction mode in the intra-prediction mode candidate list in the current block template, and then determines the intra-prediction mode for the current block based on the cost. Correspondingly, the decoding side constructs an intra-prediction mode candidate list based on the same method as the encoding side, which further includes interpolation filtering prediction modes. Next, it determines the cost of each candidate prediction mode in the intra-prediction mode candidate list in the current block template, and then determines the intra-prediction mode for the current block based on the cost. Finally, it makes a prediction for the current block using the determined intra-prediction mode for the current block and obtains the predicted value for the current block.

[0418] As can be seen from the above methods, when the encoding side determines the prediction mode for the current block, it first determines several candidate prediction modes, and then determines the prediction mode for the current block from among these several candidate prediction modes. Here, the multiple candidate prediction modes include interpolation filtering prediction modes.

[0419] Here, the specific method for determining the prediction mode for the current block from among multiple candidate prediction modes may be for the encoding side to decide on any candidate prediction mode from among the multiple candidate prediction modes as the prediction mode for the current block. That is, the encoding side makes a prediction for the current block using each of the multiple candidate prediction modes, determines the cost corresponding to each candidate prediction mode, and this cost may be RDO or SATD, etc., and further determines the candidate prediction mode with the minimum cost as the prediction mode for the current block.

[0420] The encoding side determines the prediction mode of the current block based on the above method, and if the determined prediction mode of the current block is the interpolation filtering prediction mode, it executes the step in S201 above.

[0421] In some embodiments, before determining the prediction mode for the current block from among multiple candidate prediction modes, the encoding side needs to determine whether the position of the current block in the current image satisfies a predetermined position requirement, and whether the size of the current block satisfies a predetermined block size requirement.

[0422] The embodiments of this application are determined, but are not limited to, the pre-set positions and predicted block sizes, which are specifically determined according to actual needs.

[0423] In one example, as shown in Figure 11, the position of the upper left corner of the current image is (0,0), and the position of the upper left corner of the current block is (x,y). Here, the predetermined position requirement is that the x of the current block is greater than or equal to the first predetermined value XX, and the y of the current block is greater than or equal to the second predetermined value YY.

[0424] The embodiments of this application are not limited to specific numerical values ​​for the first and second predetermined values ​​described above.

[0425] For example, the first predetermined value and the second predetermined value are the same.

[0426] For example, both the first predetermined value and the second predetermined value are 13. That is, if the distance from the top edge of the current block to the top edge of the current image is 13 rows of pixels or more, and the distance from the left edge of the current block to the left edge of the current image is 13 columns of pixels or more, then the position of the current block in the current image satisfies the predetermined position requirement.

[0427] For example, continuing to refer to Figure 11, if the current block width is W and the current block height is H, then the required block size is that the current block width W is less than or equal to the third predetermined value A, and the current block height H is less than or equal to the fourth predetermined value B.

[0428] The embodiments of this application are not limited to specific numerical values ​​for the third and fourth predetermined values ​​described above.

[0429] For example, the third predetermined value and the fourth predetermined value are the same.

[0430] For example, both the third and fourth predetermined values ​​are 32. That is, if the current block's width and height are both 32 or less, it indicates that the current block meets the predetermined block size requirement.

[0431] In the embodiment of the present invention, before determining whether the current block employs an interpolation filtering prediction mode for prediction, the decoding side first determines whether the position of the current block in the current image satisfies a predetermined position requirement, and whether the size of the current block satisfies a predetermined block size requirement. If the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the prediction mode for the current block is determined from among a plurality of candidate prediction modes, including the interpolation filtering prediction mode. For example, as shown in Figure 11, if the distance from the top edge of the current block to the top edge of the current image is 13 rows of pixels or more, the distance from the left edge of the current block to the left edge of the current image is 13 columns of pixels or more, and the width and height of the current block are both 32 or less, the prediction mode for the current block is determined from among a plurality of candidate prediction modes, including the interpolation filtering prediction mode.

[0432] In some embodiments, the first predetermined value, second predetermined value, third predetermined value, and fourth predetermined value are default values.

[0433] In some embodiments, if the position of the current block in the current image does not satisfy a predetermined position requirement, and / or the size of the current block does not satisfy a predetermined block size requirement, the encoding side determines that the prediction mode of the current block is not an interpolation filtering prediction mode. In this case, the encoding side determines the prediction mode of the current block from among candidate prediction modes that do not include the interpolation filtering prediction mode.

[0434] In some embodiments, the encoding side further includes the steps of determining whether the current sequence is permitted to perform predictions using an interpolation filtering prediction mode before determining whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement; and, if the current sequence is permitted to perform predictions using an interpolation filtering prediction mode, determining whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement.

[0435] In the embodiments of the present invention, a higher-level syntactic element indicates whether the current sequence is permitted to perform predictions using an interpolation-filtering prediction mode. If the current sequence is permitted to perform predictions using an interpolation-filtering prediction mode, the encoding side determines whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement. If it is determined that the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the encoding side determines the prediction mode for the current block from among candidate prediction modes, including the interpolation-filtering prediction mode.

[0436] In some embodiments, if the encoding side determines that the current sequence does not allow prediction using interpolation filtering prediction mode, the encoding side skips step S201 above.

[0437] In some embodiments, the encoding side writes a fourth piece of information to the bitstream. This fourth piece of information is used to indicate whether the current sequence is allowed to perform predictions using an interpolation filtering prediction mode.

[0438] The embodiments of this application may, but are not limited to, any directive information that can indicate whether the current sequence is permitted to perform predictions using an interpolation filtering prediction mode, with respect to the specific representation format of the fourth information.

[0439] For example, the fourth piece of information may be represented as sps_eip_enabled_flag, where different values ​​of sps_eip_enabled_flag indicate whether the current sequence is allowed to perform predictions using interpolation filtering prediction mode. For example, sps_eip_enabled_flag=0 indicates that the current sequence is not allowed to perform predictions using interpolation filtering prediction mode, and sps_eip_enabled_flag=1 indicates that the current sequence is allowed to perform predictions using interpolation filtering prediction mode.

[0440] For example, the fourth piece of information is transported to the sequence-level parameter set (SPS).

[0441] In some embodiments, embodiments of the present application may further include a general constraints information (GCI) identification bit that indicates whether or not to use interpolation filtering prediction techniques. For example, gci_no_eip_constraint_flag indicates whether or not the current video enables interpolation filtering prediction techniques. For example, as shown in Table 2, the gci_no_eip_constraint_flag is carried to general constraints information general_constraints_info().

[0442] In some embodiments, the encoding side writes third information to the bitstream if it decides to allow the current sequence to perform predictions using interpolation filtering prediction mode. This third information is used to indicate whether the current block should perform predictions using interpolation filtering prediction mode.

[0443] The embodiments of this application may, but are not limited to, any instruction information that can indicate whether or not the current block employs an interpolation filtering prediction mode to perform predictions, regarding the specific representation format of the third information described above.

[0444] In one example, the third piece of information may be represented as intra_eip_flag, where different assigned values ​​of intra_eip_flag indicate whether the current block is performing predictions using interpolation filtering prediction mode. For example, if intra_eip_flag=0, it indicates that the current block is performing predictions without employing interpolation filtering prediction mode, and if intra_eip_flag=1, it indicates that the current block is performing predictions using interpolation filtering prediction mode. The encoding side then writes the predetermined flag intra_eip_flag to the bitstream, and the decoding side determines the prediction mode of the current block based on the value of the decoded predetermined flag intra_eip_flag. For example, if the predetermined flag intra_eip_flag=1, it indicates that the prediction mode of the current block is interpolation filtering prediction mode, and the decoding side performs predictions for the current block using interpolation filtering prediction mode.

[0445] In some embodiments, as shown in Figure 27, the process for determining the prediction mode of the current block in embodiments of the present invention may include the following steps: First, it is determined whether the current block employs an interpolation filtering prediction mode for prediction. For example, if the fourth information of the sequence level indicates that the current sequence allows the use of an interpolation filtering prediction mode, and the position of the current block in the current image satisfies a predetermined position requirement, and the size of the current block satisfies a predetermined block size requirement, it is determined that the current block can employ an interpolation filtering prediction mode for prediction. Next, filter coefficients are obtained, and the current block is predicted based on these filter coefficients to obtain the predicted value of the current block. Simultaneously, a rough selection of prediction modes is performed together with other intra-prediction mode tools, and several prediction modes with relatively low costs are selected and further refined to determine the final intra-prediction mode, which is then used as the prediction mode for the current block. If it is determined that the current block cannot employ an interpolation filtering prediction mode for prediction, the selection of the interpolation filtering prediction mode is skipped.

[0446] For example, in the rough selection stage of the prediction modes for the current block, the encoding side calculates the cost of each candidate intra-prediction mode (including interpolation filtering prediction modes). The formula for calculating the cost is shown in equation (10).

[0447] JPEG2024192733000037.jpg50168

[0448] For example, the method for calculating the strain value D is as shown in equation (11).

[0449] JPEG2024192733000038.jpg69168

[0450] The encoding side determines the cost of each candidate prediction mode, and then performs refinement by selecting several more candidate prediction modes from among the multiple candidate prediction modes.

[0451] JPEG2024192733000039.jpg59168

[0452] The encoding side determines the candidate prediction mode that minimizes cost during the refining process as the prediction mode for the current block.

[0453] If the encoding side determines that the prediction mode for the current block is the interpolation filtering prediction mode, it executes step S101 described above.

[0454] The following describes the process by which the encoding side makes predictions for the current block using the interpolation filter prediction mode.

[0455] If the encoding side decides that the current block will perform predictions using the interpolation filter prediction mode, it first determines the reference region and interpolation filter of the current block.

[0456] The following describes the specific process by which the encoding side determines the reference region of the current block.

[0457] In the embodiments of the present application, the reference region of the current block is part or all of the already reconfigured region surrounding the current block.

[0458] For example, as shown in Figure 12, the reconfiguration region surrounding the current block may include the upper reconfiguration region of the current block, the left reconfiguration region of the current block, the upper right reconfiguration region of the current block, the lower left reconfiguration region of the current block, and the upper left reconfiguration region of the current block.

[0459] The embodiments of this application do not currently limit the specific shape and size of the reference region of the block.

[0460] For example, the reference region of the current block includes any one of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, or the upper left reconstruction region of the current block. For example, the reference region of the current block may be the upper reconstruction region of the current block, or the reference region of the current block may be the left reconstruction region of the current block.

[0461] For example, the reference region of the current block includes any two of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region and the left reconstruction region of the current block. Also, for example, the reference region of the current block includes the upper reconstruction region and the lower left reconstruction region of the current block.

[0462] For example, the reference region of the current block includes any three of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region of the current block, the upper right reconstruction region of the current block, and the upper left reconstruction region of the current block. Also, for example, the reference region of the current block includes the left reconstruction region of the current block, the upper left reconstruction region of the current block, and the lower left reconstruction region of the current block.

[0463] For example, the reference region of the current block includes any four of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region of the current block, the upper right reconstruction region of the current block, the upper left reconstruction region of the current block, and the left reconstruction region of the current block. Also, for example, the reference region of the current block includes the left reconstruction region of the current block, the upper left reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper reconstruction region of the current block.

[0464] In one example, the reference region of the current block includes all five reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block.

[0465] In the embodiments of this application, the specific methods by which the encoding side determines the reference region of the current block include, but are not limited to, those described below.

[0466] Method 1: The current block's reference region is used as the default region. For example, the encoding and decoding sides are set to default to the current block's reference region including at least one of the following: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block.

[0467] Method 2: Determine the first cost for making a prediction for the current block based on P reference regions, and determine the reference region with the minimum first cost among the P reference regions as the reference region for the current block.

[0468] In this implementation, the encoding side predicts the current block based on these P reference regions, determines a first cost corresponding to each reference region, and selects the reference region with the minimum first cost among these P reference regions as the reference region of the current block.

[0469] In some embodiments, the encoder writes first information to the bitstream. This first information is used to indicate the type of reference region of the current block. That is, in scheme 2, the encoder further indicates the determined type of reference region of the current block to the decoder via the first information.

[0470] Furthermore, these pre-configured P reference regions will each be of a different type or shape.

[0471] The embodiments of this application do not specifically limit the number and shape of the P reference regions.

[0472] In one example, the P reference regions include at least one of the first, second, and third reference regions.

[0473] Here, the first reference region shown in Figure 13A includes the reconstruction regions above, to the upper right, to the left, to the lower left, and to the upper left of the current block. The second reference region shown in Figure 13B includes the reconstruction regions above, to the upper right, and to the upper left of the current block. The third reference region shown in Figure 13C includes the reconstruction regions to the left, to the lower left, and to the upper left of the current block.

[0474] The embodiments of this application do not limit the specific representation format of the first information. Any directive information that can indicate the type of reference region of the current block is acceptable.

[0475] In one example, the first piece of information is represented as eip_ref_type. For instance, different types of reference regions are indicated depending on the value of eip_ref_type.

[0476] As a specific example, the correspondence between the three reference regions shown in Figures 13A and 13B and the value of eip_ref_type is as shown in Table 4.

[0477] Based on Table 4 above, the encoding side determines the value of the first information eip_ref_type based on the reference region type of the current block. For example, if the reference region of the current block is determined to be the first reference region, eip_ref_type is determined to be 0. If the reference region of the current block is determined to be the second reference region, eip_ref_type is determined to be 1. If the reference region of the current block is determined to be the third reference region, eip_ref_type is determined to be 2.

[0478] In the above explanation, we have used the example of the P reference regions being the three reference regions shown in Figures 13A to 13C. However, the P reference regions in the embodiment of this application may include other reference regions besides the three reference regions mentioned above, but the embodiment of this application is not limited to this. The correspondence between the reference regions shown in Table 4 above and the value of eip_ref_type can be adjusted as appropriate depending on the number of reference regions.

[0479] In some embodiments, the encoding side may employ a langue-bound binary code encoding scheme to write the first information to the bitstream.

[0480] For example, the correspondence between truncated binary code, the value of eip_ref_type, and the type of reference area is shown in Table 5.

[0481] In the embodiments of the present invention, the encoding side may employ an equiprobability coding scheme, a context model coding scheme, or codewords of truncated binary code.

[0482] In addition to determining the reference region of the current block using method 1 or method 2 described above, the encoding side may also determine the reference region of the current block using method 3 shown below.

[0483] Method 3: Based on the shape of the current block, the reference region of the current block is determined from among the P pre-set reference regions.

[0484] Method 3 improves the accuracy of predictions by using different reference regions for each current block with a different shape.

[0485] For example, if the current block shape is a square, use the first type of reference region.

[0486] Additionally, for example, if the current block shape is a rectangle where the width is greater than the height, a second type of reference area is used.

[0487] Furthermore, for example, if the current block shape is a rectangle where the width is less than the height, a third type of reference area is used.

[0488] In other words, in the embodiment of the present application, the correspondence between the P reference regions and the shape of the current block is predetermined. As a result, the encoding side may determine the reference region of the current block from among the P reference regions based on the correspondence between the P reference regions and the shape of the current block.

[0489] The following describes the process by which the encoding side determines the interpolation filter for the current block.

[0490] In the embodiments of this application, the specific shape of the interpolation filter is not limited.

[0491] Exemplary interpolation filters provided by embodiments of the present application include, but are not limited to, square interpolation filters and interpolation filters where the height is less than the width.

[0492] For example, a square interpolation filter includes, but is not limited to, the 4x4 interpolation filter shown in Figure 14A.

[0493] Furthermore, interpolation filters where the height is greater than the width include, but are not limited to, the 5×3 interpolation filter shown in Figure 14B, the 6×2 interpolation filter shown in Figure 14D, and the 7×1 interpolation filter shown in Figure 14G.

[0494] Furthermore, interpolation filters where the height is smaller than the width include, but are not limited to, the 3×5 interpolation filter shown in Figure 14C, the 2×6 interpolation filter shown in Figure 14E, and the 1×7 interpolation filter shown in Figure 14F.

[0495] JPEG2024192733000040.jpg20168

[0496] In the embodiments of this application, specific methods for determining the interpolation filter of the current block by the decoding side include, but are not limited to, those described below.

[0497] Method 1: The interpolation filter for the current block is set as the default interpolation filter. For example, the default setting for the encoding side and the decoding side is to set the interpolation filter for the current block to any one of the interpolation filters shown in Figures 14A to 14G. For example, the default interpolation filter is a 4x4 interpolation filter.

[0498] Method 2: The encoding side determines the interpolation filter for the current block from among the Q interpolation filters that have been set in advance.

[0499] For example, the encoding side randomly selects one interpolation filter from among Q interpolation filters and uses it as the interpolation filter for the current block.

[0500] Furthermore, for example, the encoding side determines the second cost for each of the Q interpolation filters used to make a prediction for the current block, and then selects the interpolation filter with the minimum second cost among the Q interpolation filters as the interpolation filter for the current block.

[0501] In some embodiments, the encoding side writes second information to the bitstream. This second information is used to indicate the shape of the interpolation filter for the current block.

[0502] In this implementation, the encoding side determines the interpolation filter for the current block from among Q pre-configured interpolation filters. For example, the encoding side determines a second cost corresponding to each of these Q interpolation filters and determines the interpolation filter with the minimum second cost as the interpolation filter for the current block. Next, the shape of the interpolation filter with the minimum determined second cost is instructed to the encoding side via second information. As a result, the decoding side decodes the bitstream to obtain the second information and then determines the interpolation filter for the current block from among the Q pre-configured interpolation filters based on the shape of the interpolation filter indicated by the second information.

[0503] Furthermore, these pre-configured Q interpolation filters will each have a different shape.

[0504] The embodiments of this application do not specifically limit the number and shape of the Q interpolation filters. For example, the Q interpolation filters include at least one of a first interpolation filter, a second interpolation filter, and a third interpolation filter, wherein the first interpolation filter is a square interpolation filter, the second interpolation filter is a rectangular interpolation filter whose width is greater than its height, and the third interpolation filter is a rectangular interpolation filter whose height is greater than its width.

[0505] In one example, Q interpolation filters include multiple interpolation filters as shown in Figures 14A to 14G.

[0506] The embodiments of this application do not limit the specific representation format of the second information. Any instruction information that can specify the shape of the interpolation filter of the current block is acceptable.

[0507] In one example, the second piece of information is represented as `eip_filter_type`. For instance, the value of `eip_filter_type` specifies an interpolation filter of a different shape.

[0508] For example, if the Q interpolation filters are the five interpolation filters shown in Figure 15, the correspondence between the five interpolation filters and the value of eip_filter_type is as shown in Table 6.

[0509] Based on Table 6 above, the encoding side determines the value of the second information eip_filter_type based on the shape of the interpolation filter of the current block that has been determined. For example, if the shape of the interpolation filter of the current block is determined to be 4x4, then eip_filter_type=0 is determined. If the shape of the interpolation filter of the current block is determined to be 3x5, then eip_filter_type=1 is determined. If the shape of the interpolation filter of the current block is determined to be 5x3, then eip_filter_type=2 is determined. If the shape of the interpolation filter of the current block is determined to be 2x6, then eip_filter_type=3 is determined. If the shape of the interpolation filter of the current block is determined to be 6x2, then eip_filter_type=4 is determined.

[0510] In some embodiments, the encoding side may employ a truncated binary code encoding scheme to incorporate the second information into the bitstream.

[0511] For example, if the pre-configured Q interpolation filters include the five interpolation filters shown in Figure 15, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 7.

[0512] In this case, combining the five types of interpolation filter shapes shown in Table 7 above with the three types of reconstruction region types shown in Table 5 above results in a total of 15 possible combinations of interpolation filters and reconstruction regions.

[0513] In some embodiments, when the embodiment of the present application includes the seven interpolation filters shown in Figure 16, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 9.

[0514] In this case, combining the seven types of interpolation filter shapes shown in Table 9 above with the three types of reconstruction region types shown in Table 5 above results in a total of 21 possible combinations of interpolation filters and reconstruction regions.

[0515] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in Figure 17, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 10.

[0516] In this case, combining the three types of interpolation filter shapes shown in Table 10 above with the three types of reconstruction region types shown in Table 5 above results in a total of nine possible combinations of interpolation filters and reconstruction regions.

[0517] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in Figure 18A, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 11.

[0518] In this case, by combining the three types of interpolation filter shapes shown in Table 11 above with the three types of reconstruction region types shown in Table 5 above, there are a total of nine possible combinations of interpolation filters and reconstruction regions.

[0519] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in Figure 18B, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 12.

[0520] In this case, combining the three types of interpolation filter shapes shown in Table 12 above with the three types of reconstruction region types shown in Table 5 above results in a total of nine possible combinations of interpolation filters and reconstruction regions.

[0521] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in Figure 19, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is as shown in Table 13.

[0522] In this case, combining the three types of interpolation filter shapes shown in Table 13 above with the three types of reconstruction region types shown in Table 5 above results in a total of nine possible combinations of interpolation filters and reconstruction regions.

[0523] In addition to determining the interpolation filter for the current block using method 1 or method 2 described above, the encoding side may also determine the interpolation filter for the current block using method 3 shown below.

[0524] Method 3: Based on the shape of the current block, the interpolation filter for the current block is determined from among Q pre-set interpolation filters.

[0525] In method 3, prediction accuracy is improved by using different interpolation filters for each current block with a different shape.

[0526] For example, if the current block shape is a square, use the interpolation filter for the first type of shape.

[0527] Additionally, for example, if the current block shape is a rectangle where the width is greater than the height, a second type of shape interpolation filter is used.

[0528] Furthermore, for example, if the current block shape is a rectangle where the width is less than the height, a third type of shape interpolation filter is used.

[0529] In other words, in the embodiment of the present invention, the correspondence between the Q interpolation filters and the shape of the current block is predetermined. As a result, the decoding side can determine the interpolation filter for the current block from among the Q interpolation filters based on the correspondence between the Q interpolation filters and the shape of the current block.

[0530] In the embodiment of the present invention, the encoding side determines the reference region and interpolation filter of the current block based on the steps described above, and then determines the predicted block of the current block based on the reference region and interpolation filter.

[0531] The following describes the process by which the encoding side determines the predicted block for the current block based on the reference region of the current block and the interpolation filter.

[0532] In the embodiment of the present invention, the encoding side determines the reference region and interpolation filter of the current block, then determines the filter coefficients of the interpolation filter by filtering the reference region using the interpolation filter, and further performs interpolation filtering on the current block based on the determined filter coefficients to obtain a predicted block of the current block.

[0533] The embodiments of this application do not limit the specific method by which the encoding side determines the predicted block of the current block based on the reference region of the current block and the interpolation filter.

[0534] In some embodiments, determining the predicted block of the current block based on the reference region and interpolation filter of the current block in S201 described above includes the following steps.

[0535] S201-A1, the filter coefficients of the interpolation filter are determined based on the reference region.

[0536] S201-A2, based on the filter coefficients, performs interpolation filtering prediction on the current block using an interpolation filter and obtains the predicted block for the current block.

[0537] The method for determining the filter coefficients of the interpolation filter in S201-A1 described above includes at least the method shown below.

[0538] Method 1: The interpolation filter determined above is slid across the reference region of the current block to construct the Wiener-Hoff equations. Next, the filter coefficients of the interpolation filter are obtained by solving the Wiener-Hoff equations.

[0539] JPEG2024192733000041.jpg68168

[0540] In one example, the Wiener-Hoff equation constructed by sliding the interpolation filter within the reference region of the current block is as shown in equation (3).

[0541] Since the reference region of the current block is the reconstruction region, all parameters in equation (3) above are known except for the interpolation filter coefficients. Therefore, by solving equation (3) above, the filter coefficients of the interpolation filter of the current block can be determined.

[0542] In one example, the encoding side employs a method of decomposing the autocorrelation matrix using Cholesky decomposition, solving the Wiener-Hoff equation shown in equation (3) above to obtain the filter coefficients of the filter.

[0543] The encoding side determines the filter coefficients of the interpolation filter based on equation (3) described above, and then uses the interpolation filter to perform interpolation filtering prediction on the current block based on these filter coefficients, and obtains the predicted block for the current block.

[0544] As a specific example, the decryption side obtains the predicted block for the current block based on the following equation (4).

[0545] Method 2: The above-mentioned S201-A1 includes the following steps S201-A11 to S201-A14.

[0546] S201-A11, currently, the first reconstruction region around the block is determined.

[0547] S201-A12 determines the pixel-averaged reconstruction value based on the reconstruction value of the first reconstruction region.

[0548] S201-A13 performs mean subtraction on the reconstructed values ​​of the pixel points in the reference region based on the pixel-average reconstructed values.

[0549] S201-A14 uses the pixel values ​​of the pixel points after averaging removal in the reference region as input to the interpolation filter, slides the interpolation filter within the reference region, and obtains the filter coefficients of the interpolation filter.

[0550] In method two, averaging removal is performed on the reference region, and the filter coefficients of the interpolation filter are determined based on the reference region after averaging removal. Since the amount of data is reduced by performing averaging removal on the reference region, the efficiency of determining the filter coefficients can be improved when the filter coefficients are determined based on the reference region after averaging removal.

[0551] Specifically, the encoding side first determines a first reconstruction region, which may be any part of the reconstruction region within the reconstruction region surrounding the current block.

[0552] In embodiments of the present invention, the method by which the encoding side determines the first reconstruction region around the current block includes at least the method shown below.

[0553] Method 1: By default, the encoding side determines one reconfiguration region surrounding the current block as the first reconfiguration region.

[0554] For example, as shown in Figure 20, the encoding side defaults to using the region consisting of the top row, leftmost column, and one pixel point in the upper left corner of the current block as the reconstruction region.

[0555] Method 2: Determine the first reconstruction region based on the current block shape.

[0556] For example, if the current block shape is square, the top row and leftmost column of the current block are determined to be the first reconstruction region.

[0557] Furthermore, for example, if the current block shape is a rectangle where the width is greater than the height, the top row of reconstructed pixels of the current block is determined as the first reconstructed region.

[0558] Furthermore, for example, if the current block shape is a rectangle where the height is greater than the width, the reconstructed pixel region of the leftmost column of the current block is determined as the first reconstructed region.

[0559] The current method for determining the first reconstruction region based on the shape of the block includes, but is not limited to, the examples described above.

[0560] The encoding side determines the first reconstruction region, and then determines the pixel-average reconstruction value m based on the reconstruction value of the first reconstruction region.

[0561] In one implementation method, the average value of the reconstruction values ​​in the first reconstruction region is determined as the pixel-average reconstruction value m.

[0562] For example, if the first reconstruction region is as shown in Figure 20, the pixel-average reconstruction value m may be calculated using the method shown in Table 14.

[0563] In one example, if the first reconstruction region is the top row and / or leftmost column of the current block, the average of the reconstruction values ​​of the top row and / or leftmost column may be determined as the pixel-average reconstruction value m. In this case, the pixel-average reconstruction value m can be calculated using the method shown in Table 15.

[0564] As shown in Table 15 above, if the first reconstruction region is the top row and / or leftmost column of the current block, the pixel average reconstruction value m can be calculated quickly by using a shift operation instead of division.

[0565] In addition to determining the average of the reconstruction values ​​in the first reconstruction region as the pixel-average reconstruction value m, the encoding side may also determine the pixel-average reconstruction value m using the method shown below.

[0566] In another method, the weighted average of the reconstruction values ​​in the first reconstruction region is determined as the pixel-average reconstruction value m.

[0567] The encoding side may also determine the pixel-average reconstruction value m using other methods.

[0568] The encoding side determines the pixel-average reconstruction value, and then performs average removal on the reconstruction values ​​of the pixel points in the reference region based on that pixel-average reconstruction value.

[0569] For example, for each pixel point in the reference region, the pixel value of the pixel point after averaging removal is obtained by dividing the reconstructed value of that pixel point by the above-mentioned average reconstructed value and then rounding it.

[0570] Furthermore, for example, the encoding side obtains the pixel value of a pixel point in the reference region after averaging removal by subtracting the pixel average reconstruction value from the reconstruction value of the pixel point in the reference region. For example, for each pixel point in the reference region, the pixel value of the pixel point in the reference region after averaging removal is obtained by subtracting the aforementioned pixel average reconstruction value from the reconstruction value of that pixel point.

[0571] The embodiments of this application are not limited to a specific method by which the decoding side performs average removal on the reconstructed values ​​of pixel points in the reference region based on the pixel-averaged reconstructed values.

[0572] The encoding side performs average removal on the reconstructed values ​​of the pixel points in the reference region based on the method described above, obtains the pixel values ​​of the pixel points after average removal in the reference region, and then executes the steps of S201-A14 described above. The pixel values ​​of the pixel points after average removal in the reference region are used as input to the interpolation filter, and the interpolation filter is slid within the reference region to obtain the filter coefficients of the interpolation filter.

[0573] As an example, Figure 21 shows the process of obtaining the filter coefficients of the interpolation filter by sliding the interpolation filter of the current block on the reference region of the current block after averaging removal, when the interpolation filter of the current block has five different shapes and the reference region of the current block has three different types. The interpolation filter may be slid horizontally one row at a time or vertically one column at a time on the reference region after averaging removal.

[0574] JPEG2024192733000042.jpg69168

[0575] In one example, the Wiener-Hoff equation constructed by sliding the interpolation filter within the reference region of the current block is as shown in equation (5).

[0576] Since the reference region of the current block is the reconstruction region, all parameters in equation (5) above are known except for the interpolation filter coefficients. Therefore, by solving equation (5) above, the filter coefficients of the interpolation filter of the current block can be determined.

[0577] In one example, the encoding side employs a method of decomposing the autocorrelation matrix using Cholesky decomposition, solving the Wiener-Hoff equation shown in equation (5) above to obtain the filter coefficients of the filter.

[0578] The encoding side determines the filter coefficients of the interpolation filter based on equation (5) described above, then executes the steps S201-A2 described above, and based on the filter coefficients, performs interpolation filtering prediction on the current block using the interpolation filter and obtains the predicted block for the current block.

[0579] In equation (5) described above, the filter coefficients are determined using the reference region after averaging removal. Therefore, when determining the predicted value of the current block based on these filter coefficients, it is necessary to consider the influence of the pixel-averaged reconstruction value m.

[0580] In one possible implementation method, the interpolation filter coefficients determined by equation (5) above are substituted into equation (4) above to obtain the predicted values ​​for each point in the current block. Then, the pixel-average reconstruction value m is added to the predicted values ​​for each point to obtain the final predicted values ​​for each point in the current block, and consequently, the predicted block for the current block is obtained.

[0581] In another possible implementation, the above-described S201-A2 includes the following steps:

[0582] S201-A21, for the r-th point in the current block, the pixel values ​​of the N positions corresponding to the r-th point are determined based on the shape of the interpolation filter, where r is a positive integer.

[0583] S201-A22 performs average removal on the pixel values ​​at N positions based on the pixel average reconstruction value, and obtains the pixel values ​​after average removal at the N positions.

[0584] S201-A23 obtains the predicted value of the r-th point based on the pixel values ​​and filter coefficients after averaging and removing N positions.

[0585] S201-A24: Based on the predicted values ​​of each point within the current block, obtain the predicted block for the current block.

[0586] As shown in Figure 22, assuming that the shape of the interpolation filter in the current block is 4x4, the encoding side sequentially performs interpolation predictions for each position in the current block using an interpolation filter with known filter coefficients. Specifically, for the r-th point in the current block, the pixel values ​​of the N positions corresponding to the r-th point are first determined based on the shape of the interpolation filter in the current block. For example, as shown in Figure 22, in a 4x4 interpolation filter, the dark positions are the positions of the r-th point to be processed, and the 15 light positions are the N positions corresponding to the r-th point.

[0587] Next, the pixel values ​​of the N positions corresponding to the r-th point are determined. For example, for any of the N positions, if that position is in the reconstruction area surrounding the current block, the reconstruction value of that position is determined as the pixel value of that position. If that position is within the current block, the predicted value of that position is determined as the pixel value of that position.

[0588] Since the filter coefficients described above are determined based on the reference region after averaging and removal, the encoding side performs averaging and removal on the pixel values ​​at N positions of the r-th point based on the pixel average reconstruction value, and obtains the pixel values ​​at N positions of the r-th point after averaging and removal. For example, the pixel values ​​at N positions of the r-th point after averaging and removal are obtained by subtracting the pixel average reconstruction value from the pixel values ​​at N positions of the r-th point.

[0589] Next, the predicted value of the r-th point is obtained based on the pixel values ​​and filter coefficients after averaging and removing the values ​​of N positions.

[0590] The embodiments of this application are not limited to a specific method for obtaining a predicted value of the r-th point based on the averaged and removed pixel values ​​and filter coefficients of N positions.

[0591] JPEG2024192733000043.jpg49168

[0592] In an alternative implementation, the above-mentioned S201-A23 includes the following steps.

[0593] S201-A231 determines the second reconstruction region around the current block, and determines the maximum and minimum reconstruction values ​​for the said second reconstruction region.

[0594] S201-A232 obtains a first predicted value based on the pixel values ​​after averaging and removing N positions, the filter coefficients, and the pixel-averaged reconstruction value.

[0595] Based on S201-A233, the first predicted value, the maximum reconstruction value, and the minimum reconstruction value, the predicted value of the r-th point is determined.

[0596] JPEG2024192733000044.jpg29168

[0597] The embodiments of this application do not limit the specific method for determining the second reconstruction region around the current block.

[0598] In one example, the second reconfiguration region of the current block coincides with the reference region of the current block.

[0599] In one example, the second reconfiguration region of the current block coincides with the first reconfiguration region of the current block.

[0600] In one example, the reconstruction areas above, to the left, to the right, to the left, and to the left of the current block are determined as the second reconstruction area. For example, the reconstruction areas of the top 13 rows, left 13 columns, top 13 rows, top 13 rows and 13 columns of the top left, and bottom 13 columns of the current block are determined as the second reconstruction area.

[0601] Furthermore, the execution order is not limited in the specific implementation process of S201-A231 and S201-A232 described above. For example, S201-A231 may be executed before S201-A232, after S201-A232, or simultaneously with S201-A232.

[0602] The embodiments of this application are not limited to a specific method by which the encoding side obtains a first predicted value based on the pixel values ​​after averaging and removal of N positions, filter coefficients, and pixel-averaged reconstruction value.

[0603] For example, the second predicted value for the r-th point is obtained by multiplying the pixel values ​​after averaging and removing the N positions of the r-th point by the filter coefficient, and the first predicted value for the r-th point is obtained by adding the second predicted value and the pixel-averaged reconstruction value.

[0604] For example, the encoding side obtains the first predicted value of the r-th point based on the following equation (6).

[0605] Furthermore, for example, the encoding side obtains one predicted value for the r-th point based on equation (6) described above, then performs a prediction process on that predicted value to obtain the first predicted value for the r-th point.

[0606] Based on the steps described above, the encoding side determines the first predicted value for the r-th point in the current block, and then determines the predicted value for the r-th point based on the first predicted value, the maximum reconstructed value, and the minimum reconstructed value.

[0607] For example, if the first predicted value is greater than the minimum reconstruction value and less than the maximum reconstruction value, the first predicted value is determined to be the predicted value for the r-th point.

[0608] Furthermore, for example, if the first predicted value is less than or equal to the minimum reconstruction value, the minimum reconstruction value is determined as the predicted value for the r-th point.

[0609] Furthermore, for example, if the first predicted value is greater than or equal to the maximum reconstruction value, the maximum reconstruction value is determined as the predicted value for the r-th point.

[0610] In one example, the encoding side determines the predicted value of the r-th point based on the following equation (7).

[0611] The above example illustrates the case of determining the predicted value of the r-th point in the current block. The encoding side refers to the above method to determine the predicted value of each point in the current block, and furthermore, the predicted values ​​of each point in the current block constitute the predicted block of the current block.

[0612] Based on the steps above, the encoding side performs interpolation filtering prediction on the current block, obtains the predicted block for the current block, and then performs the following steps.

[0613] S202 determines the intra-prediction mode corresponding to the prediction block, and based on the intra-prediction mode corresponding to the prediction block, determines the transformation kernel corresponding to the current block.

[0614] As is clear from the above explanation, when encoding the current block, the encoding side determines the predicted block of the current block based on the above steps. Next, the residual block of the current block is obtained by subtracting the predicted block of the current block from the current block. Furthermore, a transformation is performed on the residual block of the current block to obtain the transformation coefficients, quantization is performed on the transformation coefficients to obtain the quantization coefficients, encoding is performed on the quantization coefficients to obtain the bitstream.

[0615] When performing a transformation on the residual value of the current block and obtaining the transformation coefficients, it is necessary to determine the transformation kernel, and the transformation coefficients are obtained by performing a transformation on the residual value of the current block based on the transformation kernel. Currently, the encoding side employs a conventional intra-prediction mode to make predictions on the current block. The encoding side can determine the transformation kernel to be used for the current block based on the correspondence between the conventional intra-prediction mode and the transformation kernel. However, in the embodiment of the present invention, when making predictions on the current block, an interpolation filter prediction mode is used instead of the conventional intra-prediction mode. Therefore, it is not possible to directly determine the transformation kernel corresponding to the current block.

[0616] To solve this technical problem, in the embodiment of the present invention, after determining the predicted block of the current block using an interpolation filtering prediction mode, a conventional intra prediction mode corresponding to the predicted block is determined, and then a transformation kernel corresponding to the current block is determined based on the conventional intra prediction mode.

[0617] The following describes the specific process by which the encoding side determines the intra-prediction mode corresponding to the prediction block.

[0618] As an example, as shown in Figure 7, the conventional intra-prediction modes included in the current VVC are as follows: PLANAR mode: Intra prediction mode index is 0, DC mode: Intra prediction mode index is 1, Angle mode: Intra prediction mode index is 2-66.

[0619] In one example, as shown in Figure 23, the direction of the arrows in the figure indicates the direction of the angular mode prediction present in the VVC, where the prediction mode index used during coding is 2 to 66. If the block is currently a non-square block, some angular directions are replaced with wider angles, for example, -1 to -14 and 67 to 80 in Figure 23.

[0620] In some embodiments, the intra-prediction mode corresponding to the prediction block is the default intra-prediction mode. That is, if the current block employs the interpolation filtering prediction mode to perform prediction and obtains a prediction block, one of the conventional intra-prediction modes is determined by default as the intra-prediction mode corresponding to that prediction block.

[0621] In some embodiments, the encoding side determines the intra-prediction mode corresponding to the prediction block by the following steps:

[0622] S202-A1, determine the angle values ​​of M points in the prediction block, where M is a positive integer.

[0623] S202-A2 determines the intra-prediction mode corresponding to the prediction block based on the angle values ​​of M points.

[0624] In the embodiment of this invention, the intra-prediction mode corresponding to a prediction block is determined by statistically analyzing the intra-prediction modes corresponding to the angle values ​​of M points in the prediction block.

[0625] The embodiments of this application do not limit the specific location and number of M points for determining the angle value in the prediction block. For example, the M points may be one point in the prediction block or multiple points in the prediction block.

[0626] For example, if the above M points are one point, the encoding side determines the angle value of one point within the prediction block (for example, the center point of the prediction block), determines the intra-prediction mode corresponding to that point based on the angle value of that point, and further determines that intra-prediction mode as the intra-prediction mode corresponding to the prediction block.

[0627] Furthermore, for example, if the M points mentioned above are multiple points, the encoding side determines the angle values ​​of these multiple points, determines the intra-prediction mode corresponding to each of these multiple points based on the angle values ​​of these multiple points, and then determines the intra-prediction mode that has the largest number of identical intra-prediction modes among these multiple points as the intra-prediction mode corresponding to the prediction block.

[0628] In some embodiments, when the angular values ​​of M points in a prediction block are determined by a sliding window method, the selection of these M points is related to the shape and size of the sliding window. For example, each of the M points is the center point of the sliding window as it slides within the prediction block.

[0629] In the embodiments of this application, the method for determining the angle value of each of the M points is the same. For simplicity of explanation, the case of determining the angle value of the i-th point among the M points will be explained as an example.

[0630] The embodiments of this application do not limit the specific method for determining the angle value of a point.

[0631] In some embodiments, step S202-A1 includes steps S202-A11 and S202-A12.

[0632] S202-A11, for the i-th point out of M points, determine the horizontal and vertical slopes of the i-th point, where i is a positive integer less than or equal to M.

[0633] S202-A12, the angle value of the i-th point is determined based on the horizontal and vertical slopes of the i-th point.

[0634] In this embodiment, the encoding side first determines the horizontal and vertical slopes of each of the M points (for example, the i-th point), and then determines the angle value of the i-th point based on the horizontal and vertical slopes.

[0635] The embodiments of this application do not limit the specific method for determining the horizontal and vertical slopes of the i-th point.

[0636] In one example, the horizontal gradient value of the i-th point is determined based on the horizontal change between the predicted values ​​of the points surrounding the i-th point within the prediction block and the predicted value of the i-th point itself, and the vertical gradient value of the i-th point is determined based on the vertical change between the predicted values ​​of the points surrounding the i-th point within the prediction block and the predicted value of the i-th point itself.

[0637] In another example, the encoding side determines the predicted values ​​of points within a sliding window centered on the i-th point within the prediction block, and obtains the horizontal and vertical gradients of the i-th point based on the predicted values ​​of points within the sliding window, the horizontal gradient operator, and the vertical gradient operator.

[0638] In this example, first, the sliding window is determined, for example, a 3x3 sliding window as shown in Figure 24. The sliding window is slid within the prediction block, and each time it slides, the horizontal and vertical slopes of the center point of the sliding window are determined. To illustrate with an example where the current center point of the sliding window is the i-th point, first, the predicted values ​​for each point in the current sliding window are obtained, for example, 3x3 = 9 predicted values ​​can be obtained. Next, the horizontal and vertical slopes of the i-th point are determined based on the predicted values ​​of these 9 points and the pre-set horizontal and vertical slope operators.

[0639] JPEG2024192733000045.jpg27168

[0640] JPEG2024192733000046.jpg38168

[0641] The embodiments of this application do not limit the specific values ​​of the horizontal gradient operator and the vertical gradient operator.

[0642] After determining the horizontal and vertical slopes of the i-th point based on the above steps, the encoding side may then determine the angle value of the i-th point based on the horizontal and vertical slopes of the i-th point.

[0643] For example, the arctangent value of the ratio of the vertical slope to the horizontal slope at the i-th point is determined as the angle value of the i-th point. For example, the angle value of the i-th point is determined based on equation (8).

[0644] The encoding side may determine the angle value of the i-th point using a method other than equation (8) above. For example, the encoding side may obtain the angle value of the i-th point by adjusting the angle value determined by equation (8) above.

[0645] The encoding side employs the above method for each of the M points, determines the angle value for each of the M points, and then executes S202-A2 to determine the intra-prediction mode corresponding to the prediction block based on the angle values ​​of the M points.

[0646] The embodiments of this application are not limited to a specific method for determining an intra-prediction mode corresponding to a prediction block based on the angular values ​​of M points.

[0647] In some embodiments, the encoding side selects the angle value 1 with the highest frequency of occurrence from among the angle values ​​of M points, matches this angle value 1 with the predicted angle of the conventional intra prediction mode to obtain the intra prediction mode corresponding to this angle value 1, and further determines the intra prediction mode corresponding to this angle value 1 as the intra prediction mode corresponding to the prediction block.

[0648] In some embodiments, the above S202-A2 includes the following steps S202-A21 and S202-A22.

[0649] S202-A21 determines the intra-prediction mode corresponding to M points based on the angle values ​​of M points.

[0650] S202-A22 determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction modes corresponding to M points.

[0651] In this implementation method, the encoding side determines the intra-prediction mode corresponding to each of the M points based on the angle value of each point. For example, for each of the M points, the intra-prediction mode corresponding to the angle value of that point is obtained by matching the angle value of that point with the predicted angle of a conventional intra-prediction mode. This makes it possible to obtain the intra-prediction mode corresponding to each of the M points.

[0652] Next, based on the intra-prediction modes corresponding to each of these M points, the intra-prediction mode corresponding to the prediction block is determined.

[0653] In one possible implementation, among the intra-prediction modes that each of the M points corresponds to, the intra-prediction mode that is repeated most frequently is determined as the intra-prediction mode corresponding to the prediction block.

[0654] In another possible implementation, S202-A22 includes the following steps:

[0655] S202-A221 determines the gradient amplitude value corresponding to M points based on the horizontal and vertical gradients of M points.

[0656] S202-A222 determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction mode and gradient amplitude values ​​corresponding to M points.

[0657] In this implementation method, the encoding side determines the gradient amplitude value corresponding to each of the M points based on the horizontal and vertical gradients of each of the M points determined above.

[0658] In the embodiments of this application, the specific method by which the encoding side determines the gradient amplitude value corresponding to each of the M points is the same. For simplicity of explanation, the case in which the gradient amplitude value corresponding to the i-th point among the M points is determined will be explained as an example.

[0659] The embodiments of this application do not limit the specific method by which the encoding side determines the gradient amplitude value corresponding to the i-th point based on the horizontal and vertical gradients of the i-th point.

[0660] For example, the encoding side multiplies the horizontal gradient and vertical gradient of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.

[0661] Furthermore, for example, the encoding side adds the absolute values ​​of the horizontal and vertical gradients of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.

[0662] For example, the encoding side determines the gradient amplitude value corresponding to the i-th point based on equation (9) below.

[0663] Based on the above steps, the encoding side can determine the gradient amplitude value corresponding to each of the M points. Next, the encoding side performs S202-A222 and determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction mode and gradient amplitude value corresponding to the M points.

[0664] In one example, the intra-prediction mode corresponding to the point with the maximum gradient amplitude among the M points is determined as the intra-prediction mode corresponding to the prediction block.

[0665] In another example, for any of the M points, the gradient amplitude value corresponding to that point is accumulated in the intra-prediction mode corresponding to that point, and the cumulative gradient amplitude value of the intra-prediction modes corresponding to the M points is obtained. Among the intra-prediction modes corresponding to the M points, the intra-prediction mode with the largest cumulative gradient amplitude value is determined as the intra-prediction mode corresponding to the prediction block.

[0666] As an example, as shown in Figure 25, the gradient amplitude values ​​corresponding to each of the M points are accumulated in the corresponding intra-prediction mode. For example, if the intra-prediction modes corresponding to points 1 and 2 among the M points are both intra-prediction mode 1, the gradient amplitude values ​​corresponding to points 1 and 2 are accumulated and added to the gradient amplitude value corresponding to intra-prediction mode 1. By this analogy, the gradient amplitude value histogram shown in Figure 25 can be obtained. This allows the intra-prediction mode with the largest accumulated gradient amplitude value in the gradient amplitude value histogram to be determined as the intra-prediction mode corresponding to the prediction block. For example, the intra-prediction mode corresponding to the accumulated gradient amplitude value shown in dark color in Figure 25 is determined as the intra-prediction mode corresponding to the prediction block.

[0667] In some embodiments, if the gradient amplitude values ​​corresponding to M points are all 0, the first intra-prediction mode is determined as the intra-prediction mode corresponding to the prediction block. That is, if the gradient amplitude values ​​corresponding to all M points are all 0, it means that the horizontal and vertical gradients of each of the M points are both 0. In this case, a pre-set first intra-prediction mode may be determined as the intra-prediction mode corresponding to the prediction block.

[0668] The embodiments of this application are not limited to the type of the first intra-prediction mode described above.

[0669] The embodiments of this application are not limited to the first intra-prediction mode described above.

[0670] The encoding side determines the intra-prediction mode corresponding to the prediction block based on the above steps, and then determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to that prediction block.

[0671] The embodiments of this application do not limit the specific method by which the encoding side determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block.

[0672] In some embodiments, the encoding side searches for an image block among already decoded image blocks surrounding the prediction block whose intra-prediction mode is the same as the intra-prediction mode corresponding to the prediction block, based on the intra-prediction mode corresponding to the prediction block, and then determines the transformation kernel corresponding to that image block as the transformation kernel corresponding to the current block.

[0673] In some embodiments, the step of determining the translation kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block in S202 above includes the following steps.

[0674] S202-B1 obtains the correspondence between the intra prediction mode and the transformation kernel group, where a transformation kernel group includes at least one type of transformation kernel.

[0675] S202-B2 looks up the first transformation kernel group corresponding to the intra-prediction mode of the prediction block within the above correspondence.

[0676] S202-B3 determines the translation kernel corresponding to the current block from the first translation kernel group.

[0677] In the embodiment of the present invention, a correspondence exists between the intra-prediction mode and the transformation kernel group. Based on this, the encoding side determines the intra-prediction mode corresponding to the prediction block, and then obtains the pre-configured correspondence between the intra-prediction mode and the transformation kernel group.

[0678] For example, the correspondence between the intra-prediction mode and the conversion kernel group is shown in Table 16.

[0679] Table 16 above merely shows the correspondence between the intra-prediction mode and the conversion kernel group in the embodiment of the present application, and the correspondence between the intra-prediction mode and the conversion kernel group in the embodiment of the present application is not limited to those shown in Table 16.

[0680] Here, each translation kernel group includes at least one type of translation kernel.

[0681] The encoding side obtains the correspondence between intra-prediction modes and transformation kernel groups shown in Table 16. Based on the intra-prediction mode corresponding to the prediction block, it looks up the transformation kernel group corresponding to the intra-prediction mode from the correspondence between intra-prediction modes and transformation kernel groups, and denotes this transformation kernel group as the first transformation kernel group. For example, if the intra-prediction mode corresponding to the prediction block is an angle prediction mode in 64 angular directions, looking up Table 16 above will reveal that the transformation kernel group corresponding to that 64 angular direction angle prediction mode is 4. As a result, the encoding side determines the transformation kernel corresponding to the current block from at least one type of transformation kernel included in transformation kernel group 4.

[0682] For example, if the first translation kernel group contains one translation kernel, that translation kernel is determined to be the translation kernel corresponding to the current block.

[0683] Furthermore, for example, if the first transformation kernel group contains multiple types of transformation kernels, the encoding side determines the type of transformation kernel corresponding to the current block, and then determines the transformation kernel of that type in the first transformation kernel group as the transformation kernel corresponding to the current block.

[0684] Here, the method by which the encoding side determines the type of transformation kernel corresponding to the current block is not limited to those shown below.

[0685] In one example, the type of transformation kernel corresponding to the current block is the default type. In this case, the encoding side determines that the default type is the type of transformation kernel corresponding to the current block.

[0686] In another example, the encoder writes the type of transformation kernel corresponding to the current block into the bitstream. The encoder then obtains the type of transformation kernel corresponding to the current block by encoding the bitstream.

[0687] As is clear from the above description, in the embodiment of the present invention, the encoding side uses an interpolation filtering prediction mode to determine the predicted block of the current block, then determines the conventional intra-prediction mode corresponding to that predicted block, and then determines the transform kernel corresponding to the current block based on the conventional intra-prediction mode corresponding to that predicted block. That is, in the embodiment of the present invention, the conventional intra-prediction mode derived from the interpolation filtering prediction is used to select the transform kernel group of Non-separable primary transform (NSPT) and Low Frequency non-separable secondary transform (LFNST). This allows the determined transform kernel to better suit the characteristics of the current block, improving the accuracy of the transform kernel determination. Determining the reconstruction value of the current block using a transform kernel with this accuracy improves the accuracy of the reconstruction value and enhances the encoding accuracy of the current block. Furthermore, in the embodiment of the present invention, when determining the transform kernel of the current block via the conventional prediction mode corresponding to the predicted block, it is not necessary to specify the transform kernel individually, saving codewords and improving the video encoding effect.

[0688] After determining the conversion kernel corresponding to the current block based on the above steps, the encoding side executes the following steps in S203.

[0689] S203 performs a transformation on the residual block of the current block based on the transformation kernel corresponding to the current block, obtains the transformation coefficients of the current block, and performs encoding based on the transformation coefficients of the current block to obtain the bitstream.

[0690] In the embodiment of the present invention, the encoding side determines the predicted block of the current block and the transformation kernel corresponding to the current block based on the above steps. This allows the encoding side to obtain the residual block of the current block based on the predicted block of the current block and the current block. For example, the residual block of the current block is obtained by subtracting the predicted block of the current block from the current block. Next, a transformation is performed on the residual block of the current block based on the transformation kernel determined above to obtain the transformation coefficients of the current block. Furthermore, the transformation coefficients are directly encoded to obtain a bitstream, or quantization is performed on the transformation coefficients to obtain quantization coefficients, and the quantization coefficients are encoded to obtain a bitstream.

[0691] In some embodiments, the current block is either a luminance block or a chroma block. That is, in embodiments of the present application, prediction can be performed for either a luminance block or a chroma block using the interpolation filtering prediction mode provided by embodiments of the present application.

[0692] In some embodiments, when the current block is a luminance block, the prediction mode of the current block is an interpolation filtering prediction mode, and the chroma block corresponding to the current block employs a direct derivation mode DM, the PLANAR mode or the intra-prediction mode corresponding to the prediction block is determined as the prediction mode of the chroma block.

[0693] In the following, we will test the effectiveness of the interpolation filter prediction mode proposed in the embodiment of this application through experiments.

[0694] As an example, Table 17 shows the compression effect when compressing different videos under All Intra Main test conditions, using 3 × 5 = 15 combinations of filter coefficients determined by the type of reference region shown in Figures 13A to 13C and the 5 types of interpolation filters shown in Figure 15.

[0695] JPEG2024192733000047.jpg82160

[0696] As shown in Table 17, when compression was performed on different types of test data under general-purpose test conditions using the reference region types shown in Figures 13A to 13C and the five types of interpolation filters shown in Figure 15, an objective improvement in compression effect of 0.26%, 0.20%, and 0.17% was achieved for the Y, U, and V components, respectively.

[0697] As an example, Table 18 shows the compression effect when compressing different videos under All Intra Main test conditions, using 3 × 7 = 21 combinations of filter coefficients determined by the type of reference region shown in Figures 13A to 13C and the seven types of interpolation filters shown in Figures 14A to 14G.

[0698] JPEG2024192733000048.jpg67123

[0699] As shown in Table 18, when compression was performed on different types of test data under general-purpose test conditions using the reference region types shown in Figures 13A to 13C and the seven types of interpolation filters shown in Figures 14A to 14G, an objective improvement in compression effect of 0.26%, 0.19%, and 0.19% was achieved for the Y, U, and V components, respectively.

[0700] As an example, using the types of reference regions shown in Figures 13A to 13C and the interpolation filter of one shape shown in Figure 14A, a combination of 3 × 1 = 3 filter coefficients is determined. When compression is performed on different videos under All Intra Main test conditions, the compression effect is shown in Table 19.

[0701] JPEG2024192733000049.jpg68125

[0702] As shown in Table 19, when compression was performed on different types of test data under general-purpose test conditions using the reference region types shown in Figures 13A to 13C and the interpolation filter of one shape shown in Figure 14A, an objective improvement in compression effect of 0.13%, 0.11%, and 0.02% was obtained for the Y, U, and V components, respectively.

[0703] As an example, using the types of reference regions shown in Figures 13A to 13C and the three types of interpolation filters shown in Figure 17, a combination of 3 × 3 = 9 filter coefficients is determined. When compression is performed on different videos under All Intra Main test conditions, the compression effect is shown in Table 20.

[0704] JPEG2024192733000050.jpg68125

[0705] As shown in Table 20, when compression was performed on different types of test data under general-purpose test conditions using the reference region types shown in Figures 13A to 13C and the three types of interpolation filters shown in Figure 17, an objective improvement in compression effect of 0.23%, 0.20%, and 0.15% was achieved for the Y, U, and V components, respectively.

[0706] As an example, Table 21 shows the compression effect when compressing different videos under All Intra Main test conditions, using a combination of 3 × 3 = 9 types of filter coefficients determined by the type of reference region shown in Figures 13A to 13C and the three types of interpolation filters shown in Figure 18A or Figure 18B.

[0707] JPEG2024192733000051.jpg67125

[0708] As shown in Table 21, when compression was performed on different types of test data under general-purpose test conditions using the reference region types shown in Figures 13A to 13C and the three types of interpolation filters shown in Figure 18A or Figure 18B, an objective improvement in compression effect of 0.25%, 0.18%, and 0.18% was achieved for the Y, U, and V components, respectively.

[0709] As an example, when using a combination of 3 × 3 = 9 types of filter coefficients obtained by combining the types of reference regions shown in Figures 13A to 13C and the three types of interpolation filters shown in Figure 19, and compressing different videos under All Intra Main test conditions, the compression effect is shown in Table 22.

[0710] JPEG2024192733000052.jpg67125

[0711] As shown in Table 22, when compression was performed on different types of test data under general-purpose test conditions using the reference region types shown in Figures 13A to 13C and the three types of interpolation filters shown in Figure 19, an objective improvement in compression effect of 0.19%, 0.15%, and 0.14% was achieved for the Y, U, and V components, respectively.

[0712] In the video encoding method provided by the embodiment of the present application, when making a prediction for the current block, first the reference region and interpolation filter of the current block are determined, and the prediction block of the current block is determined based on the reference region and interpolation filter. Next, the prediction mode corresponding to the prediction block is determined, and further, a transformation is performed on the residual block of the current block based on the transformation kernel corresponding to the current block to obtain the transformation coefficients of the current block, and encoding is performed based on the transformation coefficients of the current block to obtain a bitstream. That is, in the embodiment of the present application, when an interpolation filter prediction method is used to make a prediction for the current block, the transformation kernel corresponding to the current block is determined by determining the conventional prediction mode corresponding to the prediction block, so that the determined transformation kernel is better suited to the characteristics of the current block, and the accuracy of the transformation kernel determination is improved. When the reconstruction value of the current block is determined using the transformation kernel determined with this accuracy, the accuracy of the reconstruction value is improved, and the encoding accuracy of the current block can be increased. Furthermore, in the embodiment of the present application, since the transformation kernel of the current block is determined through the conventional prediction mode corresponding to the prediction block, there is no need to individually specify the transformation kernel, saving the amount of code and further improving the video encoding effect.

[0713] It is understood that Figures 10 to 26 are merely illustrative examples of the present application and should not be interpreted as limiting the present application.

[0714] Although preferred embodiments of the present application have been described in detail above with reference to the drawings, the present application is not limited to the specific details of the embodiments described above. Within the scope of the technical idea of ​​the present application, several simple modifications can be made to the technical solution of the present application, and all of these simple modifications fall within the scope of protection of the present application. For example, each specific technical feature described in the specific embodiments described above can be combined in any appropriate manner, as long as no contradiction arises. To avoid unnecessary redundancy, the various possible combinations are not described separately in the present application. Furthermore, different embodiments of the present application can be combined in any way, and as long as this does not violate the idea of ​​the present application, it will be considered to be the same as the content disclosed in the present application.

[0715] It should also be understood that, in the embodiments of the methods of this application, the magnitude of the process numbers described above does not indicate the order of execution, and the execution order of each process should be determined by its function and inherent logic, and does not limit the execution processes of the embodiments of this application in any sense. Furthermore, the term "and / or" in the embodiments of this application simply describes the relationship between related objects, indicating that three types of relationships may exist. Specifically, A and / or B indicates three cases: when A exists alone, when A and B exist simultaneously, and when B exists alone. In addition, the character " / " in this application generally indicates that the preceding and succeeding related objects are in an "or" relationship.

[0716] As described above with reference to Figures 10 to 27, embodiments of the present invention's method have been explained in detail. Below, embodiments of the present invention's apparatus will be described in detail with reference to Figures 28 to 30.

[0717] Figure 28 is a schematic block diagram of a video decoding device according to one embodiment of the present invention, and the video decoding device 10 is applied to the video decoder described above.

[0718] As shown in Figure 28, the video decoding device 10 includes the following configuration.

[0719] The prediction unit 11 determines the reference region and interpolation filter of the current block, and is used to determine the predicted block of the current block based on the reference region and the interpolation filter.

[0720] The conversion unit 12 is used to determine an intra-prediction mode corresponding to the prediction block and to determine a conversion kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block.

[0721] The decoding unit 13 is used to perform an inverse transformation on the transformation coefficients of the current block based on the transformation kernel corresponding to the current block, to obtain the residual block of the current block, and to obtain the reconstructed block of the current block based on the predicted block and residual block of the current block.

[0722] In some embodiments, the conversion unit 12 is specifically used to determine the angular values ​​of M points in the prediction block, where M is a positive integer, and is used to determine the intra-prediction mode corresponding to the prediction block based on the angular values ​​of the M points.

[0723] In some embodiments, the conversion unit 12 is specifically used to determine the horizontal and vertical slopes of the i-th point among the M points, where i is a positive integer less than or equal to M, and is used to determine the angle value of the i-th point based on its horizontal and vertical slopes.

[0724] In some embodiments, the transformation unit 12 is specifically used to determine the predicted value of a point in a sliding window centered on the i-th point within the prediction block, and to obtain the horizontal and vertical gradients of the i-th point based on the predicted value of the point in the sliding window and the horizontal and vertical gradient operators.

[0725] In some embodiments, the conversion unit 12 is specifically used to determine the horizontal gradient of the i-th point by multiplying the predicted value of a point in the sliding window by the horizontal gradient operator, and to determine the vertical gradient of the i-th point by multiplying the predicted value of a point in the sliding window by the vertical gradient operator.

[0726] In some embodiments, the conversion unit 12 is specifically used to determine the arctangent value of the ratio of the vertical slope to the horizontal slope of the i-th point as the angle value corresponding to the i-th point.

[0727] In some embodiments, the conversion unit 12 is specifically used to determine an intra-prediction mode corresponding to the M points based on the angular values ​​of the M points, and to determine an intra-prediction mode corresponding to the prediction block based on the intra-prediction modes corresponding to the M points.

[0728] In some embodiments, the conversion unit 12 is specifically used to determine gradient amplitude values ​​corresponding to the M points based on the horizontal and vertical gradients of the M points, and to determine the intra-prediction mode corresponding to the prediction block based on the intra-prediction mode and gradient amplitude values ​​corresponding to the M points.

[0729] In some embodiments, the conversion unit 12 is specifically used to add the absolute value of the horizontal slope and the absolute value of the vertical slope of the i-th point to obtain a slope amplitude value corresponding to the i-th point.

[0730] In some embodiments, the conversion unit 12 specifically accumulates the gradient amplitude value corresponding to any of the M points in the intra-prediction mode corresponding to the point, obtains the accumulated gradient amplitude value of the intra-prediction modes corresponding to the M points, and uses the intra-prediction mode with the largest accumulated gradient amplitude value among the intra-prediction modes corresponding to the M points to determine as the intra-prediction mode corresponding to the prediction block.

[0731] In some embodiments, the conversion unit 12 is further used to determine the first intra-prediction mode as the intra-prediction mode corresponding to the prediction block when the gradient amplitude values ​​corresponding to the M points are all 0.

[0732] Selectively, the first intra-prediction mode is the PLANAR mode.

[0733] In some embodiments, the conversion unit 12 is specifically used to obtain a correspondence between an intra-prediction mode and a conversion kernel group, where one conversion kernel group includes at least one type of conversion kernel, and is used to look up a first conversion kernel group corresponding to the intra-prediction mode of the prediction block within the correspondence, and to determine the conversion kernel corresponding to the current block from among the first conversion kernel group.

[0734] In some embodiments, the conversion unit 12 is specifically used to determine the type of conversion kernel corresponding to the current block, and to determine the conversion kernel of that type in the first conversion kernel group as the conversion kernel corresponding to the current block.

[0735] In some embodiments, the conversion unit 12 is specifically used to decode the bitstream and obtain the type of conversion kernel corresponding to the current block.

[0736] In some embodiments, the prediction unit 11 is specifically used to decode a bitstream to obtain first information, which is used to indicate the type of reference region of the current block, which is used to determine the reference region of the current block among a set of P reference regions based on the type of reference region, where P is a positive integer greater than 1.

[0737] In some embodiments, the prediction unit 11 is specifically used to determine the reference region of the current block within a set of P reference regions based on the shape of the current block, where P is a positive integer greater than 1.

[0738] In some embodiments, the P reference regions include at least one of the first, second, and third reference regions, where the first reference region includes the reconstruction regions above, above, to the right, to the left, to the left, and to the left of the current block; the second reference region includes the reconstruction regions above, above, and to the right of the current block; and the third reference region includes the reconstruction regions to the left, to the left, and to the left of the current block.

[0739] In some embodiments, the prediction unit 11 is specifically used to decode the bitstream to obtain second information, which is used to indicate the shape of the interpolation filter for the current block, which is used to determine the interpolation filter for the current block from among a preset Q interpolation filters, where Q is a positive integer greater than 1, based on the shape of the interpolation filter.

[0740] In some embodiments, the prediction unit 11 is specifically used to determine the interpolation filter for the current block from among a set of Q interpolation filters based on the shape of the current block, where Q is a positive integer greater than 1.

[0741] In some embodiments, the Q interpolation filters include at least one of a first interpolation filter, a second interpolation filter, and a third interpolation filter, wherein the first interpolation filter is a square interpolation filter, the second interpolation filter is a rectangular interpolation filter whose width is greater than its height, and the third interpolation filter is a rectangular interpolation filter whose height is greater than its width.

[0742] In some embodiments, the prediction unit 11 further employs a truncated binary code decoding scheme and is used to decode and obtain the first information and / or the second information from the bitstream.

[0743] In some embodiments, the prediction unit 11 is used to decode the bitstream and obtain third information before the step of determining the reference region and interpolation filter of the current block, the third information being used to indicate whether the current block will perform predictions using the interpolation filter prediction mode, and if it is determined based on the third information that the current block will perform predictions using the interpolation filter prediction mode, the unit is used to determine the reference region and interpolation filter of the current block.

[0744] In some embodiments, before decoding the bitstream to obtain the third information, the prediction unit 11 further determines whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement. If it is determined that the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the bitstream is used to decode the bitstream and obtain the third information.

[0745] In some embodiments, the prediction unit 11 is further used to determine that the current block will not employ the interpolation filter prediction mode if the position of the current block in the current image does not satisfy the predetermined position requirement and / or the size of the current block does not satisfy the predetermined block size requirement.

[0746] In some embodiments, the prediction unit 11 is used to decode the bitstream and obtain fourth information before determining whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement, the fourth information being used to indicate whether the current sequence is permitted to perform predictions using an interpolation filter prediction mode, and if the fourth information indicates that the current sequence is permitted to perform predictions using the interpolation filter prediction mode, it is used to determine whether the position of the current block in the current image satisfies the predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement.

[0747] In some embodiments, the prediction unit 11 is specifically used to determine the filter coefficients of the interpolation filter based on the reference region, to perform an interpolation filter prediction on the current block using the interpolation filter based on the filter coefficients, and to obtain a predicted block of the current block.

[0748] In some embodiments, the prediction unit 11 specifically determines a first reconstruction region around the current block, determines a pixel-average reconstruction value based on the reconstruction value of the first reconstruction region, performs mean removal on the reconstruction values ​​of the pixel points in the reference region based on the pixel-average reconstruction value, uses the pixel values ​​of the pixel points after mean removal in the reference region as input to the interpolation filter, and slides the interpolation filter within the reference region to obtain the filter coefficients of the interpolation filter.

[0749] In some embodiments, the prediction unit 11 is specifically used to determine the first reconstruction region based on the shape of the current block.

[0750] In some embodiments, the prediction unit 11 is used to determine the reconstructed pixel area of ​​the top row and left column of the current block as the first reconstruction area if the shape of the current block is a square, or to determine the reconstructed pixel area of ​​the top row of the current block as the first reconstruction area if the shape of the current block is a rectangle with a width greater than its height, or to determine the reconstructed pixel area of ​​the left column of the current block as the first reconstruction area if the shape of the current block is a rectangle with a height greater than its width.

[0751] In some embodiments, the prediction unit 11 is specifically used to obtain the pixel values ​​of the pixel points in the reference region after mean removal by subtracting the pixel average reconstruction value from the reconstruction value of the pixel points in the reference region.

[0752] In some embodiments, the prediction unit 11 is specifically used to determine the pixel values ​​of N positions corresponding to the r-th point in the current block based on the shape of the interpolation filter, where r is a positive integer; to perform mean removal on the pixel values ​​of the N positions based on the pixel mean reconstruction value to obtain the pixel values ​​after mean removal of the N positions; to obtain a predicted value for the r-th point based on the pixel values ​​after mean removal of the N positions and the filter coefficients; and to obtain a predicted block of the current block based on the predicted value of the point in the current block.

[0753] In some embodiments, the prediction unit 11 is specifically used to determine the reconstructed value of any of the N positions as the pixel value of the position if the position is in the reconstructed region around the current block, or to determine the predicted value of the position as the pixel value of the position if the position is within the current block.

[0754] In some embodiments, the prediction unit 11 is specifically used to obtain the pixel values ​​of the N locations after the mean value has been removed by subtracting the pixel average reconstruction value from the pixel values ​​of the N locations.

[0755] In some embodiments, the prediction unit 11 specifically determines a second reconstruction region around the current block, determines the maximum and minimum reconstruction values ​​of the second reconstruction region, obtains a first prediction value based on the mean-removed pixel values ​​of the N locations, the filter coefficients, and the pixel-average reconstruction value, and uses the first prediction value, the maximum reconstruction value, and the minimum reconstruction value to determine the prediction value of the r-th point.

[0756] In some embodiments, the prediction unit 11 is specifically used to obtain a second predicted value for the r-th point by multiplying the pixel values ​​after mean removal of the N positions by the filter coefficient, and to obtain the first predicted value by adding the second predicted value to the pixel mean reconstruction value.

[0757] In some embodiments, the prediction unit 11 is specifically used to determine the first prediction value as the prediction value for the r-th point if the first prediction value is greater than the minimum reconstruction value and less than the maximum reconstruction value.

[0758] In some embodiments, the prediction unit 11 is specifically used to determine the minimum reconstruction value as the prediction value for the r-th point if the first prediction value is less than or equal to the minimum reconstruction value.

[0759] In some embodiments, the prediction unit 11 is specifically used to determine the maximum reconstruction value as the prediction value for the r-th point if the first prediction value is greater than or equal to the maximum reconstruction value.

[0760] In some embodiments, the prediction unit 11 is specifically used to determine the reconstruction regions above, to the left, to the right, to the left and to the left of the current block as the second reconstruction region.

[0761] In some embodiments, if the current block is a luminance block and the prediction mode of the current block is an interpolation filter prediction mode, the prediction unit 11 is further used to determine the prediction mode of the chroma block as the PLANAR mode or the intra-prediction mode corresponding to the prediction block, if the chroma block corresponding to the current block employs a direct derivation mode DM.

[0762] It is understood that the embodiments of the apparatus and the embodiments of the method are interchangeable, and that similar descriptions can be referred to in the embodiments of the method. To avoid duplication, detailed explanations are omitted here. Specifically, the apparatus 10 shown in Figure 28 can perform the decoding method on the decoding side in the embodiments of the present application, and the above and other operations and / or functions of each component within the apparatus 10 are for realizing the corresponding flows in each method, such as the above decoding method on the decoding side. For simplicity, detailed explanations are omitted here.

[0763] Figure 29 is a schematic block diagram of a video encoding device according to one embodiment of the present invention, and this video encoding device is applied to the encoder described above.

[0764] As shown in Figure 29, the video encoding device 20 includes the following configuration.

[0765] The prediction unit 21 determines the reference region and interpolation filter of the current block, and is used to determine the predicted block of the current block based on the reference region and the interpolation filter.

[0766] The conversion unit 22 is used to determine the intra-prediction mode corresponding to the prediction block and to determine the conversion kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block.

[0767] The encoding unit 23 is used to perform a transformation on the residual block of the current block based on the transformation kernel corresponding to the current block, obtain the transformation coefficients of the current block, and perform encoding based on the transformation coefficients of the current block to obtain a bitstream.

[0768] In some embodiments, the conversion unit 22 is specifically used to determine the angular values ​​of M points in the prediction block, where M is a positive integer, and is used to determine the intra-prediction mode corresponding to the prediction block based on the angular values ​​of the M points.

[0769] In some embodiments, the conversion unit 22 is specifically used to determine the horizontal and vertical slopes of the i-th point among the M points, where i is a positive integer less than or equal to M, and is used to determine the angle value of the i-th point based on its horizontal and vertical slopes.

[0770] In some embodiments, the conversion unit 22 is specifically used to determine the predicted value of a point in a sliding window centered on the i-th point within the prediction block, and to obtain the horizontal and vertical gradients of the i-th point based on the predicted value of the point in the sliding window and the horizontal and vertical gradient operators.

[0771] In some embodiments, the conversion unit 22 is specifically used to determine the horizontal gradient of the i-th point by multiplying the predicted value of a point in the sliding window by the horizontal gradient operator, and to determine the vertical gradient of the i-th point by multiplying the predicted value of a point in the sliding window by the vertical gradient operator.

[0772] In some embodiments, the conversion unit 22 is specifically used to determine the arctangent value of the ratio of the vertical slope to the horizontal slope of the i-th point as the angle value corresponding to the i-th point.

[0773] In some embodiments, the conversion unit 22 is specifically used to determine an intra-prediction mode corresponding to the M points based on the angular values ​​of the M points, and to determine an intra-prediction mode corresponding to the prediction block based on the intra-prediction modes corresponding to the M points.

[0774] In some embodiments, the conversion unit 22 is specifically used to determine gradient amplitude values ​​corresponding to the M points based on the horizontal and vertical gradients of the M points, and to determine the intra-prediction mode corresponding to the prediction block based on the intra-prediction mode and gradient amplitude values ​​corresponding to the M points.

[0775] In some embodiments, the conversion unit 22 is specifically used to add the absolute value of the horizontal slope and the absolute value of the vertical slope of the i-th point to obtain a slope amplitude value corresponding to the i-th point.

[0776] In some embodiments, the conversion unit 22 specifically accumulates the gradient amplitude value corresponding to any of the M points in the intra-prediction mode corresponding to that point, obtains the accumulated gradient amplitude value of the intra-prediction modes corresponding to the M points, and uses the intra-prediction mode with the largest accumulated gradient amplitude value among the intra-prediction modes corresponding to the M points to determine as the intra-prediction mode corresponding to the prediction block.

[0777] In some embodiments, the conversion unit 22 is further used to determine the first intra-prediction mode as the intra-prediction mode corresponding to the prediction block when the gradient amplitude values ​​corresponding to the M points are all 0.

[0778] Selectively, the first intra-prediction mode is the PLANAR mode.

[0779] In some embodiments, the conversion unit 22 is specifically used to obtain a correspondence between an intra-prediction mode and a conversion kernel group, where one conversion kernel group includes at least one type of conversion kernel, and is used to look up a first conversion kernel group corresponding to the intra-prediction mode of the prediction block within the correspondence, and is used to determine the conversion kernel corresponding to the current block from among the first conversion kernel group.

[0780] In some embodiments, the conversion unit 22 is specifically used to determine the type of conversion kernel corresponding to the current block, and to determine the conversion kernel of that type in the first conversion kernel group as the conversion kernel corresponding to the current block.

[0781] In some embodiments, the encoding unit 23 further writes the type of conversion kernel corresponding to the current block to the bitstream.

[0782] In some embodiments, the prediction unit 21 is specifically used to determine the reference region of the current block among a set of P reference regions, where P is a positive integer greater than 1.

[0783] In some embodiments, the prediction unit 21 is specifically used to determine a first cost for making a prediction for the current block based on the P reference regions, and to determine the reference region with the minimum first cost among the P reference regions as the reference region for the current block.

[0784] In some embodiments, the encoding unit 23 is further used to write first information to the bitstream, which is used to indicate the type of reference region of the current block.

[0785] In some embodiments, the prediction unit 21 is specifically used to determine the reference region of the current block from among a set of P reference regions based on the shape of the current block, where P is a positive integer greater than 1.

[0786] In some embodiments, the P reference regions include at least one of a first reference region, a second reference region, and a third reference region, wherein the first reference region includes the reconfiguration regions above, above, above, to the left, above, and above left of the current block; the second reference region includes the reconfiguration regions above, above, above, and above left of the current block; and the third reference region includes the reconfiguration regions to the left, above, and above left of the current block.

[0787] In some embodiments, the prediction unit 21 is specifically used to determine the interpolation filter for the current block among a set of Q interpolation filters, where Q is a positive integer greater than 1.

[0788] In some embodiments, the prediction unit 21 specifically determines the second cost of making a prediction for the current block using each of the Q interpolation filters, and uses the interpolation filter that minimizes the second cost among the Q interpolation filters to determine the interpolation filter for the current block.

[0789] In some embodiments, the encoding unit 23 is further used to write second information to the bitstream, which is used to indicate the shape of the interpolation filter for the current block.

[0790] In some embodiments, the prediction unit 21 is specifically used to determine the interpolation filter for the current block from among a set of Q interpolation filters based on the shape of the current block, where Q is a positive integer greater than 1.

[0791] In some embodiments, the Q interpolation filters include at least one of a first interpolation filter, a second interpolation filter, and a third interpolation filter, wherein the first interpolation filter is a square interpolation filter, the second interpolation filter is a rectangular interpolation filter whose width is greater than its height, and the third interpolation filter is a rectangular interpolation filter whose height is greater than its width.

[0792] In some embodiments, the encoding unit 23 further employs a langue-bound binary code encoding scheme and is used to write the first information and / or the second information to the bitstream.

[0793] In some embodiments, the prediction unit 21 is used to determine the prediction mode for the current block from among a plurality of candidate prediction modes before determining the reference region and interpolation filter for the current block, the plurality of candidate prediction modes including an interpolation filtering prediction mode, and is used to determine the reference region and interpolation filter for the current block when making a prediction if the prediction mode for the current block is the interpolation filtering prediction mode.

[0794] In some embodiments, the prediction unit 21 is used to determine whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement, prior to the step of determining the prediction mode of the current block from among a plurality of candidate prediction modes. If the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the prediction unit 21 is used to determine the prediction mode of the current block from among the plurality of candidate prediction modes.

[0795] In some embodiments, the prediction unit 21 is further used to determine that the current block will not employ the interpolation filter prediction mode if the position of the current block in the current image does not satisfy the predetermined position requirement and / or the size of the current block does not satisfy the predetermined block size requirement.

[0796] In some embodiments, the prediction unit 21 is used to determine whether the current sequence is permitted to perform predictions using the interpolation filter prediction mode, before determining whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement. If the current sequence is permitted to perform predictions using the interpolation filter prediction mode, the unit 21 is used to determine whether the position of the current block in the current image satisfies the predetermined position requirement and whether the size of the current block satisfies the predetermined block size requirement.

[0797] In some embodiments, the encoding unit 23 is further used to write third information to the bitstream. The third information is used to indicate whether the current block is performing predictions using an interpolation filtering prediction mode.

[0798] In some embodiments, the encoding unit 23 is further used to write a fourth piece of information to the bitstream. The fourth piece of information is used to indicate whether the current sequence is permitted to perform predictions using an interpolation filtering prediction mode.

[0799] In some embodiments, the prediction unit 21 is specifically used to determine the filter coefficients of the interpolation filter based on the reference region, to perform an interpolation filter prediction on the current block using the interpolation filter based on the filter coefficients, and to obtain a predicted block of the current block.

[0800] In some embodiments, the prediction unit 21 specifically determines a first reconstruction region around the current block, determines a pixel-average reconstruction value based on the reconstruction value of the first reconstruction region, performs mean removal on the reconstruction values ​​of the pixel points in the reference region based on the pixel-average reconstruction value, uses the pixel values ​​of the pixel points after mean removal in the reference region as input to the interpolation filter, and slides the interpolation filter within the reference region to obtain the filter coefficients of the interpolation filter.

[0801] In some embodiments, the prediction unit 21 is specifically used to determine the first reconstruction region based on the shape of the current block.

[0802] In some embodiments, the prediction unit 21 is specifically used to determine the reconstructed pixel area of ​​the top row and left column of the current block as the first reconstruction area when the shape of the current block is a square; or to determine the reconstructed pixel area of ​​the top row of the current block as the first reconstruction area when the shape of the current block is a rectangle with a width greater than its height; or to determine the reconstructed pixel area of ​​the left column of the current block as the first reconstruction area when the shape of the current block is a rectangle with a height greater than its width.

[0803] In some embodiments, the prediction unit 21 is specifically used to obtain the pixel values ​​of the pixel points in the reference region after the mean value has been removed by subtracting the pixel average reconstruction value from the reconstruction value of the pixel points in the reference region.

[0804] In some embodiments, the prediction unit 21 is specifically used to determine the pixel values ​​of N positions corresponding to the r-th point in the current block based on the shape of the interpolation filter, where r is a positive integer; to perform mean removal on the pixel values ​​of the N positions based on the pixel mean reconstruction value to obtain the pixel values ​​after mean removal of the N positions; to obtain a predicted value for the r-th point based on the pixel values ​​after mean removal of the N positions and the filter coefficients; and to obtain a predicted block of the current block based on the predicted value of the point in the current block.

[0805] In some embodiments, the prediction unit 21 is specifically used to determine the reconstructed value of any of the N positions as the pixel value of the position if the position is in the reconstructed region around the current block, or to determine the predicted value of the position as the pixel value of the position if the position is within the current block.

[0806] In some embodiments, the prediction unit 21 is specifically used to obtain the pixel values ​​of the N locations after the mean value has been removed by subtracting the pixel average reconstruction value from the pixel values ​​of the N locations.

[0807] In some embodiments, the prediction unit 21 specifically determines a second reconstruction region around the current block, determines the maximum and minimum reconstruction values ​​of the second reconstruction region, obtains a first prediction value based on the mean-removed pixel values ​​of the N locations, the filter coefficients, and the pixel mean reconstruction value, and uses the first prediction value, the maximum reconstruction value, and the minimum reconstruction value to determine the prediction value of the r-th point.

[0808] In some embodiments, the prediction unit 21 is specifically used to obtain a second predicted value for the r-th point by multiplying the pixel values ​​after average removal of the N positions by the filter coefficient, and to obtain the first predicted value by adding the second predicted value to the pixel average reconstruction value.

[0809] In some embodiments, the prediction unit 21 is specifically used to determine the first prediction value as the prediction value for the r-th point if the first prediction value is greater than the minimum reconstruction value and less than the maximum reconstruction value.

[0810] In some embodiments, the prediction unit 21 is specifically used to determine the minimum reconstruction value as the prediction value for the r-th point if the first prediction value is less than or equal to the minimum reconstruction value.

[0811] In some embodiments, the prediction unit 21 is specifically used to determine the maximum reconstruction value as the prediction value for the r-th point if the first prediction value is greater than or equal to the maximum reconstruction value.

[0812] In some embodiments, the prediction unit 21 is specifically used to determine the reconstruction regions above, to the left, to the right, to the left and to the left of the current block as the second reconstruction region.

[0813] In some embodiments, when the current block is a luminance block and the prediction mode of the current block is an interpolation filter prediction mode, the prediction unit 21 is further used to determine the prediction mode of the chroma block as the PLANAR mode or the intra-prediction mode corresponding to the prediction block, if the chroma block corresponding to the current block employs a direct derivation mode DM.

[0814] It is understood that the embodiments of the apparatus and the embodiments of the method are interchangeable, and that similar descriptions can be referred to in the embodiments of the method. To avoid redundancy, detailed explanations are omitted here. Specifically, the apparatus 20 shown in Figure 29 can perform the decoding method on the decoding side in the embodiments of the present application, and the above and other operations and / or functions of each component within the apparatus 20 are for realizing the corresponding flows in each method, such as the above decoding method on the decoding side. For simplicity, detailed explanations are omitted here.

[0815] The above description, with reference to the drawings, describes the apparatus and system according to the embodiment of the present application from the perspective of functional units. It is understood that the functional units may be implemented in hardware form, in software form instructions, or in combination of hardware and software units. Specifically, each step of the embodiment of the method in the embodiment of the present application can be performed by hardware integrated logic circuits and / or software form instructions in a processor. The steps of the method disclosed in the embodiment of the present application may be completed by being performed directly by a hardware decode processor, or by being completed by a combination of hardware and software units in a decode processor. Optionally, the software units may be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable read-only memory, or registers. The storage medium is located in memory, and the processor reads the information in memory and combines it with its own hardware to complete the steps of the embodiment of the method described above.

[0816] Figure 30 is a schematic block diagram of an electronic device according to an embodiment of the present invention.

[0817] As shown in Figure 30, the electronic device 30 may be a video encoder or video decoder as described in the embodiments of the present application, and the electronic device 30 may include a memory 33 and a processor 32.

[0818] The memory 31 is used to store the computer program 34 and to transmit the computer program 34 to the processor 32. In other words, the processor 32 can realize the method in the embodiment of the present invention by calling and executing the computer program 34 from the memory 31.

[0819] For example, the processor 32 is used to execute the steps in the video encoding method or video decoding method described above based on instructions in the computer program 34.

[0820] In some embodiments of the present application, the processor 32 is

[0821] This includes, but is not limited to, general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0822] In some embodiments of the present application, the memory 31 includes, but is not limited to, volatile memory and / or non-volatile memory.

[0823] Here, non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (Erasable PROM, EPROM), electrically erasable programmable read-only memory (Electrically EPROM, EEPROM), or flash memory. Volatile memory is random access memory (RAM), which functions as an external high-speed cache. Many forms of RAM are available, for example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch-link dynamic random access memory (SLDRAM), and direct Rambus random access memory (DR RAM).

[0824] In some embodiments of the present application, the computer program 34 may be divided into one or more units. The one or more units are stored in the memory 31 and executed by the processor 32 to complete the method provided by the present application. The one or more units are a set of computer program instruction segments capable of performing a specific function, and the instruction segments are used to describe the process by which the computer program 34 is executed in the electronic device 30.

[0825] As shown in Figure 30, the electronic device 30 may further include a transceiver 33.

[0826] The transceiver 33 can be connected to the processor 32 or the memory 31.

[0827] Here, the processor 32 can control the transceiver 31 to communicate with other devices, specifically by transmitting information or data to other devices or receiving information or data transmitted from other devices. The transceiver 31 may include a transmitter and a receiver. The transceiver 31 may further include an antenna, and the number of antennas may be one or more.

[0828] Each component of the electronic device 30 is connected via a bus system. This bus system includes a data bus, a power bus, a control bus, and a status signal bus.

[0829] Figure 31 is a schematic block diagram of a video encoding and decoding system according to an embodiment of the present invention.

[0830] As shown in Figure 31, the video encoding / decoding system 40 may include a video encoder 41 and a video decoder 42. Here, the video encoder 41 is used to perform the video encoding method according to the embodiment of the present application, and the video decoder 42 is used to perform the video decoding method according to the embodiment of the present application.

[0831] The present invention further provides a non-temporary computer storage medium. This computer storage medium stores a computer program, and when the computer program is executed by a computer, the computer is made to execute the method according to the embodiment of the method described above. In other words, embodiments of the present invention further provide a computer program product including instructions. When these instructions are executed by a computer, the computer is made to execute the method according to the embodiment of the method described above.

[0832] The present invention further provides a bitstream which is generated based on the encoding method described above.

[0833] When implemented by software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded into a computer and executed, a flow or function relating to the embodiments of the present application is generated, in whole or in part. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted by wired means (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless means (e.g., infrared, radio, microwave, etc.) from one website, computer, server, or data center to another. The computer-readable storage medium may be any available medium accessible to the computer, or a data storage device such as a server or data center that integrates one or more available media. The usable media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid state drives (SSDs)), etc.

[0834] As those skilled in the art will understand, each example unit and algorithmic step described in the embodiments of this application may be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the invention. Skilled technicians may implement the described functions using different methods for each specific application, but such implementation will not be considered beyond the scope of this application.

[0835] In some embodiments provided herein, it should be understood that the disclosed systems, apparatus, and methods may be implemented in other ways. For example, the embodiments of the apparatus described above are merely schematic. For example, the division of the units is merely a division of logical functions, and other division methods may exist in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not performed. Also, the mutual coupling, direct coupling, or communication connection indicated or discussed may be via some interface. Indirect coupling or communication connection between apparatus or units can be implemented in electrical, mechanical or other forms.

[0836] Units described as separate components may or may not be physically separate. Components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Depending on the actual needs, some or all of the units can be selected to achieve the objectives of this embodiment. For example, each functional unit in each embodiment of this application may be integrated into a single processing unit, each unit may exist individually in physical form, or two or more units may be integrated into a single unit.

[0837] The above describes only the specific implementation of the present application, but the scope of protection is not limited thereto. Any modification or substitution that a person skilled in the art can easily conceive within the technical scope disclosed in this application should be included in the scope of protection. Therefore, the scope of protection of this application shall be equivalent to the scope of protection of the claims.

Claims

1. A video decoding method, The steps include determining the reference region and interpolation filter of the current block, and determining the predicted block of the current block based on the reference region and the interpolation filter, The steps include determining an intra-prediction mode corresponding to the prediction block, and determining a conversion kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block, The process includes the steps of: performing an inverse transformation on the transformation coefficients of the current block based on the transformation kernel corresponding to the current block; obtaining the residual block of the current block; and obtaining the reconstructed block of the current block based on the predicted block and the residual block of the current block. A method characterized by the following:

2. The step of determining the intra prediction mode corresponding to the prediction block is: A step of determining the angle values ​​of M points in the prediction block, wherein M is a positive integer; The step includes determining an intra-prediction mode corresponding to the prediction block based on the angle values ​​of the M points, The step of determining the angle values ​​of M points in the prediction block is: A step of determining the horizontal and vertical slopes of the i-th point among the M points, wherein i is a positive integer less than or equal to M. The steps include determining the angle value of the i-th point based on the horizontal and vertical slopes of the i-th point, The method according to feature 1.

3. The step of determining the horizontal and vertical slopes of the i-th point is as follows: Within the prediction block, the steps include determining the predicted value of a point in a sliding window centered on the i-th point, The step of obtaining the horizontal and vertical gradients of the i-th point based on the predicted values ​​of the points in the sliding window, the horizontal gradient operator, and the vertical gradient operator, is included. The step of obtaining the horizontal and vertical gradients of the i-th point based on the predicted values ​​of the points in the sliding window, the horizontal gradient operator, and the vertical gradient operator is: The steps include determining the horizontal gradient of the i-th point by multiplying the predicted value of a point within the sliding window by the horizontal gradient operator, The step includes determining the vertical gradient of the i-th point by determining the product of the predicted value of a point in the sliding window and the vertical gradient operator. The method according to feature 2.

4. The step of determining the intra-prediction mode corresponding to the prediction block based on the angle values ​​of the M points is: The steps include determining an intra-prediction mode corresponding to the M points based on the angle values ​​of the M points, The step of determining an intra-prediction mode corresponding to the prediction block based on the intra-prediction modes corresponding to the M points, includes: The method according to feature 2.

5. The step of determining the intra-prediction mode corresponding to the prediction block based on the intra-prediction modes corresponding to the M points is: A step of determining the gradient amplitude value corresponding to the M points based on the horizontal and vertical gradients of the M points, The process includes the step of determining an intra-prediction mode corresponding to a prediction block based on the intra-prediction modes and gradient amplitude values ​​corresponding to the M points, The method according to feature 4.

6. The step of determining the gradient amplitude value corresponding to the M points based on the horizontal and vertical gradients of the M points is: The step includes adding the absolute value of the horizontal slope and the absolute value of the vertical slope of the i-th point among the M points to obtain the slope amplitude value corresponding to the i-th point. Or, The step of determining the intra-prediction mode corresponding to the prediction block based on the intra-prediction mode and gradient amplitude value corresponding to the M points is: For any of the M points, the gradient amplitude value corresponding to the point is accumulated in the intra-prediction mode corresponding to the point, and the accumulated gradient amplitude value of the intra-prediction mode corresponding to the M points is obtained. The process includes the step of determining, among the intra-prediction modes corresponding to the M points, the intra-prediction mode with the largest cumulative gradient amplitude value as the intra-prediction mode corresponding to the prediction block. The method according to specification 5.

7. The aforementioned method, If the gradient amplitude values ​​corresponding to the M points are all 0, the method further includes determining the first intra-prediction mode as the intra-prediction mode corresponding to the prediction block. The first intra prediction mode is the PLANAR mode. The method according to specification 5.

8. The step of determining the transformation kernel corresponding to the current block based on the intra prediction mode corresponding to the prediction block is: A step of obtaining a correspondence between an intra prediction mode and a transformation kernel group, wherein one transformation kernel group includes at least one type of transformation kernel. Among the aforementioned correspondences, the step of looking up the first transformation kernel group corresponding to the intra-prediction mode of the prediction block, The step includes determining the conversion kernel corresponding to the current block from among the first conversion kernel group, The step of determining the translation kernel corresponding to the current block from the first translation kernel group is: The steps include determining the type of conversion kernel corresponding to the current block, The step includes determining the conversion kernel of the conversion kernel type in the first conversion kernel group as the conversion kernel corresponding to the current block, The step of determining the type of translation kernel corresponding to the current block is: The steps include decrypting the bitstream and obtaining the type of transformation kernel corresponding to the current block, The method according to feature 1.

9. The step of determining the reference region of the current block is: A step of decoding a bitstream to obtain first information, wherein the first information is used to indicate the type of reference region of the current block, The step of determining the reference region of the current block from among P predetermined reference regions based on the type of the reference region, wherein P is a positive integer greater than 1, Or, The step of determining the reference region of the current block is: A step of determining the reference region of the current block from among P predetermined reference regions based on the shape of the current block, wherein P is a positive integer greater than 1. The method according to feature 1.

10. The P reference regions include at least one of the first, second, and third reference regions, wherein the first reference region includes the reconstruction regions above, above, above, to the left, below, and above left of the current block, the second reference region includes the reconstruction regions above, above, above, and above left of the current block, and the third reference region includes the reconstruction regions to the left, below, and above left of the current block. The method according to feature 9.

11. The step of determining the interpolation filter for the current block is: A step of decoding a bitstream to obtain second information, wherein the second information is used to indicate the shape of the interpolation filter of the current block, The step of determining the interpolation filter for the current block from among Q pre-set interpolation filters based on the shape of the interpolation filter, wherein Q is a positive integer greater than 1, Or, The step of determining the interpolation filter for the current block is: A step of determining the interpolation filter for the current block from among Q pre-set interpolation filters based on the shape of the current block, wherein Q is a positive integer greater than 1. The method according to feature 1.

12. The aforementioned method, The method further includes employing a decryption method for truncated binary code and decrypting and obtaining first information and / or second information from the bitstream, The method according to feature 9.

13. Before the step of determining the reference region and interpolation filter of the current block, the method: A step of decoding a bitstream to obtain third information, the third information being used to indicate whether the current block employs an interpolation filter prediction mode to perform a prediction, further comprising: The step of determining the current block's reference region and interpolation filter is: If it is determined that the current block will perform a prediction using the interpolation filter prediction mode based on the third information, the steps include determining the reference region and interpolation filter of the current block. Before the step of decrypting the bitstream and obtaining third information, the method, The method further includes the steps of determining whether the position of the current block in the current image satisfies a predetermined position requirement, and whether the size of the current block satisfies a predetermined block size requirement, The step of decrypting the bitstream to obtain third information is: If it is determined that the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the bitstream is decoded to obtain the third information, the process includes the step of obtaining the third information. The method according to feature 1.

14. The aforementioned method, The further step includes determining that if the position of the current block in the current image does not satisfy the predetermined position requirement, and / or if the size of the current block does not satisfy the predetermined block size requirement, the current block will perform the prediction without employing the interpolation filter prediction mode. The method according to the present invention, characterized by the present invention.

15. Before the step of determining whether the position of the current block in the current image satisfies a predetermined position requirement, and whether the size of the current block satisfies a predetermined block size requirement, the method: A step of decoding the bitstream to obtain fourth information, the fourth information being used to indicate whether the current sequence is permitted to perform predictions by employing the interpolation filter prediction mode, The steps of determining whether the position of the current block in the current image satisfies a predetermined position requirement, and whether the size of the current block satisfies a predetermined block size requirement, If the fourth piece of information indicates that the current sequence is permitted to perform predictions using the interpolation filter prediction mode, the process includes the steps of determining whether the position of the current block in the current image satisfies the predetermined position requirement and whether the size of the current block satisfies the predetermined block size requirement. The method according to the present invention, characterized by the present invention.

16. The step of determining the predicted block of the current block based on the reference region and the interpolation filter is: The steps include determining the filter coefficients of the interpolation filter based on the aforementioned reference region, The process includes the steps of: performing an interpolation filter prediction on the current block using the interpolation filter based on the filter coefficients, and obtaining a predicted block for the current block; The method according to feature 1.

17. The step of determining the filter coefficients of the interpolation filter based on the reference region is: The steps include determining a first reconstruction region around the current block, The steps include determining the pixel average reconstruction value based on the reconstruction value of the first reconstruction region, The steps include performing mean removal on the reconstruction values ​​of the pixel points in the reference region based on the aforementioned pixel average reconstruction values, The step includes taking the pixel values ​​of the pixel points after average value removal in the reference region as input to the interpolation filter, sliding the interpolation filter within the reference region, and obtaining the filter coefficients of the interpolation filter, The step of determining the first reconstruction region around the current block is: The step includes determining the first reconstruction region based on the shape of the current block, The method according to 16, characterized by...

18. If the current block is a luminance block and the prediction mode of the current block is an interpolation filter prediction mode, the method is: If the chroma block corresponding to the current block employs direct derivation mode DM, the further step includes determining the PLANAR mode or the intra-prediction mode corresponding to the prediction block as the prediction mode of the chroma block. The method according to feature 1.

19. A video encoding method, The steps include determining the reference region and interpolation filter of the current block, and determining the predicted block of the current block based on the reference region and the interpolation filter, The steps include determining an intra-prediction mode corresponding to the prediction block, and determining a conversion kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block, The process includes the steps of: performing a transformation on the residual block of the current block based on the transformation kernel corresponding to the current block; obtaining the transformation coefficients of the current block; performing encoding based on the transformation coefficients of the current block; and obtaining a bitstream. A method characterized by the following:

20. A computer-readable storage medium A computer program and a bitstream stored in the computer-readable storage medium, When the aforementioned computer program is executed by the processor, The steps include determining the reference region and interpolation filter of the current block, and determining the predicted block of the current block based on the reference region and the interpolation filter, The steps include determining an intra-prediction mode corresponding to the prediction block, and determining a conversion kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block, The processor is made to perform the following steps to generate the bitstream: a transformation is performed on the residual block of the current block based on the transformation kernel corresponding to the current block, the transformation coefficients of the current block are obtained, encoding is performed based on the transformation coefficients of the current block is performed, and a bitstream is obtained; A computer-readable storage medium characterized by the following features.