Video codec methods, apparatus, devices, systems, and storage media

Parallel prediction on multiple pixel points using interpolation filter coefficients addresses the inefficiencies in current video codecs, enhancing prediction speed and codec performance.

JP2026512534A5Pending Publication Date: 2026-05-01GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
Filing Date
2023-04-21
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Current interpolation filter prediction methods in video codecs suffer from low prediction efficiency and poor video codec performance due to sequential processing of pixel points, which affects overall codec performance.

Method used

Perform parallel prediction on at least two pixel points in a current block using determined filter coefficients of an interpolation filter, improving prediction efficiency and codec performance.

Benefits of technology

Enhances prediction speed and codec efficiency by performing parallel predictions on multiple pixel points within a block, thereby improving overall video decoding and encoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

This application provides a video codec method, apparatus, device, system, and storage medium. When predicting the current block, first a reference region and interpolation filter of the current block are determined, filter coefficients are determined based on the reference region, and parallel prediction is performed on at least two pixel points in the current block using the interpolation filter based on the filter coefficients to obtain a predicted block of the current block. A transformation kernel corresponding to the current block is determined, and the reconstruction value of the current block is determined based on the transformation kernel and the predicted block. That is, in the embodiments of this application, when performing interpolation filter prediction on the current block using the interpolation filter, parallel prediction is performed on at least two points in the current block to improve prediction speed and further improve codec efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video codec technology, and in particular, to a video codec method, apparatus, device, system, and storage medium.

Background Art

[0002] Digital video technology can be incorporated into various video devices such as, for example, digital televisions, smartphones, computers, e - book readers, or video players. With the development of video technology, the amount of data contained in video data has increased. To facilitate the transmission of video data, video devices perform video compression technology to transmit or store video data more efficiently.

[0003] Since there is temporal or spatial redundancy in video, the redundancy of the video can be removed or reduced by prediction, and the compression efficiency can be improved. To enhance the prediction effect, prediction compression may be performed using an interpolation filter prediction method. However, current interpolation filter prediction has a problem of low prediction efficiency and poor video codec performance.

Summary of the Invention

Means for Solving the Problems

[0004] Embodiments of this application provide a video decoding method, apparatus, equipment, system, and storage medium. When performing prediction using interpolation filter prediction, parallel prediction can be performed, further improving the prediction efficiency of the interpolation filter prediction mode and improving codec performance.

[0005] According to a first aspect, in this application, A video decoding method applied to a decoder, comprising: Determining a reference area and an interpolation filter of a current block, and determining filter coefficients of the interpolation filter based on the reference area; The steps include: determining the predicted block of the current block by performing parallel predictions for at least two pixel points in the current block using the interpolation filter based on the filter coefficients; A video decoding method is provided, which includes the steps of determining a transformation kernel corresponding to the current block, and determining a reconstruction block of the current block based on the transformation kernel corresponding to the current block and the predicted block.

[0006] According to a second aspect, in the embodiments of this application, A video encoding method applied to an encoder, The steps include determining the reference region and interpolation filter of the current block, and determining the filter coefficients of the interpolation filter based on the reference region, The steps include: determining the predicted block of the current block by performing parallel predictions for at least two pixel points in the current block using the interpolation filter based on the filter coefficients; A video encoding method is provided, which includes the steps of determining a transformation kernel corresponding to the current block, encoding the current block based on the transformation kernel corresponding to the current block and the predicted block, and obtaining a code stream.

[0007] According to a third aspect, the present application provides a video decoding apparatus for performing the method of the first aspect or each embodiment thereof. Specifically, the apparatus includes a functional unit for performing the method of the first aspect or each embodiment thereof.

[0008] According to a fourth aspect, the present application provides a video encoding apparatus for performing the method of the second aspect or each embodiment thereof. Specifically, the apparatus includes a functional unit for performing the method of the second aspect or each embodiment thereof.

[0009] According to a fifth aspect, a video decoder including a processor and memory is provided. The memory is for storing computer programs, and the processor is for calling and executing the computer programs stored in the memory in order to perform the methods of the first aspect or each of its embodiments.

[0010] According to a sixth aspect, a video encoder is provided which includes a processor and memory. The memory is for storing computer programs, and the processor is for calling and executing the computer programs stored in the memory in order to perform the methods of the second aspect or each of its embodiments.

[0011] According to the seventh aspect, a video codec system is provided which includes a video encoder and a video decoder. The video decoder is for performing the method of the first aspect or each embodiment thereof, and the video encoder is for performing the method of the second aspect or each embodiment thereof.

[0012] According to the eighth aspect, a chip is provided for implementing any one of the first to second aspects or the method of each embodiment described above. Specifically, the chip includes a processor that calls and executes a computer program from memory and causes a device on which the chip is mounted to execute any one of the first to second aspects or the method of each embodiment described above.

[0013] According to the ninth aspect, a computer-readable storage medium for storing a computer program is provided. The computer program causes a computer to execute any one of the first to second aspects or the method of each embodiment thereof.

[0014] According to the tenth aspect, a computer program product is provided which includes computer program instructions. The computer program instructions cause a computer to execute any one of the first to second aspects or the method of each embodiment thereof.

[0015] According to the eleventh aspect, a computer program is provided that, when executed on a computer, causes the computer to execute one of the first to second aspects or the method of each embodiment thereof.

[0016] Based on the above technical solutions, this application proposes an interpolation filter prediction method in which, when predicting the current block, first the reference region and interpolation filter of the current block are determined, the filter coefficients are determined based on the reference region, and based on the filter coefficients, parallel prediction is performed on at least two pixel points in the current block using the interpolation filter to obtain a predicted block of the current block. A transformation kernel corresponding to the current block is determined, and the reconstruction value of the current block is determined based on the transformation kernel and the predicted block. In other words, in the embodiment of this application, when performing interpolation filter prediction on the current block using the interpolation filter, parallel prediction is performed on at least two points in the current block to improve prediction speed and further improve codec efficiency. [Brief explanation of the drawing]

[0017] [Figure 1] This is a schematic block diagram of a video codec system according to an embodiment of this application. [Figure 2] This is a schematic block diagram of a video encoder according to an embodiment of the present application. [Figure 3] This is a schematic block diagram of a video decoder according to an embodiment of the present application. [Figure 4A] This is a schematic diagram of intranet prediction. [Figure 4B] This is a schematic diagram of intranet prediction. [Figure 5A] This is a schematic diagram of in-frame prediction. [Figure 5B] It is a schematic diagram of intra-frame prediction. [Figure 5C] It is a schematic diagram of intra-frame prediction. [Figure 5D] It is a schematic diagram of intra-frame prediction. [Figure 5E] It is a schematic diagram of intra-frame prediction. [Figure 5F] It is a schematic diagram of intra-frame prediction. [Figure 5G] It is a schematic diagram of intra-frame prediction. [Figure 5H] It is a schematic diagram of intra-frame prediction. [Figure 5I] It is a schematic diagram of intra-frame prediction. [Figure 6] It is a schematic diagram of the intra prediction mode. [Figure 7] It is a schematic diagram of the intra prediction mode. [Figure 8] It is a schematic diagram of the intra prediction mode. [Figure 9] It is a schematic diagram showing the principle of CCCM. [Figure 10] It is a schematic flowchart of a video decoding method according to an embodiment of the present application. [Figure 11] It is a schematic diagram showing the position of the current block in the current image. [Figure 12] It is a schematic diagram showing the reconstruction area. [Figure 13A] It is a schematic diagram showing the reference area. [Figure 13B] It is a schematic diagram showing the reference area. [Figure 13C] It is a schematic diagram showing the reference area. [Figure 14A] It is a schematic diagram showing the interpolation filter shape. [Figure 14B] It is a schematic diagram showing the interpolation filter shape. [Figure 14C] It is a schematic diagram showing the interpolation filter shape. [Figure 14D] It is a schematic diagram showing the interpolation filter shape. [Figure 14E] It is a schematic diagram showing the interpolation filter shape. [Figure 14F] This is a schematic diagram showing the interpolation filter shape. [Figure 14G] This is a schematic diagram showing the interpolation filter shape. [Figure 15] This is a schematic diagram showing the shapes of several interpolation filters according to embodiments of this application. [Figure 16] This is a schematic diagram showing the shapes of several interpolation filters according to embodiments of this application. [Figure 17] This is a schematic diagram showing the shapes of several interpolation filters according to embodiments of this application. [Figure 18A] This is a schematic diagram showing the shape of the interpolation filter according to an embodiment of this application. [Figure 18B] This is a schematic diagram showing the shape of the interpolation filter according to an embodiment of this application. [Figure 19] This is a schematic diagram showing the shapes of several interpolation filters according to embodiments of this application. [Figure 20A] This is the sliding step of the interpolation filter. [Figure 20B] This is a schematic diagram of the first reconstruction region. [Figure 21] This is a schematic diagram illustrating the movement of interpolation filters of different shapes in different types of reference regions. [Figure 22A] This is a schematic diagram showing how interpolation filters are used to predict the current block. [Figure 22B] This is a schematic diagram showing how interpolation filters are used to predict the current block. [Figure 23] This is a schematic diagram showing the interpolation prediction of the current block along the diagonal direction according to an embodiment of this application. [Figure 24A] This is a schematic diagram showing the direction of the diagonals. [Figure 24B] This is a schematic diagram showing the direction of the diagonals. [Figure 24C] This is a schematic diagram showing the direction of the diagonals. [Figure 25] This is a schematic diagram of the intra-prediction mode. [Figure 26] This is a schematic diagram for determining the horizontal and vertical slopes. [Figure 27]This is a histogram of gradient amplitude values. [Figure 28] This is a schematic flowchart of a video encoding method according to one embodiment of this application. [Figure 29] This is a schematic flowchart of the prediction mode determination according to one embodiment of this application. [Figure 30] This is a schematic block diagram of a video decoding device according to one embodiment of the present application. [Figure 31] This is a schematic block diagram of a video encoding device according to one embodiment of the present application. [Figure 32] This is a schematic block diagram of an electronic device according to an embodiment of this application. [Figure 33] This is a schematic block diagram of a video codec system according to an embodiment of this application. [Modes for carrying out the invention]

[0018] This application can be applied to the fields of image codecs, video codecs, hardware video codecs, dedicated circuit video codecs, real-time video codecs, and the like. For example, the technical solutions of this application can be incorporated into audio-video coding standards (AVS) such as the H.264 / Audio Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), and H.266 / Versatile Video Coding (VVC). Alternatively, the technical solutions of this application may be operated in combination with other exclusive or industry standards, including ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, and ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), which includes Scalable Video Codec (SVC) and Multiview Video Codec (MVC) extensions. It should be understood that the technology of this application is not limited to any specific codec standard or technology.

[0019] To facilitate understanding, we will first describe the video codec system according to the embodiment of this application with reference to Figure 1.

[0020] Figure 1 is a schematic block diagram of a video codec system according to an embodiment of the present application. Note that Figure 1 is merely an example, and the video codec system of the embodiment of the present application includes, but is not limited to, the one shown in Figure 1. As shown in Figure 1, the video codec system 100 includes an encoding device 110 and a decoding device 120. The encoding device encodes (may be understood as compression) video data to generate a code stream and transmits the code stream to the decoding device. The decoding device decodes the code stream encoded by the encoding device to obtain decoded video data.

[0021] The encoding device 110 in the embodiments of this application may be understood as a device having video encoding capabilities, and the decoding device 120 may be understood as a device having video decoding capabilities; that is, the encoding device 110 and decoding device 120 in the embodiments of this application include a broader range of devices, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, and in-vehicle computers.

[0022] In some embodiments, the encoding device 110 can transmit encoded video data (e.g., a code stream) to the decoding device 120 via channel 130. Channel 130 may include one or more media and / or devices that can transmit the encoded video data from the encoding device 110 to the decoding device 120.

[0023] In one example, channel 130 includes one or more communication media that enable the encoding device 110 to directly transmit encoded video data to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to the decoding device 120. The communication media includes, for example, a radio frequency spectrum, and optionally, the communication media may also include a wired communication medium, for example, one or more physical transmission lines.

[0024] In another example, channel 130 includes a storage medium capable of storing video data encoded by the encoding device 110. The storage medium includes various local-access data storage media such as optical discs, DVDs, and flash memory. In this example, the decoding device 120 can retrieve the encoded video data from the storage medium.

[0025] In another example, channel 130 may include a storage server that can store video data encoded by the encoding device 110. In this example, the decoding device 120 can download the encoded video data stored from the storage server. Optionally, the storage server can store the encoded video data and transmit the encoded video data to the decoding device 120, for example, a web server (e.g., for a website), a File Transfer Protocol (FTP) server, etc.

[0026] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113, where the output interface 113 may include a modulator / demodulator (modem) and / or transmitter.

[0027] In some embodiments, the encoding device 110 includes a video encoder 112 and output In addition to interface 113, it may also include video source 111.

[0028] The video source 111 may include at least one of the following: a video acquisition device (e.g., a video camera), a video archive, a video input interface for receiving video data from a video content provider, and a computer graphics system for generating video data.

[0029] The video encoder 112 encodes video data from the video source 111 and generates a code stream. The video data may include one or more pictures or sequences of pictures. The code stream contains the encoding information of the pictures or sequences of pictures as a bitstream. The encoding information may include encoded image data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. An SPS may include parameters that apply to one or more sequences. A PPS may include parameters that apply to one or more pictures. A syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the code stream.

[0030] The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data may also be stored in a storage medium or storage server for subsequent reading by the decoding device 120.

[0031] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122.

[0032] In some embodiments, the decoding device 120 may include a display device 123 in addition to the input interface 121 and the video decoder 122. Here, the input interface 121 includes a receiver and / or modem. The input interface 121 can receive encoded video data via channel 130.

[0033] The video decoder 122 decodes the encoded video data to obtain the decoded video data, and then transfers the decoded video data to the display device 123.

[0034] The display device 123 displays the decoded video data. The display device 123 may be integrated with the decoding device 120 or may be located outside the decoding device 120. The display device 123 can include various types of displays, such as liquid crystal displays (LCDs), plasma displays, organic light-emitting diode (OLED) displays, or other types of displays.

[0035] Furthermore, Figure 1 is merely an example, and the technical solution of the embodiment of this application is not limited to Figure 1. For example, the technology of this application can also be applied to one-sided video encoding and one-sided video decoding.

[0036] Next, a video encoding framework according to an embodiment of this application will be described.

[0037] Figure 2 is a schematic block diagram of a video encoder according to an embodiment of the present application. It should be understood that the video encoder 200 may be used for lossy compression of an image or for lossless compression of an image. This lossless compression may be visually lossless compression or mathematically lossless compression.

[0038] This video encoder 200 can be applied to image data in luminance-chromaticity (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents luminance (Luma), Cb (U) represents blue chromaticity, Cr (V) represents red chromaticity, and U and V represent hue and saturation as chromaticity (Chroma). For example, in the color format, 4:2:0 represents four luminance components and two chromaticity components (YYYYCbCr) for every four pixels, 4:2:2 represents four luminance components and four chromaticity components (YYYYCbCrCbCr) for every four pixels, and 4:4:4 represents full pixel display (YYYYCbCrCbCrCbCrCbCr).

[0039] For example, this video encoder 200 reads video data and, for each frame of image in the video data, divides the image of one frame into several coding tree units (CTUs), which in some examples may be called "tree blocks," "largest coding unit" (LCU), or "coding tree blocks" (CTB). Each CTU may be associated with a pixel block of equal size in the image. Each pixel may correspond to one luminance (or luma) sample and two chrominance (or chroma) samples. Thus, each CTU may be associated with one luminance sample block and two chrominance sample blocks. The size of a CTU may be, for example, 128×128, 64×64, 32×32, etc. A single CTU may then be further divided and coded into several coding units (CUs), which may be rectangular blocks or square blocks. A CU can be further divided into a prediction unit (PU) and a transform unit (TU), which allows for greater flexibility in encoding, prediction, and transform separation during processing. For example, a CTU can be divided into CUs using a quadtree scheme, and each CU can then be divided into TUs and PUs using a quadtree scheme.

[0040] Video encoders and video decoders can support a variety of PU sizes. Assuming a particular CU size is 2N×2N, video encoders and video decoders can support 2N×2N or N×N PU sizes for intra-prediction, and 2N×2N, 2N×N, N×2N, N×N, or similarly sized symmetric PUs for inter-prediction. Video encoders and video decoders can also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-prediction.

[0041] In some embodiments, as shown in Figure 2, the video encoder 200 may include a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filtering unit 260, a decoded image buffer 270, and an entropy coding unit 280. The video encoder 200 may also include more, fewer, or different functional components.

[0042] Optionally, in this application, the current block may be referred to as the current coding unit (CU) or current prediction unit (PU), etc. The prediction block is also called the prediction image block or image prediction block, and the reconstructed image block is also called the reconstruction block or image reconstructed image block.

[0043] In some embodiments, the prediction unit 210 includes an inter-prediction unit 211 and an intra-prediction unit 212. Because there is a strong correlation between adjacent pixels within a single frame of video, video codec techniques use intra-prediction methods to eliminate spatial redundancy between adjacent pixels. Because there is a strong similarity between adjacent frames of video, video codec techniques use inter-prediction methods to eliminate temporal redundancy between adjacent frames, thereby improving encoding efficiency.

[0044] The interprediction unit 211 can be used for interprediction, which can include motion estimation and motion compensation, and can reference image information from different frames. Interprediction uses motion information to find reference blocks from reference frames and generates prediction blocks to eliminate temporal redundancy based on the reference blocks. The frames used for interprediction may be P frames pointing to forward prediction frames and / or B frames pointing to bidirectional prediction frames. Interprediction uses motion information to find reference blocks from reference frames and generates prediction blocks based on the reference blocks. Motion information includes a list of reference frames in which the reference frames exist, a reference frame index, and a motion vector. The motion vector may be in integer pixel units or sub-pixel units. If the motion vector is in sub-pixel units, interpolation filtering must be used in the reference frames to create blocks of the desired sub-pixel units. The integer pixel or sub-pixel units in the reference frames found based on the motion vector are called reference blocks. Some techniques use the reference blocks directly as prediction blocks, while others process the reference blocks to generate prediction blocks. The process of generating prediction blocks by processing reference blocks can be understood as using the reference block as a prediction block and generating new prediction blocks based on the prediction blocks.

[0045] The intra-prediction unit 212 predicts pixel information in the current encoded image block by referring only to image information from the same frame, thereby eliminating spatial redundancy. The frame used for intra-prediction may be an I-frame.

[0046] Intra-prediction has multiple prediction modes. Taking the international digital video coding standard H-series as an example, the H.264 / AVC standard has 8 angle prediction modes and 1 non-angle prediction mode, while H.265 / HEVC extends this to 33 angle prediction modes and 2 non-angle prediction modes. HEVC uses Planar mode, DC, and 33 angle modes, for a total of 35 prediction modes. VVC uses Planar mode, DC, and 65 angle modes, for a total of 67 prediction modes.

[0047] Furthermore, as the number of angle modes increases, intra-prediction becomes more accurate, better meeting the requirements of high-resolution and ultra-high-resolution digital video development.

[0048] The residual unit 220 can generate residual blocks of the CU based on the pixel blocks of the CU and the predicted blocks of the PU of the CU. For example, the residual unit 220 can generate residual blocks of the CU in which each sample in the residual block has a value equal to the difference between the sample in the pixel blocks of the CU and the corresponding sample in the predicted blocks of the PU of the CU.

[0049] The conversion / quantization unit 230 can quantize the conversion coefficients. The conversion / quantization unit 230 can quantize the conversion coefficients associated with the TU of the CU based on the quantization parameter (QP) value associated with the CU. The video encoder 200 can adjust the degree of quantization applied to the conversion coefficients associated with the CU by adjusting the QP value associated with the CU.

[0050] The inverse transform / quantization unit 240 can reconstruct residual blocks from quantized transform coefficients by applying inverse quantization and inverse transform, respectively, to the quantized transform coefficients.

[0051] The reconstruction unit 250 can generate a reconstructed image block associated with the TU by adding the samples of the reconstructed residual block to the corresponding samples of one or more prediction blocks generated by the prediction unit 210. By reconstructing the sample block of each TU in the CU in this way, the video encoder 200 can reconstruct the pixel block of the CU.

[0052] The loop filtering unit 260 is for processing inversely transformed and inversely quantized pixels, complementing distortion information, and providing a better reference for subsequent encoded pixels. For example, it can perform a deblocking filtering operation to reduce the blocking effect of pixel blocks associated with the CU.

[0053] In some embodiments, the loop filtering unit 260 includes a deblocking filtering unit and a sample adaptive offset / adaptive loop filtering (SAO / ALF) unit, where the deblocking filtering unit is for removing blocking effects and the SAO / ALF unit is for removing ringing effects.

[0054] The decoded image buffer 270 can store reconstructed pixel blocks. The inter-prediction unit 211 can perform inter-prediction on other PUs of other images using a reference image containing the reconstructed pixel blocks. The intra-prediction unit 212 can perform intra-prediction on other PUs of the same image as the CU using the reconstructed pixel blocks in the decoded image buffer 270.

[0055] The entropy coding unit 280 can receive quantized transformation coefficients from the transformation / quantization unit 230. The entropy coding unit 280 can perform one or more entropy coding operations on the quantized transformation coefficients to generate entropy coded data.

[0056] Figure 3 is a schematic block diagram of a video decoder according to an embodiment of this application.

[0057] As shown in Figure 3, the video decoder 300 includes an entropy decoding unit 310, a prediction unit 320, an inverse quantization / conversion unit 330, a reconstruction unit 340, a loop filter unit 350, and a decoded image buffer 360. The video decoder 300 may include more, fewer, or different functional components.

[0058] The video decoder 300 can receive a code stream. The entropy decoding unit 310 can analyze the code stream and extract syntax elements from it. As part of the code stream analysis, the entropy decoding unit 310 can analyze the entropy-encoded syntax elements in the code stream. The prediction unit 320, the inverse quantization / conversion unit 330, the reconstruction unit 340, and the loop filtering unit 350 can decode the video data according to the syntax elements extracted from the code stream, i.e., generate decoded video data.

[0059] In some embodiments, the prediction unit 320 includes an intra-prediction unit 322 and an inter-prediction unit 321.

[0060] The intra-prediction unit 322 can perform intra-prediction to generate prediction blocks for the PU. The intra-prediction unit 322 can use an intra-prediction mode to generate prediction blocks for the PU based on spatially adjacent pixel blocks of the PU. The intra-prediction unit 322 can also determine the intra-prediction mode for the PU based on one or more syntax elements parsed from the code stream.

[0061] The interprediction unit 321 can construct a first reference image list (list 0) and a second reference image list (list 1) from the syntax elements analyzed from the code stream. Furthermore, if the PU uses interprediction coding, the entropy decoding unit 310 can analyze the motion information of the PU. The interprediction unit 321 can determine one or more reference blocks of the PU from the motion information of the PU. The interprediction unit 321 can generate prediction blocks of the PU from one or more reference blocks of the PU.

[0062] The inverse quantization / conversion unit 330 can dequantize the conversion coefficients associated with the TU. The inverse quantization / conversion unit 330 can determine the degree of quantization using the QP value associated with the CU of the TU.

[0063] After inverse quantization of the conversion coefficients, the inverse quantization / conversion unit 330 can apply one or more inverse transformations to the inverse quantization conversion coefficients to generate residual blocks associated with the TU.

[0064] The reconstruction unit 340 reconstructs the pixel blocks of the CU using the residual blocks associated with the TU of the CU and the predicted blocks of the PU of the CU. For example, the reconstruction unit 340 can reconstruct the pixel blocks of the CU by adding the samples of the residual blocks to the corresponding samples of the predicted blocks, thereby obtaining a reconstructed image block.

[0065] The loop filtering unit 350 can perform a deblocking filtering operation to reduce the blocking effect of pixel blocks associated with the CU.

[0066] The video decoder 300 can store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 can use the reconstructed image in the decoded image buffer 360 as a reference image for later prediction, or it can send the reconstructed image to a display device for rendering.

[0067] The basic flow of the video codec is as follows: On the encoding side, the image of one frame is divided into blocks, and for the current block, the prediction unit 210 generates a predicted block of the current block using intra-prediction or inter-prediction. The residual unit 220 can calculate a residual block, i.e., the difference between the predicted block and the original block of the current block, based on the predicted block and the original block of the current block. This residual block is also called residual information. This residual block can be processed to remove information that is insensitive to the human eye and to eliminate visual redundancy through processes such as transformation and quantization by the transformation / quantization unit 230. Optionally, the residual block before transformation and quantization by the transformation / quantization unit 230 may be called a time-domain residual block, and the time-domain residual block after transformation and quantization by the transformation / quantization unit 230 may be called a frequency residual block or frequency-domain residual block. The entropy coding unit 280 can receive the quantized change coefficients output from the transformation / quantization unit 230, entropy encode these quantized change coefficients, and output a code stream. For example, the entropy coding unit 280 can remove character redundancy based on the target context model and the probabilistic information of the binary code stream.

[0068] On the decoding side, the entropy decoding unit 310 analyzes the code stream to obtain prediction information and quantization coefficient matrix for the current block. Based on the prediction information, the prediction unit 320 generates a predicted block for the current block using intra-prediction or inter-prediction. The inverse quantization / transformation unit 330 uses the quantization coefficient matrix obtained from the code stream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block. The reconstruction unit 340 adds the predicted block and the residual block to obtain a reconstructed block. The reconstructed block constitutes a reconstructed image, and the loop filter unit 350 performs loop filtering on the reconstructed image based on the image or block to obtain a decoded image. The encoding side requires the same operation as the decoding side to obtain a decoded image. This decoded image is also called a reconstructed image, and the reconstructed image can be used as a reference frame for inter-prediction for subsequent frames.

[0069] Furthermore, the block partitioning information determined on the encoding side, as well as mode information or parameter information such as prediction, transformation, quantization, entropy coding, and loop filtering, are transported by the code stream as needed. The decoding side analyzes the code stream and, based on the existing information, determines the same block partitioning information as the encoding side, and determines mode information or parameter information such as prediction, transformation, quantization, entropy coding, and loop filtering, thereby ensuring that the decoded image obtained on the encoding side and the decoded image obtained on the decoding side are identical.

[0070] The above is a basic flow of a video codec under a block-based hybrid coding framework, and as technology advances, some modules or steps of this framework or flow can be optimized. This application applies to, but is not limited to, the basic flow of a video codec under a block-based hybrid coding framework.

[0071] In the embodiments of this application, the current block may be the current coding unit (CU) or the current prediction unit (PU), etc. Due to the need for parallel processing, the image can be divided into slices, etc., and slices within the same image can be processed in parallel, meaning they are data-independent. On the other hand, "frame" is a general term and is generally understood to mean one image. The frames described in this application may be replaced with images or slices, etc.

[0072] Intra prediction typically involves making predictions on the current coded block using both angular and non-angular modes to obtain a predicted block. Based on rate distortion information calculated from the predicted block and the original block, the optimal prediction mode for the current coded unit is selected, and this prediction mode is then transmitted to the decoding side via the code stream. The decoding side analyzes the prediction mode, predicts the predicted image of the current coded block, and obtains a reconstructed image by superimposing the residual pixels transmitted via the code stream. The intra prediction method uses the reconstructed pixels encoded and decoded around the current block as reference pixels to predict the current block. Figure 4A is a schematic diagram of intra prediction, and as shown in Figure 4A, the size of the current block is 4x4, and the left side of the current block is column and above line The pixels are the reference pixels of the current block, and intra-prediction uses these reference pixels to predict the current block. These reference pixels may all be available, i.e., all encoded / decoded, or some may be unavailable, for example, if the current block is the leftmost block in the entire frame, the reference pixels to the left of the current block may not be available. Alternatively, when encoding / decoding the current block, the bottom-left portion of the current block may not yet be encoded / decoded, and the bottom-left reference pixels may also be unavailable. If reference pixels are unavailable, they may be filled with available reference pixels, some values, or some methods, or they may not be filled at all.

[0073] Figure 4B is a schematic diagram of intraprediction. As shown in Figure 4B, the multiple reference line (MRL) intraprediction method can improve codec efficiency by using more reference pixels, for example, by using reference pixels of the current block with four reference rows / columns.

[0074] Furthermore, intraprediction has multiple prediction modes, and Figures 5A to 5I are schematic diagrams of intraprediction. As shown in Figures 5A to 5I, intraprediction for a 4x4 block in H.264 can mainly include nine modes. Here, pattern 0 shown in Figure 5A copies the pixels above the current block vertically to the current block as predicted values, pattern 1 shown in Figure 5B copies the left reference pixels horizontally to the current block as predicted values, DC in pattern 2 shown in Figure 5C uses the average of eight points A to D and I to L as predicted values ​​for all points, and patterns 3 to 8 shown in Figures 5D to 5I copy the reference pixels to the corresponding positions in the current block at a certain angle. Because some positions in the current block do not exactly correspond to the reference pixels, it is necessary to use the weighted average of the reference pixels, i.e., the interpolated subpixels of the reference pixels.

[0075] In addition to these, there are other models such as Plane and Planar, but with technological advancements and the expansion of blocks, angle prediction models are becoming more numerous. Figure 6 is a schematic diagram of intra-prediction modes, and as shown in Figure 6, the intra-prediction modes used in HEVC include Planar, DC, and 33 angle modes, for a total of 35 prediction modes. Figure 7 is a schematic diagram of intra-prediction modes, and as shown in Figure 7, the intra-prediction modes used by VVC include Planar, DC, and 65 types of angle modes, for a total of 67 prediction modes. Figure 8 is a schematic diagram of intra-prediction modes, and as shown in Figure 8, A VS3 uses a total of 66 prediction modes, including DC, Plane, Bilinear, PCM, and 62 different angular modes.

[0076] Furthermore, there are techniques to improve prediction, such as improving sub-pixel interpolation of reference pixels and filtering of predicted pixels. In AVS3, the multiple intraprediction filter (MIPF) generates predicted values ​​using different filters for different block sizes. For pixels at different positions within the same block, pixels close to the reference pixel use one type of filter to generate predicted values, while pixels far from the reference pixel use a different type of filter to generate predicted values. Techniques for filtering predicted pixels, such as the intraprediction filter (IPF) in AVS3, can filter predicted values ​​using reference pixels.

[0077] In some embodiments, adaptive loop filtering (ALF) techniques are currently used in the loop filter unit of video codecs. For example, the reconstructed image is filtered using ALF techniques to obtain the final decoded image.

[0078] Next, we will discuss adaptive loop filter (ALF) technology.

[0079] ALF is a type of loop filter designed based on the Wiener filter principle to minimize the error between a target sample and an input sample. In a loop filter, the target sample is the original image, and the input is the reconstructed image.

[0080] Before performing filtering using ALF, first determine the filter coefficients.

[0081] For example, by creating a Wiener-Hoff equation as shown in equation (1) and solving this Wiener-Hoff equation, the filter coefficients of the interpolation filter can be obtained.

number

[0082] JPEG2024216632000003.jpg15167

[0083] For example, the filter coefficients can be obtained by solving the above Wiener-Hoff equation using the Cholesky decomposition autocorrelation matrix.

[0084] After determining the filter coefficients of the filter based on equation (1) above, the samples to be filtered are filtered using the following equation (2) to obtain the filtered samples.

number

[0085] The Convolutional Cross Component Model (CCCM) is a process for predicting chromaticity pixels from reconstructed luminance pixels. Its advantage is that the coefficients of the CCCM filter can be obtained from the decoder side using the reconstructed pixels, thereby eliminating the overhead of storing filter coefficients in the code stream, as in ALF. As shown in Figure 9, the CCCM coefficients are calculated from the reconstructed pixels around the chromaticity block that is currently to be predicted and the reconstructed pixels around the luminance block at the position corresponding to this chromaticity block.

[0086] To improve video compression performance, we propose an interpolation filter prediction mode in which the filter coefficients of the interpolation filter are determined by the reconstruction region around the current block, and based on these filter coefficients, the interpolation filter is used to predict each point in the current block, obtaining the predicted value for each point in the current block, and thus obtaining the predicted block of the current block.

[0087] In related technologies, when interpolation filter prediction is performed for each pixel point in the current block using an interpolation filter, the interpolation filter is applied to each pixel point one by one. That is, after the interpolation filter prediction for the previous pixel point is completed, the interpolation filter prediction for the next pixel point is performed, and when performing the interpolation filter prediction for the next pixel point, the predicted value of the previous pixel point is used. From this, it can be seen that in conventional technologies, when interpolation filter prediction is performed for the current block, pixel points are predicted sequentially, and only one pixel point can be predicted at a time, so the prediction efficiency decreases and, consequently, affects the overall codec performance of the video.

[0088] To solve the above technical problems, the embodiment of this application improves prediction efficiency and enhances video codec performance by performing parallel prediction on pixel points in the current block when the current block performs prediction using interpolation filter prediction mode.

[0089] Next, with reference to Figure 10, the video decoding method according to the embodiment of this application will be described using the decoding side as an example.

[0090] Figure 10 is a schematic flowchart of a video decoding method according to one embodiment of the present application, applied to the video decoders shown in Figures 1 and 3. As shown in Figure 10, the method of the embodiment of the present application includes the following steps. S101 determines the reference region and interpolation filter of the current block, and determines the filter coefficients of the interpolation filter based on the reference region.

[0091] When the decryption side decrypts the current block, it decrypts the code stream, obtains the quantization coefficients of the current block, performs inverse quantization on the quantization coefficients to obtain the transformation coefficients of the current block, performs inverse transformation on the transformation coefficients to obtain the residual value of the current block. Next, it determines the prediction mode of the current block, determines the predicted value of the current block based on the prediction mode, and obtains the reconstructed value of the current block based on the predicted value and residual value of the current block.

[0092] In some embodiments, the current block is also a block that should be predicted.

[0093] In the embodiments of this application, the decoding side first determines the prediction mode of the current block.

[0094] In some embodiments, the method by which the decryption side determines the prediction mode of the current block includes at least the following:

[0095] In Method 1, the encoding side determines the prediction mode for the current block. For example, among the candidate prediction modes consisting of the conventional prediction mode and the interpolation filter prediction mode shown in Figure 6 or Figure 7, the candidate prediction mode with the minimum cost is determined as the prediction mode for the current block. Next, the encoding side adds instruction information for the prediction mode of the current block to the code stream. In this way, the decoding side decodes the code stream to obtain instruction information for the prediction mode of the current block, further determines the prediction mode of the current block based on that instruction information, further predicts the current block using that intra-prediction mode, and obtains the predicted value of the current block.

[0096] For example, if the terminal device determines that the prediction mode of the current block is a conventional prediction mode, it writes the index of the current block's prediction mode to the code stream as instruction information for that prediction mode. The decoding side obtains the prediction mode index by decoding the code stream, and then determines the prediction mode of the current block from the conventional prediction modes shown in Figure 6 or Figure 7 based on this index.

[0097] In Method 2, the encoding side creates an intra-prediction mode candidate list and selects the intra-prediction mode for the current block from this list. Note that the interpolation filter prediction mode is included in this intra-prediction mode candidate list. Next, the encoding side writes the number (or index number) of the current block's intra-prediction mode in the intra-prediction mode candidate list to the code stream. In this way, the decoding side decodes the code stream to determine the number of the current block's intra-prediction mode in the intra-prediction mode candidate list, and, similar to the encoding side, creates an intra-prediction mode candidate list (note that the created intra-prediction mode candidate list includes the interpolation filter prediction mode). Furthermore, based on the number of the current block's intra-prediction mode in the intra-prediction mode candidate list, the decoding side determines the current block's intra-prediction mode from the created intra-prediction mode candidate list. Finally, the current block is predicted using the determined intra-prediction mode of the current block, and the predicted value of the current block is obtained.

[0098] In method 3, the encoding side creates an intra-prediction mode candidate list including interpolation filter prediction modes, then selects an intra-prediction mode for the current block from the intra-prediction mode candidate list, determines the cost of each candidate prediction mode in the intra-prediction mode candidate list on the current block's template, and then determines the intra-prediction mode for the current block based on the cost. Correspondingly, the decoding side, similar to the encoding side, creates an intra-prediction mode candidate list including interpolation filter prediction modes, then determines the cost of each candidate prediction mode in the intra-prediction mode candidate list on the current block's template, and determines the intra-prediction mode for the current block based on the cost. Finally, the current block is predicted using the determined intra-prediction mode for the current block, and the predicted value for the current block is obtained.

[0099] In Method 4, the encoding and decoding sides, by default, use the interpolation filter prediction mode for the current block to perform predictions.

[0100] The decoding side can determine whether the current block uses interpolation filter prediction mode using methods 1 to 4 described above, and can also determine whether the current block uses interpolation filter prediction mode using method 5 described below.

[0101] In method 5, the decoding side decodes the code stream and obtains third information indicating whether the current block will perform predictions using the interpolation filter prediction mode. If the decoding side determines, based on the third information, that the current block will perform predictions using the interpolation filter prediction mode, it determines the reference region and interpolation filter of the current block.

[0102] In this method 5, when the encoding side decides that the current block uses interpolation filter prediction mode, it writes third information to the code stream, and the decoding side obtains the third information by decoding the code stream, and then decides whether or not the current block makes a prediction using interpolation filter prediction mode based on this third information. If the third information indicates that the current block makes a prediction using interpolation filter prediction mode, the decoding side makes a prediction of the current block using the interpolation filter prediction mode and obtains a predicted block of the current block. If the third information indicates that the current block makes a prediction without using interpolation filter prediction mode, the decoding side skips the step of the current block making a prediction using interpolation filter prediction mode, and instead determines the prediction mode of the current block, makes a prediction of the current block using the determined prediction mode, and obtains a predicted block of the current block.

[0103] The embodiments of this application are not limited to the specific representation of the third piece of information described above, but may also be arbitrary indicator information indicating whether or not the current block performs a prediction using the interpolation filter prediction mode.

[0104] In one example, the third piece of information can be represented as intra_eip_flag. In this way, whether the current block makes a prediction using the interpolation filter prediction mode can be determined by assigning different values ​​to intra_eip_flag. For example, if intra_eip_flag=0, it indicates that the current block makes a prediction without using the interpolation filter prediction mode, and if intra_eip_flag=1, it indicates that the current block makes a prediction using the interpolation filter prediction mode. In this way, the encoding side writes this preset flag intra_eip_flag to the code stream, and the decoding side determines the prediction mode of the current block based on the value of this decoded preset flag intra_eip_flag. For example, if this preset flag intra_eip_flag=1, it indicates that the prediction mode of the current block is the interpolation filter prediction mode, and the decoding side further determines that the current block makes a prediction using the interpolation filter prediction mode.

[0105] In some embodiments, the conditions for using the interpolation filter prediction mode are limited. Based on this, a decision is made whether or not the current image block is allowed to perform predictions using the interpolation filter prediction mode before determining the reference region and interpolation filter of the current block.

[0106] The embodiments of this application do not limit the specific method for determining whether the current image block is permitted to perform predictions using the interpolation filter prediction mode. In other words, there are no restrictions on the specific conditions under which the interpolation filter prediction mode can be used.

[0107] In some embodiments, in order to improve the prediction accuracy of the interpolation filter prediction mode, the interpolation filter prediction mode is used for some blocks that meet the requirements and not for some blocks that do not meet the requirements. Based on this, before decoding the code stream and obtaining the third information, the decoding side determines whether the position of the current block in the current image meets a predetermined position requirement and whether the size of the current block meets a predetermined block size. If the position of the current block in the current image meets the predetermined position requirement and the size of the current block meets the predetermined block size, the decoding side decodes the code stream and obtains the third information.

[0108] The embodiments of this application are not limited to predetermined positional requirements and predetermined block sizes, but are specifically determined according to actual needs.

[0109] In one example, as shown in Figure 11, if the position of the upper left corner of the current image is (0,0) and the position of the upper left corner of the current block is (x,y), then the predetermined position requirements are that the x of the current block is greater than or equal to a first predetermined value XX, and the y of the current block is greater than or equal to a second predetermined value YY.

[0110] The embodiments of this application do not limit the specific values ​​of the first and second predetermined values ​​described above.

[0111] For example, the first predetermined value and the second predetermined value are the same.

[0112] For example, if both the first predetermined value and the second predetermined value are 13, that is, if the distance from the top edge of the current block to the top edge of the current image is 13 rows of pixels or more, and the distance from the left edge of the current block to the left edge of the current image is 13 columns of pixels or more, then the position of the current block in the current image satisfies the predetermined position requirement.

[0113] In one example, continuing to refer to Figure 11, if the current block width is W and the current block height is H, then the predetermined block size requirement is that the current block width W is less than or equal to the third predetermined value A, and the current block height H is less than or equal to the fourth predetermined value B.

[0114] The embodiments of this application do not limit the specific values ​​of the third and fourth predetermined values ​​described above.

[0115] For example, the third predetermined value and the fourth predetermined value are the same.

[0116] For example, if both the third and fourth predetermined values ​​are 32, that is, if the current block's width and height are both 32 or less, it means that the current block meets the predetermined block size requirement.

[0117] In the embodiments of this application, before determining whether the current block performs prediction using the interpolation filter prediction mode, the decoding side first determines whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement. If the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the decoding side decodes the code stream, obtains third information, and determines whether the current block performs prediction using the interpolation filter prediction mode based on the third information. For example, as shown in Figure 11, if the distance from the top edge of the current block to the top edge of the current image is 13 rows of pixels or more, the distance from the left edge of the current block to the left edge of the current image is 13 columns of pixels or more, and the width and height of the current block are both 32 or less, the decoding side decodes the code stream and obtains third information.

[0118] In some embodiments, the first predetermined value, the second predetermined value, the third predetermined value, and the fourth predetermined value are default values.

[0119] In some embodiments, the first predetermined value, the second predetermined value, the third predetermined value, and the fourth predetermined value are values ​​decoded from the code stream by the decoder.

[0120] In some embodiments, if the position of the current block in the current image does not meet a predetermined position requirement, and / or the size of the current block does not meet a predetermined block size requirement, it is determined that the current block will be predicted without using an interpolation filter prediction mode.

[0121] In some embodiments, the decoding side further includes the steps of decoding a code stream and obtaining second information indicating whether the current sequence is permitted to perform predictions using an interpolation filter prediction mode, before determining whether the position of the current block in the current image satisfies predetermined positional requirements and whether the size of the current block satisfies predetermined block size, and if the second information indicates that the current sequence is permitted to perform predictions using an interpolation filter prediction mode, determining whether the position of the current block in the current image satisfies predetermined positional requirements and whether the size of the current block satisfies predetermined block size.

[0122] In embodiments of this application, a high-level syntax element, such as a second piece of information at the sequence level, indicates whether the current sequence is permitted to perform predictions using the interpolation filter prediction mode. If the second piece of information indicates that the current sequence is permitted to perform predictions using the interpolation filter prediction mode, the decoding side determines whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size. Furthermore, if it is determined that the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies a predetermined block size requirement, the decoding side decodes the third piece of information and determines whether the current block is permitted to perform predictions using the interpolation filter prediction mode.

[0123] In some embodiments, if the second information indicates that the current sequence is not permitted to be predicted using the interpolation filter prediction mode, the decoding side skips the steps described above: determining whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement, and skips the step of decoding the third information.

[0124] The embodiments of this application are not limited to specific representations of the second information, and may include any directive information that can indicate whether the current sequence is permitted to perform predictions using the interpolation filter prediction mode.

[0125] For example, the second piece of information can be represented as sps_eip_enabled_flag, which allows determining whether the current sequence is permitted to perform predictions using interpolation filter prediction mode by assigning different values ​​to sps_eip_enabled_flag. For instance, sps_eip_enabled_flag=0 indicates that the current sequence is not permitted to perform predictions using interpolation filter prediction mode, while sps_eip_enabled_flag=1 indicates that the current sequence is permitted to perform predictions using interpolation filter prediction mode.

[0126] For example, the second piece of information is conveyed by a sequence-level parameter set (SPS), as shown in Table 1. [Table 1] Here, sps_eip_enabled_flag represents second information, which is carried by seq_parameter_set_rbsp(). For example, sps_eip_enabled_flag=0 indicates that the current sequence is not allowed to perform predictions using interpolation filter prediction mode, and sps_eip_enabled_flag=1 indicates that the current sequence is allowed to perform predictions using interpolation filter prediction mode.

[0127] In some embodiments, embodiments of this application may also include general constraints information (GCI) identifier bits to indicate whether interpolation filter prediction techniques are used. For example, whether the current video enables interpolation filter prediction techniques is indicated by gci_no_eip_constraint_flag. For example, as shown in Table 2, this gci_no_eip_constraint_flag is carried by general constraints information general_constraints_info(). [Table 2]

[0128] As shown in Table 2, gci_no_eip_constraint_flag=1 indicates that the current video does not enable interpolation filter prediction, meaning that the sequence-level interpolation filter intra-prediction must be 0 for all images, meaning that the use of interpolation filter intra-prediction is not permitted for any sequence in the current video. gci_no_eip_constraint_flag=0 indicates that the current video enables interpolation filter prediction, meaning that the sequence-level interpolation filter intra-prediction is not restricted to being 0 for all images.

[0129] From the above, it is conceivable that the syntax elements of the embodiment of this application include high-level syntax elements gci_no_eip_constraint_flag and sps_eip_enabled_flag, and block-level intra_eip_flag. The decoding side first decodes the high-level syntax elements, that is, first decodes gci_no_eip_constraint_flag, and if gci_no_eip_constraint_flag=0, then decodes sps_eip_enabled_flag, and if sps_eip_enabled_flag=1, analyzes the block syntax elements.

[0130] As an example, block-level syntax elements are shown in Table 3. [Table 3]

[0131] In Table 3, cbWidth and cbHeight may be understood as the width and height of the current block, SIZE_A as the third predetermined value, SIZE_B as the fourth predetermined value, XX as the first predetermined value, YY as the second predetermined value, and x0 and y0 as the coordinate difference between the top-left corner of the current block and the top-left corner of the current image.

[0132] As can be seen from Table 3 above, when the second sequence-level piece of information sps_eip_enabled_flag=1, that is, when it is indicated that the current sequence is permitted to use interpolation filter prediction mode, it is determined whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement. If it is determined that the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the third piece of information intra_eip_flag is decoded, and based on the decoded third piece of information intra_eip_flag, it is determined whether the current block will perform prediction using interpolation filter prediction mode.

[0133] From the above, determining whether the current block uses interpolation filter prediction mode can be limited by higher-level syntax elements, such as GCI, sequence level, frame level, slice level, and block level. It can also be limited by the size and position of the current block.

[0134] In some embodiments, the computational cost and complexity increase when the interpolation filter prediction mode is used for several smaller blocks. This is because the interpolation filter prediction mode is computationally complex in this application, and using it for several smaller blocks increases the number of times the interpolation filter prediction mode is used throughout the image decoding process, further increasing the computational cost and complexity of the image. Based on this, in embodiments of this application, the use of the interpolation filter prediction mode is permitted only for slightly larger blocks. For example, the use of the interpolation filter prediction mode is permitted only if the size of the current block is greater than or equal to a predetermined size. If the size of the current block is smaller than the predetermined size, the current block is not permitted to use the interpolation filter prediction mode. Embodiments of this application do not limit the specific value of the predetermined size. For example, the current block being greater than or equal to a predetermined size may mean that the number of pixel points in the current block is greater than or equal to a predetermined number, or that at least one of the length and width of the current block is greater than or equal to a predetermined value, or that the ratio of the length and width of the current block is greater than or equal to a predetermined ratio, etc.

[0135] In some embodiments, if the current block is in the first row of the current CTU, it is determined that the current block does not allow the use of the interpolation filter prediction mode. That is, if the current block is to perform predictions using the interpolation filter prediction mode, the current block is not located in the first row of the current CTU.

[0136] In some embodiments, the decision of whether the current block allows the use of interpolation filter prediction mode also depends on the type of the current image. For example, for intra-prediction images (i.e., images that use intra-prediction when predicting), it is specified that prediction can be made using interpolation filter prediction mode, while for inter-prediction images (i.e., images that use inter-prediction when predicting), prediction using interpolation filter prediction mode is not permitted. Based on this, if the current image in which the current block exists is an intra-prediction image, it is decided that the current block allows prediction using interpolation filter prediction mode. If the current image is not an intra-prediction image (e.g., an inter-prediction image), it is decided that the current block does not allow prediction using interpolation filter prediction mode.

[0137] In some embodiments, several complex intra-prediction modes are introduced in the ECM reference software to improve codec performance, such as template-based intra-prediction derivation (TIMD), decoder-side intra-prediction derivation (DIMD), template-based multiple reference line intra-prediction (TMRL), spatial geometrical partitioning mode (SGPM), and convolutional cross-component model (CCCM). All of these complex intra-prediction modes are based on template matching techniques, and the interpolation filter prediction mode in the embodiments of this application also uses information from the reconstructed region (which may be understood as the template region) during use. Therefore, in the embodiments of this application, the interpolation filtering prediction mode can also be classified as an intra-prediction mode based on template matching techniques. Based on this, in the embodiments of this application, the intra-prediction modes based on the template matching technology are uniformly indicated using unified identification information (e.g., first information). For example, if the first information indicates that the template matching technology is not enabled, it means that none of the intra-prediction modes based on the template matching technology (i.e., TIMD, DIMD, TMRL, SGPM, TMRL, CCCM, and interpolation filtering prediction modes) are permitted to be used. TaIf this is indicated, then the use of the intra-predictive modes of the above-mentioned template matching-based techniques is permitted, and it is explained that the intra-predictive mode to be specifically used by the current block will be further determined based on other information.

[0138] Based on the above description, in embodiments of the present application, the step of determining whether the current block is permitted to use the interpolation filter prediction mode includes the steps of decoding the code stream and obtaining first information indicating whether the template matching-based technique is enabled, and determining whether the current block is permitted to use the interpolation filter prediction mode based on the first information. For example, if the first information indicates that the template matching-based technique is not enabled, it is determined that the current block is not permitted to make predictions using the interpolation filter prediction mode. Alternatively, for example, if the first information indicates that the template matching-based technique is enabled, the decoding side determines, based on other information, whether the current block makes predictions using the interpolation filter prediction mode.

[0139] The embodiments of this application do not limit the specific forms of expression of the first information described above.

[0140] For example, the first piece of information described above may be GCI, sequence-level, frame-level, slice-level, or block-level instruction information.

[0141] In one example, if the first piece of information is sequence-level instruction information, the decoder decodes the code stream to obtain the first piece of information. If the first piece of information indicates that the template matching technique is enabled, the decoder continues to decode the code stream to obtain the second piece of information (sps_eip_enabled_flag), and then, based on the second piece of information, decides whether the current block is allowed to use the interpolation filter prediction mode. If the first piece of information indicates that the template matching technique is not enabled, the decoder directly decides that the current block will make predictions without applying the interpolation filter prediction mode and skips the step of decoding the second piece of information.

[0142] In some embodiments, the decoding side determines whether the current block can be predicted using the interpolation filter prediction mode, including at least one of the following conditions: 1) Is the current image an intra-predicted image or not? 2) Whether or not higher-level syntax allows it. This is optional, and higher-level syntax includes sequence level, frame level, slice level, block level, etc. See the explanation above for details. 3) Whether the current block size and shape are permitted. See the explanation above for details. 4) Whether the current block's location is permitted. See the explanation above for details.

[0143] The above explains the specific process for determining whether the current block will perform predictions using the interpolation filter prediction mode.

[0144] In the embodiments of this application, when the decoding side determines that the current block is to be predicted using the interpolation filter prediction mode, it predicts the current block using the interpolation filter prediction mode and obtains the predicted value of the current block.

[0145] Next, we will describe the process by which the decoding side predicts the current block using the interpolation filter prediction mode.

[0146] When the decoding side decides that the current block will be predicted using the interpolation filter prediction mode, it first determines the reference region and interpolation filter for the current block.

[0147] Next, we will explain the specific process by which the decryption side determines the reference region of the current block.

[0148] In the embodiments of this application, the reference region of the current block is part or all of the reconfigured region surrounding the current block.

[0149] For example, as shown in Figure 12, the reconfiguration region surrounding the current block may include the upper reconfiguration region of the current block, the left reconfiguration region of the current block, the upper right reconfiguration region of the current block, the lower left reconfiguration region of the current block, and the upper left reconfiguration region of the current block. Here, the block to be predicted in Figure 12 is the current block.

[0150] The embodiments of this application do not limit the specific shape and size of the reference region of the current block.

[0151] For example, the reference region of the current block includes one of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For instance, the reference region of the current block is the upper reconstruction region of the current block, or the reference region of the current block is the left reconstruction region of the current block.

[0152] For example, the reference region of the current block includes any two of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region and the left reconstruction region of the current block. Alternatively, for example, the reference region of the current block includes the upper reconstruction region and the lower left reconstruction region of the current block.

[0153] For example, the reference region of the current block includes any three of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region of the current block, the upper right reconstruction region of the current block, and the upper left reconstruction region of the current block. Also, for example, the reference of the current block region This includes the left reconfiguration area of ​​the current block, the upper left reconfiguration area of ​​the current block, and the lower left reconfiguration area of ​​the current block.

[0154] For example, the reference region of the current block includes any four of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region of the current block, the upper right reconstruction region of the current block, the upper left reconstruction region of the current block, and the left reconstruction region of the current block. Also, for example, the reference of the current block region This includes the left reconfiguration region of the current block, the upper left reconfiguration region of the current block, the lower left reconfiguration region of the current block, and the upper reconfiguration region of the current block.

[0155] In one example, the reference region of the current block includes five reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block.

[0156] In the embodiment of this application, the decoding side determines the reference region of the current block from a predetermined P reference regions.

[0157] In the embodiments of this application, the specific method by which the decoding side determines the reference region of the current block from a predetermined P reference regions includes, but is not limited to, the following methods.

[0158] In Method 1, the reference region of the current block is the default region. For example, the encoding and decoding sides assume, by default, that the reference region of the current block includes at least one of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block.

[0159] In method 2, the decryption side decrypts the code stream to obtain a fourth piece of information indicating the type of reference region of the current block, and based on the type of reference region, determines the reference region of the current block from a predetermined P reference regions, where P is a positive integer greater than 1.

[0160] In this embodiment, the encoding side determines the reference region of the current block from a predetermined P reference regions. For example, the encoding side determines the encoding cost corresponding to each of these P reference regions and determines the reference region with the minimum encoding cost as the reference region of the current block. The type of the reference region with the minimum encoding cost is then indicated to the decoding side by the fourth piece of information. In this way, the decoding side obtains the fourth piece of information by decoding the code stream and determines the reference region of the current block from the predetermined P reference regions based on the type of reference region indicated by this fourth piece of information.

[0161] Note that the type and shape of these predetermined P reference regions are all different.

[0162] The embodiments of this application do not particularly limit the specific number and shape of the P reference regions.

[0163] In one example, the P reference regions include at least one of the first reference region, the second reference region, and the third reference region.

[0164] Here, as shown in Figure 13A, the first reference region is above, to the right of, to the left of the current block. under , and the upper left reconstruction area. As shown in Figure 13B, the second reference area includes the reconstruction areas above, to the right of, and to the upper left of the current block. As shown in Figure 13C, the third reference area includes the left side of the current block, and to the left under , and the reconstruction region in the upper left. The block to be predicted in Figures 13A to 13C is the current block.

[0165] The embodiments of this application do not limit the specific form of representation of the fourth information, and any indicating information that can indicate the type of reference region of the current block is acceptable.

[0166] For example, eip_ref_type is used to represent a fourth piece of information, and for instance, different values ​​of eip_ref_type indicate different types of reference regions.

[0167] For example, see Figures 13A to 13 C The correspondence between the three reference regions shown and the value of eip_ref_type is shown in Table 4. [Table 4]

[0168] Based on Table 4 above, the decryption side decrypts the code stream, obtains the fourth piece of information eip_ref_type, and determines the reference region of the current block based on the value of the fourth piece of information eip_ref_type. For example, if eip_ref_type=0, the reference region of the current block is determined as the first reference region, and as shown in Figure 13A, the first reference region is above, to the right of, to the left of, and to the left of the current block. under , and the upper left reconstruction region. If eip_ref_type=1, the reference region of the current block is determined as the second reference region, and as shown in Figure 13B, the second reference region includes the top, upper right, and upper left reconstruction regions of the current block. If eip_ref_type=2, the reference region of the current block is determined as the third reference region, and as shown in Figure 13C, the third reference region includes the current block's left, bottom left, and includes the reconstruction area in the upper left.

[0169] In the above explanation, the P reference regions were described as the three reference regions shown in Figures 13A to 13C. While the P reference regions in the embodiments of this application may include reference regions other than the three mentioned above, the embodiments of this application are not limited to these. The correspondence between the reference regions shown in Table 4 and the values ​​of eip_ref_type can be adaptively adjusted according to the number of reference regions.

[0170] In some embodiments, the decryption side can decrypt a fourth piece of information from the code stream using a truncated binary code decryption scheme.

[0171] For example, the correspondence between truncated binary code, the value of eip_ref_type, and the type of reference region is shown in Table 5. [Table 5]

[0172] In the embodiments of this application, the decryption side can decrypt the codeword of the truncated binary code using either an equal-probability decoding scheme or a context-model decoding scheme.

[0173] The decryption side can determine the reference region of the current block using either method 1 or method 2 described above, or it can determine the reference region of the current block using method 3 described below.

[0174] In method 3, the reference region of the current block is determined from a predetermined number of P reference regions based on the shape of the current block.

[0175] In this method 3, predictions are made using different reference regions for current blocks of different shapes, thereby improving the accuracy of the predictions.

[0176] For example, if the current block is square in shape, the first type of reference region is used.

[0177] For example, if the current block is a rectangle with a width greater than its height, a second type of reference area is used.

[0178] For example, if the current block is a rectangle with a width smaller than its height, a third type of reference area is used.

[0179] In other words, in the embodiment of this application, the correspondence between P reference regions and the shape of the current block is predetermined. Thus, the decoding side can determine the reference region of the current block from among the P reference regions based on the correspondence between the P reference regions and the shape of the current block, according to the shape of the current block.

[0180] Next, we will describe the process by which the decoding side determines the interpolation filter for the current block.

[0181] In the embodiments of this application, the specific shape of the interpolation filter is not limited.

[0182] For example, the interpolation filters according to the embodiments of this application include, but are not limited to, square interpolation filters and interpolation filters where the height is less than the width.

[0183] For example, square interpolation filters include, but are not limited to, the 4x4 interpolation filter shown in Figure 14A.

[0184] Furthermore, interpolation filters where the height is greater than the width include, but are not limited to, the 5×3 interpolation filter shown in Figure 14B, the 6×2 interpolation filter shown in Figure 14D, and the 7×1 interpolation filter shown in Figure 14G.

[0185] Furthermore, interpolation filters where the height is smaller than the width include, but are not limited to, the 3×5 interpolation filter shown in Figure 14C, the 2×6 interpolation filter shown in Figure 14E, and the 1×7 interpolation filter shown in Figure 14F.

[0186] JPEG2024216632000011.jpg15169

[0187] In the embodiments of this application, the decoding side determines the interpolation filter for the current block from a predetermined Q interpolation filters.

[0188] In the embodiments of this application, the decoding side determines the interpolation filter for the current block from a predetermined Q interpolation filters, and the specific methods include, but are not limited to, the following.

[0189] In Method 1, the interpolation filter for the current block is the default interpolation filter. For example, the encoding and decoding sides default to using one of the Q interpolation filters shown in Figures 14A to 14G as the interpolation filter for the current block. For example, the default interpolation filter is a 4x4 interpolation filter.

[0190] In method 2, the decoding side decodes the code stream and obtains a fifth piece of information indicating the shape of the interpolation filter for the current block. Based on the shape of the interpolation filter for the current block, it determines the interpolation filter for the current block from a predetermined Q number of interpolation filters, where Q is a positive integer greater than 1.

[0191] In this embodiment, the encoding side determines the interpolation filter for the current block from a predetermined Q interpolation filters. For example, the encoding side determines the encoding cost corresponding to each of these Q interpolation filters and determines the interpolation filter with the minimum encoding cost as the interpolation filter for the current block. Next, the shape of the interpolation filter with the minimum determined encoding cost is shown to the decoding side by the fifth piece of information. In this way, the decoding side obtains the fifth piece of information by decoding the code stream, and further determines the interpolation filter for the current block from the predetermined Q interpolation filters based on the shape of the interpolation filter shown by the fifth piece of information.

[0192] Note that the shapes of these predetermined Q interpolation filters are all different.

[0193] The embodiments of this application do not particularly limit the specific number and shape of the Q interpolation filters. For example, the Q interpolation filters include at least one of a first interpolation filter which is a square interpolation filter, a second interpolation filter which is a rectangle whose width is greater than its height, and a third interpolation filter which is a rectangle whose height is greater than its width.

[0194] In one example, Q interpolation filters are shown in Figures 14A to 14. G Includes multiple interpolation filters.

[0195] The embodiments of this application do not limit the specific representation format of the fifth piece of information, and any instruction information capable of indicating the shape of the interpolation filter of the current block is acceptable.

[0196] In one example, the fifth piece of information is represented by eip_filter_type, and different interpolation filters of different shapes are indicated by the value of eip_filter_type.

[0197] For example, if the Q interpolation filters are the five interpolation filters shown in Figure 15, the correspondence between the five interpolation filters and the value of eip_filter_type is shown in Table 6. [Table 6]

[0198] Based on Table 5 above, the decoder decodes the code stream, obtains the fifth piece of information, eip_filter_type, and then determines the interpolation filter for the current block based on the value of the fifth piece of information, eip_filter_type. For example, if eip_filter_type=0, the shape of the interpolation filter for the current block is determined to be 4x4. If eip_filter_type=1, the shape of the interpolation filter for the current block is determined to be 3x5. If eip_filter_type=2, the shape of the interpolation filter for the current block is determined to be 5x3. If eip_filter_type=3, the shape of the interpolation filter for the current block is determined to be 2x6. If eip_filter_type=4, the shape of the interpolation filter for the current block is determined to be 6x2.

[0199] In some embodiments, the decryption side can decrypt a fifth piece of information from the code stream using a truncated binary code decryption scheme.

[0200] For example, if a given number of interpolation filters includes the five interpolation filters shown in Figure 15, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 7. [Table 7]

[0201] In this case, the five types of interpolation filter shapes shown in Table 7 and the three types of reconstruction region types shown in Table 5 result in a total of 15 combinations of interpolation filters and reconstruction regions.

[0202] In some embodiments, the decryption side can obtain the fifth piece of information, eip_filter_type, by decoding the code stream, and then determine the interpolation filter for the current block from Table 7 based on the shape of the interpolation filter indicated by the fifth piece of information, eip_filter_type. Similarly, the fourth piece of information is obtained by decoding the code stream, and then the reference region of the current block is determined from Table 5 based on the value of the fourth piece of information, eip_ref_type.

[0203] As an example, the syntax elements of the embodiment of this application are shown in Table 8. [Table 8]

[0204] As shown in Table 8, when the decoding side decodes the code stream, it first obtains a fifth piece of sequence-level information, sps_eip_enabled_flag, which indicates whether the current sequence is allowed to perform predictions using interpolation filter prediction mode. Next, it determines whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement. If it is determined that the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, it decodes a third piece of information, intra_eip_flag, which indicates whether the current block performs predictions using interpolation filter prediction mode. If the third piece of information intra_eip_flag=1 indicates that the current block performs predictions using interpolation filter prediction mode, it decodes the code stream and obtains a fourth piece of information, eip_ref_type, and a fifth piece of information, eip_filter_type. Here, since the fourth piece of information, eip_ref_type, indicates the type of reference region of the current block, the decoding side can obtain the reference region of the current block by table lookup based on the value of the fourth piece of information, eip_ref_type. Here, the fifth piece of information, eip_filter_type, indicates the shape of the interpolation filter for the current block. Furthermore, based on the value of the fifth piece of information, eip_filter_type, the interpolation filter for the current block is obtained by a table lookup.

[0205] In some embodiments, when the embodiments of this application include seven interpolation filters as shown in Figure 16, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 9. [Table 9]

[0206] In this case, the seven types of interpolation filter shapes shown in Table 9 and the three types of reconstruction region types shown in Table 5 result in a total of 21 combinations of interpolation filters and reconstruction regions.

[0207] Similarly, the decoder can obtain the reference region and interpolation filter of the current block by decoding the syntax shown in Table 8 and looking up Table 5 and Table 8 above.

[0208] In some embodiments, when the embodiments of this application include the three interpolation filters shown in Figure 17, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 10. [Table 10]

[0209] In this case, the three types of interpolation filter shapes shown in Table 10 and the three types of reconstruction region types shown in Table 5 result in a total of nine combinations of interpolation filters and reconstruction regions.

[0210] Similarly, the decoder can obtain the reference region and interpolation filter of the current block by decoding the syntax shown in Table 8 and looking up Table 5 and Table 10 above.

[0211] In some embodiments, when the embodiments of this application include the three interpolation filters shown in Figure 18A, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 11. [Table 11]

[0212] In this case, the three types of interpolation filter shapes shown in Table 11 and the three types of reconstruction region types shown in Table 5 result in a total of nine combinations of interpolation filters and reconstruction regions.

[0213] Similarly, the decoder can obtain the reference region and interpolation filter of the current block by decoding the syntax shown in Table 8 and looking up Table 5 and Table 11 above.

[0214] In some embodiments, when the embodiments of this application include the three interpolation filters shown in Figure 18B, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 12. [Table 12]

[0215] In this case, the three types of interpolation filter shapes shown in Table 12 and the three types of reconstruction region types shown in Table 5 result in a total of nine combinations of interpolation filters and reconstruction regions.

[0216] Similarly, the decryption side decrypts the syntax shown in Table 8, thereby obtaining Table 5 and the above table. 12 By looking up this, you can obtain the reference region and interpolation filter of the current block.

[0217] Generally, using a filter with more taps for the same number of samples yields a better interpolation effect. Figure 18B shows an increased number of multi-taps in the filter compared to the 2x6 and 6x2 interpolation filters shown in Figure 18A, expanding to, for example, 2x8 and 8x2 interpolation filters. However, in reality, the 2x8 and 8x2 filters and the 4x4 filter all take 15 samples as input and have one output, so their complexity is similar. Therefore, the interpolation filters shown in Figure 18B increase the interpolation effect without increasing complexity.

[0218] In some embodiments, when the embodiments of this application include three interpolation filters as shown in Figure 19, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 13.

Table 13

[0219] In this case, the three types of interpolation filter shapes shown in Table 13 above and the three types of reconstruction region types shown in Table 5 above have a total of nine combinations of interpolation filters and reconstruction regions.

[0220] Similarly, the decoding side can obtain the reference region and interpolation filter of the current block by decoding the syntax shown in Table 8 and looking up Tables 5 and Table 13 above.

[0221] The decoding side can determine the interpolation filter of the current block using the method of Method 1 or Method 2 above. In addition, the interpolation filter of the current block can be determined using the following Method 3. [[ID=...]] [[ID=...]]

[0222] [[ID=...]] Method 3 determines the interpolation filter of the current block from a predetermined Q interpolation filters based on the shape of the current block.

[0223] In this Method 3, for current blocks of different shapes, prediction is performed using different interpolation filters to improve the prediction accuracy.

[0224] For example, when the shape of the current block is square, an interpolation filter of the first shape is used.

[0225] Also, for example, when the shape of the current block is a rectangle with a width larger than the height, an interpolation filter of the second shape is used.

[0226] Also, for example, when the shape of the current block is a rectangle with a width smaller than the height, an interpolation filter of the third shape is used.

[0227] That is, in the embodiments of the present application, the correspondence between the Q interpolation filters and the shape of the current block is preset. Thus, on the decoding side, according to the shape of the current block, the interpolation filter for the current block can be determined from the Q interpolation filters based on the correspondence between the Q interpolation filters and the shape of the current block.

[0228] Next, the determination of the filter coefficients of the interpolation filter based on the reference region will be described.

[0229] In the embodiments of the present application, the methods by which the decoding side determines the filter coefficients of the interpolation filter include at least some of the following methods.

[0230] In Method 1, the determined interpolation filter is slid in the reference region of the current block to create the Wiener-Hopf equation. Next, the Wiener-Hopf equation is solved to obtain the filter coefficients of the interpolation filter.

[0231] JPEG2024216632000020.jpg45167

[0232] In one example, the interpolation filter is slid within the reference region of the current block to create the Wiener-Hopf equation as shown in Equation (3).

Equation

[0233] JPEG2024216632000023.jpg29168

[0234] As an example, the decoding side can obtain the filter coefficients of the filter by solving the Wiener-Hopf equation shown in Equation (3) using the Cholesky decomposition autocorrelation coefficient matrix.

[0235] The embodiments of the present application do not limit the sliding step of the interpolation filter within the reference region.

[0236] In one example, as shown in Figure 20A, the horizontal and vertical slide steps of the interpolation filter within the reference region are equal, and both are 1 pixel.

[0237] In one example, the horizontal and vertical slide steps of the interpolation filter within the reference region are not equal. For instance, the horizontal slide step is 2 pixels and the vertical slide step is 1 pixel. Alternatively, for example, the horizontal slide step is 1 pixel and the vertical slide step is 2 pixels.

[0238] In one example, at least one of the horizontal slide step and vertical slide step of the interpolation filter within the reference region is greater than a predetermined step. For example, the horizontal slide step is greater than a predetermined step. Also, for example, the vertical slide step is greater than a predetermined step. Also, for example, both the horizontal slide step and the vertical slide step are greater than a predetermined step. The embodiments of this application do not limit the specific value of the predetermined step. For example, it can be a number such as 1, 2, or 3.

[0239] In method 2, the decoding side determines the filter coefficients by following steps S101-A1 to S101-A4. S101-A1, determine the first reconstruction region around the current block. S101-A2 determines the average reconstruction value of pixels based on the reconstruction value of the first reconstruction region. S101-A3 removes the average value from the reconstruction values ​​of the pixel points in the reference region, based on the pixel average reconstruction value. S101-A4: The pixel values ​​of the pixel points with the mean value removed in the reference region are used as input to the interpolation filter, and the interpolation filter is slid within the reference region to obtain the filter coefficients of the interpolation filter.

[0240] In this method 2, the reference region is subjected to mean value removal, and the filter coefficients of the interpolation filter are determined based on the reference region after mean value removal. Since the amount of data decreases when the reference region is subjected to mean value removal, the determination efficiency of the filter coefficients can be improved when determining the filter coefficients based on the reference region after mean value removal.

[0241] Specifically, on the decoding side, first, a first reconstruction region that can be any partial reconstruction region in the reconstruction region around the current block is determined.

[0242] In the embodiments of this application, the method by which the decoding side determines the first reconstruction region around the current block includes at least some of the following methods.

[0243] In method 1, on the decoding side, by default, one reconstruction region around the current block is determined as the first reconstruction region.

[0244] For example, as shown in FIG. 20B, on the decoding side, by default, a region consisting of one row above the current block, one column on the left side, and one pixel point at the upper left corner is determined as the first reconstruction region.

[0245] In method 2, the first reconstruction region is determined based on the shape of the current block.

[0246] For example, when the shape of the current block is square, a reconstruction pixel region of one row above the current block and one column on the left side is determined as the first reconstruction region.

[0247] Also, for example, when the shape of the current block is a rectangle with a width greater than the height, a reconstruction pixel region of one row above the current block is determined as the first reconstruction region.

[0248] Also, for example, when the shape of the current block is a rectangle with a height greater than the width, a reconstruction pixel region of one column on the left side of the current block is reconstructed as the first reconstruction region.

[0249] The method for determining the first reconstruction region based on the current block shape includes, but is not limited to, the examples described above.

[0250] After the decoding side determines the first reconstruction region, it determines the pixel-average reconstruction value m based on the reconstruction value of the first reconstruction region.

[0251] The embodiments of this application are not limited to a specific method for determining the pixel average reconstruction value m based on the reconstruction value of the first reconstruction region in S101-A2 described above.

[0252] In method 1, steps S101-A2 include determining the average value of the reconstruction values ​​in the first reconstruction region as the pixel average reconstruction value m.

[0253] In one example of Method 1, if the first reconstruction region is as shown in Figure 20B, the pixel-average reconstruction value m is calculated using the method shown in Table 14. [Table 14]

[0254] In another example of Method 1, if the first reconstruction region is the top row and / or leftmost column of the current block, the average of the reconstruction values ​​of the top row and / or leftmost column can be determined as the pixel-average reconstruction value m. In this case, the pixel-average reconstruction value m can be calculated using the method shown in Table 15. [Table 15]

[0255] As shown in Table 15 above, if the first reconstruction region is the row above and / or the column to the left of the current block, the pixel-average reconstruction value m can be quickly calculated using a shift operation instead of division.

[0256] In method 2, steps S101-A2 include determining the pixel average reconstruction value based on the current block shape and the reconstruction value of the first reconstruction region.

[0257] For example, if the current block is a square, the average value of the entire first reconstruction region determined above is determined as the pixel-average reconstruction value m.

[0258] In one example of method 2, the first reconstruction region includes an upper reconstruction region and a left reconstruction region of the current block, in which case the step of determining the pixel average reconstruction value based on the shape of the current block and the reconstruction value of the first reconstruction region includes the step of determining the first region from the upper reconstruction region and the left reconstruction region based on the shape of the current block, the step of determining the average reconstruction value of the first region based on the reconstruction value of the first region, and the step of determining the pixel average reconstruction value based on the average reconstruction value of the first region.

[0259] In this example, if the first reconstruction region includes the upper and left-side reconstruction regions of the current block, the mean value is determined, similar to the DC prediction mode, to reduce computational complexity. Specifically, the first region is determined from the upper and left-side reconstruction regions included in the first reconstruction region, based on the shape of the current block.

[0260] For example, if the width of the current block shape is greater than its height, the upper reconstruction region is determined to be the first region.

[0261] For example, if the height of the current block shape is greater than its width, the left-side reconstruction region is determined to be the first region.

[0262] For example, if the height of the current block shape is equal to its width, the upper reconstruction region and the left-side reconstruction region are determined as the first region.

[0263] Next, the average reconstruction value of the selected first region is determined based on the reconstruction value of the first region, and further, the pixel average reconstruction value m is determined based on the average reconstruction value of the first region. For example, this average reconstruction value of the first region is determined as the pixel average reconstruction value m.

[0264] In other words, in this example, if the width of the current block shape is greater than the height, the average reconstruction value of the upper reconstruction region of the current block is determined as the pixel average reconstruction value m. If the height of the current block shape is greater than the width, the average reconstruction value of the left reconstruction region of the current block is determined as the pixel average reconstruction value m. If the height of the current block shape is equal to the width, the average reconstruction value of both the upper and left reconstruction regions of the current block is determined as the pixel average reconstruction value m.

[0265] For example, the decoding side can calculate the pixel-average reconstruction value m using the method shown in Table 16. [Table 16]

[0266] As shown in Table 16 above, a reconstruction region (i.e., a first region) for calculating the pixel average reconstruction value m is determined based on the current block shape, and then the pixel average reconstruction value m is quickly calculated based on the selected first region. In the process of calculating the pixel average reconstruction value m, division is achieved by a simple shift, avoiding the problem of large division calculations when calculating the average value due to the difference in size between the left and upper reconstruction regions caused by the different lengths and widths of the current blocks, and further improving the calculation speed of the pixel average reconstruction value m, and predicting the current block. Efficiency and Improve speed.

[0267] The decoding side determines the pixel-average reconstruction value, and then removes the average value from the reconstruction value of the pixel points in the reference region based on that pixel-average reconstruction value.

[0268] For example, for each pixel point in the reference region, the reconstructed value of that pixel point is divided by the above-mentioned average reconstructed value and rounded to obtain the pixel value of the pixel point in the reference region after the mean value has been removed.

[0269] For example, the decoding side subtracts the pixel average reconstruction value from the reconstruction value of the pixel point in the reference region to obtain the pixel value of the pixel point in the reference region from which the mean value has been removed. For example, for each pixel point in the reference region, the above pixel average reconstruction value is subtracted from the reconstruction value of that pixel point to obtain the pixel value of the pixel point in the reference region from which the mean value has been removed.

[0270] The embodiments of this application do not limit the specific method by which the decoding side removes the average value from the reconstructed values ​​of the pixel points in the reference region based on the pixel average reconstructed value.

[0271] The decoding side removes the mean value from the reconstructed pixel points of the reference region based on the above method, obtains the pixel values ​​of the pixel points from which the mean value has been removed in the reference region, then executes step S101-A4 above, uses the pixel values ​​of the pixel points from which the mean value has been removed in the reference region as input to the interpolation filter, slides the interpolation filter within the reference region, and obtains the filter coefficients of the interpolation filter.

[0272] For example, as shown in Figure 21, if the interpolation filter for the current block consists of five different shapes, and the reference region for the current block consists of three different reference regions, the interpolation filter for the current block is slid across the current block's reference region (which has had its mean value removed) to obtain the filter coefficients of the interpolation filter. The interpolation filter may be slid one row horizontally or one column vertically across the reference region (which has had its mean value removed). Here, the block to be predicted in Figure 21 is the current block.

[0273] JPEG2024216632000027.jpg45168

[0274] In one example, the Wiener-Hoff equation created by sliding the interpolation filter within the reference region of the current block is given by equation (4).

number

[0275] JPEG2024216632000029.jpg23168

[0276] JPEG2024216632000030.jpg30168

[0277] In one example, the decoding side can obtain the filter coefficients of the filter by solving the Wiener-Hoff equation shown in equation (4) above using the Cholesky decomposition autocorrelation matrix.

[0278] Based on the steps described above, the decoding side determines the filter coefficients of the interpolation filter and then performs the next step S102.

[0279] S102, based on the filter coefficients, an interpolation filter is used to predict at least two pixel points in the current block in parallel, and the predicted block for the current block is determined.

[0280] Based on the above steps, the decoding side determines the filter coefficients of the interpolation filter, and then, using the interpolation filter based on these filter coefficients, performs an interpolation filter prediction for the current block and obtains the predicted block for the current block.

[0281] In related technologies, when performing interpolation filter prediction for the current block using an interpolation filter, the next pixel point is predicted after the prediction of one pixel point in the current block is completed. For example, as shown in Figure 22A, the interpolation filter performs interpolation filter prediction one by one along the horizontal direction for each pixel point in the current block. During prediction, after the prediction of the previous pixel point in the horizontal direction is completed, the predicted value of that previous pixel point is used as the input pixel value of the interpolation filter for the next pixel point and used for the prediction of the next pixel point. As shown in Figure 22, if the shape of the interpolation filter for the current block is 4x4, the decoding side performs interpolation prediction one by one for each position in the current block using an interpolation filter with known filter coefficients. Specifically, for the r-th point in the current block, first, the pixel values ​​of the N positions corresponding to the r-th point are determined based on the shape of the interpolation filter for the current block. For example, as shown in Figure 22, in a 4x4 interpolation filter, the dark positions are the positions of the r-th point to be processed, and the 15 light positions are the N positions corresponding to the r-th point. Here, the block to be predicted in Figure 22 is the current block. Next, the pixel values ​​of the N positions corresponding to the r-th point are determined. For example, for any of the N positions, if the position lies in the reconstruction region surrounding the current block, the reconstructed value of the position is determined as the pixel value of that position. If the position lies within the current block, the predicted value of the position is determined as the pixel value of that position.

[0282] For example, as shown in Figure 22B, the interpolation filter performs one interpolation filter prediction for each pixel point in the current block along the vertical direction. During prediction, after the prediction of the previous pixel point in the vertical direction is completed, the predicted value of the previous pixel point is used as the input pixel value for the interpolation filter of the next pixel point and used for the prediction of the next pixel point. In other words, in the conventional technique, when performing interpolation filter prediction for the current block using an interpolation filter, only one pixel prediction can be performed at a time, so the prediction takes time, the prediction efficiency decreases, and it affects the decoding efficiency.

[0283] To solve the above technical problems, the embodiment of this application predicts at least two points in the current block in parallel when performing interpolation filter prediction for the current block using an interpolation filter, that is, the decoding side can perform interpolation filter prediction for at least two pixel points in the current block using an interpolation filter at the same time.

[0284] In the embodiments of this application, the decoding side performs interpolation filter predictions for at least two pixel points in the current block using an interpolation filter at the same time, which includes at least two embodiments.

[0285] In the first implementation method, for at least two pixel points in the current block, these at least two pixel points are adjacent pixel points. The decoding side first determines the input information for the interpolation filter corresponding to these at least two pixel points, inputs this input information into the interpolation filter to perform interpolation filter prediction, obtains one predicted value, and determines the predicted values ​​for these at least two pixel points based on this predicted value. For example, the predicted values ​​are processed based on the associated feature information of the at least two pixel points, and predicted values ​​corresponding to each of these at least two pixel points are obtained. Alternatively, for example, these predicted values ​​are determined as the predicted values ​​corresponding to each of these at least two pixel points.

[0286] In this first implementation method, the specific method by which the decoding side determines the input information for the interpolation filter corresponding to these at least two pixel points is not limited.

[0287] For example, based on the shape of the interpolation filter, the same interpolation filter input value is determined for at least two of these pixel points, and this same interpolation filter input value is used as the input information for the interpolation filter. For example, these at least two pixel points include pixel point 1 and pixel point 2, and based on the shape of the interpolation filter, N input values ​​corresponding to pixel point 1 and N input values ​​corresponding to pixel point 2 are determined. The same input value is determined from the N input values ​​corresponding to pixel point 1 and the N input values ​​corresponding to pixel point 2, and this same input value is used as the input value for the interpolation filter. Note that if any of the N input values ​​corresponding to pixel point 2 or pixel point 1 contain undecoded values, those undecoded values ​​are discarded.

[0288] For example, based on the shape of the interpolation filter, the input value corresponding to the pixel point with the most decoded input values ​​among these at least two pixel points is determined as the input information for the interpolation filter. For example, these at least two pixel points include pixel point 1 and pixel point 2, and based on the shape of the interpolation filter, N input values ​​corresponding to pixel point 1 and N input values ​​corresponding to pixel point 2 are determined. Here, all N input values ​​corresponding to pixel point 1 have been decoded, and the N input values ​​corresponding to pixel point 2 include undecoded input values, so the N input values ​​corresponding to pixel point 1 are used as the input information for the interpolation filter.

[0289] Method 1 described above explains the specific process in which the decoding side inputs the same input value to the interpolation filter at once, and then uses the interpolation filter to predict the values ​​of at least two pixel points in the current block.

[0290] In the second method, at the same time, an interpolation filter is applied to each of at least two pixels in the current block. For example, at time t, the decoding side uses an interpolation filter to perform an interpolation filter prediction on pixel point 1 in the current block and obtains the predicted value of pixel point 1, while simultaneously using an interpolation filter to perform an interpolation filter prediction on pixel point 2 in the current block and obtains the predicted value of pixel point 2.

[0291] The embodiments of this application do not limit the prediction direction when the decoding side uses an interpolation filter to predict at least two pixel points in the current block in parallel.

[0292] In some embodiments, the decoding side can use an interpolation filter to predict at least two pixels in the current block in parallel along the horizontal direction. For example, when predicting at least two pixel points in parallel, if one or more input values ​​of the interpolation filter corresponding to the pixel points have not been decoded, the undecoded input values ​​are discarded and the decoded input values ​​are used as input values ​​for the interpolation filter.

[0293] In some embodiments, the decoding side can use an interpolation filter to predict at least two pixels in the current block in parallel along the vertical direction. For example, when predicting at least two pixel points in parallel, if one or more input values ​​of the interpolation filter corresponding to the pixel points have not been decoded, the undecoded input values ​​are discarded and the decoded input values ​​are used as input values ​​for the interpolation filter.

[0294] In some embodiments, the decoding side can use an interpolation filter to predict at least two pixels in the current block in parallel along the diagonal. Based on this, as an example, step S102 above includes the following step S102-A.

[0295] S102-A, based on the filter coefficients, parallel interpolation filter prediction is performed on pixel points on the same diagonal of the current block using an interpolation filter along the diagonal direction to obtain a predicted block for the current block.

[0296] As an example, as shown in Figure 23, when the decoding side predicts a pixel point in the current block using an interpolation filter, the pixel point to be predicted is located in one corner of the selection area of ​​the interpolation filter (for example, the lower right corner or the upper left corner). In this way, for any pixel point on the same diagonal (i.e., the pixel point indicated by the dark block in Figure 23), the N positions corresponding to the selected pixel point, based on the shape of the interpolation filter, do not include the positions of any other pixel points on that diagonal. That is, none of the N positions corresponding to each pixel point on the same diagonal include any pixel points on the diagonal. For example, taking two adjacent pixel points a and b on the same diagonal of the current block as an example, as shown in Figure 23, based on the shape of the interpolation filter, N positions corresponding to pixel point a and N positions corresponding to pixel point b are determined, and neither the N positions corresponding to pixel point a nor the N positions corresponding to the determined pixel point b include any pixel points on the diagonal. Based on this, when the decoding side performs interpolation filter prediction for the current block using the interpolation filter, it can perform parallel interpolation filter prediction along the diagonal direction for pixel points located on the same diagonal in the current block. For example, parallel interpolation filter prediction can be performed for pixel points a and b located on the same diagonal. Note that the shape of the interpolation filter shown in Figure 23 is just one example, and the shape of the interpolation filter in this embodiment is not limited to this.

[0297] The results of the investigation in this application show the left side, lower left, top, and left of the current block. above Since the upper right region has already been decoded, we can determine the starting point of the prediction along the diagonal direction of the current block based on the decoded region and the shape of the interpolation filter.

[0298] In one example, as shown in Figure 23, the decoding side performs interpolation filter prediction along the diagonal direction relative to the current block, starting from the upper left corner of the current block. Step S102-A above includes the following step S102-A1.

[0299] In S102-A1, the decoding side uses an interpolation filter to perform parallel interpolation filter predictions on pixel points on the same diagonal of the current block, starting from the upper left corner of the current block and moving diagonally, based on the filter coefficients, to obtain the predicted block of the current block. In this example, the pixel point to be predicted is located in the lower right corner of the interpolation filter selection region.

[0300] The embodiments of this application do not limit the specific orientation of the diagonal direction.

[0301] In some embodiments, as shown in Figure 23, when the decoding side starts from the upper left corner of the current block and predicts pixel points located on the same diagonal in the current block in parallel, the diagonal direction includes at least one of the directions from upper right to lower left and from lower left to upper right.

[0302] In one example, as shown in Figure 24A, the diagonal direction includes the direction from the upper right to the lower left, and in this case, as indicated by the arrows in Figure 24A, all of the diagonal directions of the current block lower left from upper right It is the direction towards.

[0303] In one example, as shown in Figure 24B, the diagonal direction includes the direction from the bottom left to the top right, and in this case, as indicated by the arrow in Figure 24B, the diagonal direction of the current block is upper right from lower left It is the direction towards.

[0304] In one example, as shown in Figure 24C, the diagonal direction includes the direction from the upper right to the lower left and the direction from the lower left to the upper right. In this case, as indicated by the arrows in Figure 24C, the diagonal direction of the current block includes two types of directions: the direction from the upper right to the lower left and the direction from the lower left to the upper right.

[0305] In the embodiments of this application, the decoding side predicts pixel points located on the same diagonal in the current block in parallel, so the specific orientation of the diagonal direction does not limit the technical means of the embodiments of this application.

[0306] In the embodiments of this application, the decoding side makes predictions for each prediction, using a pixel point on one diagonal of the current block as the unit. The process by which the decoding side predicts each pixel point on each diagonal of the current block in parallel is similar, and for the sake of explanation, the k-th diagonal of the current block is used as an example. In this case, S102-A1 above includes the following steps S102-A11 and S102-A12. In S102-A11, for M pixels on the k-th diagonal of the current block, the predicted values ​​of the M pixels are determined in parallel using an interpolation filter based on the filter coefficients, where both k and M are positive integers. S102-A12 obtains the predicted value of the current block based on the predicted values ​​of the pixel points on each diagonal in the current block.

[0307] This k-th diagonal may be understood as one of the diagonals of the current block shown in Figure 23, and there are M pixel points on this k-th diagonal. The decoding side determines the predicted values ​​of these M pixel points in parallel using an interpolation filter based on the filter coefficients. That is, the decoding side can simultaneously determine the predicted values ​​of the M pixel points located on the k-th diagonal, significantly increasing the prediction speed.

[0308] To illustrate with an example, as shown in Figure 24A, if three pixel points are located on the k-th diagonal of the current block, the decoding side determines the predicted values ​​of these three pixel points in parallel. For example, let's call these three pixel points pixel 1, pixel 2, and pixel 3. At the same time, the decoding side performs interpolation filter prediction on pixel 1 using an interpolation filter based on the filter coefficients to obtain the predicted value of pixel 1, simultaneously performing interpolation filter prediction on pixel 2 using an interpolation filter based on the filter coefficients to obtain the predicted value of pixel 2, and simultaneously performing interpolation filter prediction on pixel 3 using an interpolation filter based on the filter coefficients to obtain the predicted value of pixel 3. In this way, the decoding side determines the predicted values ​​of three pixel points located on the k-th diagonal of the current block in parallel in a single interpolation filter prediction process, significantly improving the speed of interpolation filter prediction. The decoding side refers to a method for determining the predicted value of the k-th diagonal pixel point, which allows it to determine the predicted values ​​of other diagonal pixel points in the current block, and further obtains the predicted block for the current block, thereby improving the prediction speed of the current block and improving the decoding efficiency.

[0309] The embodiments of this application do not limit the specific method by which the decoding side determines the predicted values ​​of M pixel points in parallel using an interpolation filter based on filter coefficients.

[0310] In some embodiments, these M pixel points are points on the k-th diagonal of the current block, so these M pixel points may be understood as adjacent pixel points, and their features are relatively similar. To reduce the complexity of the calculation, the input values ​​of the interpolation filter corresponding to one or more of these M pixel points are determined based on the shape of the interpolation filter. Then, based on the input values ​​of the interpolation filter corresponding to this one or more pixel points, calculations such as averaging and weighting are performed to determine the input values ​​of the interpolation filter corresponding to the other pixel points among the M pixel points. Finally, the predicted values ​​of the M pixel points are determined in parallel based on the filter coefficients and the input values ​​of the interpolation filter corresponding to each of the M pixel points.

[0311] In some embodiments, the step of determining the predicted values ​​of M pixel points in parallel using an interpolation filter based on the filter coefficients in S102-A11 above includes the following steps. S102-A11-a1, based on the shape of the interpolation filter, determines the pixel values ​​of N positions corresponding to each of the M pixel points in parallel. S102-A11-a2: Based on the filter coefficients and the pixel values ​​of N positions corresponding to each of the M pixel points, the predicted values ​​for the M pixel points are determined in parallel.

[0312] In this embodiment, when the decoding side determines the predicted values ​​of M pixel points on the k-th diagonal of the current block in parallel, it determines the pixel values ​​of N positions corresponding to each of these M pixel points in parallel based on the shape of the interpolation filter, and the pixel values ​​of N positions corresponding to each pixel point can be understood as the input values ​​of the interpolation filter corresponding to that pixel point. Next, the decoding side determines the predicted values ​​of the M pixel points in parallel based on the filter coefficients and the pixel values ​​of N positions corresponding to each of the M pixel points.

[0313] To illustrate with an example, see Figure 24. A or Figure 24B or Figure 24CAs shown, the current block contains three pixel points on the k-th diagonal, and these three pixel points are designated as pixel point 1, pixel point 2, and pixel point 3. Assume that the interpolation filter for the current block is a 4x4 interpolation filter. When the decoding side determines the predicted values ​​of these three pixel points in parallel, it determines 15 pixel values ​​corresponding to pixel point 1 based on the shape of the interpolation filter, uses these 15 pixel values ​​as input to the interpolation filter, and determines the predicted value of pixel point 1 based on the filter coefficients determined above. At the same time, the decoding side determines 15 pixel values ​​corresponding to pixel point 2 based on the shape of the interpolation filter, uses these 15 pixel values ​​as input to the interpolation filter, and determines the predicted value of pixel point 2 based on the filter coefficients determined above. At the same time, the decoding side determines 15 pixel values ​​corresponding to pixel point 3 based on the shape of the interpolation filter, uses these 15 pixel values ​​as input to the interpolation filter, and determines the predicted value of pixel point 3 based on the filter coefficients determined above. In other words, in this embodiment, the decoding side determines the predicted values ​​for three pixels of the current block in parallel at the same time, thereby significantly improving the prediction speed and decoding efficiency.

[0314] Next, we will explain the specific process for determining the predicted values ​​of the M pixel points in parallel in step S102-A11-a2 above, based on the filter coefficients and the pixel values ​​of the N positions corresponding to each of the M pixel points.

[0315] In some embodiments, for each of the M pixels, the decoding side directly multiplies the pixel values ​​at the N positions corresponding to that pixel point by a filter coefficient in order to obtain the predicted value of that pixel.

[0316] For example, the decoding side obtains the predicted value for each of the M pixel points based on the following equation (5).

number

[0317] Based on equation (5), the decoding side can determine the predicted values ​​of each pixel point located on the same diagonal in the current block in parallel.

[0318] In some embodiments, the decoding side determines the filter coefficients of the interpolation filter based on equation (4) above. Since equation (4) determines the filter coefficients using the reference region after the mean has been removed, the influence of the pixel mean reconstruction value m must be considered in order to determine the predicted value of the current block based on these filter coefficients.

[0319] In one possible implementation of this embodiment, the interpolation filter coefficients determined in equation (4) are substituted into equation (5) to obtain the predicted values ​​for each point in the current block. Then, the pixel-average reconstruction value m is added to the predicted values ​​for each point to obtain the final predicted values ​​for each point in the current block, and consequently, the predicted block for the current block.

[0320] In another possible implementation of this embodiment, step S102-A11-a2 includes the following steps: S102-A11-a21, based on the pixel average reconstruction value, average value removal is performed in parallel on the pixel values ​​at N positions corresponding to each of the M pixel points, and the average value removed pixel values ​​at N positions corresponding to each of the M pixel points are obtained. S102-A11-a22, based on the filter coefficients and the average value of the N positions corresponding to each of the M pixel points, the predicted values ​​for M pixel points are determined in parallel.

[0321] Since the above filter coefficients are determined based on the reference region after mean removal, the decoding side performs mean removal on the pixel values ​​at N positions corresponding to each of the M pixel points on the k-th diagonal of the current block, based on the pixel mean reconstruction value, and obtains the mean-removed pixel values ​​at N positions corresponding to each of the M pixel points. For example, for any of the M pixel points, the pixel mean reconstruction value is subtracted from the pixel values ​​at N positions of that pixel point to obtain the mean-removed pixel values ​​at N positions of that pixel point.

[0322] Next, the predicted values ​​for the M pixel points are determined in parallel based on the filter coefficients and the average removed pixel values ​​of the N positions corresponding to each of the M pixel points.

[0323] The embodiments of this application do not limit the specific method for determining the predicted values ​​of M pixel points in parallel based on filter coefficients and the average value of N positions corresponding to each of the M pixel points after removing the filter coefficients.

[0324] JPEG2024216632000033.jpg37168

[0325] In an alternative implementation, step S102-A11-a22 includes the following steps. S102-A11-a221 determines a second reconstruction region around the current block, and determines the maximum and minimum reconstruction values ​​for the second reconstruction region. S102-A11-a222, based on the average-removed pixel values, filter coefficients, and pixel-average reconstruction values ​​of N positions corresponding to each of the M pixel points, a first predicted value for M pixel points is obtained in parallel. S102-A11-a223, based on the first predicted value, maximum reconstructed value, and minimum reconstructed value of M pixel points, the predicted values ​​of M pixel points are determined in parallel.

[0326] JPEG2024216632000034.jpg21168

[0327] The embodiments of this application do not limit the specific method for determining the second reconfiguration region around the current block.

[0328] In one example, the second reconfiguration region of the current block coincides with the reference region of the current block.

[0329] In one example, the second reconfiguration region of the current block coincides with the first reconfiguration region of the current block.

[0330] In one example, the reconstruction areas above, to the left, to the right, to the left, and to the left of the current block are determined as the second reconstruction area. For example, the reconstruction areas of the top 13 rows, the left 13 columns, the top 13 rows, the top 13 rows and 13 columns of the top left block, and the bottom 13 columns of the bottom left block are determined as the second reconstruction area.

[0331] Furthermore, there is no specific order in which steps S102-A11-a221 and S102-A11-a222 are performed in the actual implementation process. For example, step S102-A11-a221 may be performed before step S102-A11-a222, after step S102-A11-a222, or in sync with step S102-A11-a222.

[0332] The embodiments of this application do not limit the specific method by which the decoding side obtains first predicted values ​​for M pixel points in parallel based on the average value of the N positions corresponding to each of the M pixel points, the filter coefficient, and the pixel average reconstruction value.

[0333] For example, for any of the M pixel points, a second predicted value for the pixel point is obtained by multiplying the average value of the N positions of the pixel point (after removing the filter coefficient) by the pixel point, and a first predicted value for the pixel point is obtained by adding the second predicted value to the pixel average reconstruction value.

[0334] For example, the decoding side obtains the first predicted value for this pixel point based on the following equation (6).

number

[0335] JPEG2024216632000036.jpg25168

[0336] For example, the decoding side obtains a predicted value for the pixel point based on equation (6) above, then performs a prediction process on the predicted value to obtain a first predicted value for the pixel point.

[0337] Based on the steps described above, the decoding side determines a first predicted value for M pixel points, and then, based on the first predicted value, the maximum reconstructed value, and the minimum reconstructed value, determines the predicted values ​​for M pixel points in parallel.

[0338] For example, if, for any of the M pixel points, the first predicted value for that pixel point is greater than the minimum reconstruction value and less than the maximum reconstruction value, then the first predicted value is determined to be the predicted value for that pixel point.

[0339] For example, if the first predicted value of the pixel point is less than or equal to the minimum reconstruction value, the minimum reconstruction value is determined as the predicted value of the pixel point.

[0340] For example, if the first predicted value of the pixel point is greater than or equal to the maximum reconstruction value, the maximum reconstruction value is determined as the predicted value of the pixel point.

[0341] In one example, the decoding side determines the predicted value of the pixel point using the following equation (7).

number

[0342] The above is an example of determining the predicted values ​​of M pixel points on the k-th diagonal in the current block. The decoding side can then refer to the above method to determine the predicted values ​​of each pixel point on each diagonal in the current block in parallel, and further obtain the predicted values ​​of each point in the current block to construct a predicted block for the current block.

[0343] Based on the steps above, the decoding side performs interpolation filter prediction for the current block, obtains the predicted block for the current block, and then performs the following steps.

[0344] S103, determine the transformation kernel corresponding to the current block, and determine the reconstructed block of the current block based on the transformation kernel corresponding to the current block and the predicted block.

[0345] Therefore, when decoding the current block, the decoding side decodes the code stream to obtain the quantization coefficients of the current block, then dequantizes the quantization coefficients to obtain the transformation coefficients of the current block, and then dequantizes the transformation coefficients of the current block to obtain the residual block (or residual value) of the current block. At the same time, the prediction mode of the current block is determined, the current block is predicted using this prediction mode to obtain the predicted block of the current block, and the predicted block and the residual block are added together to obtain the reconstructed block of the current block.

[0346] When inversely transforming the transformation coefficients of the current block, it is necessary to determine the transformation kernel, inversely transform the transformation coefficients of the current block based on the transformation kernel, and obtain the residual value of the current block.

[0347] The embodiments of this application do not limit the specific method by which the decoding side determines the transformation kernel corresponding to the current block.

[0348] In some embodiments, the encoding and decoding sides use a default transformation kernel as the transformation kernel for the current block.

[0349] In some embodiments, the encoding side determines the transformation kernel for the current block and then writes the instruction information for the transformation kernel to the code stream. In this way, the decoding side determines the transformation kernel for the current block by decoding the code stream.

[0350] In some embodiments, the decryption side performs the following steps S103-A and S103- B This determines the translation kernel for the current block. S103-A, the intra-prediction mode corresponding to the prediction block is determined. S103-B, based on the intra-prediction mode corresponding to the predicted block, the transformation kernel corresponding to the current block is determined.

[0351] In the embodiments of this application, the predicted block of the current block is determined using an interpolation filter prediction mode, then the conventional intra-prediction mode corresponding to the predicted block is determined, and further, the transformation kernel corresponding to the current block is determined based on the conventional intra-prediction mode.

[0352] Next, we will describe the specific process by which the decoding side determines the intra-prediction mode corresponding to the predicted block.

[0353] As an example, as shown in Figure 7, the conventional intra-prediction modes included in the current VVC are as follows: PLANAR mode: Intra prediction mode index is 0. DC mode: Intra prediction mode index is 1. Angle mode: Intra prediction mode index is 2-66.

[0354] In one example, as shown in Figure 25, the direction of the arrows in the figure indicates the direction of the angle pattern prediction present in the VVC, and the prediction pattern indices used during decoding are 2 to 66. If the current block is a non-square block, some angle directions are replaced with wider angles such as -1 to -14 and 67 to 80 in Figure 25.

[0355] In some embodiments, the intra-prediction mode corresponding to the prediction block is the default intra-prediction mode. That is, if the current block makes a prediction using the interpolation filter prediction mode and obtains a prediction block, one of the conventional intra-prediction modes is determined by default as the intra-prediction mode corresponding to that prediction block.

[0356] In some embodiments, the decoding side determines the intra-prediction mode corresponding to the prediction block by following steps S103-A1 and S103-A2. S103-A1, determine the angle values ​​of R points in the prediction block, where R is a positive integer. S103-A2 determines the intra-prediction mode corresponding to the prediction block based on the angle values ​​of R points.

[0357] In the embodiments of this application, the intra-prediction mode corresponding to a prediction block is determined by aggregating the intra-prediction modes corresponding to the angle values ​​of R points in the prediction block.

[0358] The embodiments of this application do not limit the specific locations and number of R points in the prediction block used to determine the angle value. For example, the R points may be one point in the prediction block or multiple points in the prediction block.

[0359] For example, if the above R points are one point, the decoding side determines the angle value of one point in the prediction block (for example, the center point of the prediction block), determines the intra-prediction mode corresponding to that point based on the angle value of that point, and further determines that intra-prediction mode as the intra-prediction mode corresponding to the prediction block.

[0360] For example, if the R points mentioned above are multiple points, the decoding side determines the angle values ​​of these multiple points, determines the intra-prediction mode corresponding to each of these multiple points based on the angle values ​​of these multiple points, and further determines the intra-prediction mode with the most identical intra-prediction modes among these multiple points as the intra-prediction mode corresponding to the prediction block.

[0361] In some embodiments, when the angular values ​​of R points in a prediction block are determined by a sliding window, the selection of these R points is related to the shape and size of the sliding window. For example, each of the R points is the center point within the sliding window as it slides through the prediction block.

[0362] In the embodiments of this application, the method for determining the angle value of each of the R points is the same, and for the sake of explanation, an example of determining the angle value of the i-th point among the R points will be given.

[0363] The embodiments of this application do not limit the specific method for determining the angle value of a point.

[0364] In some embodiments, step S103-A1 includes steps S103-A11 and S103-A12. S103-A11, for the i-th point out of R points, determine the horizontal and vertical slopes of the i-th point, where i is a positive integer less than or equal to R. S103-A12, the angle value of the i-th point is determined based on the horizontal and vertical slopes of the i-th point.

[0365] In this embodiment, the decoding side first determines the horizontal and vertical slopes of each of the R points, for example, the i-th point, and then determines the angle value of the i-th point based on the horizontal and vertical slopes.

[0366] The embodiments of this application do not limit the specific method for determining the horizontal and vertical slopes of the i-th point.

[0367] In one example, the horizontal gradient value of the i-th point is determined based on the horizontal change between the predicted values ​​of the points surrounding the i-th point in the prediction block and the predicted value of the i-th point itself, and the vertical gradient value of the i-th point is determined based on the vertical change between the predicted values ​​of the points surrounding the i-th point in the prediction block and the predicted value of the i-th point itself.

[0368] In another example, the decryption side determines the predicted values ​​of points within a sliding window centered on the i-th point in the prediction block, and obtains the horizontal and vertical gradients of the i-th point based on the predicted values ​​of points within the sliding window, the horizontal gradient operator, and the vertical gradient operator.

[0369] In this example, first, a sliding window is determined, for example, a sliding window with a size of 3x3 as shown in Figure 26. This sliding window is then slid within the prediction block, and the horizontal and vertical slopes of the center point of this sliding window are determined each time it is slid. If the current center point of the sliding window is the i-th point, then the predicted values ​​for each point within the current sliding window are obtained, for example, 3x3 = 9 predicted values ​​can be obtained. Then, based on these 9 predicted values ​​and the pre-set horizontal and vertical slope operators, the horizontal and vertical slopes of the i-th point are determined.

[0370] JPEG2024216632000039.jpg23168

[0371] JPEG2024216632000040.jpg23168

[0372] The embodiments of this application do not limit the specific values ​​of the horizontal gradient operator and the vertical gradient operator.

[0373] JPEG2024216632000041.jpg31168

[0374] Based on the steps described above, the decoding side can determine the horizontal and vertical slopes of the i-th point, and then determine the angle value of the i-th point based on the horizontal and vertical slopes of the i-th point.

[0375] For example, the arctangent value of the ratio of the vertical slope to the horizontal slope at the i-th point is determined as the angle value of the i-th point. This is shown in equation (8) as an example.

number

[0376] The decoding side may determine the angle value of the i-th point by a method other than using equation (8) above. For example, the decoding side may adjust the angle value determined by equation (8) above to obtain the angle value of the i-th point.

[0377] On the decoding side, the angle value of each of the R points is determined using the method described above, and then step S103-A2 is performed to determine the intra prediction mode corresponding to the prediction block based on the angle values ​​of the R points.

[0378] The embodiments of this application do not limit the specific method for determining the intra-prediction mode corresponding to a prediction block based on the angular values ​​of R points.

[0379] In some embodiments, the decoding side selects angle value 1 with the highest number of recurrences from the angle values ​​of R points, matches this angle value 1 with the predicted angle of a conventional intra-prediction mode to obtain the intra-prediction mode corresponding to angle value 1, and determines the intra-prediction mode corresponding to angle value 1 as the intra-prediction mode corresponding to the prediction block.

[0380] In some embodiments, step S103-A2 above includes the following steps S103-A21 and S103-A22. S103-A21 determines the intra-prediction mode corresponding to R points based on the angle values ​​of R points. S103-A22 determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction modes corresponding to R points.

[0381] In this implementation method, the decoding side determines the intra-prediction mode corresponding to each of the R points based on the angle value of each point. For example, for each of the R points, the angle value of that point is matched with the predicted angle of a conventional intra-prediction mode, and the intra-prediction mode corresponding to the angle value of that point is obtained. In this way, the intra-prediction mode corresponding to each of the R points can be obtained.

[0382] Then, based on the intra-prediction mode corresponding to each of these R points, the intra-prediction mode corresponding to the prediction block is determined.

[0383] In one possible implementation, the intra-prediction mode with the highest recurrence rate among the intra-prediction modes corresponding to each of the R points is determined as the intra-prediction mode corresponding to the prediction block.

[0384] In another possible implementation, steps S103-A22 above include the following steps: S103-A221 determines the gradient amplitude value corresponding to R points based on the horizontal and vertical gradients of R points. S103-A222 determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction mode and gradient amplitude values ​​corresponding to R points.

[0385] In this implementation method, the decoding side determines the gradient amplitude value corresponding to each of the R points based on the horizontal and vertical gradients of each of the R points determined above.

[0386] In the embodiments of this application, the specific method by which the decoding side determines the gradient amplitude value corresponding to each of the R points is the same. For the sake of explanation, we will take the example of determining the gradient amplitude value corresponding to the i-th point among the R points.

[0387] The embodiments of this application do not limit the specific method by which the decoding side determines the gradient amplitude value corresponding to the i-th point based on the horizontal and vertical gradients of the i-th point.

[0388] For example, the decoding side multiplies the horizontal gradient and vertical gradient of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.

[0389] For example, the decoding side adds the absolute value of the horizontal gradient and the absolute value of the vertical gradient of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.

[0390] As an example, the decoding side determines the gradient amplitude value corresponding to the i-th point based on the following equation (9).

number

[0391] Based on the steps described above, the decoding side can determine the gradient amplitude value corresponding to each of the R points. Next, the decoding side performs steps S103-A222 above and determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction modes and gradient amplitude values ​​corresponding to the R points.

[0392] In one example, the intra-prediction mode corresponding to the point with the maximum gradient amplitude among the R points is determined as the intra-prediction mode corresponding to the prediction block.

[0393] In another example, for any of the R points, the gradient amplitude value corresponding to that point is accumulated in the corresponding intra-prediction mode, and the cumulative gradient amplitude value of the intra-prediction modes corresponding to the R points is obtained. Among the intra-prediction modes corresponding to the R points, the intra-prediction mode with the largest cumulative gradient amplitude value is determined as the intra-prediction mode corresponding to the prediction block.

[0394] For example, as shown in Figure 27, the gradient amplitude values ​​corresponding to each of the R points are accumulated in the corresponding intra-prediction mode. For instance, if the intra-prediction modes corresponding to points 1 and 2 among the R points are both intra-prediction mode 1, the gradient amplitude values ​​corresponding to points 1 and 2 are accumulated in the gradient amplitude value corresponding to intra-prediction mode 1. The gradient amplitude value histogram shown in Figure 27 can be obtained in the same manner. In this way, the intra-prediction mode with the maximum accumulated gradient amplitude value in this gradient amplitude value histogram can be determined as the intra-prediction mode corresponding to the prediction block. For example, in Figure 27, the intra-prediction mode corresponding to the dark-colored accumulated gradient amplitude value is determined as the intra-prediction mode corresponding to the prediction block.

[0395] In some embodiments, if the gradient amplitude values ​​corresponding to R points are all 0, the first intra-prediction mode is determined as the intra-prediction mode corresponding to the prediction block. That is, if the gradient amplitude values ​​corresponding to all of the R points are 0, it means that the horizontal and vertical gradients of each of the R points are both 0, and in this case, the pre-set first intra-prediction mode can be determined as the intra-prediction mode corresponding to the prediction block.

[0396] The embodiments of this application do not limit the type of the first intra-prediction mode described above.

[0397] As an example, the first intra prediction mode described above is the PLANAR mode.

[0398] Based on the steps described above, the decoding side determines the intra-prediction mode corresponding to the prediction block, and then determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block.

[0399] The embodiments of this application do not limit the specific method by which the decoding side determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block.

[0400] In some embodiments, the decoding side searches for an image block with the same intra-prediction mode as the prediction block from the decoded image blocks surrounding the prediction block, based on the intra-prediction mode corresponding to the prediction block, and then determines the transformation kernel corresponding to that image block as the transformation kernel corresponding to the current block.

[0401] In some embodiments, the step in step S103-B above, which determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block, includes the following steps: S103-B1 obtains the correspondence between the intra prediction mode and the transformation kernel group, and one transformation kernel group contains at least one type of transformation kernel. S103-B2, in the correspondence, the first transformation kernel group corresponding to the intra-prediction mode of the prediction block is searched. S103-B3: Determine the translation kernel corresponding to the current block from the first translation kernel group.

[0402] In the embodiments of this application, there is a correspondence between the intra-prediction mode and the conversion kernel group. Based on this, the decoding side determines the intra-prediction mode corresponding to the prediction block, and then obtains the pre-configured correspondence between the intra-prediction mode and the conversion kernel group.

[0403] As an example, Table 17 shows the correspondence between the intra prediction mode and the conversion kernel group. [Table 17]

[0404] Table 17 above represents only one type of correspondence between the intra-prediction mode and the conversion kernel group in the embodiment of this application. The correspondence between the intra-prediction mode and the conversion kernel group in the embodiment of this application includes, but is not limited to, those shown in Table 17.

[0405] Each translation kernel group includes at least one type of translation kernel.

[0406] The decoding side obtains the correspondence between intra-prediction modes and transformation kernel groups as shown in Table 17, and then, based on the intra-prediction mode corresponding to the prediction block, searches for the transformation kernel group corresponding to the intra-prediction mode corresponding to the prediction block in the correspondence between intra-prediction modes and transformation kernel groups, and sets this transformation kernel group as the first transformation kernel group. For example, if the intra-prediction mode corresponding to the prediction block is an angle prediction mode in 64 angular directions, querying Table 17 reveals that the transformation kernel group corresponding to this angle prediction mode in 64 angular directions is 4. In this way, the decoding side determines the transformation kernel corresponding to the current block from at least one type of transformation kernel included in transformation kernel group 4.

[0407] For example, if the first group of translation kernels contains one translation kernel, that translation kernel is determined to be the translation kernel corresponding to the current block.

[0408] For example, if the first group of transformation kernels includes multiple types of transformation kernels, the decryption side determines the type of transformation kernel corresponding to the current block, and further determines the transformation kernel of that type in the first group of transformation kernels as the transformation kernel corresponding to the current block.

[0409] Here, the following methods are possible, but are not limited to, for determining the type of transformation kernel that corresponds to the current block on the decoding side.

[0410] In one example, the type of transformation kernel corresponding to the current block is the default type. In this way, the decryption side determines the default type as the type of transformation kernel corresponding to the current block.

[0411] In another example, the encoder writes the transformation kernel type corresponding to the current block into the code stream. The decoder then decodes the code stream to obtain the transformation kernel type corresponding to the current block.

[0412] Based on the above, in the embodiment of this application, the decoding side determines the predicted block of the current block using an interpolation filter prediction mode, further determines the conventional intra-prediction mode corresponding to the predicted block, and determines the transformation kernel corresponding to the current block based on the conventional intra-prediction mode corresponding to the predicted block. That is, in the embodiment of this application, the conventional intra-prediction mode derived based on the interpolation filter prediction is used to select the NSPT (Non-separable primary transform) and LFNST (Low Frequency non-separable secondary transform) transformation kernel groups, and the determined transformation kernel is matched to the characteristics of the current block to improve the accuracy of transformation kernel determination. When the reconstruction value of the current block is determined using this accurately determined transformation kernel, the accuracy of reconstruction value determination is improved, and the decoding accuracy of the current block can be improved. In addition, in the embodiment of this application, when determining the transformation kernel of the current block using the conventional prediction mode corresponding to the predicted block, it is not necessary to individually specify the transformation kernel, saving codewords and further improving the video codec effect.

[0413] Based on the steps above, the decoding side determines the transformation kernel corresponding to the current block, then inversely transforms the transformation coefficients of the current block based on the transformation kernel corresponding to the current block to obtain the residual block of the current block, and obtains the reconstructed block of the current block based on the predicted block and residual block of the current block.

[0414] In the embodiments of this application, the decoding side determines the predicted block of the current block and the transformation kernel corresponding to the current block based on the steps described above. In this way, the decoding side decodes the code stream to obtain the quantization coefficients of the current block, then dequantizes the quantization coefficients to obtain the transformation coefficients of the current block, and uses the transformation kernel corresponding to the current block determined above to detransform the transformation coefficients of the current block inversely to obtain the residual block (or residual value) of the current block. Finally, the decoding side adds the predicted block and the residual block of the current block to obtain the reconstructed block of the current block.

[0415] In some embodiments, the current block described above is either a luminance block or a chromaticity block, that is, in the implementations of this application, predictions can be made for both luminance blocks and chromaticity blocks using the interpolation filter prediction mode according to the implementations of this application.

[0416] In some embodiments, if the current block is a luminance block, the prediction mode for this current block is an interpolation filter prediction mode, and the chromaticity block corresponding to the current block employs a direct derivation mode DM, then the PLANAR mode or the intra-prediction mode corresponding to the prediction block is determined as the prediction mode for the chromaticity block.

[0417] The video decoding method according to the embodiment of this application first determines the reference region and interpolation filter of the current block when predicting the current block, determines the filter coefficients based on the reference region, and performs parallel prediction on at least two pixel points in the current block using the interpolation filter based on the filter coefficients to obtain a predicted block of the current block. A transformation kernel corresponding to the current block is determined, and the reconstruction value of the current block is determined based on the transformation kernel and the predicted block. In other words, in the embodiment of this application, when performing interpolation filter prediction on the current block using the interpolation filter, parallel prediction is performed on at least two points in the current block to improve the prediction speed and further improve the decoding efficiency.

[0418] The prediction method of this application has been explained using the decoding side as an example. The following explanation will use the encoding side as an example.

[0419] Figure 28 is a schematic flowchart of a prediction method according to one embodiment of the present application, applied to the video encoder shown in Figures 1 and 2. As shown in Figure 28, the method of the embodiment of the present application includes the following steps.

[0420] S201 determines the reference region and interpolation filter of the current block, and determines the filter coefficients of the interpolation filter based on the reference region.

[0421] The encoding process, when encoding the current block, first determines the prediction mode for the current block, uses this prediction mode to predict the current block, and obtains the predicted block (or predicted value) of the current block. The predicted block of the current block is subtracted from the current block to obtain the residual block (or residual value) of the current block. Next, the residual block of the current block is transformed to obtain transformation coefficients, the transformation coefficients are quantized to obtain quantization coefficients, and the quantization coefficients are encoded to obtain the code stream.

[0422] In the embodiments of this application, the encoding side first determines the prediction mode of the current block.

[0423] In some embodiments, the method by which the encoding side determines the prediction mode of the current block includes at least the following:

[0424] In Method 1, the encoding side determines the candidate prediction mode with the minimum cost from among several candidate prediction modes, which consist of the conventional prediction mode and the interpolation filter prediction mode shown in Figure 6 or Figure 7, as the prediction mode for the current block. Next, the encoding side adds instruction information for the prediction mode of the current block to the code stream. In this way, the decoding side obtains instruction information for the prediction mode of the current block by decoding the code stream, and further determines the prediction mode for the current block based on that instruction information.

[0425] In Method 2, the encoding side creates an intra-prediction mode candidate list and selects the intra-prediction mode for the current block from this list. This intra-prediction mode candidate list includes the interpolation filter prediction mode. Next, the encoding side writes the number (or index number) of the current block's intra-prediction mode in the intra-prediction mode candidate list to the code stream.

[0426] In method 3, the encoding side creates an intra-prediction mode candidate list that includes interpolation filter prediction modes, then selects an intra-prediction mode for the current block from the intra-prediction mode candidate list, determines the cost of each candidate prediction mode in the intra-prediction mode candidate list on the template of the current block, and then determines the intra-prediction mode for the current block based on the cost.

[0427] As is clear from the above methods, when the encoding side determines the prediction mode of the current block, it first determines several candidate prediction modes, including the interpolation filter prediction mode, and then determines the prediction mode of the current block from these multiple candidate prediction modes.

[0428] Here, the specific method for determining the prediction mode of the current block from multiple candidate prediction modes may be that the encoding side decides one of the multiple candidate prediction modes as the prediction mode of the current block. That is, the encoding side predicts the current block using multiple candidate prediction modes, determines the cost corresponding to each candidate prediction mode, and this cost may be RDO or SATD, and the candidate prediction mode with the minimum cost is determined as the prediction mode of the current block.

[0429] The encoding side determines the prediction mode of the current block based on the method described above, and if the determined prediction mode of the current block is the interpolation filter prediction mode, it executes the step in S201 above.

[0430] In some embodiments, the conditions for using the interpolation filter prediction mode are limited. Based on this, a decision is made whether the current image block is allowed to perform predictions using the interpolation filter prediction mode before determining the reference region and interpolation filter of the current block.

[0431] The embodiments of this application do not limit the specific method for determining whether the current image block is permitted to perform predictions using the interpolation filter prediction mode. In other words, there are no restrictions on the specific conditions under which the interpolation filter prediction mode can be used.

[0432] In some embodiments, the encoding side needs to determine whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement before determining the prediction mode of the current block from a plurality of candidate prediction modes.

[0433] The embodiments of this application are not limited to a predetermined location and block size, but are specifically determined according to the actual needs.

[0434] In one example, as shown in Figure 11, if the position of the upper left corner of the current image is (0,0) and the position of the upper left corner of the current block is (x,y), then the predetermined position requirements are that the x of the current block is greater than or equal to a first predetermined value XX, and the y of the current block is greater than or equal to a second predetermined value YY.

[0435] The embodiments of this application do not limit the specific values ​​of the first and second predetermined values ​​described above.

[0436] For example, the first predetermined value and the second predetermined value are the same.

[0437] For example, if both the first predetermined value and the second predetermined value are 13, that is, if the distance from the top edge of the current block to the top edge of the current image is 13 rows of pixels or more, and the distance from the left edge of the current block to the left edge of the current image is 13 columns of pixels or more, then the position of the current block in the current image satisfies the predetermined position requirement.

[0438] In one example, continuing to refer to Figure 11, if the current block width is W and the current block height is H, then the predetermined block size requirement is that the current block width W is less than or equal to the third predetermined value A, and the current block height H is less than or equal to the fourth predetermined value B.

[0439] The embodiments of this application do not limit the specific values ​​of the third and fourth predetermined values ​​described above.

[0440] For example, the third predetermined value and the fourth predetermined value are the same.

[0441] For example, if both the third and fourth predetermined values ​​are 32, that is, if the current block's width and height are both 32 or less, it means that the current block meets the predetermined block size requirement.

[0442] In the embodiments of this application, before determining whether the current block will perform prediction using an interpolation filter prediction mode, the encoding side first determines whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement. If the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the prediction mode for the current block is determined from the above-mentioned plurality of candidate prediction modes, including the interpolation filter prediction mode. For example, as shown in Figure 11, if the distance from the top edge of the current block to the top edge of the current image is 13 rows of pixels or more, the distance from the left edge of the current block to the left edge of the current image is 13 columns of pixels or more, and the width and height of the current block are both 32 or less, the prediction mode for the current block is determined from the above-mentioned plurality of candidate prediction modes, including the interpolation filter prediction mode.

[0443] In some embodiments, the first predetermined value, the second predetermined value, the third predetermined value, and the fourth predetermined value are default values.

[0444] In some embodiments, if the position of the current block in the current image does not satisfy a predetermined position requirement, and / or the size of the current block does not satisfy a predetermined block size requirement, the encoding side determines that the prediction mode of the current block is not an interpolation filter prediction mode, and as a result, the encoding side determines the prediction mode of the current block from candidate prediction modes that do not include the interpolation filter prediction mode.

[0445] In some embodiments, the encoding side further includes the steps of determining whether to allow the current sequence to perform predictions using an interpolation filter prediction mode before determining whether the position of the current block in the current image satisfies predetermined positional requirements and whether the size of the current block satisfies predetermined block size, and if the current sequence is allowed to perform predictions using an interpolation filter prediction mode, determining whether the position of the current block in the current image satisfies predetermined positional requirements and whether the size of the current block satisfies predetermined block size.

[0446] In embodiments of this application, a high-level syntax element indicates whether the current sequence is permitted to perform predictions using an interpolation filter prediction mode. If it indicates that the current sequence is permitted to perform predictions using an interpolation filter prediction mode, the encoding side determines whether the position of the current block in the current image satisfies a predetermined position requirement and whether the size of the current block satisfies a predetermined block size requirement. If it is determined that the position of the current block in the current image satisfies the predetermined position requirement and the size of the current block satisfies the predetermined block size requirement, the encoding side determines the prediction mode for the current block from candidate prediction modes that do not include an interpolation filter prediction mode.

[0447] In some embodiments, if the encoding side determines that the current sequence does not allow prediction using the interpolation filter prediction mode, the encoding side skips step S201 above.

[0448] In some embodiments, the encoding side writes second information to the code stream indicating whether the current sequence is allowed to be predicted using the interpolation filter prediction mode.

[0449] The embodiments of this application are not limited to specific representations of the second information, and may include any directive information that can indicate whether the current sequence is permitted to perform predictions using the interpolation filter prediction mode.

[0450] For example, the second piece of information can be represented as sps_eip_enabled_flag, which indicates whether the current sequence is allowed to perform predictions using interpolation filter prediction mode by assigning different values ​​to sps_eip_enabled_flag. For example, sps_eip_enabled_flag=0 indicates that the current sequence is not allowed to perform predictions using interpolation filter prediction mode, while sps_eip_enabled_flag=1 indicates that the current sequence is allowed to perform predictions using interpolation filter prediction mode.

[0451] For example, the second piece of information is conveyed by a sequence-level parameter set (SPS).

[0452] In some embodiments, embodiments of this application may also include general constraints information (GCI) identifier bits to indicate whether interpolation filter prediction techniques are used. For example, whether the current video enables interpolation filter prediction techniques is indicated by gci_no_eip_constraint_flag. For example, as shown in Table 2, this gci_no_eip_constraint_flag is carried by general constraints information general_constraints_info().

[0453] From the above, determining whether the current block uses interpolation filter prediction mode can be limited by higher-level syntax elements, such as GCI, sequence level, frame level, slice level, and block level. It can also be limited by the size and position of the current block.

[0454] In some embodiments, the computational cost and complexity increase when the interpolation filter prediction mode is used for several smaller blocks. This is because the interpolation filter prediction mode is computationally complex in this application, and using it for several smaller blocks increases the number of times the interpolation filter prediction mode is used throughout the image decoding process, further increasing the computational cost and complexity of the image. Based on this, in embodiments of this application, the use of the interpolation filter prediction mode is permitted only for slightly larger blocks. For example, the use of the interpolation filter prediction mode is permitted only if the size of the current block is greater than or equal to a predetermined size. If the size of the current block is smaller than the predetermined size, the current block is not permitted to use the interpolation filter prediction mode. Embodiments of this application do not limit the specific value of the predetermined size. For example, the current block being greater than or equal to a predetermined size may mean that the number of pixel points in the current block is greater than or equal to a predetermined number, or that at least one of the length and width of the current block is greater than or equal to a predetermined value, or that the ratio of the length and width of the current block is greater than or equal to a predetermined ratio, etc.

[0455] In some embodiments, if the current block is in the first row of the current CTU, it is determined that the current block does not allow the use of the interpolation filter prediction mode. That is, if the current block is to perform predictions using the interpolation filter prediction mode, the current block is not located in the first row of the current CTU.

[0456] In some embodiments, the decision of whether the current block uses interpolation filter prediction mode also depends on the type of the current image. For example, it is specified that prediction can be made using interpolation filter prediction mode for intra-prediction images, while prediction using interpolation filter prediction mode is not permitted for inter-prediction images. Based on this, if the current image in which the current block exists is an intra-prediction image, it is decided that the current block will allow prediction using interpolation filter prediction mode. If the current image is not an intra-prediction image (e.g., an inter-prediction image), it is decided that the current block will not allow prediction using interpolation filter prediction mode.

[0457] In some embodiments, several complex intra-prediction modes are introduced in the ECM reference software to improve codec performance, such as template-based intra-prediction derivation (TIMD), decoder-side intra-prediction derivation (DIMD), template-based multiple reference line intra-prediction (TMRL), spatial geometrical partitioning mode (SGPM), and convolutional cross-component model (CCCM). All of these complex intra-prediction modes are based on template matching techniques, and the interpolation filter prediction mode in the embodiments of this application also uses information from the reconstructed region (which may be understood as the template region) during use. Therefore, in the embodiments of this application, the interpolation filtering prediction mode can also be classified as an intra-prediction mode based on template matching techniques. Based on this, in the embodiments of this application, the intra-prediction modes based on the template matching technology are uniformly indicated using unified identification information (e.g., first information). For example, if the first information indicates that the template matching technology is not enabled, it means that none of the intra-prediction modes based on the template matching technology (i.e., TIMD, DIMD, TMRL, SGPM, TMRL, CCCM, and interpolation filtering prediction modes) are permitted to be used. TaIf this is indicated, then the use of the intra-predictive modes of the above-mentioned template matching-based techniques is permitted, and it is explained that the intra-predictive mode to be specifically used by the current block will be further determined based on other information.

[0458] Based on the above description, in embodiments of the present application, the step of determining whether the current block is permitted to use the interpolation filter prediction mode includes the step of determining whether to obtain first information indicating whether the template matching-based technology is enabled, and the step of determining whether the current block is permitted to use the interpolation filter prediction mode based on the first information. For example, if the first information indicates that the template matching-based technology is not enabled, it is determined that the current block is not permitted to make predictions using the interpolation filter prediction mode. Alternatively, for example, if the first information indicates that the template matching-based technology is enabled, other information is used to determine whether the current block makes predictions using the interpolation filter prediction mode.

[0459] The embodiments of this application do not limit the specific forms of expression of the first information described above.

[0460] For example, the first piece of information described above may be GCI, sequence-level, frame-level, slice-level, or block-level instruction information.

[0461] In one example, if the first information is sequence-level instruction information, the first information is determined, and if the first information indicates that the template matching-based technique is enabled, the second information (sps_eip_enabled_flag) is then determined, and based on the second information, it is determined whether the current block is allowed to use the interpolation filter prediction mode. If the first information indicates that the template matching-based technique is not enabled, it is directly determined that the current block will perform the prediction without applying the interpolation filter prediction mode, and the step of determining the second information is skipped.

[0462] In some embodiments, the encoding side determines whether the current block can be predicted using the interpolation filter prediction mode, including at least one of the following conditions: 1) Is the current image an intra-predicted image or not? 2) Whether or not higher-level syntax allows it. This is optional, and higher-level syntax includes sequence level, frame level, slice level, block level, etc. See the explanation above for details. 3) Whether the current block size and shape are permitted. See the explanation above for details. 4) Whether the current block's location is permitted. See the explanation above for details.

[0463] In some embodiments, if the encoding side determines that the current sequence is permitted to perform predictions using interpolation filter prediction mode, it writes third information to the code stream to indicate whether the current block will perform predictions using interpolation filter prediction mode.

[0464] The embodiments of this application are not limited to the specific representation of the third piece of information described above, but may also be any instruction information that indicates whether or not the current block is performing a prediction using the interpolation filter prediction mode.

[0465] For example, the third piece of information can be represented as intra_eip_flag, which indicates whether the current block makes predictions using interpolation filter prediction mode by assigning different values ​​to intra_eip_flag. For example, if intra_eip_flag=0, it indicates that the current block makes predictions without using interpolation filter prediction mode, and if intra_eip_flag=1, it indicates that the current block makes predictions using interpolation filter prediction mode. In this way, the encoding side writes this preset flag intra_eip_flag to the code stream, and the decoding side determines the prediction mode of the current block based on the value of this decoded preset flag intra_eip_flag. For example, if this preset flag intra_eip_flag=1, it indicates that the prediction mode of the current block is interpolation filter prediction mode, and the decoding side further determines that the current block makes predictions using interpolation filter prediction mode.

[0466] In some embodiments, as shown in Figure 29, the process for determining the prediction mode of the current block in embodiments of the present application may include the step of first determining whether the current block can be predicted using an interpolation filter prediction mode, for example, if the second information at the sequence level indicates that the current sequence allows the use of an interpolation filter prediction mode, and it is determined that the position of the current block in the current image satisfies a predetermined position requirement, and the size of the current block satisfies a predetermined block size requirement, then determining that the current block can be predicted using an interpolation filter prediction mode. Next, filter coefficients are obtained, the current block is predicted based on these filter coefficients, and the predicted value of the current block is obtained. Simultaneously, a rough selection of prediction modes is performed with other intra-prediction mode tools, several prediction modes with low cost are selected, and then a fine selection is performed to determine the final intra-prediction mode, which is to be used as the prediction mode for the current block. If it is determined that the current block cannot be predicted using an interpolation filter prediction mode, the selection of interpolation filter prediction modes is skipped.

[0467] As an example, at the stage when the prediction modes for the current block are roughly selected, the encoding side calculates the cost of each candidate intra prediction mode (including the interpolation filter prediction mode), and the formula for calculating the cost is shown in equation (10).

number

[0468] In one example, the calculation of the strain value D is as shown in equation (11).

number

[0469] The encoding side determines the cost of each candidate prediction mode, and then selects and refines several candidate prediction modes from among the multiple candidate modes.

[0470] JPEG2024216632000051.jpg46169

[0471] The encoding side determines the candidate prediction mode that minimizes the cost of selection as the prediction mode for the current block.

[0472] When the encoding side determines that the prediction mode for the current block is the interpolation filter prediction mode, it executes the step in step S201 above.

[0473] Next, we will explain the process by which the encoding side makes predictions for the current block using the interpolation filter prediction mode.

[0474] When the encoding side decides that the current block will perform predictions using the interpolation filter prediction mode, it first determines the reference region and interpolation filter for the current block.

[0475] Next, we will describe the specific process by which the encoding side determines the reference region of the current block.

[0476] In the embodiments of this application, the reference region of the current block is part or all of the reconfigured region surrounding the current block.

[0477] For example, as shown in Figure 12, the reconfiguration region around the current block may include the upper reconfiguration region of the current block, the left reconfiguration region of the current block, the upper right reconfiguration region of the current block, the lower left reconfiguration region of the current block, and the upper left reconfiguration region of the current block.

[0478] The embodiments of this application do not limit the specific shape and size of the reference region of the current block.

[0479] For example, the reference region of the current block includes one of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For instance, the reference region of the current block is the upper reconstruction region of the current block, or the reference region of the current block is the left reconstruction region of the current block.

[0480] For example, the reference region of the current block includes any two of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region and the left reconstruction region of the current block. Alternatively, for example, the reference region of the current block includes the upper reconstruction region and the lower left reconstruction region of the current block.

[0481] For example, the reference region of the current block includes any three of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region of the current block, the upper right reconstruction region of the current block, and the upper left reconstruction region of the current block. Also, for example, the reference of the current block region This includes the left reconfiguration area of ​​the current block, the upper left reconfiguration area of ​​the current block, and the lower left reconfiguration area of ​​the current block.

[0482] For example, the reference region of the current block includes any four of the following reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block. For example, the reference region of the current block includes the upper reconstruction region of the current block, the upper right reconstruction region of the current block, the upper left reconstruction region of the current block, and the left reconstruction region of the current block. Also, for example, the reference of the current block region This includes the left reconfiguration region of the current block, the upper left reconfiguration region of the current block, the lower left reconfiguration region of the current block, and the upper reconfiguration region of the current block.

[0483] In one example, the reference region of the current block includes five reconstruction regions: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block.

[0484] In the embodiments of this application, specific methods by which the encoding side determines the reference region of the current block include, but are not limited to, the following:

[0485] In Method 1, the reference region of the current block is the default region, for example, the encoding and decoding sides, by default, include at least one of the following in the reference region of the current block: the upper reconstruction region of the current block, the left reconstruction region of the current block, the upper right reconstruction region of the current block, the lower left reconstruction region of the current block, and the upper left reconstruction region of the current block.

[0486] In method 2, a first cost is determined for predicting the current block based on P reference regions, and the reference region with the smallest first cost among the P reference regions is determined as the reference region of the current block.

[0487] In this implementation, the encoding side predicts the current block based on P reference regions, determines a first cost corresponding to each reference region, and then determines the reference region with the smallest first cost among the P reference regions as the reference region of the current block.

[0488] In some embodiments, the encoder writes a fourth piece of information to the code stream indicating the type of reference region of the current block. That is, in this scheme 2, the encoder further indicates the type of the determined reference region of the current block to the decoder using the fourth piece of information.

[0489] Note that the type and shape of these predetermined P reference regions are all different.

[0490] The embodiments of this application do not particularly limit the specific number and shape of the P reference regions.

[0491] In one example, the P reference regions include at least one of the first reference region, the second reference region, and the third reference region.

[0492] Here, as shown in Figure 13A, the first reference region is above, to the right of, to the left of the current block. under, and the upper left reconstruction area. As shown in Figure 13B, the second reference area includes the reconstruction areas above, to the right of, and to the upper left of the current block. As shown in Figure 13C, the third reference area includes the left side of the current block, and to the left under , and including the reconstruction area in the upper left.

[0493] The embodiments of this application do not limit the specific form of representation of the fourth information, and any indicating information that can indicate the type of reference region of the current block is acceptable.

[0494] For example, eip_ref_type is used to represent a fourth piece of information, and for instance, different values ​​of eip_ref_type indicate different types of reference regions.

[0495] For example, see Figures 13A to 13 C The correspondence between the three reference regions shown and the value of eip_ref_type is shown in Table 4.

[0496] Based on Table 4 above, the encoding side determines the value of the fourth piece of information, eip_ref_type, based on the reference region type of the current block. For example, if the current block's reference region is determined to be the first reference region, eip_ref_type is set to 0. If the current block's reference region is determined to be the second reference region, eip_ref_type is set to 1. If the current block's reference region is determined to be the third reference region, eip_ref_type is set to 2.

[0497] In the above explanation, the P reference regions were described as the three reference regions shown in Figures 13A to 13C. While the P reference regions in the embodiments of this application may include reference regions other than the three mentioned above, the embodiments of this application are not limited to these. The correspondence between the reference regions shown in Table 4 and the values ​​of eip_ref_type can be adaptively adjusted according to the number of reference regions.

[0498] In some embodiments, the encoding side can write the fourth piece of information to the code stream using a truncated binary code encoding scheme.

[0499] For example, the correspondence between truncated binary code, the value of eip_ref_type, and the type of reference region is shown in Table 5.

[0500] In the embodiments of this application, the encoding side can encode the codeword of the truncated binary code using either an equiprobability coding scheme or a context model coding scheme.

[0501] The encoding side can determine the reference region of the current block using either method 1 or method 2 described above, or it can determine the reference region of the current block using method 3 described below.

[0502] In method 3, the reference region of the current block is determined from a predetermined number of P reference regions based on the shape of the current block.

[0503] In this method 3, predictions are made using different reference regions for current blocks of different shapes, thereby improving the accuracy of the predictions.

[0504] For example, if the current block is square in shape, the first type of reference region is used.

[0505] For example, if the current block is a rectangle with a width greater than its height, a second type of reference area is used.

[0506] For example, if the current block is a rectangle with a width smaller than its height, a third type of reference area is used.

[0507] In other words, in the embodiment of this application, the correspondence between P reference regions and the shape of the current block is predetermined. In this way, the encoding side can determine the reference region of the current block from among the P reference regions based on the correspondence between the P reference regions and the shape of the current block, according to the shape of the current block.

[0508] Next, we will describe the process by which the encoding side determines the interpolation filter for the current block.

[0509] In the embodiments of this application, the specific shape of the interpolation filter is not limited.

[0510] For example, the interpolation filters according to the embodiments of this application include, but are not limited to, square interpolation filters and interpolation filters where the height is less than the width.

[0511] For example, square interpolation filters include, but are not limited to, the 4x4 interpolation filter shown in Figure 14A.

[0512] Furthermore, interpolation filters where the height is greater than the width include, but are not limited to, the 5×3 interpolation filter shown in Figure 14B, the 6×2 interpolation filter shown in Figure 14D, and the 7×1 interpolation filter shown in Figure 14G.

[0513] Furthermore, interpolation filters where the height is smaller than the width include, but are not limited to, the 3×5 interpolation filter shown in Figure 14C, the 2×6 interpolation filter shown in Figure 14E, and the 1×7 interpolation filter shown in Figure 14F.

[0514] JPEG2024216632000052.jpg14169

[0515] In the embodiments of this application, the specific methods by which the encoding side determines the interpolation filter for the current block include, but are not limited to, the following methods.

[0516] Method 1 uses the interpolation filter of the current block as the default interpolation filter. For example, the encoding and decoding sides default to one of the interpolation filters shown in Figures 14A to 14G for the current block. For example, the default interpolation filter is a 4x4 interpolation filter.

[0517] In method 2, the encoding side determines the interpolation filter for the current block from a predetermined Q interpolation filters.

[0518] For example, the encoding side randomly selects one interpolation filter from Q interpolation filters to be used as the interpolation filter for the current block.

[0519] For example, the encoding side determines the second cost for predicting the current block using Q interpolation filters, and then selects the interpolation filter with the smallest second cost among the Q interpolation filters as the interpolation filter for the current block.

[0520] In some embodiments, the encoding side writes a fifth piece of information to the code stream that indicates the shape of the interpolation filter for the current block.

[0521] In this embodiment, the encoding side determines the interpolation filter for the current block from a predetermined Q interpolation filters, for example, the encoding side determines a second cost corresponding to each of these Q interpolation filters, and determines the interpolation filter with the minimum second cost as the interpolation filter for the current block. Next, the shape of the interpolation filter with the minimum determined second cost is determined by the fifth information. decrypt This is shown on the side. In this way, the decoding side obtains a fifth piece of information by decoding the code stream, and further determines the interpolation filter for the current block from a predetermined Q interpolation filters based on the shape of the interpolation filter indicated by the fifth piece of information.

[0522] Note that the shapes of these predetermined Q interpolation filters are all different.

[0523] The embodiments of this application do not particularly limit the specific number and shape of the Q interpolation filters. For example, the Q interpolation filters include at least one of a first interpolation filter which is a square interpolation filter, a second interpolation filter which is a rectangle whose width is greater than its height, and a third interpolation filter which is a rectangle whose height is greater than its width.

[0524] In one example, Q interpolation filters are shown in Figures 14A to 14. G Includes multiple interpolation filters.

[0525] The embodiments of this application do not limit the specific representation format of the fifth piece of information, and any instruction information capable of indicating the shape of the interpolation filter of the current block is acceptable.

[0526] In one example, the fifth piece of information is represented by eip_filter_type, and different interpolation filters of different shapes are indicated depending on the value of eip_filter_type.

[0527] For example, if the Q interpolation filters are the five interpolation filters shown in Figure 15, the correspondence between the five interpolation filters and the value of eip_filter_type is shown in Table 6.

[0528] Based on Table 5 above, the encoding side determines the value of the fifth piece of information, eip_filter_type, based on the shape of the interpolation filter for the current block that has been determined. For example, if the shape of the interpolation filter for the current block is determined to be 4x4, then eip_filter_type=0 is determined. If the shape of the interpolation filter for the current block is determined to be 3x5, then eip_filter_type=1 is determined. If the shape of the interpolation filter for the current block is determined to be 5x3, then eip_filter_type=2 is determined. If the shape of the interpolation filter for the current block is determined to be 2x6, then eip_filter_type=3 is determined. If the shape of the interpolation filter for the current block is determined to be 6x2, then eip_filter_type=4 is determined.

[0529] In some embodiments, the encoding side can write the fifth piece of information to the code stream using a truncated binary code encoding mode.

[0530] For example, if a given number of interpolation filters includes the five interpolation filters shown in Figure 15, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 7.

[0531] In this case, the five types of interpolation filter shapes shown in Table 7 and the three types of reconstruction region types shown in Table 5 result in a total of 15 combinations of interpolation filters and reconstruction regions.

[0532] In some embodiments, when the embodiments of this application include seven interpolation filters as shown in Figure 16, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 9.

[0533] In this case, the seven types of interpolation filter shapes shown in Table 9 and the three types of reconstruction region types shown in Table 5 result in a total of 21 combinations of interpolation filters and reconstruction regions.

[0534] In some embodiments, when the embodiments of this application include the three interpolation filters shown in Figure 17, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 10.

[0535] In this case, the three types of interpolation filter shapes shown in Table 10 and the three types of reconstruction region types shown in Table 5 result in a total of nine combinations of interpolation filters and reconstruction regions.

[0536] In some embodiments, when the embodiments of this application include the three interpolation filters shown in Figure 18A, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 11.

[0537] In this case, the three types of interpolation filter shapes shown in Table 11 and the three types of reconstruction region types shown in Table 5 result in a total of nine combinations of interpolation filters and reconstruction regions.

[0538] In some embodiments, when the embodiments of this application include the three interpolation filters shown in Figure 18B, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 12.

[0539] In this case, the three types of interpolation filter shapes shown in Table 12 and the three types of reconstruction region types shown in Table 5 result in a total of nine combinations of interpolation filters and reconstruction regions.

[0540] In some embodiments, when the embodiments of this application include three interpolation filters as shown in Figure 19, the correspondence between the truncated binary code, the value of eip_filter_type, and the shape of the interpolation filter is shown in Table 13.

[0541] In this case, the three types of interpolation filter shapes shown in Table 13 and the three types of reconstruction region types shown in Table 5 result in a total of nine combinations of interpolation filters and reconstruction regions.

[0542] The encoding side can determine the interpolation filter for the current block using either method 1 or method 2 described above, or it can determine the interpolation filter for the current block using method 3 described below.

[0543] Method 3 determines the interpolation filter for the current block from a predetermined number of Q interpolation filters based on the shape of the current block.

[0544] In this method 3, predictions are made using different interpolation filters for current blocks of different shapes, thereby improving prediction accuracy.

[0545] For example, if the current block shape is a square, the interpolation filter for the first shape is used.

[0546] For example, if the current block shape is a rectangle with a width greater than its height, a second shape interpolation filter is used.

[0547] For example, if the current block shape is a rectangle with a width smaller than its height, a third shape interpolation filter is used.

[0548] In other words, in the embodiment of this application, the correspondence between Q interpolation filters and the shape of the current block is predetermined. Thus, the encoding side can determine the interpolation filter for the current block from the Q interpolation filters based on the correspondence between these Q interpolation filters and the shape of the current block, according to the shape of the current block.

[0549] In the embodiments of this application, the encoding side determines the reference region and interpolation filter of the current block based on the above steps, and then determines the predicted block of the current block based on the reference region and interpolation filter.

[0550] Next, we will explain how to determine the filter coefficients of an interpolation filter based on the reference region.

[0551] In the embodiments of this application, the method by which the encoding side determines the filter coefficients of the interpolation filter includes at least the following methods.

[0552] In Method 1, the interpolation filter determined above is slid across the reference region of the current block to create the Wiener-Hoff equation. Next, the Wiener-Hoff equation is solved to obtain the filter coefficients of the interpolation filter.

[0553] JPEG2024216632000053.jpg45168

[0554] In one example, the interpolation filter is slid within the reference region of the current block to create the Wiener-Hoff equation, as shown in equation (3).

[0555] Since the reference region of the current block is the reconstruction region, all parameters other than the interpolation filter coefficients are known in equation (3) above. Therefore, by solving equation (3) above, the filter coefficients of the interpolation filter for the current block can be determined.

[0556] For example, the encoding side can obtain the filter coefficients of the filter by solving the Wiener-Hoff equation shown in equation (3) above using the Cholesky decomposition autocorrelation matrix.

[0557] The embodiments of this application do not limit the sliding steps of the interpolation filter within the reference region.

[0558] In one example, as shown in Figure 20A, the horizontal and vertical slide steps of the interpolation filter within the reference region are equal, and both are 1 pixel.

[0559] In one example, the horizontal and vertical slide steps of the interpolation filter within the reference region are not equal. For instance, the horizontal slide step is 2 pixels and the vertical slide step is 1 pixel. Alternatively, for example, the horizontal slide step is 1 pixel and the vertical slide step is 2 pixels.

[0560] In one example, at least one of the horizontal slide step and vertical slide step of the interpolation filter within the reference region is greater than a predetermined step. For example, the horizontal slide step is greater than a predetermined step. Also, for example, the vertical slide step is greater than a predetermined step. Also, for example, both the horizontal slide step and the vertical slide step are greater than a predetermined step. The embodiments of this application do not limit the specific value of the predetermined step. For example, it can be a number such as 1, 2, or 3.

[0561] In method 2, the encoding side determines the filter coefficients by following steps S201-A1 to S201-A4. S201-A1 determines the first reconstruction region around the current block. S201-A2 determines the average reconstruction value of pixels based on the reconstruction value of the first reconstruction region. S201-A3 removes the average value from the reconstruction values ​​of the pixel points in the reference region, based on the pixel-average reconstruction value. S201-A4 uses the pixel values ​​of pixel points with the mean value removed in the reference region as input to the interpolation filter, slides the interpolation filter within the reference region to obtain the filter coefficients of the interpolation filter.

[0562] In this method 2, the reference region is subjected to mean removal, and the filter coefficients of the interpolation filter are determined based on the mean-removed reference region. Since the amount of data is reduced when the mean is removed from the reference region, the efficiency of determining the filter coefficients can be improved when determining the filter coefficients based on the mean-removed reference region.

[0563] Specifically, the encoding side first determines a first reconfiguration region, which can be any part of the reconfiguration region around the current block.

[0564] In the embodiments of this application, the method by which the encoding side determines a first reconstruction region around the current block includes at least some of the following methods:

[0565] In Method 1, the encoding side, by default, determines one reconfiguration region surrounding the current block as the first reconfiguration region.

[0566] For example, as shown in Figure 20B, the encoding side, by default, defines the first reconstruction region as the area consisting of the top row, leftmost column, and top-left corner pixel of the current block.

[0567] In method 2, the first reconstruction region is determined based on the current block shape.

[0568] For example, if the current block is square, the reconstructed pixel region of the top row and leftmost column of the current block is determined as the first reconstructed region.

[0569] For example, if the current block is a rectangle with a width greater than its height, the top row of reconstructed pixels of the current block is determined as the first reconstructed region.

[0570] For example, if the current block is a rectangle with a height greater than its width, the leftmost column of reconstructed pixels of the current block is reconstructed as the first reconstructed region.

[0571] The method for determining the first reconstruction region based on the current block shape includes, but is not limited to, the examples described above.

[0572] After the encoding side determines the first reconstruction region, the pixel-average reconstruction value m is determined based on the reconstruction value of the first reconstruction region.

[0573] The embodiments of this application are not limited to a specific method for determining the pixel average reconstruction value m based on the reconstruction value of the first reconstruction region in S201-A2 described above.

[0574] In method 1, step S201-A2 includes determining the average value of the reconstruction values ​​of the first reconstruction region as the pixel average reconstruction value m.

[0575] In one example of Method 1, if the first reconstruction region is as shown in Figure 20B, then Table 1 4 The pixel-average reconstruction value m is calculated using the method shown.

[0576] In one example of Method 1, if the first reconstruction region is the top row and / or leftmost column of the current block, the average of the reconstruction values ​​of the top row and / or leftmost column can be determined as the pixel-average reconstruction value m. In this case, the pixel-average reconstruction value m can be calculated using the method shown in Table 15.

[0577] As shown in Table 15 above, if the first reconstruction region is the row above and / or the column to the left of the current block, the pixel-average reconstruction value m can be quickly calculated using a shift operation instead of division.

[0578] In method 2, steps S201-A2 include determining the pixel average reconstruction value based on the current block shape and the reconstruction value of the first reconstruction region.

[0579] For example, if the current block is a square, the average value of the entire first reconstruction region determined above is determined as the pixel-average reconstruction value m.

[0580] In one example of method 2, the first reconstruction region includes an upper reconstruction region and a left reconstruction region of the current block, in which case the step of determining the pixel average reconstruction value based on the shape of the current block and the reconstruction value of the first reconstruction region includes the step of determining the first region from the upper reconstruction region and the left reconstruction region based on the shape of the current block, the step of determining the average reconstruction value of the first region based on the reconstruction value of the first region, and the step of determining the pixel average reconstruction value based on the average reconstruction value of the first region.

[0581] In this example, if the first reconstruction region includes the upper and left-side reconstruction regions of the current block, the mean value is determined, similar to the DC prediction mode, to reduce computational complexity. Specifically, the first region is determined from the upper and left-side reconstruction regions included in the first reconstruction region, based on the shape of the current block.

[0582] For example, if the width of the current block shape is greater than its height, the upper reconstruction region is determined to be the first region.

[0583] For example, if the height of the current block shape is greater than its width, the left-side reconstruction region is determined to be the first region.

[0584] For example, if the height of the current block shape is equal to its width, the upper reconstruction region and the left-side reconstruction region are determined as the first region.

[0585] Next, the average reconstruction value of the selected first region is determined based on the reconstruction value of the first region, and further, the pixel average reconstruction value m is determined based on the average reconstruction value of the first region. For example, this average reconstruction value of the first region is determined as the pixel average reconstruction value m.

[0586] In other words, in this example, if the width of the current block shape is greater than the height, the average reconstruction value of the upper reconstruction region of the current block is determined as the pixel average reconstruction value m. If the height of the current block shape is greater than the width, the average reconstruction value of the left reconstruction region of the current block is determined as the pixel average reconstruction value m. If the height of the current block shape is equal to the width, the average reconstruction value of both the upper and left reconstruction regions of the current block is determined as the pixel average reconstruction value m.

[0587] The encoding side determines the pixel-average reconstruction value, and then removes the average value from the reconstruction values ​​of the pixel points in the reference region based on that pixel-average reconstruction value.

[0588] For example, for each pixel point in the reference region, the reconstructed value of that pixel point is divided by the above-mentioned average reconstructed value and rounded to obtain the pixel value of the pixel point in the reference region after the mean value has been removed.

[0589] For example, the encoding side subtracts the pixel average reconstruction value from the reconstruction value of a pixel point in the reference region to obtain the pixel value of the pixel point in the reference region from which the mean value has been removed. For example, for each pixel point in the reference region, the above pixel average reconstruction value is subtracted from the reconstruction value of that pixel point to obtain the pixel value of the pixel point in the reference region from which the mean value has been removed.

[0590] The embodiments of this application do not limit the specific method by which the encoding side removes the average value from the reconstructed values ​​of the pixel points in the reference region based on the pixel average reconstructed values.

[0591] The encoding side removes the mean value from the reconstructed values ​​of the pixel points in the reference region based on the method described above, obtains the pixel values ​​of the pixel points from which the mean value has been removed in the reference region, then executes the steps in S201-A4 above, uses the pixel values ​​of the pixel points from which the mean value has been removed in the reference region as input to the interpolation filter, slides the interpolation filter within the reference region, and obtains the filter coefficients of the interpolation filter.

[0592] For example, as shown in Figure 21, if the interpolation filters of the current block are of five different shapes, and the reference regions of the current block are of three different shapes, the interpolation filters of the current block are slid across the current block's mean-removed reference regions to obtain the filter coefficients of the interpolation filters. The interpolation filters may be slid one row horizontally or one column vertically across the mean-removed reference regions.

[0593] JPEG2024216632000054.jpg45168

[0594] In one example, the Wiener-Hoff equation created by sliding the interpolation filter within the reference region of the current block is given by equation (4).

[0595] Since the reference region of the current block is the reconstruction region, all parameters other than the interpolation filter coefficients are known in equation (4) above. Therefore, by solving equation (4) above, the filter coefficients of the interpolation filter for the current block can be determined.

[0596] In one example, the encoding side can obtain the filter coefficients of the filter by solving the Wiener-Hoff equation shown in equation (4) above using the Cholesky decomposition autocorrelation matrix.

[0597] Based on the above steps, the encoding side determines the filter coefficients of the interpolation filter and then performs the next step S202.

[0598] S202, based on the filter coefficients, an interpolation filter is used to predict at least two pixel points in the current block in parallel, and the predicted block for the current block is determined.

[0599] Based on the above steps, the encoding side determines the filter coefficients of the interpolation filter, and then, using the interpolation filter based on these filter coefficients, performs an interpolation filter prediction for the current block and obtains the predicted block for the current block.

[0600] In related technologies, when performing interpolation filter prediction for the current block using an interpolation filter, the next pixel point is predicted after the prediction of one pixel point in the current block is completed. For example, as shown in Figure 22A, the interpolation filter performs interpolation filter prediction one by one along the horizontal direction for each pixel point in the current block. During prediction, after the prediction of the previous pixel point in the horizontal direction is completed, the predicted value of that previous pixel point is used as the input pixel value of the interpolation filter for the next pixel point and is used to predict the next pixel point. As shown in Figure 22, if the shape of the interpolation filter for the current block is 4x4, the encoding side performs interpolation prediction one by one for each position in the current block using an interpolation filter with known filter coefficients. Specifically, for the r-th point in the current block, first, the pixel values ​​of the N positions corresponding to the r-th point are determined based on the shape of the interpolation filter for the current block. For example, as shown in Figure 22, in a 4x4 interpolation filter, the dark positions are the positions of the r-th point to be processed, and the 15 light positions are the N positions corresponding to the r-th point. Here, the block to be predicted in Figure 22 is the current block. Next, the pixel values ​​of the N positions corresponding to the r-th point are determined. For example, for any of the N positions, if the position lies in the reconstruction region surrounding the current block, the reconstructed value of the position is determined as the pixel value of that position. If the position lies within the current block, the predicted value of the position is determined as the pixel value of that position.

[0601] For example, as shown in Figure 22B, the interpolation filter performs one interpolation filter prediction for each pixel point in the current block along the vertical direction. During prediction, after the prediction of the previous pixel point in the vertical direction is completed, the predicted value of the previous pixel point is used as the input pixel value for the interpolation filter of the next pixel point and used for the prediction of the next pixel point. In other words, in conventional techniques, when performing interpolation filter prediction for the current block using an interpolation filter, only one pixel prediction can be performed at a time, so prediction takes time, prediction efficiency decreases, and coding efficiency is affected.

[0602] To solve the above technical problems, the embodiment of this application predicts at least two points in the current block in parallel when performing interpolation filter prediction for the current block using an interpolation filter, that is, the encoding side can perform interpolation filter prediction for at least two pixel points in the current block using an interpolation filter at the same time.

[0603] In the embodiments of this application, the encoding side performs interpolation filter predictions for at least two pixel points in the current block using an interpolation filter at the same time, which includes at least two embodiments.

[0604] In the first implementation method, for at least two pixel points in the current block, these at least two pixel points are adjacent pixel points. The encoding side first determines the input information for the interpolation filter corresponding to these at least two pixel points, inputs this input information into the interpolation filter to perform interpolation filter prediction, obtains one predicted value, and determines the predicted values ​​for these at least two pixel points based on this predicted value. For example, the predicted values ​​are processed based on the associated feature information of the at least two pixel points, and predicted values ​​corresponding to each of these at least two pixel points are obtained. Alternatively, for example, these predicted values ​​are determined as the predicted values ​​corresponding to each of these at least two pixel points.

[0605] In this first implementation method, the specific method by which the encoding side determines the input information for the interpolation filter corresponding to these at least two pixel points is not limited.

[0606] For example, based on the shape of the interpolation filter, the same interpolation filter input value is determined for at least two of these pixel points, and this same interpolation filter input value is used as the input information for the interpolation filter. For example, these at least two pixel points include pixel point 1 and pixel point 2, and based on the shape of the interpolation filter, N input values ​​corresponding to pixel point 1 and N input values ​​corresponding to pixel point 2 are determined. The same input value is determined from the N input values ​​corresponding to pixel point 1 and the N input values ​​corresponding to pixel point 2, and this same input value is used as the input value for the interpolation filter. Note that if any of the N input values ​​corresponding to pixel point 2 or pixel point 1 contain unencoded values, those unencoded values ​​are discarded.

[0607] For example, based on the shape of the interpolation filter, the input value corresponding to the pixel point with the most encoded input values ​​among these at least two pixel points is determined as the input information for the interpolation filter. For example, these at least two pixel points include pixel point 1 and pixel point 2, and based on the shape of the interpolation filter, N input values ​​corresponding to pixel point 1 and N input values ​​corresponding to pixel point 2 are determined. Here, all N input values ​​corresponding to pixel point 1 are encoded, and the N input values ​​corresponding to pixel point 2 include unencoded input values, so the N input values ​​corresponding to pixel point 1 are used as the input information for the interpolation filter.

[0608] Method 1 described above explains the specific process in which the encoding side inputs the same input value to the interpolation filter at once, and then the interpolation filter predicts the predicted values ​​of at least two pixel points in the current block.

[0609] In the second method, at the same time, an interpolation filter is applied to each of at least two pixels in the current block. For example, at time t, the encoding side uses an interpolation filter to perform an interpolation filter prediction on pixel point 1 in the current block to obtain the predicted value of pixel point 1, and at the same time uses an interpolation filter to perform an interpolation filter prediction on pixel point 2 in the current block to obtain the predicted value of pixel point 2.

[0610] The embodiments of this application do not limit the prediction direction when the encoding side uses an interpolation filter to predict at least two pixel points in the current block in parallel.

[0611] In some embodiments, the encoding side can use an interpolation filter to predict at least two pixels in the current block in parallel along the horizontal direction. For example, when predicting at least two pixel points in parallel, if one or more input values ​​of the interpolation filter corresponding to the pixel points are not encoded, the unencoded input values ​​are discarded and the encoded input values ​​are used as input values ​​for the interpolation filter.

[0612] In some embodiments, the encoding side can use an interpolation filter to predict at least two pixels in the current block in parallel along the vertical direction. For example, when predicting at least two pixel points in parallel, if one or more input values ​​of the interpolation filter corresponding to the pixel points are not encoded, the unencoded input values ​​are discarded and the encoded input values ​​are used as input values ​​for the interpolation filter.

[0613] In some embodiments, the encoding side can use an interpolation filter to predict at least two pixels in the current block in parallel along the diagonal. Based on this, as an example, S202 above includes the following step S202-A.

[0614] S202-A, based on the filter coefficients, parallel interpolation filter predictions are performed on pixel points on the same diagonal of the current block using an interpolation filter along the diagonal direction to obtain a predicted block for the current block.

[0615] As an example, as shown in Figure 23, when the encoding side predicts a pixel point in the current block using an interpolation filter, the pixel point to be predicted is located in one corner of the selection area of ​​the interpolation filter (for example, the lower right corner or the upper left corner). In this way, for any pixel point on the same diagonal, the N positions corresponding to that selected pixel point based on the shape of the interpolation filter do not include the positions of any other pixel points on that diagonal. That is, none of the N positions corresponding to each pixel point on the same diagonal include any pixel points on the diagonal. For example, taking two adjacent pixel points a and b on the same diagonal of the current block as an example, as shown in Figure 23, based on the shape of the interpolation filter, N positions corresponding to pixel point a and N positions corresponding to pixel point b are determined, where the N positions corresponding to pixel point a and the N positions corresponding to the determined pixel point b are... Place , including pixel points on the diagonal. do not have Based on this, when the encoding side performs interpolation filter prediction for the current block using an interpolation filter, it can perform parallel interpolation filter prediction along the diagonal direction for pixel points located on the same diagonal in the current block. For example, parallel interpolation filter prediction can be performed for pixel points a and b located on the same diagonal.

[0616] The results of the investigation in this application show the left side, lower left, top, and left of the current block. above Since the upper right region has already been encoded, we can determine the starting point of the prediction along the diagonal direction of the current block based on the encoded region and the shape of the interpolation filter.

[0617] In one example, as shown in Figure 23, the encoding side performs interpolation filter prediction along the diagonal direction relative to the current block, starting from the upper left corner of the current block. Step S202-A above includes the following step S202-A1.

[0618] In S202-A1, the encoding side uses an interpolation filter to perform parallel interpolation filter predictions on pixel points on the same diagonal of the current block, starting from the upper left corner of the current block and moving diagonally, based on the filter coefficients, to obtain the predicted block for the current block. In this example, the pixel point to be predicted is located in the lower right corner of the interpolation filter selection region.

[0619] The embodiments of this application do not limit the specific orientation of the diagonal direction.

[0620] In some embodiments, as shown in Figure 23, when the encoding side starts from the upper-left corner of the current block and predicts pixel points located on the same diagonal in the current block in parallel, the diagonal direction includes at least one of the directions from upper-right to lower-left and from lower-left to upper-right.

[0621] In one example, as shown in Figure 24A, the diagonal direction includes the direction from the upper right to the lower left, and in this case, as indicated by the arrows in Figure 24A, all of the diagonal directions of the current block lower left from upper right It is the direction towards.

[0622] In one example, as shown in Figure 24B, the diagonal direction includes the direction from the bottom left to the top right, and in this case, as indicated by the arrow in Figure 24B, the diagonal direction of the current block is upper right from lower left It is the direction towards.

[0623] In one example, as shown in Figure 24C, the diagonal direction includes the direction from the upper right to the lower left and the direction from the lower left to the upper right. In this case, as indicated by the arrows in Figure 24C, the diagonal direction of the current block includes two types of directions: the direction from the upper right to the lower left and the direction from the lower left to the upper right.

[0624] In the embodiments of this application, the encoding side predicts pixel points located on the same diagonal in the current block in parallel, so the specific orientation of the diagonal direction does not limit the technical means of the embodiments of this application.

[0625] In the embodiments of this application, the encoding side makes predictions for each prediction, using a pixel point on one diagonal of the current block as the unit. The process by which the encoding side predicts pixel points on each diagonal of the current block in parallel is similar, and for the sake of explanation, the k-th diagonal of the current block is used as an example. In this case, S202-A1 above includes the following steps S202-A11 and S202-A12. In S202-A11, for M pixels on the k-th diagonal of the current block, the predicted values ​​of the M pixels are determined in parallel using an interpolation filter based on the filter coefficients, where k and M are both positive integers. S202-A12 obtains the predicted value of the current block based on the predicted values ​​of the pixel points on each diagonal in the current block.

[0626] This k-th diagonal may be understood as one of the diagonals of the current block shown in Figure 23, and there are M pixel points on this k-th diagonal. The encoding side determines the predicted values ​​of these M pixel points in parallel using an interpolation filter based on the filter coefficients. That is, the encoding side can simultaneously determine the predicted values ​​of the M pixel points located on the k-th diagonal, significantly increasing the prediction speed.

[0627] To illustrate with an example, as shown in Figure 23, if there are three pixel points on the k-th diagonal of the current block, the encoding side determines the predicted values ​​of these three pixel points in parallel. For example, let's call these three pixel points pixel 1, pixel 2, and pixel 3. At the same time, the encoding side performs interpolation filter prediction on pixel 1 using an interpolation filter based on the filter coefficients to obtain the predicted value of pixel 1, simultaneously performing interpolation filter prediction on pixel 2 using an interpolation filter based on the filter coefficients to obtain the predicted value of pixel 2, and simultaneously performing interpolation filter prediction on pixel 3 using an interpolation filter based on the filter coefficients to obtain the predicted value of pixel 3. In this way, the encoding side determines the predicted values ​​of the three pixel points located on the k-th diagonal of the current block in parallel in a single interpolation filter prediction process, significantly improving the speed of interpolation filter prediction. The encoding side refers to a method for determining the predicted value of the k-th diagonal pixel point, which allows it to determine the predicted values ​​of other diagonal pixel points in the current block, and further obtains the predicted block for the current block, thereby improving the prediction speed of the current block and improving encoding efficiency.

[0628] The embodiments of this application do not limit the specific method by which the encoding side determines the predicted values ​​of M pixel points in parallel using an interpolation filter based on filter coefficients.

[0629] In some embodiments, these M pixel points are points on the k-th diagonal of the current block, so these M pixel points may be understood as adjacent pixel points, and their features are relatively similar. To reduce the complexity of the calculation, the input values ​​of the interpolation filter corresponding to one or more of these M pixel points are determined based on the shape of the interpolation filter. Then, based on the input values ​​of the interpolation filter corresponding to this one or more pixel points, calculations such as averaging and weighting are performed to determine the input values ​​of the interpolation filter corresponding to the other pixel points among the M pixel points. Finally, the predicted values ​​of the M pixel points are determined in parallel based on the filter coefficients and the input values ​​of the interpolation filter corresponding to each of the M pixel points.

[0630] In some embodiments, the step of determining the predicted values ​​of M pixel points in parallel using an interpolation filter based on the filter coefficients in S202-A11 above includes the following steps: S202-A11-a1 determines the pixel values ​​of N positions corresponding to each of the M pixel points in parallel, based on the shape of the interpolation filter. S202-A11-a2 determines the predicted values ​​of M pixel points in parallel, based on the filter coefficients and the pixel values ​​of N positions corresponding to each of the M pixel points.

[0631] In this embodiment, when the encoding side determines the predicted values ​​of M pixel points on the k-th diagonal of the current block in parallel, it determines the pixel values ​​of N positions corresponding to each of these M pixel points in parallel based on the shape of the interpolation filter, and the pixel values ​​of N positions corresponding to each pixel point can be understood as the input values ​​of the interpolation filter corresponding to that pixel point. Next, the encoding side determines the predicted values ​​of the M pixel points in parallel based on the filter coefficients and the pixel values ​​of N positions corresponding to each of the M pixel points.

[0632] To illustrate with an example, as shown in Figure 24B, the current block contains three pixel points on the k-th diagonal, which we will call pixel point 1, pixel point 2, and pixel point 3. We assume that the interpolation filter for the current block is a 4x4 interpolation filter. When the encoding side determines the predicted values ​​of these three pixel points in parallel, it determines 15 pixel values ​​corresponding to pixel point 1 based on the shape of the interpolation filter, uses these 15 pixel values ​​as input to the interpolation filter, and determines the predicted value of pixel point 1 based on the filter coefficients determined above. At the same time, the encoding side determines 15 pixel values ​​corresponding to pixel point 2 based on the shape of the interpolation filter, uses these 15 pixel values ​​as input to the interpolation filter, and determines the predicted value of pixel point 2 based on the filter coefficients determined above. At the same time, the encoding side determines 15 pixel values ​​corresponding to pixel point 3 based on the shape of the interpolation filter, uses these 15 pixel values ​​as input to the interpolation filter, and determines the predicted value of pixel point 3 based on the filter coefficients determined above. In other words, in this embodiment, the coding side determines the predicted values ​​for three pixels of the current block in parallel at the same time, thereby significantly improving the prediction speed and coding efficiency.

[0633] Next, we will explain the specific process for determining the predicted values ​​of the M pixel points in parallel in step S202-A11-a2 above, based on the filter coefficients and the pixel values ​​of the N positions corresponding to each of the M pixel points.

[0634] In some embodiments, for each of the M pixels, the encoding side directly multiplies the pixel values ​​at the N positions corresponding to that pixel point by a filter coefficient in order to obtain a predicted value for that pixel.

[0635] For example, the encoding side obtains the predicted value for each of the M pixel points based on equation (5) above.

[0636] Based on equation (5), the encoding side can determine the predicted values ​​of each pixel point located on the same diagonal in the current block in parallel.

[0637] In some embodiments, the encoding side determines the filter coefficients of the interpolation filter based on equation (4) above. Since equation (4) determines the filter coefficients using the reference region after the mean has been removed, the influence of the pixel mean reconstruction value m must be considered in order to determine the predicted value of the current block based on these filter coefficients.

[0638] In one possible implementation of this embodiment, the interpolation filter coefficients determined in equation (4) are substituted into equation (5) to obtain the predicted values ​​for each point in the current block. Then, the pixel-average reconstruction value m is added to the predicted values ​​for each point to obtain the final predicted values ​​for each point in the current block, and consequently, the predicted block for the current block.

[0639] In another possible implementation of this embodiment, step S202-A11-a2 includes the following steps: S202-A11-a21, based on the pixel average reconstruction value, average value removal is performed in parallel on the pixel values ​​at N positions corresponding to each of the M pixel points, and the average value removed pixel values ​​at N positions corresponding to each of the M pixel points are obtained. S202-A11-a22 determines the predicted values ​​of M pixel points in parallel based on the filter coefficients and the average value of the N positions corresponding to each of the M pixel points after removing the filter coefficients.

[0640] Since the above filter coefficients are determined based on the mean-removed reference region, the encoding side performs mean-removal on the pixel values ​​at N positions corresponding to each of the M pixel points on the k-th diagonal of the current block, based on the pixel mean-reconstruction value, and obtains the mean-removed pixel values ​​at N positions corresponding to each of the M pixel points. For example, for any of the M pixel points, the pixel mean-reconstruction value is subtracted from the pixel values ​​at N positions of that pixel point to obtain the mean-removed pixel values ​​at N positions of that pixel point.

[0641] Next, the predicted values ​​for the M pixel points are determined in parallel based on the filter coefficients and the average removed pixel values ​​of the N positions corresponding to each of the M pixel points.

[0642] The embodiments of this application do not limit the specific method for determining the predicted values ​​of M pixel points in parallel based on filter coefficients and the average value of N positions corresponding to each of the M pixel points after removing the filter coefficients.

[0643] JPEG2024216632000055.jpg37168

[0644] In an alternative implementation, step S202-A11-a22 includes the following steps. S202-A11-a221 determines a second reconfiguration region around the current block, and determines the maximum and minimum reconfiguration values ​​for the second reconfiguration region. S202-A11-a222 obtains first predicted values ​​for M pixel points in parallel, based on the average removed pixel values, filter coefficients, and pixel-average reconstruction values ​​for N positions corresponding to each of the M pixel points. S202-A11-a223 determines the predicted values ​​of M pixel points in parallel based on the first predicted value, maximum reconstruction value, and minimum reconstruction value of M pixel points.

[0645] JPEG2024216632000056.jpg22168

[0646] The embodiments of this application do not limit the specific method for determining the second reconfiguration region around the current block.

[0647] In one example, the second reconfiguration region of the current block coincides with the reference region of the current block.

[0648] In one example, the second reconfiguration region of the current block coincides with the first reconfiguration region of the current block.

[0649] In one example, the reconstruction areas above, to the left, to the right, to the left, and to the left of the current block are determined as the second reconstruction area. For example, the reconstruction areas of the top 13 rows, the left 13 columns, the top 13 rows, the top 13 rows and 13 columns of the top left block, and the bottom 13 columns of the bottom left block are determined as the second reconstruction area.

[0650] Note that there is no specific order in which steps S202-A11-a221 and S202-A11-a222 are performed in the actual implementation process. For example, step S202-A11-a221 may be performed before step S202-A11-a222, after step S202-A11-a222, or in sync with step S202-A11-a222.

[0651] The embodiments of this application do not limit the specific method by which the encoding side obtains first predicted values ​​for M pixel points in parallel based on the average-removed pixel values, filter coefficients, and pixel-average reconstruction values ​​for N positions corresponding to each of the M pixel points.

[0652] For example, for any of the M pixel points, a second predicted value for the pixel point is obtained by multiplying the average value of the N positions of the pixel point (after removing the filter coefficient) by the pixel point, and a first predicted value for the pixel point is obtained by adding the second predicted value to the pixel average reconstruction value.

[0653] For example, the encoding side obtains a first predicted value for this pixel point based on equation (6) above.

[0654] For example, the encoding side obtains a predicted value for the pixel point based on equation (6) above, then performs a prediction process on the predicted value to obtain a first predicted value for the pixel point.

[0655] Based on the above steps, the encoding side determines a first predicted value for M pixel points, and then determines the predicted values ​​for M pixel points in parallel based on the first predicted value, the maximum reconstructed value, and the minimum reconstructed value.

[0656] For example, if, for any of the M pixel points, the first predicted value for that pixel point is greater than the minimum reconstruction value and less than the maximum reconstruction value, then the first predicted value is determined to be the predicted value for that pixel point.

[0657] For example, if the first predicted value of the pixel point is less than or equal to the minimum reconstruction value, the minimum reconstruction value is determined as the predicted value of the pixel point.

[0658] For example, if the first predicted value of the pixel point is greater than or equal to the maximum reconstruction value, the maximum reconstruction value is determined as the predicted value of the pixel point.

[0659] In one example, the encoding side determines the predicted value of the pixel point using the above equation (7).

[0660] The above is an example of determining the predicted values ​​of M pixels on the k-th diagonal in the current block. The encoding side can refer to the above method to determine the predicted values ​​of each pixel on each diagonal in the current block in parallel, and further obtain the predicted values ​​of each point in the current block to construct a predicted block for the current block.

[0661] Based on the steps above, the encoding side performs interpolation filter prediction for the current block, obtains the predicted block for the current block, and then performs the following steps.

[0662] S203, determine the translation kernel corresponding to the current block, encode the current block based on the translation kernel corresponding to the current block and the predicted block, and obtain the code stream.

[0663] Based on the above, when encoding the current block, the encoding side identifies the predicted block of the current block based on the steps above. Next, the current block is subtracted from the predicted block of the current block to obtain the residual block of the current block. Next, the residual block of the current block is transformed to obtain transformation coefficients, the transformation coefficients are quantized to obtain quantization coefficients, and the quantization coefficients are encoded to obtain a code stream.

[0664] When converting the residual values ​​of the current block to obtain conversion coefficients, it is necessary to determine the conversion kernel and then convert the residual values ​​of the current block based on the conversion kernel to obtain the conversion coefficients.

[0665] The embodiments of this application do not limit the specific method by which the encoding side determines the transformation kernel corresponding to the current block.

[0666] In some embodiments, the encoding and decoding sides use a default transformation kernel as the transformation kernel for the current block.

[0667] In some embodiments, the encoding side proceeds to the next steps S203-A and S203- B This determines the translation kernel for the current block. S203-A determines the intra-prediction mode corresponding to the prediction block. S203-B determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to the predicted block.

[0668] Next, we will describe the specific process by which the encoding side determines the intra-prediction mode corresponding to the prediction block.

[0669] As an example, as shown in Figure 7, the conventional intra-prediction modes included in the current VVC are as follows: PLANAR mode: Intra prediction mode index is 0. DC mode: Intra prediction mode index is 1. Angle mode: Intra prediction mode index is 2-66.

[0670] In one example, as shown in Figure 25, the direction of the arrows in the figure indicates the direction of the angle pattern prediction present in the VVC, and the prediction pattern indices used during coding are 2 to 66. If the current block is a non-square block, some angle directions are replaced with wider angles such as -1 to -14 and 67 to 80 in Figure 25.

[0671] In some embodiments, the intra-prediction mode corresponding to the prediction block is the default intra-prediction mode. That is, if the current block makes a prediction using the interpolation filter prediction mode and obtains a prediction block, one of the conventional intra-prediction modes is determined by default as the intra-prediction mode corresponding to that prediction block.

[0672] In some embodiments, the encoding side determines the intra-prediction mode corresponding to the prediction block by the following steps S203-A1 and S203-A2. S203-A1 determines the angle values ​​of R points in the prediction block, where R is a positive integer. S203-A2 determines the intra-prediction mode corresponding to the prediction block based on the angle values ​​of R points.

[0673] In the embodiments of this application, the intra-prediction mode corresponding to a prediction block is determined by aggregating the intra-prediction modes corresponding to the angle values ​​of R points in the prediction block.

[0674] The embodiments of this application do not limit the specific locations and number of R points in the prediction block used to determine the angle value. For example, the R points may be one point in the prediction block or multiple points in the prediction block.

[0675] For example, if the above R points are one point, the encoding side determines the angle value of one point in the prediction block (for example, the center point of the prediction block), determines the intra-prediction mode corresponding to that point based on the angle value of that point, and further determines that intra-prediction mode as the intra-prediction mode corresponding to the prediction block.

[0676] For example, if the R points mentioned above are multiple points, the encoding side determines the angle values ​​of these multiple points, determines the intra-prediction mode corresponding to each of these multiple points based on the angle values ​​of these multiple points, and further determines the intra-prediction mode with the most identical intra-prediction modes among these multiple points as the intra-prediction mode corresponding to the prediction block.

[0677] In some embodiments, when the angular values ​​of R points in a prediction block are determined by a sliding window, the selection of these R points is related to the shape and size of the sliding window. For example, each of the R points is the center point within the sliding window as it slides through the prediction block.

[0678] In the embodiments of this application, the method for determining the angle value of each of the R points is the same, and for the sake of explanation, an example of determining the angle value of the i-th point among the R points will be given.

[0679] The embodiments of this application do not limit the specific method for determining the angle value of a point.

[0680] In some embodiments, step S203-A1 is replaced by step S20 3 -Includes steps S203-A12 and A11. S203-A11, for the i-th point out of R points, determine the horizontal and vertical slopes of the i-th point, where i is a positive integer less than or equal to R. S203-A12, the angle value of the i-th point is determined based on the horizontal and vertical slopes of the i-th point.

[0681] In this embodiment, the encoding side first determines the horizontal and vertical slopes of each of the R points, for example, the i-th point, and then determines the angle value of the i-th point based on the horizontal and vertical slopes.

[0682] The embodiments of this application do not limit the specific method for determining the horizontal and vertical slopes of the i-th point.

[0683] In one example, the horizontal gradient value of the i-th point is determined based on the horizontal change between the predicted values ​​of the points surrounding the i-th point in the prediction block and the predicted value of the i-th point itself, and the vertical gradient value of the i-th point is determined based on the vertical change between the predicted values ​​of the points surrounding the i-th point in the prediction block and the predicted value of the i-th point itself.

[0684] In another example, the encoding side determines the predicted values ​​of points within a sliding window centered on the i-th point in the prediction block, and obtains the horizontal and vertical gradients of the i-th point based on the predicted values ​​of points within the sliding window, the horizontal gradient operator, and the vertical gradient operator.

[0685] In this example, first, a sliding window is determined, for example, a sliding window with a size of 3x3 as shown in Figure 26. This sliding window is then slid within the prediction block, and the horizontal and vertical slopes of the center point of this sliding window are determined each time it is slid. If the current center point of the sliding window is the i-th point, then the predicted values ​​for each point within the current sliding window are obtained, for example, 3x3 = 9 predicted values ​​can be obtained. Then, based on these 9 predicted values ​​and the pre-set horizontal and vertical slope operators, the horizontal and vertical slopes of the i-th point are determined.

[0686] JPEG2024216632000057.jpg21168

[0687] JPEG2024216632000058.jpg21168

[0688] The embodiments of this application do not limit the specific values ​​of the horizontal gradient operator and the vertical gradient operator.

[0689] Based on the above steps, the encoding side can determine the horizontal and vertical slopes of the i-th point, and then determine the angle value of the i-th point based on the horizontal and vertical slopes of the i-th point.

[0690] For example, the arctangent value of the ratio of the vertical slope to the horizontal slope at the i-th point is determined as the angle value of the i-th point. For example, the angle value of the i-th point is determined by equation (8).

[0691] The encoding side may determine the angle value of the i-th point by a method other than determining the angle value of the i-th point using the above equation (8). For example, the encoding side may adjust the angle value determined by the above equation (8) to obtain the angle value of the i-th point.

[0692] On the encoding side, the angle value of each of the R points is determined using the method described above, and then step S203-A2 is performed to determine the intra prediction mode corresponding to the prediction block based on the angle values ​​of the R points.

[0693] The embodiments of this application do not limit the specific method for determining the intra-prediction mode corresponding to a prediction block based on the angular values ​​of R points.

[0694] In some embodiments, the encoding side selects angle value 1 with the highest number of recurrences from the angle values ​​of R points, matches this angle value 1 with the predicted angle of a conventional intra-prediction mode to obtain the intra-prediction mode corresponding to angle value 1, and determines the intra-prediction mode corresponding to angle value 1 as the intra-prediction mode corresponding to the prediction block.

[0695] In some embodiments, step S203-A2 above includes the following steps S203-A21 and S203-A22. S203-A21 determines the intra-prediction mode corresponding to R points based on the angle values ​​of R points. S203-A22 determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction modes corresponding to R points.

[0696] In this implementation method, the encoding side determines the intra-prediction mode corresponding to each of the R points based on the angle value of each point. For example, for each of the R points, the angle value of that point is matched with the predicted angle of a conventional intra-prediction mode, and the intra-prediction mode corresponding to the angle value of that point is obtained. In this way, the intra-prediction mode corresponding to each of the R points can be obtained.

[0697] Then, based on the intra-prediction mode corresponding to each of these R points, the intra-prediction mode corresponding to the prediction block is determined.

[0698] In one possible implementation, the intra-prediction mode with the highest recurrence rate among the intra-prediction modes corresponding to each of the R points is determined as the intra-prediction mode corresponding to the prediction block.

[0699] In another possible implementation, steps S203-A22 above include the following steps: S203-A221 determines the gradient amplitude value corresponding to R points based on the horizontal and vertical gradients of R points. S203-A222 determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction mode and gradient amplitude values ​​corresponding to R points.

[0700] In this implementation method, the encoding side determines the gradient amplitude value corresponding to each of the R points based on the horizontal and vertical gradients of each of the R points determined above.

[0701] In the embodiments of this application, the specific method by which the encoding side determines the gradient amplitude value corresponding to each of the R points is the same. For the sake of explanation, we will take the example of determining the gradient amplitude value corresponding to the i-th point among the R points.

[0702] The embodiments of this application do not limit the specific method by which the encoding side determines the gradient amplitude value corresponding to the i-th point based on the horizontal and vertical gradients of the i-th point.

[0703] For example, the encoding side multiplies the horizontal gradient and vertical gradient of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.

[0704] For example, the encoding side adds the absolute value of the horizontal gradient and the absolute value of the vertical gradient of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.

[0705] As an example, the encoding side determines the gradient amplitude value corresponding to the i-th point based on the following equation (9).

[0706] Based on the steps above, the encoding side can determine the gradient amplitude value corresponding to each of the R points. Next, the encoding side performs steps S203-A222 above and determines the intra-prediction mode corresponding to the prediction block based on the intra-prediction mode and gradient amplitude value corresponding to the R points.

[0707] In one example, the intra-prediction mode corresponding to the point with the maximum gradient amplitude among the R points is determined as the intra-prediction mode corresponding to the prediction block.

[0708] In another example, for any of the R points, the gradient amplitude value corresponding to that point is accumulated in the corresponding intra-prediction mode, and the cumulative gradient amplitude value of the intra-prediction modes corresponding to the R points is obtained. Among the intra-prediction modes corresponding to the R points, the intra-prediction mode with the largest cumulative gradient amplitude value is determined as the intra-prediction mode corresponding to the prediction block.

[0709] For example, as shown in Figure 27, the gradient amplitude values ​​corresponding to each of the R points are accumulated in the corresponding intra-prediction mode. For instance, if the intra-prediction modes corresponding to points 1 and 2 among the R points are both intra-prediction mode 1, the gradient amplitude values ​​corresponding to points 1 and 2 are accumulated in the gradient amplitude value corresponding to intra-prediction mode 1. The gradient amplitude value histogram shown in Figure 27 can be obtained in the same manner. In this way, the intra-prediction mode with the maximum accumulated gradient amplitude value in this gradient amplitude value histogram can be determined as the intra-prediction mode corresponding to the prediction block. For example, in Figure 27, the intra-prediction mode corresponding to the dark-colored accumulated gradient amplitude value is determined as the intra-prediction mode corresponding to the prediction block.

[0710] In some embodiments, if the gradient amplitude values ​​corresponding to R points are all 0, the first intra-prediction mode is determined as the intra-prediction mode corresponding to the prediction block. That is, if the gradient amplitude values ​​corresponding to all of the R points are 0, it means that the horizontal and vertical gradients of each of the R points are both 0, and in this case, the pre-set first intra-prediction mode can be determined as the intra-prediction mode corresponding to the prediction block.

[0711] The embodiments of this application do not limit the type of the first intra-prediction mode described above.

[0712] As an example, the first intra prediction mode described above is the PLANAR mode.

[0713] Based on the above steps, the encoding side determines the intra-prediction mode corresponding to the prediction block, and then determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block.

[0714] The embodiments of this application do not limit the specific method by which the encoding side determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block.

[0715] In some embodiments, the encoding side searches for an image block with the same intra-prediction mode as the prediction block from the encoded image blocks surrounding the prediction block, based on the intra-prediction mode corresponding to the prediction block, and then determines the transformation kernel corresponding to that image block as the transformation kernel corresponding to the current block.

[0716] In some embodiments, the step in step S203-B above, which determines the transformation kernel corresponding to the current block based on the intra-prediction mode corresponding to the prediction block, includes the following steps: S203-B1 obtains the correspondence between the intra-prediction mode and the transformation kernel group, and one transformation kernel group contains at least one type of transformation kernel. S203-B2 searches for the first transformation kernel group corresponding to the intra-prediction mode of the prediction block in the correspondence relationship. S203-B3 determines the translation kernel corresponding to the current block from the first translation kernel group.

[0717] In the embodiments of this application, there is a correspondence between the intra-prediction mode and the transformation kernel group. Based on this, the encoding side determines the intra-prediction mode corresponding to the prediction block, and then obtains the correspondence between the pre-configured intra-prediction mode and the transformation kernel group.

[0718] As an example, Table 17 shows the correspondence between the intra prediction mode and the conversion kernel group.

[0719] Note that Table 17 above represents only one type of correspondence between the intra-prediction mode and the conversion kernel group in the embodiment of this application, and the correspondence between the intra-prediction mode and the conversion kernel group in the embodiment of this application is shown in Table 1. 7 This includes, but is not limited to, the items shown.

[0720] Each translation kernel group includes at least one type of translation kernel.

[0721] The encoding side obtains the correspondence between intra-prediction modes and transformation kernel groups as shown in Table 17, and then, based on the intra-prediction mode corresponding to the prediction block, searches for the transformation kernel group corresponding to the intra-prediction mode corresponding to the prediction block in the correspondence between intra-prediction modes and transformation kernel groups, and sets this transformation kernel group as the first transformation kernel group. For example, the intra-prediction mode corresponding to the prediction block is an angle prediction mode in 64 angular directions, as shown in Table 17 above. 7 Upon querying, it is found that there are 4 transformation kernel groups corresponding to the angle prediction modes in these 64 angular directions. In this way, the encoding side determines the transformation kernel corresponding to the current block from at least one transformation kernel included in transformation kernel group 4.

[0722] For example, if the first group of translation kernels contains one translation kernel, that translation kernel is determined to be the translation kernel corresponding to the current block.

[0723] For example, if the first transformation kernel group includes multiple types of transformation kernels, the encoding side determines the type of transformation kernel corresponding to the current block, and further determines the transformation kernel of that type in the first transformation kernel group as the transformation kernel corresponding to the current block.

[0724] Here, the following methods are possible, but are not limited to, for determining the type of transformation kernel that corresponds to the current block on the encoding side.

[0725] In one example, the type of transformation kernel corresponding to the current block is the default type. In this way, the encoding side determines the default type as the type of transformation kernel corresponding to the current block.

[0726] In another example, the encoder writes the transformation kernel type corresponding to the current block into the code stream. The encoder then obtains the transformation kernel type corresponding to the current block by encoding the code stream.

[0727] Based on the above, in the embodiment of this application, the encoding side determines the predicted block of the current block using an interpolation filter prediction mode, further determines the conventional intra-prediction mode corresponding to the predicted block, and determines the transform kernel corresponding to the current block based on the conventional intra-prediction mode corresponding to the predicted block. That is, in the embodiment of this application, the conventional intra-prediction mode derived based on the interpolation filter prediction is used to select the NSPT (Non-separable primary transform) and LFNST (Low Frequency non-separable secondary transform) transform kernel groups, and the determined transform kernel is matched to the characteristics of the current block to improve the accuracy of transform kernel determination. When the reconstruction value of the current block is determined using this accurately determined transform kernel, the accuracy of reconstruction value determination is improved, and the encoding accuracy of the current block can be improved. In addition, in the embodiment of this application, when determining the transform kernel of the current block by the conventional prediction mode corresponding to the predicted block, it is not necessary to individually specify the transform kernel, saving codewords and further improving the video encoding effect.

[0728] In the embodiments of this application, the encoding side determines the predicted block of the current block and the transformation kernel corresponding to the current block based on the above steps. In this way, the encoding side can obtain the residual block of the current block based on the predicted block of the current block and the current block. For example, it can obtain the residual block of the current block by subtracting the predicted block of the current block from the current block. Next, the residual block of the current block is transformed based on the determined transformation kernel to obtain the transformation coefficients of the current block. Next, the transformation coefficients are directly encoded to obtain a code stream. Alternatively, the transformation coefficients are quantized to obtain quantization coefficients, and the quantization coefficients are encoded to obtain a code stream.

[0729] In some embodiments, the current block described above is either a luminance block or a chromaticity block, that is, in the implementations of this application, predictions can be made for both luminance blocks and chromaticity blocks using the interpolation filter prediction mode according to the implementations of this application.

[0730] In some embodiments, if the current block is a luminance block, the prediction mode for this current block is an interpolation filter prediction mode, and the chromaticity block corresponding to the current block employs a direct derivation mode DM, then the PLANAR mode or the intra-prediction mode corresponding to the prediction block is determined as the prediction mode for the chromaticity block.

[0731] The video encoding method according to the embodiment of this application first determines a reference region and interpolation filter for the current block when predicting the current block, determines filter coefficients based on the reference region, and performs parallel prediction on at least two pixel points in the current block using the interpolation filter based on the filter coefficients to obtain a predicted block for the current block. A transformation kernel corresponding to the current block is determined, and the current block is encoded based on the transformation kernel and the predicted block to obtain a code stream. In other words, in the embodiment of this application, when performing interpolation filter prediction on the current block using the interpolation filter, parallel prediction is performed on at least two points in the current block to improve prediction speed and further improve encoding efficiency.

[0732] Figures 10-29 are merely examples of this application and should not be interpreted as limiting this application.

[0733] While preferred embodiments of this application have been described in detail above with reference to the attached drawings, this application is not limited to the specific details of the above embodiments. Within the scope of the technical idea of ​​this application, various simple modifications are possible to the technical solutions of this application, and these simple modifications fall within the scope of protection of this application. For example, the individual specific technical features described in the above specific embodiments can be combined by any suitable means without contradiction, and in order to avoid unnecessary repetition, this application does not describe any further possible combinations. Also, for example, various different embodiments of this application can be combined in any way as long as they do not contradict the idea of ​​this application, and should likewise be considered to be within the disclosure of this application.

[0734] Furthermore, in the various method embodiments of this application, the magnitude of the number of each process does not indicate the order of execution, and the order of execution of each process should be determined by its function and internal logic, and it should be understood that no limitation is intended in any way to the implementation processes of the embodiments of this application. Also, in the embodiments of this application, the term "and / or" is merely a relation to describe the related objects, and it means that three relationships may exist. Specifically, A and / or B can mean three cases: A exists alone, A and B exist simultaneously, or B exists alone. Also, in this application, the symbol " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0735] The above describes in detail the method embodiment of this application with reference to Figures 10 to 29. The following describes in detail the apparatus embodiment of this application with reference to Figures 30 to 31.

[0736] Figure 30 is a schematic block diagram of a video decoding device according to one embodiment of the present application, which is applied to the video decoder described above.

[0737] As shown in Figure 30, the video decoding device 10 is A coefficient determination unit 11 determines the reference region and interpolation filter of the current block, and determines the filter coefficients of the interpolation filter based on the reference region, A prediction unit 12 determines the predicted block of the current block by performing parallel predictions on at least two pixel points in the current block using the interpolation filter based on the filter coefficients, The system includes a reconstruction unit 13 that determines a transformation kernel corresponding to the current block and determines a reconstruction block for the current block based on the transformation kernel corresponding to the current block and the predicted block.

[0738] In some embodiments, the prediction unit 12 specifically performs parallel interpolation filter predictions on pixel points on the same diagonal of the current block using the interpolation filter along the diagonal direction, based on the filter coefficients, in order to obtain a predicted block of the current block.

[0739] In some embodiments, the prediction unit 12 specifically obtains a predicted block of the current block by performing parallel interpolation filter predictions on pixel points on the same diagonal of the current block, starting from the upper left corner of the current block and moving along the diagonal direction, using the interpolation filter, based on the filter coefficients.

[0740] In some embodiments, the diagonal direction includes at least one of the directions from the upper right to the lower left and from the lower left to the upper right.

[0741] In some embodiments, the prediction unit 12 specifically determines the predicted values ​​of M pixel points on the k-th diagonal of the current block in parallel using the interpolation filter based on the filter coefficients, where k and M are both positive integers, and obtains the predicted value of the current block based on the predicted values ​​of the pixel points on each diagonal of the current block.

[0742] In some embodiments, the prediction unit 12 specifically determines in parallel the pixel values ​​of N positions corresponding to each of the M pixel points based on the shape of the interpolation filter, and determines in parallel the predicted values ​​of the M pixel points based on the filter coefficients and the pixel values ​​of N positions corresponding to each of the M pixel points.

[0743] In some embodiments, the coefficient determination unit 11 specifically determines a first reconstruction region around the current block, determines a pixel-average reconstruction value based on the reconstruction value of the first reconstruction region, removes the mean value from the reconstruction value of the pixel points in the reference region based on the pixel-average reconstruction value, uses the pixel values ​​of the pixel points with the mean value removed in the reference region as input to the interpolation filter, slides the interpolation filter within the reference region, and obtains the filter coefficients of the interpolation filter.

[0744] In some embodiments, the coefficient determination unit 11 is specifically for determining the pixel average reconstruction value based on the current block shape and the reconstruction value of the first reconstruction region.

[0745] In some embodiments, the coefficient determination unit 11 specifically determines a first region from the upper reconstruction region and the left reconstruction region based on the shape of the current block, determines the average reconstruction value of the first region based on the reconstruction value of the first region, and determines the pixel average reconstruction value based on the average reconstruction value of the first region.

[0746] In some embodiments, the coefficient determination unit 11 specifically determines the upper reconstruction region as the first region when the width of the current block shape is greater than the height, or determines the left reconstruction region as the first region when the height of the current block shape is greater than the width, or determines the upper reconstruction region and the left reconstruction region as the first region when the height of the current block shape is equal to the width.

[0747] In some embodiments, the horizontal and vertical sliding steps of the interpolation filter in the reference region are different.

[0748] In some embodiments, at least one of the horizontal sliding step and vertical sliding step of the interpolation filter in the reference region is greater than a predetermined step.

[0749] In some embodiments, the prediction unit 12 specifically performs mean removal in parallel on the pixel values ​​at N positions corresponding to each of the M pixel points based on the pixel mean reconstruction value to obtain mean-removed pixel values ​​at N positions corresponding to each of the M pixel points, and then determines the predicted values ​​of the M pixel points in parallel based on the filter coefficients and the mean-removed pixel values ​​at N positions corresponding to each of the M pixel points.

[0750] In some embodiments, the prediction unit 12 specifically subtracts the average pixel reconstruction value from the pixel values ​​at N positions of any of the M pixel points to obtain the average-removed pixel values ​​at N positions of the pixel point.

[0751] In some embodiments, the prediction unit 12 specifically determines a second reconstruction region around the current block, determines the maximum and minimum reconstruction values ​​of the second reconstruction region, obtains first predicted values ​​for the M pixel points in parallel based on the average filtered pixel values ​​of N positions corresponding to each of the M pixel points, the filter coefficients, and the pixel average reconstruction value, and determines the predicted values ​​for the M pixel points in parallel based on the first predicted values ​​for the M pixel points, the maximum reconstruction value, and the minimum reconstruction value.

[0752] In some embodiments, the prediction unit 12 specifically obtains a second predicted value for any of the M pixel points by multiplying the average removed pixel value of N positions of the pixel point by the filter coefficient, and then obtains a first predicted value for the pixel point by adding the second predicted value to the pixel average reconstruction value.

[0753] In some embodiments, the prediction unit 12 specifically determines, for any of the M pixel points, the first predicted value of the pixel point if the first predicted value of the pixel point is greater than the minimum reconstruction value and less than the maximum reconstruction value; or determines the minimum reconstruction value of the pixel point if the first predicted value of the pixel point is less than or equal to the minimum reconstruction value; or determines the maximum reconstruction value of the pixel point if the first predicted value of the pixel point is greater than or equal to the maximum reconstruction value.

[0754] In some embodiments, the coefficient determination unit 11 further determines whether the current block allows the use of the interpolation filter prediction mode before determining the reference region and interpolation filter of the current block, and if the current block allows the use of the interpolation filter prediction mode, it determines the reference region and interpolation filter of the current block.

[0755] In some embodiments, the coefficient determination unit 11 is specifically for determining that if the current block is in the first row of the current CTU, the current block does not allow the use of the interpolation filter prediction mode.

[0756] In some embodiments, the coefficient determination unit 11 specifically determines, based on the type of the current image, whether the current block allows the use of the interpolation filter prediction mode.

[0757] In some embodiments, the coefficient determination unit 11 is specifically for determining that if the current image is not an intra-prediction image, the current block does not allow the use of the interpolation filter prediction mode.

[0758] In some embodiments, the coefficient determination unit 11 specifically determines that if the size of the current block is smaller than a predetermined size, the current block does not allow the use of the interpolation filter prediction mode.

[0759] In some embodiments, the coefficient determination unit 11 specifically decodes the code stream to obtain first information indicating whether or not a template matching-based technique is enabled, and, based on the first information, determines whether or not the current block allows the use of the interpolation filter prediction mode.

[0760] In some embodiments, the coefficient determination unit 11 is specifically for determining that the current block does not allow the use of the interpolation filter prediction mode if the first information indicates that the template matching-based technique is not enabled.

[0761] In some embodiments, the coefficient determination unit 11 specifically decodes the code stream to obtain second information indicating whether the current sequence is permitted to make predictions using the interpolation filter prediction mode, if the first information indicates that the template matching-based technique is enabled, and determines, based on the second information, whether the current block is permitted to use the interpolation filter prediction mode.

[0762] In some embodiments, the reconstruction unit 13 specifically determines an intra-prediction mode corresponding to the prediction block and, based on the intra-prediction mode corresponding to the prediction block, determines a transformation kernel corresponding to the current block.

[0763] In some embodiments, the reconstruction unit 13 specifically determines the angular values ​​of R (a positive integer) points in the prediction block and determines the intra-prediction mode corresponding to the prediction block based on the angular values ​​of the R points.

[0764] In some embodiments, the reconfiguration unit 13 specifically obtains a correspondence between intra-prediction modes and transformation kernel groups, where each transformation kernel group includes at least one type of transformation kernel, searches for a first transformation kernel group corresponding to the intra-prediction mode of the prediction block in the correspondence, and determines the transformation kernel corresponding to the current block from the first transformation kernel group.

[0765] In some embodiments, the coefficient determination unit 11 is specifically for determining the reference region of the current block from a predetermined P (a positive integer greater than 1) reference regions.

[0766] In some embodiments, the coefficient determination unit 11 is specifically for determining the interpolation filter for the current block from a predetermined Q (a positive integer greater than 1) interpolation filters.

[0767] It should be understood that the apparatus embodiment can correspond to the method embodiment, and similar descriptions can refer to the method embodiment. To avoid redundancy, the explanation is omitted here. Specifically, the apparatus 10 shown in Figure 30 can perform the decoding-side decoding method of the embodiment of this application, and the above and other operations and / or functions of each unit in the apparatus 10 are for realizing the corresponding flows in each method, such as the above decoding-side decoding method, respectively, and for the sake of brevity, the explanation is omitted here.

[0768] Figure 31 is a schematic block diagram of a video encoding device according to one embodiment of the present invention, which is applied to the encoder described above.

[0769] As shown in Figure 31, the video encoding device 20 is A coefficient determination unit 21 determines the reference region and interpolation filter of the current block, and determines the filter coefficients of the interpolation filter based on the reference region, A prediction unit 22 determines the predicted block of the current block by performing parallel predictions on at least two pixel points in the current block using the interpolation filter based on the filter coefficients, The encoding unit 23 may include a unit that determines a conversion kernel corresponding to the current block, and then encodes the current block to obtain a code stream based on the conversion kernel corresponding to the current block and the predicted block.

[0770] In some embodiments, the prediction unit 22 specifically performs parallel interpolation filter predictions on pixel points on the same diagonal of the current block using the interpolation filter along the diagonal direction, based on the filter coefficients, in order to obtain a predicted block of the current block.

[0771] In some embodiments, the prediction unit ...

Claims

1. The steps include determining the reference region and interpolation filter of the current block, and determining the filter coefficients of the interpolation filter based on the reference region, The steps include: determining the predicted block of the current block by making a prediction for at least one pixel point in the current block using the interpolation filter based on the filter coefficients; A video decoding method characterized by comprising the steps of determining a transformation kernel corresponding to the current block, and determining a reconstruction block of the current block based on the transformation kernel corresponding to the current block and the predicted block.

2. The step of determining the predicted block of the current block by making a prediction for at least one pixel point in the current block using the interpolation filter based on the filter coefficients is as follows: The method according to claim 1, characterized by including the step of performing interpolation filter predictions for pixel points on the same diagonal of the current block using the interpolation filter along the diagonal direction based on the filter coefficients, in order to obtain a predicted block of the current block.

3. The step of obtaining a predicted block of the current block by performing interpolation filter predictions on pixel points on the same diagonal of the current block using the interpolation filter along the diagonal direction based on the filter coefficients is as follows: The step of obtaining a predicted block of the current block is to perform interpolation filter prediction on pixel points on the same diagonal of the current block using the interpolation filter, starting from the upper left corner of the current block and moving along the diagonal direction, based on the filter coefficients, The method according to claim 2, characterized in that the diagonal direction includes at least one of the directions from the upper right to the lower left and from the lower left to the upper right.

4. The step of obtaining a predicted block of the current block by performing interpolation filter predictions on pixel points on the same diagonal of the current block using the interpolation filter, starting from the upper left corner of the current block and moving along the diagonal direction, is as follows: A step in which, for M pixel points on the k-th diagonal of the current block, the interpolation filter is used to determine the predicted values ​​of the M pixel points based on the filter coefficients, wherein k and M are both positive integers. The method according to claim 3, characterized by comprising the step of obtaining a predicted value of the current block based on the predicted values ​​of the pixel points on each diagonal in the current block.

5. The step of determining the predicted values ​​of the M pixel points using the interpolation filter based on the filter coefficients is as follows: The steps include determining the pixel values ​​of N positions corresponding to each of the M pixel points based on the shape of the interpolation filter, The method according to claim 4, characterized by comprising the step of determining a predicted value for the M pixel points based on the filter coefficients and the pixel values ​​of N positions corresponding to each of the M pixel points.

6. Before the step of determining the reference region and interpolation filter of the current block, The process further includes determining whether the current block allows the use of the interpolation filter prediction mode, The step of determining the reference region and interpolation filter of the current block is: The method according to claim 1, further comprising the step of determining the reference region and interpolation filter of the current block if the current block allows the use of the interpolation filter prediction mode.

7. The step of determining whether the current block allows the use of the interpolation filter prediction mode is: If the current block is in the first row of the current CTU, the step includes determining that the current block does not allow the use of the interpolation filter prediction mode, or The step of determining whether the current block allows the use of the interpolation filter prediction mode is: The method according to claim 6, comprising the step of determining whether the current block allows the use of the interpolation filter prediction mode based on the type of the current image.

8. The step of determining whether the current block allows the use of the interpolation filter prediction mode based on the type of the current image is: The method according to claim 7, further comprising the step of determining that if the current image is not an intra-prediction image, the current block does not allow the use of the interpolation filter prediction mode.

9. The step of determining whether the current block allows the use of the interpolation filter prediction mode is: The method according to claim 6, characterized in that, if the size of the current block is smaller than a predetermined size, it is determined that the current block does not allow the use of the interpolation filter prediction mode.

10. The step of determining the translation kernel corresponding to the current block is: The steps include determining the intra-prediction mode corresponding to the prediction block, The method according to claim 1, characterized by comprising the step of determining a transformation kernel corresponding to the current block based on an intra prediction mode corresponding to the prediction block.

11. The step of determining the intra prediction mode corresponding to the prediction block is: The steps include determining the angle values ​​of R (positive integer) points in the prediction block, The steps include determining an intra-prediction mode corresponding to the prediction block based on the angle values ​​of the R points, or The step of determining the transformation kernel corresponding to the current block based on the intra prediction mode corresponding to the prediction block is: A step of obtaining the correspondence between intra prediction modes and transformation kernel groups, wherein one transformation kernel group includes at least one type of transformation kernel, In the aforementioned correspondence, the steps include searching for a first transformation kernel group corresponding to the intra-prediction mode of the prediction block, The method according to claim 10, further comprising the step of determining a translation kernel corresponding to the current block from the first translation kernel group.

12. The step of determining the reference region of the current block is: The method according to claim 1, characterized by comprising the step of determining the reference region of the current block from a predetermined P (a positive integer greater than 1) reference regions.

13. The step of determining the interpolation filter for the current block is: The method according to claim 1, characterized by comprising the step of determining the interpolation filter for the current block from a predetermined Q (a positive integer greater than 1) interpolation filters.

14. The steps include determining the reference region and interpolation filter of the current block, and determining the filter coefficients of the interpolation filter based on the reference region, The steps include: determining the predicted block of the current block by making a prediction for at least one pixel point in the current block using the interpolation filter based on the filter coefficients; A video encoding method characterized by comprising the steps of determining a transformation kernel corresponding to the current block, encoding the current block based on the transformation kernel corresponding to the current block and the predicted block, and obtaining a code stream.

15. A computer-readable storage medium for storing computer programs and code streams, When the aforementioned computer program is executed by the processor, The steps include determining the reference region and interpolation filter of the current block, and determining the filter coefficients of the interpolation filter based on the reference region, The steps include: determining the predicted block of the current block by making a prediction for at least one pixel point in the current block using the interpolation filter based on the filter coefficients; A computer-readable storage medium characterized by having the processor perform the steps of determining a translation kernel corresponding to the current block, encoding the current block based on the translation kernel corresponding to the current block and the predicted block, and obtaining a code stream to generate the code stream.