Video coding and decoding method, device, equipment, system and storage medium
Patent Information
- Application Number
- CN202380097396.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2025-12-09
AI Technical Summary
Existing interpolation filter prediction methods have low prediction efficiency in video encoding and decoding, which affects encoding and decoding performance.
A parallel prediction method is used to simultaneously predict at least two pixels in the current block, an interpolation filter is used to determine the filter coefficient, and prediction is performed based on the coefficient to improve the efficiency of interpolation filter prediction.
Improved video encoding and decoding performance, improved prediction speed and efficiency, and improved encoding and decoding performance.
Smart Images

Figure CN121100525A_ABST
Abstract
Description
Video encoding and decoding method, device, equipment, system, and storage medium Technical Field
[0001] The present application relates to the field of video coding and decoding technology, and in particular to a video coding and decoding method, apparatus, device, system, and storage medium. Background Art
[0002] Digital video technology can be incorporated into a variety of video devices, such as digital televisions, smartphones, computers, e-readers, or video players. With the development of video technology, the amount of video data included has increased. To facilitate the transmission of video data, video devices implement video compression technology to enable more efficient transmission or storage of video data.
[0003] Due to the temporal or spatial redundancy in videos, prediction can eliminate or reduce it, improving compression efficiency. To improve prediction, interpolation filtering prediction methods are sometimes used for predictive compression. However, current interpolation filtering prediction methods suffer from low prediction efficiency, resulting in poor video encoding and decoding performance.
[0004] Summary of the Invention
[0005] The embodiments of the present application provide a video encoding and decoding method, apparatus, device, system, and storage medium. When using interpolation filter prediction for prediction, parallel prediction can be performed, thereby improving the prediction efficiency of the interpolation filter prediction mode and enhancing the encoding and decoding performance.
[0006] In a first aspect, the present application provides a video decoding method, applied to a decoder, comprising:
[0007] Determining a reference area and an interpolation filter of a current block, and determining a filter coefficient of the interpolation filter based on the reference area;
[0008] Based on the filter coefficients, use the interpolation filter to perform parallel prediction on at least two pixels in the current block to determine a prediction block of the current block;
[0009] A transform kernel corresponding to the current block is determined, and a reconstructed block of the current block is determined based on the transform kernel corresponding to the current block and the prediction block.
[0010] In a second aspect, an embodiment of the present application provides a video encoding method, applied to an encoder, comprising:
[0011] Determining a reference area and an interpolation filter of a current block, and determining a filter coefficient of the interpolation filter based on the reference area;
[0012] Based on the filter coefficients, use the interpolation filter to perform parallel prediction on at least two pixels in the current block to determine a prediction block of the current block;
[0013] A transform kernel corresponding to the current block is determined, and based on the transform kernel corresponding to the current block and the prediction block, the current block is encoded to obtain a code stream.
[0014] In a third aspect, the present application provides a video decoding device for executing the method of the first aspect or its respective implementations. Specifically, the device includes a functional unit for executing the method of the first aspect or its respective implementations.
[0015] In a fourth aspect, the present application provides a video encoding device for executing the method of the second aspect or its respective implementations. Specifically, the device includes a functional unit for executing the method of the second aspect or its respective implementations.
[0016] In a fifth aspect, a video decoder is provided, comprising a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the first aspect or its respective implementations.
[0017] In a sixth aspect, a video encoder is provided, comprising a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the second aspect or its respective implementations.
[0018] In a seventh aspect, a video encoding and decoding system is provided, comprising a video encoder and a video decoder. The video decoder is configured to execute the method of the first aspect or its respective implementations, and the video encoder is configured to execute the method of the second aspect or its respective implementations.
[0019] In an eighth aspect, a chip is provided for implementing the method described in any one of the first and second aspects above, or their respective implementations. Specifically, the chip includes a processor configured to load and execute a computer program from a memory, causing a device equipped with the chip to perform the method described in any one of the first and second aspects above, or their respective implementations.
[0020] In a ninth aspect, a computer-readable storage medium is provided for storing a computer program, which enables a computer to execute the method of any one of the first to second aspects or their respective implementations.
[0021] In a tenth aspect, a computer program product is provided, comprising computer program instructions, which enable a computer to execute the method of any one of the first to second aspects or their respective implementations.
[0022] In an eleventh aspect, a computer program is provided, which, when executed on a computer, enables the computer to execute the method in any one of the first to second aspects or their respective implementations.
[0023] Based on the above technical solution, this application proposes an interpolation filtering prediction method. When predicting a current block, the reference area and interpolation filter of the current block are first determined. Based on the reference area, the filter coefficients are determined. Based on the filter coefficients, the interpolation filter is used to perform parallel prediction on at least two pixels in the current block to obtain a prediction block for the current block. The transform kernel corresponding to the current block is determined, and based on the transform kernel and the prediction block, the reconstructed value of the current block is determined. In other words, in an embodiment of the present application, when using the interpolation filter to perform interpolation filtering prediction on the current block, at least two points in the current block are predicted in parallel, which improves the prediction speed and thus improves the encoding and decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] FIG1 is a schematic block diagram of a video encoding and decoding system according to an embodiment of the present application;
[0025] FIG2 is a schematic block diagram of a video encoder according to an embodiment of the present application;
[0026] FIG3 is a schematic block diagram of a video decoder according to an embodiment of the present application;
[0027] FIG4A is a schematic diagram of intra-frame prediction;
[0028] FIG4B is a schematic diagram of intra-frame prediction;
[0029] 5A-5I are schematic diagrams of intra-frame prediction;
[0030] FIG6 is a schematic diagram of an intra-frame prediction mode;
[0031] FIG7 is a schematic diagram of an intra-frame prediction mode;
[0032] FIG8 is a schematic diagram of an intra-frame prediction mode;
[0033] Figure 9 is a schematic diagram of the CCCM principle;
[0034] FIG10 is a schematic flow chart of a video decoding method according to an embodiment of the present application;
[0035] FIG11 is a schematic diagram showing the position of the current block in the current image;
[0036] Figure 12 is a schematic diagram of the reconstruction area;
[0037] 13A to 13C are schematic diagrams of several reference areas;
[0038] 14A to 14G are schematic diagrams of several interpolation filter shapes;
[0039] FIG15 is a schematic diagram of shapes of several interpolation filters involved in embodiments of the present application;
[0040] FIG16 is a schematic diagram of shapes of several interpolation filters involved in an embodiment of the present application;
[0041] FIG17 is a schematic diagram of shapes of several interpolation filters involved in embodiments of the present application;
[0042] 18A and 18B are schematic diagrams of shapes of several interpolation filters involved in embodiments of the present application;
[0043] FIG19 is a schematic diagram of shapes of several interpolation filters involved in an embodiment of the present application;
[0044] FIG20A shows a sliding step size of an interpolation filter;
[0045] FIG20B is a schematic diagram of a first reconstruction area;
[0046] FIG21 is a schematic diagram showing the movement of interpolation filters of different shapes within different types of reference areas;
[0047] 22A and 22B are schematic diagrams of using an interpolation filter to perform interpolation prediction on a current block;
[0048] FIG23 is a schematic diagram of performing interpolation prediction on a current block along a diagonal direction according to an embodiment of the present application;
[0049] 24A to 24C are schematic diagrams of several directions of diagonal lines;
[0050] FIG25 is a schematic diagram of an intra-frame prediction mode;
[0051] FIG26 is a schematic diagram of determining horizontal gradient and vertical gradient;
[0052] FIG27 is a histogram of gradient amplitude values;
[0053] FIG28 is a flow chart of a video encoding method according to an embodiment of the present application;
[0054] FIG29 is a schematic diagram of a process for determining a prediction mode according to an embodiment of the present application;
[0055] FIG30 is a schematic block diagram of a video decoding device provided in an embodiment of the present application;
[0056] FIG31 is a schematic block diagram of a video encoding device provided in an embodiment of the present application;
[0057] FIG32 is a schematic block diagram of an electronic device provided in an embodiment of the present application;
[0058] Figure 33 is a schematic block diagram of the video encoding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] The present application can be applied to the field of image coding and decoding, the field of video coding and decoding, the field of hardware video coding and decoding, the field of dedicated circuit video coding and decoding, the field of real-time video coding and decoding, etc. For example, the solution of the present application can be combined with an audio and video coding standard (AVS), such as the H.264 / audio and video coding (AVC) standard, the H.265 / high efficiency video coding (HEVC) standard, and the H.266 / versatile video coding (VVC) standard. Alternatively, the solution of the present application can be combined with other proprietary or industry standards and operated, and the standards include ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), including scalable video coding (SVC) and multi-view video coding (MVC) extensions. It should be understood that the technology of this application is not limited to any specific coding standard or technology.
[0060] For ease of understanding, the video encoding and decoding system involved in the embodiment of the present application is first introduced with reference to FIG1 .
[0061] FIG1 is a schematic block diagram of a video encoding and decoding system involved in an embodiment of the present application. It should be noted that FIG1 is only an example, and the video encoding and decoding system of the embodiment of the present application includes but is not limited to that shown in FIG1. As shown in FIG1, the video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compressing) the video data to generate a code stream, and transmit the code stream to the decoding device. The decoding device decodes the code stream generated by the encoding device to obtain decoded video data.
[0062] The encoding device 110 of the embodiment of the present application can be understood as a device with a video encoding function, and the decoding device 120 can be understood as a device with a video decoding function, that is, the embodiment of the present application includes a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, car computers, etc.
[0063] In some embodiments, the encoding device 110 may transmit the encoded video data (eg, a code stream) to the decoding device 120 via a channel 130. The channel 130 may include one or more media and / or devices capable of transmitting the encoded video data from the encoding device 110 to the decoding device 120.
[0064] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to transmit the encoded video data directly to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media may also include wired communication media, such as one or more physical transmission lines.
[0065] In another example, channel 130 includes a storage medium that can store the video data encoded by encoding device 110. The storage medium includes various locally accessible data storage media, such as optical disks, DVDs, and flash memories. In this example, decoding device 120 can retrieve the encoded video data from the storage medium.
[0066] In another example, the channel 130 may include a storage server that can store the video data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded video data from the storage server. Alternatively, the storage server can store the encoded video data and transmit the encoded video data to the decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.
[0067] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113. The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.
[0068] In some embodiments, the encoding device 110 may further include a video source 111 in addition to the video encoder 112 and the input interface 113 .
[0069] The video source 111 may include at least one of a video acquisition device (eg, a video camera), a video archive, a video input interface, and a computer graphics system, wherein the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data.
[0070] The video encoder 112 encodes the video data from the video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream contains the coding information of the picture or picture sequence in the form of a bitstream. The coding information may include the coded picture data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. The SPS may contain parameters that apply to one or more sequences. The PPS may contain parameters that apply to one or more pictures. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.
[0071] The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data may also be stored in a storage medium or a storage server for subsequent reading by the decoding device 120.
[0072] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122 .
[0073] In some embodiments, the decoding device 120 may further include a display device 123 in addition to the input interface 121 and the video decoder 122 .
[0074] The input interface 121 includes a receiver and / or a modem and can receive the encoded video data via the channel 130 .
[0075] The video decoder 122 is configured to decode the encoded video data to obtain decoded video data, and transmit the decoded video data to the display device 123 .
[0076] The decoded video data is displayed on the display device 123. The display device 123 may be integrated with the decoding device 120 or external to the decoding device 120. The display device 123 may include various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0077] In addition, Figure 1 is only an example, and the technical solution of the embodiment of the present application is not limited to Figure 1. For example, the technology of the present application can also be applied to unilateral video encoding or unilateral video decoding.
[0078] The following is an introduction to the video encoding framework involved in the embodiments of the present application.
[0079] FIG2 is a schematic block diagram of a video encoder according to an embodiment of the present application. It should be understood that the video encoder 200 can be used to perform lossy compression or lossless compression on an image. The lossless compression can be visually lossless or mathematically lossless.
[0080] The video encoder 200 can be applied to image data in a luminance and chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents brightness (Luma), Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance (Chroma) for describing color and saturation. For example, in terms of color format, 4:2:0 means that every 4 pixels have 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that every 4 pixels have 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 represents full pixel display (YYYYCbCrCbCrCbCrCbCr).
[0081] For example, the video encoder 200 reads video data, and for each frame of the video data, divides the frame into a number of coding tree units (CTUs). In some examples, CTB may be referred to as a "tree block", "largest coding unit" (LCU) or "coding tree block" (CTB). Each CTU may be associated with a pixel block of equal size within the image. Each pixel may correspond to a luminance (luminance or luma) sample and two chrominance (chroma) samples. Therefore, each CTU may be associated with a luminance sample block and two chroma sample blocks. The size of a CTU is, for example, 128×128, 64×64, 32×32, etc. A CTU can be further divided into a number of coding units (CUs) for encoding. The CU may be a rectangular block or a square block. A CU can be further divided into prediction units (PUs) and transform units (TUs), allowing for separation of coding, prediction, and transform, and greater flexibility in processing. In one example, a CTU is divided into CUs using a quadtree, and a CU is divided into TUs and PUs using a quadtree.
[0082] The video encoder and video decoder can support various PU sizes. Assuming that the size of a particular CU is 2N×2N, the video encoder and video decoder can support PU sizes of 2N×2N or N×N for intra-frame prediction, and support symmetric PUs of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-frame prediction. The video encoder and video decoder can also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-frame prediction.
[0083] In some embodiments, as shown in FIG2 , the video encoder 200 may include a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filter unit 260, a decoded image buffer 270, and an entropy coding unit 280. It should be noted that the video encoder 200 may include more, fewer, or different functional components.
[0084] Optionally, in this application, the current block may be referred to as the current coding unit (CU) or the current prediction unit (PU), etc. The prediction block may also be referred to as a predicted image block or an image prediction block, and the reconstructed image block may also be referred to as a reconstructed block or an image reconstructed image block.
[0085] In some embodiments, the prediction unit 210 includes an inter-frame prediction unit 211 and an intra-frame prediction unit 212. Because there is a strong correlation between adjacent pixels in a video frame, intra-frame prediction is used in video coding and decoding technologies to eliminate spatial redundancy between adjacent pixels. Because there is a strong similarity between adjacent frames in a video, inter-frame prediction is used in video coding and decoding technologies to eliminate temporal redundancy between adjacent frames, thereby improving coding efficiency.
[0086] The inter-frame prediction unit 211 is used for inter-frame prediction. Inter-frame prediction includes motion estimation and motion compensation. It can refer to image information from different frames. Inter-frame prediction uses motion information to find a reference block from a reference frame and generates a prediction block based on the reference block to eliminate temporal redundancy. The frames used for inter-frame prediction can be P-frames and / or B-frames. P-frames refer to forward-predicted frames, while B-frames refer to bidirectionally predicted frames. Inter-frame prediction uses motion information to find a reference block from a reference frame and generates a prediction block based on the reference block. Motion information includes the reference frame list in which the reference frame is located, the reference frame index, and the motion vector. The motion vector can be integer pixel or fractional pixel. If the motion vector is fractional pixel, interpolation filtering is required to generate the required fractional pixel block in the reference frame. Here, the integer pixel or fractional pixel block in the reference frame found based on the motion vector is called a reference block. Some technologies directly use the reference block as the prediction block, while others further process the reference block to generate a prediction block. Reprocessing the reference block to generate a prediction block can also be understood as using the reference block as the prediction block and then processing the prediction block to generate a new prediction block.
[0087] The intra-frame prediction unit 212 only refers to the information of the same frame image to predict the pixel information in the current code image block to eliminate spatial redundancy. The frame used for intra-frame prediction can be an I frame.
[0088] Intra-frame prediction has multiple prediction modes. For example, the H-series international digital video coding standard H.264 / AVC has eight angular prediction modes and one non-angular prediction mode. H.265 / HEVC expands this to 33 angular prediction modes and two non-angular prediction modes. HEVC uses planar, DC, and 33 angular modes for a total of 35 intra-frame prediction modes. VVC uses planar, DC, and 65 angular modes for a total of 67 intra-frame prediction modes.
[0089] It should be noted that with the increase of angle modes, intra-frame prediction will be more accurate and more in line with the needs of high-definition and ultra-high-definition digital video development.
[0090] The residual unit 220 may generate a residual block for the CU based on the pixel blocks of the CU and the prediction blocks of the PUs of the CU. For example, the residual unit 220 may generate the residual block for the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel blocks of the CU and the corresponding sample in the prediction blocks of the PUs of the CU.
[0091] The transform / quantization unit 230 may quantize the transform coefficients. The transform / quantization unit 230 may quantize the transform coefficients associated with the TUs of a CU based on a quantization parameter (QP) value associated with the CU. The video encoder 200 may adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.
[0092] The inverse transform / quantization unit 240 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficients to reconstruct a residual block from the quantized transform coefficients.
[0093] Reconstruction unit 250 may add samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by prediction unit 210 to generate a reconstructed image block associated with the TU. By reconstructing the sample blocks of each TU of a CU in this manner, video encoder 200 can reconstruct the pixel blocks of the CU.
[0094] The loop filter unit 260 is used to process the inverse transformed and inverse quantized pixels to compensate for distortion information and provide a better reference for subsequent coded pixels. For example, it can perform a deblocking filtering operation to reduce the blocking effect of pixel blocks associated with the CU.
[0095] In some embodiments, the loop filtering unit 260 includes a deblocking filtering unit and a sample adaptive offset / adaptive loop filtering (SAO / ALF) unit, wherein the deblocking filtering unit is used to remove blocking effects, and the SAO / ALF unit is used to remove ringing effects.
[0096] The decoded image buffer 270 may store the reconstructed pixel blocks. The inter-prediction unit 211 may use a reference image containing the reconstructed pixel blocks to perform inter-prediction on PUs of other images. In addition, the intra-prediction unit 212 may use the reconstructed pixel blocks in the decoded image buffer 270 to perform intra-prediction on other PUs in the same image as the CU.
[0097] The entropy coding unit 280 may receive the quantized transform coefficients from the transform / quantization unit 230. The entropy coding unit 280 may perform one or more entropy coding operations on the quantized transform coefficients to generate entropy-coded data.
[0098] FIG3 is a schematic block diagram of a video decoder according to an embodiment of the present application.
[0099] 3 , the video decoder 300 includes an entropy decoding unit 310, a prediction unit 320, an inverse quantization / transformation unit 330, a reconstruction unit 340, a loop filter unit 350, and a decoded picture buffer 360. It should be noted that the video decoder 300 may include more, fewer, or different functional components.
[0100] The video decoder 300 may receive a bitstream. The entropy decoding unit 310 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, the entropy decoding unit 310 may parse the entropy-encoded syntax elements in the bitstream. The prediction unit 320, the inverse quantization / transform unit 330, the reconstruction unit 340, and the loop filter unit 350 may decode the video data based on the syntax elements extracted from the bitstream, thereby generating decoded video data.
[0101] In some embodiments, the prediction unit 320 includes an intra-frame prediction unit 322 and an inter-frame prediction unit 321 .
[0102] The intra-prediction unit 322 may perform intra-prediction to generate a prediction block for the PU. The intra-prediction unit 322 may use an intra-prediction mode to generate a prediction block for the PU based on pixel blocks of spatially neighboring PUs. The intra-prediction unit 322 may also determine the intra-prediction mode for the PU based on one or more syntax elements parsed from the codestream.
[0103] The inter-frame prediction unit 321 may construct a first reference picture list (List 0) and a second reference picture list (List 1) based on syntax elements parsed from the codestream. In addition, if a PU is encoded using inter-frame prediction, the entropy decoding unit 310 may parse the motion information of the PU. The inter-frame prediction unit 321 may determine one or more reference blocks for the PU based on the motion information of the PU. The inter-frame prediction unit 321 may generate a prediction block for the PU based on the one or more reference blocks of the PU.
[0104] The inverse quantization / transform unit 330 may inversely quantize (ie, dequantize) the transform coefficients associated with the TU. The inverse quantization / transform unit 330 may use the QP value associated with the CU of the TU to determine the degree of quantization.
[0105] After inverse quantizing the transform coefficients, the inverse quantization / transform unit 330 may apply one or more inverse transforms to the inverse quantized transform coefficients in order to generate a residual block associated with the TU.
[0106] The reconstruction unit 340 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, the reconstruction unit 340 can add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU to obtain a reconstructed image block.
[0107] The loop filtering unit 350 may perform a deblocking filtering operation to reduce blocking artifacts of pixel blocks associated with a CU.
[0108] The video decoder 300 may store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 may use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.
[0109] The basic process of video encoding and decoding is as follows: At the encoder end, a frame of image is divided into blocks. For the current block, the prediction unit 210 uses intra-frame prediction or inter-frame prediction to generate a prediction block for the current block. The residual unit 220 calculates a residual block based on the predicted block and the original block of the current block. This residual block is the difference between the predicted block and the original block of the current block. This residual block is also called residual information. This residual block undergoes transformation and quantization by the transform / quantization unit 230, removing information that is insensitive to the human eye and eliminating visual redundancy. Optionally, the residual block before transformation and quantization by the transform / quantization unit 230 is called a time-domain residual block, and the time-domain residual block after transformation and quantization by the transform / quantization unit 230 is called a frequency residual block or a frequency-domain residual block. The entropy coding unit 280 receives the quantized change coefficients output by the change quantization unit 230 and performs entropy coding on these quantized change coefficients to output a bitstream. For example, the entropy coding unit 280 can eliminate character redundancy based on the target context model and probability information of the binary bitstream.
[0110] At the decoding end, the entropy decoding unit 310 can parse the code stream to obtain the prediction information, quantization coefficient matrix, etc. of the current block. The prediction unit 320 uses intra-frame prediction or inter-frame prediction for the current block based on the prediction information to generate a prediction block for the current block. The inverse quantization / transformation unit 330 uses the quantization coefficient matrix obtained from the code stream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block. The reconstruction unit 340 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed blocks constitute a reconstructed image, and the loop filtering unit 350 performs loop filtering on the reconstructed image based on the image or block to obtain a decoded image. The encoding end also requires similar operations as the decoding end to obtain a decoded image. The decoded image can also be called a reconstructed image, and the reconstructed image can be used as a reference frame for inter-frame prediction for subsequent frames.
[0111] It should be noted that the block division information determined by the encoder, as well as mode information or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering, etc., are carried in the bitstream when necessary. The decoder parses the bitstream and analyzes the existing information to determine the same block division information, prediction, transform, quantization, entropy coding, loop filtering, etc. mode information or parameter information as the encoder, thereby ensuring that the decoded image obtained by the encoder and the decoder are identical.
[0112] The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. This application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.
[0113] In the embodiments of the present application, the current block can be the current coding unit (CU) or the current prediction unit (PU). Due to the need for parallel processing, an image can be divided into slices, and slices within the same image can be processed in parallel, meaning that there is no data dependency between them. "Frame" is a commonly used term, generally understood to mean that a frame is an image. In the application, the term "frame" can also be replaced with "image" or "slice."
[0114] Intra-frame prediction typically uses both angular and non-angular modes to predict the current coding block to obtain a predicted block. Based on the rate-distortion information calculated between the predicted block and the original block, the optimal prediction mode for the current coding unit is selected and transmitted to the decoder via the bitstream. The decoder parses the prediction mode, generates a predicted image for the current decoding block, and superimposes it with the residual pixels transmitted via the bitstream to obtain a reconstructed image. Intra-frame prediction uses the already coded and decoded reconstructed pixels surrounding the current block as reference pixels to predict the current block. Figure 4A illustrates intra-frame prediction. As shown in Figure 4A, the current block is 4x4 in size. The pixels in the row to the left and the column above the current block serve as reference pixels for the current block. Intra-frame prediction uses these reference pixels to predict the current block. These reference pixels may all be available, meaning they have already been coded and decoded. However, some may not be available. For example, if the current block is at the leftmost edge of the frame, the reference pixels to the left of the current block may not be available. Alternatively, when encoding and decoding the current block, the lower left portion of the current block has not yet been coded and decoded, making the reference pixels to the lower left unavailable. In the case where reference pixels are unavailable, available reference pixels or certain values or methods may be used for filling, or no filling may be performed.
[0115] FIG4B is a schematic diagram of intra prediction. As shown in FIG4B , the multiple reference line (MRL) intra prediction method can use more reference pixels to improve encoding and decoding efficiency. For example, four reference rows / columns are used as reference pixels of the current block.
[0116] Furthermore, intra-frame prediction has multiple prediction modes. Figures 5A-5I are schematic diagrams of intra-frame prediction. As shown in Figures 5A-5I, intra-frame prediction for 4x4 blocks in H.264 mainly includes nine modes. Among them, mode 0, shown in Figure 5A, copies the pixels above the current block vertically to the current block as the prediction value. Mode 1, shown in Figure 5B, copies the reference pixels to the left of the current block horizontally as the prediction value. Mode 2 (DC), shown in Figure 5C, uses the average of the eight points A to D and I to L as the prediction value for all points. Modes 3 to 8, shown in Figures 5D-5I, copy the reference pixels to the corresponding positions in the current block at a specific angle. Because some positions in the current block cannot correspond exactly to the reference pixels, it may be necessary to use a weighted average of the reference pixels, or interpolated sub-pixels of the reference pixels.
[0117] In addition, there are Plane, Planar and other modes. With the development of technology and the expansion of blocks, there are more and more angular prediction modes. Figure 6 is a schematic diagram of intra-frame prediction modes. As shown in Figure 6, the intra-frame prediction modes used by HEVC include Planar, DC and 33 angular modes, totaling 35 prediction modes. Figure 7 is a schematic diagram of intra-frame prediction modes. As shown in Figure 7, the intra-frame modes used by VVC include Planar, DC and 65 angular modes, totaling 67 prediction modes. Figure 8 is a schematic diagram of intra-frame prediction modes. As shown in Figure 8, VS3 uses DC, Plane, Bilinear, PCM and 62 angular modes, totaling 66 prediction modes.
[0118] There are also some technologies that improve prediction, such as improving the pixel-by-pixel interpolation of reference pixels and filtering the predicted pixels. For example, the multiple intra-frame prediction filter (MIPF) in AVS3 uses different filters to generate prediction values for different block sizes. For pixels at different positions in the same block, one filter is used to generate prediction values for pixels closer to the reference pixel, and another filter is used to generate prediction values for pixels farther from the reference pixel. Technologies that filter predicted pixels, such as the intra-frame prediction filter (IPF) in AVS3, can use reference pixels to filter the predicted values.
[0119] In some embodiments, in current video encoding and decoding, an adaptive loop filter (ALF) technology is used in a loop filter unit, for example, ALF technology is used to filter the reconstructed image to obtain a final decoded image.
[0120] The following is an introduction to the adaptive loop filter (ALF) technology.
[0121] ALF is a filter in the loop filter. It is designed based on the principle of Wiener filter and is a filter that minimizes the error between the target sample and the input sample. In the loop filter, the target sample is the original image and the input is the reconstructed image.
[0122] Before using ALF for filtering, the filter coefficients must be determined first.
[0123] For example, by constructing the Wienerhof equation shown in formula (1), and solving the Wienerhof equation, the filter coefficients of the interpolation filter can be obtained:
[0124] in, Represents a range of the current 2D image. For example, if the input of the filter is a reconstructed image, then Represents the reconstruction area around the current block in the reconstructed image. r is The position of a sample within, for example, the coordinates of the sample at position r can be expressed as (x, y). o[r] is the original pixel value of the sample at position r, t[r] is the pixel value to be filtered at position r, for example, if the input of the filter is a reconstructed image, then t[r] is also called the reconstructed value of the pixel at position r in the reconstructed image. c=[c0,c1,…,c N-1 ] T is the filter coefficient of the adaptive filter, {p0,p1,…,p N-1} is the relative position difference between the N positions corresponding to position r and position r.
[0125] In the above formula (1), except for the filter coefficients c=[c0,c1,…,c N-1 ] T Except for , all other data are known, so the filter coefficients of the filter can be obtained by solving the above formula (1).
[0126] In one example, the filter coefficients of the filter can be obtained by solving the above Wienerhof equation through Cholesky decomposition of the autocorrelation coefficient matrix.
[0127] After the filter coefficients are determined based on the above formula (1), the samples to be filtered are filtered using the following formula (2) to obtain the filtered samples:
[0128] Among them, t[r]′ is the pixel value after filtering at position r, p nis the relative position difference between the nth position and position r in the N positions corresponding to position r, t[r+p n ] means r+p n The pixel value to be filtered at position .
[0129] The convolutional cross component model (CCCM) predicts chrominance pixels using reconstructed luminance pixels. Its advantage is that the CCCM filter coefficients can be obtained from reconstructed pixels at the decoder, eliminating the overhead of storing filter coefficients in the bitstream as with ALF. As shown in Figure 9, the CCCM coefficients are calculated from the reconstructed pixels surrounding the chrominance block to be predicted and the reconstructed pixels surrounding the luminance block at the corresponding position of the chrominance block.
[0130] In order to improve the compression performance of video, an interpolation filtering prediction mode is proposed. The interpolation filtering prediction mode determines the filter coefficient of the interpolation filter through the reconstructed area around the current block. Based on the filter coefficient, the interpolation filter is used to perform interpolation filtering prediction on each point in the current block to obtain the predicted value of each point in the current block, and then obtain the predicted block of the current block.
[0131] In related technologies, when using an interpolation filter to predict each pixel in the current block, interpolation filtering is performed on each pixel one by one. That is, after the interpolation filtering prediction for the previous pixel is completed, the interpolation filtering prediction for the next pixel is performed, and the prediction value of the previous pixel is used when performing the interpolation filtering prediction for the next pixel. Therefore, it can be seen that when performing interpolation filtering prediction on the current block, the related technology performs prediction point by point, and can only predict one point at a time, resulting in low prediction efficiency, which in turn affects the overall encoding and decoding performance of the video.
[0132] In order to solve the above technical problems, the embodiment of the present application performs parallel prediction on the pixels in the current block when using the interpolation filter prediction mode to predict the current block, thereby improving the prediction efficiency and enhancing the encoding and decoding performance of the video.
[0133] 10 , the video decoding method provided in the embodiment of the present application is introduced by taking the decoding end as an example.
[0134] FIG10 is a flow chart of a video decoding method according to an embodiment of the present application, which is applied to the video decoders shown in FIG1 and FIG3. As shown in FIG10, the method according to the embodiment of the present application includes:
[0135] S101 : Determine a reference area and an interpolation filter of a current block, and determine a filter coefficient of the interpolation filter based on the reference area.
[0136] When decoding the current block, the decoder decodes the bitstream to obtain the quantization coefficients for the current block, dequantizes the quantization coefficients to obtain the transform coefficients for the current block, and de-transforms the transform coefficients to obtain the residual value for the current block. Next, the prediction mode for the current block is determined, and based on the prediction mode, the predicted value for the current block is determined. Based on the predicted value and the residual value for the current block, the reconstructed value for the current block is obtained.
[0137] In some embodiments, the current block also becomes a block to be predicted.
[0138] In the embodiment of the present application, the decoding end first determines the prediction mode of the current block.
[0139] In some embodiments, the decoding end determines the prediction mode of the current block in at least the following ways:
[0140] In Method 1, the encoder determines the prediction mode for the current block. For example, from the candidate prediction modes consisting of the traditional prediction mode and the interpolation filter prediction mode shown in Figure 6 or Figure 7, the candidate prediction mode with the lowest cost is selected as the prediction mode for the current block. Next, the encoder adds information indicating the prediction mode for the current block to the bitstream. The decoder then decodes the bitstream to obtain the information indicating the prediction mode for the current block. Based on this information, the decoder determines the prediction mode for the current block and uses the intra-frame prediction mode to predict the current block, obtaining a predicted value for the current block.
[0141] For example, if the terminal device determines that the prediction mode for the current block is the traditional prediction mode, the prediction mode index for the current block is written into the bitstream as the indication of the prediction mode. The decoder obtains the prediction mode index by decoding the bitstream and, based on the index, determines the prediction mode for the current block from the traditional prediction modes shown in Figures 6 or 7.
[0142] In method 2, the encoder constructs a candidate list of intra-frame prediction modes and selects the intra-frame prediction mode for the current block from the candidate list. It should be noted that the candidate list includes the interpolation filter prediction mode. Next, the encoder writes the sequence number (or index number) of the intra-frame prediction mode of the current block in the candidate list into the bitstream. In this way, the decoder determines the sequence number of the intra-frame prediction mode of the current block in the candidate list of intra-frame prediction modes by decoding the bitstream. Simultaneously, based on the same method as the encoder, it constructs a candidate list of intra-frame prediction modes (it should be noted that the constructed candidate list of intra-frame prediction modes includes the interpolation filter prediction mode). Then, based on the sequence number of the intra-frame prediction mode of the current block in the candidate list of intra-frame prediction modes, it determines the intra-frame prediction mode for the current block from the constructed candidate list of intra-frame prediction modes. Finally, the determined intra-frame prediction mode of the current block is used to predict the current block to obtain a predicted value for the current block.
[0143] Method 3: The encoding end constructs a candidate list of intra-frame prediction modes, which includes an interpolation filter prediction mode. Then, the intra-frame prediction mode of the current block is selected from the candidate list of intra-frame prediction modes. For example, the cost of each candidate prediction mode in the candidate list of intra-frame prediction modes on the template of the current block is determined, and then the intra-frame prediction mode of the current block is determined based on the cost. Correspondingly, the decoding end constructs a candidate list of intra-frame prediction modes based on the same method as the encoding end. The constructed candidate list of intra-frame prediction modes also includes an interpolation filter prediction mode. Then, the cost of each candidate prediction mode in the candidate list of intra-frame prediction modes on the template of the current block is determined, and then the intra-frame prediction mode of the current block is determined based on the cost. Finally, the intra-frame prediction mode of the current block is used to predict the current block to obtain the predicted value of the current block.
[0144] In mode 4, the encoder and decoder use the interpolation filter prediction mode by default to predict the current block.
[0145] In addition to determining whether the current block adopts the interpolation filtering prediction mode for prediction through the above-mentioned methods 1 to 4, the decoding end can also determine whether the current block adopts the interpolation filtering prediction mode through the following method 5.
[0146] In mode 5, the decoder decodes the bitstream to obtain third information indicating whether the current block is predicted using the interpolation filter prediction mode. If the decoder determines, based on the third information, that the current block is predicted using the interpolation filter prediction mode, it determines a reference area and an interpolation filter for the current block.
[0147] In this method 5, if the encoder determines that the current block adopts the interpolation filter prediction mode, the third information is written into the bitstream. In this way, the decoder obtains the third information by decoding the bitstream, and then determines whether the current block adopts the interpolation filter prediction mode for prediction based on the third information. If the third information indicates that the current block adopts the interpolation filter prediction mode for prediction, the decoder uses the interpolation filter prediction mode to predict the current block and obtains the prediction block of the current block. If the third information indicates that the current block does not adopt the interpolation filter prediction mode for prediction, the decoder skips the step of predicting the current block using the interpolation filter prediction mode, further determines the prediction mode of the current block, and predicts the current block using the determined prediction mode to obtain the prediction block of the current block.
[0148] The embodiment of the present application does not limit the specific form of expression of the above-mentioned third information, which can be any indication information that can indicate whether the current block adopts the interpolation filtering prediction mode for prediction.
[0149] In one example, the third information can be represented as intra_eip_flag, so that different values of intra_eip_flag can be used to determine whether the current block is predicted using the interpolation filtering prediction mode. For example, when intra_eip_flag = 0, it indicates that the current block is not predicted using the interpolation filtering prediction mode, and when intra_eip_flag = 1, it indicates that the current block is predicted using the interpolation filtering prediction mode. In this way, the encoder writes the preset flag intra_eip_flag into the bitstream, and the decoder determines the prediction mode of the current block by decoding the value of the preset flag intra_eip_flag. For example, when the preset flag intra_eip_flag = 1, it indicates that the prediction mode of the current block is the interpolation filtering prediction mode, and the decoder then uses the interpolation filtering prediction mode to predict the current block.
[0150] In some embodiments, the use conditions of the interpolation filter prediction mode are limited. Based on this, before determining the reference area and interpolation filter of the current block, it is determined whether the current image block is allowed to be predicted using the interpolation filter prediction mode.
[0151] The embodiment of the present application does not limit the specific method of determining whether the current image block is allowed to be predicted using the interpolation filter prediction mode, that is, it does not limit the specific use conditions of the interpolation filter prediction mode.
[0152] In some embodiments, to improve the prediction accuracy of the interpolation filter prediction mode, the interpolation filter prediction mode is used for some blocks that meet the requirements, while the interpolation filter prediction mode is not used for some blocks that do not meet the requirements. Based on this, before decoding the bitstream and obtaining the third information, the decoding end needs to determine whether the position of the current block in the current image meets the preset position requirements and whether the size of the current block meets the preset block size. If it is determined that the position of the current block in the current image meets the preset position requirements and the size of the current block meets the preset block size, the bitstream is decoded to obtain the third information.
[0153] The embodiment of the present application does not impose any restrictions on the preset position requirements and the prediction block size, which are determined based on actual needs.
[0154] In one example, as shown in FIG11 , assuming that the position of the upper left corner of the current image is (0, 0), and the position of the upper left corner of the current block is (x, y), where the preset position requirements are that the x value of the current block is greater than or equal to a first preset value XX, and the y value of the current block is greater than or equal to a second preset value YY.
[0155] The embodiment of the present application does not limit the specific values of the first preset value and the second preset value.
[0156] Exemplarily, the first preset value and the second preset value are the same.
[0157] Exemplarily, the first preset value and the second preset value are both 13, that is, when the distance from the upper edge line of the current block to the upper edge line of the current image is greater than or equal to 13 pixel rows, and the distance from the left edge line of the current block to the left edge line of the current image is greater than or equal to 13 pixel columns, it indicates that the position of the current block in the current image meets the preset position requirements.
[0158] In one example, continuing to refer to Figure 11, assuming that the width of the current block is W and the height of the current block is H, the preset block size requirement is that the width W of the current block is less than or equal to the third preset value A, and the height H of the current block is less than or equal to the fourth preset value B.
[0159] The embodiment of the present application does not limit the specific values of the third preset value and the fourth preset value.
[0160] Exemplarily, the third preset value and the fourth preset value are the same.
[0161] Exemplarily, the third preset value and the fourth preset value are both 32, that is, when the width and height of the current block are both less than or equal to 32, it indicates that the current block meets the preset block size requirement.
[0162] In an embodiment of the present application, before determining whether the current block is predicted using the interpolation filter prediction mode, the decoding end first determines whether the position of the current block in the current image meets the preset position requirement, and determines whether the size of the current block meets the preset block size requirement. If the position of the current block in the current image meets the preset position requirement, and the size of the current block meets the preset block size requirement, the code stream is decoded to obtain third information, and based on the third information, it is determined whether the current block is predicted using the interpolation filter prediction mode. For example, as shown in Figure 11, the distance from the upper edge of the current block to the upper edge of the current image is greater than or equal to 13 pixel rows, the distance from the left edge of the current block to the left edge of the current image is greater than or equal to 13 pixel columns, and the width and height of the current block are both less than or equal to 32, then the decoding end decodes the code stream to obtain the third information.
[0163] In some embodiments, the first preset value, the second preset value, the third preset value, and the fourth preset value are default values.
[0164] In some embodiments, the first preset value, the second preset value, the third preset value, and the fourth preset value are values decoded from the bitstream by the decoding end.
[0165] In some embodiments, if the position of the current block in the current image does not meet a preset position requirement, and / or the size of the current block does not meet a preset block size requirement, it is determined that the current block is not predicted using the interpolation filtering prediction mode.
[0166] In some embodiments, before determining whether the position of the current block in the current image meets the preset position requirement and determining whether the size of the current block meets the preset block size, the decoding end further includes: decoding the code stream to obtain second information, where the second information is used to indicate whether the current sequence is allowed to be predicted using the interpolation filtering prediction mode; if the second information indicates that the current sequence is allowed to be predicted using the interpolation filtering prediction mode, then determining whether the position of the current block in the current image meets the preset position requirement and determining whether the size of the current block meets the preset block size.
[0167] In an embodiment of the present application, a high-level syntax element, such as second information at the sequence level, is used to indicate whether the current sequence is allowed to be predicted using the interpolation filtering prediction mode. If the second information indicates that the current sequence is allowed to be predicted using the interpolation filtering prediction mode, the decoding end determines whether the position of the current block in the current image meets the preset position requirement and whether the size of the current block meets the preset block size requirement. Then, when it is determined that the position of the current block in the current image meets the preset position requirement and the size of the current block meets the preset block size requirement, the decoding end decodes the third information to determine whether the current block is predicted using the interpolation filtering prediction mode.
[0168] In some embodiments, if the second information indicates that the current sequence is not allowed to be predicted using the interpolation filtering prediction mode, the decoding end skips the above-mentioned steps of determining whether the position of the current block in the current image meets the preset position requirements, and determining whether the size of the current block meets the preset block size requirements, and skips the step of decoding the third information.
[0169] The embodiment of the present application does not limit the specific form of the second information, which can be any indication information that can indicate whether the current sequence is allowed to be predicted using the interpolation filtering prediction mode.
[0170] In one example, the second information may be represented by sps_eip_enabled_flag, so that different values of sps_eip_enabled_flag can be used to determine whether the current sequence is allowed to be predicted using the interpolation filtering prediction mode. For example, when sps_eip_enabled_flag = 0, it indicates that the current sequence is not allowed to be predicted using the interpolation filtering prediction mode, and when sps_eip_enabled_flag = 1, it indicates that the current sequence is allowed to be predicted using the interpolation filtering prediction mode.
[0171] Exemplarily, the second information is carried in a sequence parameter set (SPS), for example, as shown in Table 1:
[0172] Table 1
[0173] Wherein, sps_eip_enabled_flag represents the second information, and the sps_eip_enabled_flag is carried in seq_parameter_set_rbsp(). For example, when sps_eip_enabled_flag=0, it indicates that the current sequence is not allowed to be predicted using the interpolation filter prediction mode, and when sps_eip_enabled_flag=1, it indicates that the current sequence is allowed to be predicted using the interpolation filter prediction mode.
[0174] In some embodiments, embodiments of the present application may further include a general constraints information (GCI) flag to indicate whether interpolation filter prediction technology is used. Exemplarily, gci_no_eip_constraint_flag is used to indicate whether interpolation filter prediction technology is enabled for the current video. Exemplarily, as shown in Table 2, the gci_no_eip_constraint_flag is carried in the general constraints information general_constraints_info().
[0175] Table 2
[0176] As shown in Table 2, if gci_no_eip_constraint_flag = 1, it means that the interpolation filter prediction technology is not enabled for the current video, that is, the interpolation filter intra prediction technology at the sequence level must be 0 in all images, that is, it means that the interpolation filter intra prediction technology is not allowed in all sequences in the current video. If gci_no_eip_constraint_flag = 0, it means that the interpolation filter prediction technology is enabled for the current video, that is, the interpolation filter intra prediction technology at the sequence level must be 0 in all images.
[0177] As can be seen from the above, if the syntax elements of the embodiment of the present application include the high-level syntax elements gci_no_eip_constraint_flag and sps_eip_enabled_flag, as well as the block-level intra_eip_flag, the decoding end first decodes the high-level syntax elements, that is, first decodes gci_no_eip_constraint_flag. If gci_no_eip_constraint_flag = 0, it continues to decode sps_eip_enabled_flag. If sps_eip_enabled_flag = 1, it parses the syntax elements of the block.
[0178] Exemplary block-level syntax elements are shown in Table 3:
[0179] Table 3
[0180] In Table 3, cbWidth and cbHeight are the width and height of the current block, SIZE_A can be understood as the third preset value mentioned above, SIZE_B can be understood as the fourth preset value, XX can be understood as the first preset value, YY can be understood as the second preset value, and x0 and y0 represent the coordinate difference between the upper left corner of the current block and the upper left corner of the current image.
[0181] As can be seen from Table 3 above, if the second sequence-level information sps_eip_enabled_flag = 1, indicating that the current sequence allows the use of the interpolation filter prediction mode, then it is determined whether the position of the current block in the current image meets the preset position requirement, and whether the size of the current block meets the preset block size requirement. If it is determined that the position of the current block in the current image meets the preset position requirement and the size of the current block meets the preset block size requirement, the third information intra_eip_flag is decoded, and based on the decoded third information intra_eip_flag, it is determined whether the current block is predicted using the interpolation filter prediction mode.
[0182] As can be seen from the above, whether the current block adopts the interpolation filter prediction mode can be determined by high-level syntax, such as GCI, sequence level, frame level, slice level, block level, etc. It can also be determined by the size and position of the current block.
[0183] In some embodiments, using the interpolation filter prediction mode for smaller blocks increases computational cost and complexity. This is because the interpolation filter prediction mode in this application has a high computational complexity. If the interpolation filter prediction mode is also used for smaller blocks, this will increase the number of times the interpolation filter prediction mode is used during the entire image decoding, thereby increasing the computational cost and complexity of the image. Based on this, in this embodiment of the present application, the interpolation filter prediction mode is only allowed for slightly larger blocks. For example, the interpolation filter prediction mode is only allowed if the size of the current block is greater than or equal to a preset size. If the size of the current block is less than the preset size, the interpolation filter prediction mode is not allowed for the current block. This embodiment of the present application does not limit the specific value of the preset size. For example, the size of the current block being greater than or equal to the preset size can mean that the number of pixels in the current block is greater than or equal to a preset number, or at least one of the length and width of the current block is greater than or equal to a preset value, or the ratio of the length and width of the current block is greater than or equal to a preset ratio, etc.
[0184] In some embodiments, if the current block is in the first row of the current CTU, it is determined that the current block is not allowed to use the interpolation filter prediction mode. In other words, if the current block is predicted using the interpolation filter prediction mode, the current block is not in the first row of the current CTU.
[0185] In some embodiments, determining whether the current block is allowed to use the interpolation filter prediction mode is also related to the type of the current image. For example, for intra-frame prediction images (i.e., images that use intra-frame prediction during prediction), it is stipulated that the interpolation filter prediction mode can be used for prediction, while for inter-frame prediction images (i.e., images that use inter-frame prediction during prediction), the interpolation filter prediction mode is not allowed for prediction. Based on this, if the current image in which the current block is located is an intra-frame prediction image, it is determined that the current block is allowed to use the interpolation filter prediction mode for prediction. If the current image is not an intra-frame prediction image (e.g., an inter-frame prediction image), it is determined that the current block is not allowed to use the interpolation filter prediction mode for prediction.
[0186] In some embodiments, in the ECM reference software, in order to improve the encoding and decoding performance, a series of complex intra-frame prediction modes are introduced, such as: template-based intra prediction derivation mode (TIMD), decoder-side intra prediction derivation mode (DIMD), template-based multiple reference line intra prediction (TMRL), spatial geometrical partitioning mode (SGPM) and convolutional cross component model (CCCM). These copied intra-frame prediction modes are all intra-frame prediction modes based on template matching technology. The interpolation filter prediction mode of the embodiment of the present application also uses the information of the reconstructed area (which can be understood as the template area) during use. Therefore, in the embodiment of the present application, the interpolation filter prediction mode can also be classified as an intra-frame prediction mode based on template matching technology. Based on this, in the embodiment of the present application, unified identification information (such as the first information) is used to uniformly indicate the above-mentioned intra-frame prediction modes based on template matching technology. For example, if the first information indicates that the template matching-based technology is not enabled, it means that the aforementioned intra-frame prediction modes based on the template matching technology (i.e., TIMD, DIMD, TMRL, SGPM, TMRL, CCCM, and interpolation filter prediction mode) are not allowed to be used. If the first information indicates that the template matching-based technology is not enabled, it means that the aforementioned intra-frame prediction modes based on the template matching technology are allowed to be used, and then based on other information, the specific intra-frame prediction mode used for the current block is further determined.
[0187] Based on the above description, in an embodiment of the present application, determining whether the current block is allowed to use the interpolation filter prediction mode includes: decoding the code stream to obtain first information, the first information is used to indicate whether the template matching-based technology is enabled; and based on the first information, determining whether the current block is allowed to use the interpolation filter prediction mode. For example, if the first information indicates that the template matching-based technology is not enabled, then it is determined that the current block is not allowed to use the interpolation filter prediction mode for prediction. For another example, if the first information indicates that the template matching-based technology is enabled, the decoding end determines whether the current block is predicted using the interpolation filter prediction mode based on other information.
[0188] The embodiment of the present application does not limit the specific form of expression of the above-mentioned first information.
[0189] Exemplarily, the first information may be GCI, sequence-level, frame-level, slice-level, or block-level indication information.
[0190] In one example, if the first information is sequence-level indication information, the decoder decodes the bitstream to obtain the first information. If the first information indicates that the template matching technology is enabled, the decoder continues to decode the bitstream to obtain the second information (sps_eip_enabled_flag), and then determines whether the interpolation filter prediction mode is allowed for the current block based on the second information. If the first information indicates that the template matching technology is not enabled, the decoder directly determines that the interpolation filter prediction mode is not applicable for prediction of the current block, and skips the step of decoding the second information.
[0191] In some embodiments, the conditions for the decoder to determine whether the current block can be predicted using the interpolation filtering prediction mode include at least one of the following:
[0192] 1) Whether the current image is an intra-frame prediction image;
[0193] 2) Whether high-level syntax is allowed; optional, high-level syntax includes sequence level, frame level, slice level, block level, etc., refer to the above description for details;
[0194] 3) Whether the size and shape of the current block are allowed; refer to the above description for details;
[0195] 4) Whether the position of the current block is allowed; refer to the above description for details.
[0196] The above describes the specific process of determining whether the current block is predicted using the interpolation filtering prediction mode.
[0197] In an embodiment of the present application, if the decoding end determines that the current block is predicted using the interpolation filtering prediction mode, the interpolation filtering prediction mode is used to predict the current block to obtain a predicted value of the current block.
[0198] The following describes the process of using the interpolation filter prediction mode at the decoding end to predict the current block.
[0199] When the decoding end determines that the current block is predicted using the interpolation filtering prediction mode, it first determines the reference area and interpolation filter of the current block.
[0200] The following describes the specific process of determining the reference area of the current block at the decoding end.
[0201] In the embodiment of the present application, the reference area of the current block is part or all of the reconstructed area around the current block.
[0202] For example, as shown in FIG12 , the reconstructed area around the current block may include: an upper reconstructed area of the current block, a left reconstructed area of the current block, an upper right reconstructed area of the current block, a lower left reconstructed area of the current block, and an upper left reconstructed area of the current block. The block to be predicted in FIG12 is the current block.
[0203] The embodiment of the present application does not limit the specific shape and size of the reference area of the current block.
[0204] In one example, the reference area of the current block includes any one of an upper reconstruction area of the current block, a left reconstruction area of the current block, an upper right reconstruction area of the current block, a lower left reconstruction area of the current block, and an upper left reconstruction area of the current block. For example, the reference area of the current block is the upper reconstruction area of the current block, or the reference area of the current block is the left reconstruction area of the current block.
[0205] In one example, the reference area of the current block includes any two of the reconstruction areas above the current block, the reconstruction area to the left of the current block, the reconstruction area to the upper right of the current block, the reconstruction area to the lower left of the current block, and the reconstruction area to the upper left of the current block. For example, the reference area of the current block includes the reconstruction area above the current block and the reconstruction area to the left of the current block. For another example, the reference area of the current block includes the reconstruction area above the current block and the reconstruction area to the lower left of the current block.
[0206] In one example, the reference area of the current block includes any three reconstruction areas of the upper reconstruction area of the current block, the left reconstruction area of the current block, the upper right reconstruction area of the current block, the lower left reconstruction area of the current block, and the upper left reconstruction area of the current block. For example, the reference area of the current block includes the upper reconstruction area of the current block, the upper right reconstruction area of the current block, and the upper left reconstruction area of the current block. For another example, the reference area of the current block includes the left reconstruction area of the current block, the upper left reconstruction area of the current block, and the lower left reconstruction area of the current block.
[0207] In one example, the reference area of the current block includes any four reconstruction areas of the upper reconstruction area of the current block, the left reconstruction area of the current block, the upper right reconstruction area of the current block, the lower left reconstruction area of the current block, and the upper left reconstruction area of the current block. For example, the reference area of the current block includes the upper reconstruction area of the current block, the upper right reconstruction area of the current block, the upper left reconstruction area of the current block, and the left reconstruction area of the current block. For another example, the reference area of the current block includes the left reconstruction area of the current block, the upper left reconstruction area of the current block, the lower left reconstruction area of the current block, and the upper reconstruction area of the current block.
[0208] In one example, the reference area of the current block includes five reconstruction areas: an upper reconstruction area of the current block, a left reconstruction area of the current block, an upper right reconstruction area of the current block, a lower left reconstruction area of the current block, and an upper left reconstruction area of the current block.
[0209] In the embodiment of the present application, the decoding end determines a reference area of the current block from among the preset P reference areas.
[0210] In the embodiment of the present application, the specific manner in which the decoding end determines the reference region of the current block from the preset P reference regions includes but is not limited to the following:
[0211] In method 1, the reference area of the current block is a default area. For example, the encoding end and the decoding end default that the reference area of the current block includes at least one reconstruction area of the P reference areas: the upper reconstruction area of the current block, the left reconstruction area of the current block, the upper right reconstruction area of the current block, the lower left reconstruction area of the current block, and the upper left reconstruction area of the current block.
[0212] Method 2: The decoding end decodes the code stream to obtain fourth information, which is used to indicate the type of the reference area of the current block; based on the type of the reference area, the reference area of the current block is determined in the preset P reference areas, where P is a positive integer greater than 1.
[0213] In this implementation, the encoder determines a reference region for the current block from among P preset reference regions. For example, the encoder determines the coding costs corresponding to each of the P reference regions and selects the reference region with the lowest coding cost as the reference region for the current block. The encoder then indicates the type of the reference region with the lowest coding cost to the decoder via fourth information. The decoder then decodes the bitstream to obtain the fourth information and, based on the reference region type indicated by the fourth information, determines a reference region for the current block from among the P preset reference regions.
[0214] It should be noted that the types or shapes of the preset P reference areas are different.
[0215] The embodiment of the present application does not impose any specific limitation on the specific number and shape of the P reference areas.
[0216] In one example, the P reference regions include at least one of a first reference region, a second reference region, and a third reference region.
[0217] As shown in FIG13A , the first reference region includes the reconstruction regions above, to the upper right, to the left, to the upper left, and to the upper left of the current block. As shown in FIG13B , the second reference region includes the reconstruction regions above, to the upper right, and to the upper left of the current block. As shown in FIG13C , the third reference region includes the reconstruction regions to the left, to the upper left, and to the upper left of the current block. The block to be predicted in FIG13A through FIG13C is the current block.
[0218] The embodiment of the present application does not limit the specific form of the fourth information, as long as it is any indication information that can indicate the type of the reference area of the current block.
[0219] In an example, eip_ref_type is used to represent the fourth information. For example, different types of reference areas are indicated by the value of eip_ref_type.
[0220] For example, as shown in Table 4, the correspondence between the three reference areas shown in FIG. 13A and FIG. 13B and the eip_ref_type value is as follows:
[0221] Table 4
[0222] Based on Table 4 above, the decoder decodes the bitstream to obtain the fourth information eip_ref_type, and then determines the reference area of the current block based on the value of the fourth information eip_ref_type. For example, if eip_ref_type = 0, the reference area of the current block is determined to be the first reference area. As shown in FIG13A, the first reference area includes the reconstruction area above, to the upper right, to the left, to the upper left, and to the upper left of the current block. If eip_ref_type = 1, the reference area of the current block is determined to be the second reference area. As shown in FIG13B, the second reference area includes the reconstruction area above, to the upper right, and to the upper left of the current block. If eip_ref_type = 2, the reference area of the current block is determined to be the third reference area. As shown in FIG13C, the third reference area includes the reconstruction area above, to the upper right, and to the upper left of the current block.
[0223] It should be noted that the above description uses the three reference regions shown in Figures 13A to 13C as an example. The P reference regions in this embodiment of the present application also include other reference regions in addition to the three reference regions described above, and this embodiment of the present application does not limit this. The correspondence between the reference regions and the eip_ref_type values shown in Table 4 above can be adaptively adjusted according to the number of reference regions.
[0224] In some embodiments, the decoding end may adopt a truncated binary code decoding method to decode the code stream to obtain the fourth information.
[0225] For example, the correspondence between the truncated binary code, the eip_ref_type value, and the type of the reference area is shown in Table 5:
[0226] Table 5
[0227] In the embodiment of the present application, the decoding end may adopt an equal probability decoding method or a context model decoding method to decode the codeword of the truncated binary code.
[0228] In addition to using the above-mentioned method 1 or method 2 to determine the reference area of the current block, the decoding end can also use the following method 3 to determine the reference area of the current block.
[0229] Mode 3: Based on the shape of the current block, a reference area of the current block is determined from among P preset reference areas.
[0230] In this method 3, different reference areas are used to perform prediction on current blocks of different shapes to improve the accuracy of prediction.
[0231] For example, if the shape of the current block is a square, the first type of reference region is used.
[0232] For another example, if the shape of the current block is a rectangle with a width greater than a height, the second type of reference region is used.
[0233] For another example, if the shape of the current block is a rectangle with a width smaller than a height, the third type of reference region is used.
[0234] That is, in this embodiment of the present application, the correspondence between the P reference regions and the shape of the current block is preset. Thus, the decoding end can determine the reference region of the current block from the P reference regions based on the shape of the current block and the correspondence between the P reference regions and the shape of the current block.
[0235] The following describes the process of determining the interpolation filter of the current block at the decoding end.
[0236] In the embodiment of the present application, there is no limitation on the specific shape of the interpolation filter.
[0237] Illustratively, the interpolation filters provided in the embodiments of the present application include, but are not limited to, square interpolation filters and interpolation filters with a height smaller than a width.
[0238] For example, the square interpolation filter includes but is not limited to the 4X4 interpolation filter shown in Figure 14A.
[0239] For another example, interpolation filters that are taller than they are wide include, but are not limited to, the 5X3 interpolation filter shown in FIG. 14B , the 6X2 interpolation filter shown in FIG. 14D , and the 7X1 interpolation filter shown in FIG. 14G .
[0240] For another example, interpolation filters with a height smaller than a width include but are not limited to the 3X5 interpolation filter shown in FIG. 14C , the 2X6 interpolation filter shown in FIG. 14E , and the 1X7 interpolation filter shown in FIG. 14F .
[0241] It should be noted that in the above filter, the dark gray position represents the current position to be predicted, and the light gray position represents the input position of the interpolation filter, that is, {p0, p1, ..., p N-1}Location.
[0242] In the embodiment of the present application, the decoding end determines the interpolation filter of the current block from among the preset Q interpolation filters.
[0243] In the embodiment of the present application, the specific manner in which the decoding end determines the interpolation filter for the current block from among the preset Q interpolation filters includes but is not limited to the following:
[0244] In mode 1, the interpolation filter of the current block is a default interpolation filter. For example, the encoder and decoder default the interpolation filter of the current block to any one of the Q interpolation filters shown in Figures 14A to 14G. For example, the default interpolation filter is a 4x4 interpolation filter.
[0245] Method 2: The decoding end decodes the code stream to obtain fifth information, which is used to indicate the shape of the interpolation filter of the current block; based on the shape of the interpolation filter of the current block, the interpolation filter of the current block is determined from the preset Q interpolation filters, where Q is a positive integer greater than 1.
[0246] In this implementation, the encoder determines the interpolation filter for the current block from among Q preset interpolation filters. For example, the encoder determines the coding costs corresponding to each of the Q interpolation filters and selects the interpolation filter with the lowest coding cost as the interpolation filter for the current block. The shape of the interpolation filter with the lowest coding cost is then indicated to the decoder via fifth information. The decoder then decodes the bitstream to obtain the fifth information and, based on the shape of the interpolation filter indicated by the fifth information, determines the interpolation filter for the current block from among the Q preset interpolation filters.
[0247] It should be noted that the shapes of the preset Q interpolation filters are different.
[0248] The embodiments of the present application do not impose any specific restrictions on the specific number and shape of the Q interpolation filters. For example, the Q interpolation filters include at least one of a first interpolation filter, a second interpolation filter, and a third interpolation filter, where the first interpolation filter is a square interpolation filter, the second interpolation filter is a rectangular interpolation filter with a width greater than a height, and the third interpolation filter is a rectangular interpolation filter with a height greater than a width.
[0249] In one example, the Q interpolation filters include the plurality of interpolation filters in FIG. 14A to FIG. 14H .
[0250] The embodiment of the present application does not limit the specific form of the fifth information, as long as it is any indication information that can indicate the shape of the interpolation filter of the current block.
[0251] In an example, eip_filter_type is used to represent the fifth information. For example, the value of eip_filter_type is used to indicate interpolation filters of different shapes.
[0252] For example, if the Q interpolation filters are the five interpolation filters shown in FIG15 , as shown in Table 6, the correspondence between the five interpolation filters and the eip_filter_type values is:
[0253] Table 6
[0254] Based on Table 5 above, the decoder decodes the bitstream to obtain the fifth information eip_filter_type, and then determines the interpolation filter of the current block based on the value of the fifth information eip_filter_type. For example, if eip_filter_type = 0, the shape of the interpolation filter of the current block is determined to be 4x4. If eip_filter_type = 1, the shape of the interpolation filter of the current block is determined to be 3x5. If eip_filter_type = 2, the shape of the interpolation filter of the current block is determined to be 5x3. If eip_filter_type = 3, the shape of the interpolation filter of the current block is determined to be 2x6. If eip_filter_type = 4, the shape of the interpolation filter of the current block is determined to be 6x2.
[0255] In some embodiments, the decoding end may adopt a truncated binary code decoding method to decode the code stream to obtain the fifth information.
[0256] For example, if the preset Q interpolation filters include the five interpolation filters shown in FIG15 , the correspondence between the truncated binary code, the eip_filter_type value, and the shape of the interpolation filter is shown in Table 7:
[0257] Table 7
[0258] At this time, the five interpolation filter shapes shown in Table 7 and the three reconstruction region types shown in Table 5 provide a total of 15 combinations of interpolation filters and reconstruction regions.
[0259] In some embodiments, the decoding end can obtain the fifth information eip_filter_type by decoding the bitstream, and then determine the interpolation filter for the current block based on the shape of the interpolation filter indicated by the fifth information eip_filter_type, as shown in Table 7. Similarly, the decoding end can obtain the fourth information by decoding the bitstream, and then determine the reference area for the current block based on the value of the fourth information eip_ref_type, as shown in Table 5.
[0260] In an example, the syntax elements of the embodiment of the present application are shown in Table 8:
[0261] Table 8
[0262] As shown in Table 8, the decoder decodes the bitstream and first obtains the fifth sequence-level information, sps_eip_enabled_flag, which indicates whether the current sequence allows prediction using the interpolation filter prediction mode. Next, it determines whether the position of the current block in the current image meets preset position requirements and whether the size of the current block meets preset block size requirements. If the position and size of the current block in the current image meet the preset position and block size requirements, it decodes the third information, intra_eip_flag, which indicates whether the current block is predicted using the interpolation filter prediction mode. If the third information, intra_eip_flag = 1, indicates that the current block is predicted using the interpolation filter prediction mode, it decodes the bitstream to obtain the fourth information, eip_ref_type, and the fifth information, eip_filter_type. The fourth information, eip_ref_type, indicates the type of the reference region for the current block. Based on the value of the fourth information, eip_ref_type, the decoder can obtain the reference region for the current block by looking up the table. The fifth information eip_filter_type indicates the shape of the interpolation filter of the current block, and based on the value of the fifth information eip_filter_type, the interpolation filter of the current block is obtained by looking up a table.
[0263] In some embodiments, when the embodiment of the present application includes the seven interpolation filters shown in FIG16 , the correspondence between the truncated binary code, the eip_filter_type value, and the shape of the interpolation filter is shown in Table 9:
[0264] Table 9
[0265] At this time, the 7 interpolation filter shapes shown in Table 9 and the 3 reconstruction region types shown in Table 5 provide a total of 21 combinations of interpolation filters and reconstruction regions.
[0266] Similarly, the decoding end can obtain the reference area and interpolation filter of the current block by decoding the syntax shown in Table 8 and looking up Table 5 and the above Table 8.
[0267] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in FIG17 , the correspondence between the truncated binary code, the eip_filter_type value, and the shape of the interpolation filter is shown in Table 10:
[0268] Table 10
[0269] At this time, the three interpolation filter shapes shown in Table 10 and the three reconstruction region types shown in Table 5 provide a total of nine combinations of interpolation filters and reconstruction regions.
[0270] Similarly, the decoding end can obtain the reference area and interpolation filter of the current block by decoding the syntax shown in Table 8 and looking up Table 5 and the above Table 10.
[0271] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in FIG18A , the correspondence between the truncated binary code, the eip_filter_type value, and the shape of the interpolation filter is shown in Table 11:
[0272] Table 11
[0273] At this time, the three interpolation filter shapes shown in Table 11 and the three reconstruction region types shown in Table 5 provide a total of nine combinations of interpolation filters and reconstruction regions.
[0274] Similarly, the decoding end can obtain the reference area and interpolation filter of the current block by decoding the syntax shown in Table 8 and looking up Table 5 and the above Table 11.
[0275] In some embodiments, when the embodiment of the present application includes three interpolation filters as shown in FIG18B , the correspondence between the truncated binary code, the eip_filter_type value, and the shape of the interpolation filter is shown in Table 12:
[0276] Table 12
[0277] At this time, the three interpolation filter shapes shown in Table 12 and the three reconstruction region types shown in Table 5 provide a total of nine combinations of interpolation filters and reconstruction regions.
[0278] Similarly, the decoding end can obtain the reference area and interpolation filter of the current block by decoding the syntax shown in Table 8 and looking up Table 5 and the above Table 11.
[0279] Generally, using a filter with more taps for the same number of samples can achieve better interpolation effects. Compared to the 2x6 and 6x2 interpolation filters shown in Figure 18A, Figure 18B increases the number of taps in the filter, for example, expanding it to 2x8 and 8x2 interpolation filters. In fact, the 2x8 and 8x2 filters and the 4x4 filter all use 15 samples as input and 1 output, and their complexity is similar. Therefore, the interpolation filter shown in Figure 18B improves the interpolation effect without increasing complexity.
[0280] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in FIG19 , the correspondence between the truncated binary code, the eip_filter_type value, and the shape of the interpolation filter is shown in Table 13:
[0281] Table 13
[0282] At this time, the three interpolation filter shapes shown in Table 13 and the three reconstruction region types shown in Table 5 provide a total of nine combinations of interpolation filters and reconstruction regions.
[0283] Similarly, the decoding end can obtain the reference area and interpolation filter of the current block by decoding the syntax shown in Table 8 and looking up Table 5 and the above Table 13.
[0284] In addition to using the above-mentioned method 1 or method 2 to determine the interpolation filter of the current block, the decoding end may also use the following method 3 to determine the interpolation filter of the current block.
[0285] Mode 3: Based on the shape of the current block, an interpolation filter for the current block is determined from among Q preset interpolation filters.
[0286] In this method 3, different interpolation filters are used to perform prediction on current blocks of different shapes to improve the accuracy of prediction.
[0287] For example, if the shape of the current block is a square, an interpolation filter of the first shape is used.
[0288] For another example, if the shape of the current block is a rectangle with a width greater than a height, the interpolation filter of the second shape is used.
[0289] For another example, if the shape of the current block is a rectangle with a width smaller than a height, the interpolation filter of the third shape is used.
[0290] That is, in the embodiment of the present application, the correspondence between the Q interpolation filters and the shape of the current block is preset. In this way, the decoding end can determine the interpolation filter for the current block from the Q interpolation filters based on the shape of the current block and the correspondence between the Q interpolation filters and the shape of the current block.
[0291] The following describes how to determine the filter coefficients of the interpolation filter based on the reference area.
[0292] In the embodiment of the present application, the decoding end determines the filter coefficients of the interpolation filter in at least the following ways:
[0293] Method 1: Use the interpolation filter determined above to slide over the reference area of the current block to construct the Wienerhof equation. Then, solve the Wienerhof equation to obtain the filter coefficients of the interpolation filter.
[0294] In the process of sliding the interpolation filter in the reference area of the current block, the N positions corresponding to each position in the reference area are determined based on the shape of the interpolation filter. For example, for position r in the reference area, based on the shape of the interpolation filter, the N positions corresponding to position r are determined in the reference area. The pixel reconstruction values of these N positions are the input of the interpolation filter. The phase position difference between these N positions and position r is {p0, p1, ..., p N-1}, p N is a two-dimensional representation. In which {c0,c1,…,c N-1} is {p0,p1,…,p N-1}Interpolation filter coefficients at the position.
[0295] In one example, the interpolation filter is slid in the reference area of the current block to construct the Wienerhof equation, as shown in formula (3):
[0296] in, is the reference area of the current block, t[r+p n ] is r+p in the reference area n The pixel reconstructed value of the pixel at position r in the reference area is t[r].
[0297] Since the reference area of the current block is the reconstruction area, the above formula (3) contains the following: Except for , all other parameters are known, so the filter coefficients of the interpolation filter of the current block can be determined by solving the above formula (3).
[0298] In one example, the decoding end may solve the Wienerhof equation shown in the above formula (3) by Cholesky decomposition of the autocorrelation coefficient matrix to obtain the filter coefficients of the filter.
[0299] The embodiment of the present application does not limit the sliding step size of the interpolation filter within the reference area.
[0300] In one example, as shown in FIG20A , the horizontal sliding step size and the vertical sliding step size of the interpolation filter in the reference area are equal, both being 1 pixel.
[0301] In one example, the horizontal sliding step size and the vertical sliding step size of the interpolation filter in the reference area are not equal. For example, the horizontal sliding step size is 2 pixels and the vertical sliding step size is 1 pixel. For another example, the horizontal sliding step size is 1 pixel and the vertical sliding step size is 2 pixels.
[0302] In one example, at least one of the horizontal sliding step size and the vertical sliding step size of the interpolation filter within the reference region is greater than a preset step size. For example, the horizontal sliding step size is greater than the preset step size. For another example, the vertical sliding step size is greater than the preset step size. For another example, both the horizontal sliding step size and the vertical sliding step size are greater than the preset step size. This embodiment of the present application does not limit the specific value of the preset step size. For example, it can be 1, 2, 3, etc.
[0303] In the second method, the decoding end determines the filter coefficients through the following steps S101-A1 to S101-A4:
[0304] S101-A1, determining a first reconstruction area around a current block;
[0305] S101-A2, determining a pixel average reconstruction value based on the reconstruction value of the first reconstruction area;
[0306] S101-A3, based on the pixel average reconstruction value, removing the mean of the reconstructed values of the pixels in the reference area;
[0307] S101-A4: Using the pixel values of the pixels in the reference area after averaging as inputs of the interpolation filter, sliding the interpolation filter in the reference area to obtain filter coefficients of the interpolation filter.
[0308] In this second approach, the reference region is de-averaged, and the interpolation filter coefficients are determined based on the de-averaged reference region. Since the amount of data in the de-averaged reference region is reduced, determining the filter coefficients based on the de-averaged reference region can improve efficiency.
[0309] Specifically, the decoding end first determines a first reconstruction area, and the first reconstruction area can be any part of the reconstruction area around the current block.
[0310] In the embodiment of the present application, the decoding end determines the first reconstruction area around the current block in at least the following ways:
[0311] In mode 1, the decoding end determines a reconstruction area around the current block as the first reconstruction area by default.
[0312] For example, as shown in FIG20B , the decoding end defaults to determining an area consisting of a row above, a column to the left, and a pixel point in the upper left corner of the current block as the first reconstruction area.
[0313] Method 2: Determine the first reconstruction area based on the shape of the current block.
[0314] For example, if the shape of the current block is a square, the reconstructed pixel area in one row above and one column on the left of the current block is determined as the first reconstructed area.
[0315] For another example, if the shape of the current block is a rectangle with a width greater than a height, a row of reconstructed pixel areas above the current block is determined as the first reconstructed area.
[0316] For another example, if the shape of the current block is a rectangle with a height greater than a width, a left column of reconstructed pixel areas of the current block is determined as the first reconstructed area.
[0317] It should be noted that, based on the shape of the current block, the manner of determining the first reconstruction area includes but is not limited to the above examples.
[0318] After determining the first reconstruction area, the decoding end determines the pixel average reconstruction value m based on the reconstruction value of the first reconstruction area.
[0319] The embodiment of the present application does not limit the specific method of determining the pixel average reconstruction value m based on the reconstruction value of the first reconstruction area in the above S101-A2.
[0320] In mode 1, the above S101 - A2 includes: determining the average value of the reconstruction values of the first reconstruction area as the pixel average reconstruction value m.
[0321] In an example of method 1, if the first reconstruction area is as shown in FIG20B , the pixel average reconstruction value m can be calculated using the method shown in Table 14:
[0322] Table 14
[0323] In another example of method 1, if the first reconstruction area is a row above and / or a column to the left of the current block, the average of the reconstruction values in the row above and / or the column to the left can be determined as the pixel average reconstruction value m. In this case, the pixel average reconstruction value m can be calculated using the method shown in Table 15:
[0324] Table 15
[0325] As shown in Table 15 above, if the first reconstruction area is a row above and / or a column to the left of the current block, shift calculation can be used instead of division to quickly calculate the pixel average reconstruction value m.
[0326] Mode 2, the above S101-A2 includes: determining a pixel average reconstruction value based on the shape of the current block and the reconstruction value of the first reconstruction area.
[0327] For example, if the current block is a square, the average value of the entire first reconstruction area determined above is determined as the pixel average reconstruction value m.
[0328] In an example of method 2, the first reconstructed area includes an upper reconstructed area and a left reconstructed area of the current block. At this time, determining the pixel average reconstruction value based on the shape of the current block and the reconstruction value of the first reconstructed area includes: determining the first area from the upper reconstructed area and the left reconstructed area based on the shape of the current block; determining the average reconstruction value of the first area based on the reconstruction value of the first area; and determining the pixel average reconstruction value based on the average reconstruction value of the first area.
[0329] In this example, if the first reconstructed region includes the reconstructed region above and to the left of the current block, to reduce computational duplication, the mean is determined in the same manner as in the DC prediction mode. Specifically, based on the shape of the current block, the first region is determined from the reconstructed region above and to the left of the first reconstructed region.
[0330] For example, if the shape of the current block is that the width is greater than the height, the upper reconstruction area is determined as the first area.
[0331] For example, if the shape of the current block is that the height is greater than the width, the left reconstructed area is determined as the first area.
[0332] For example, if the shape of the current block is that the height is equal to the width, the upper reconstruction area and the left reconstruction area are determined as the first area.
[0333] Next, based on the reconstruction value of the selected first area, the average reconstruction value of the first area is determined, and then based on the average reconstruction value of the first area, the pixel average reconstruction value m is determined, for example, the average reconstruction value of the first area is determined as the pixel average reconstruction value m.
[0334] That is, in this example, if the current block is wider than it is tall, the average reconstructed value of the reconstructed area above the current block is determined as the pixel average reconstructed value m. If the current block is taller than it is wide, the average reconstructed value of the reconstructed area to the left of the current block is determined as the pixel average reconstructed value m. If the current block is taller than it is wide, the average reconstructed value of the reconstructed areas above and to the left of the current block is determined as the pixel average reconstructed value m.
[0335] For example, the decoding end may calculate the pixel average reconstruction value m by the method shown in Table 16:
[0336] Table 16
[0337] As shown in Table 16, the reconstruction region (i.e., the first region) used to calculate the average pixel reconstruction value m is determined based on the shape of the current block. The average pixel reconstruction value m is then quickly calculated based on the selected first region. During the calculation of the average pixel reconstruction value m, a simple shift is used to implement division. This avoids the problem of large division operations during average value calculation caused by the different lengths and widths of the current block, resulting in different sizes of the left and upper reconstruction regions. This improves the calculation speed of the average pixel reconstruction value m and enhances the prediction efficiency of the current block.
[0338] After the decoding end determines the pixel average reconstruction value, it removes the average of the reconstructed values of the pixels in the reference area based on the pixel average reconstruction value.
[0339] For example, for each pixel in the reference area, the reconstructed value of the pixel is divided by the pixel average reconstructed value and then rounded to the integer to obtain the pixel value of the pixel in the reference area after removing the average value.
[0340] For another example, the decoding end subtracts the pixel average reconstructed value from the reconstructed value of the pixel in the reference area to obtain the pixel value of the pixel in the reference area after the average is removed. For example, for each pixel in the reference area, the pixel average reconstructed value is subtracted from the reconstructed value of the pixel to obtain the pixel value of the pixel in the reference area after the average is removed.
[0341] The embodiment of the present application does not limit the specific manner in which the decoding end performs de-averaging on the reconstructed values of the pixels in the reference area based on the pixel average reconstruction value.
[0342] Based on the above method, the decoding end de-averages the reconstructed values of the pixel points in the reference area, obtains the pixel values of the de-averaged pixel points in the reference area, and then executes the above steps S101-A4, uses the pixel values of the de-averaged pixel points in the reference area as the input of the interpolation filter, slides the interpolation filter in the reference area, and obtains the filter coefficients of the interpolation filter.
[0343] For example, as shown in FIG21, if the interpolation filter of the current block is an interpolation filter of five different shapes and the reference area of the current block is a reference area of three different types, the interpolation filter of the current block is slid on the reference area after the current block is averaged to obtain the filter coefficient of the interpolation filter. The interpolation filter can slide horizontally row by row or vertically column by column on the reference area after the current block is averaged. The block to be predicted in FIG21 is the current block.
[0344] As shown in Figure 21, when using the interpolation filter to slide in the reference area of the current block, first, according to the shape of the interpolation filter, the N positions corresponding to each position in the reference area are determined. For example, for position r in the reference area, based on the shape of the interpolation filter, the N positions corresponding to position r are determined in the reference area, and the pixel reconstruction values of these N positions are the input of the interpolation filter. The phase position difference between these N positions and position r is {p0, p1, ..., p N-1}, p N is a two-dimensional representation. In which {c0,c1,…,c N-1} is {p0,p1,…,p N-1}Interpolation filter coefficients at the position.
[0345] In one example, the interpolation filter is slid in the reference area of the current block, and the constructed Wienerhof equation is shown in formula (4):
[0346] in, is the reference area of the current block, t[r+p n ]-m is r+p in the reference area n The pixel reconstructed value after removing the mean of the pixel point at position r in the reference area is t[r]-m, and the pixel reconstructed value after removing the mean of the pixel point at position r in the reference area is t[r]-m.
[0347] Since the reference area of the current block is the reconstruction area, the above formula (4) contains the following: Except for , all other parameters are known, so the filter coefficients of the interpolation filter of the current block can be determined by solving the above formula (4).
[0348] In one example, the decoding end may solve the Wienerhof equation shown in the above formula (4) by means of Cholesky decomposition of the autocorrelation coefficient matrix to obtain the filter coefficients of the filter.
[0349] After the decoding end determines the filter coefficients of the interpolation filter based on the above steps, it executes the following step S102.
[0350] S102 : Based on the filter coefficients, use an interpolation filter to perform parallel prediction on at least two pixels in the current block to determine a prediction block of the current block.
[0351] After determining the filter coefficients of the interpolation filter based on the above steps, the decoder uses the interpolation filter to perform interpolation filtering prediction on the current block based on the filter coefficients to obtain a predicted block of the current block.
[0352] In related art, when using an interpolation filter to perform interpolation filtering prediction on the current block, after the prediction of one pixel in the current block is completed, the next pixel is predicted. For example, as shown in FIG22A , the interpolation filter performs interpolation filtering prediction on each pixel in the current block one by one along the horizontal direction. During prediction, after the prediction of the previous pixel in the horizontal direction is completed, the predicted value of the previous pixel is used as an input pixel value of the interpolation filter of the next pixel to predict the next pixel. As shown in FIG22 , assuming that the shape of the interpolation filter of the current block is 4×4, the decoder uses an interpolation filter with known filter coefficients to perform interpolation prediction on each position in the current block one by one. Specifically, for the rth point in the current block, the pixel values of the N positions corresponding to the rth point are first determined based on the shape of the interpolation filter of the current block. For example, as shown in FIG22 , in the 4×4 interpolation filter, the dark position is the position of the rth point to be processed, and the 15 light positions are the N positions corresponding to the rth point. The block to be predicted in FIG22 is the current block. Next, the pixel values of the N positions corresponding to the rth point are determined. For example, for any of the N positions, if the position is located in the reconstructed area surrounding the current block, the reconstructed value of the position is determined as the pixel value of the position. If the position is located within the current block, the predicted value of the position is determined as the pixel value of the position.
[0353] For another example, as shown in Figure 22B, the interpolation filter performs interpolation filtering prediction on each pixel in the current block one by one along the vertical direction. During prediction, after the prediction of the previous pixel in the vertical direction is completed, the predicted value of the previous pixel is used as an input pixel value of the interpolation filter of the next pixel to predict the next pixel. In other words, when the related art uses the interpolation filter to perform interpolation filtering prediction on the current block, it can only complete the prediction of one pixel at a time, resulting in long prediction time and low prediction efficiency, which in turn affects decoding efficiency.
[0354] In order to solve the above technical problems, the embodiment of the present application uses an interpolation filter to perform interpolation filtering prediction on the current block, and performs parallel prediction on at least two points in the current block. That is, the decoding end can use the interpolation filter to perform interpolation filtering prediction on at least two pixel points in the current block at the same time.
[0355] In the embodiment of the present application, the decoding end uses the interpolation filter to perform interpolation filtering prediction on at least two pixels in the current block at the same time, including at least two implementation methods:
[0356] In a first implementation, for at least two pixels in a current block, these at least two pixels are considered adjacent pixels. The decoder first determines interpolation filter input information corresponding to the at least two pixels, inputs the input information into the interpolation filter for interpolation filter prediction, obtains a prediction value, and then determines the prediction value of the at least two pixels based on the prediction value. For example, based on relevant feature information of the at least two pixels, the prediction value is processed to obtain the prediction value corresponding to the at least two pixels. In another example, the prediction value is determined as the prediction value corresponding to the at least two pixels.
[0357] In the first implementation manner, the decoding end does not impose any restrictions on the specific method of determining the interpolation filter input information corresponding to the at least two pixel points.
[0358] For example, based on the shape of the interpolation filter, the interpolation filter input value that is the same for the at least two pixels is determined, and the same interpolation filter input value is used as the interpolation filter input information. For example, the at least two pixels include pixel 1 and pixel 2. Based on the shape of the interpolation filter, the N input values corresponding to pixel 1 and the N input values corresponding to pixel 2 are determined. The same input value among the N input values corresponding to pixel 1 and the N input values corresponding to pixel 2 is determined, and the same input value is used as the input value of the interpolation filter. It should be noted that if the N input values corresponding to pixel 2 or pixel 1 include an undecoded value, the undecoded value is discarded.
[0359] For another example, based on the shape of the interpolation filter, the input value corresponding to the pixel with the most decoded input values among the at least two pixels is determined as the interpolation filter input information. For example, the at least two pixels include pixel 1 and pixel 2. Based on the shape of the interpolation filter, the N input values corresponding to pixel 1 and the N input values corresponding to pixel 2 are determined. The N input values corresponding to pixel 1 are all decoded, and the N input values corresponding to pixel 2 include undecoded input values. Therefore, the N input values corresponding to pixel 1 are used as the interpolation filter input information.
[0360] In the above-mentioned method 1, the specific process of the interpolation filter predicting the predicted values of at least two pixels in the current block after the decoding end inputs the same input value to the interpolation filter at one time is introduced.
[0361] The second method is to use an interpolation filter to perform interpolation filtering on at least two pixels in the current block at the same time. For example, at time t, the decoder uses an interpolation filter to perform interpolation filtering prediction on pixel 1 in the current block to obtain the predicted value of pixel 1. At the same time, the decoder uses an interpolation filter to perform interpolation filtering prediction on pixel 2 in the current block to obtain the predicted value of pixel 2.
[0362] The embodiment of the present application does not restrict the prediction direction when the decoding end uses the interpolation filter to perform parallel prediction on at least two pixels in the current block.
[0363] In some embodiments, the decoding end may use an interpolation filter to perform parallel prediction on at least two pixels in the current block along a horizontal direction. Exemplarily, when performing parallel prediction on at least two pixels, if one or more input values of the interpolation filter corresponding to the pixel are not decoded, the undecoded input values are discarded and the decoded input values are used as input values of the interpolation filter.
[0364] In some embodiments, the decoding end may use an interpolation filter to perform parallel prediction on at least two pixels in the current block along a vertical direction. Exemplarily, when performing parallel prediction on at least two pixels, if one or more input values of the interpolation filter corresponding to the pixel are not decoded, the undecoded input values are discarded and the decoded input values are used as input values of the interpolation filter.
[0365] In some embodiments, the decoding end may use an interpolation filter to perform parallel prediction on at least two pixels in the current block along a diagonal direction. Based on this, in one example, the above S102 includes the following step S102-A:
[0366] S102-A: Based on the filter coefficients, use an interpolation filter along the diagonal direction to perform parallel interpolation filtering prediction on the pixel points on the same diagonal line of the current block to obtain a prediction block of the current block.
[0367] Exemplarily, as shown in FIG23, when the decoding end uses the interpolation filter to predict the pixel points in the current block, the pixel point to be predicted is located at a corner of the area selected by the interpolation filter (for example, the lower right corner or the upper left corner). In this way, for any pixel point on the same angle line (i.e., the pixel point indicated by the depth block in FIG23), based on the shape of the interpolation filter, the N positions corresponding to the selected pixel point do not include other pixel points on the diagonal line. In other words, the N positions corresponding to each pixel point on the same diagonal line do not include the pixel points on the diagonal line. For example, taking two adjacent pixel points a and pixel point b on the same diagonal line of the current block as an example, as shown in FIG23, based on the shape of the interpolation filter, the N positions corresponding to pixel point a and the N positions corresponding to pixel point b are determined respectively, wherein the N positions corresponding to pixel point a and the N positions corresponding to the determined pixel point b do not include the pixel points on the diagonal line. Based on this, when the decoder uses an interpolation filter to perform interpolation filtering prediction on the current block, it can perform parallel interpolation filtering prediction on the pixels on the same diagonal line in the current block along the diagonal direction. For example, parallel interpolation filtering prediction can be performed on pixel a and pixel b on the same diagonal line. It should be noted that the shape of the interpolation filter shown in Figure 23 is an example, and the shape of the interpolation filter in the embodiments of the present application is not limited to this.
[0368] In an embodiment of the present application, the left, lower left, upper left, left and upper right areas of the current block have been decoded. Therefore, the starting point of the prediction of the current block along the diagonal direction can be determined based on the decoded areas and the shape of the interpolation filter.
[0369] In one example, as shown in FIG23 , the decoding end performs interpolation filtering prediction on the current block along the diagonal direction starting from the upper left corner of the current block. The above S102-A includes the following steps S102-A1:
[0370] S102-A1: The decoder uses the interpolation filter to perform parallel interpolation filtering prediction on the pixels on the same diagonal line of the current block, starting from the upper left corner of the current block, based on the filter coefficients, to obtain a predicted block for the current block. In this example, the pixel to be predicted is located in the lower right corner of the area selected by the interpolation filter.
[0371] The embodiment of the present application does not limit the specific direction of the diagonal line.
[0372] In some embodiments, as shown in FIG23 , when the decoding end starts from the upper left corner of the current block and performs parallel prediction on the pixel points on the same diagonal line in the current block, the diagonal direction includes at least one of the following: a direction from the upper right to the lower left, and a direction from the lower left to the upper right.
[0373] In one example, as shown in FIG24A , the diagonal direction includes a direction from the upper right to the lower left. At this time, as shown by the arrow in FIG24A , the diagonal direction of the current block is from the upper right to the lower left.
[0374] In one example, as shown in FIG24B , the diagonal direction includes a direction from the lower left to the upper right. At this time, as shown by the arrows in FIG24B , the diagonal directions of the current block are all from the lower left to the upper right.
[0375] In one example, as shown in FIG24C , the diagonal direction includes the direction from the upper right to the lower left and the direction from the lower left to the upper right. At this time, as shown by the arrow in FIG24C , the diagonal direction of the current block includes two directions: the direction from the upper right to the lower left and the direction from the lower left to the upper right.
[0376] It should be noted that, since in the embodiment of the present application, the decoding end performs parallel prediction on the pixel points located on the same diagonal line in the current block, the specific direction of the diagonal line does not constitute a limitation on the technical solution of the embodiment of the present application.
[0377] In the embodiment of the present application, the decoding end performs prediction on a pixel point on a diagonal line of the current block at each prediction, and the decoding end performs the same parallel prediction process on each pixel point on a diagonal line of the current block. For ease of description, the k-th diagonal line of the current block is used as an example. In this case, the above-mentioned S102-A1 includes the following steps S102-A11 and S102-A12:
[0378] S102-A11. For M pixels on the k-th diagonal line of the current block, use an interpolation filter in parallel to determine predicted values of the M pixels based on the filter coefficients, where k and M are both positive integers.
[0379] S102-A12: Obtain a predicted value of the current block based on the predicted values of the pixels on each diagonal line in the current block.
[0380] The kth diagonal line can be understood as any diagonal line in the current block shown in Figure 23 , which includes M pixels. The decoder uses an interpolation filter based on the filter coefficients to concurrently determine the predicted values for these M pixels. In other words, the decoder can simultaneously determine the predicted values for the M pixels on the kth diagonal line, significantly increasing prediction speed.
[0381] For example, as shown in Figure 24A, assume that the kth diagonal of the current block includes three pixels. The decoder determines the predicted values of these three pixels in parallel. For example, these three pixels are denoted as pixel 1, pixel 2, and pixel 3. At the same time, the decoder uses an interpolation filter based on the filter coefficients to perform interpolation filtering prediction on pixel 1 to obtain the predicted value of pixel 1. Simultaneously, the interpolation filter is used to perform interpolation filtering prediction on pixel 2 based on the filter coefficients to obtain the predicted value of pixel 2. Simultaneously, the interpolation filter is used to perform interpolation filtering prediction on pixel 3 based on the filter coefficients to obtain the predicted value of pixel 3. In this way, the decoder determines the predicted values of the three pixels on the kth diagonal of the current block in parallel during a single interpolation filtering prediction process, significantly improving the speed of interpolation filtering prediction. By referring to the method for determining the predicted value of the pixel on the kth diagonal, the decoder can determine the predicted values of the pixels on other diagonals in the current block, thereby obtaining the predicted block for the current block, thereby improving the prediction speed of the current block and enhancing decoding efficiency.
[0382] The embodiment of the present application does not limit the specific method in which the decoding end uses an interpolation filter to determine the predicted values of M pixels in parallel based on the filter coefficients.
[0383] In some embodiments, since these M pixel points are points on the k-th diagonal of the current block, these M pixel points can be understood as adjacent pixel points with relatively close features. Therefore, in order to reduce the computational complexity, the input value of the interpolation filter corresponding to one or several pixel points among the M pixel points is determined based on the shape of the interpolation filter. Then, based on the input value of the interpolation filter corresponding to the one or several pixel points, for example, an average value calculation, a weighted calculation, or other calculation method is performed to determine the input value of the interpolation filter corresponding to the other pixel points among the M pixel points except for the one or several pixel points. Finally, based on the filter coefficient and the input value of the interpolation filter corresponding to each pixel point among the M pixel points, the predicted values of the M pixel points are determined in parallel.
[0384] In some embodiments, the above-mentioned S102-A11, based on the filter coefficients, uses the interpolation filter to determine the predicted values of M pixels in parallel, including the following steps:
[0385] S102-A11-a1. Based on the shape of the interpolation filter, determine the pixel values of N positions corresponding to the M pixel points in parallel;
[0386] S102-A11-a2. Based on the filter coefficients and the pixel values at N positions corresponding to the M pixel points, the predicted values of the M pixel points are determined in parallel.
[0387] In this embodiment, when the decoder concurrently determines the predicted values for the M pixels on the kth diagonal line of the current block, it concurrently determines the pixel values at the N positions corresponding to each of the M pixels based on the shape of the interpolation filter. The pixel values at the N positions corresponding to each pixel can be understood as the input value of the interpolation filter corresponding to that pixel. Next, the decoder concurrently determines the predicted values for the M pixels based on the filter coefficients and the pixel values at the N positions corresponding to each of the M pixels.
[0388] For example, as shown in Figure 24, assume that the kth diagonal line of the current block includes three pixels, which are denoted as pixel 1, pixel 2, and pixel 3. Assume that the interpolation filter for the current block is a 4x4 interpolation filter. When the decoder determines the predicted values for these three pixels in parallel, it determines the pixel values at the 15 positions corresponding to pixel 1 based on the shape of the interpolation filter. These 15 pixel values are input to the interpolation filter, and the predicted value for pixel 1 is determined based on the filter coefficients determined above. Simultaneously, the decoder determines the pixel values at the 15 positions corresponding to pixel 2 based on the shape of the interpolation filter. These 15 pixel values are input to the interpolation filter, and the predicted value for pixel 2 is determined based on the filter coefficients determined above. Simultaneously, the decoder determines the pixel values at the 15 positions corresponding to pixel 3 based on the shape of the interpolation filter. These 15 pixel values are input to the interpolation filter, and the predicted value for pixel 3 is determined based on the filter coefficients determined above. That is, in this embodiment, the decoding end determines the predicted values of three pixels in the current block in parallel at the same time, which greatly improves the prediction speed and thus improves the decoding efficiency.
[0389] The following describes the specific process of determining the predicted values of M pixel points in parallel based on the filter coefficients and the pixel values of N positions corresponding to the M pixel points in the above S102-A11-a2.
[0390] In some embodiments, for each of the M pixels, the decoding end directly multiplies the pixel values of the N positions corresponding to the pixel by the filter coefficient to obtain the predicted value of the pixel.
[0391] For example, the decoding end obtains the predicted value of each pixel in the M pixels based on the following formula (5):
[0392] Among them, p n is the relative position difference between the nth position and position ri in the N positions corresponding to position ri in the current block, c n is the nth filter coefficient among the filter coefficients. n ] is the position ri+p nThe pixel value at position ri+p n In the current block, t[ri+p n ] is the position ri+p n The predicted value of the pixel at position ri+p n In the reconstruction area around the current block, t[ri+p n ] is the position ri+p n The reconstructed value of the pixel at pred ri is the predicted value of the pixel at position ri in the current block.
[0393] Based on formula (5), the decoding end can determine the predicted values of each pixel located on the same diagonal line in the current block in parallel.
[0394] In some embodiments, if the decoding end determines the filter coefficient of the interpolation filter based on the above formula (4), the filter coefficient is determined by the reference area after the mean value is removed in the above formula (4). Therefore, when determining the prediction value of the current block based on the filter coefficient, the influence of the pixel average reconstruction value m needs to be considered.
[0395] In a possible implementation of this embodiment, the interpolation filter coefficient determined by the above formula (4) is substituted into the above formula (5) to obtain the predicted value of each point in the current block. Then, the predicted value of each point is added to the pixel average reconstruction value m to obtain the final predicted value of each point in the current block, thereby obtaining the predicted block of the current block.
[0396] In another possible implementation of this embodiment, the above S102-A11-a2 includes the following steps:
[0397] S102-A11-a21, based on the pixel average reconstruction value, performing de-averaging on the pixel values of the N positions corresponding to the M pixel points in parallel to obtain the de-averaged pixel values of the N positions corresponding to the M pixel points;
[0398] S102-A11-a22. Based on the filter coefficients and the pixel values after averaging at N positions corresponding to the M pixel points, the predicted values of the M pixel points are determined in parallel.
[0399] Because the filter coefficients are determined based on the reference area after averaging, the decoder performs averaging on the pixel values at the N positions corresponding to the M pixels on the k-th diagonal of the current block based on the pixel average reconstruction value to obtain the pixel values at the N positions after averaging. For example, for any pixel among the M pixels, the pixel value at the N positions of the pixel is subtracted from the pixel average reconstruction value to obtain the pixel value at the N positions after averaging.
[0400] Next, based on the filter coefficients and the pixel values after averaging at the N positions corresponding to the M pixels, the predicted values of the M pixels are determined in parallel.
[0401] The embodiment of the present application does not limit the specific method of determining the predicted values of M pixel points in parallel based on the filter coefficients and the pixel values after averaging at N positions corresponding to the M pixel points.
[0402] In one implementation, for each of the M pixels, the decoding end substitutes the pixel value and filter coefficient of the N positions of the pixel after removing the mean value into the above formula (5). At this time, t[ri+p n ] is the position ri+p n The pixel value after removing the mean value of the pixel at . After determining a predicted value of the r-th point based on the above formula (5), add the pixel average reconstruction value m to the predicted value to obtain the final predicted value of the r-th point.
[0403] In another implementation, the above S102-A11-a22 includes the following steps:
[0404] S102-A11-a221, determine a second reconstruction area around the current block, and determine a maximum reconstruction value and a minimum reconstruction value of the second reconstruction area;
[0405] S102-A11-a222, based on the pixel values after de-averaging at N positions corresponding to the M pixel points, the filter coefficients, and the pixel average reconstruction value, obtain first prediction values of the M pixel points in parallel;
[0406] S102-A11-a223. Based on the first predicted values, the maximum reconstructed value and the minimum reconstructed value of the M pixel points, the predicted values of the M pixel points are determined in parallel.
[0407] In this implementation, the decoding end limits the prediction value of the current block to a range. Specifically, a second reconstruction area is determined, and the maximum reconstruction value (max) and the minimum reconstruction value (min) of the pixels in the second reconstruction area are determined.
[0408] The embodiment of the present application does not limit the specific method of determining the second reconstruction area around the current block.
[0409] In an example, the second reconstructed region of the current block is consistent with the reference region of the current block.
[0410] In one example, the second reconstructed area of the current block is consistent with the first reconstructed area of the current block.
[0411] In one example, the reconstruction areas above, to the left, to the upper right, to the upper left, and to the lower left of the current block are determined as the second reconstruction area. For example, the reconstruction areas 13 rows above, 13 columns to the left, 13 rows to the upper right, 13 rows and 13 columns to the upper left, and 13 columns to the lower left of the current block are determined as the second reconstruction area.
[0412] It should be noted that there is no order of priority between the above S102-A11-a221 and the above S102-A11-a222 in the specific implementation process. For example, the above S102-A11-a221 can be executed before the above S102-A11-a222, or after the above S102-A11-a222, or synchronously with the above S102-A11-a222.
[0413] The embodiment of the present application does not limit the specific method for the decoding end to obtain the first predicted values of M pixel points in parallel based on the pixel values, filter coefficients and pixel average reconstruction values after de-averaging N positions corresponding to the M pixel points.
[0414] For example, for any pixel among the M pixels, the pixel value after de-averaging the N positions of the pixel is multiplied by the filter coefficient to obtain the second predicted value of the pixel; the second predicted value and the pixel average reconstruction value are added to obtain the first predicted value of the pixel.
[0415] For example, the decoding end obtains the first predicted value of the pixel based on the following formula (6):
[0416] in, is the position r+p n The pixel value after removing the mean value of the pixel point at pred r is the first predicted value of the rth pixel among M pixels, is the second predicted value of the r-th point.
[0417] For another example, after the decoding end obtains a predicted value of the pixel point based on the above formula (6), the predicted value is processed in a preset manner to obtain a first predicted value of the pixel point.
[0418] After determining the first predicted values of the M pixels based on the above steps, the decoding end determines the predicted values of the M pixels in parallel based on the first predicted values, the maximum reconstructed value, and the minimum reconstructed value.
[0419] For example, for any pixel among the M pixels, if the first predicted value of the pixel is greater than the minimum reconstruction value and less than the maximum reconstruction value, the first predicted value is determined as the predicted value of the pixel.
[0420] For another example, if the first predicted value of the pixel point is less than or equal to the minimum reconstructed value, the minimum reconstructed value is determined as the predicted value of the pixel point.
[0421] For another example, if the first predicted value of the pixel point is greater than or equal to the maximum reconstructed value, the maximum reconstructed value is determined as the predicted value of the pixel point.
[0422] In one example, the decoding end determines the predicted value of the pixel point using the following formula (7):
[0423] Where Clip represents the first predicted value of the rth point among M pixels The value is limited to the maximum reconstruction value max and the minimum reconstruction value min.
[0424] Taking the above method of determining the predicted values of M pixels on the kth diagonal line in the current block as an example, the decoding end can refer to the above method to determine the predicted values of each pixel on the diagonal line in the current block in parallel, and then obtain the predicted value of each point in the current block to form a predicted block of the current block.
[0425] Based on the above steps, the decoding end performs interpolation filtering prediction on the current block, obtains the predicted block of the current block, and then performs the following steps.
[0426] S103 : Determine a transformation kernel corresponding to the current block, and determine a reconstructed block of the current block based on the transformation kernel corresponding to the current block and the prediction block.
[0427] As can be seen from the above, when decoding the current block, the decoder decodes the bitstream to obtain the quantization coefficients of the current block. Then, it dequantizes the quantization coefficients to obtain the transform coefficients of the current block. The transform coefficients of the current block are then inversely transformed to obtain the residual block (or residual value) of the current block. At the same time, the prediction mode of the current block is determined, and the current block is predicted using this prediction mode to obtain the predicted block of the current block. The predicted block and the residual block are added together to obtain the reconstructed block of the current block.
[0428] When inversely transforming the transform coefficients of the current block, it is necessary to determine a transform kernel, and inversely transform the transform coefficients of the current block based on the transform kernel to obtain a residual value of the current block.
[0429] The embodiment of the present application does not limit the specific method by which the decoding end determines the transform kernel corresponding to the current block.
[0430] In some embodiments, the encoding end and the decoding end use a default transform kernel as the transform kernel of the current block.
[0431] In some embodiments, after determining the transformation core of the current block, the encoder writes the indication information of the transformation core into the bitstream, so that the decoder can determine the transformation core of the current block by decoding the bitstream.
[0432] In some embodiments, the decoding end determines the transform kernel of the current block through the following steps S103-A and S103-A:
[0433] S103-A, determining the intra prediction mode corresponding to the prediction block;
[0434] S103-B. Determine a transform kernel corresponding to the current block based on the intra prediction mode corresponding to the prediction block.
[0435] In the embodiment of the present application, after determining the prediction block of the current block using the interpolation filter prediction mode, the traditional intra prediction mode corresponding to the prediction block is determined, and then based on the traditional intra prediction mode, the transform kernel corresponding to the current block is determined.
[0436] The following describes the specific process of determining the intra-frame prediction mode corresponding to the prediction block at the decoding end.
[0437] In one example, as shown in FIG7 , the conventional intra prediction modes currently included in VVC are:
[0438] PLANAR mode: intra prediction mode index is 0,
[0439] DC mode: intra prediction mode index is 1,
[0440] Angle mode: The intra prediction mode index is 2 to 66.
[0441] In one example, as shown in Figure 25 , the arrows in the figure point to the directions predicted by the angle modes in VVC, and the prediction mode indexes used during decoding are 2 to 66. When the current block is a non-square block, some angle directions will be replaced with wide angles, such as -1 to -14 and 67 to 80 in Figure 25 .
[0442] In some embodiments, the intra-frame prediction mode corresponding to the above-mentioned prediction block is a default intra-frame prediction mode. That is, if the current block is predicted using the interpolation filtering prediction mode, when the prediction block is obtained, one of the traditional intra-frame prediction modes is determined as the default intra-frame prediction mode corresponding to the prediction block.
[0443] In some embodiments, the decoding end determines the intra prediction mode corresponding to the prediction block through the following steps S103-A1 and S103-A2:
[0444] S103-A1, determining angle values of R points in the prediction block, where R is a positive integer;
[0445] S103-A2: Determine the intra-frame prediction mode corresponding to the prediction block based on the angle values of the R points.
[0446] In the embodiment of the present application, the intra-frame prediction mode corresponding to the prediction block is determined by counting the intra-frame prediction modes corresponding to the angle values of R points in the prediction block.
[0447] The embodiment of the present application does not limit the specific position and number of the R points in the prediction block used to determine the angle value. For example, the R points can be one point in the prediction block, or multiple points in the prediction block.
[0448] For example, if the above R points are one point, the decoding end determines the angle value of a point in the prediction block (for example, the center point of the prediction block), and based on the angle value of the point, determines the intra-frame prediction mode corresponding to the point, and then determines the intra-frame prediction mode as the intra-frame prediction mode corresponding to the prediction block.
[0449] For another example, if the above-mentioned R points are multiple points, the decoding end determines the angle values of these multiple points, and based on the angle values of these multiple points, determines the intra-frame prediction mode corresponding to each of these multiple points, and then determines the intra-frame prediction mode with the largest number of identical intra-frame prediction modes among these multiple points as the intra-frame prediction mode corresponding to the prediction block.
[0450] In some embodiments, when determining the angle values of R points in a prediction block using a sliding window approach, the selection of these R points is related to the shape and size of the sliding window. For example, each of the R points is the center point of the sliding window as it slides across the prediction block.
[0451] In the embodiment of the present application, the method for determining the angle value of each of the R points is the same. For the convenience of description, the method of determining the angle value of the i-th point among the R points is used as an example for explanation.
[0452] The embodiment of the present application does not limit the specific method of determining the angle value of the point.
[0453] In some embodiments, the above S103-A1 includes steps S103-A11 and S103-A12:
[0454] S103-A11, for the i-th point among the R points, determine the horizontal gradient and vertical gradient of the i-th point, where i is a positive integer less than or equal to R;
[0455] S103-A12: Determine the angle value of the i-th point based on the horizontal gradient and the vertical gradient of the i-th point.
[0456] In this embodiment, for each of the R points, for example, the i-th point, the decoding end first determines the horizontal gradient and vertical gradient of the i-th point, and then determines the angle value of the i-th point based on the horizontal gradient and the vertical gradient.
[0457] The embodiment of the present application does not limit the specific method of determining the horizontal gradient and vertical gradient of the i-th point.
[0458] In one example, the horizontal gradient value of the i-th point is determined based on the predicted values of the points around the i-th point in the prediction block and the change in the predicted value of the i-th point in the horizontal direction, and the vertical gradient value of the i-th point is determined based on the predicted values of the points around the i-th point in the prediction block and the change in the predicted value of the i-th point in the vertical direction.
[0459] In another example, the decoding end determines the prediction value of the point in the sliding window centered on the i-th point in the prediction block; based on the prediction value of the point in the sliding window and the horizontal gradient operator, as well as the vertical gradient operator, the horizontal gradient and vertical gradient of the i-th point are obtained.
[0460] In this example, a sliding window is first determined. For example, as shown in Figure 26, a 3×3 sliding window is determined. This sliding window is then slid across the prediction block. With each slide, the horizontal and vertical gradients at the center point of the sliding window are determined. For example, with the center point of the current sliding window as the i-th point, the predicted values for each point within the current sliding window are first obtained. For example, predicted values for 3×3 = 9 points can be obtained. Next, based on the predicted values for these 9 points and the preset horizontal and vertical gradient operators, the horizontal and vertical gradients at the i-th point are determined.
[0461] For example, the product of the predicted value of the point in the sliding window and the horizontal gradient operator is determined as the horizontal gradient G of the i-th point x ; The product of the predicted value of the point in the sliding window and the vertical gradient operator is determined as the vertical gradient of the i-th point.
[0462] For another example, the predicted value of the point in the sliding window is multiplied by the horizontal gradient operator and then the preset operation is performed with the preset value to obtain the horizontal gradient G of the i-th point. x ; Multiply the predicted value of the point in the sliding window by the vertical gradient operator and then perform a preset operation with the preset value to obtain the vertical gradient of the i-th point.
[0463] The embodiment of the present application does not limit the specific values of the horizontal gradient operator and the vertical gradient operator.
[0464] For example, the horizontal gradient operator M x and the vertical gradient operator M y for:
[0465] After the decoding end determines the horizontal gradient and vertical gradient of the i-th point based on the above steps, it can determine the angle value of the i-th point according to the horizontal gradient and vertical gradient of the i-th point.
[0466] For example, the inverse tangent value of the ratio of the vertical gradient to the horizontal gradient at the i-th point is determined as the angle value of the i-th point. For example, as shown in formula (8):
[0467] Among them, G x is the horizontal gradient of the i-th point, G y is the vertical gradient of the i-th point, O is the angle value of the i-th point, and atan() is the inverse tangent function.
[0468] In addition to using the above formula (8) to determine the angle value of the i-th point, the decoding end may also use other methods to determine the angle value of the i-th point. For example, the decoding end adjusts the angle value determined by the above formula (8) to obtain the angle value of the i-th point.
[0469] The decoding end uses the above method for each of the R points to determine the angle value of each of the R points, and then executes the above S103-A2 to determine the intra-frame prediction mode corresponding to the prediction block based on the angle values of the R points.
[0470] The embodiment of the present application does not limit the specific method of determining the intra-frame prediction mode corresponding to the prediction block based on the angle values of R points.
[0471] In some embodiments, the decoding end selects the angle value 1 that is the same the most times from the angle values of the R points, matches the angle value 1 with the prediction angle of the traditional intra-frame prediction mode, obtains the intra-frame prediction mode corresponding to the angle value 1, and determines the intra-frame prediction mode corresponding to the angle value 1 as the intra-frame prediction mode corresponding to the prediction block.
[0472] In some embodiments, the above S103-A2 includes the following steps S103-A21 and S103-A22:
[0473] S103-A21, determining the intra-frame prediction modes corresponding to the R points based on the angle values of the R points;
[0474] S103-A22. Determine the intra-frame prediction mode corresponding to the prediction block based on the intra-frame prediction modes corresponding to the R points.
[0475] In this implementation, the decoder determines the intra-frame prediction mode corresponding to each of the R points based on the angle value of each point. For example, for each of the R points, the angle value of that point is matched with the prediction angle of the traditional intra-frame prediction mode to obtain the intra-frame prediction mode corresponding to that angle value. In this way, the intra-frame prediction mode corresponding to each of the R points can be obtained.
[0476] Next, based on the intra-frame prediction mode corresponding to each of the R points, the intra-frame prediction mode corresponding to the prediction block is determined.
[0477] In a possible implementation, the intra-frame prediction mode with the greatest number of repetitions among the intra-frame prediction modes corresponding to the R points is determined as the intra-frame prediction mode corresponding to the prediction block.
[0478] In another possible implementation, the above S103-A22 includes the following steps:
[0479] S103-A221, based on the horizontal gradients and vertical gradients of the R points, determine the gradient amplitude values corresponding to the R points;
[0480] S103-A222: Determine the intra-frame prediction mode corresponding to the prediction block based on the intra-frame prediction modes and gradient magnitude values corresponding to the R points.
[0481] In this implementation, the decoding end determines the gradient amplitude value corresponding to each of the R points based on the horizontal gradient and vertical gradient of each of the R points determined above.
[0482] In the embodiment of the present application, the specific manner in which the decoding end determines the gradient amplitude value corresponding to each of the R points is the same. For ease of description, the example of determining the gradient amplitude value corresponding to the i-th point among the R points is taken as an example.
[0483] The embodiment of the present application does not limit the specific manner in which the decoding end determines the gradient amplitude value corresponding to the i-th point based on the horizontal gradient and vertical gradient of the i-th point.
[0484] For example, the decoding end multiplies the horizontal gradient and the vertical gradient of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.
[0485] For another example, the decoding end adds the absolute value of the horizontal gradient and the absolute value of the vertical gradient of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.
[0486] For example, the decoding end determines the gradient amplitude value corresponding to the i-th point based on the following formula (9): G = |G x |+|G y | (9)
[0487] Among them, G is the gradient amplitude value corresponding to the i-th point, G x is the horizontal gradient of the i-th point, G y is the vertical gradient of the i-th point.
[0488] The decoder can determine the gradient magnitude value corresponding to each of the R points based on the above steps. Next, the decoder performs the above steps S103-A222 to determine the intra prediction mode corresponding to the prediction block based on the intra prediction modes and gradient magnitude values corresponding to the R points.
[0489] In one example, the intra-frame prediction mode corresponding to the point with the largest gradient magnitude value among the R points is determined as the intra-frame prediction mode corresponding to the prediction block.
[0490] In another example, for any point among the R points, the gradient amplitude value corresponding to the point is accumulated on the intra-frame prediction mode corresponding to the point to obtain the accumulated gradient amplitude values of the intra-frame prediction modes corresponding to the R points; and the intra-frame prediction mode with the largest accumulated gradient amplitude value among the intra-frame prediction modes corresponding to the R points is determined as the intra-frame prediction mode corresponding to the prediction block.
[0491] For example, as shown in FIG27, the gradient amplitude value corresponding to each of the R points is accumulated on the corresponding intra-frame prediction mode. For example, the intra-frame prediction modes corresponding to point 1 and point 2 of the R points are both intra-frame prediction mode 1, so the gradient amplitude values corresponding to point 1 and point 2 are accumulated to the gradient amplitude value corresponding to intra-frame prediction mode 1. Similarly, the gradient amplitude value histogram shown in FIG27 can be obtained. In this way, the intra-frame prediction mode with the largest cumulative gradient amplitude value in the gradient amplitude value histogram can be determined as the intra-frame prediction mode corresponding to the prediction block. For example, the intra-frame prediction mode corresponding to the dark cumulative gradient amplitude value in FIG27 is determined as the intra-frame prediction mode corresponding to the prediction block.
[0492] In some embodiments, if the gradient magnitude values corresponding to the R points are all 0, the first intra-frame prediction mode is determined as the intra-frame prediction mode corresponding to the prediction block. In other words, if the gradient magnitude values corresponding to all of the R points are 0, it means that the horizontal gradient and vertical gradient of each of the R points are both 0. In this case, the preset first intra-frame prediction mode can be determined as the intra-frame prediction mode corresponding to the prediction block.
[0493] The embodiment of the present application does not limit the type of the first intra-frame prediction mode.
[0494] Exemplarily, the first intra-frame prediction mode is the PLANAR mode.
[0495] After the decoding end determines the intra-frame prediction mode corresponding to the prediction block based on the above steps, it determines the transform kernel corresponding to the current block based on the intra-frame prediction mode corresponding to the prediction block.
[0496] The embodiment of the present application does not limit the specific manner in which the decoding end determines the transform kernel corresponding to the current block based on the intra-frame prediction mode corresponding to the prediction block.
[0497] In some embodiments, the decoding end searches for an image block whose intra-frame prediction mode is the same as the intra-frame prediction mode corresponding to the prediction block in the decoded image blocks around the prediction block based on the intra-frame prediction mode corresponding to the prediction block, and then determines the transform kernel corresponding to the image block as the transform kernel corresponding to the current block.
[0498] In some embodiments, determining the transform kernel corresponding to the current block based on the intra prediction mode corresponding to the prediction block in S103-B includes the following steps:
[0499] S103-B1, obtaining a correspondence between an intra prediction mode and a transform core group, wherein one transform core group includes at least one type of transform core;
[0500] S103-B2, searching the corresponding relationship for the first transform kernel group corresponding to the intra prediction mode of the prediction block;
[0501] S103-B3: Determine a transform core corresponding to the current block from the first transform core group.
[0502] In the embodiment of the present application, there is a correspondence between the intra prediction mode and the transform core group. Based on this, after determining the intra prediction mode corresponding to the prediction block, the decoder obtains the preset correspondence between the intra prediction mode and the transform core group.
[0503] In one example, the correspondence between the intra prediction modes and the transform kernel groups is shown in Table 17:
[0504] Table 17
[0505] It should be noted that the above Table 17 is only a correspondence between an intra-frame prediction mode and a transform core group involved in an embodiment of the present application. The correspondence between the intra-frame prediction mode and the transform core group in the embodiment of the present application includes but is not limited to that shown in Table 17.
[0506] Each transformation core group includes at least one type of transformation core.
[0507] After obtaining the correspondence between intra-frame prediction modes and transform kernel groups as shown in Table 17, the decoder searches the correspondence between intra-frame prediction modes and transform kernel groups based on the intra-frame prediction mode corresponding to the prediction block, and records this transform kernel group as the first transform kernel group. For example, if the intra-frame prediction mode corresponding to the prediction block is an angular prediction mode in the 64-angle direction, searching Table 17 shows that the transform kernel group corresponding to this angular prediction mode in the 64-angle direction is 4. In this way, the decoder determines the transform kernel corresponding to the current block from the at least one type of transform kernel included in transform kernel group 4.
[0508] For example, if the first transform core group includes one transform core, the transform core is determined as the transform core corresponding to the current block.
[0509] For another example, if the first transform core group includes transform cores of multiple categories, the decoder determines the transform core category corresponding to the current block, and then determines the transform core of the transform core category in the first transform core group as the transform core corresponding to the current block.
[0510] The methods for the decoder to determine the transform kernel type corresponding to the current block include but are not limited to the following:
[0511] In one example, the transform kernel category corresponding to the current block is a default category, so the decoding end determines the default category as the transform kernel category corresponding to the current block.
[0512] In another example, the encoder writes the transform kernel type corresponding to the current block into the bitstream, so that the decoder obtains the transform kernel type corresponding to the current block by decoding the bitstream.
[0513] As can be seen from the above, in an embodiment of the present application, the decoding end uses an interpolation filter prediction mode to determine the prediction block of the current block, and then determines the traditional intra-frame prediction mode corresponding to the prediction block, and based on the traditional intra-frame prediction mode corresponding to the prediction block, determines the transform kernel corresponding to the current block. That is to say, the embodiment of the present application is based on the traditional intra-frame prediction mode derived from the interpolation filter prediction, and is used for the selection of transform kernel groups of the non-separable primary transform (NSPT) and the non-separable secondary transform (LFNST), so that the determined transform kernel is more consistent with the characteristics of the current block, and the accuracy of determining the transform kernel is improved. When the accurately determined transform kernel is used to determine the reconstruction value of the current block, the accuracy of determining the reconstruction value can be improved, and the decoding accuracy of the current block can be improved. In addition, when the embodiment of the present application determines the transform kernel of the current block through the traditional prediction mode corresponding to the prediction block, there is no need to indicate the transform kernel separately, which saves codewords and further improves the video encoding and decoding effect.
[0514] Based on the above steps, the decoding end determines the transform kernel corresponding to the current block, and then performs an inverse transform on the transform coefficients of the current block based on the transform kernel corresponding to the current block to obtain the residual block of the current block, and obtains the reconstructed block of the current block based on the prediction block and the residual block of the current block.
[0515] In the embodiment of the present application, the decoding end determines the prediction block of the current block and the transform kernel corresponding to the current block based on the above steps. In this way, the decoding end can decode the bitstream to obtain the quantization coefficients of the current block. Then, the quantization coefficients are dequantized to obtain the transform coefficients of the current block. Using the transform kernel corresponding to the current block determined above, the transform coefficients of the current block are detransformed to obtain the residual block (or residual value) of the current block. Finally, the decoding end adds the prediction block and the residual block of the current block to obtain the reconstructed block of the current block.
[0516] In some embodiments, the above-mentioned current block is a bright color block or a chroma block, that is, in the embodiment of the present application, the interpolation filtering prediction mode provided by the embodiment of the present application can be used to predict both the luminance block and the chroma block.
[0517] In some embodiments, if the current block is a luminance block, the prediction mode of the current block is an interpolation filtering prediction mode, and the chrominance block corresponding to the current block adopts a direct derivation mode DM, then the PLANAR mode or the intra-frame prediction mode corresponding to the above prediction block is determined as the prediction mode of the chrominance block.
[0518] The video decoding method provided in the embodiments of the present application, when predicting a current block, first determines a reference area and an interpolation filter for the current block, and based on the reference area, determines a filter coefficient. Based on the filter coefficient, the interpolation filter is used to perform a parallel prediction on at least two pixels in the current block to obtain a predicted block for the current block. A transform kernel corresponding to the current block is determined, and based on the transform kernel and the predicted block, a reconstructed value for the current block is determined. In other words, in the embodiments of the present application, when using an interpolation filter to perform interpolation filtering prediction on the current block, the at least two points in the current block are predicted in parallel, thereby improving prediction speed and, in turn, decoding efficiency.
[0519] The above describes the prediction method of the present application using the decoding end as an example, and the following describes it using the encoding end as an example.
[0520] FIG28 is a flow chart of a prediction method according to an embodiment of the present application, which is applied to the video encoders shown in FIG1 and FIG2. As shown in FIG28, the method according to the embodiment of the present application includes:
[0521] S201 : Determine a reference area and an interpolation filter of a current block, and determine a filter coefficient of the interpolation filter based on the reference area.
[0522] When encoding the current block, the encoder first determines the prediction mode for the current block and uses it to predict the current block, obtaining the predicted block (or predicted value) for the current block. The current block is then subtracted from the predicted block to obtain the residual block (or residual value) for the current block. The residual block is then transformed to obtain transform coefficients, which are then quantized to obtain quantized coefficients. These quantized coefficients are then encoded to produce the bitstream.
[0523] In the embodiment of the present application, the encoding end first determines the prediction mode of the current block.
[0524] In some embodiments, the encoder determines the prediction mode of the current block in at least the following ways:
[0525] In Method 1, the encoder selects the candidate prediction mode with the lowest cost from among multiple candidate prediction modes, including the traditional prediction mode and the interpolation filter prediction mode, as shown in Figures 6 or 7. The encoder then adds information indicating the prediction mode for the current block to the bitstream. The decoder then decodes the bitstream to obtain the information indicating the prediction mode for the current block and, based on this information, determines the prediction mode for the current block.
[0526] In Method 2, the encoder constructs a candidate list of intra-prediction modes and selects the intra-prediction mode for the current block from this list. It should be noted that this list includes the interpolation filter prediction mode. The encoder then writes the sequence number (or index number) of the intra-prediction mode for the current block in the candidate list into the bitstream.
[0527] Method 3: The encoding end constructs an intra-frame prediction mode candidate list, which includes an interpolation filtering prediction mode. Then, the intra-frame prediction mode of the current block is selected from the intra-frame prediction mode candidate list. For example, the cost of each candidate prediction mode in the intra-frame prediction mode candidate list on the template of the current block is determined, and then the intra-frame prediction mode of the current block is determined based on the cost.
[0528] It can be seen from the above methods that when determining the prediction mode of the current block, the encoder first determines multiple candidate prediction modes, and then determines the prediction mode of the current block from these multiple candidate prediction modes, where the multiple candidate prediction modes include the interpolation filtering prediction mode.
[0529] A specific method for determining the prediction mode of the current block from the multiple candidate prediction modes may be that the encoder determines any one of the multiple candidate prediction modes as the prediction mode of the current block. In other words, the encoder predicts the current block using the multiple candidate prediction modes, determines a cost corresponding to each candidate prediction mode, which may be RDO or SATD, and then determines the candidate prediction mode with the smallest cost as the prediction mode of the current block.
[0530] The encoder determines the prediction mode of the current block based on the above method. If the prediction mode of the current block is the interpolation filtering prediction mode, the above step S201 is executed.
[0531] In some embodiments, the use conditions of the interpolation filter prediction mode are limited. Based on this, before determining the reference area and interpolation filter of the current block, it is determined whether the current image block is allowed to be predicted using the interpolation filter prediction mode.
[0532] The embodiment of the present application does not limit the specific method of determining whether the current image block is allowed to be predicted using the interpolation filter prediction mode, that is, it does not limit the specific use conditions of the interpolation filter prediction mode.
[0533] In some embodiments, before determining the prediction mode of the current block from multiple candidate prediction modes, the encoding end also needs to determine whether the position of the current block in the current image meets the preset position requirement and whether the size of the current block meets the preset block size requirement.
[0534] The embodiment of the present application does not limit the preset position and prediction block size, which are determined based on actual needs.
[0535] In one example, as shown in FIG11 , assuming that the position of the upper left corner of the current image is (0, 0), and the position of the upper left corner of the current block is (x, y), where the preset position requirements are that the x value of the current block is greater than or equal to a first preset value XX, and the y value of the current block is greater than or equal to a second preset value YY.
[0536] The embodiment of the present application does not limit the specific values of the first preset value and the second preset value.
[0537] Exemplarily, the first preset value and the second preset value are the same.
[0538] Exemplarily, the first preset value and the second preset value are both 13, that is, when the distance from the upper edge line of the current block to the upper edge line of the current image is greater than or equal to 13 pixel rows, and the distance from the left edge line of the current block to the left edge line of the current image is greater than or equal to 13 pixel columns, it indicates that the position of the current block in the current image meets the preset position requirements.
[0539] In one example, continuing to refer to Figure 11, assuming that the width of the current block is W and the height of the current block is H, the preset block size requirement is that the width W of the current block is less than or equal to the third preset value A, and the height H of the current block is less than or equal to the fourth preset value B.
[0540] The embodiment of the present application does not limit the specific values of the third preset value and the fourth preset value.
[0541] Exemplarily, the third preset value and the fourth preset value are the same.
[0542] Exemplarily, the third preset value and the fourth preset value are both 32, that is, when the width and height of the current block are both less than or equal to 32, it indicates that the current block meets the preset block size requirement.
[0543] In an embodiment of the present application, before determining whether the current block is predicted using the interpolation filter prediction mode, the encoding end first determines whether the position of the current block in the current image meets the preset position requirement, and determines whether the size of the current block meets the preset block size requirement. If the position of the current block in the current image meets the preset position requirement, and the size of the current block meets the preset block size requirement, then the prediction mode of the current block is determined from the above-mentioned multiple candidate prediction modes including the interpolation filter prediction mode. For example, as shown in Figure 11, the distance from the upper edge of the current block to the upper edge of the current image is greater than or equal to 13 pixel rows, the distance from the left edge of the current block to the left edge of the current image is greater than or equal to 13 pixel columns, and the width and height of the current block are both less than or equal to 32, then the prediction mode of the current block is determined from the above-mentioned multiple candidate prediction modes including the interpolation filter prediction mode.
[0544] In some embodiments, the first preset value, the second preset value, the third preset value, and the fourth preset value are default values.
[0545] In some embodiments, if the position of the current block in the current image does not meet the preset position requirement, and / or the size of the current block does not meet the preset block size requirement, the encoding end determines that the prediction mode of the current block is not the interpolation filtering prediction mode, so that the encoding end determines the prediction mode of the current block from the candidate prediction modes that do not include the interpolation filtering prediction mode.
[0546] In some embodiments, before determining whether the position of the current block in the current image meets the preset position requirement and determining whether the size of the current block meets the preset block size, the encoding end also includes: determining whether the current sequence allows the use of an interpolation filtering prediction mode for prediction; if the current sequence allows the use of an interpolation filtering prediction mode for prediction, then determining whether the position of the current block in the current image meets the preset position requirement and determining whether the size of the current block meets the preset block size.
[0547] In an embodiment of the present application, a high-level syntax element is used to indicate whether the current sequence is allowed to be predicted using the interpolation filtering prediction mode. If the current sequence is predicted using the interpolation filtering prediction mode, the encoder determines whether the position of the current block in the current image meets the preset position requirement and whether the size of the current block meets the preset block size requirement. When it is determined that the position of the current block in the current image meets the preset position requirement and the size of the current block meets the preset block size requirement, the encoder determines the prediction mode for the current block from candidate prediction modes that do not include the interpolation filtering prediction mode.
[0548] In some embodiments, if the encoder determines that the current sequence is not allowed to be predicted using the interpolation filtering prediction mode, the encoder skips step S201 .
[0549] In some embodiments, the encoder writes second information into the bitstream, where the second information is used to indicate whether the current sequence is allowed to be predicted using the interpolation filtering prediction mode.
[0550] The embodiment of the present application does not limit the specific form of the second information, which can be any indication information that can indicate whether the current sequence is allowed to be predicted using the interpolation filtering prediction mode.
[0551] In one example, the second information may be represented by sps_eip_enabled_flag, so that different values of sps_eip_enabled_flag can be assigned to indicate whether the current sequence is allowed to be predicted using the interpolation filtering prediction mode. For example, when sps_eip_enabled_flag = 0, it indicates that the current sequence is not allowed to be predicted using the interpolation filtering prediction mode, and when sps_eip_enabled_flag = 1, it indicates that the current sequence is allowed to be predicted using the interpolation filtering prediction mode.
[0552] Exemplarily, the second information is carried in a sequence parameter set (SPS).
[0553] In some embodiments, embodiments of the present application may further include a general constraints information (GCI) flag to indicate whether interpolation filter prediction technology is used. Exemplarily, gci_no_eip_constraint_flag is used to indicate whether interpolation filter prediction technology is enabled for the current video. Exemplarily, as shown in Table 2, the gci_no_eip_constraint_flag is carried in the general constraints information general_constraints_info().
[0554] As can be seen from the above, whether the current block adopts the interpolation filter prediction mode can be determined by high-level syntax, such as GCI, sequence level, frame level, slice level, block level, etc. It can also be determined by the size and position of the current block.
[0555] In some embodiments, using the interpolation filter prediction mode for smaller blocks increases computational cost and complexity. This is because the interpolation filter prediction mode in this application has a high computational complexity. If the interpolation filter prediction mode is also used for smaller blocks, this will increase the number of times the interpolation filter prediction mode is used during the entire image decoding, thereby increasing the computational cost and complexity of the image. Based on this, in this embodiment of the present application, the interpolation filter prediction mode is only allowed for slightly larger blocks. For example, the interpolation filter prediction mode is only allowed if the size of the current block is greater than or equal to a preset size. If the size of the current block is less than the preset size, the interpolation filter prediction mode is not allowed for the current block. This embodiment of the present application does not limit the specific value of the preset size. For example, the size of the current block being greater than or equal to the preset size can mean that the number of pixels in the current block is greater than or equal to a preset number, or at least one of the length and width of the current block is greater than or equal to a preset value, or the ratio of the length and width of the current block is greater than or equal to a preset ratio, etc.
[0556] In some embodiments, if the current block is in the first row of the current CTU, it is determined that the current block is not allowed to use the interpolation filter prediction mode. In other words, if the current block is predicted using the interpolation filter prediction mode, the current block is not in the first row of the current CTU.
[0557] In some embodiments, determining whether the current block uses the interpolation filter prediction mode is also related to the type of the current image. For example, for intra-frame prediction images, it is stipulated that the interpolation filter prediction mode can be used for prediction, but for inter-frame prediction images, the interpolation filter prediction mode is not allowed. Based on this, if the current image in which the current block is located is an intra-frame prediction image, it is determined that the current block can be predicted using the interpolation filter prediction mode. If the current image is not an intra-frame prediction image (for example, an inter-frame prediction image), it is determined that the current block cannot be predicted using the interpolation filter prediction mode.
[0558] In some embodiments, in the ECM reference software, in order to improve the encoding and decoding performance, a series of complex intra-frame prediction modes are introduced, such as: template-based intra prediction derivation mode (TIMD), decoder-side intra prediction derivation mode (DIMD), template-based multiple reference line intra prediction (TMRL), spatial geometrical partitioning mode (SGPM) and convolutional cross component model (CCCM). These copied intra-frame prediction modes are all intra-frame prediction modes based on template matching technology. The interpolation filter prediction mode of the embodiment of the present application also uses the information of the reconstructed area (which can be understood as the template area) during use. Therefore, in the embodiment of the present application, the interpolation filter prediction mode can also be classified as an intra-frame prediction mode based on template matching technology. Based on this, in the embodiment of the present application, unified identification information (such as the first information) is used to uniformly indicate the above-mentioned intra-frame prediction modes based on template matching technology. For example, if the first information indicates that the template matching-based technology is not enabled, it means that the aforementioned intra-frame prediction modes based on the template matching technology (i.e., TIMD, DIMD, TMRL, SGPM, TMRL, CCCM, and interpolation filter prediction mode) are not allowed to be used. If the first information indicates that the template matching-based technology is not enabled, it means that the aforementioned intra-frame prediction modes based on the template matching technology are allowed to be used, and then based on other information, the specific intra-frame prediction mode used for the current block is further determined.
[0559] Based on the above description, in an embodiment of the present application, determining whether the current block is allowed to use the interpolation filter prediction mode includes: determining whether to obtain first information, the first information being used to indicate whether the template matching-based technology is enabled; and determining whether the current block is allowed to use the interpolation filter prediction mode based on the first information. For example, if the first information indicates that the template matching-based technology is not enabled, then determining that the current block is not allowed to use the interpolation filter prediction mode for prediction. For another example, if the first information indicates that the template matching-based technology is enabled, then determining whether the current block is allowed to use the interpolation filter prediction mode for prediction based on other information.
[0560] The embodiment of the present application does not limit the specific form of expression of the above-mentioned first information.
[0561] Exemplarily, the first information may be GCI, sequence-level, frame-level, slice-level, or block-level indication information.
[0562] In one example, if the first information is sequence-level indication information, the first information is determined. If the first information indicates that the template matching technology is enabled, the second information (sps_eip_enabled_flag) is determined, and then, based on the second information, whether the interpolation filter prediction mode is allowed for the current block is determined. If the first information indicates that the template matching technology is not enabled, it is directly determined that the interpolation filter prediction mode is not applicable for prediction of the current block, and the step of determining the second information is skipped.
[0563] In some embodiments, the condition for the encoder to determine whether the current block can be predicted using the interpolation filtering prediction mode includes at least one of the following:
[0564] 1) Whether the current image is an intra-frame prediction image;
[0565] 2) Whether high-level syntax is allowed; optional, high-level syntax includes sequence level, frame level, slice level, block level, etc., refer to the above description for details;
[0566] 3) Whether the size and shape of the current block are allowed; refer to the above description for details;
[0567] 4) Whether the position of the current block is allowed; refer to the above description for details.
[0568] In some embodiments, if the encoder determines that the current series allows prediction using the interpolation filtering prediction mode, third information is written into the bitstream, where the third information is used to indicate whether the current block is predicted using the interpolation filtering prediction mode.
[0569] The embodiment of the present application does not limit the specific form of expression of the above-mentioned third information, which can be any indication information that can indicate whether the current block adopts the interpolation filtering prediction mode for prediction.
[0570] In one example, the third information can be represented as intra_eip_flag, so that different values of intra_eip_flag can be assigned to indicate whether the current block is predicted using the interpolation filtering prediction mode. For example, when intra_eip_flag = 0, it indicates that the current block is not predicted using the interpolation filtering prediction mode, and when intra_eip_flag = 1, it indicates that the current block is predicted using the interpolation filtering prediction mode. In this way, the encoder writes the preset flag intra_eip_flag into the bitstream, and the decoder determines the prediction mode of the current block by decoding the value of the preset flag intra_eip_flag. For example, when the preset flag intra_eip_flag = 1, it indicates that the prediction mode of the current block is the interpolation filtering prediction mode, and the decoder then uses the interpolation filtering prediction mode to predict the current block.
[0571] In some embodiments, as shown in FIG29 , the process of determining the prediction mode of the current block in an embodiment of the present application may include: first, determining whether the current block is predicted using the interpolation filter prediction mode. For example, if the second information at the sequence level indicates that the current sequence allows the use of the interpolation filter prediction mode, and if it is determined that the position of the current block in the current image meets the preset position requirement, and if it is determined that the size of the current block meets the preset block size requirement, then it is determined that the current block can be predicted using the interpolation filter prediction mode. Next, the filter coefficient is obtained, and the current block is predicted based on the filter coefficient to obtain the predicted value of the current block. At the same time, a coarse screening of the prediction mode is performed with other intra-frame prediction mode tools, and several prediction modes with a smaller cost are selected for fine screening to determine the final intra-frame prediction mode as the prediction mode for the current block. If it is determined that the current block cannot be predicted using the interpolation filter prediction mode, the screening of the interpolation filter prediction mode is skipped.
[0572] For example, during the coarse screening of the prediction mode of the current block, the encoder calculates the cost of each candidate intra prediction mode (including the interpolation filter prediction mode). The cost calculation formula is shown in Formula (10): cost = D + λR (10)
[0573] Where R represents the bit overhead expected for the intra-frame prediction mode, λ is the Lagrange multiplier, which is related to the quantization parameter used in the current encoding, and D represents the distortion value between the predicted block and the original block in the current prediction mode.
[0574] In one example, the calculation of the distortion value D is as shown in formula (11): D = min (SAD × 2, SATD) (11)
[0575] Among them, SAD (The sum of absolute difference) and SATD (The sum of transformed difference) represent the absolute error sum algorithm and Hadamard transform error sum algorithm between the prediction and the original block, respectively.
[0576] After the encoder determines the cost of each candidate prediction mode, it selects several candidate prediction modes from multiple candidate prediction modes for detailed screening.
[0577] The prediction modes that pass the rough screening described above undergo a complete residual transform, quantization, inverse quantization, inverse transform, and reconstruction. The rate-distortion cost of each mode combination (prediction mode + transform mode + quantization mode) is compared to determine the final prediction mode, transform mode, and quantized residual value. The rate-distortion cost calculation is still D + λR, but here D represents the SSE (sum of squared error) between the reconstructed block and the original block, and R represents the total bit overhead of encoding the current block's mode identifier, coefficients, and so on.
[0578] The encoder determines the candidate prediction mode with the lowest cost during the fine screening process as the prediction mode for the current block.
[0579] If the encoder determines that the prediction mode of the current block is the interpolation filtering prediction mode, the above step S201 is executed.
[0580] The following describes the process of using the interpolation filter prediction mode at the encoder to predict the current block.
[0581] When the encoder determines that the current block is predicted using the interpolation filtering prediction mode, it first determines the reference area and interpolation filter of the current block.
[0582] The following describes the specific process of the encoder determining the reference area of the current block.
[0583] In the embodiment of the present application, the reference area of the current block is part or all of the reconstructed area around the current block.
[0584] Exemplarily, as shown in FIG12 , the reconstruction area around the current block may include: an upper reconstruction area of the current block, a left reconstruction area of the current block, an upper right reconstruction area of the current block, a lower left reconstruction area of the current block, and an upper left reconstruction area of the current block.
[0585] The embodiment of the present application does not limit the specific shape and size of the reference area of the current block.
[0586] In one example, the reference area of the current block includes any one of an upper reconstruction area of the current block, a left reconstruction area of the current block, an upper right reconstruction area of the current block, a lower left reconstruction area of the current block, and an upper left reconstruction area of the current block. For example, the reference area of the current block is the upper reconstruction area of the current block, or the reference area of the current block is the left reconstruction area of the current block.
[0587] In one example, the reference area of the current block includes any two of the reconstruction areas above the current block, the reconstruction area to the left of the current block, the reconstruction area to the upper right of the current block, the reconstruction area to the lower left of the current block, and the reconstruction area to the upper left of the current block. For example, the reference area of the current block includes the reconstruction area above the current block and the reconstruction area to the left of the current block. For another example, the reference area of the current block includes the reconstruction area above the current block and the reconstruction area to the lower left of the current block.
[0588] In one example, the reference area of the current block includes any three reconstruction areas of the upper reconstruction area of the current block, the left reconstruction area of the current block, the upper right reconstruction area of the current block, the lower left reconstruction area of the current block, and the upper left reconstruction area of the current block. For example, the reference area of the current block includes the upper reconstruction area of the current block, the upper right reconstruction area of the current block, and the upper left reconstruction area of the current block. For another example, the reference area of the current block includes the left reconstruction area of the current block, the upper left reconstruction area of the current block, and the lower left reconstruction area of the current block.
[0589] In one example, the reference area of the current block includes any four reconstruction areas of the upper reconstruction area of the current block, the left reconstruction area of the current block, the upper right reconstruction area of the current block, the lower left reconstruction area of the current block, and the upper left reconstruction area of the current block. For example, the reference area of the current block includes the upper reconstruction area of the current block, the upper right reconstruction area of the current block, the upper left reconstruction area of the current block, and the left reconstruction area of the current block. For another example, the reference area of the current block includes the left reconstruction area of the current block, the upper left reconstruction area of the current block, the lower left reconstruction area of the current block, and the upper reconstruction area of the current block.
[0590] In one example, the reference area of the current block includes five reconstruction areas: an upper reconstruction area of the current block, a left reconstruction area of the current block, an upper right reconstruction area of the current block, a lower left reconstruction area of the current block, and an upper left reconstruction area of the current block.
[0591] In the embodiment of the present application, the specific manners in which the encoder determines the reference area of the current block include but are not limited to the following:
[0592] In method 1, the reference area of the current block is a default area. For example, the encoding end and the decoding end default that the reference area of the current block includes at least one reconstruction area of the upper reconstruction area of the current block, the left reconstruction area of the current block, the upper right reconstruction area of the current block, the lower left reconstruction area of the current block, and the upper left reconstruction area of the current block.
[0593] Method 2: determining first costs for predicting the current block based on P reference regions respectively; determining the reference region with the smallest first cost among the P reference regions as the reference region for the current block.
[0594] In this implementation, the encoding end predicts the current block based on the P reference areas, determines the first cost corresponding to each reference area, and then determines the reference area with the smallest first cost among the P reference areas as the reference area of the current block.
[0595] In some embodiments, the encoder writes fourth information into the bitstream, where the fourth information indicates the type of the reference region of the current block. That is, in this approach 2, the encoder also indicates the determined type of the reference region of the current block to the decoder via the fourth information.
[0596] It should be noted that the types or shapes of the preset P reference areas are different.
[0597] The embodiment of the present application does not impose any specific limitation on the specific number and shape of the P reference areas.
[0598] In one example, the P reference regions include at least one of a first reference region, a second reference region, and a third reference region.
[0599] As shown in FIG13A , the first reference area includes the reconstruction areas above, to the upper right, to the left, to the upper left, and to the upper left of the current block. As shown in FIG13B , the second reference area includes the reconstruction areas above, to the upper right, and to the upper left of the current block. As shown in FIG13C , the third reference area includes the reconstruction areas to the left, to the upper left, and to the upper left of the current block.
[0600] The embodiment of the present application does not limit the specific form of the fourth information, as long as it is any indication information that can indicate the type of the reference area of the current block.
[0601] In an example, eip_ref_type is used to represent the fourth information. For example, different types of reference areas are indicated by the value of eip_ref_type.
[0602] Exemplarily, the correspondence between the three reference areas shown in FIG. 13A and FIG. 13B and the eip_ref_type values is shown in Table 4.
[0603] Based on Table 4 above, the encoder determines the value of the fourth information eip_ref_type based on the reference region type of the current block. For example, if the reference region of the current block is determined to be the first reference region, eip_ref_type is determined to be 0. If the reference region of the current block is determined to be the second reference region, eip_ref_type is determined to be 1. If the reference region of the current block is determined to be the third reference region, eip_ref_type is obtained to be 2.
[0604] It should be noted that the above description uses the three reference regions shown in Figures 13A to 13C as an example. The P reference regions in this embodiment of the present application also include other reference regions in addition to the three reference regions described above, and this embodiment of the present application does not limit this. The correspondence between the reference regions and the eip_ref_type values shown in Table 4 above can be adaptively adjusted according to the number of reference regions.
[0605] In some embodiments, the encoding end may use a truncated binary code encoding method to write the fourth information into the code stream.
[0606] For example, the correspondence between the truncated binary code, the eip_ref_type value, and the type of the reference area is shown in Table 5.
[0607] In an embodiment of the present application, the encoding end may use an equal probability encoding method or a context model encoding method to encode the codeword of the truncated binary code.
[0608] In addition to using the above-mentioned method 1 or method 2 to determine the reference area of the current block, the encoder can also use the following method 3 to determine the reference area of the current block.
[0609] Mode 3: Based on the shape of the current block, a reference area of the current block is determined from among P preset reference areas.
[0610] In this method 3, different reference areas are used to perform prediction on current blocks of different shapes to improve the accuracy of prediction.
[0611] For example, if the shape of the current block is a square, the first type of reference region is used.
[0612] For another example, if the shape of the current block is a rectangle with a width greater than a height, the second type of reference region is used.
[0613] For another example, if the shape of the current block is a rectangle with a width smaller than a height, the third type of reference region is used.
[0614] That is, in this embodiment of the present application, the correspondence between the P reference regions and the shape of the current block is preset. Thus, the encoder can determine the reference region of the current block from among the P reference regions based on the shape of the current block and the correspondence between the P reference regions and the shape of the current block.
[0615] The following describes the process of determining the interpolation filter of the current block at the encoder end.
[0616] In the embodiment of the present application, there is no limitation on the specific shape of the interpolation filter.
[0617] Illustratively, the interpolation filters provided in the embodiments of the present application include, but are not limited to, square interpolation filters and interpolation filters with a height smaller than a width.
[0618] For example, the square interpolation filter includes but is not limited to the 4X4 interpolation filter shown in Figure 14A.
[0619] For another example, interpolation filters that are taller than they are wide include, but are not limited to, the 5X3 interpolation filter shown in FIG. 14B , the 6X2 interpolation filter shown in FIG. 14D , and the 7X1 interpolation filter shown in FIG. 14G .
[0620] For another example, interpolation filters with a height smaller than a width include but are not limited to the 3X5 interpolation filter shown in FIG. 14C , the 2X6 interpolation filter shown in FIG. 14E , and the 1X7 interpolation filter shown in FIG. 14F .
[0621] It should be noted that in the above filter, the dark gray position represents the current position to be predicted, and the light gray position represents the input position of the interpolation filter, that is, {p0, p1, ..., p N-1}Location.
[0622] In the embodiment of the present application, the specific manners in which the encoder determines the interpolation filter for the current block include but are not limited to the following:
[0623] In mode 1, the interpolation filter of the current block is a default interpolation filter. For example, the encoder and decoder default the interpolation filter of the current block to any one of the interpolation filters in Figures 14A to 14G. For example, the default interpolation filter is a 4x4 interpolation filter.
[0624] In mode 2, the encoder determines an interpolation filter for the current block from among Q preset interpolation filters.
[0625] For example, the encoder randomly selects an interpolation filter from Q interpolation filters as the interpolation filter of the current block.
[0626] For another example, the encoder determines the second costs when using Q interpolation filters to predict the current block respectively; and determines the interpolation filter with the smallest second cost among the Q interpolation filters as the interpolation filter for the current block.
[0627] In some embodiments, the encoding end writes fifth information into the bitstream, where the fifth information is used to indicate the shape of the interpolation filter of the current block.
[0628] In this implementation, the encoder determines the interpolation filter for the current block from among the Q preset interpolation filters. For example, the encoder determines the second costs corresponding to each of the Q interpolation filters and selects the interpolation filter with the smallest second cost as the interpolation filter for the current block. The shape of the interpolation filter with the smallest second cost is then indicated to the encoder via fifth information. The decoder then decodes the bitstream to obtain the fifth information and, based on the shape of the interpolation filter indicated by the fifth information, determines the interpolation filter for the current block from among the Q preset interpolation filters.
[0629] It should be noted that the shapes of the preset Q interpolation filters are different.
[0630] The embodiments of the present application do not impose any specific restrictions on the specific number and shape of the Q interpolation filters. For example, the Q interpolation filters include at least one of a first interpolation filter, a second interpolation filter, and a third interpolation filter, where the first interpolation filter is a square interpolation filter, the second interpolation filter is a rectangular interpolation filter with a width greater than a height, and the third interpolation filter is a rectangular interpolation filter with a height greater than a width.
[0631] In one example, the Q interpolation filters include the plurality of interpolation filters in FIG. 14A to FIG. 14H .
[0632] The embodiment of the present application does not limit the specific form of the fifth information, as long as it is any indication information that can indicate the shape of the interpolation filter of the current block.
[0633] In an example, eip_filter_type is used to represent the fifth information. For example, the value of eip_filter_type is used to indicate interpolation filters of different shapes.
[0634] Exemplarily, if the Q interpolation filters are the five interpolation filters shown in FIG. 15 , the correspondence between the five interpolation filters and the eip_filter_type values is as shown in Table 6.
[0635] Based on Table 5 above, the encoder determines the value of the fifth information eip_filter_type based on the shape of the interpolation filter of the current block. For example, if the shape of the interpolation filter of the current block is determined to be 4x4, eip_filter_type is determined to be 0. If the shape of the interpolation filter of the current block is determined to be 3x5, eip_filter_type is determined to be 1. If the shape of the interpolation filter of the current block is determined to be 5x3, eip_filter_type is determined to be 2. If the shape of the interpolation filter of the current block is determined to be 2x6, eip_filter_type is determined to be 3. If the shape of the interpolation filter of the current block is determined to be 6x2, eip_filter_type is determined to be 4.
[0636] In some embodiments, the encoding end may use a truncated binary code encoding method to encode the fifth information into the code stream.
[0637] Exemplarily, if the preset Q interpolation filters include the five interpolation filters shown in FIG. 15 , the correspondence between the truncated binary code, the eip_filter_type value, and the shape of the interpolation filter is shown in Table 7.
[0638] At this time, the five interpolation filter shapes shown in Table 7 and the three reconstruction region types shown in Table 5 provide a total of 15 combinations of interpolation filters and reconstruction regions.
[0639] In some embodiments, when the embodiment of the present application includes the seven interpolation filters shown in Figure 16, the correspondence between the truncated binary code, the eip_filter_type value and the shape of the interpolation filter is shown in Table 9.
[0640] At this time, the 7 interpolation filter shapes shown in Table 9 and the 3 reconstruction region types shown in Table 5 provide a total of 21 combinations of interpolation filters and reconstruction regions.
[0641] In some embodiments, when the embodiment of the present application includes three interpolation filters as shown in Figure 17, the correspondence between the truncated binary code, the eip_filter_type value and the shape of the interpolation filter is shown in Table 10.
[0642] At this time, the three interpolation filter shapes shown in Table 10 and the three reconstruction region types shown in Table 5 provide a total of nine combinations of interpolation filters and reconstruction regions.
[0643] In some embodiments, when the embodiment of the present application includes three interpolation filters as shown in Figure 18A, the correspondence between the truncated binary code, the eip_filter_type value and the shape of the interpolation filter is shown in Table 11.
[0644] At this time, the three interpolation filter shapes shown in Table 11 and the three reconstruction region types shown in Table 5 provide a total of nine combinations of interpolation filters and reconstruction regions.
[0645] In some embodiments, when the embodiment of the present application includes three interpolation filters as shown in Figure 18B, the correspondence between the truncated binary code, the eip_filter_type value and the shape of the interpolation filter is shown in Table 12.
[0646] At this time, the three interpolation filter shapes shown in Table 12 and the three reconstruction region types shown in Table 5 provide a total of nine combinations of interpolation filters and reconstruction regions.
[0647] In some embodiments, when the embodiment of the present application includes the three interpolation filters shown in Figure 19, the correspondence between the truncated binary code, the eip_filter_type value and the shape of the interpolation filter is shown in Table 13.
[0648] At this time, the three interpolation filter shapes shown in Table 13 and the three reconstruction region types shown in Table 5 provide a total of nine combinations of interpolation filters and reconstruction regions.
[0649] In addition to using the above-mentioned method 1 or method 2 to determine the interpolation filter of the current block, the encoder can also use the following method 3 to determine the interpolation filter of the current block.
[0650] Mode 3: Based on the shape of the current block, an interpolation filter for the current block is determined from among Q preset interpolation filters.
[0651] In this method 3, different interpolation filters are used to perform prediction on current blocks of different shapes to improve the accuracy of prediction.
[0652] For example, if the shape of the current block is a square, an interpolation filter of the first shape is used.
[0653] For another example, if the shape of the current block is a rectangle with a width greater than a height, the interpolation filter of the second shape is used.
[0654] For another example, if the shape of the current block is a rectangle with a width smaller than a height, the interpolation filter of the third shape is used.
[0655] That is, in the embodiment of the present application, the correspondence between the Q interpolation filters and the shape of the current block is preset. In this way, the encoder can determine the interpolation filter for the current block from the Q interpolation filters based on the shape of the current block and the correspondence between the Q interpolation filters and the shape of the current block.
[0656] In an embodiment of the present application, after the encoder determines the reference area and interpolation filter of the current block based on the above steps, it determines the prediction block of the current block based on the reference area and the interpolation filter.
[0657] The following describes how to determine the filter coefficients of the interpolation filter based on the reference area.
[0658] In the embodiment of the present application, the methods for determining the filter coefficients of the interpolation filter include at least the following:
[0659] Method 1: Use the interpolation filter determined above to slide over the reference area of the current block to construct the Wienerhof equation. Then, solve the Wienerhof equation to obtain the filter coefficients of the interpolation filter.
[0660] In the process of sliding the interpolation filter in the reference area of the current block, the N positions corresponding to each position in the reference area are determined based on the shape of the interpolation filter. For example, for position r in the reference area, based on the shape of the interpolation filter, the N positions corresponding to position r are determined in the reference area. The pixel reconstruction values of these N positions are the input of the interpolation filter. The phase position difference between these N positions and position r is {p0, p1, ..., p N-1}, p N is a two-dimensional representation. In which {c0,c1,…,c N-1} is {p0,p1,…,p N-1}Interpolation filter coefficients at the position.
[0661] In one example, the interpolation filter is slid in the reference area of the current block to construct the Wienerhof equation, as shown in formula (3).
[0662] Since the reference area of the current block is the reconstructed area, all parameters except the interpolation filter coefficient in the above formula (3) are known. Therefore, the filter coefficient of the interpolation filter of the current block can be determined by solving the above formula (3).
[0663] In one example, the encoding end may solve the Wienerhof equation shown in the above formula (3) by Cholesky decomposition of the autocorrelation coefficient matrix to obtain the filter coefficients of the filter.
[0664] The embodiment of the present application does not limit the sliding step size of the interpolation filter within the reference area.
[0665] In one example, as shown in FIG20A , the horizontal sliding step size and the vertical sliding step size of the interpolation filter in the reference area are equal, both being 1 pixel.
[0666] In one example, the horizontal sliding step size and the vertical sliding step size of the interpolation filter in the reference area are not equal. For example, the horizontal sliding step size is 2 pixels and the vertical sliding step size is 1 pixel. For another example, the horizontal sliding step size is 1 pixel and the vertical sliding step size is 2 pixels.
[0667] In one example, at least one of the horizontal sliding step size and the vertical sliding step size of the interpolation filter within the reference region is greater than a preset step size. For example, the horizontal sliding step size is greater than the preset step size. For another example, the vertical sliding step size is greater than the preset step size. For another example, both the horizontal sliding step size and the vertical sliding step size are greater than the preset step size. This embodiment of the present application does not limit the specific value of the preset step size. For example, it can be 1, 2, 3, etc.
[0668] In the second method, the encoder determines the filter coefficients through the following steps S201-A1 to S201-A4:
[0669] S201-A1, determining a first reconstruction area around the current block;
[0670] S201-A2, determining a pixel average reconstruction value based on the reconstruction value of the first reconstruction area;
[0671] S201-A3, based on the pixel average reconstruction value, removing the mean of the reconstructed values of the pixels in the reference area;
[0672] S201-A4: Using the pixel values of the pixels in the reference area after averaging as inputs of the interpolation filter, sliding the interpolation filter in the reference area to obtain filter coefficients of the interpolation filter.
[0673] In this second approach, the reference region is de-averaged, and the interpolation filter coefficients are determined based on the de-averaged reference region. Since the amount of data in the de-averaged reference region is reduced, determining the filter coefficients based on the de-averaged reference region can improve efficiency.
[0674] Specifically, the encoding end first determines a first reconstruction area, where the first reconstruction area may be any part of the reconstruction area surrounding the current block.
[0675] In the embodiment of the present application, the encoder determines the first reconstruction area around the current block in at least the following ways:
[0676] In mode 1, the encoder determines a reconstruction area around the current block as the first reconstruction area by default.
[0677] For example, as shown in FIG20B , the encoder uses, by default, an area consisting of a row above, a column to the left, and a pixel point in the upper left corner of the current block as the first reconstruction area.
[0678] Method 2: Determine the first reconstruction area based on the shape of the current block.
[0679] For example, if the shape of the current block is a square, the reconstructed pixel area in one row above and one column on the left of the current block is determined as the first reconstructed area.
[0680] For another example, if the shape of the current block is a rectangle with a width greater than a height, a row of reconstructed pixel areas above the current block is determined as the first reconstructed area.
[0681] For another example, if the shape of the current block is a rectangle with a height greater than a width, a left column of reconstructed pixel areas of the current block is determined as the first reconstructed area.
[0682] It should be noted that, based on the shape of the current block, the manner of determining the first reconstruction area includes but is not limited to the above examples.
[0683] After determining the first reconstruction area, the encoder determines the pixel average reconstruction value m based on the reconstruction value of the first reconstruction area.
[0684] The embodiment of the present application does not limit the specific method of determining the pixel average reconstruction value m based on the reconstruction value of the first reconstruction area in the above S201-A2.
[0685] In mode 1, the above S201 - A2 includes: determining the average value of the reconstruction values of the first reconstruction area as the pixel average reconstruction value m.
[0686] In an example of method 1, if the first reconstruction area is as shown in FIG. 20B , the pixel average reconstruction value m can be calculated using the method shown in Table 13.
[0687] In one example of Method 1, if the first reconstructed area is the row above and / or the column to the left of the current block, the average of the reconstructed values in the row above and / or the column to the left can be determined as the pixel average reconstructed value m. In this case, the pixel average reconstructed value m can be calculated using the method shown in Table 15.
[0688] As shown in Table 15 above, if the first reconstruction area is a row above and / or a column to the left of the current block, shift calculation can be used instead of division to quickly calculate the pixel average reconstruction value m.
[0689] Mode 2, the above S201-A2 includes: determining a pixel average reconstruction value based on the shape of the current block and the reconstruction value of the first reconstruction area.
[0690] For example, if the current block is a square, the average value of the entire first reconstruction area determined above is determined as the pixel average reconstruction value m.
[0691] In an example of method 2, the first reconstructed area includes an upper reconstructed area and a left reconstructed area of the current block. At this time, determining the pixel average reconstruction value based on the shape of the current block and the reconstruction value of the first reconstructed area includes: determining the first area from the upper reconstructed area and the left reconstructed area based on the shape of the current block; determining the average reconstruction value of the first area based on the reconstruction value of the first area; and determining the pixel average reconstruction value based on the average reconstruction value of the first area.
[0692] In this example, if the first reconstructed region includes the reconstructed region above and to the left of the current block, to reduce computational duplication, the mean is determined in the same manner as in the DC prediction mode. Specifically, based on the shape of the current block, the first region is determined from the reconstructed region above and to the left of the first reconstructed region.
[0693] For example, if the shape of the current block is that the width is greater than the height, the upper reconstruction area is determined as the first area.
[0694] For example, if the shape of the current block is that the height is greater than the width, the left reconstructed area is determined as the first area.
[0695] For example, if the shape of the current block is that the height is equal to the width, the upper reconstruction area and the left reconstruction area are determined as the first area.
[0696] Next, based on the reconstruction value of the selected first area, the average reconstruction value of the first area is determined, and then based on the average reconstruction value of the first area, the pixel average reconstruction value m is determined, for example, the average reconstruction value of the first area is determined as the pixel average reconstruction value m.
[0697] That is, in this example, if the current block is wider than it is tall, the average reconstructed value of the reconstructed area above the current block is determined as the pixel average reconstructed value m. If the current block is taller than it is wide, the average reconstructed value of the reconstructed area to the left of the current block is determined as the pixel average reconstructed value m. If the current block is taller than it is wide, the average reconstructed value of the reconstructed areas above and to the left of the current block is determined as the pixel average reconstructed value m.
[0698] After determining the pixel average reconstruction value, the encoder performs de-averaging on the reconstruction values of the pixels in the reference area based on the pixel average reconstruction value.
[0699] For example, for each pixel in the reference area, the reconstructed value of the pixel is divided by the pixel average reconstructed value and then rounded to the integer to obtain the pixel value of the pixel in the reference area after removing the average value.
[0700] For another example, the encoder subtracts the average pixel reconstruction value from the reconstructed value of the pixel in the reference area to obtain the averaged pixel value of the pixel in the reference area. For example, for each pixel in the reference area, the average pixel reconstruction value is subtracted from the reconstructed value of the pixel to obtain the averaged pixel value of the pixel in the reference area.
[0701] The embodiment of the present application does not limit the specific method of de-averaging the reconstructed values of the pixels in the reference area based on the pixel average reconstruction value at the encoding end.
[0702] Based on the above method, the encoding end de-averages the reconstructed values of the pixel points in the reference area, obtains the pixel values of the de-averaged pixel points in the reference area, and then executes the above steps S201-A4, uses the pixel values of the de-averaged pixel points in the reference area as the input of the interpolation filter, slides the interpolation filter in the reference area, and obtains the filter coefficients of the interpolation filter.
[0703] For example, as shown in FIG21, when the interpolation filter of the current block is an interpolation filter of five different shapes and the reference area of the current block is a reference area of three different types, the interpolation filter of the current block is slid on the reference area after the current block is averaged to obtain the filter coefficients of the interpolation filter. The interpolation filter can slide horizontally row by row or vertically column by column on the reference area after the current block is averaged.
[0704] As shown in Figure 21, when using the interpolation filter to slide in the reference area of the current block, first, according to the shape of the interpolation filter, the N positions corresponding to each position in the reference area are determined. For example, for position r in the reference area, based on the shape of the interpolation filter, the N positions corresponding to position r are determined in the reference area, and the pixel reconstruction values of these N positions are the input of the interpolation filter. The phase position difference between these N positions and position r is {p0, p1, ..., p N-1}, p N is a two-dimensional representation. In which {c0,c1,…,c N-1} is {p0,p1,…,p N-1}Interpolation filter coefficients at the position.
[0705] In one example, the interpolation filter is slid in the reference area of the current block, and the constructed Wienerhof equation is shown in formula (4).
[0706] Since the reference area of the current block is the reconstructed area, all parameters except the interpolation filter coefficient in the above formula (4) are known. Therefore, the filter coefficient of the interpolation filter of the current block can be determined by solving the above formula (4).
[0707] In one example, the encoding end may solve the Wienerhof equation shown in the above formula (4) by Cholesky decomposition of the autocorrelation coefficient matrix to obtain the filter coefficients of the filter.
[0708] After the encoder determines the filter coefficients of the interpolation filter based on the above steps, it executes the following step S202.
[0709] S202 : Based on the filter coefficients, use an interpolation filter to perform parallel prediction on at least two pixels in the current block to determine a prediction block of the current block.
[0710] After determining the filter coefficients of the interpolation filter based on the above steps, the encoder uses the interpolation filter to perform interpolation filtering prediction on the current block based on the filter coefficients to obtain a predicted block of the current block.
[0711] In related art, when using an interpolation filter to perform interpolation filtering prediction on the current block, after the prediction of one pixel in the current block is completed, the next pixel is predicted. For example, as shown in FIG22A , the interpolation filter performs interpolation filtering prediction on each pixel in the current block one by one along the horizontal direction. During prediction, after the prediction of the previous pixel in the horizontal direction is completed, the predicted value of the previous pixel is used as an input pixel value of the interpolation filter of the next pixel to predict the next pixel. As shown in FIG22 , assuming that the interpolation filter of the current block has a 4×4 shape, the encoder uses an interpolation filter with known filter coefficients to perform interpolation prediction on each position in the current block one by one. Specifically, for the rth point in the current block, the pixel values of the N positions corresponding to the rth point are first determined based on the shape of the interpolation filter of the current block. For example, as shown in FIG22 , in the 4×4 interpolation filter, the dark position is the position of the rth point to be processed, and the 15 light positions are the N positions corresponding to the rth point. The block to be predicted in FIG22 is the current block. Next, the pixel values of the N positions corresponding to the rth point are determined. For example, for any of the N positions, if the position is located in the reconstructed area surrounding the current block, the reconstructed value of the position is determined as the pixel value of the position. If the position is located within the current block, the predicted value of the position is determined as the pixel value of the position.
[0712] For another example, as shown in Figure 22B, the interpolation filter performs interpolation filtering prediction on each pixel in the current block one by one along the vertical direction. During prediction, after the prediction of the previous pixel in the vertical direction is completed, the predicted value of the previous pixel is used as an input pixel value of the interpolation filter of the next pixel to predict the next pixel. In other words, when the related art uses the interpolation filter to perform interpolation filtering prediction on the current block, it can only complete the prediction of one pixel at a time, resulting in long prediction time and low prediction efficiency, which in turn affects coding efficiency.
[0713] In order to solve the above technical problems, the embodiment of the present application uses an interpolation filter to perform interpolation filtering prediction on the current block, and performs parallel prediction on at least two points in the current block. That is, the encoding end can use the interpolation filter to perform interpolation filtering prediction on at least two pixel points in the current block at the same time.
[0714] In the embodiment of the present application, the encoder uses an interpolation filter to perform interpolation filtering prediction on at least two pixels in the current block at the same time, including at least two implementation methods:
[0715] In a first implementation, for at least two pixels in a current block, these at least two pixels are considered adjacent pixels. The encoder first determines interpolation filter input information corresponding to the at least two pixels, inputs the input information into the interpolation filter for interpolation filter prediction, obtains a prediction value, and then determines the prediction value of the at least two pixels based on the prediction value. For example, based on relevant feature information of the at least two pixels, the prediction value is processed to obtain the prediction value corresponding to the at least two pixels. In another example, the prediction value is determined as the prediction value corresponding to the at least two pixels.
[0716] In the first implementation manner, the encoding end does not impose any restrictions on the specific method of determining the interpolation filter input information corresponding to the at least two pixel points.
[0717] For example, based on the shape of the interpolation filter, determine the interpolation filter input value that is the same for the at least two pixels, and use the same interpolation filter input value as the interpolation filter input information. For example, the at least two pixels include pixel 1 and pixel 2. Based on the shape of the interpolation filter, determine the N input values corresponding to pixel 1 and the N input values corresponding to pixel 2. Determine the same input value among the N input values corresponding to pixel 1 and the N input values corresponding to pixel 2, and use the same input value as the input value of the interpolation filter. It should be noted that if the N input values corresponding to pixel 2 or pixel 1 include an uncoded value, the uncoded value is discarded.
[0718] For another example, based on the shape of the interpolation filter, the input value corresponding to the pixel with the most encoded input values among the at least two pixels is determined as the interpolation filter input information. For example, the at least two pixels include pixel 1 and pixel 2. Based on the shape of the interpolation filter, the N input values corresponding to pixel 1 and the N input values corresponding to pixel 2 are determined. The N input values corresponding to pixel 1 are all encoded, and the N input values corresponding to pixel 2 include unencoded input values. Therefore, the N input values corresponding to pixel 1 are used as the interpolation filter input information.
[0719] In the above-mentioned method 1, the specific process of the interpolation filter predicting the predicted values of at least two pixels in the current block after the encoder inputs the same input value to the interpolation filter at one time is introduced.
[0720] The second method is to use an interpolation filter to perform interpolation filtering on at least two pixels in the current block at the same time. For example, at time t, the encoder uses an interpolation filter to perform interpolation filtering and prediction on pixel 1 in the current block to obtain the predicted value of pixel 1. At the same time, the encoder uses an interpolation filter to perform interpolation filtering and prediction on pixel 2 in the current block to obtain the predicted value of pixel 2.
[0721] The embodiment of the present application does not restrict the prediction direction when the encoder uses the interpolation filter to perform parallel prediction on at least two pixels in the current block.
[0722] In some embodiments, the encoder may use an interpolation filter to perform parallel prediction on at least two pixels in the current block along a horizontal direction. For example, when performing parallel prediction on at least two pixels, if one or more input values of the interpolation filter corresponding to the pixel are not encoded, the unencoded input values are discarded and the encoded input values are used as input values of the interpolation filter.
[0723] In some embodiments, the encoder may use an interpolation filter to perform parallel prediction on at least two pixels in the current block along a vertical direction. For example, when performing parallel prediction on at least two pixels, if one or more input values of the interpolation filter corresponding to the pixel are not encoded, the unencoded input values are discarded and the encoded input values are used as input values of the interpolation filter.
[0724] In some embodiments, the encoding end may use an interpolation filter to perform parallel prediction on at least two pixels in the current block along a diagonal direction. Based on this, in one example, the above S202 includes the following step S202-A:
[0725] S202-A: Based on the filter coefficients, use an interpolation filter along the diagonal direction to perform parallel interpolation filtering prediction on the pixel points on the same diagonal line of the current block to obtain a prediction block of the current block.
[0726] Exemplarily, as shown in FIG23, when the encoder uses an interpolation filter to predict a pixel in the current block, the pixel to be predicted is located at a corner of the interpolation filter selected area (e.g., the lower right corner or the upper left corner). Thus, for any pixel on the same angle line, based on the shape of the interpolation filter, the N positions corresponding to the selected pixel do not include the positions of other pixels on the diagonal line. In other words, the N positions corresponding to each pixel on the same diagonal line do not include the pixels on the diagonal line. For example, taking two adjacent pixels a and b on the same diagonal line of the current block as an example, as shown in FIG23, based on the shape of the interpolation filter, the N positions corresponding to pixel a and the N positions corresponding to pixel b are determined respectively, wherein the N positions corresponding to pixel a and the N positions corresponding to pixel b both include the pixels on the diagonal line. Based on this, when the encoder uses an interpolation filter to perform interpolation filtering prediction on the current block, it can perform parallel interpolation filtering prediction on the pixels on the same diagonal line in the current block along the diagonal direction. For example, parallel interpolation filtering prediction is performed on pixel point a and pixel point b located on the same diagonal line.
[0727] In an embodiment of the present application, the left, lower left, upper left, left and upper right areas of the current block have been encoded. Therefore, the starting point of the prediction of the current block along the diagonal direction can be determined based on the encoded area and the shape of the interpolation filter.
[0728] In one example, as shown in FIG23 , the encoder performs interpolation filtering prediction on the current block along the diagonal direction starting from the upper left corner of the current block. The above S202-A includes the following steps S202-A1:
[0729] S202-A1: The encoder uses the interpolation filter to perform parallel interpolation filtering prediction on pixels on the same diagonal line of the current block, starting from the upper left corner of the current block, based on the filter coefficients, to obtain a predicted block for the current block. In this example, the pixel to be predicted is located in the lower right corner of the area selected by the interpolation filter.
[0730] The embodiment of the present application does not limit the specific direction of the diagonal line.
[0731] In some embodiments, as shown in FIG23 , when the encoding end starts from the upper left corner of the current block and performs parallel prediction on the pixel points on the same diagonal line in the current block, the diagonal direction includes at least one of the following: a direction from the upper right to the lower left, and a direction from the lower left to the upper right.
[0732] In one example, as shown in FIG24A , the diagonal direction includes a direction from the upper right to the lower left. At this time, as shown by the arrow in FIG24A , the diagonal direction of the current block is from the upper right to the lower left.
[0733] In one example, as shown in FIG24B , the diagonal direction includes a direction from the lower left to the upper right. At this time, as shown by the arrows in FIG24B , the diagonal directions of the current block are all from the lower left to the upper right.
[0734] In one example, as shown in FIG24C , the diagonal direction includes the direction from the upper right to the lower left and the direction from the lower left to the upper right. At this time, as shown by the arrow in FIG24C , the diagonal direction of the current block includes two directions: the direction from the upper right to the lower left and the direction from the lower left to the upper right.
[0735] It should be noted that, since in the embodiment of the present application, the encoding end performs parallel prediction on the pixel points located on the same diagonal line in the current block, the specific direction of the diagonal line does not constitute a limitation on the technical solution of the embodiment of the present application.
[0736] In this embodiment of the present application, the encoder performs prediction on a diagonal pixel of the current block as a unit during each prediction. The encoder also performs the same parallel prediction process for each diagonal pixel of the current block. For ease of description, the k-th diagonal pixel of the current block is used as an example. In this case, the above-mentioned S202-A1 includes the following steps S202-A11 and S202-A12:
[0737] S202-A11. For M pixels on the k-th diagonal line of the current block, use an interpolation filter in parallel to determine predicted values of the M pixels based on the filter coefficients, where k and M are both positive integers.
[0738] S202-A12: Obtain a prediction value of the current block based on the prediction values of the pixels on each diagonal line in the current block.
[0739] The kth diagonal line can be understood as any diagonal line in the current block shown in Figure 23 , which includes M pixels. The encoder uses an interpolation filter based on the filter coefficients to concurrently determine the predicted values for these M pixels. This means the encoder can simultaneously determine the predicted values for the M pixels on the kth diagonal line, significantly increasing prediction speed.
[0740] For example, as shown in Figure 23, assume that the kth diagonal of the current block includes three pixels. The encoder determines the predicted values of these three pixels in parallel. For example, these three pixels are denoted as pixel 1, pixel 2, and pixel 3. At the same time, the encoder uses an interpolation filter based on the filter coefficients to perform interpolation filtering prediction on pixel 1 to obtain the predicted value of pixel 1. Simultaneously, the encoder uses an interpolation filter based on the filter coefficients to perform interpolation filtering prediction on pixel 2 to obtain the predicted value of pixel 2. Simultaneously, the encoder uses an interpolation filter based on the filter coefficients to perform interpolation filtering prediction on pixel 3 to obtain the predicted value of pixel 3. In this way, the encoder determines the predicted values of the three pixels on the kth diagonal of the current block in parallel during a single interpolation filtering prediction process, significantly improving the speed of interpolation filtering prediction. By referring to the method for determining the predicted value of the pixel on the kth diagonal, the encoder can determine the predicted values of the pixels on other diagonals in the current block, thereby obtaining the predicted block for the current block, improving the prediction speed of the current block and enhancing coding efficiency.
[0741] The embodiment of the present application does not limit the specific method in which the encoding end uses an interpolation filter to determine the predicted values of M pixels in parallel based on the filter coefficients.
[0742] In some embodiments, since these M pixel points are points on the k-th diagonal of the current block, these M pixel points can be understood as adjacent pixel points with relatively close features. Therefore, in order to reduce the computational complexity, the input value of the interpolation filter corresponding to one or several pixel points among the M pixel points is determined based on the shape of the interpolation filter. Then, based on the input value of the interpolation filter corresponding to the one or several pixel points, for example, an average value calculation, a weighted calculation, or other calculation method is performed to determine the input value of the interpolation filter corresponding to the other pixel points among the M pixel points except for the one or several pixel points. Finally, based on the filter coefficient and the input value of the interpolation filter corresponding to each pixel point among the M pixel points, the predicted values of the M pixel points are determined in parallel.
[0743] In some embodiments, the above-mentioned step S202-A11 of using an interpolation filter to determine the predicted values of M pixels in parallel based on the filter coefficients includes the following steps:
[0744] S202-A11-a1. Based on the shape of the interpolation filter, determine the pixel values of N positions corresponding to the M pixel points in parallel;
[0745] S202-A11-a2. Based on the filter coefficients and the pixel values at N positions corresponding to the M pixel points, the predicted values of the M pixel points are determined in parallel.
[0746] In this embodiment, when the encoder determines the predicted values of the M pixels on the kth diagonal line of the current block in parallel, it also determines the pixel values of the N positions corresponding to these M pixels in parallel based on the shape of the interpolation filter. The pixel values of the N positions corresponding to each pixel can be understood as the input value of the interpolation filter corresponding to the pixel. Next, the encoder determines the predicted values of the M pixels in parallel based on the filter coefficients and the pixel values of the N positions corresponding to the M pixels.
[0747] For example, as shown in FIG24B , assume that the kth diagonal line of the current block includes three pixels, denoted as pixel 1, pixel 2, and pixel 3. Assume that the interpolation filter for the current block is a 4×4 interpolation filter. When the encoder determines the predicted values for these three pixels in parallel, it determines the pixel values at the 15 positions corresponding to pixel 1 based on the shape of the interpolation filter. These 15 pixel values are used as input to the interpolation filter. Based on the determined filter coefficients, the encoder determines the predicted value for pixel 1. Simultaneously, the encoder determines the pixel values at the 15 positions corresponding to pixel 2 based on the shape of the interpolation filter. These 15 pixel values are used as input to the interpolation filter. Based on the determined filter coefficients, the encoder determines the predicted value for pixel 2. Simultaneously, the encoder determines the pixel values at the 15 positions corresponding to pixel 3 based on the shape of the interpolation filter. These 15 pixel values are used as input to the interpolation filter. Based on the determined filter coefficients, the encoder determines the predicted value for pixel 3. That is, in this embodiment, the encoding end determines the prediction values of three pixels in the current block in parallel at the same time, which greatly improves the prediction speed and thus improves the encoding efficiency.
[0748] The following describes the specific process of determining the predicted values of M pixel points in parallel based on the filter coefficients and the pixel values of N positions corresponding to the M pixel points in the above S202-A11-a2.
[0749] In some embodiments, for each of the M pixels, the encoding end directly multiplies the pixel values of the N positions corresponding to the pixel by the filter coefficient to obtain a predicted value of the pixel.
[0750] For example, the encoding end obtains the predicted value of each of the M pixels based on the above formula (5).
[0751] Based on formula (5), the encoder can determine the predicted values of the pixels on the same diagonal line in the current block in parallel.
[0752] In some embodiments, if the encoder determines the filter coefficients of the interpolation filter based on the above formula (4), the filter coefficients are determined by the reference area after the mean value is removed in the above formula (4). Therefore, when determining the prediction value of the current block based on the filter coefficients, the influence of the pixel average reconstruction value m needs to be considered.
[0753] In a possible implementation of this embodiment, the interpolation filter coefficient determined by the above formula (4) is substituted into the above formula (5) to obtain the predicted value of each point in the current block. Then, the predicted value of each point is added to the pixel average reconstruction value m to obtain the final predicted value of each point in the current block, thereby obtaining the predicted block of the current block.
[0754] In another possible implementation of this embodiment, the above S202-A11-a2 includes the following steps:
[0755] S202-A11-a21, based on the pixel average reconstruction value, performing de-averaging on the pixel values of the N positions corresponding to the M pixel points in parallel to obtain the de-averaged pixel values of the N positions corresponding to the M pixel points;
[0756] S202-A11-a22. Based on the filter coefficients and the pixel values after averaging at N positions corresponding to the M pixel points, the predicted values of the M pixel points are determined in parallel.
[0757] Because the filter coefficients are determined based on the reference area after averaging, the encoder performs averaging on the pixel values at the N positions corresponding to the M pixels on the k-th diagonal of the current block based on the pixel average reconstruction value to obtain the pixel values at the N positions after averaging. For example, for any pixel among the M pixels, the pixel average reconstruction value is subtracted from the pixel values at the N positions of the pixel to obtain the pixel value at the N positions after averaging.
[0758] Next, based on the filter coefficients and the pixel values after averaging at the N positions corresponding to the M pixels, the predicted values of the M pixels are determined in parallel.
[0759] The embodiment of the present application does not limit the specific method of determining the predicted values of M pixel points in parallel based on the filter coefficients and the pixel values after averaging at N positions corresponding to the M pixel points.
[0760] In one implementation, for each of the M pixels, the encoder takes the pixel value and filter coefficient of the N positions of the pixel after removing the mean, and puts them into the above formula (5). At this time, t[ri+p n ] is the position ri+p nThe pixel value after removing the mean value of the pixel at . After determining a predicted value of the r-th point based on the above formula (5), add the pixel average reconstruction value m to the predicted value to obtain the final predicted value of the r-th point.
[0761] In another implementation, the above S202-A11-a22 includes the following steps:
[0762] S202-A11-a221, determine a second reconstruction area around the current block, and determine a maximum reconstruction value and a minimum reconstruction value of the second reconstruction area;
[0763] S202-A11-a222, based on the pixel values after de-averaging at N positions corresponding to the M pixels, the filter coefficients, and the pixel average reconstruction values, obtain first prediction values of the M pixels in parallel;
[0764] S202-A11-a223. Based on the first predicted values, the maximum reconstructed value and the minimum reconstructed value of the M pixel points, the predicted values of the M pixel points are determined in parallel.
[0765] In this implementation, the encoder limits the prediction value of the current block to a range. Specifically, a second reconstruction area is determined, and the maximum reconstruction value (max) and the minimum reconstruction value (min) of the pixels in the second reconstruction area are determined.
[0766] The embodiment of the present application does not limit the specific method of determining the second reconstruction area around the current block.
[0767] In an example, the second reconstructed region of the current block is consistent with the reference region of the current block.
[0768] In one example, the second reconstructed area of the current block is consistent with the first reconstructed area of the current block.
[0769] In one example, the reconstruction areas above, to the left, to the upper right, to the upper left, and to the lower left of the current block are determined as the second reconstruction area. For example, the reconstruction areas 13 rows above, 13 columns to the left, 13 rows to the upper right, 13 rows and 13 columns to the upper left, and 13 columns to the lower left of the current block are determined as the second reconstruction area.
[0770] It should be noted that there is no order of priority between the above S202-A11-a221 and the above S202-A11-a222 in the specific implementation process. For example, the above S202-A11-a221 can be executed before the above S202-A11-a222, or after the above S202-A11-a222, or synchronously with the above S202-A11-a222.
[0771] The embodiment of the present application does not limit the specific method for the encoding end to obtain the first predicted values of M pixel points in parallel based on the pixel values, filter coefficients and pixel average reconstruction values after de-averaging N positions corresponding to M pixel points.
[0772] For example, for any pixel among the M pixels, the pixel value after de-averaging the N positions of the pixel is multiplied by the filter coefficient to obtain the second predicted value of the pixel; the second predicted value and the pixel average reconstruction value are added to obtain the first predicted value of the pixel.
[0773] Exemplarily, the encoding end obtains the first predicted value of the pixel based on the above formula (6).
[0774] For another example, after the encoding end obtains a predicted value of the pixel point based on the above formula (6), it performs preset processing on the predicted value to obtain a first predicted value of the pixel point.
[0775] After determining the first predicted values of the M pixels based on the above steps, the encoder determines the predicted values of the M pixels in parallel based on the first predicted values, the maximum reconstructed value, and the minimum reconstructed value.
[0776] For example, for any pixel among the M pixels, if the first predicted value of the pixel is greater than the minimum reconstruction value and less than the maximum reconstruction value, the first predicted value is determined as the predicted value of the pixel.
[0777] For another example, if the first predicted value of the pixel point is less than or equal to the minimum reconstructed value, the minimum reconstructed value is determined as the predicted value of the pixel point.
[0778] For another example, if the first predicted value of the pixel point is greater than or equal to the maximum reconstructed value, the maximum reconstructed value is determined as the predicted value of the pixel point.
[0779] In one example, the encoding end determines the predicted value of the pixel point using the above formula (7).
[0780] Taking the above method of determining the predicted values of M pixels on the kth diagonal line in the current block as an example, the encoding end can refer to the above method to determine the predicted values of each pixel on the diagonal line in the current block in parallel, and then obtain the predicted value of each point in the current block to form a predicted block of the current block.
[0781] Based on the above steps, the encoder performs interpolation filtering prediction on the current block, obtains the predicted block of the current block, and then performs the following steps.
[0782] S203 : Determine a transformation kernel corresponding to the current block, and encode the current block based on the transformation kernel corresponding to the current block and the prediction block to obtain a code stream.
[0783] As can be seen from the above, when encoding the current block, the encoder determines the prediction block for the current block based on the above steps. Next, the prediction block of the current block is subtracted from the current block to obtain the residual block of the current block. Next, the residual block of the current block is transformed to obtain transform coefficients, which are quantized to obtain quantized coefficients, and these quantized coefficients are encoded to obtain the bitstream.
[0784] When transforming the residual value of the current block to obtain the transform coefficient, it is necessary to determine a transform kernel, and transform the residual value of the current block based on the transform kernel to obtain the transform coefficient.
[0785] The embodiment of the present application does not limit the specific method by which the encoding end determines the transform kernel corresponding to the current block.
[0786] In some embodiments, the encoding end and the decoding end use a default transform kernel as the transform kernel of the current block.
[0787] In some embodiments, the encoder determines the transform kernel of the current block through the following steps S203-A and S203-A:
[0788] S203-A, determining the intra prediction mode corresponding to the prediction block;
[0789] S203-B: Determine a transform kernel corresponding to the current block based on the intra prediction mode corresponding to the prediction block.
[0790] The following describes the specific process by which the encoder determines the intra-frame prediction mode corresponding to the prediction block.
[0791] In one example, as shown in FIG7 , the conventional intra prediction modes currently included in VVC are:
[0792] PLANAR mode: intra prediction mode index is 0,
[0793] DC mode: intra prediction mode index is 1,
[0794] Angle mode: The intra prediction mode index is 2 to 66.
[0795] In one example, as shown in Figure 25 , the arrows in the figure point to the directions predicted by the angle modes in VVC, and the prediction mode indexes used during encoding are 2 to 66. When the current block is a non-square block, some angle directions will be replaced with wide angles, such as -1 to -14 and 67 to 80 in Figure 25 .
[0796] In some embodiments, the intra-frame prediction mode corresponding to the above-mentioned prediction block is a default intra-frame prediction mode. That is, if the current block is predicted using the interpolation filtering prediction mode, when the prediction block is obtained, one of the traditional intra-frame prediction modes is determined as the default intra-frame prediction mode corresponding to the prediction block.
[0797] In some embodiments, the encoder determines the intra prediction mode corresponding to the prediction block through the following steps S203-A1 and S203-A2:
[0798] S203-A1, determining angle values of R points in the prediction block, where R is a positive integer;
[0799] S203-A2: Determine the intra-frame prediction mode corresponding to the prediction block based on the angle values of the R points.
[0800] In the embodiment of the present application, the intra-frame prediction mode corresponding to the prediction block is determined by counting the intra-frame prediction modes corresponding to the angle values of R points in the prediction block.
[0801] The embodiment of the present application does not limit the specific position and number of the R points in the prediction block used to determine the angle value. For example, the R points can be one point in the prediction block, or multiple points in the prediction block.
[0802] For example, if the above R points are one point, the encoding end determines the angle value of a point in the prediction block (for example, the center point of the prediction block), and based on the angle value of the point, determines the intra-frame prediction mode corresponding to the point, and then determines the intra-frame prediction mode as the intra-frame prediction mode corresponding to the prediction block.
[0803] For another example, if the above-mentioned R points are multiple points, the encoding end determines the angle values of these multiple points, and based on the angle values of these multiple points, determines the intra-frame prediction mode corresponding to each of these multiple points, and then determines the intra-frame prediction mode with the largest number of identical intra-frame prediction modes among these multiple points as the intra-frame prediction mode corresponding to the prediction block.
[0804] In some embodiments, when determining the angle values of R points in a prediction block using a sliding window approach, the selection of these R points is related to the shape and size of the sliding window. For example, each of the R points is the center point of the sliding window as it slides across the prediction block.
[0805] In the embodiment of the present application, the method for determining the angle value of each of the R points is the same. For the convenience of description, the method of determining the angle value of the i-th point among the R points is used as an example for explanation.
[0806] The embodiment of the present application does not limit the specific method of determining the angle value of the point.
[0807] In some embodiments, the above S203-A1 includes steps S202-A11 and S203-A12:
[0808] S203-A11. For the i-th point among the R points, determine the horizontal gradient and vertical gradient of the i-th point, where i is a positive integer less than or equal to R;
[0809] S203-A12: Determine the angle value of the i-th point based on the horizontal gradient and the vertical gradient of the i-th point.
[0810] In this embodiment, the encoding end first determines the horizontal gradient and vertical gradient of each of the R points, such as the i-th point, and then determines the angle value of the i-th point based on the horizontal gradient and the vertical gradient.
[0811] The embodiment of the present application does not limit the specific method of determining the horizontal gradient and vertical gradient of the i-th point.
[0812] In one example, the horizontal gradient value of the i-th point is determined based on the predicted values of the points around the i-th point in the prediction block and the change in the predicted value of the i-th point in the horizontal direction, and the vertical gradient value of the i-th point is determined based on the predicted values of the points around the i-th point in the prediction block and the change in the predicted value of the i-th point in the vertical direction.
[0813] In another example, the encoding end determines the prediction value of the point in the sliding window centered on the i-th point in the prediction block; based on the prediction value of the point in the sliding window and the horizontal gradient operator, as well as the vertical gradient operator, the horizontal gradient and vertical gradient of the i-th point are obtained.
[0814] In this example, a sliding window is first determined. For example, as shown in Figure 26, a 3×3 sliding window is determined. This sliding window is then slid across the prediction block. With each slide, the horizontal and vertical gradients at the center point of the sliding window are determined. For example, with the center point of the current sliding window as the i-th point, the predicted values for each point within the current sliding window are first obtained. For example, predicted values for 3×3 = 9 points can be obtained. Next, based on the predicted values for these 9 points and the preset horizontal and vertical gradient operators, the horizontal and vertical gradients at the i-th point are determined.
[0815] For example, the product of the predicted value of the point in the sliding window and the horizontal gradient operator is determined as the horizontal gradient G of the i-th point x ; The product of the predicted value of the point in the sliding window and the vertical gradient operator is determined as the vertical gradient of the i-th point.
[0816] For another example, the predicted value of the point in the sliding window is multiplied by the horizontal gradient operator and then the preset operation is performed with the preset value to obtain the horizontal gradient G of the i-th point.x ; Multiply the predicted value of the point in the sliding window by the vertical gradient operator and then perform a preset operation with the preset value to obtain the vertical gradient of the i-th point.
[0817] The embodiment of the present application does not limit the specific values of the horizontal gradient operator and the vertical gradient operator.
[0818] After determining the horizontal gradient and vertical gradient of the i-th point based on the above steps, the encoder can determine the angle value of the i-th point based on the horizontal gradient and vertical gradient of the i-th point.
[0819] For example, the inverse tangent value of the ratio of the vertical gradient to the horizontal gradient at the i-th point is determined as the angle value of the i-th point. For example, the angle value of the i-th point is determined according to formula (8).
[0820] In addition to using the above formula (8) to determine the angle value of the i-th point, the encoder can also use other methods to determine the angle value of the i-th point. For example, the encoder adjusts the angle value determined by the above formula (8) to obtain the angle value of the i-th point.
[0821] The encoding end uses the above method for each of the R points to determine the angle value of each of the R points, and then executes the above S203-A2 to determine the intra-frame prediction mode corresponding to the prediction block based on the angle values of the R points.
[0822] The embodiment of the present application does not limit the specific method of determining the intra-frame prediction mode corresponding to the prediction block based on the angle values of R points.
[0823] In some embodiments, the encoding end selects the angle value 1 that is the same the most times from the angle values of R points, matches the angle value 1 with the prediction angle of the traditional intra-frame prediction mode, obtains the intra-frame prediction mode corresponding to the angle value 1, and determines the intra-frame prediction mode corresponding to the angle value 1 as the intra-frame prediction mode corresponding to the prediction block.
[0824] In some embodiments, the above S203-A2 includes the following steps S203-A21 and S203-A22:
[0825] S203-A21, determining the intra-frame prediction modes corresponding to the R points based on the angle values of the R points;
[0826] S203-A22: Determine the intra-frame prediction mode corresponding to the prediction block based on the intra-frame prediction modes corresponding to the R points.
[0827] In this implementation, the encoder determines the intra-frame prediction mode corresponding to each of the R points based on the angle value of each point. For example, for each of the R points, the angle value of the point is matched with the prediction angle of the traditional intra-frame prediction mode to obtain the intra-frame prediction mode corresponding to the angle value of the point. In this way, the intra-frame prediction mode corresponding to each of the R points can be obtained.
[0828] Next, based on the intra-frame prediction mode corresponding to each of the R points, the intra-frame prediction mode corresponding to the prediction block is determined.
[0829] In a possible implementation, the intra-frame prediction mode with the greatest number of repetitions among the intra-frame prediction modes corresponding to the R points is determined as the intra-frame prediction mode corresponding to the prediction block.
[0830] In another possible implementation, the above S203-A22 includes the following steps:
[0831] S203-A221, based on the horizontal gradients and vertical gradients of the R points, determine the gradient amplitude values corresponding to the R points;
[0832] S203-A222: Determine the intra-frame prediction mode corresponding to the prediction block based on the intra-frame prediction modes and gradient magnitude values corresponding to the R points.
[0833] In this implementation, the encoding end determines the gradient amplitude value corresponding to each of the R points based on the horizontal gradient and vertical gradient of each of the R points determined above.
[0834] In the embodiment of the present application, the specific manner in which the encoder determines the gradient amplitude value corresponding to each of the R points is the same. For ease of description, the example of determining the gradient amplitude value corresponding to the i-th point among the R points is taken as an example.
[0835] The embodiment of the present application does not limit the specific manner in which the encoding end determines the gradient amplitude value corresponding to the i-th point based on the horizontal gradient and vertical gradient of the i-th point.
[0836] For example, the encoder multiplies the horizontal gradient and the vertical gradient of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.
[0837] For another example, the encoding end adds the absolute value of the horizontal gradient and the absolute value of the vertical gradient of the i-th point to obtain the gradient amplitude value corresponding to the i-th point.
[0838] Exemplarily, the encoder determines the gradient amplitude value corresponding to the i-th point based on the following formula (9).
[0839] The encoder can determine the gradient magnitude value corresponding to each of the R points based on the above steps. Next, the encoder performs S203-A222 above to determine the intra prediction mode corresponding to the prediction block based on the intra prediction modes and gradient magnitude values corresponding to the R points.
[0840] In one example, the intra-frame prediction mode corresponding to the point with the largest gradient magnitude value among the R points is determined as the intra-frame prediction mode corresponding to the prediction block.
[0841] In another example, for any point among the R points, the gradient amplitude value corresponding to the point is accumulated on the intra-frame prediction mode corresponding to the point to obtain the accumulated gradient amplitude values of the intra-frame prediction modes corresponding to the R points; and the intra-frame prediction mode with the largest accumulated gradient amplitude value among the intra-frame prediction modes corresponding to the R points is determined as the intra-frame prediction mode corresponding to the prediction block.
[0842] For example, as shown in FIG27, the gradient amplitude value corresponding to each of the R points is accumulated on the corresponding intra-frame prediction mode. For example, the intra-frame prediction modes corresponding to point 1 and point 2 of the R points are both intra-frame prediction mode 1, so the gradient amplitude values corresponding to point 1 and point 2 are accumulated to the gradient amplitude value corresponding to intra-frame prediction mode 1. Similarly, the gradient amplitude value histogram shown in FIG27 can be obtained. In this way, the intra-frame prediction mode with the largest cumulative gradient amplitude value in the gradient amplitude value histogram can be determined as the intra-frame prediction mode corresponding to the prediction block. For example, the intra-frame prediction mode corresponding to the dark cumulative gradient amplitude value in FIG27 is determined as the intra-frame prediction mode corresponding to the prediction block.
[0843] In some embodiments, if the gradient magnitude values corresponding to the R points are all 0, the first intra-frame prediction mode is determined as the intra-frame prediction mode corresponding to the prediction block. In other words, if the gradient magnitude values corresponding to all of the R points are 0, it means that the horizontal gradient and vertical gradient of each of the R points are both 0. In this case, the preset first intra-frame prediction mode can be determined as the intra-frame prediction mode corresponding to the prediction block.
[0844] The embodiment of the present application does not limit the type of the first intra-frame prediction mode.
[0845] Exemplarily, the first intra-frame prediction mode is the PLANAR mode.
[0846] After the encoder determines the intra-frame prediction mode corresponding to the prediction block based on the above steps, it determines the transform kernel corresponding to the current block based on the intra-frame prediction mode corresponding to the prediction block.
[0847] The embodiment of the present application does not limit the specific manner in which the encoder determines the transform kernel corresponding to the current block based on the intra-frame prediction mode corresponding to the prediction block.
[0848] In some embodiments, the encoding end searches for an image block whose intra-frame prediction mode is the same as the intra-frame prediction mode corresponding to the prediction block in the encoded image blocks around the prediction block based on the intra-frame prediction mode corresponding to the prediction block, and then determines the transform kernel corresponding to the image block as the transform kernel corresponding to the current block.
[0849] In some embodiments, determining the transform kernel corresponding to the current block based on the intra prediction mode corresponding to the prediction block in S203-B above includes the following steps:
[0850] S203-B1, obtaining a correspondence between an intra prediction mode and a transform core group, wherein one transform core group includes at least one type of transform core;
[0851] S203-B2, searching the corresponding relationship for the first transform kernel group corresponding to the intra prediction mode of the prediction block;
[0852] S203-B3: Determine a transform core corresponding to the current block from the first transform core group.
[0853] In the embodiment of the present application, there is a correspondence between the intra prediction mode and the transform core group. Based on this, after determining the intra prediction mode corresponding to the prediction block, the encoder obtains the preset correspondence between the intra prediction mode and the transform core group.
[0854] In an example, the correspondence between intra prediction modes and transform kernel groups is shown in Table 17.
[0855] It should be noted that the above Table 17 is only a correspondence between an intra-frame prediction mode and a transform core group involved in an embodiment of the present application. The correspondence between the intra-frame prediction mode and the transform core group in the embodiment of the present application includes but is not limited to that shown in Table 16.
[0856] Each transformation core group includes at least one type of transformation core.
[0857] After obtaining the correspondence between intra-frame prediction modes and transform kernel groups shown in Table 17, the encoder searches the correspondence between intra-frame prediction modes and transform kernel groups based on the intra-frame prediction mode corresponding to the prediction block, and records this transform kernel group as the first transform kernel group. For example, if the intra-frame prediction mode corresponding to the prediction block is an angular prediction mode in the 64-angle direction, searching Table 16 above shows that the transform kernel group corresponding to this angular prediction mode in the 64-angle direction is 4. In this way, the encoder determines the transform kernel corresponding to the current block from the at least one type of transform kernel included in transform kernel group 4.
[0858] For example, if the first transform core group includes one transform core, the transform core is determined as the transform core corresponding to the current block.
[0859] For another example, if the first transform core group includes transform cores of multiple categories, the encoder determines the transform core category corresponding to the current block, and then determines the transform core of the transform core category in the first transform core group as the transform core corresponding to the current block.
[0860] The methods for the encoder to determine the transform kernel type corresponding to the current block include but are not limited to the following:
[0861] In one example, the transform kernel category corresponding to the current block is a default category, so the encoder determines the default category as the transform kernel category corresponding to the current block.
[0862] In another example, the encoder writes the transform kernel type corresponding to the current block into the bitstream, so that the encoder obtains the transform kernel type corresponding to the current block by encoding the bitstream.
[0863] As can be seen from the above, in an embodiment of the present application, the encoding end uses an interpolation filter prediction mode to determine the prediction block of the current block, and then determines the traditional intra-frame prediction mode corresponding to the prediction block, and based on the traditional intra-frame prediction mode corresponding to the prediction block, determines the transform kernel corresponding to the current block. That is to say, the embodiment of the present application is based on the traditional intra-frame prediction mode derived from the interpolation filter prediction, and is used for the selection of transform kernel groups of the non-separable primary transform (NSPT) and the non-separable secondary transform (LFNST), so that the determined transform kernel is more consistent with the characteristics of the current block, and the accuracy of determining the transform kernel is improved. When the accurately determined transform kernel is used to determine the reconstruction value of the current block, the accuracy of determining the reconstruction value can be improved, and the encoding accuracy of the current block can be improved. In addition, when the embodiment of the present application determines the transform kernel of the current block through the traditional prediction mode corresponding to the prediction block, there is no need to indicate the transform kernel separately, which saves codewords and further improves the video encoding effect.
[0864] In an embodiment of the present application, the encoder determines the prediction block of the current block and the transform kernel corresponding to the current block based on the above steps. In this way, the encoder can obtain the residual block of the current block based on the prediction block of the current block and the current block, for example, by subtracting the current block from the prediction block of the current block to obtain the residual block of the current block. Next, the residual block of the current block is transformed based on the above-determined transform kernel to obtain the transform coefficients of the current block. Next, the transform coefficients are directly encoded to obtain a bitstream. Alternatively, the transform coefficients are quantized to obtain quantized coefficients, and the quantized coefficients are encoded to obtain a bitstream.
[0865] In some embodiments, the above-mentioned current block is a bright color block or a chroma block, that is, in the embodiment of the present application, the interpolation filtering prediction mode provided by the embodiment of the present application can be used to predict both the luminance block and the chroma block.
[0866] In some embodiments, if the current block is a luminance block, the prediction mode of the current block is an interpolation filtering prediction mode, and the chrominance block corresponding to the current block adopts a direct derivation mode DM, then the PLANAR mode or the intra-frame prediction mode corresponding to the above prediction block is determined as the prediction mode of the chrominance block.
[0867] The video encoding method provided in the embodiments of the present application, when predicting a current block, first determines a reference area and an interpolation filter for the current block, and based on the reference area, determines a filter coefficient. Based on the filter coefficient, the interpolation filter is used to perform parallel prediction on at least two pixels in the current block to obtain a prediction block for the current block. A transform kernel corresponding to the current block is determined, and based on the transform kernel and the prediction block, the current block is encoded to obtain a bitstream. In other words, in the embodiments of the present application, when using an interpolation filter to perform interpolation filtering prediction on the current block, parallel prediction is performed on at least two points in the current block, thereby increasing prediction speed and, in turn, improving encoding efficiency.
[0868] It should be understood that Figures 10 to 29 are merely examples of the present application and should not be understood as limiting the present application.
[0869] The preferred embodiments of the present application are described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the technical concept of the present application, a variety of simple modifications can be made to the technical solution of the present application, and these simple modifications all fall within the scope of protection of the present application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. In order to avoid unnecessary repetition, the present application will not further explain various possible combinations. For another example, the various different embodiments of the present application can also be arbitrarily combined, and as long as they do not violate the ideas of the present application, they should also be regarded as the contents disclosed in the present application.
[0870] It should also be understood that in the various method embodiments of the present application, the size of the sequence numbers of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, in the embodiments of the present application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. Specifically, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this application generally indicates that the related objects before and after are in an "or" relationship.
[0871] The above describes in detail the method embodiment of the present application in conjunction with Figures 10 to 29, and the following describes in detail the device embodiment of the present application in conjunction with Figures 30 to 31.
[0872] FIG30 is a schematic block diagram of a video decoding device according to an embodiment of the present application. The video decoding device 10 is applied to the above-mentioned video decoder.
[0873] As shown in FIG30 , the video decoding apparatus 10 includes:
[0874] a coefficient determination unit 11, configured to determine a reference area and an interpolation filter of a current block, and determine a filter coefficient of the interpolation filter based on the reference area;
[0875] A prediction unit 12 is configured to perform parallel prediction on at least two pixels in the current block using the interpolation filter based on the filter coefficients to determine a prediction block of the current block;
[0876] The reconstruction unit 13 is configured to determine a transformation kernel corresponding to the current block, and determine a reconstructed block of the current block based on the transformation kernel corresponding to the current block and the prediction block.
[0877] In some embodiments, the prediction unit 12 is specifically configured to perform parallel interpolation filtering prediction on pixel points on the same diagonal line of the current block using the interpolation filter along the diagonal direction based on the filter coefficient to obtain a prediction block of the current block.
[0878] In some embodiments, the prediction unit 12 is specifically used to perform parallel interpolation filtering prediction on pixel points on the same diagonal line of the current block using the interpolation filter based on the filter coefficient, starting from the upper left corner of the current block and along the diagonal direction, to obtain a predicted block of the current block.
[0879] In some embodiments, the diagonal direction includes at least one of the following: a direction from the upper right to the lower left, and a direction from the lower left to the upper right.
[0880] In some embodiments, the prediction unit 12 is specifically used to determine the predicted values of the M pixel points on the kth diagonal line of the current block in parallel using the interpolation filter based on the filter coefficient, where k and M are both positive integers; and obtain the predicted value of the current block based on the predicted values of the pixel points on each diagonal line in the current block.
[0881] In some embodiments, the prediction unit 12 is specifically used to determine the pixel values of the N positions corresponding to the M pixel points in parallel based on the shape of the interpolation filter; and determine the predicted values of the M pixel points in parallel based on the filter coefficients and the pixel values of the N positions corresponding to the M pixel points.
[0882] In some embodiments, the coefficient determination unit 11 is specifically used to determine a first reconstruction area around the current block; determine a pixel average reconstruction value based on the reconstruction value of the first reconstruction area; de-average the reconstruction values of the pixel points in the reference area based on the pixel average reconstruction value; use the de-averaged pixel values of the pixel points in the reference area as the input of the interpolation filter, slide the interpolation filter within the reference area, and obtain the filter coefficients of the interpolation filter.
[0883] In some embodiments, the coefficient determination unit 11 is specifically configured to determine the pixel average reconstruction value based on the shape of the current block and the reconstruction value of the first reconstruction area.
[0884] In some embodiments, the coefficient determination unit 11 is specifically used to determine a first area from the upper reconstruction area and the left reconstruction area based on the shape of the current block; determine an average reconstruction value of the first area based on the reconstruction value of the first area; and determine the pixel average reconstruction value based on the average reconstruction value of the first area.
[0885] In some embodiments, the coefficient determination unit 11 is specifically used to determine the upper reconstruction area as the first area if the shape of the current block is wider than the height; or, if the shape of the current block is higher than the width, determine the left reconstruction area as the first area; or, if the shape of the current block is higher than the width, determine the upper reconstruction area and the left reconstruction area as the first area.
[0886] In some embodiments, the horizontal sliding step size and the vertical sliding step size of the interpolation filter in the reference area are different.
[0887] In some embodiments, at least one of a horizontal sliding step size and a vertical sliding step size of the interpolation filter within the reference area is greater than a preset step size.
[0888] In some embodiments, the prediction unit 12 is specifically used to perform de-averaging on the pixel values of the N positions corresponding to the M pixel points in parallel based on the pixel average reconstruction value to obtain the de-averaged pixel values of the N positions corresponding to the M pixel points; and determine the predicted values of the M pixel points in parallel based on the filter coefficients and the de-averaged pixel values of the N positions corresponding to the M pixel points.
[0889] In some embodiments, the prediction unit 12 is specifically configured to, for any pixel point among the M pixel points, subtract the pixel average reconstruction value from the pixel values at the N positions of the pixel point to obtain the pixel values at the N positions of the pixel point after removing the mean.
[0890] In some embodiments, the prediction unit 12 is specifically used to determine a second reconstruction area around the current block, and determine the maximum reconstruction value and the minimum reconstruction value of the second reconstruction area; based on the pixel values after de-averaging of the N positions corresponding to the M pixel points, the filter coefficient and the pixel average reconstruction value, obtain the first prediction values of the M pixel points in parallel; based on the first prediction values of the M pixel points, the maximum reconstruction value and the minimum reconstruction value, determine the prediction values of the M pixel points in parallel.
[0891] In some embodiments, the prediction unit 12 is specifically used to multiply the pixel value after d...
Claims
1. A video decoding method, characterized in that: include: Determining a reference area and an interpolation filter of a current block, and determining a filter coefficient of the interpolation filter based on the reference area; Based on the filter coefficients, use the interpolation filter to perform parallel prediction on at least two pixels in the current block to determine a prediction block of the current block; A transform kernel corresponding to the current block is determined, and a reconstructed block of the current block is determined based on the transform kernel corresponding to the current block and the prediction block.
2. The method according to claim 1, characterized in that The step of performing parallel interpolation filtering prediction on at least two pixels in the current block using the interpolation filter based on the filter coefficient to obtain a prediction block of the current block includes: Based on the filter coefficients, along the diagonal direction, the interpolation filter is used to perform parallel interpolation filtering prediction on the pixel points on the same diagonal line of the current block to obtain a prediction block of the current block.
3. The method according to claim 2, characterized in that The method of performing parallel interpolation filtering prediction on pixel points on the same diagonal line of the current block using the interpolation filter along the diagonal direction based on the filter coefficient to obtain a prediction block of the current block includes: Based on the filter coefficient, starting from the upper left corner of the current block and along the diagonal direction, the interpolation filter is used to perform parallel interpolation filtering prediction on the pixel points on the same diagonal line of the current block to obtain a prediction block of the current block.
4. The method according to claim 3, characterized in that The diagonal direction includes at least one of the following: a direction from the upper right to the lower left, and a direction from the lower left to the upper right.
5. The method according to claim 3, characterized in that: The method of performing parallel interpolation filtering prediction on pixel points on the same diagonal line of the current block using the interpolation filter based on the filter coefficient, starting from the upper left corner of the current block and along the diagonal direction, to obtain a prediction block of the current block includes: For M pixel points on the kth diagonal line of the current block, based on the filter coefficients, using the interpolation filter to determine the predicted values of the M pixel points in parallel, where k and M are both positive integers; The predicted value of the current block is obtained based on the predicted values of the pixels on each diagonal line in the current block.
6. The method according to claim 5, characterized in that The method of determining the predicted values of the M pixels in parallel using the interpolation filter based on the filter coefficients includes: Based on the shape of the interpolation filter, determining in parallel the pixel values of the N positions respectively corresponding to the M pixel points; Based on the filter coefficients and the pixel values of the N positions respectively corresponding to the M pixel points, the predicted values of the M pixel points are determined in parallel.
7. The method according to claim 6, characterized in that The determining, based on the reference area, a filter coefficient of the interpolation filter comprises: Determining a first reconstruction area around the current block; Determining a pixel average reconstruction value based on the reconstruction value of the first reconstruction area; Based on the pixel average reconstruction value, de-averaging the reconstruction values of the pixel points in the reference area; The pixel values of the pixels in the reference area after averaging are used as the input of the interpolation filter, and the interpolation filter is slid in the reference area to obtain the filter coefficients of the interpolation filter.
8. The method according to claim 7, characterized in that The determining the pixel average reconstruction value based on the reconstruction value of the first reconstruction area includes: The pixel average reconstruction value is determined based on the shape of the current block and the reconstruction value of the first reconstruction area.
9. The method according to claim 8, characterized in that The first reconstruction area includes an upper reconstruction area and a left reconstruction area of the current block, and determining a pixel average reconstruction value based on the shape of the current block and the reconstruction value of the first reconstruction area includes: Based on the shape of the current block, determine a first area from the upper reconstruction area and the left reconstruction area; determining an average reconstruction value of the first region based on the reconstruction value of the first region; The pixel average reconstruction value is determined based on the average reconstruction value of the first area.
10. The method according to claim 9, characterized in that The determining, based on the shape of the current block, a first area from the upper reconstruction area and the left reconstruction area comprises: If the shape of the current block is that the width is greater than the height, the upper reconstruction area is determined as the first area; or, If the shape of the current block is that the height is greater than the width, the left reconstruction area is determined as the first area; or, If the shape of the current block is that the height is equal to the width, the upper reconstruction area and the left reconstruction area are determined as the first area.
11. The method according to claim 7, characterized in that The horizontal sliding step length and the vertical sliding step length of the interpolation filter in the reference area are different.
12. The method according to claim 7, characterized in that At least one of a horizontal sliding step size and a vertical sliding step size of the interpolation filter in the reference area is greater than a preset step size.
13. The method according to claim 7, characterized in that The step of determining the predicted values of the M pixels in parallel based on the filter coefficients and the pixel values of the N positions respectively corresponding to the M pixels comprises: Based on the pixel average reconstruction value, the pixel values of the N positions corresponding to the M pixel points are de-averaged in parallel to obtain the pixel values of the N positions corresponding to the M pixel points after de-averaging; Based on the filter coefficients and the pixel values after averaging at N positions respectively corresponding to the M pixel points, the predicted values of the M pixel points are determined in parallel.
14. The method according to claim 13, characterized in that The method of performing de-averaging on the pixel values of the N positions respectively corresponding to the M pixel points in parallel based on the pixel average reconstruction value to obtain the de-averaged pixel values of the N positions respectively corresponding to the M pixel points includes: For any pixel point among the M pixel points, the pixel average reconstruction value is subtracted from the pixel values at the N positions of the pixel point to obtain the pixel values of the N positions of the pixel point after removing the mean.
15. The method according to claim 13, characterized in that The step of determining the predicted values of the M pixels in parallel based on the filter coefficients and the pixel values of the N positions respectively corresponding to the M pixels after averaging, comprises: Determine a second reconstruction area around the current block, and determine a maximum reconstruction value and a minimum reconstruction value of the second reconstruction area; Based on the pixel values after averaging at the N positions respectively corresponding to the M pixel points, the filter coefficient and the pixel average reconstruction value, obtaining first prediction values of the M pixel points in parallel; Based on the first predicted values of the M pixels, the maximum reconstruction value and the minimum reconstruction value, the predicted values of the M pixels are determined in parallel.
16. The method according to claim 15, characterized in that The method of obtaining the first predicted values of the M pixels in parallel based on the pixel values after averaging the N positions respectively corresponding to the M pixels, the filter coefficients and the pixel average reconstruction values comprises: For any pixel point among the M pixels, multiply the pixel value after removing the mean value at N positions of the pixel point by the filter coefficient to obtain a second predicted value of the pixel point; The second predicted value and the pixel average reconstruction value are added to obtain a first predicted value of the pixel point.
17. The method according to claim 15, characterized in that The step of determining the predicted values of the M pixels in parallel based on the first predicted values of the M pixels, the maximum reconstructed value, and the minimum reconstructed value includes: For any pixel point among the M pixels, if the first predicted value of the pixel point is greater than the minimum reconstruction value and less than the maximum reconstruction value, the first predicted value is determined as the predicted value of the pixel point; or, If the first predicted value of the pixel point is less than or equal to the minimum reconstruction value, the minimum reconstruction value is determined as the predicted value of the pixel point; or, If the first predicted value of the pixel point is greater than or equal to the maximum reconstruction value, the maximum reconstruction value is determined as the predicted value of the pixel point.
18. The method according to claim 1, characterized in that Before determining the reference area and the interpolation filter of the current block, the method further includes: Determining whether the current block allows the use of an interpolation filtering prediction mode; The determining of the reference area and the interpolation filter of the current block comprises: If the current block allows the interpolation filter prediction mode, a reference area and an interpolation filter of the current block are determined.
19. The method according to claim 18, characterized in that The determining whether the current block is allowed to use the interpolation filtering prediction mode comprises: If the current block is in the first row of the current CTU, it is determined that the current block is not allowed to use the interpolation filtering prediction mode.
20. The method according to claim 18, characterized in that The determining whether the current block is allowed to use the interpolation filtering prediction mode comprises: Based on the type of the current image, it is determined whether the current block is allowed to use the interpolation filtering prediction mode.
21. The method according to claim 20, characterized in that The determining, based on the type of the current image, whether the current block is allowed to use the interpolation filtering prediction mode comprises: If the current image is not an intra-prediction image, it is determined that the current block is not allowed to use the interpolation filtering prediction mode.
22. The method according to claim 18, characterized in that The determining whether the current block is allowed to use the interpolation filtering prediction mode comprises: If the size of the current block is smaller than a preset size, it is determined that the current block is not allowed to use the interpolation filtering prediction mode.
23. The method according to claim 18, characterized in that The determining whether the current block is allowed to use the interpolation filtering prediction mode comprises: Decoding the bitstream to obtain first information, where the first information is used to indicate whether the template matching-based technology is enabled; Based on the first information, it is determined whether the current block is allowed to use the interpolation filtering prediction mode.
24. The method according to claim 23, characterized in that The determining, based on the first information, whether the current block is allowed to use the interpolation filtering prediction mode comprises: If the first information indicates that the template matching-based technology is not enabled, it is determined that the current block is not allowed to use the interpolation filtering prediction mode.
25. The method according to claim 23, characterized in that The determining, based on the first information, whether the current block is allowed to use the interpolation filtering prediction mode comprises: If the first information indicates that the template matching-based technology is turned on, the bitstream is decoded to obtain second information, where the second information is used to indicate whether the current sequence is allowed to be predicted using the interpolation filtering prediction mode; Based on the second information, it is determined whether the current block is allowed to use the interpolation filtering prediction mode.
26. The method according to claim 1, characterized in that The determining the transformation kernel corresponding to the current block includes: Determining an intra prediction mode corresponding to the prediction block; Based on the intra prediction mode corresponding to the prediction block, a transform kernel corresponding to the current block is determined.
27. The method according to claim 26, characterized in that The determining the intra prediction mode corresponding to the prediction block includes: Determine the angle values of R points in the prediction block, where R is a positive integer; Based on the angle values of the R points, an intra-frame prediction mode corresponding to the prediction block is determined.
28. The method according to claim 1, characterized in that The determining, based on the intra prediction mode corresponding to the prediction block, a transform kernel corresponding to the current block comprises: Acquire a correspondence between an intra-frame prediction mode and a transform core group, wherein one transform core group includes at least one type of transform core; In the corresponding relationship, searching for a first transform core group corresponding to the intra prediction mode of the prediction block; A transform core corresponding to the current block is determined from the first transform core group.
29. The method according to claim 1, characterized in that The determining of the reference area of the current block includes: A reference area of the current block is determined in preset P reference areas, where P is a positive integer greater than 1.
30. The method according to claim 1, characterized in that The determining of the interpolation filter of the current block comprises: An interpolation filter for the current block is determined among preset Q interpolation filters, where Q is a positive integer greater than 1.
31. A video encoding method, characterized in that: include: Determining a reference area and an interpolation filter of a current block, and determining a filter coefficient of the interpolation filter based on the reference area; Based on the filter coefficients, use the interpolation filter to perform parallel prediction on at least two pixels in the current block to determine a prediction block of the current block; A transform kernel corresponding to the current block is determined, and based on the transform kernel corresponding to the current block and the prediction block, the current block is encoded to obtain a code stream.
32. The method according to claim 31, characterized in that The step of performing parallel interpolation filtering prediction on at least two pixels in the current block using the interpolation filter based on the filter coefficient to obtain a prediction block of the current block includes: Based on the filter coefficients, along the diagonal direction, the interpolation filter is used to perform parallel interpolation filtering prediction on the pixel points on the same diagonal line of the current block to obtain a prediction block of the current block.
33. The method according to claim 32, characterized in that The method of performing parallel interpolation filtering prediction on pixel points on the same diagonal line of the current block using the interpolation filter along the diagonal direction based on the filter coefficient to obtain a prediction block of the current block includes: Based on the filter coefficient, starting from the upper left corner of the current block and along the diagonal direction, the interpolation filter is used to perform parallel interpolation filtering prediction on the pixel points on the same diagonal line of the current block to obtain a prediction block of the current block.
34. The method according to claim 33, characterized in that The diagonal direction includes at least one of the following: a direction from the upper right to the lower left, and a direction from the lower left to the upper right.
35. The method according to claim 33, characterized in that The method of performing parallel interpolation filtering prediction on pixel points on the same diagonal line of the current block using the interpolation filter based on the filter coefficient, starting from the upper left corner of the current block and along the diagonal direction, to obtain a prediction block of the current block includes: For M pixel points on the kth diagonal line of the current block, based on the filter coefficients, using the interpolation filter to determine the predicted values of the M pixel points in parallel, where k and M are both positive integers; The predicted value of the current block is obtained based on the predicted values of the pixels on each diagonal line in the current block.
36. The method according to claim 35, characterized in that The method of determining the predicted values of the M pixels in parallel using the interpolation filter based on the filter coefficients includes: Based on the shape of the interpolation filter, determining in parallel the pixel values of the N positions respectively corresponding to the M pixel points; Based on the filter coefficients and the pixel values of the N positions respectively corresponding to the M pixel points, the predicted values of the M pixel points are determined in parallel.
37. The method according to claim 36, characterized in that The determining, based on the reference area, a filter coefficient of the interpolation filter comprises: Determining a first reconstruction area around the current block; Determining a pixel average reconstruction value based on the reconstruction value of the first reconstruction area; Based on the pixel average reconstruction value, de-averaging the reconstruction values of the pixel points in the reference area; The pixel values of the pixels in the reference area after averaging are used as the input of the interpolation filter, and the interpolation filter is slid in the reference area to obtain the filter coefficients of the interpolation filter.
38. The method according to claim 37, characterized in that The determining the pixel average reconstruction value based on the reconstruction value of the first reconstruction area includes: The pixel average reconstruction value is determined based on the shape of the current block and the reconstruction value of the first reconstruction area.
39. The method according to claim 38, characterized in that The first reconstruction area includes an upper reconstruction area and a left reconstruction area of the current block, and determining a pixel average reconstruction value based on the shape of the current block and the reconstruction value of the first reconstruction area includes: Based on the shape of the current block, determine a first area from the upper reconstruction area and the left reconstruction area; determining an average reconstruction value of the first region based on the reconstruction value of the first region; The pixel average reconstruction value is determined based on the average reconstruction value of the first area.
40. The method according to claim 39, characterized in that The determining, based on the shape of the current block, a first area from the upper reconstruction area and the left reconstruction area comprises: If the shape of the current block is that the width is greater than the height, the upper reconstruction area is determined as the first area; or, If the shape of the current block is that the height is greater than the width, the left reconstruction area is determined as the first area; or, If the shape of the current block is that the height is equal to the width, the upper reconstruction area and the left reconstruction area are determined as the first area.
41. The method according to claim 37, characterized in that The horizontal sliding step size and the vertical sliding step size of the interpolation filter in the reference area are different.
42. The method according to claim 37, characterized in that At least one of a horizontal sliding step size and a vertical sliding step size of the interpolation filter in the reference area is greater than a preset step size.
43. The method according to claim 37, characterized in that The step of determining the predicted values of the M pixels in parallel based on the filter coefficients and the pixel values of the N positions respectively corresponding to the M pixels comprises: Based on the pixel average reconstruction value, the pixel values of the N positions corresponding to the M pixel points are de-averaged in parallel to obtain the pixel values of the N positions corresponding to the M pixel points after de-averaging; Based on the filter coefficients and the pixel values after averaging at N positions respectively corresponding to the M pixel points, the predicted values of the M pixel points are determined in parallel.
44. The method according to claim 43, characterized in that The method of performing de-averaging on the pixel values of the N positions respectively corresponding to the M pixel points in parallel based on the pixel average reconstruction value to obtain the de-averaged pixel values of the N positions respectively corresponding to the M pixel points includes: For any pixel point among the M pixel points, the pixel average reconstruction value is subtracted from the pixel values at the N positions of the pixel point to obtain the pixel values of the N positions of the pixel point after removing the mean.
45. The method according to claim 43, characterized in that The step of determining the predicted values of the M pixels in parallel based on the filter coefficients and the pixel values of the N positions respectively corresponding to the M pixels after averaging, comprises: Determine a second reconstruction area around the current block, and determine a maximum reconstruction value and a minimum reconstruction value of the second reconstruction area; Based on the pixel values after averaging at the N positions respectively corresponding to the M pixel points, the filter coefficient and the pixel average reconstruction value, obtaining first prediction values of the M pixel points in parallel; Based on the first predicted values of the M pixels, the maximum reconstruction value and the minimum reconstruction value, the predicted values of the M pixels are determined in parallel.
46. The method according to claim 45, characterized in that The method of obtaining the first predicted values of the M pixels in parallel based on the pixel values after averaging the N positions respectively corresponding to the M pixels, the filter coefficients and the pixel average reconstruction values comprises: For any pixel point among the M pixels, multiply the pixel value after removing the mean value at N positions of the pixel point by the filter coefficient to obtain a second predicted value of the pixel point; The second predicted value and the pixel average reconstruction value are added to obtain a first predicted value of the pixel point.
47. The method according to claim 45, characterized in that The step of determining the predicted values of the M pixels in parallel based on the first predicted values of the M pixels, the maximum reconstructed value, and the minimum reconstructed value includes: For any pixel point among the M pixels, if the first predicted value of the pixel point is greater than the minimum reconstruction value and less than the maximum reconstruction value, the first predicted value is determined as the predicted value of the pixel point; or, If the first predicted value of the pixel point is less than or equal to the minimum reconstruction value, the minimum reconstruction value is determined as the predicted value of the pixel point; or, If the first predicted value of the pixel point is greater than or equal to the maximum reconstruction value, the maximum reconstruction value is determined as the predicted value of the pixel point.
48. The method according to claim 31, characterized in that Before determining the reference area and the interpolation filter of the current block, the method further includes: Determining whether the current block allows the use of an interpolation filtering prediction mode; The determining of the reference area and the interpolation filter of the current block comprises: If the current block allows the interpolation filter prediction mode, a reference area and an interpolation filter of the current block are determined.
49. The method according to claim 48, characterized in that The determining whether the current block is allowed to use the interpolation filtering prediction mode comprises: If the current block is in the first row of the current CTU, it is determined that the current block is not allowed to use the interpolation filtering prediction mode.
50. The method according to claim 48, characterized in that The determining whether the current block is allowed to use the interpolation filtering prediction mode comprises: Based on the type of the current image, it is determined whether the current block is allowed to use the interpolation filtering prediction mode.
51. The method according to claim 50, characterized in that The determining, based on the type of the current image, whether the current block is allowed to use the interpolation filtering prediction mode comprises: If the current image is not an intra-prediction image, it is determined that the current block is not allowed to use the interpolation filtering prediction mode.
52. The method of claim 48, wherein: The determining whether the current block is allowed to use the interpolation filtering prediction mode comprises: If the size of the current block is smaller than a preset size, it is determined that the current block is not allowed to use the interpolation filtering prediction mode.
53. The method according to claim 48, characterized in that The determining whether the current block is allowed to use the interpolation filtering prediction mode comprises: Determining first information, where the first information is used to indicate whether a technique based on template matching is enabled; Based on the first information, it is determined whether the current block is allowed to use the interpolation filtering prediction mode.
54. The method according to claim 53, characterized in that The determining, based on the first information, whether the current block is allowed to use the interpolation filtering prediction mode comprises: If the first information indicates that the template matching-based technology is not enabled, it is determined that the current block is not allowed to use the interpolation filtering prediction mode.
55. The method according to claim 53, characterized in that The determining, based on the first information, whether the current block is allowed to use the interpolation filtering prediction mode comprises: If the first information indicates that the template matching-based technology is turned on, determining second information, where the second information is used to indicate whether the current sequence is allowed to be predicted using the interpolation filtering prediction mode; Based on the second information, it is determined whether the current block is allowed to use the interpolation filtering prediction mode.
56. The method according to claim 31, characterized in that The determining the transformation kernel corresponding to the current block includes: Determining an intra prediction mode corresponding to the prediction block; Based on the intra prediction mode corresponding to the prediction block, a transform kernel corresponding to the current block is determined.
57. The method according to claim 56, characterized in that The determining the intra prediction mode corresponding to the prediction block includes: Determine the angle values of R points in the prediction block, where R is a positive integer; Based on the angle values of the R points, an intra-frame prediction mode corresponding to the prediction block is determined.
58. The method according to claim 56, characterized in that The determining, based on the intra prediction mode corresponding to the prediction block, a transform kernel corresponding to the current block comprises: Acquire a correspondence between an intra-frame prediction mode and a transform core group, wherein one transform core group includes at least one type of transform core; In the corresponding relationship, searching for a first transform core group corresponding to the intra prediction mode of the prediction block; A transform core corresponding to the current block is determined from the first transform core group.
59. The method according to claim 31, characterized in that The determining of the reference area of the current block includes: A reference area of the current block is determined in preset P reference areas, where P is a positive integer greater than 1.
60. The method according to claim 31, characterized in that The determining of the interpolation filter of the current block comprises: An interpolation filter for the current block is determined among preset Q interpolation filters, where Q is a positive integer greater than 1.
61. A video decoding device, characterized in that: include: A coefficient determination unit, configured to determine a reference area and an interpolation filter of a current block, and determine a filter coefficient of the interpolation filter based on the reference area; A prediction unit, configured to perform parallel prediction on at least two pixels in the current block using the interpolation filter based on the filter coefficient to determine a prediction block of the current block; The reconstruction unit is used to determine a transformation kernel corresponding to the current block, and determine a reconstructed block of the current block based on the transformation kernel corresponding to the current block and the prediction block.
62. A video encoding device, characterized in that: include: A coefficient determination unit, configured to determine a reference area and an interpolation filter of a current block, and determine a filter coefficient of the interpolation filter based on the reference area; A prediction unit, configured to perform parallel prediction on at least two pixels in the current block using the interpolation filter based on the filter coefficient to determine a prediction block of the current block; The encoding unit is used to determine a transformation kernel corresponding to the current block, and based on the transformation kernel corresponding to the current block and the prediction block, encode the current block to obtain a code stream.
63. An electronic device, characterized in that: including a processor and a memory; The memory shown is used to store computer programs; The processor is used to call and run the computer program stored in the memory to implement the method described in any one of claims 1 to 30 or 41 to 60 above.
64. A video encoding and decoding system, characterized in that: include: Video encoders and video decoders; The video decoder is used to implement the method described in any one of claims 1 to 30 above; The video encoder is used to implement the method described in any one of claims 31 to 60.
65. A computer-readable storage medium, characterized in that For storing computer programs; The computer program enables a computer to execute the method according to any one of claims 1 to 30 or 41 to 60 above.