Video coding and decoding method, device and equipment

By implicitly exporting the target interpolation filter, the problem of excessive bit usage in the existing technology is solved, and the video encoding and codec performance is improved.

CN120128731APending Publication Date: 2025-06-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311692288.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-09
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing video encoding and decoding methods based on interpolation filters occupy too many bits, resulting in unsatisfactory video encoding and decoding performance.

Method used

By implicitly exporting the target interpolation filter, no separate indication is required, bit savings and improved video encoding and codec performance.

Benefits of technology

It realizes bit saving, improves video encoding and codec performance, and avoids bit waste caused by explicit indication interpolation filters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128731A_ABST
    Figure CN120128731A_ABST
Patent Text Reader

Abstract

The invention provides a video coding and decoding method, device and equipment, which can be applied to the calculation fields of image coding and decoding, video coding and decoding and the like, and the decoding method comprises the following steps: decoding a code stream to obtain at least one of prediction information and a decoding coefficient of a to-be-decoded region, and further decoding the to-be-decoded region on the basis of at least one of the prediction information and the decoding coefficient. And implicitly deriving a target interpolation filter corresponding to the to-be-decoded region, and then interpolating a reference region of the to-be-decoded region by using the target interpolation filter to determine a predicted value of the to-be-decoded region. The index of the used target interpolation filter is transmitted in various implicit indication modes, so that bits are saved, and the video decoding performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to a video encoding and decoding method, apparatus, and device. Background Art

[0002] Digital video technology can be incorporated into various video devices, such as digital televisions, smartphones, computers, e-readers, or video players. With the development of video technology, the amount of data included in video data is large. To facilitate the transmission of video data, video devices perform video compression technology to make the video data more effectively transmitted or stored.

[0003] Due to the temporal redundancy in video, the temporal redundancy between video frames is eliminated through inter-frame prediction to improve the compression efficiency. In inter-frame prediction, in some cases, an interpolation filter is required to interpolate the reference image. However, current related prediction methods based on interpolation filters occupy too many bits, resulting in unsatisfactory video encoding and decoding performance. Summary of the Invention

[0004] The present application provides a video encoding and decoding method, apparatus, and device, which implicitly derive a target interpolation filter without separate indication, save bits, and thus improve video encoding and decoding performance.

[0005] In a first aspect, the present application provides a video decoding method applied to a decoder, including:

[0006] Decoding a bitstream to obtain at least one of prediction information and decoding coefficients of a region to be decoded;

[0007] Based on at least one of the prediction information and decoding coefficients, implicitly derive a target interpolation filter corresponding to the region to be decoded;

[0008] Determine a reference region of the region to be decoded, and based on the target interpolation filter, interpolate the reference region to determine a predicted value of the region to be decoded.

[0009] In a second aspect, an embodiment of the present application provides a video encoding method applied to an encoder, including:

[0010] Determine prediction information of a region to be encoded;

[0011] Based on the prediction information, determine a target interpolation filter corresponding to the region to be encoded;

[0012] Determine a reference region of the region to be encoded, and based on the target interpolation filter, interpolate the reference region to determine a predicted value of the region to be encoded.

[0013] In a third aspect, the present application provides a video decoding device, including:

[0014] A decoding unit configured to decode a bitstream to obtain at least one of prediction information and decoding coefficients of a region to be decoded;

[0015] An export unit configured to implicitly export a target interpolation filter corresponding to the region to be decoded based on at least one of the prediction information and decoding coefficients;

[0016] An interpolation unit configured to determine a reference region of the region to be decoded and perform interpolation on the reference region based on the target interpolation filter to determine a predicted value of the region to be decoded.

[0017] In a fourth aspect, the present application provides a video encoding device, including:

[0018] An acquisition unit configured to determine prediction information of a region to be encoded;

[0019] A determination unit configured to determine a target interpolation filter corresponding to the region to be encoded based on the prediction information;

[0020] An interpolation unit configured to determine a reference region of the region to be encoded and perform interpolation on the reference region based on the target interpolation filter to determine a predicted value of the region to be encoded.

[0021] In a fifth aspect, a video decoder is provided, including a processor and a memory. The memory is configured to store a computer program, and the processor is configured to call and run the computer program stored in the memory to execute the method in the first aspect or its various implementation manners above.

[0022] In a sixth aspect, a video encoder is provided, including a processor and a memory. The memory is configured to store a computer program, and the processor is configured to call and run the computer program stored in the memory to execute the method in the second aspect or its various implementation manners above.

[0023] In a seventh aspect, a video encoding and decoding system is provided, including a video encoder and a video decoder. The video decoder is configured to execute the method in the first aspect or its various implementation manners above, and the video encoder is configured to execute the method in the second aspect or its various implementation manners above.

[0024] In an eighth aspect, a chip is provided for implementing the method in any one of the first aspect to the second aspect or its various implementation manners. Specifically, the chip includes: a processor configured to call and run a computer program from a memory, such that a device installed with the chip executes the method in any one of the first aspect to the second aspect or its various implementation manners.

[0025] In a ninth aspect, a computer-readable storage medium is provided for storing a computer program, which causes a computer to execute the method in any one of the above first aspect to second aspect or its various implementation manners.

[0026] In a tenth aspect, a computer program product is provided, including computer program instructions, which cause a computer to execute the method in any one of the above first aspect to second aspect or its various implementation manners.

[0027] In an eleventh aspect, a computer program is provided, which when running on a computer, causes the computer to execute the method in any one of the above first aspect to second aspect or its various implementation manners.

[0028] Based on the above technical solutions, the present application implicitly indicates a target interpolation filter. Specifically, the decoding end decodes the bitstream to obtain at least one of prediction information and decoding coefficients of the area to be decoded, and then implicitly derives the target interpolation filter corresponding to the area to be decoded based on at least one of the prediction information and decoding coefficients, and then uses the target interpolation filter to interpolate the reference area of the area to be decoded to determine the predicted value of the area to be decoded. That is to say, the embodiments of the present application transmit the index of the target interpolation filter used through various implicit indication methods, saving bits and improving the performance of video decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0030] Figure 1 It is a schematic block diagram of a video encoding and decoding system according to an embodiment of the present application;

[0031] Figure 2 It is a schematic block diagram of a video encoder according to an embodiment of the present application;

[0032] Figure 3 It is a schematic block diagram of a video decoder according to an embodiment of the present application;

[0033] Figure 4A It is a schematic diagram of inter-frame prediction;

[0034] Figure 4B It is a schematic diagram of pixel interpolation;

[0035] Figure 5 It is a schematic flowchart of a video decoding method provided by an embodiment of the present application;

[0036] Figure 6 Schematic flowchart of a video encoding method provided by an embodiment of the present application;

[0037] Figure 7 Schematic block diagram of a video decoding device provided by an embodiment of the present application;

[0038] Figure 8 Schematic block diagram of a video encoding device provided by an embodiment of the present application;

[0039] Figure 9 Schematic block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0040] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0041] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In the embodiments of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, but B can also be determined according to A and / or other information. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In the description of the present application, unless otherwise specified, "a plurality of" means two or more than two.

[0042] This application can be applied to the fields of image coding and decoding, video coding and decoding, hardware video coding and decoding, dedicated circuit video coding and decoding, real-time video coding and decoding, etc. For example, the solution of this application can be combined with audio video coding standards (AVS for short), such as the H.264 / Audio Video Coding (AVC for short) standard, the H.265 / High Efficiency Video Coding (HEVC for short) standard, and the H.266 / Versatile Video Coding (VVC for short) standard. Alternatively, the solution of this application can be combined with other proprietary or industry standards for operation, and the standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), including scalable video coding (SVC) and multi-view video coding (MVC) extensions. It should be understood that the technology of this application is not limited to any specific coding and decoding standard or technology.

[0043] For ease of understanding, first, in combination with Figure 1 the video coding and decoding system involved in the embodiments of this application will be introduced.

[0044] Figure 1 It is a schematic block diagram of a video coding and decoding system involved in the embodiments of this application. It should be noted that Figure 1 it is just an example. The video coding and decoding system in the embodiments of this application includes but is not limited to Figure 1 as shown. As Figure 1 shown, the video coding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compressing) video data to generate a bitstream and transmit the bitstream to the decoding device. The decoding device decodes the bitstream generated by the encoding device to obtain the decoded video data.

[0045] The encoding device 110 in the embodiments of this application can be understood as a device with video encoding function, and the decoding device 120 can be understood as a device with video decoding function. That is, the embodiments of this application include a wider range of devices for the encoding device 110 and the decoding device 120, such as including smart phones, desktop computers, mobile computing devices, notebooks (e.g., laptops) computers, tablet computers, set-top boxes, TVs, cameras, display devices, digital media players, video game consoles, in-vehicle computers, etc.

[0046] In some embodiments, the encoding device 110 may transmit the encoded video data (such as a bitstream) to the decoding device 120 via the channel 130. The channel 130 may include one or more media and / or devices capable of transmitting the encoded video data from the encoding device 110 to the decoding device 120.

[0047] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to directly transmit the encoded video data to the decoding device 120 in real time. In this example, the encoding device 110 may modulate the encoded video data according to a communication standard and transmit the modulated video data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media may also include wired communication media, such as one or more physical transmission lines.

[0048] In another example, the channel 130 includes a storage medium that can store the encoded video data of the encoding device 110. The storage medium includes various locally accessible data storage media, such as optical discs, DVDs, flash memories, etc. In this example, the decoding device 120 may obtain the encoded video data from the storage medium.

[0049] In another example, the channel 130 may include a storage server that can store the encoded video data of the encoding device 110. In this example, the decoding device 120 may download the stored encoded video data from the storage server. Optionally, the storage server may store the encoded video data and may transmit the encoded video data to the decoding device 120, such as a web server (e.g., for a website), a File Transfer Protocol (FTP) server, etc.

[0050] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113. Among them, the output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.

[0051] In some embodiments, in addition to including the video encoder 112 and the input interface 113, the encoding device 110 may further include a video source 111.

[0052] The video source 111 may include at least one of a video capture device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system. Among them, the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data.

[0053] Video encoder 112 encodes the video data from video source 111 to generate a bitstream. The video data may include one or more pictures or sequences of pictures. The bitstream contains the encoded information of the pictures or sequences of pictures in the form of a bit stream. The encoded information may include encoded picture data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. The SPS may contain parameters applied to one or more sequences. The PPS may contain parameters applied to one or more pictures. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.

[0054] Video encoder 112 directly transmits the encoded video data to decoding device 120 via output interface 113. The encoded video data may also be stored on a storage medium or a storage server for subsequent reading by decoding device 120.

[0055] In some embodiments, decoding device 120 includes input interface 121 and video decoder 122.

[0056] In some embodiments, decoding device 120 may further include display device 123 in addition to input interface 121 and video decoder 122.

[0057] Among them, input interface 121 includes a receiver and / or a modem. Input interface 121 may receive the encoded video data through channel 130.

[0058] Video decoder 122 is used to decode the encoded video data to obtain the decoded video data and transmit the decoded video data to display device 123.

[0059] Display device 123 displays the decoded video data. Display device 123 may be integrated with decoding device 120 or external to decoding device 120. Display device 123 may include various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0060] In addition, Figure 1 For example only, the technical solutions of the embodiments of the present application are not limited to Figure 1 , for example, the technology of the present application can also be applied to unilateral video encoding or unilateral video decoding.

[0061] The video coding framework involved in the embodiments of the present application will be introduced below.

[0062] Figure 2 It is a schematic block diagram of a video encoder according to an embodiment of the present application. It should be understood that the video encoder 200 can be used for lossy compression of images or lossless compression of images. The lossless compression can be visually lossless compression or mathematically lossless compression.

[0063] The video encoder 200 can be applied to image data in the luminance chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4. Y represents luminance, Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance used to describe color and saturation. For example, in terms of color format, 4:2:0 means that for every 4 pixels, there are 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that for every 4 pixels, there are 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 means full pixel display (YYYYCbCrCbCrCbCrCbCr).

[0064] For example, the video encoder 200 reads video data. For each frame of image in the video data, a frame of image is divided into several coding tree units (CTUs). In some examples, the CTB can be referred to as a "tree block", "Largest Coding Unit" (LCU for short), or "coding tree block" (CTB for short). Each CTU can be associated with a pixel block of equal size within the image. Each pixel can correspond to one luminance sample and two chrominance samples. Therefore, each CTU can be associated with one luminance sample block and two chrominance sample blocks. The size of a CTU is, for example, 128×128, 64×64, 32×32, etc. A CTU can be further divided into several coding units (CUs) for encoding. The CU can be a rectangular block or a square block. The CU can be further divided into a prediction unit (PU for short) and a transform unit (TU for short), so that encoding, prediction, and transformation are separated, making the processing more flexible. In one example, the CTU is divided into CUs in a quadtree manner, and the CU is divided into TUs and PUs in a quadtree manner.

[0065] Video encoders and video decoders can support various PU sizes. Assuming that the size of a specific CU is 2N×2N, the video encoder and video decoder can support a PU size of 2N×2N or N×N for intra prediction, and support symmetric PUs of 2N×2N, 2N×N, N×2N, N×N or similar sizes for inter prediction. The video encoder and video decoder can also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter prediction.

[0066] In some embodiments, as Figure 2 shown, the video encoder 200 may include: a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filter unit 260, a decoded image buffer 270, and an entropy encoding unit 280. It should be noted that the video encoder 200 may include more, fewer, or different functional components.

[0067] Optionally, in this application, the current block can be referred to as the current coding unit (CU) or the current prediction unit (PU), etc. The prediction block can also be referred to as the prediction image block or the image prediction block, and the reconstructed image block can also be referred to as the reconstruction block or the image reconstruction image block.

[0068] In some embodiments, the prediction unit 210 includes an inter prediction unit 211 and an intra prediction unit 212. Since there is a strong correlation between adjacent pixels in a frame of video, the method of intra prediction is used in video coding and decoding technologies to eliminate the spatial redundancy between adjacent pixels. Since there is a strong similarity between adjacent frames in video, the method of inter prediction is used in video coding and decoding technologies to eliminate the temporal redundancy between adjacent frames, thereby improving the coding efficiency.

[0069] The inter-frame prediction unit 211 can be used for inter-frame prediction. Inter-frame prediction can include motion estimation and motion compensation, and can refer to the image information of different frames. Inter-frame prediction uses motion information to find a reference block from a reference frame, and generates a prediction block based on the reference block to eliminate temporal redundancy. The frames used for inter-frame prediction can be P-frames and / or B-frames. A P-frame refers to a forward prediction frame, and a B-frame refers to a bi-directional prediction frame. Inter-frame prediction uses motion information to find a reference block from a reference frame and generates a prediction block based on the reference block. The motion information includes the reference frame list where the reference frame is located, the reference frame index, and the motion vector. The motion vector can be an integer pixel or a fractional pixel. If the motion vector is a fractional pixel, then an interpolation filter needs to be used in the reference frame to create the required fractional pixel block. Here, the integer pixel or fractional pixel block found in the reference frame according to the motion vector is called the reference block. In some technologies, the reference block is directly used as the prediction block, and in some technologies, the prediction block is further processed based on the reference block to generate a prediction block. Processing the reference block to generate a prediction block can also be understood as using the reference block as the prediction block and then processing the prediction block to generate a new prediction block.

[0070] The intra-frame prediction unit 212 only refers to the information of the same frame image and predicts the pixel information within the current coded image block to eliminate spatial redundancy. The frame used for intra-frame prediction can be an I-frame.

[0071] There are multiple intra-frame prediction modes. Taking the international digital video coding standard H series as an example, the H.264 / AVC standard has 8 angular prediction modes and 1 non-angular prediction mode, and the H.265 / HEVC is extended to 33 angular prediction modes and 2 non-angular prediction modes. The intra-frame prediction modes used in HEVC include the Planar mode, DC, and 33 angular modes, for a total of 35 prediction modes. The intra-frame modes used in VVC include Planar, DC, and 65 angular modes, for a total of 67 prediction modes.

[0072] It should be noted that with the increase in the angular mode, the intra-frame prediction will be more accurate and more in line with the requirements for the development of high-definition and ultra-high-definition digital videos.

[0073] The residual unit 220 can generate a residual block of the CU based on the pixel block of the CU and the prediction block of the PU of the CU. For example, the residual unit 220 can generate a residual block of the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel block of the CU and the corresponding sample in the prediction block of the PU of the CU.

[0074] The transform / quantization unit 230 may quantize the transform coefficients. The transform / quantization unit 230 may quantize the transform coefficients associated with the TUs of a CU based on the quantization parameter (QP) value associated with the CU. The video encoder 200 may adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.

[0075] The inverse transform / quantization unit 240 may apply inverse quantization and inverse transform to the quantized transform coefficients respectively to reconstruct the residual block from the quantized transform coefficients.

[0076] The reconstruction unit 250 may add the samples of the reconstructed residual block to the corresponding samples of one or more prediction blocks generated by the prediction unit 210 to generate a reconstructed image block associated with the TU. By reconstructing the sampled blocks of each TU of the CU in this way, the video encoder 200 may reconstruct the pixel block of the CU.

[0077] The loop filter unit 260 is used to process the pixels after inverse transform and inverse quantization, compensate for the distorted information, and provide a better reference for subsequent encoded pixels. For example, it may perform a deblocking filter operation to reduce the blocking effect of the pixel block associated with the CU.

[0078] In some embodiments, the loop filter unit 260 includes a deblocking filter unit and a sample adaptive offset / adaptive loop filter (SAO / ALF) unit, where the deblocking filter unit is used to remove the blocking effect and the SAO / ALF unit is used to remove the ringing effect.

[0079] The decoded picture buffer 270 may store the reconstructed pixel blocks. The inter prediction unit 211 may perform inter prediction on the PUs of other pictures using the reference pictures containing the reconstructed pixel blocks. Additionally, the intra prediction unit 212 may perform intra prediction on the PUs in the same picture as the CU using the reconstructed pixel blocks in the decoded picture buffer 270.

[0080] The entropy coding unit 280 may receive the quantized transform coefficients from the transform / quantization unit 230. The entropy coding unit 280 may perform one or more entropy coding operations on the quantized transform coefficients to generate the entropy-coded data.

[0081] Figure 3 It is a schematic block diagram of a video decoder according to an embodiment of the present application.

[0082] As Figure 3 shown, the video decoder 300 includes: an entropy decoding unit 310, a prediction unit 320, an inverse quantization / transform unit 330, a reconstruction unit 340, a loop filter unit 350, and a decoded picture buffer 360. It should be noted that the video decoder 300 may include more, fewer, or different functional components.

[0083] Video decoder 300 may receive a bitstream. Entropy decoding unit 310 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, entropy decoding unit 310 may parse the entropy-coded syntax elements in the bitstream. Prediction unit 320, inverse quantization / transformation unit 330, reconstruction unit 340, and loop filter unit 350 may decode video data according to the syntax elements extracted from the bitstream, that is, generate decoded video data.

[0084] In some embodiments, prediction unit 320 includes an intra prediction unit 322 and an inter prediction unit 321.

[0085] Intra prediction unit 322 may perform intra prediction to generate a prediction block of the PU. Intra prediction unit 322 may use an intra prediction mode to generate a prediction block of the PU based on pixel blocks of spatially adjacent PUs. Intra prediction unit 322 may also determine the intra prediction mode of the PU according to one or more syntax elements parsed from the bitstream.

[0086] Inter prediction unit 321 may construct a first reference picture list (list 0) and a second reference picture list (list 1) according to the syntax elements parsed from the bitstream. In addition, if the PU is encoded using inter prediction, entropy decoding unit 310 may parse the motion information of the PU. Inter prediction unit 321 may determine one or more reference blocks of the PU according to the motion information of the PU. Inter prediction unit 321 may generate a prediction block of the PU according to one or more reference blocks of the PU.

[0087] Inverse quantization / transformation unit 330 may inverse quantize (i.e., dequantize) the transform coefficients associated with the TU. Inverse quantization / transformation unit 330 may use the QP value associated with the CU of the TU to determine the degree of quantization.

[0088] After inverse quantizing the transform coefficients, inverse quantization / transformation unit 330 may apply one or more inverse transforms to the inverse quantized transform coefficients to generate a residual block associated with the TU.

[0089] Reconstruction unit 340 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, reconstruction unit 340 may add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the pixel block of the CU, obtaining a reconstructed image block.

[0090] Loop filter unit 350 may perform a deblocking filter operation to reduce the blocking effect of the pixel block associated with the CU.

[0091] The video decoder 300 may store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 may use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.

[0092] The basic process of video coding and decoding is as follows: At the encoding end, a frame of image is divided into blocks. For the current block, the prediction unit 210 generates a predicted block of the current block using intra prediction or inter prediction. The residual unit 220 may calculate a residual block based on the predicted block and the original block of the current block, that is, the difference between the predicted block and the original block of the current block. This residual block may also be referred to as residual information. The residual block undergoes processes such as transformation and quantization by the transform / quantization unit 230, which can remove information that is insensitive to the human eye to eliminate visual redundancy. Optionally, the residual block before being transformed and quantized by the transform / quantization unit 230 may be referred to as a temporal residual block, and the temporal residual block after being transformed and quantized by the transform / quantization unit 230 may be referred to as a frequency residual block or a frequency-domain residual block. The entropy coding unit 280 receives the quantized transform coefficients output by the transform quantization unit 230 and may perform entropy coding on the quantized transform coefficients to output a bitstream. For example, the entropy coding unit 280 may eliminate character redundancy according to the target context model and the probability information of the binary bitstream.

[0093] At the decoding end, the entropy decoding unit 310 may parse the bitstream to obtain prediction information, a quantized coefficient matrix, etc. of the current block. The prediction unit 320 performs intra prediction or inter prediction on the current block based on the prediction information to generate a predicted block of the current block. The inverse quantization / transformation unit 330 uses the quantized coefficient matrix obtained from the bitstream to perform inverse quantization and inverse transformation on the quantized coefficient matrix to obtain a residual block. The reconstruction unit 340 adds the predicted block and the residual block to obtain a reconstructed block. The reconstructed blocks form a reconstructed image. The loop filter unit 350 performs loop filtering on the reconstructed image based on the image or based on blocks to obtain a decoded image. The encoding end also needs to perform operations similar to those at the decoding end to obtain a decoded image. This decoded image may also be referred to as a reconstructed image, and the reconstructed image may be used as a reference frame for inter prediction for subsequent frames.

[0094] It should be noted that the block partitioning information determined at the encoding end, as well as mode information or parameter information such as prediction, transformation, quantization, entropy coding, and loop filtering, are carried in the bitstream when necessary. The decoding end determines the same block partitioning information, mode information or parameter information such as prediction, transformation, quantization, entropy coding, and loop filtering as the encoding end by parsing the bitstream and analyzing based on the existing information, so as to ensure that the decoded image obtained at the encoding end is the same as the decoded image obtained at the decoding end.

[0095] The above is the basic process of a video codec under a block-based hybrid coding framework. With the development of technology, some modules or steps of this framework or process may be optimized. This application is applicable to the basic process of a video codec under this block-based hybrid coding framework, but is not limited to this framework and process.

[0096] In the embodiments of this application, the current block can be the current coding unit (CU) or the current prediction unit (PU), etc. Due to the need for parallel processing, an image can be divided into slices, etc. Slices in the same image can be processed in parallel, that is, there is no data dependency between them. And "frame" is a common term. Generally, it can be understood that one frame is an image. The frame mentioned in the application can also be replaced by an image or a slice, etc.

[0097] Inter-frame prediction refers to the process of predicting a reference block for a block to be encoded (such as the current block) in the current image from a neighboring encoded image (such as a reference image). The purpose is to remove the temporal redundancy of the video signal. A schematic diagram of inter-frame prediction is as Figure 4A shown. According to the block matching criterion, the best matching block of the current block is searched for in the reference frame.

[0098] Inter-frame prediction uses motion information to represent "motion". The basic motion information includes information about the reference frame (or reference picture) and information about the motion vector (MV). Currently, the commonly used bidirectional prediction uses two reference blocks to predict the current block. The two reference blocks can use one forward reference block and one backward reference block. Later, it is also allowed that both are forward or both are backward. The so-called forward means that the time corresponding to the reference frame is before the current frame, and backward means that the time corresponding to the reference frame is after the current frame. Or it can be said that forward means that the position of the reference frame in the video is before the current frame, and backward means that the position of the reference frame in the video is after the current frame. Or it can be said that forward means that the POC (picture order count) of the reference frame is less than the POC of the current frame, and backward means that the POC of the reference frame is greater than the POC of the current frame. In order to be able to use bidirectional prediction, naturally, two reference blocks need to be found, so two sets of information about the reference frame and motion vector information are required. Each set of them can be understood as a unidirectional motion information, and combining these two sets together forms a bidirectional motion information. In specific implementation, the unidirectional motion information and the bidirectional motion information can use the same data structure, except that the information about the reference frame and motion vector in both sets of the bidirectional motion information is valid, while the information about the reference frame and motion vector in one of the sets of the unidirectional motion information is invalid.

[0099] VVC supports two reference picture lists, denoted as RPL0 and RPL1, where RPL is the abbreviation of Reference Picture List. In VVC, P slices can only use RPL0, and B slices can use RPL0 and RPL1. For a slice, there are several reference pictures in each reference picture list, and the codec finds a certain reference picture through the reference picture index. VVC represents motion information using the reference picture index and the motion vector. For the above bidirectional motion information, VVC uses the reference picture index refIdxL0 corresponding to reference picture list 0, the motion vector mvL0 corresponding to reference picture list 0, the reference picture index refIdxL1 corresponding to reference picture list 1, and the motion vector mvL0 corresponding to reference picture list 1. The reference picture index corresponding to reference picture list 0 and the reference picture index corresponding to reference picture list 1 here can be understood as the information of the above reference pictures. VVC uses two flag bits to represent whether to use the motion information corresponding to reference picture list 0 and whether to use the motion information corresponding to reference picture list 0, denoted as predFlagL0 and predFlagL1 respectively. It can also be understood that predFlagL0 and predFlagL1 represent whether the above unidirectional motion information is "valid". Therefore, although there is no explicit mention of the data structure of motion information in VVC, it uses the reference picture index, motion vector, and the flag bit of "whether it is valid" corresponding to each reference picture list to represent motion information together. In the VVC standard text, motion information does not appear, but motion vectors are used. It can also be considered that the reference picture index and the flag bit of whether to use the corresponding motion information are appendages of the motion vector. In this article, for the convenience of description, "motion information" is still used, but it should be understood that "motion vector" can also be used to describe.

[0100] As can be seen from the above, inter-frame prediction needs to use the motion vector to obtain the pixel prediction value in the reference picture. In actual situations, the motion of objects between adjacent images does not necessarily take the whole pixel as the basic unit. Therefore, in order to improve the prediction accuracy, it is necessary to improve the accuracy of motion estimation to the sub-pixel level, interpolate the reference picture to improve the accuracy of motion compensation, and then improve the coding efficiency. As Figure 4B shown, the pixel points where the capital A is located are whole pixels, such as A 0,0 、A 1,0 etc., and the pixel points where the lowercase a, b, c, etc. are located are sub-pixels, such as a 0,0 is a 1 / 4 pixel point in the X direction, b 0,0It is 1 / 2 pixel point in the X direction, etc. Exemplarily, the common inter-frame prediction modes in AVS3 support five motion vector precisions of 1 / 4, 1 / 2, 1, 2, and 4. When the MV precision is 1 / 4 or 1 / 2, exemplary 8-tap interpolation filters can be used to interpolate the luminance prediction value and 4-tap interpolation filters can be used to interpolate the chrominance prediction value.

[0101] Traditional bidirectional prediction obtains the prediction value of the current block by weighted averaging of two reconstructed blocks, where the two reconstructed blocks come from the forward reference frame and the backward reference frame respectively. Further, in order to improve the prediction effect, the bidirectional optical flow (BIO) technology can be adopted to compensate for the motion after bidirectional prediction, so as to reduce the motion deviation and improve the coding efficiency. Specifically, BIO in AVS3 first calculates the gradient values in the x direction and the y direction (in the full text, the x direction represents the horizontal direction and the y direction represents the vertical direction), then obtains the calculation factor of each pixel according to the pixel value and the gradient value, derives the motion vector, and finally calculates a more accurate prediction value. Exemplarily, when calculating the gradient value in the x direction, BIO first uses an 8-tap gradient filter for gradient calculation, and then uses an 8-tap interpolation filter for interpolation. When calculating the gradient value in the y direction, BIO first uses an 8-tap interpolation filter for interpolation, and then uses an 8-tap gradient filter for gradient calculation. In addition, the BIO technology is only applied to the luminance component.

[0102] Adaptive Motion Vector Resolution (AMVR), traditional video coding adopts motion vectors with different precisions to obtain more accurate motion estimation results, thereby improving the coding performance. AVS3 adopts 5 motion vector precisions of 1 / 4, 1 / 2, 1, 2, and 4 as shown in Table 1. These 5 motion vector precisions are encoded at the coding end, and the optimal motion vector precision is selected through RDO, and then the corresponding index is transmitted to the decoding end.

[0103] Table 1

[0104] MVR Index 0 1 2 3 4 Resolution 1 / 4 1 / 2 1 2 4

[0105] The object motion between adjacent images is relatively complex, not just the translation between pixels. In order to express more complex motions, traditional video coding standards such as VVC and AVS3 both adopt the Affine motion compensation (AFFINE) technology to obtain more accurate motion compensation results. Among them, the affine motion compensation supports 3 motion vector precisions of 1 / 16, 1 / 4, and 1.

[0106] The History based Motion Vector Prediction (HMVP) technology copies 8 motion information candidates from previously encoded blocks into the FIFO and continuously updates them in a first-in, first-out manner. The candidate motion information (such as MVP) selected by HMVP can be directly used as the motion information of the current encoded block.

[0107] The Implicit selection of transforms (IST) is a new intra-frame coding transform tool in AVS3. IST provides two separable transform kernels and selects the optimal transform kernel through RDO. To reduce bit consumption, IST does not transmit the index of the selected transform kernel in the bitstream, but binds it to the parity of the number of even coefficients among the non-zero transform coefficients, as shown in Table 2. In this way, the decoding end derives the corresponding transform kernel based on this implicit indication.

[0108] Table 2

[0109]

[0110] Currently, in the related encoding and decoding technologies based on interpolation filters, the encoding end determines the target interpolation filter and writes the index of the target interpolation filter into the bitstream, so that the decoding end obtains the index of the target interpolation filter by decoding the bitstream, and then obtains the target interpolation filter based on this index. In this way, explicitly indicating the target interpolation filter will waste bits and thus reduce the performance of video encoding and decoding.

[0111] To solve the above technical problems, the embodiments of the present application indicate the target interpolation filter in an implicit manner. Specifically, the decoding end decodes the bitstream to obtain at least one of the prediction information and decoding coefficients of the area to be decoded, and then implicitly derives the target interpolation filter corresponding to the area to be decoded based on at least one of the prediction information and decoding coefficients, and then uses the target interpolation filter to interpolate the reference area of the area to be decoded to determine the predicted value of the area to be decoded. That is to say, the embodiments of the present application transmit the index of the target interpolation filter used in various implicit indication ways, saving bits and improving the performance of video encoding and decoding.

[0112] The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0113] First, taking the decoding end as an example, the video decoding method provided by the embodiments of the present application will be introduced.

[0114] Figure 5Schematic diagram of the video decoding method provided by an embodiment of the present application. The embodiment of the present application is applied to Figure 1 and Figure 3 the video decoders shown. As shown in Figure 5 , the method of the embodiment of the present application includes:

[0115] S101. Decode the bitstream to obtain at least one of the prediction information and decoding coefficients of the area to be decoded.

[0116] The embodiment of the present application does not limit the specific size and shape of the area to be decoded.

[0117] In some embodiments, the area to be decoded may be one or several image blocks to be decoded in the current frame to be decoded. For example, it may be one or several CTUs, or one or several CUs, or one or several PUs, etc. In one example, the image block to be decoded may be referred to as the current block, the current image block to be decoded, and so on.

[0118] In some embodiments, the area to be decoded may be the current frame to be decoded, that is, the area to be decoded is a whole-frame image.

[0119] During the video decoding process, when the decoding end decodes the area to be decoded, it decodes the bitstream to obtain the quantization coefficients of the area to be decoded, inverse-quantizes the quantization coefficients to obtain the transform coefficients of the area to be decoded, inverse-transforms the transform coefficients to obtain the residual values of the area to be decoded. Then, it determines the prediction mode of the area to be decoded, determines the predicted value of the area to be decoded based on the prediction mode, and obtains the reconstructed value of the area to be decoded based on the predicted value and the residual value of the area to be decoded.

[0120] The embodiment of the present application mainly relates to the prediction process of the area to be decoded.

[0121] In the embodiment of the present application, in order to improve the prediction accuracy of the area to be decoded, an interpolation filter is used to interpolate the reference area of the area to be decoded. However, if the encoding end explicitly indicates the determined target interpolation filter of the area to be decoded to the decoding end, for example, writing the index of the target interpolation filter into the bitstream, it will waste bits and thus reduce the performance of video coding and decoding.

[0122] To solve the above technical problem, in the embodiment of the present application, the target interpolation filter is indicated in an implicit manner. Specifically, the encoding end does not need to write the index of the determined target interpolation filter into the bitstream, and the decoding end can implicitly derive the target interpolation filter corresponding to the area to be decoded based on at least one of the prediction information and decoding coefficients of the area to be decoded, thereby saving bits and improving the performance of video coding and decoding.

[0123] The embodiments of the present application do not limit the specific content of the prediction information of the area to be decoded, which can be understood as all information required for predicting the predicted value of the area to be decoded, such as at least one of a prediction mode, component information, a reference frame list, motion vector precision, and the like.

[0124] In some embodiments, the prediction information of the area to be decoded includes at least one of the prediction mode of the area to be decoded and the motion vector precision corresponding to the area to be decoded.

[0125] The embodiments of the present application do not limit the specific content of the decoding coefficients of the area to be decoded. Exemplarily, the decoding coefficients may include transform coefficients, quantization coefficients, and the like.

[0126] In one example, when the decoding coefficient is a quantization coefficient, the decoding end decodes the bitstream to obtain the quantization coefficient of the area to be decoded.

[0127] In one example, when the decoding coefficient is a transform coefficient, the decoding end decodes the bitstream to obtain the quantization coefficient of the area to be decoded, and performs an inverse transform on the quantization coefficient to obtain the transform coefficient of the area to be decoded.

[0128] After the decoding end obtains at least one of the prediction information and the decoding coefficients of the area to be decoded, it performs the following steps of S102.

[0129] S102: Implicitly derive a target interpolation filter corresponding to the area to be decoded based on at least one of the prediction information and the decoding coefficients.

[0130] In the embodiments of the present application, the decoding end implicitly derives a target interpolation filter corresponding to the area to be decoded from interpolation filters with different numbers of taps based on at least one of the prediction information and the decoding coefficients, saving bits and thus improving the performance of video coding and decoding.

[0131] The embodiments of the present application do not limit the specific manner in which the decoding end implicitly derives a target interpolation filter corresponding to the area to be decoded based on at least one of the prediction information and the decoding coefficients.

[0132] In some embodiments, the decoding end implicitly derives a target interpolation filter corresponding to the area to be decoded based on the prediction information.

[0133] In some embodiments, the decoding end implicitly derives a target interpolation filter corresponding to the area to be decoded based on the decoding coefficients.

[0134] In some embodiments, the decoding end implicitly derives a target interpolation filter corresponding to the area to be decoded based on the prediction information and the decoding coefficients.

[0135] In some embodiments, the prediction information of the region to be decoded includes at least one of an index of the motion vector precision corresponding to the region to be decoded and a prediction mode. In this case, the above S102 includes the following steps S102-A1 to S102-A3:

[0136] S102-A1. Determine the motion vector precision corresponding to the region to be decoded based on the index of the motion vector precision.

[0137] S102-A2. Implicitly derive a target interpolation filter based on at least one of the motion vector precision, the prediction mode, and the decoding coefficient.

[0138] In this implementation, if the prediction information of the region to be decoded includes at least one of an index of the motion vector precision corresponding to the region to be decoded and a prediction mode, the decoding end determines the motion vector precision corresponding to the region to be decoded based on the index of the motion vector precision. Furthermore, a target interpolation filter is implicitly derived based on at least one of the motion vector precision, the prediction mode, and the decoding coefficient.

[0139] In the embodiments of the present application, the specific manner for the decoding end to obtain the index of the motion vector precision corresponding to the region to be decoded includes, but is not limited to, the following several types:

[0140] In one example, the encoding end selects the motion vector precision corresponding to the region to be decoded from multiple candidate motion vector precisions. For example, based on the rate-distortion cost, the motion vector precision corresponding to the region to be decoded is selected from multiple candidate motion vector precisions, and then the index of the selected motion vector precision corresponding to the region to be decoded is written into the code stream. In this way, the decoding end can obtain the index of the motion vector precision corresponding to the region to be decoded by decoding the code stream.

[0141] In another example, the decoding end can implicitly derive the index of the motion vector precision corresponding to the region to be decoded through some decoding information obtained by decoding. Alternatively, both the encoding and decoding ends default to use a certain motion vector precision as the motion vector precision corresponding to the region to be decoded.

[0142] In some embodiments of the embodiments of the present application, when encoding the region to be decoded, the encoding end selects the prediction mode of the region to be decoded from multiple candidate prediction modes, and then writes the index of the prediction mode into the code stream. In this way, the decoding end obtains the prediction mode of the region to be decoded by decoding the code stream.

[0143] Then, the decoding end implicitly derives a target interpolation filter based on at least one of the motion vector precision, the prediction mode, and the decoding coefficient.

[0144] The embodiments of the present application do not limit the specific manner for the decoding end to implicitly derive a target interpolation filter based on at least one of the motion vector precision, the prediction mode, and the decoding coefficient.

[0145] In Case 1, the decoding end implicitly derives a target interpolation filter based on any one of the motion vector precision, prediction mode, and decoding coefficients. For example, it is stipulated at both the encoding and decoding ends that interpolation filter 1 is used in prediction mode 1. Another example is that for high-precision motion vectors, an interpolation filter with more taps is used, and for low-precision motion vectors, an interpolation filter with fewer taps is used.

[0146] In Case 2, the decoding end implicitly derives a target interpolation filter based on any two of the motion vector precision, prediction mode, and decoding coefficients.

[0147] In an example of Case 2, the above S102-A2 includes the following steps of S102-A2-a:

[0148] S102-A2-a: Implicitly derive a target interpolation filter based on the motion vector precision and prediction mode.

[0149] In this implementation, the decoding end can implicitly derive a target interpolation filter based on the motion vector precision and prediction mode corresponding to the area to be decoded. For example, the candidate motion vector precisions corresponding to different prediction modes may be different. Furthermore, based on the candidate motion vector precision corresponding to the prediction mode of the area to be decoded and the motion vector precision corresponding to the area to be decoded, a target interpolation filter can be implicitly derived. That is to say, the embodiments of the present application can reduce bit consumption and improve decoding efficiency by binding the index of the motion vector and the index of the non-fixed-length interpolation filter.

[0150] In the embodiments of the present application, generally, if the motion vector precision is sub-pixel, interpolation is performed; if the motion vector precision is integer pixel, no interpolation is performed. Therefore, before the decoding end implicitly derives a target interpolation filter based on the motion vector precision and prediction mode, it is first necessary to determine whether the motion vector precision corresponding to the area to be decoded is sub-pixel. If it is sub-pixel, a target interpolation filter is implicitly derived based on the motion vector precision and prediction mode corresponding to the area to be decoded. If the motion vector precision corresponding to the area to be decoded is integer pixel, the step of implicitly deriving the target interpolation filter is skipped. Based on this, the specific implementation of the above S102-A2-a includes at least the following examples:

[0151] In a possible implementation of S102-A2-a, if the motion vector precision is sub-pixel motion vector precision and the prediction mode is non-affine motion compensation mode, then based on the motion vector precision, in the first relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the non-affine motion compensation mode, a target interpolation filter is derived.

[0152] The embodiments of the present application do not limit the specific quantity and type of the sub-pixel motion vector precision included in the first relationship list.

[0153] In a possible implementation, the first relationship list includes the sub-pixel motion vector precision among the n motion vector precisions corresponding to Adaptive Motion Vector Resolution (AMVR).

[0154] Exemplarily, the n motion vector precisions may include five motion vector precisions: 1 / 4, 1 / 2, 1, 2, and 4. Of course, other motion vector precisions may also be included, and the embodiments of the present application do not limit this.

[0155] Exemplarily, the first relationship list between different sub-pixel motion vector precisions corresponding to the non-affine motion compensation mode and different interpolation filters is shown in Table 3 as follows:

[0156] Table 3

[0157]

[0158] In Table 3, AmvrIndex represents the index of the motion vector precision, Resolution represents the motion vector precision, and Interpolation Filter represents the number of taps of the interpolation filter. As shown in Table 3, the embodiments of the present application bind the index of the non-fixed-length interpolation filter to the index of the motion vector precision and the prediction mode, so that the decoding end can implicitly derive the target interpolation filter corresponding to the area to be decoded through the index of the motion vector precision corresponding to the area to be decoded and the prediction mode.

[0159] In one example, in the non-affine motion compensation mode, the encoding end uses 3 bits to encode the index of the motion vector precision. In this way, the decoding end decodes the code stream to obtain the 3-bit index of the motion vector precision, and then based on the 3-bit index of the motion vector precision, the target interpolation filter is derived in Table 4 below.

[0160] Table 4

[0161]

[0162] It should be noted that Table 4 can be understood as another form of Table 3, and the above Table 3 and the above Table 4 are only examples of the first relationship list. The first relationship list of the embodiments of the present application includes but is not limited to those shown in Table 3 and Table 4 above.

[0163] In another possible implementation of S102-A2-a, if the motion vector precision is sub-pixel motion vector precision and the prediction mode is affine motion compensation mode, then based on the motion vector precision, in the second relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the affine motion compensation mode, the target interpolation filter is derived.

[0164] The embodiments of the present application do not limit the specific quantity and type of the sub-pixel motion vector precision included in the second relationship list.

[0165] In a possible implementation, the second relationship list includes the sub-pixel motion vector precision among the m motion vector precisions corresponding to the affine motion compensation mode.

[0166] Exemplarily, the m motion vector precisions may include 3 kinds of motion vector precisions of 1 / 16, 1 / 4, and 1. Of course, other motion vector precisions may also be included, and the embodiments of the present application do not limit this.

[0167] Exemplarily, the second relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the affine motion compensation mode is shown in Table 5 as follows:

[0168] Table 5

[0169]

[0170] In Table 5, AffineAmvrIndex represents the index of the motion vector precision, Resolution represents the motion vector precision, and Interpolation Filter represents the number of taps of the interpolation filter. As shown in Table 5, the embodiments of the present application bind the index of the variable-length interpolation filter to the motion vector precision and the affine motion compensation mode, so that the decoding end can implicitly derive the target interpolation filter corresponding to the area to be decoded through the motion vector precision and the prediction mode corresponding to the area to be decoded.

[0171] In one example, in the affine motion compensation mode, the encoding end uses 2 bits to encode the index of the motion vector precision. In this way, the decoding end decodes the bitstream to obtain the 2-bit index of the motion vector precision, and then based on the 2-bit index of the motion vector precision, the target interpolation filter is derived in Table 6 below.

[0172] Table 6

[0173]

[0174]

[0175] It should be noted that Table 6 can be understood as another form of Table 5, and the above Table 5 and the above Table 6 are only examples of the second relationship list. The second relationship list in the embodiments of the present application includes but is not limited to those shown in the above Table 5 and Table 6.

[0176] In some embodiments, in the above first relationship list and second relationship list, the number of taps of the interpolation filter corresponding to the sub-pixel motion vector accuracy with high precision is greater than the number of taps of the interpolation filter corresponding to the sub-pixel motion vector accuracy with low precision. As shown in Table 3 for example, the 1 / 4 motion vector accuracy corresponds to an interpolation filter with 12 taps, while the 1 / 2 motion vector accuracy corresponds to an interpolation filter with 8 taps. Another example is shown in Table 5, where the 1 / 16 motion vector accuracy corresponds to an interpolation filter with 12 taps, while the 1 / 4 motion vector accuracy corresponds to an interpolation filter with 8 taps. In this way, using an interpolation filter with more taps for the sub-pixel motion vector accuracy with high precision for interpolation can improve the accuracy of motion estimation and thus enhance the prediction effect.

[0177] In another example of Case 2, the above S102 - A2 includes the following steps of S102 - A2 - b and S102 - A2 - c:

[0178] S102 - A2 - b: If the prediction mode is motion vector prediction based on historical information, then decode the code stream to obtain the index of the motion vector prediction MVP corresponding to the region to be decoded.

[0179] S102 - A2 - c: Derive the target interpolation filter based on the index of the MVP and the motion vector accuracy.

[0180] In this embodiment, when the decoding end determines the target interpolation filter corresponding to the region to be decoded based on the above prediction mode and motion vector accuracy, it first judges the prediction mode of the region to be decoded. If the prediction mode of the region to be decoded is the motion vector prediction mode based on historical information, then the decoding end decodes the code stream to obtain the index of the motion vector prediction MVP corresponding to the region to be decoded, and then derives the target interpolation filter based on the index of the MVP and the motion vector accuracy. That is to say, in this implementation manner, by binding the index of the MVP in the HMVP and the index of the variable - length interpolation filter, the bit consumption can be reduced and the coding efficiency can be improved.

[0181] In the embodiments of the present application, generally, interpolation is performed when the motion vector precision is sub-pixel, and no interpolation is performed when the motion vector precision is integer pixel. Therefore, before the decoding end implicitly derives the target interpolation filter based on the motion vector precision and the prediction mode, it is first necessary to determine whether the motion vector precision corresponding to the area to be decoded is sub-pixel. If it is sub-pixel, the target interpolation filter is implicitly derived based on the motion vector precision and the prediction mode corresponding to the area to be decoded. If the motion vector precision corresponding to the area to be decoded is integer pixel, the step of implicitly deriving the target interpolation filter is skipped. Based on this, the above S102-A2-c includes the step of S102-A2-c1:

[0182] S102-A2-c1. If the motion vector precision is sub-pixel motion vector precision, the target interpolation filter is derived based on the index of the MVP.

[0183] The embodiments of the present application do not limit the specific manner of deriving the target interpolation filter based on the index of the MVP.

[0184] In some embodiments, as can be seen from the above, the HMVP technology copies k (for example, 8) pieces of motion information from the previously encoded blocks as candidates. The encoding and decoding ends can default that these k candidate MVPs correspond to interpolation filters with different tap numbers. In this way, the decoding end can derive the target interpolation filter corresponding to the area to be decoded from the interpolation filters with different tap numbers corresponding to these k candidate MVPs based on the index of the MVP corresponding to the area to be decoded.

[0185] In one example, assuming k is equal to 8, the correspondence between the indices of the 8 candidate MVPs and the interpolation filters with different tap numbers is shown in Table 7:

[0186] Table 7

[0187]

[0188]

[0189] In Table 7 above, hmvpIndex represents the index of the candidate MVP, and Interpolation Filter represents the number of taps of the interpolation filter. As shown in Table 7, in the motion vector prediction mode based on historical information, the index of the MVP in the HMVP and the index of the non-fixed-length interpolation filter are bound, so that the decoding end can implicitly derive the target interpolation filter corresponding to the area to be decoded through the index of the MVP corresponding to the area to be decoded.

[0190] In one example, in the HMVP mode, the encoding end uses 3 bits to encode the index of the MVP. In this way, the decoding end decodes the bitstream to obtain the 3-bit index of the MVP corresponding to the area to be decoded, and then based on the 3-bit index of the MVP, in Table 8 below, derives the target interpolation filter.

[0191] Table 8

[0192]

[0193] It should be noted that Table 8 can be understood as another form of Table 7, and the above Table 7 and Table 8 are only a corresponding relationship between the index of the MVP and the interpolation filters with different numbers of taps. The corresponding relationship between the index of the MVP and the interpolation filters with different numbers of taps in the embodiments of the present application includes but is not limited to those shown in Table 8 and Table 7 above.

[0194] As can be seen from the above, in the HMVP mode, the encoding end does not need to add an additional flag bit to indicate the target interpolation filter. The decoding end only needs to decode the index of the MVP corresponding to the area to be decoded, and then based on the index of the MVP, implicitly derives the target interpolation filter.

[0195] In some embodiments, the number of taps of the interpolation filter corresponding to the index of the MVP closer to the area to be decoded is greater than the number of taps of the interpolation filter corresponding to the index of the MVP farther from the area to be decoded. The HMVP candidate list is composed of the MVs of several previously decoded blocks. The closer to the area to be decoded, the larger the index assigned to the MV in the candidate list. Assuming that the closer the MVP is to the area to be decoded, the greater the motion correlation with the area to be decoded, it can be considered to use an interpolation filter with more taps for the closer MVP to interpolate a more accurate prediction result. As shown in Table 7 for example, the decoded block corresponding to the MVP with index 7 is the closest to the area to be decoded. Therefore, an interpolation filter with more taps, such as a 12-tap interpolation filter, is configured for the MVP with index 7, and an interpolation filter with fewer taps, such as a 6-tap interpolation filter, is configured for the MVP with index 3.

[0196] The above introduces the process in which the decoding end implicitly derives the target interpolation filter corresponding to the area to be decoded based on the prediction mode of the area to be decoded and the motion vector accuracy of the area to be decoded.

[0197] In some embodiments, the implicit derivation of the target interpolation filter based on at least one of the motion vector accuracy, prediction mode, and decoding coefficient in S102-A2 above further includes the implementation manner of the following Case 3:

[0198] Case 3, S102-A2 above includes the following steps of S102-A2-d and S102-A2-e:

[0199] S102 - A2 - d. If the motion vector accuracy is sub - pixel motion vector accuracy, the reference region is sub - pixel divided based on the motion vector accuracy;

[0200] S102 - A2 - e. For each sub - pixel, based on the position information of the sub - pixel, the target interpolation filter corresponding to the sub - pixel is derived.

[0201] Although traditional video coding uses different pixel precisions for motion estimation and motion compensation to obtain accurate prediction results, at the same pixel precision, the importance of sub - pixels at different positions is not necessarily the same. Therefore, it is possible to implicitly derive the target interpolation filter based on the position information of the sub - pixels.

[0202] Specifically, the decoding end first determines whether interpolation is required based on the motion vector accuracy corresponding to the region to be decoded. If the motion vector accuracy corresponding to the region to be decoded is sub - pixel motion vector accuracy, the decoding end sub - pixel divides the reference region corresponding to the region to be decoded based on this sub - pixel motion vector accuracy. The specific division method refers to the above Figure 4B as shown, and will not be elaborated here.

[0203] For each sub - pixel obtained by the division, based on the position information of the sub - pixel, the target interpolation filter corresponding to the sub - pixel is derived.

[0204] The embodiments of the present application do not limit the specific method for the decoding end to derive the target interpolation filter corresponding to the sub - pixel based on the position information of the sub - pixel.

[0205] In a possible implementation, if the sub - pixel is a key pixel point of the region to be decoded, such as a point in the central region, an interpolation filter with more taps is assigned to the sub - pixel to improve the prediction accuracy. On the contrary, if the sub - pixel is not a key point of the region to be decoded, an interpolation filter with fewer taps can be assigned to the sub - pixel to improve the preset speed.

[0206] In a possible implementation, the above S102 - A2 - e includes the following steps of S102 - A2 - e1:

[0207] S102 - A2 - e1. Based on the parity of the position of the sub - pixel, the target interpolation filter corresponding to the sub - pixel is derived.

[0208] In this implementation, for each sub - pixel, the target interpolation filter corresponding to the sub - pixel can be implicitly derived based on the parity of the position of the sub - pixel. For example, the target interpolation filter corresponding to the sub - pixel can be implicitly indicated based on the parity of the last bit in the binary representation of the position of the sub - pixel.

[0209] In some embodiments, the two candidate interpolation filters corresponding to different motion vector precisions are the same, that is, the two candidate interpolation filters do not change with different motion vector precisions. In this way, for each sub-pixel to be interpolated, based on the parity of the position of the sub-pixel, one candidate interpolation filter is selected from these two candidate interpolation filters as the target interpolation filter for the sub-pixel. The number of taps of these two candidate interpolation filters is different.

[0210] In some embodiments, the two candidate interpolation filters corresponding to different motion vector precisions may be different.

[0211] In one example, when the motion vector precision is 1 / 4 pixel precision, the interpolation filters corresponding to the positions of different sub-pixels are shown in Table 9:

[0212] Table 9

[0213]

[0214]

[0215] In one example, when the motion vector precision is 1 / 2 pixel precision, the interpolation filters corresponding to the positions of different sub-pixels are shown in Table 10:

[0216] Table 10

[0217]

[0218] In one example, when the motion vector precision is 1 / 8 pixel precision, the interpolation filters corresponding to the positions of different sub-pixels are shown in Table 11:

[0219] Table 11

[0220]

[0221] In one example, when the motion vector precision is 1 / 16 pixel precision, the interpolation filters corresponding to the positions of different sub-pixels are shown in Table 12:

[0222] Table 12

[0223]

[0224]

[0225] In one example, when the motion vector precision is 1 / 32 pixel precision, the interpolation filters corresponding to the positions of different sub-pixels are shown in Table 13:

[0226] Table 13

[0227]

[0228] Tables 9 to 13 above show that when the last digit is even, an interpolation filter with a longer number of taps is used, and when the last digit is odd, an interpolation filter with a shorter number of taps is used. It can also be reversed, that is, when the last digit is odd, an interpolation filter with a longer number of taps is used, and when the last digit is even, an interpolation filter with a shorter number of taps is used.

[0229] Based on Tables 9 to 13, when the decoding end derives the target interpolation filter corresponding to the sub-pixel based on the parity of the position of the sub-pixel, it first obtains the first interpolation filter and the second interpolation filter corresponding to the motion vector accuracy of the area to be decoded, where the number of taps of the first interpolation filter and the second interpolation filter is different. For each sub-pixel, if the position of the sub-pixel is odd (for example, the last digit in the binary representation of the position of the sub-pixel is odd), the first interpolation filter is determined as the target interpolation filter corresponding to the sub-pixel. If the position of the sub-pixel is even (for example, the last digit in the binary representation of the position of the sub-pixel is even), the second interpolation filter is determined as the target interpolation filter corresponding to the sub-pixel. For example, the motion vector accuracy corresponding to the area to be decoded is 1 / 16 pixel accuracy. As shown in Table 11, the first interpolation filter corresponding to 1 / 16 pixel accuracy is an interpolation filter with 4 taps, and the second interpolation filter is an interpolation filter with 6 taps. For the sub-pixel at the 1 / 8 position, the corresponding interpolation filter is an interpolation filter with 4 taps, and then the interpolation filter with 4 taps is determined as the target interpolation filter corresponding to the sub-pixel at the 1 / 8 position.

[0230] The embodiments of the present application do not limit the specific number of taps of the first interpolation filter and the second interpolation filter.

[0231] In one example, the number of taps of the first interpolation filter is greater than the number of taps of the second interpolation filter.

[0232] In one example, the number of taps of the first interpolation filter is less than the number of taps of the second interpolation filter.

[0233] The above introduces the process of the decoding end implicitly deriving the target interpolation filter based on the motion vector accuracy corresponding to the area to be decoded and the position information of the sub-pixel.

[0234] In some embodiments, the implicit derivation of the target interpolation filter based on at least one of the motion vector accuracy, prediction mode, and decoding coefficient in S102-A2 above further includes the implementation manner of case 4, that is, implicitly deriving the target interpolation filter based on the motion vector accuracy and the decoding coefficient.

[0235] In one example of case 4, S102-A2 above includes the following steps of S102-A2-f:

[0236] S102-A2-f. If the motion vector accuracy corresponding to the region to be decoded is sub-pixel motion vector accuracy, an object interpolation filter is implicitly derived based on relevant quantities of the decoding coefficients.

[0237] In this implementation manner, an object interpolation filter can be implicitly derived based on relevant quantities of the decoding coefficients, and then the index of the interpolation filter is bound to a similar quantity of the decoding coefficients to reduce bit consumption and improve decoding efficiency.

[0238] Specifically, the decoding end first determines whether interpolation is required based on the motion vector accuracy corresponding to the region to be decoded. If the motion vector accuracy corresponding to the region to be decoded is sub-pixel motion vector accuracy, the decoding end implicitly derives an object interpolation filter based on relevant quantities of the decoding coefficients. If the motion vector accuracy corresponding to the region to be decoded is integer-pixel motion vector accuracy, the step of the object interpolation filter is skipped.

[0239] The embodiments of the present application do not limit the specific manner in which the decoding end implicitly derives an object interpolation filter based on relevant quantities of the decoding coefficients.

[0240] In a possible implementation manner, if the relevant quantity of the decoding coefficients is large, an interpolation filter with more taps is allocated to improve prediction accuracy. On the contrary, if the relevant quantity of the decoding coefficients is small, an interpolation filter with fewer taps is allocated to increase the preset speed.

[0241] In a possible implementation manner, the above S102-A2-f includes the following steps of S102-A2-f1:

[0242] S102-A2-f1. Implicitly derive an object interpolation filter based on the parity of relevant quantities of the decoding coefficients.

[0243] In this implementation manner, the index of the variable-length interpolation filter can be bound to the parity of the relevant quantity of the decoding coefficients in the region to be decoded, without additional bit consumption.

[0244] The embodiments of the present application do not limit the specific manifestation form of the relevant quantity of the decoding coefficients. In some examples, the relevant quantity of the decoding coefficients may include the number of non-zero coefficients in the decoding coefficients, the number of zero coefficients in the decoding coefficients, the number of even coefficients among the non-zero coefficients, and the number of odd coefficients among the non-zero coefficients. Based on this, the above S102-A2-f1 may include the following several examples:

[0245] Example 1. The decoding end implicitly derives an object interpolation filter based on the parity of the number of non-zero coefficients or zero coefficients in the decoding coefficients.

[0246] Exemplarily, the interpolation filters corresponding to the parity of the number of non-zero coefficients are shown in Table 14:

[0247] Table 14

[0248] Parity of the number of non - zero coefficients x - direction y - direction Even (odd) 12 - tap 12 - tap Odd (even) 8 - tap 8 - tap

[0249] As shown in Table 14, in the embodiments of the present application, the interpolation filter is bound to the parity of the number of non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the decoding coefficients. For example, when the number of non-zero coefficients is odd, a 12-tap interpolation filter is bound; when the number of non-zero coefficients is even, an 8-tap interpolation filter is bound. In this way, when the decoding end determines the target interpolation filter corresponding to the area to be decoded, it decodes the bitstream to obtain the decoding coefficients of the area to be decoded, and then determines the target interpolation filter corresponding to the area to be decoded based on the parity of the number of non-zero coefficients in the decoding coefficients. For example, if the number of non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the decoding coefficients of the area to be decoded is odd, the 12-tap interpolation filter is determined as the target interpolation filter corresponding to the area to be decoded. If the number of non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the decoding coefficients of the area to be decoded is even, the 8-tap interpolation filter is determined as the target interpolation filter corresponding to the area to be decoded. The above 12 taps and 8 taps are just examples and can be replaced according to actual needs.

[0250] Example 2: The decoding end implicitly derives the target interpolation filter based on the parity of the number of even coefficients in the non-zero coefficients.

[0251] Exemplarily, the interpolation filters corresponding to the parity of the number of even coefficients in the non-zero coefficients are shown in Table 15:

[0252] Table 15

[0253]

[0254] As shown in Table 15, in the embodiments of the present application, the interpolation filter is bound to the parity of the number of even coefficients among the non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the decoding coefficients. For example, when the number of even coefficients among the non-zero coefficients is odd, a 12-tap interpolation filter is bound; when the number of even coefficients among the non-zero coefficients is even, an 8-tap interpolation filter is bound. In this way, when the decoding end determines the target interpolation filter corresponding to the area to be decoded, it decodes the code stream to obtain the decoding coefficients of the area to be decoded, and then determines the target interpolation filter corresponding to the area to be decoded based on the parity of the number of even coefficients among the non-zero coefficients in the decoding coefficients. For example, if the number of even coefficients among the non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the decoding coefficients of the area to be decoded is odd, the 12-tap interpolation filter is determined as the target interpolation filter corresponding to the area to be decoded. If the number of even coefficients among the non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the decoding coefficients of the area to be decoded is even, the 8-tap interpolation filter is determined as the target interpolation filter corresponding to the area to be decoded. The above 12 taps and 8 taps are just examples and can be replaced according to actual needs.

[0255] Example 3: The decoding end implicitly derives the target interpolation filter based on the parity of the number of odd coefficients among the non-zero coefficients.

[0256] Exemplarily, the interpolation filters corresponding to the parity of the number of odd coefficients among the non-zero coefficients are shown in Table 16:

[0257] Table 16

[0258]

[0259] As shown in Table 16, in the embodiments of the present application, the interpolation filter is bound to the parity of the number of odd coefficients among the non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the decoding coefficients. For example, when the number of odd coefficients among the non-zero coefficients is odd, a 12-tap interpolation filter is bound; when the number of odd coefficients among the non-zero coefficients is even, an 8-tap interpolation filter is bound. In this way, when the decoding end determines the target interpolation filter corresponding to the area to be decoded, it decodes the code stream to obtain the decoding coefficients of the area to be decoded, and then determines the target interpolation filter corresponding to the area to be decoded based on the parity of the number of odd coefficients among the non-zero coefficients in the decoding coefficients. For example, if the number of odd coefficients among the non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the decoding coefficients of the area to be decoded is odd, the 12-tap interpolation filter is determined as the target interpolation filter corresponding to the area to be decoded. If the number of odd coefficients among the non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the decoding coefficients of the area to be decoded is even, the 8-tap interpolation filter is determined as the target interpolation filter corresponding to the area to be decoded. The above 12 taps and 8 taps are just examples and can be replaced according to actual needs.

[0260] The above text introduces the process by which the decoding end implicitly derives the target interpolation filter based on the motion vector accuracy and decoding coefficients corresponding to the area to be decoded.

[0261] In some embodiments, the decoding end can also implicitly derive the target interpolation filter corresponding to the area to be decoded in other ways.

[0262] After the decoding end implicitly derives the target interpolation filter corresponding to the area to be decoded based on the above steps, it executes the following step S103.

[0263] S103: Determine the reference area of the area to be decoded, and based on the target interpolation filter, perform interpolation on the reference area to determine the predicted value of the area to be decoded.

[0264] After the decoding end determines the target interpolation filter corresponding to the area to be decoded based on the above steps, it uses the target interpolation filter to perform interpolation on the reference area of the area to be decoded to determine the predicted value of the area to be decoded.

[0265] The embodiments of the present application do not limit the order of the decoding end to determine the reference area of the area to be decoded and the target interpolation filter corresponding to the area to be decoded. That is to say, the decoding end can first determine the target interpolation filter corresponding to the area to be decoded, and then determine the reference area of the area to be decoded. Or, the decoding end can first determine the reference area of the area to be decoded, and then determine the target interpolation filter corresponding to the area to be decoded. Or, the decoding end can determine the target interpolation filter corresponding to the area to be decoded while determining the reference area of the area to be decoded.

[0266] In the embodiments of the present application, the reference region of the region to be decoded may be a reference block or a reference frame, and the embodiments of the present application do not limit this. Among them, the specific manner for the decoding end to determine the reference region of the region to be decoded may refer to the description of related technologies and will not be elaborated herein.

[0267] Next, the decoding end interpolates the reference region of the region to be decoded based on the determined target interpolation filter, and then determines the predicted value of the region to be decoded.

[0268] In one example, when the above interpolation target is to interpolate the pixel predicted value, the decoding end uses the target interpolation filter for interpolation to obtain the interpolated reference region, and then based on the motion vector corresponding to the region to be decoded, obtains the predicted block of the region to be decoded in the interpolated reference region.

[0269] In one example, if the prediction mode corresponding to the region to be decoded in the embodiments of the present application is BIO, the decoding end first uses the bidirectional prediction mode to determine a forward reference block and a backward reference block of the region to be decoded. Then, the target interpolation filter is sampled to interpolate the forward reference block and the backward reference block, determines the gradients of the forward reference block and the backward reference block, and then corrects the motion vectors corresponding to the forward reference block and the backward reference block based on the gradients to obtain the predicted value of the region to be decoded.

[0270] In some embodiments, the bit widths of the filter coefficients of the interpolation filters with different tap numbers in the embodiments of the present application may be different. And / or the bit widths of the filter coefficients of the interpolation filters with the same tap number may also be different.

[0271] In some embodiments, the embodiments of the present application also provide the coefficients of some interpolation filters with different tap numbers. Specifically as follows, where MV position can be understood as the position of a pixel point, and coefficients are the filter coefficients. In some embodiments, the interpolation filter for determining the gradient is also referred to as a gradient filter or a gradient interpolation filter.

[0272] Table 17: Coefficients of the 8-tap gradient filter used in the BIO mode

[0273] MV position coefficients 0 –4,11,–39,–1,41,–14,8,–2 1 / 4 –2,6,–19,–31,53,–12,7,–2 1 / 2 0,–1,0,–50,50,0,1,0 3 / 4 2,–7,12,–53,31,19,–6,2

[0274] Table 18: Coefficients of the 12-tap filter used in the BIO and non-BIO modes other than the AFFINE mode

[0275] MV position coefficients 0 0,0,0,0,0,256,0,0,0,0,0,0 1 / 4 -2,5,-11,21,-43,230,75,-29,15,-8,4,-1 1 / 2 -2,6,-13,25,-50,162,162,-50,25,-13,6,-2 3 / 4 -1,4,-8,15,-29,75,230,-43,21,-11,5,-2

[0276] Table 19: Coefficients of the 12-tap filter used in the AFFINE mode

[0277]

[0278]

[0279] Table 20: 6-tap filter coefficients used in BIO and non-BIO modes other than the AFFINE mode

[0280] MV position coefficients 0 0,0,256,0,0,0 1 / 8 4,-21,248,33,-10,2 2 / 8 8,-35,227,73,-22,5 3 / 8 10,-42,196,117,-34,9 4 / 8 10,-40,158,158,-40,10 5 / 8 9,-34,117,196,-42,10 6 / 8 5,-22,73,227,-35,8 7 / 8 2,-10,33,248,-21,4

[0281] Table 21: 6-tap filter coefficients used in the AFFINE mode

[0282]

[0283]

[0284] The video decoding method provided by the embodiments of the present application implicitly indicates a target interpolation filter. Specifically, the decoding end decodes the code stream to obtain at least one of the prediction information and decoding coefficients of the area to be decoded, and then implicitly derives the target interpolation filter corresponding to the area to be decoded based on at least one of the prediction information and decoding coefficients, and then uses the target interpolation filter to interpolate the reference area of the area to be decoded to determine the predicted value of the area to be decoded. That is to say, the embodiments of the present application transmit the index of the target interpolation filter used through various implicit indication methods, saving bits and improving the performance of video decoding.

[0285] The above introduces the video decoding method of the present application by taking the decoding end as an example, and the following will be described by taking the encoding end as an example.

[0286] Figure 6 It is a schematic flowchart of the video encoding method provided by an embodiment of the present application. The embodiments of the present application are applied to Figure 1 and Figure 2 the video encoder shown. As Figure 6 shown, the method of the embodiments of the present application includes:

[0287] S201. Determine the prediction information of the area to be encoded.

[0288] The embodiments of the present application do not limit the specific size and shape of the area to be encoded.

[0289] In some embodiments, the area to be encoded may be one or several image blocks to be encoded in the current frame to be encoded. For example, it may be one or several CTUs, or one or several CUs, or one or several PUs, etc. In one example, the image block to be encoded may be referred to as the current block, the current image block to be encoded, etc.

[0290] In some embodiments, the area to be encoded may be the current frame to be encoded, that is, the area to be encoded is an entire frame of image.

[0291] During video encoding, when the encoding end encodes the area to be encoded, it determines the prediction mode of the area to be encoded. Based on the prediction mode, it determines the predicted value of the area to be encoded, and then based on the predicted value of the area to be encoded, it determines the residual value of the area to be encoded. Then, the residual value is transformed to obtain transform coefficients, the transform coefficients are quantized to obtain quantization coefficients, and finally the quantization coefficients are encoded to obtain a bitstream.

[0292] The embodiments of the present application mainly relate to the prediction process of the area to be encoded.

[0293] In the embodiments of the present application, in order to improve the prediction accuracy of the area to be encoded, an interpolation filter is used to interpolate the reference area of the area to be encoded. However, currently, the encoding end explicitly indicates the determined target interpolation filter of the area to be encoded to the decoding end. For example, the index of the target interpolation filter is written in the bitstream, which will waste bits and thus reduce the performance of video encoding and decoding.

[0294] To solve the above technical problems, in the embodiments of the present application, the target interpolation filter is indicated in an implicit manner. Specifically, the encoding end does not need to write the index of the determined target interpolation filter into the bitstream, and the decoding end can implicitly derive the target interpolation filter corresponding to the area to be encoded based on at least one of the prediction information of the area to be encoded and the decoding coefficients, thereby saving bits and improving the performance of video encoding and decoding.

[0295] The embodiments of the present application do not limit the specific content of the prediction information of the area to be encoded, which can be understood as all information required for predicting the predicted value of the area to be encoded, such as at least one of the prediction mode, component information, reference frame list, motion vector accuracy, etc.

[0296] The embodiments of the present application do not limit the specific manner in which the encoding end determines the prediction mode of the area to be encoded. For example, the encoding end determines the prediction mode of the module to be encoded from multiple candidate prediction modes based on the rate-distortion cost.

[0297] In some embodiments, the prediction information of the area to be encoded includes at least one of the prediction mode of the area to be encoded and the motion vector accuracy corresponding to the area to be encoded.

[0298] After the encoding end obtains at least one of the prediction information of the area to be encoded and the encoding coefficients, it executes the following step S202.

[0299] S202: Based on the prediction information, determine the target interpolation filter corresponding to the area to be encoded.

[0300] In the embodiments of the present application, based on prediction information, the encoding end implicitly derives a target interpolation filter corresponding to the region to be encoded from interpolation filters with different numbers of taps, without indicating the target interpolation filter to the decoding end, thereby saving bits and improving the performance of video coding.

[0301] The embodiments of the present application do not limit the specific manner in which the encoding end determines the target interpolation filter corresponding to the region to be encoded based on prediction information.

[0302] In some embodiments, the prediction information includes at least one of an index of the motion vector accuracy corresponding to the region to be encoded and the prediction mode. In this case, step S202 above includes the following steps of S202-A:

[0303] S202-A. Implicitly derive the target interpolation filter based on at least one of the motion vector accuracy and the prediction mode.

[0304] In this implementation manner, when the prediction information of the region to be encoded includes at least one of the motion vector accuracy corresponding to the region to be encoded and the prediction mode, the encoding end implicitly derives the target interpolation filter based on at least one of the motion vector accuracy and the prediction mode. That is to say, in the embodiments of the present application, the index of the interpolation filter is bound to at least one of the motion vector accuracy and the prediction mode, and the encoding end does not need to separately indicate the target interpolation filter, thereby saving bits and improving the video coding performance.

[0305] The embodiments of the present application do not limit the specific manner in which the encoding end implicitly derives the target interpolation filter based on at least one of the motion vector accuracy and the prediction mode.

[0306] Case 1. The encoding end implicitly derives the target interpolation filter based on any one of the motion vector accuracy and the prediction mode. For example, it is stipulated at both the encoding and decoding ends that interpolation filter 1 is used in prediction mode 1 and interpolation filter 2 is used in prediction mode 2. Another example is that for high-precision motion vectors, an interpolation filter with more taps is used, and for low-precision motion vectors, an interpolation filter with fewer taps is used.

[0307] Case 2. The above S202-A includes the following steps of S202-A-a:

[0308] S202-A-a. Implicitly derive the target interpolation filter based on the motion vector accuracy and the prediction mode.

[0309] In this implementation manner, the encoding end can implicitly derive a target interpolation filter based on the motion vector precision and prediction mode corresponding to the area to be encoded. For example, the candidate motion vector precisions corresponding to different prediction modes may be different. Thus, the target interpolation filter can be implicitly derived based on the candidate motion vector precision corresponding to the prediction mode of the area to be encoded and the motion vector precision corresponding to the area to be encoded. That is to say, the embodiments of the present application can reduce bit consumption and improve encoding efficiency by binding the index of the motion vector and the index of the variable-length interpolation filter.

[0310] In the embodiments of the present application, generally, interpolation is performed when the motion vector precision is sub-pixel, and no interpolation is performed when the motion vector precision is integer pixel. Therefore, before the encoding end implicitly derives the target interpolation filter based on the motion vector precision and prediction mode, it is first necessary to determine whether the motion vector precision corresponding to the area to be encoded is sub-pixel. If it is sub-pixel, the target interpolation filter is implicitly derived based on the motion vector precision and prediction mode corresponding to the area to be encoded. If the motion vector precision corresponding to the area to be encoded is integer pixel, the step of implicitly deriving the target interpolation filter is skipped. Based on this, the specific implementation manner of the above S202-A-a at least includes the following examples:

[0311] In a possible implementation manner of S202-A-a, if the motion vector precision is sub-pixel motion vector precision and the prediction mode is non-affine motion compensation mode, the target interpolation filter is derived based on the motion vector precision from a first relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the non-affine motion compensation mode.

[0312] The embodiments of the present application do not limit the specific quantity and type of the sub-pixel motion vector precisions included in the first relationship list.

[0313] In a possible implementation manner, the sub-pixel motion vector precisions among the n motion vector precisions corresponding to Adaptive Motion Vector Resolution (AMVR) are included in the first relationship list.

[0314] Exemplarily, the n motion vector precisions may include 5 motion vector precisions: 1 / 4, 1 / 2, 1, 2, and 4. Of course, other motion vector precisions may also be included, and the embodiments of the present application do not limit this.

[0315] Exemplarily, a first relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the non-affine motion compensation mode is shown in Table 3.

[0316] As shown in Table 3, in the embodiments of the present application, the index of the variable-length interpolation filter is bound to the index of the motion vector precision and the prediction mode, so that the encoding end can implicitly derive the target interpolation filter corresponding to the area to be encoded through the index of the motion vector precision and the prediction mode corresponding to the area to be encoded.

[0317] In one example, in the non-affine motion compensation mode, the encoding end uses 3 bits to encode the index of the motion vector precision. In this way, the decoding end decodes the bitstream to obtain the 3-bit index of the motion vector precision, and then based on the 3-bit index of the motion vector precision, the target interpolation filter is derived in Table 4 above.

[0318] It should be noted that Table 4 can be understood as another form of Table 3, and the above Table 3 and Table 4 are only examples of the first relationship list. The first relationship list in the embodiments of the present application includes but is not limited to those shown in Table 3 and Table 4 above.

[0319] In another possible implementation manner of S202-A-a, if the motion vector precision is sub-pixel motion vector precision and the prediction mode is affine motion compensation mode, then based on the motion vector precision, in the second relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the affine motion compensation mode, the target interpolation filter is derived.

[0320] The embodiments of the present application do not limit the specific quantity and type of the sub-pixel motion vector precision included in the second relationship list.

[0321] In one possible implementation manner, the second relationship list includes the sub-pixel motion vector precision among the m motion vector precisions corresponding to the affine motion compensation mode.

[0322] Exemplarily, the m motion vector precisions may include 3 kinds of motion vector precisions: 1 / 16, 1 / 4, and 1. Of course, other motion vector precisions may also be included, and the embodiments of the present application do not limit this.

[0323] Exemplarily, the second relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the affine motion compensation mode is shown in Table 5.

[0324] As shown in Table 5, in the embodiments of the present application, the index of the variable-length interpolation filter is bound to the motion vector precision and the affine motion compensation mode, so that the encoding end can implicitly derive the target interpolation filter corresponding to the area to be encoded through the motion vector precision and the prediction mode corresponding to the area to be encoded.

[0325] In one example, in the affine motion compensation mode, the encoding end uses 2 bits to encode the index of the motion vector precision. In this way, the encoding end decodes the bitstream to obtain the index of the motion vector precision of 2 bits, and then based on the index of the motion vector precision of 2 bits, in Table 6 above, derives the target interpolation filter.

[0326] It should be noted that Table 6 can be understood as another form of Table 5, and the above Table 5 and the above Table 6 are only examples of the second relationship list. The second relationship list of the embodiments of the present application includes but is not limited to those shown in the above Table 5 and Table 6.

[0327] In some embodiments, in the above first relationship list and the second relationship list, the number of taps of the interpolation filter corresponding to the sub-pixel motion vector precision with higher precision is greater than the number of taps of the interpolation filter corresponding to the sub-pixel motion vector precision with lower precision. As shown in Table 3 for example, the 1 / 4 motion vector precision corresponds to an interpolation filter with 12 taps, while the 1 / 2 motion vector precision corresponds to an interpolation filter with 8 taps. Another example is as shown in Table 5, the 1 / 16 motion vector precision corresponds to an interpolation filter with 12 taps, while the 1 / 4 motion vector precision corresponds to an interpolation filter with 8 taps. In this way, using an interpolation filter with more taps for the sub-pixel motion vector precision with higher precision for interpolation can improve the accuracy of motion estimation, and thus improve the prediction effect.

[0328] In another example of Case 2, the above S202-A includes the following steps of S202-A-b and S202-A-c:

[0329] S202-A-b, if the prediction mode is motion vector prediction based on historical information, then determine the MVP corresponding to the region to be encoded from the MVP candidate list;

[0330] S202-A-c, derive the target interpolation filter based on the index of the MVP and the motion vector precision.

[0331] In this embodiment, when the encoding end determines the target interpolation filter corresponding to the region to be encoded based on the above prediction mode and the motion vector precision, it first judges the prediction mode of the region to be encoded. If the prediction mode of the region to be encoded is the motion vector prediction mode based on historical information, then the encoding end determines the MVP corresponding to the region to be encoded from multiple candidate MVPs (for example, 8 MVPs) included in the MVP candidate list. For example, based on the rate-distortion cost, determine the MVP corresponding to the region to be encoded from multiple candidate MVPs included in the MVP candidate list. Then, based on the index of the MVP and the motion vector precision, derive the target interpolation filter. That is to say, in this implementation manner, by binding the index of the MVP in the HMVP and the index of the non-fixed-length interpolation filter, the bit consumption can be reduced and the encoding efficiency can be improved.

[0332] In the embodiments of the present application, generally, when the motion vector accuracy is sub-pixel, interpolation is performed; when the motion vector accuracy is integer pixel, interpolation is not performed. Therefore, before implicitly deriving the target interpolation filter based on the motion vector accuracy and the prediction mode at the encoding end, it is first necessary to determine whether the motion vector accuracy corresponding to the region to be encoded is sub-pixel. If it is sub-pixel, the target interpolation filter is implicitly derived based on the motion vector accuracy and the prediction mode corresponding to the region to be encoded. If the motion vector accuracy corresponding to the region to be encoded is integer pixel, the step of implicitly deriving the target interpolation filter is skipped. Based on this, the above S202-A-c includes the step of S202-A-c1:

[0333] S202-A-c1: When the motion vector accuracy is sub-pixel motion vector accuracy, the target interpolation filter is derived based on the index of the MVP.

[0334] The embodiments of the present application do not limit the specific manner of deriving the target interpolation filter based on the index of the MVP.

[0335] In some embodiments, as can be seen from the above, the HMVP technology copies k (for example, 8) pieces of motion information from the previously encoded blocks as candidates. Both the encoding and decoding ends can default that these k candidate MVPs correspond to interpolation filters with different tap numbers. In this way, the encoding end can derive the target interpolation filter corresponding to the region to be encoded from the interpolation filters with different tap numbers corresponding to these k candidate MVPs based on the index of the MVP corresponding to the region to be encoded.

[0336] In one example, assuming k is equal to 8, the correspondence between the indices of the 8 candidate MVPs and the interpolation filters with different tap numbers is shown in Table 7.

[0337] As shown in Table 7, in the motion vector prediction mode based on historical information, the index of the MVP in the HMVP and the index of the non-fixed-length interpolation filter are bound, so that the encoding end can implicitly derive the target interpolation filter corresponding to the region to be encoded through the index of the MVP corresponding to the region to be encoded.

[0338] In one example, in the HMVP mode, the encoding end uses 3 bits to encode the index of the MVP. In this way, the decoding end decodes the bitstream to obtain the 3-bit index of the MVP corresponding to the region to be decoded, and then derives the target interpolation filter in the above Table 8 based on the 3-bit index of the MVP.

[0339] It should be noted that Table 8 can be understood as another form of Table 7, and the above Table 7 and the above Table 8 are only a corresponding relationship between the index of MVP and the interpolation filters with different tap numbers. The corresponding relationship between the index of MVP and the interpolation filters with different tap numbers in the embodiments of the present application includes, but is not limited to, those shown in the above Table 8 and Table 7.

[0340] As can be seen from the above, in the HMVP mode, no additional flag bits need to be added at the encoding end to indicate the target interpolation filter. The decoding end only needs to decode the index of the MVP corresponding to the area to be encoded, and then implicitly derive the target interpolation filter based on the index of the MVP.

[0341] In some embodiments, the number of taps of the interpolation filter corresponding to the index of the MVP closer to the area to be encoded is greater than the number of taps of the interpolation filter corresponding to the index of the MVP farther from the area to be encoded. The HMVP candidate list is composed of the MVs of several previously encoded blocks. The closer to the area to be encoded, the larger the index assigned to the MV in the candidate list. Assuming that the closer the MVP is to the area to be encoded, the greater the motion correlation with the area to be encoded, more taps of the interpolation filter can be considered for the closer MVP to interpolate a more accurate prediction result. As shown in Table 7, for example, the encoded block corresponding to the MVP with index 7 is the closest to the area to be encoded. Therefore, a 12-tap interpolation filter is configured for the MVP with index 7, and a 6-tap interpolation filter is configured for the MVP with index 3.

[0342] The above introduces the process of implicitly deriving the target interpolation filter corresponding to the area to be encoded at the encoding end based on the prediction mode of the area to be encoded and the accuracy of the motion vector corresponding to the area to be encoded.

[0343] In some embodiments, the implicit derivation of the target interpolation filter based on at least one of the motion vector accuracy and the prediction mode in S202-A above further includes the implementation manner of Case 3 as follows:

[0344] Case 3, S202-A includes the following steps of S202-A-d and S202-A-e:

[0345] S202-A-d: If the motion vector accuracy is sub-pixel motion vector accuracy, then the reference area is sub-pixel divided based on the motion vector accuracy;

[0346] S202-A-e: For each sub-pixel, the target interpolation filter corresponding to the sub-pixel is derived based on the position information of the sub-pixel.

[0347] Although traditional video coding uses different pixel precisions for motion estimation and motion compensation to obtain accurate prediction results, at the same pixel precision, the importance of sub-pixels at different positions is not the same. Therefore, it is possible to implicitly derive the target interpolation filter based on the position information of the sub-pixels.

[0348] Specifically, the encoding end first determines whether interpolation is required based on the motion vector precision corresponding to the area to be encoded. If the motion vector precision corresponding to the area to be encoded is sub-pixel motion vector precision, the encoding end performs sub-pixel division on the reference area corresponding to the area to be encoded based on this sub-pixel motion vector precision. The specific division method is as shown above Figure 4B and will not be elaborated here.

[0349] For each sub-pixel obtained by the division, based on the position information of this sub-pixel, the target interpolation filter corresponding to this sub-pixel is derived.

[0350] The embodiments of the present application do not limit the specific method for the encoding end to derive the target interpolation filter corresponding to the sub-pixel based on the position information of the sub-pixel.

[0351] In a possible implementation manner, if the sub-pixel is a key pixel point of the area to be encoded, such as a point in the central area, an interpolation filter with more taps is assigned to this sub-pixel to improve the prediction accuracy. On the contrary, if the sub-pixel is not a key point of the area to be encoded, an interpolation filter with fewer taps can be assigned to this sub-pixel to improve the preset speed.

[0352] In a possible implementation manner, the above S202-A-e includes the following steps of S202-A-e1:

[0353] S202-A-e1: Derive the target interpolation filter corresponding to the sub-pixel based on the parity of the position of the sub-pixel.

[0354] In this implementation manner, for each sub-pixel, the target interpolation filter corresponding to this sub-pixel can be implicitly derived based on the parity of the position of the sub-pixel. For example, the target interpolation filter corresponding to the sub-pixel can be implicitly indicated based on the parity of the last bit in the binary representation of the position of the sub-pixel.

[0355] In some embodiments, the two candidate interpolation filters corresponding to different motion vector precisions are the same, that is, the two candidate interpolation filters do not change with different motion vector precisions. In this way, for each sub-pixel to be interpolated, based on the parity of the position of the sub-pixel, one candidate interpolation filter is selected from these two candidate interpolation filters as the target interpolation filter for this sub-pixel. The number of taps of these two candidate interpolation filters is different.

[0356] In some embodiments, the two candidate interpolation filters corresponding to different motion vector precisions may be different.

[0357] In one example, when the motion vector precision is 1 / 4 pixel precision, the interpolation filters corresponding to different sub-pixel positions are shown in Table 9.

[0358] In one example, when the motion vector precision is 1 / 2 pixel precision, the interpolation filters corresponding to different sub-pixel positions are shown in Table 10.

[0359] In one example, when the motion vector precision is 1 / 8 pixel precision, the interpolation filters corresponding to different sub-pixel positions are shown in Table 11.

[0360] In one example, when the motion vector precision is 1 / 16 pixel precision, the interpolation filters corresponding to different sub-pixel positions are shown in Table 12.

[0361] In one example, when the motion vector precision is 1 / 32 pixel precision, the interpolation filters corresponding to different sub-pixel positions are shown in Table 13.

[0362] The above Tables 9 to 13 show that when the last digit is even, an interpolation filter with a longer tap number is used, and when the last digit is odd, an interpolation filter with a shorter tap number is used. It can also be reversed, that is, when the last digit is odd, an interpolation filter with a longer tap number is used, and when the last digit is even, an interpolation filter with a shorter tap number is used.

[0363] Based on Tables 9 to 13, when the encoding end derives the target interpolation filter corresponding to the sub-pixel based on the parity of the sub-pixel position, it first obtains the first interpolation filter and the second interpolation filter corresponding to the motion vector precision of the area to be encoded, where the tap numbers of the first interpolation filter and the second interpolation filter are different. For each sub-pixel, if the position of the sub-pixel is odd (for example, the last digit in the binary representation of the position of the sub-pixel is odd), the first interpolation filter is determined as the target interpolation filter corresponding to the sub-pixel. If the position of the sub-pixel is even (for example, the last digit in the binary representation of the position of the sub-pixel is even), the second interpolation filter is determined as the target interpolation filter corresponding to the sub-pixel. For example, the motion vector precision of the area to be encoded is 1 / 16 pixel precision. As shown in Table 11, the first interpolation filter corresponding to 1 / 16 pixel precision is a 4-tap interpolation filter, and the second interpolation filter is a 6-tap interpolation filter. For the sub-pixel at the 1 / 8 position, the corresponding interpolation filter is a 4-tap interpolation filter, and then the 4-tap interpolation filter is determined as the target interpolation filter corresponding to the sub-pixel at the 1 / 8 position.

[0364] The embodiments of the present application do not limit the specific number of taps of the first interpolation filter and the second interpolation filter.

[0365] In one example, the number of taps of the first interpolation filter is greater than the number of taps of the second interpolation filter.

[0366] In one example, the number of taps of the first interpolation filter is less than the number of taps of the second interpolation filter.

[0367] The above text introduces the process of implicitly deriving the target interpolation filter at the encoding end based on the motion vector accuracy and sub-pixel position information corresponding to the region to be encoded.

[0368] In some embodiments, if the prediction information includes the motion vector accuracy corresponding to the region to be encoded, then determining the target interpolation filter corresponding to the region to be encoded based on the prediction information in S202 above includes the following steps S202-B to S202-D:

[0369] S202-B: If the motion vector accuracy is sub-pixel motion vector accuracy, obtain the third interpolation filter and the fourth interpolation filter corresponding to the motion vector accuracy, and the number of taps of the third interpolation filter and the fourth interpolation filter is different;

[0370] S202-C: Determine the rate-distortion cost corresponding to the third interpolation filter and the fourth interpolation filter respectively;

[0371] S202-D: Based on the rate-distortion cost, select the target interpolation filter from the third interpolation filter and the fourth interpolation filter.

[0372] In this implementation, the encoding end first determines whether interpolation is required based on the motion vector accuracy corresponding to the region to be encoded. If the motion vector accuracy corresponding to the region to be encoded is sub-pixel motion vector accuracy, the encoding end obtains the third interpolation filter and the fourth interpolation filter corresponding to the motion vector accuracy, and the number of taps of the third interpolation filter and the fourth interpolation filter is different.

[0373] In some embodiments, the two candidate interpolation filters corresponding to different motion vector accuracies may be the same. For example, the two candidate interpolation filters corresponding to motion vector accuracy 1 are the same as the two candidate interpolation filters corresponding to motion vector accuracy 2.

[0374] In some embodiments, the two candidate interpolation filters corresponding to different motion vector accuracies are not the same. For example, the two candidate interpolation filters corresponding to motion vector accuracy 1 are not completely the same as the two candidate interpolation filters corresponding to motion vector accuracy 2.

[0375] The encoding end determines the rate-distortion costs corresponding to the third interpolation filter and the fourth interpolation filter based on the motion vector accuracy obtained. For example, the third interpolation filter is used to interpolate the reference region of the region to be encoded, and the rate-distortion cost corresponding to the third interpolation filter is determined. At the same time, the fourth interpolation filter is used to interpolate the reference region of the region to be encoded, and the rate-distortion cost corresponding to the fourth interpolation filter is determined. Then, based on the rate-distortion costs, the target interpolation filter is selected from the third interpolation filter and the fourth interpolation filter. For example, the interpolation filter with the smallest rate-distortion cost is selected from the third interpolation filter and the fourth interpolation filter and determined as the target interpolation filter.

[0376] In some embodiments, when the target interpolation filter is indicated by the relevant quantity of the coding coefficients of the region to be encoded, the embodiments of the present application further include the following steps 1 to 3:

[0377] Step 1: Determine the coding coefficients of the region to be encoded based on the predicted value of the region to be encoded;

[0378] Step 2: Determine the preset parity corresponding to the target interpolation filter;

[0379] Step 3: Determine whether to adjust the coding coefficients based on the preset parity.

[0380] In this implementation manner, after the encoding end selects the target interpolation filter from the third interpolation filter and the fourth interpolation filter, it interpolates the reference region of the region to be encoded to determine the predicted value of the region to be encoded, and then determines the coding coefficients of the region to be encoded based on the predicted value of the region to be encoded.

[0381] In the embodiments of the present application, the selected interpolation filter can be indicated by the parity of the relevant quantity of the coding coefficients. Based on this, after the encoding end determines the target interpolation filter and the coding coefficients of the region to be encoded, it can determine whether to adjust the coding coefficients of the region to be encoded based on the preset parity of the relevant quantity of the coding coefficients corresponding to the target interpolation filter.

[0382] For example, if the preset parity corresponding to the target interpolation filter is the same as the parity of the relevant quantity of the coding coefficients, the adjustment of the coding coefficients is skipped.

[0383] For another example, if the preset parity corresponding to the target interpolation filter is different from the parity of the relevant quantity of the coding coefficients, the coding coefficients are adjusted so that the parity of the relevant quantity of the coding coefficients is the same as the preset parity corresponding to the target interpolation filter.

[0384] The embodiments of the present application do not limit the specific type of the relevant quantity of the coding coefficients.

[0385] In some embodiments, the parity of the relevant quantity of coding coefficients includes: the parity of the number of non-zero coefficients or zero coefficients in the coding coefficients, or the parity of the number of even coefficients among the non-zero coefficients, or the parity of the number of odd coefficients among the non-zero coefficients.

[0386] In one example, as shown in Table 14, in the embodiments of the present application, the interpolation filter is bound to the parity of the number of non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the coding coefficients. For example, when the number of non-zero coefficients is odd, a 12-tap interpolation filter is bound; when the number of non-zero coefficients is even, an 8-tap interpolation filter is bound. In this way, when the coding end determines the target interpolation filter corresponding to the area to be coded, it looks up the preset parity of the number of non-zero coefficients corresponding to the target interpolation filter in Table 14, and then based on this preset parity, determines whether to adjust the coding coefficients of the area to be coded. For example, if the target interpolation filter corresponding to the area to be coded is a 12-tap interpolation filter, the number of non-zero coefficients in the coding coefficients corresponding to this 12-tap interpolation filter is odd. If the number of non-zero coefficients in the coding coefficients of the area to be coded obtained by the above coding is odd, then the coding coefficients of the area to be coded are not adjusted. If the number of non-zero coefficients in the coding coefficients of the area to be coded obtained by the above coding is even, then the coding coefficients of the area to be coded are adjusted so that the number of non-zero coefficients in the coding coefficients of the area to be coded is odd.

[0387] In one example, as shown in Table 15, in the embodiments of the present application, the interpolation filter is bound to the parity of the number of even coefficients among the non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the coding coefficients. For example, when the number of even coefficients among the non-zero coefficients is odd, a 12-tap interpolation filter is bound; when the number of even coefficients among the non-zero coefficients is even, an 8-tap interpolation filter is bound. In this way, when the coding end determines the target interpolation filter corresponding to the area to be coded, it looks up the preset parity of the number of non-zero coefficients corresponding to the target interpolation filter in Table 15, and then based on this preset parity, determines whether to adjust the coding coefficients of the area to be coded. For example, if the target interpolation filter corresponding to the area to be coded is a 12-tap interpolation filter, the number of even coefficients among the non-zero coefficients in the coding coefficients corresponding to this 12-tap interpolation filter is odd. If the number of even coefficients among the non-zero coefficients in the coding coefficients of the area to be coded obtained by the above coding is odd, then the coding coefficients of the area to be coded are not adjusted. If the number of even coefficients among the non-zero coefficients in the coding coefficients of the area to be coded obtained by the above coding is even, then the coding coefficients of the area to be coded are adjusted so that the number of even coefficients among the non-zero coefficients in the coding coefficients of the area to be coded is odd.

[0388] In one example, as shown in Table 16, in the embodiments of the present application, the interpolation filter is bound to the parity of the number of odd coefficients among the non-zero coefficients (such as non-zero quantization coefficients or non-zero transform coefficients) in the coding coefficients. For example, when the number of odd coefficients among the non-zero coefficients is odd, a 12-tap interpolation filter is bound; when the number of odd coefficients among the non-zero coefficients is even, an 8-tap interpolation filter is bound. In this way, when the coding end determines the target interpolation filter corresponding to the area to be coded, it looks up the preset parity of the number of non-zero coefficients corresponding to the target interpolation filter in Table 16, and then determines whether to adjust the coding coefficients of the area to be coded based on this preset parity. For example, if the target interpolation filter corresponding to the area to be coded is a 12-tap interpolation filter, the number of odd coefficients among the non-zero coefficients in the coding coefficients corresponding to this 12-tap interpolation filter is odd. If the number of odd coefficients among the non-zero coefficients in the coding coefficients of the area to be coded obtained by the above coding is odd, then the coding coefficients of the area to be coded are not adjusted. If the number of odd coefficients among the non-zero coefficients in the coding coefficients of the area to be coded obtained by the above coding is even, then the coding coefficients of the area to be coded are adjusted so that the number of odd coefficients among the non-zero coefficients in the coding coefficients of the area to be coded is odd.

[0389] In some embodiments, the coding end can also derive the target interpolation filter corresponding to the area to be coded in other ways.

[0390] After the coding end determines the target interpolation filter corresponding to the area to be coded based on the above steps, it executes the following step S203.

[0391] S203: Determine the reference area of the area to be coded, and based on the target interpolation filter, interpolate the reference area to determine the predicted value of the area to be coded.

[0392] After the coding end determines the target interpolation filter corresponding to the area to be coded based on the above steps, it uses the target interpolation filter to interpolate the reference area of the area to be coded to determine the predicted value of the area to be coded.

[0393] The embodiments of the present application do not limit the order of the coding end determining the reference area of the area to be coded and determining the target interpolation filter corresponding to the area to be coded. That is to say, the coding end can first determine the target interpolation filter corresponding to the area to be coded, and then determine the reference area of the area to be coded. Or, the coding end can first determine the reference area of the area to be coded, and then determine the target interpolation filter corresponding to the area to be coded. Or, the coding end can determine the target interpolation filter corresponding to the area to be coded while determining the reference area of the area to be coded.

[0394] In the embodiments of the present application, the reference region of the region to be encoded may be a reference block or a reference frame, and the embodiments of the present application do not limit this. Among them, the specific manner for the encoding end to determine the reference region of the region to be encoded may refer to the description of related technologies and will not be elaborated here.

[0395] Next, the encoding end interpolates the reference region of the region to be encoded based on the determined target interpolation filter, and then determines the predicted value of the region to be encoded.

[0396] In one example, when the above interpolation target is to interpolate the pixel predicted value, the encoding end uses the target interpolation filter for interpolation to obtain the interpolated reference region, and then based on the motion vector corresponding to the region to be encoded, obtains the prediction block of the region to be encoded in the interpolated reference region.

[0397] In one example, if the prediction mode corresponding to the region to be encoded in the embodiments of the present application is BIO, the encoding end first uses the bidirectional prediction mode to determine a forward reference block and a backward reference block of the region to be encoded. Then, the target interpolation filter is sampled to interpolate the forward reference block and the backward reference block, determine the gradients of the forward reference block and the backward reference block, and then correct the motion vectors corresponding to the forward reference block and the backward reference block based on the gradients to obtain the predicted value of the region to be encoded.

[0398] The video encoding method provided by the embodiments of the present application indicates the target interpolation filter in an implicit manner. Specifically, the encoding end determines the prediction information of the region to be encoded, and then based on the prediction information, determines the target interpolation filter corresponding to the region to be encoded, and then uses the target interpolation filter to interpolate the reference region of the region to be encoded to determine the predicted value of the region to be encoded. That is to say, the embodiments of the present application transmit the index of the target interpolation filter used through various implicit indication methods, saving bits and improving the performance of video encoding.

[0399] As described above in conjunction with Figures 5 to 6 , the embodiments of the video encoding and decoding method of the present application are described in detail. Below in conjunction with Figures 7 to 8 , the device embodiments of the present application are described in detail.

[0400] Figure 7 FIG. is a schematic block diagram of a video decoding device provided by an embodiment of the present application. The device 10 may be applied to a decoding device.

[0401] As Figure 7 shown, the video decoding device 10 includes:

[0402] A decoding unit 11, configured to decode a bitstream to obtain at least one of prediction information and decoding coefficients of the region to be decoded;

[0403] An export unit 12 for implicitly exporting a target interpolation filter corresponding to the region to be decoded based on at least one of the prediction information and the decoding coefficients;

[0404] An interpolation unit 13 for determining a reference region of the region to be decoded and performing interpolation on the reference region based on the target interpolation filter to determine a predicted value of the region to be decoded.

[0405] In some embodiments, the prediction information includes at least one of an index of a motion vector precision corresponding to the region to be decoded and a prediction mode. The export unit 12 is specifically configured to determine the motion vector precision corresponding to the region to be decoded based on the index of the motion vector precision, and implicitly export the target interpolation filter based on at least one of the motion vector precision, the prediction mode, and the decoding coefficients.

[0406] In some embodiments, the export unit 12 is specifically configured to implicitly export the target interpolation filter based on the motion vector precision and the prediction mode.

[0407] In some embodiments, the export unit 12 is specifically configured to, when the motion vector precision is a sub-pixel motion vector precision and the prediction mode is a non-affine motion compensation mode, export the target interpolation filter based on the motion vector precision in a first relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the non-affine motion compensation mode; when the motion vector precision is a sub-pixel motion vector precision and the prediction mode is an affine motion compensation mode, export the target interpolation filter based on the motion vector precision in a second relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the affine motion compensation mode.

[0408] In some embodiments, the first relationship list includes sub-pixel motion vector precisions among n motion vector precisions corresponding to an adaptive motion vector precision, and the second relationship list includes sub-pixel motion vector precisions among m motion vector precisions corresponding to the affine motion compensation mode, where both n and m are positive integers.

[0409] In some embodiments, in the first relationship list and the second relationship list, the number of taps of the interpolation filter corresponding to a sub-pixel motion vector precision with a higher precision is greater than the number of taps of the interpolation filter corresponding to a sub-pixel motion vector precision with a lower precision.

[0410] In some embodiments, the derivation unit 12 is specifically configured to, when the prediction mode is motion vector prediction based on historical information, decode the code stream to obtain an index of a motion vector prediction (MVP) corresponding to the region to be decoded; and derive the target interpolation filter based on the index of the MVP and the motion vector precision.

[0411] In some embodiments, the derivation unit 12 is specifically configured to, when the motion vector precision is sub-pixel motion vector precision, derive the target interpolation filter based on the index of the MVP.

[0412] In some embodiments, the number of taps of the interpolation filter corresponding to the index of the MVP closer to the region to be decoded is greater than the number of taps of the interpolation filter corresponding to the index of the MVP farther from the region to be decoded.

[0413] In some embodiments, the derivation unit 12 is specifically configured to, when the motion vector precision is sub-pixel motion vector precision, perform sub-pixel partitioning on the reference region based on the motion vector precision; and for each sub-pixel, derive the target interpolation filter corresponding to the sub-pixel based on the position information of the sub-pixel.

[0414] In some embodiments, the derivation unit 12 is specifically configured to derive the target interpolation filter corresponding to the sub-pixel based on the parity of the position of the sub-pixel.

[0415] In some embodiments, the derivation unit 12 is specifically configured to obtain a first interpolation filter and a second interpolation filter corresponding to the motion vector precision, where the number of taps of the first interpolation filter and the second interpolation filter is different; when the position of the sub-pixel is odd, determine the first interpolation filter as the target interpolation filter corresponding to the sub-pixel; and when the position of the sub-pixel is even, determine the second interpolation filter as the target interpolation filter corresponding to the sub-pixel.

[0416] In some embodiments, the derivation unit 12 is specifically configured to, when the motion vector precision is sub-pixel motion vector precision, implicitly derive the target interpolation filter based on the relevant quantity of the decoding coefficients.

[0417] In some embodiments, the derivation unit 12 is specifically configured to implicitly derive the target interpolation filter based on the parity of the relevant quantity of the decoding coefficients.

[0418] In some embodiments, the derivation unit 12 is specifically configured to implicitly derive the target interpolation filter based on the parity of the number of non-zero coefficients or zero coefficients in the decoding coefficients; or, implicitly derive the target interpolation filter based on the parity of the number of even coefficients among the non-zero coefficients; or, implicitly derive the target interpolation filter based on the parity of the number of odd coefficients among the non-zero coefficients.

[0419] In some embodiments, the decoding coefficients include transform coefficients or quantization coefficients.

[0420] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 7 The illustrated apparatus can execute the embodiments of the above video decoding method, and the foregoing and other operations and / or functions of each module in the apparatus are respectively for implementing the above method embodiments. For the sake of brevity, they will not be elaborated here.

[0421] Figure 8 is a schematic block diagram of a video encoding apparatus provided by an embodiment of the present application. The apparatus 20 can be applied to an encoding device.

[0422] As Figure 8 shown, the video encoding apparatus 20 includes:

[0423] An acquisition unit 21, configured to determine prediction information of a region to be encoded;

[0424] A determination unit 22, configured to determine a target interpolation filter corresponding to the region to be encoded based on the prediction information;

[0425] An interpolation unit 23, configured to determine a reference region of the region to be encoded, and perform interpolation on the reference region based on the target interpolation filter to determine a predicted value of the region to be encoded.

[0426] In some embodiments, the prediction information is at least one of an index of the motion vector accuracy corresponding to the region to be encoded and a prediction mode. The determination unit 22 is specifically configured to implicitly derive the target interpolation filter based on at least one of the motion vector accuracy and the prediction mode.

[0427] In some embodiments, the determining unit 22 is specifically configured to, when the motion vector precision is sub-pixel motion vector precision and the prediction mode is a non-affine motion compensation mode, based on the motion vector precision, derive the target interpolation filter from a first relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the non-affine motion compensation mode; when the motion vector precision is sub-pixel motion vector precision and the prediction mode is an affine motion compensation mode, based on the motion vector precision, derive the target interpolation filter from a second relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the affine motion compensation mode.

[0428] In some embodiments, the first relationship list includes the sub-pixel motion vector precision among the n motion vector precisions corresponding to the adaptive motion vector precision, and the second relationship list includes the sub-pixel motion vector precision among the m motion vector precisions corresponding to the affine motion compensation mode, where both n and m are positive integers.

[0429] In some embodiments, in the first relationship list and the second relationship list, the number of taps of the interpolation filter corresponding to the sub-pixel motion vector precision with higher precision is greater than the number of taps of the interpolation filter corresponding to the sub-pixel motion vector precision with lower precision.

[0430] In some embodiments, the determining unit 22 is specifically configured to, when the prediction mode is motion vector prediction based on historical information, determine the MVP corresponding to the to-be-coded region from a motion vector prediction MVP candidate list; and derive the target interpolation filter based on the index of the MVP and the motion vector precision.

[0431] In some embodiments, the determining unit 22 is specifically configured to, when the motion vector precision is sub-pixel motion vector precision, derive the target interpolation filter based on the index of the MVP.

[0432] In some embodiments, the number of taps of the interpolation filter corresponding to the index of the MVP closer to the to-be-coded region is greater than the number of taps of the interpolation filter corresponding to the index of the MVP farther from the to-be-coded region.

[0433] In some embodiments, the determining unit 22 is specifically configured to, when the motion vector precision is sub-pixel motion vector precision, perform sub-pixel partitioning on the reference region based on the motion vector precision; for each sub-pixel, derive the target interpolation filter corresponding to the sub-pixel based on the position information of the sub-pixel.

[0434] In some embodiments, the determining unit 22 is specifically configured to derive a target interpolation filter corresponding to the sub-pixel based on the parity of the position of the sub-pixel.

[0435] In some embodiments, the determining unit 22 is specifically configured to obtain a first interpolation filter and a second interpolation filter corresponding to the motion vector accuracy, where the number of taps of the first interpolation filter and the second interpolation filter is different; if the position of the sub-pixel is odd, determine the first interpolation filter as the target interpolation filter corresponding to the sub-pixel; if the position of the sub-pixel is even, determine the second interpolation filter as the target interpolation filter corresponding to the sub-pixel.

[0436] In some embodiments, the prediction information includes the motion vector accuracy corresponding to the region to be encoded. The determining unit 22 is specifically configured to, if the motion vector accuracy is sub-pixel motion vector accuracy, obtain a third interpolation filter and a fourth interpolation filter corresponding to the motion vector accuracy, where the number of taps of the third interpolation filter and the fourth interpolation filter is different; determine the rate-distortion costs corresponding to the third interpolation filter and the fourth interpolation filter respectively; and select the target interpolation filter from the third interpolation filter and the fourth interpolation filter based on the rate-distortion cost.

[0437] In some embodiments, the determining unit 22 is further configured to determine an encoding coefficient of the region to be encoded based on a predicted value of the region to be encoded; determine a preset parity corresponding to the target interpolation filter; and determine whether to adjust the encoding coefficient based on the preset parity.

[0438] In some embodiments, the determining unit 22 is specifically configured to, if the preset parity is consistent with the parity of a relevant quantity of the encoding coefficient, skip adjusting the encoding coefficient; if the preset parity is inconsistent with the parity of the relevant quantity of the encoding coefficient, adjust the encoding coefficient so that the parity of the relevant quantity of the encoding coefficient is consistent with the preset parity.

[0439] In some embodiments, the parity of the relevant quantity of the encoding coefficient includes: the parity of the number of non-zero coefficients or zero coefficients in the encoding coefficient, or the parity of the number of even coefficients among the non-zero coefficients, or the parity of the number of odd coefficients among the non-zero coefficients.

[0440] In some embodiments, the encoding coefficient includes a transform coefficient or a quantization coefficient.

[0441] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, details are not described here again. Specifically, Figure 8The device shown can execute the embodiments of the above video encoding method, and the foregoing and other operations and / or functions of each module in the device respectively implement the above method embodiments. For the sake of brevity, they will not be elaborated here.

[0442] In the foregoing, the device of the embodiments of the present application has been described from the perspective of functional modules in combination with the accompanying drawings. It should be understood that the functional module can be implemented in the form of hardware, can also be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0443] Figure 9 is a schematic block diagram of an electronic device provided by an embodiment of the present application, Figure 9 The electronic device can be the above encoding device or a decoding device.

[0444] As Figure 9 shown, the electronic device 30 may include:

[0445] A memory 31 and a processor 32. The memory 31 is used to store a computer program 33 and transmit the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiments of the present application.

[0446] For example, the processor 32 can be used to execute the steps in the above method 200 according to the instructions in the computer program 33.

[0447] In some embodiments of the present application, the processor 32 may include, but is not limited to:

[0448] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.

[0449] In some embodiments of the present application, the memory 31 includes, but is not limited to:

[0450] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double DataRate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0451] In some embodiments of the present application, the computer program 33 can be divided into one or more modules, and the one or more modules are stored in the memory 31 and executed by the processor 32 to complete the method for recording a page provided by the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 33 in the electronic device.

[0452] As Figure 9 shown, the electronic device 30 may further include:

[0453] A transceiver 34, and the transceiver 34 can be connected to the processor 32 or the memory 31.

[0454] Among them, the processor 32 can control the transceiver 34 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 34 can include a transmitter and a receiver. The transceiver 34 can further include an antenna, and the number of antennas can be one or more.

[0455] It should be understood that the various components in the computing device 30 are connected through a bus system. Among them, in addition to the data bus, the bus system also includes a power bus, a control bus, and a status signal bus.

[0456] According to one aspect of the present application, there is provided a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods of the above method embodiments. Or, the embodiments of the present application further provide a computer program product including instructions. When the instructions are executed by a computer, the computer executes the methods of the above method embodiments.

[0457] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods of the above method embodiments.

[0458] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0459] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0460] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical, or other form.

[0461] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0462] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A video decoding method, characterized in that, comprising: decoding a bitstream to obtain at least one of prediction information and decoding coefficients of a region to be decoded; implicitly deriving a target interpolation filter corresponding to the region to be decoded based on at least one of the prediction information and the decoding coefficients; determining a reference region of the region to be decoded, and performing interpolation on the reference region based on the target interpolation filter to determine a predicted value of the region to be decoded.

2. The method according to claim 1, characterized in that, the prediction information includes at least one of an index of a motion vector precision corresponding to the region to be decoded and a prediction mode, and the implicitly deriving a target interpolation filter corresponding to the region to be decoded based on at least one of the prediction information and the decoding coefficients includes: determining a motion vector precision corresponding to the region to be decoded based on the index of the motion vector precision; implicitly deriving the target interpolation filter based on at least one of the motion vector precision, the prediction mode, and the decoding coefficients.

3. The method according to claim 2, characterized in that, the implicitly deriving the target interpolation filter based on at least one of the motion vector precision, the prediction mode, and the decoding coefficients includes: implicitly deriving the target interpolation filter based on the motion vector precision and the prediction mode.

4. The method according to claim 3, characterized in that, the implicitly deriving the target interpolation filter based on the motion vector precision and the prediction mode includes: if the motion vector precision is a sub-pixel motion vector precision and the prediction mode is a non-affine motion compensation mode, then based on the motion vector precision, in a first relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the non-affine motion compensation mode, deriving the target interpolation filter; if the motion vector precision is a sub-pixel motion vector precision and the prediction mode is an affine motion compensation mode, then based on the motion vector precision, in a second relationship list corresponding to different sub-pixel motion vector precisions and different interpolation filters in the affine motion compensation mode, deriving the target interpolation filter.

5. The method according to claim 4, characterized in that, the first relationship list includes sub-pixel motion vector precisions among n motion vector precisions corresponding to an adaptive motion vector precision, the second relationship list includes sub-pixel motion vector precisions among m motion vector precisions corresponding to the affine motion compensation mode, and both n and m are positive integers.

6. The method according to claim 4 or 5, characterized in that, in the first relationship list and the second relationship list, the number of taps of the interpolation filter corresponding to a sub-pixel motion vector precision with a higher precision is greater than the number of taps of the interpolation filter corresponding to a sub-pixel motion vector precision with a lower precision.

7. The method according to claim 3, characterized in that, the implicitly deriving the target interpolation filter based on the motion vector precision and the prediction mode includes: When the prediction mode is motion vector prediction based on historical information, decode the code stream to obtain the index of the motion vector prediction (MVP) corresponding to the area to be decoded. Derive the target interpolation filter based on the index of the MVP and the motion vector precision.

8. The method according to claim 7, wherein, the deriving the target interpolation filter based on the index of the MVP and the motion vector precision includes: when the motion vector precision is sub-pixel motion vector precision, derive the target interpolation filter based on the index of the MVP.

9. The method according to claim 2, wherein, the implicitly deriving the target interpolation filter based on at least one of the motion vector precision, the prediction mode, and the decoding coefficient includes: when the motion vector precision is sub-pixel motion vector precision, perform sub-pixel partitioning on the reference area based on the motion vector precision; for each sub-pixel, derive the target interpolation filter corresponding to the sub-pixel based on the position information of the sub-pixel.

10. The method according to claim 9, wherein, the deriving the target interpolation filter corresponding to the sub-pixel based on the position information of the sub-pixel includes: derive the target interpolation filter corresponding to the sub-pixel based on the parity of the position of the sub-pixel.

11. The method according to claim 10, wherein, the deriving the target interpolation filter corresponding to the sub-pixel based on the parity of the position of the sub-pixel includes: obtain a first interpolation filter and a second interpolation filter corresponding to the motion vector precision, where the number of taps of the first interpolation filter and the second interpolation filter is different; when the position of the sub-pixel is odd, determine the first interpolation filter as the target interpolation filter corresponding to the sub-pixel; when the position of the sub-pixel is even, determine the second interpolation filter as the target interpolation filter corresponding to the sub-pixel.

12. The method according to claim 2, wherein, the implicitly deriving the target interpolation filter based on at least one of the motion vector precision, the prediction mode, and the decoding coefficient includes: when the motion vector precision is sub-pixel motion vector precision, implicitly derive the target interpolation filter based on the relevant quantity of the decoding coefficient.

13. The method according to claim 12, wherein, the implicitly deriving the target interpolation filter based on the relevant quantity of the decoding coefficient includes: implicitly derive the target interpolation filter based on the parity of the relevant quantity of the decoding coefficient.

14. The method according to claim 13, wherein, the implicitly deriving the target interpolation filter based on the parity of the relevant quantity of the decoding coefficient includes: implicitly derive the target interpolation filter based on the parity of the number of non-zero coefficients or zero coefficients in the decoding coefficient; or, implicitly derive the target interpolation filter based on the parity of the number of even coefficients among the non-zero coefficients; or, The target interpolation filter is implicitly derived based on the parity of the number of odd coefficients among the non-zero coefficients.

15. A video encoding method, characterized in that, comprising: determining prediction information of a region to be encoded; determining a target interpolation filter corresponding to the region to be encoded based on the prediction information; determining a reference region of the region to be encoded, and performing interpolation on the reference region based on the target interpolation filter to determine a predicted value of the region to be encoded.

16. The method according to claim 15, characterized in that, the prediction information includes the motion vector precision corresponding to the region to be encoded, and the determining the target interpolation filter corresponding to the region to be encoded based on the prediction information includes: if the motion vector precision is sub-pixel motion vector precision, obtaining a third interpolation filter and a fourth interpolation filter corresponding to the motion vector precision, the number of taps of the third interpolation filter and the fourth interpolation filter being different; determining rate-distortion costs corresponding to the third interpolation filter and the fourth interpolation filter respectively; selecting the target interpolation filter from the third interpolation filter and the fourth interpolation filter based on the rate-distortion cost.

17. The method according to claim 16, characterized in that, the method further comprises: determining encoding coefficients of the region to be encoded based on the predicted value of the region to be encoded; determining a preset parity corresponding to the target interpolation filter; determining whether to adjust the encoding coefficients based on the preset parity.

18. A video decoding apparatus, characterized in that, comprising: a decoding unit configured to decode a bitstream to obtain at least one of prediction information of a region to be decoded and decoding coefficients; a derivation unit configured to implicitly derive a target interpolation filter corresponding to the region to be decoded based on at least one of the prediction information and the decoding coefficients; an interpolation unit configured to determine a reference region of the region to be decoded, and perform interpolation on the reference region based on the target interpolation filter to determine a predicted value of the region to be decoded.

19. A video encoding apparatus, characterized in that, comprising: an acquisition unit configured to determine prediction information of a region to be encoded; a determination unit configured to determine a target interpolation filter corresponding to the region to be encoded based on the prediction information; an interpolation unit configured to determine a reference region of the region to be encoded, and perform interpolation on the reference region based on the target interpolation filter to determine a predicted value of the region to be encoded.

20. An electronic device, comprising a processor and a memory; the memory is configured to store a computer program; the processor is configured to execute the computer program to implement the method according to any one of claims 1 to 14 or 15 to 17 above.