Video coding method and apparatus, video decoding method and apparatus, and device and storage medium

By adaptively selecting the interpolation filter based on the decoding information in video encoding and decoding, the problem of inaccurate selection of interpolation filters in the prior art is solved, and the video prediction effect and codec performance are improved.

WO2025118652A1PCT designated stage expired Publication Date: 2025-06-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/109317
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-09
Filing Date
2024-08-01
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The prior art is inaccurate when selecting an interpolation filter, resulting in poor video prediction effect.

Method used

The decoding information of the area to be decoded is obtained by decoding the code stream, including prediction mode, interpolation direction, decoding component and interpolation target, and adaptively select the target interpolation filter to improve the selection accuracy of the interpolation filter.

Benefits of technology

Improves video prediction effect and improves video encoding and codec performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109317_12062025_PF_FP_ABST
    Figure CN2024109317_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a video coding method and apparatus, a video decoding method and apparatus, and a device and a storage medium, which can be applied to computing fields such as image coding and decoding and video coding and decoding. The decoding method comprises: decoding a code stream, so as to obtain decoding information of a region to be decoded, wherein the decoding information comprises at least one of a prediction mode, an interpolation direction, a decoding component and an interpolation target; on the basis of the decoding information, determining a target interpolation filter corresponding to said region; and determining a reference region for said region, performing interpolation on the reference region on the basis of the target interpolation filter, and determining a predicted value of said region. In the present application, on the basis of information such as a prediction mode, an interpolation direction, a decoding component and an interpolation target, a target interpolation filter is adaptively selected from among interpolation filters with different numbers of taps, such that the accuracy of selection of the target interpolation filter is improved, thereby improving the prediction effect for a region to be to coded and decoded, and improving the performance of video coding and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding and decoding method, device, equipment and storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 9, 2023, with application number 202311694332X and invention name “Video Coding and Decoding Method, Device, Equipment and Storage Medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of video coding and decoding technology, and in particular to a video coding and decoding method, apparatus, device, and storage medium. Background Art

[0003] Digital video technology can be incorporated into a variety of video devices, such as digital televisions, smartphones, computers, e-readers, or video players. With the development of video technology, the amount of data included in video data has increased. To facilitate the transmission of video data, video devices implement video compression technology to enable more efficient transmission or storage of video data.

[0004] Because temporal redundancy exists in video, inter-frame prediction can be used to eliminate it and improve compression efficiency. In some cases, interpolation filters are required to interpolate the reference image. However, current interpolation filter selection methods often lead to inaccurate predictions, resulting in poor prediction performance.

[0005] Summary of the Invention

[0006] The present application provides a video encoding and decoding method, apparatus, device and storage medium for improving the accuracy of selecting interpolation filters, thereby improving the prediction effect of the video.

[0007] In a first aspect, the present application provides a video decoding method, which is applied to a decoder or executed by a processor, comprising:

[0008] Decoding the bitstream to obtain decoding information of the to-be-decoded region, wherein the decoding information includes at least one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target;

[0009] Determining a target interpolation filter corresponding to the to-be-decoded area based on the decoded information;

[0010] A reference area of ​​the area to be decoded is determined, and based on the target interpolation filter, the reference area is interpolated to determine a prediction value of the area to be decoded.

[0011] In a second aspect, an embodiment of the present application provides a video encoding method, which is applied to an encoder or executed by a processor, including:

[0012] Obtaining encoding information of a to-be-encoded region and a reference region corresponding to the to-be-encoded region, wherein the encoding information includes at least one of a prediction mode, an interpolation direction, an encoding component, and an interpolation target;

[0013] Determining a target interpolation filter corresponding to the to-be-encoded area based on the encoding information;

[0014] Based on the target interpolation filter, interpolation filtering is performed on the reference area to determine a prediction value of the area to be encoded.

[0015] In a third aspect, the present application provides a video decoding device, applied to a decoder, comprising:

[0016] A decoding unit, configured to decode the bitstream to obtain decoding information of the area to be decoded, wherein the decoding information includes at least one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target;

[0017] a determining unit, configured to determine a target interpolation filter corresponding to the to-be-decoded area based on the decoding information;

[0018] The interpolation unit is configured to determine a reference area of ​​the area to be decoded, and interpolate the reference area based on the target interpolation filter to determine a predicted value of the area to be decoded.

[0019] In a fourth aspect, the present application provides a video encoding device, applied to an encoder, comprising:

[0020] an acquiring unit, configured to acquire encoding information of a to-be-encoded region and a reference region corresponding to the to-be-encoded region, wherein the encoding information includes at least one of a prediction mode, an interpolation direction, an encoding component, and an interpolation target;

[0021] a determining unit, configured to determine a target interpolation filter corresponding to the to-be-encoded area based on the encoding information;

[0022] An interpolation unit is configured to perform interpolation filtering on the reference area based on the target interpolation filter to determine a prediction value of the area to be encoded.

[0023] In a fifth aspect, the present application provides a video decoder comprising a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the first aspect or its respective implementations.

[0024] In a sixth aspect, the present application provides a video encoder comprising a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the second aspect or its respective implementations.

[0025] In a seventh aspect, the present application provides a video encoding and decoding system, comprising a video encoder and a video decoder. The video decoder is configured to perform the method of the first aspect or its respective implementations, and the video encoder is configured to perform the method of any one of the first to second aspects or their respective implementations.

[0026] In an eighth aspect, the present application provides a chip for implementing the method of any one of the first and second aspects above, or their respective implementations. Specifically, the chip includes: a processor for calling and executing a computer program from a memory, so that a device equipped with the chip executes the method of any one of the first and second aspects above, or their respective implementations.

[0027] In a ninth aspect, the present application provides a computer-readable storage medium for storing a computer program, which enables a computer to execute the method of any one of the first to second aspects above or any of their implementations.

[0028] In a tenth aspect, the present application provides a computer program product, comprising computer program instructions, which enable a computer to execute the method of any one of the above-mentioned first to second aspects or their respective implementations.

[0029] In an eleventh aspect, the present application provides a computer program which, when executed on a computer, enables the computer to execute the method in any one of the first to second aspects or their respective implementations.

[0030] Based on the above technical solution, the present application obtains the decoding information of the area to be decoded by decoding the code stream, and the decoding information includes at least one of the prediction mode, interpolation direction, decoding component, and interpolation target; based on the decoding information, the target interpolation filter corresponding to the area to be decoded is determined; the reference area of ​​the area to be decoded is determined, and based on the target interpolation filter, the reference area is interpolated to determine the prediction value of the area to be decoded. That is to say, in an embodiment of the present application, the decoding end adaptively selects the target interpolation filter from the interpolation filters with different tap numbers based on the encoding and decoding information of the area to be encoded and decoded, such as the prediction mode, interpolation direction, decoding component, interpolation target and other information, thereby improving the selection accuracy of the target interpolation filter. In this way, based on the accurately selected target interpolation filter, the reference area of ​​the area to be encoded and decoded is interpolated to determine the prediction value of the area to be encoded and decoded, which can improve the prediction effect of the area to be encoded and decoded, thereby improving the performance of video encoding and decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] FIG1 is a schematic block diagram of a video encoding and decoding system according to an embodiment of the present application;

[0032] FIG2 is a schematic block diagram of a video encoder according to an embodiment of the present application;

[0033] FIG3 is a schematic block diagram of a video decoder according to an embodiment of the present application;

[0034] FIG4A is a schematic diagram of inter-frame prediction;

[0035] FIG4B is a schematic diagram of pixel interpolation;

[0036] FIG5 is a flow chart of a video decoding method according to an embodiment of the present application;

[0037] FIG6 is a flow chart of a video encoding method according to an embodiment of the present application;

[0038] FIG7 is a schematic diagram of determining multiple candidate interpolation filters;

[0039] FIG8 is a schematic diagram of determining a target interpolation filter based on rate-distortion cost;

[0040] FIG9 is a schematic block diagram of a video decoding device provided in an embodiment of the present application;

[0041] FIG10 is a schematic block diagram of a video encoding apparatus according to an embodiment of the present application;

[0042] FIG11 is a schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0044] It should be noted that the terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In an embodiment of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or devices. In the description of this application, unless otherwise specified, "plurality" means two or more than two.

[0045] The present application can be applied to the field of image coding and decoding, the field of video coding and decoding, the field of hardware video coding and decoding, the field of dedicated circuit video coding and decoding, the field of real-time video coding and decoding, etc. For example, the solution of the present application can be combined with an audio and video coding standard (AVS), such as the H.264 / audio and video coding (AVC) standard, the H.265 / high efficiency video coding (HEVC) standard, and the H.266 / versatile video coding (VVC) standard. Alternatively, the solution of the present application can be combined with other proprietary or industry standards and operated, and the standards include ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), including scalable video coding (SVC) and multi-view video coding (MVC) extensions. It should be understood that the technology of this application is not limited to any specific coding standard or technology.

[0046] For ease of understanding, the video encoding and decoding system involved in the embodiment of the present application is first introduced with reference to FIG1 .

[0047] FIG1 is a schematic block diagram of a video encoding and decoding system involved in an embodiment of the present application. It should be noted that FIG1 is only an example, and the video encoding and decoding system of the embodiment of the present application includes but is not limited to that shown in FIG1. ​​As shown in FIG1, the video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compressing) the video data to generate a code stream, and transmit the code stream to the decoding device. The decoding device decodes the code stream generated by the encoding device to obtain decoded video data.

[0048] The encoding device 110 of the embodiment of the present application can be understood as a device with a video encoding function, and the decoding device 120 can be understood as a device with a video decoding function, that is, the embodiment of the present application includes a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, car computers, etc.

[0049] In some embodiments, the encoding device 110 may transmit the encoded video data (eg, a code stream) to the decoding device 120 via a channel 130. The channel 130 may include one or more media and / or devices capable of transmitting the encoded video data from the encoding device 110 to the decoding device 120.

[0050] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to transmit the encoded video data directly to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media may also include wired communication media, such as one or more physical transmission lines.

[0051] In another example, channel 130 includes a storage medium that can store the video data encoded by encoding device 110. The storage medium includes various locally accessible data storage media, such as optical disks, DVDs, and flash memories. In this example, decoding device 120 can retrieve the encoded video data from the storage medium.

[0052] In another example, the channel 130 may include a storage server that can store the video data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded video data from the storage server. Alternatively, the storage server can store the encoded video data and transmit the encoded video data to the decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.

[0053] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113. The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.

[0054] In some embodiments, the encoding device 110 may further include a video source 111 in addition to the video encoder 112 and the input interface 113 .

[0055] The video source 111 may include at least one of a video acquisition device (eg, a video camera), a video archive, a video input interface, and a computer graphics system, wherein the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data.

[0056] The video encoder 112 encodes the video data from the video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream contains the coding information of the picture or picture sequence in the form of a bitstream. The coding information may include the coded picture data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. The SPS may contain parameters that apply to one or more sequences. The PPS may contain parameters that apply to one or more pictures. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.

[0057] The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data may also be stored in a storage medium or a storage server for subsequent reading by the decoding device 120.

[0058] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122 .

[0059] In some embodiments, the decoding device 120 may further include a display device 123 in addition to the input interface 121 and the video decoder 122 .

[0060] The input interface 121 includes a receiver and / or a modem and can receive the encoded video data via the channel 130 .

[0061] The video decoder 122 is configured to decode the encoded video data to obtain decoded video data, and transmit the decoded video data to the display device 123 .

[0062] The decoded video data is displayed on the display device 123. The display device 123 may be integrated with the decoding device 120 or external to the decoding device 120. The display device 123 may include various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0063] In addition, Figure 1 is only an example, and the technical solution of the embodiment of the present application is not limited to Figure 1. For example, the technology of the present application can also be applied to unilateral video encoding or unilateral video decoding.

[0064] The following is an introduction to the video encoding framework involved in the embodiments of the present application.

[0065] FIG2 is a schematic block diagram of a video encoder according to an embodiment of the present application. It should be understood that the video encoder 200 can be used to perform lossy compression or lossless compression on an image. The lossless compression can be visually lossless or mathematically lossless.

[0066] The video encoder 200 can be applied to image data in a luminance and chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents brightness (Luma), Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance (Chroma) for describing color and saturation. For example, in terms of color format, 4:2:0 means that every 4 pixels have 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that every 4 pixels have 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 represents full pixel display (YYYYCbCrCbCrCbCrCbCr).

[0067] For example, the video encoder 200 reads video data, and for each frame of the video data, divides the frame into a number of coding tree units (CTUs). In some examples, CTB may be referred to as a "tree block", "largest coding unit" (LCU) or "coding tree block" (CTB). Each CTU may be associated with a pixel block of equal size within the image. Each pixel may correspond to a luminance (luminance or luma) sample and two chrominance (chroma) samples. Therefore, each CTU may be associated with a luminance sample block and two chroma sample blocks. The size of a CTU is, for example, 128×128, 64×64, 32×32, etc. A CTU can be further divided into a number of coding units (CUs) for encoding. The CU may be a rectangular block or a square block. A CU can be further divided into prediction units (PUs) and transform units (TUs), allowing for separation of coding, prediction, and transform, and greater flexibility in processing. In one example, a CTU is divided into CUs using a quadtree, and a CU is divided into TUs and PUs using a quadtree.

[0068] The video encoder and video decoder can support various PU sizes. Assuming that the size of a particular CU is 2N×2N, the video encoder and video decoder can support PU sizes of 2N×2N or N×N for intra-frame prediction, and support symmetric PUs of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-frame prediction. The video encoder and video decoder can also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-frame prediction.

[0069] In some embodiments, as shown in FIG2 , the video encoder 200 may include a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filter unit 260, a decoded image buffer 270, and an entropy coding unit 280. It should be noted that the video encoder 200 may include more, fewer, or different functional components.

[0070] Optionally, in this application, the current block may be referred to as the current coding unit (CU) or the current prediction unit (PU), etc. The prediction block may also be referred to as a predicted image block or an image prediction block, and the reconstructed image block may also be referred to as a reconstructed block or an image reconstructed image block.

[0071] In some embodiments, the prediction unit 210 includes an inter-frame prediction unit 211 and an intra-frame prediction unit 212. Because there is a strong correlation between adjacent pixels in a video frame, intra-frame prediction is used in video coding and decoding technologies to eliminate spatial redundancy between adjacent pixels. Because there is a strong similarity between adjacent frames in a video, inter-frame prediction is used in video coding and decoding technologies to eliminate temporal redundancy between adjacent frames, thereby improving coding efficiency.

[0072] The inter-frame prediction unit 211 is used for inter-frame prediction. Inter-frame prediction includes motion estimation and motion compensation. It can refer to image information from different frames. Inter-frame prediction uses motion information to find a reference block from a reference frame and generates a prediction block based on the reference block to eliminate temporal redundancy. The frames used for inter-frame prediction can be P-frames and / or B-frames. P-frames refer to forward-predicted frames, while B-frames refer to bidirectionally predicted frames. Inter-frame prediction uses motion information to find a reference block from a reference frame and generates a prediction block based on the reference block. Motion information includes the reference frame list in which the reference frame is located, the reference frame index, and the motion vector. The motion vector can be integer pixel or fractional pixel. If the motion vector is fractional pixel, interpolation filtering is required to generate the required fractional pixel block in the reference frame. Here, the integer pixel or fractional pixel block in the reference frame found based on the motion vector is called a reference block. Some technologies directly use the reference block as the prediction block, while others further process the reference block to generate a prediction block. Reprocessing the reference block to generate a prediction block can also be understood as using the reference block as the prediction block and then processing the prediction block to generate a new prediction block.

[0073] The intra-frame prediction unit 212 only refers to the information of the same frame image to predict the pixel information in the current code image block to eliminate spatial redundancy. The frame used for intra-frame prediction can be an I frame.

[0074] Intra-frame prediction has multiple prediction modes. For example, the H-series international digital video coding standard H.264 / AVC has eight angular prediction modes and one non-angular prediction mode. H.265 / HEVC expands this to 33 angular prediction modes and two non-angular prediction modes. HEVC uses planar, DC, and 33 angular modes for a total of 35 intra-frame prediction modes. VVC uses planar, DC, and 65 angular modes for a total of 67 intra-frame prediction modes.

[0075] It should be noted that with the increase of angle modes, intra-frame prediction will be more accurate and more in line with the needs of high-definition and ultra-high-definition digital video development.

[0076] The residual unit 220 may generate a residual block for the CU based on the pixel blocks of the CU and the prediction blocks of the PUs of the CU. For example, the residual unit 220 may generate the residual block for the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel blocks of the CU and the corresponding sample in the prediction blocks of the PUs of the CU.

[0077] The transform / quantization unit 230 may quantize the transform coefficients. The transform / quantization unit 230 may quantize the transform coefficients associated with the TUs of a CU based on a quantization parameter (QP) value associated with the CU. The video encoder 200 may adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.

[0078] The inverse transform / quantization unit 240 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficients to reconstruct a residual block from the quantized transform coefficients.

[0079] Reconstruction unit 250 may add samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by prediction unit 210 to generate a reconstructed image block associated with the TU. By reconstructing the sample blocks of each TU of a CU in this manner, video encoder 200 can reconstruct the pixel blocks of the CU.

[0080] The loop filter unit 260 is used to process the inverse transformed and inverse quantized pixels to compensate for distortion information and provide a better reference for subsequent coded pixels. For example, it can perform a deblocking filtering operation to reduce the blocking effect of pixel blocks associated with the CU.

[0081] In some embodiments, the loop filtering unit 260 includes a deblocking filtering unit and a sample adaptive offset / adaptive loop filtering (SAO / ALF) unit, wherein the deblocking filtering unit is used to remove blocking effects, and the SAO / ALF unit is used to remove ringing effects.

[0082] The decoded image buffer 270 may store the reconstructed pixel blocks. The inter-prediction unit 211 may use a reference image containing the reconstructed pixel blocks to perform inter-prediction on PUs of other images. In addition, the intra-prediction unit 212 may use the reconstructed pixel blocks in the decoded image buffer 270 to perform intra-prediction on other PUs in the same image as the CU.

[0083] The entropy coding unit 280 may receive the quantized transform coefficients from the transform / quantization unit 230. The entropy coding unit 280 may perform one or more entropy coding operations on the quantized transform coefficients to generate entropy-coded data.

[0084] FIG3 is a schematic block diagram of a video decoder according to an embodiment of the present application.

[0085] 3 , the video decoder 300 includes an entropy decoding unit 310, a prediction unit 320, an inverse quantization / transformation unit 330, a reconstruction unit 340, a loop filter unit 350, and a decoded picture buffer 360. It should be noted that the video decoder 300 may include more, fewer, or different functional components.

[0086] The video decoder 300 may receive a bitstream. The entropy decoding unit 310 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, the entropy decoding unit 310 may parse the entropy-encoded syntax elements in the bitstream. The prediction unit 320, the inverse quantization / transform unit 330, the reconstruction unit 340, and the loop filter unit 350 may decode the video data based on the syntax elements extracted from the bitstream, thereby generating decoded video data.

[0087] In some embodiments, the prediction unit 320 includes an intra-frame prediction unit 322 and an inter-frame prediction unit 321 .

[0088] The intra-prediction unit 322 may perform intra-prediction to generate a prediction block for the PU. The intra-prediction unit 322 may use an intra-prediction mode to generate a prediction block for the PU based on pixel blocks of spatially neighboring PUs. The intra-prediction unit 322 may also determine the intra-prediction mode for the PU based on one or more syntax elements parsed from the codestream.

[0089] The inter-frame prediction unit 321 may construct a first reference picture list (List 0) and a second reference picture list (List 1) based on syntax elements parsed from the codestream. In addition, if a PU is encoded using inter-frame prediction, the entropy decoding unit 310 may parse the motion information of the PU. The inter-frame prediction unit 321 may determine one or more reference blocks for the PU based on the motion information of the PU. The inter-frame prediction unit 321 may generate a prediction block for the PU based on the one or more reference blocks of the PU.

[0090] The inverse quantization / transform unit 330 may inversely quantize (ie, dequantize) the transform coefficients associated with the TU. The inverse quantization / transform unit 330 may use the QP value associated with the CU of the TU to determine the degree of quantization.

[0091] After inverse quantizing the transform coefficients, the inverse quantization / transform unit 330 may apply one or more inverse transforms to the inverse quantized transform coefficients in order to generate a residual block associated with the TU.

[0092] The reconstruction unit 340 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, the reconstruction unit 340 can add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU to obtain a reconstructed image block.

[0093] The loop filtering unit 350 may perform a deblocking filtering operation to reduce blocking artifacts of pixel blocks associated with a CU.

[0094] The video decoder 300 may store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 may use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.

[0095] The basic process of video encoding and decoding is as follows: At the encoder end, a frame of image is divided into blocks. For the current block, the prediction unit 210 uses intra-frame prediction or inter-frame prediction to generate a prediction block for the current block. The residual unit 220 calculates a residual block based on the predicted block and the original block of the current block. This residual block is the difference between the predicted block and the original block of the current block. This residual block is also called residual information. This residual block undergoes transformation and quantization by the transform / quantization unit 230, removing information that is insensitive to the human eye and eliminating visual redundancy. Optionally, the residual block before transformation and quantization by the transform / quantization unit 230 is called a time-domain residual block, and the time-domain residual block after transformation and quantization by the transform / quantization unit 230 is called a frequency residual block or a frequency-domain residual block. The entropy coding unit 280 receives the quantized change coefficients output by the change quantization unit 230 and performs entropy coding on these quantized change coefficients to output a bitstream. For example, the entropy coding unit 280 can eliminate character redundancy based on the target context model and probability information of the binary bitstream.

[0096] At the decoding end, the entropy decoding unit 310 can parse the code stream to obtain the prediction information, quantization coefficient matrix, etc. of the current block. The prediction unit 320 uses intra-frame prediction or inter-frame prediction on the current block based on the prediction information to generate a prediction block for the current block. The inverse quantization / transformation unit 330 uses the quantization coefficient matrix obtained from the code stream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block. The reconstruction unit 340 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed blocks constitute a reconstructed image, and the loop filtering unit 350 performs loop filtering on the reconstructed image based on the image or block to obtain a decoded image. The encoding end also requires similar operations as the decoding end to obtain a decoded image. The decoded image can also be called a reconstructed image, and the reconstructed image can be used as a reference frame for inter-frame prediction for subsequent frames.

[0097] It should be noted that the block division information determined by the encoder, as well as mode information or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering, etc., are carried in the bitstream when necessary. The decoder parses the bitstream and analyzes the existing information to determine the same block division information, prediction, transform, quantization, entropy coding, loop filtering, etc. mode information or parameter information as the encoder, thereby ensuring that the decoded image obtained by the encoder and the decoder are identical.

[0098] The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. This application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.

[0099] In the embodiments of the present application, the current block can be the current coding unit (CU) or the current prediction unit (PU). Due to the need for parallel processing, an image can be divided into slices, and slices within the same image can be processed in parallel, meaning that there is no data dependency between them. "Frame" is a commonly used term, generally understood to mean that a frame is an image. In the application, the term "frame" can also be replaced with "image" or "slice."

[0100] Inter-frame prediction is the process of predicting a block to be coded (e.g., the current block) in the current image from a neighboring coded image (e.g., a reference image) to obtain a reference block. The goal is to remove temporal redundancy in the video signal. Figure 4A shows an example of inter-frame prediction. Based on a block matching criterion, the best matching block for the current block is searched in the reference frame.

[0101] Inter-frame prediction uses motion information to represent motion. Basic motion information includes information about the reference frame (or reference picture) and motion vectors (MVs). Bidirectional prediction, commonly used today, uses two reference blocks to predict the current block. These two reference blocks can be a forward reference block and a backward reference block. Later, it was also allowed that both reference blocks be forward or backward. Forward means the reference frame corresponds to a time before the current frame, while backward means the reference frame corresponds to a time after the current frame. In other words, forward means the reference frame's position in the video is before the current frame, while backward means the reference frame's position in the video is after the current frame. In other words, forward means the reference frame's POC (picture order count) is less than the current frame's POC, while backward means the reference frame's POC is greater than the current frame's POC. To use bidirectional prediction, two reference blocks must be found, which requires two sets of reference frame information and motion vector information. Each of these groups can be considered as a unidirectional motion information, and combining these two groups together forms a bidirectional motion information. In actual implementation, unidirectional motion information and bidirectional motion information can use the same data structure, except that both reference frame information and motion vector information of the bidirectional motion information are valid, while the reference frame information and motion vector information of one set of the unidirectional motion information are invalid.

[0102] VVC supports two reference frame lists, denoted as RPL0 and RPL1, where RPL stands for Reference Picture List. P slices in VVC can only use RPL0, while B slices can use both RPL0 and RPL1. For a slice, each reference frame list contains several reference frames, and the codec uses the reference frame index to find a specific reference frame. VVC uses reference frame indices and motion vectors to represent motion information. For example, for the bidirectional motion information described above, VVC uses the reference frame index refIdxL0 corresponding to reference frame list 0 and the motion vector mvL0 corresponding to reference frame list 0, and the reference frame index refIdxL1 corresponding to reference frame list 1 and the motion vector mvL0 corresponding to reference frame list 1. The reference frame index corresponding to reference frame list 0 and the reference frame index corresponding to reference frame list 1 can be understood as the reference frame information described above. VVC uses two flags, denoted as predFlagL0 and predFlagL1, to indicate whether the motion information corresponding to reference frame list 0 and reference frame list 1 is used, respectively. It can also be understood that predFlagL0 and predFlagL1 indicate whether the above-mentioned unidirectional motion information is "valid". Therefore, although VVC does not explicitly mention the data structure of motion information, it uses the reference frame index corresponding to each reference frame list, the motion vector, and the "valid" flag to represent the motion information. Motion information does not appear in the standard text of VVC, but the motion vector is used. It can also be considered that the reference frame index and the flag of whether to use the corresponding motion information are attached to the motion vector. For the sake of convenience in description, this article still uses "motion information", but it should be understood that "motion vector" can also be used to describe it.

[0103] As can be seen from the above, inter-frame prediction requires the use of motion vectors to obtain pixel prediction values ​​in the reference frame. In actual situations, the motion of objects between adjacent images is not necessarily based on integer pixels. Therefore, in order to improve the prediction accuracy, it is necessary to improve the accuracy of motion estimation to the sub-pixel level and interpolate the reference image to improve the accuracy of motion compensation, thereby improving the coding efficiency. As shown in Figure 4B, the pixel where the capital A is located is an integer pixel, for example, A 0,0 、A 1,0 The pixels where lowercase a, b, c, etc. are located are sub-pixel points, for example, a 0,0 is 1 / 4 pixel point in the X direction, b 0,0 For example, the common inter-frame prediction mode in AVS3 supports five motion vector precisions: 1 / 4, 1 / 2, 1, 2, and 4. When the MV precision is 1 / 4 or 1 / 2, an 8-tap interpolation filter can be used to interpolate the luma prediction value and a 4-tap interpolation filter can be used to interpolate the chroma prediction value.

[0104] Traditional bidirectional prediction calculates the predicted value of the current block by weighted averaging two reconstructed blocks, one from the forward reference frame and the other from the backward reference frame. Furthermore, to improve prediction performance, bidirectional optical flow (BIO) can be used to compensate for motion after bidirectional prediction, reducing motion deviation and improving coding efficiency. Specifically, BIO in AVS3 first calculates the gradient values ​​in the x and y directions (x represents the horizontal direction and y represents the vertical direction throughout this text). It then obtains a calculation factor for each pixel based on the pixel value and gradient value, derives the motion vector, and ultimately calculates a more accurate predicted value. For example, when calculating the gradient value in the x direction, BIO first uses an 8-tap gradient filter for gradient calculation, and then uses an 8-tap interpolation filter for interpolation. When calculating the gradient value in the y direction, BIO first uses an 8-tap interpolation filter for interpolation, and then uses an 8-tap gradient filter for gradient calculation. Furthermore, BIO technology is only applied to the luminance component.

[0105] As can be seen from the above, when interpolating prediction values ​​or gradients, the selected interpolation filter is a pre-defined one. For example, when interpolating predictions, an 8-tap interpolation filter is used to interpolate the luma prediction value, and a 4-tap interpolation filter is used to interpolate the chroma prediction value. When calculating gradients in the BIO, an 8-tap interpolation filter is selected. Currently, codec-related information is not considered when determining the interpolation filter, resulting in inaccurate interpolation filter selection and poor video prediction performance.

[0106] To address the above technical issues, the embodiments of the present application adaptively select an interpolation filter based on the codec information of the area to be coded, such as the prediction mode, interpolation direction, decoding component, and interpolation target. This allows the selection of the interpolation filter to be combined with the codec information, thereby improving the accuracy of the interpolation filter selection. In this way, based on the accurately selected target interpolation filter, the reference area of ​​the area to be coded is interpolated to determine the predicted value of the area to be coded, which can improve the prediction effect of the area to be coded, thereby improving the performance of video coding and decoding.

[0107] The following describes the technical solutions of the embodiments of the present application in detail through some embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0108] First, taking the decoding end as an example, the video decoding method provided in the embodiment of the present application is introduced.

[0109] FIG5 is a flow chart of a video decoding method according to an embodiment of the present application, which is applied to the video decoders shown in FIG1 and FIG3 or executed by a processor. As shown in FIG5 , the method according to the embodiment of the present application includes:

[0110] S101: Decode the code stream to obtain decoding information of the area to be decoded.

[0111] The decoding information includes at least one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target.

[0112] The embodiment of the present application does not limit the specific size and shape of the area to be decoded.

[0113] In some embodiments, the to-be-decoded region may be one or more to-be-decoded image blocks in the current frame to be decoded, for example, one or more CTUs, one or more CUs, or one or more PUs. Of course, the to-be-decoded region may also include an incomplete CTU, an incomplete CU, or an incomplete PU. In one example, the to-be-decoded image block may be referred to as a current block, an image block currently to be decoded, or the like.

[0114] In some embodiments, the area to be decoded may be the current frame to be decoded, that is, the area to be decoded may be a whole frame of image.

[0115] During the video decoding process, the decoder decodes the code stream to obtain the quantization coefficients of the area to be decoded. These coefficients are then dequantized to obtain the transform coefficients of the area to be decoded. These transform coefficients are then inversely transformed to obtain the residual value of the area to be decoded. Next, a prediction mode is determined for the area to be decoded. Based on the prediction mode, a predicted value for the area to be decoded is determined. Based on the predicted value and the residual value of the area to be decoded, a reconstructed value for the area to be decoded is obtained.

[0116] The embodiments of the present application mainly relate to the prediction process of the area to be decoded.

[0117] In the embodiment of the present application, in order to improve the prediction accuracy of the area to be decoded, an interpolation filter is used to interpolate the reference area of ​​the area to be decoded. However, the decoding information of the area to be decoded is not considered when selecting the interpolation filter, which makes the selection of the interpolation filter inaccurate.

[0118] In order to solve the above technical problems, an embodiment of the present application obtains decoding information of the area to be decoded by decoding the code stream, and then determines the target interpolation filter of the area to be decoded based on the decoding information to improve the accuracy of selecting the interpolation filter.

[0119] The embodiments of the present application do not limit the specific content of the decoding information of the area to be decoded, which can be understood as all decoding information related to decoding the area to be decoded, such as prediction mode, inverse transformation method, inverse quantization method, interpolation direction, decoding component and interpolation target, etc.

[0120] In some embodiments, the decoding information of the area to be decoded in the embodiments of the present application includes at least one of a prediction mode of the area to be decoded, an interpolation direction to be interpolated, a decoding component to be decoded, and an interpolation target.

[0121] Exemplarily, the prediction mode of the area to be decoded may include various modes in the intra-frame prediction mode, various modes in the inter-frame prediction mode, various modes in the hybrid prediction mode, and the like.

[0122] Exemplarily, the interpolation direction includes an x ​​direction and a y direction.

[0123] Exemplarily, the decoded components include a luminance component and a chrominance component.

[0124] Exemplarily, the interpolation targets include determining a prediction value (e.g., performing interpolation filtering on a reference area of ​​an area to be decoded to determine a prediction value of the area to be decoded), and determining a gradient (e.g., determining the gradient of a forward reference block and a backward reference block in the x direction or y direction in BIO).

[0125] In one possible implementation, the encoder may write the index or indication information of the prediction mode of the area to be decoded into the bitstream, so that the decoder can obtain the prediction mode of the area to be decoded by decoding the bitstream.

[0126] In one possible implementation, the encoder can write the interpolation direction of the area to be decoded into the code stream. For example, if only the x-direction of the area to be decoded is interpolated and the y-direction is not interpolated, the encoder can write interpolation direction indication information into the code stream to indicate that only the x-direction is interpolated and the y-direction is not interpolated. Conversely, if only the y-direction of the area to be decoded is interpolated and the x-direction is not interpolated, the encoder can write interpolation direction indication information into the code stream to indicate that only the y-direction is interpolated and the x-direction is not interpolated. For another example, if both the x-direction and the y-direction of the area to be decoded need to be interpolated, the encoder can write interpolation direction indication information into the code stream to indicate that both the x-direction and the y-direction need to be interpolated. Optionally, the encoder and decoder can predefine whether to interpolate in the x-direction and / or the y-direction.

[0127] In one possible implementation, the encoder encodes the luma and chroma components separately. Based on this, the encoder indicates in the bitstream whether the information included in the bitstream corresponds to the luma or chroma components. This allows the decoder to determine whether the current area to be decoded is the luma or chroma component by decoding the bitstream.

[0128] In one possible implementation, the interpolation target can be derived from the prediction mode of the area to be decoded. For example, when the decoder decodes the code stream and determines that the prediction mode of the area to be decoded is an inter-frame prediction mode and not a BIO mode, and at the same time, when the decoder obtains the sub-pixel accuracy of the motion vector corresponding to the area to be decoded through the decoded code stream, the decoder can determine that the interpolation target is to determine the prediction value. If the decoder decodes the code stream and determines that the prediction mode of the area to be decoded is a BIO mode, the decoder can determine that the interpolation target is to determine the gradient.

[0129] In one example, the decoding information of the area to be decoded may include only any one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target. For example, if the decoding information of the area to be decoded includes a prediction mode, and the prediction mode is prediction mode 1, the decoding end may determine the target interpolation filter from the interpolation filter corresponding to prediction mode 1.

[0130] In one example, the decoding information of the area to be decoded may include any two or three of a prediction mode, an interpolation direction, a decoding component, and an interpolation target. For example, if the decoding information of the area to be decoded includes a prediction mode and an interpolation direction, and the prediction mode is prediction mode 1 and the interpolation direction is the x-direction, the decoding end may determine a target interpolation filter based on the interpolation filter corresponding to prediction mode 1 and the x-direction.

[0131] In one example, the decoding information of the area to be decoded may include a prediction mode, an interpolation direction, a decoding component, and an interpolation target. For example, the decoding information of the area to be decoded may include a prediction mode of 1, an interpolation direction of the x direction, a decoding component of the luminance component, and an interpolation target of determining a predicted value of the area to be decoded. Thus, the decoding end may determine a target interpolation filter based on the interpolation filters corresponding to the prediction mode, the interpolation direction, the decoding component, and the interpolation target.

[0132] In some embodiments, the decoding end may also obtain at least one of the prediction mode, interpolation direction, decoding component, and interpolation target corresponding to the to-be-decoded region in other ways.

[0133] After the decoding end determines the decoding information of the area to be decoded, the following step S102 is executed.

[0134] S102: Determine a target interpolation filter corresponding to the area to be decoded based on the decoded information.

[0135] In this embodiment of the present application, the decoding end determines the target interpolation filter corresponding to the area to be decoded from interpolation filters with different tap numbers based on the decoding information of the area to be decoded. This fully considers the performance differences of interpolation filters with different tap numbers in different coding modes, different decoding components, different interpolation directions, and different interpolation targets, thereby improving the accuracy of the target interpolation filter selection.

[0136] The embodiment of the present application does not limit the specific manner in which the decoding end determines the target interpolation filter corresponding to the to-be-decoded area based on at least one of the prediction mode, the interpolation direction, the decoding component, and the interpolation target.

[0137] For example, in the embodiments of the present application, interpolation filters corresponding to different prediction modes, interpolation filters corresponding to different interpolation directions, interpolation filters corresponding to different decoding components, and interpolation filters corresponding to different interpolation targets are preset. In this way, the decoding end can determine the target interpolation filter corresponding to the area to be decoded based on the lookup filters corresponding to the prediction mode, interpolation direction, decoding component, and interpolation target of the area to be decoded.

[0138] In some embodiments, the above S102 includes the following steps S102-A1 to S102-A3:

[0139] S102-A1, determining multiple candidate interpolation filters based on the decoded information;

[0140] S102-A2: Decode the bitstream to obtain first indication information of a target interpolation filter, where the first indication information is used to indicate index information of the target interpolation filter among multiple candidate interpolation filters.

[0141] S102-A3: Based on the first indication information, select a target interpolation filter from multiple candidate interpolation filters.

[0142] In this implementation, when the decoding end determines the target interpolation filter corresponding to the area to be decoded based on the decoding information of the area to be decoded determined above, it first determines multiple candidate interpolation filters based on the decoding information, and then selects an interpolation filter from these candidate interpolation filters as the target interpolation filter. In other words, the embodiment of the present application determines multiple candidate interpolation filters with different numbers of taps based on the decoding information, and then selects the target interpolation filter from these multiple candidate interpolation filters with different numbers of taps, fully considering the performance differences of interpolation filters with different numbers of taps in different coding modes, different decoding components, different interpolation directions, and different interpolation targets, thereby improving the accuracy of determining the target interpolation filter.

[0143] The following describes a specific process of determining multiple candidate interpolation filters at the decoding end based on the decoding information.

[0144] In the embodiment of the present application, the implementation of determining multiple candidate interpolation filters based on the decoded information in S102-A1 includes at least the following:

[0145] In the first approach, the decoder determines multiple candidate interpolation filters through the following steps S102-A1-a1 and S102-A1-a2:

[0146] S102-A1-a1, obtaining at least one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target, and corresponding preset candidate interpolation filters;

[0147] S102-A1-a2: Determine multiple candidate interpolation filters based on preset candidate interpolation filters.

[0148] In this implementation, at least one of the prediction mode, interpolation direction, decoding component, and interpolation target in the decoding information corresponds to a preset candidate interpolation filter.

[0149] In one example, different prediction modes may correspond to different candidate interpolation filters. For example, prediction mode 1 corresponds to m1 candidate interpolation filters, and prediction mode 2 corresponds to m2 candidate interpolation filters. The number of taps of the m1 candidate interpolation filters is not completely consistent with the number of taps of the m2 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the BIO mode include an 8-tap interpolation filter and a 12-tap interpolation filter. The candidate interpolation filters corresponding to the non-BIO mode include a 6-tap interpolation filter and an 8-tap interpolation filter.

[0150] In one example, different interpolation directions may correspond to different candidate interpolation filters. For example, the x-direction corresponds to m3 candidate interpolation filters, and the y-direction corresponds to m4 candidate interpolation filters. The number of taps of the m3 candidate interpolation filters is not completely consistent with the number of taps of the m4 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the x (horizontal) direction include an 8-tap interpolation filter and a 12-tap interpolation filter, and the candidate interpolation filters corresponding to the y (vertical) direction include a 6-tap interpolation filter and an 8-tap interpolation filter.

[0151] In one example, different decoding components may correspond to different candidate interpolation filters. For example, the luminance component corresponds to m5 candidate interpolation filters, and the chrominance component corresponds to m6 candidate interpolation filters. The number of taps of the m5 candidate interpolation filters is not completely consistent with the number of taps of the m6 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the luminance component include an 8-tap interpolation filter and a 12-tap interpolation filter, and the candidate interpolation filters corresponding to the chrominance component include a 4-tap interpolation filter and a 6-tap interpolation filter.

[0152] For example, different interpolation targets may correspond to different candidate interpolation filters. For example, m7 candidate interpolation filters may be used when determining the predicted value, while m8 candidate interpolation filters may be used when determining the gradient. The number of taps in the m7 candidate interpolation filters may not be exactly the same as the number of taps in the m8 candidate interpolation filters.

[0153] In this way, the decoding end can determine multiple candidate interpolation filters based on the candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, decoding component, and interpolation target of the area to be decoded.

[0154] For example, the decoding end selects a preset number of candidate interpolation filters from the candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, decoding component, and interpolation target of the area to be decoded, as multiple candidate interpolation filters. For example, assuming that the decoding information of the area to be decoded includes a prediction mode and a decoding component, the decoding end selects a preset number (for example, n) of candidate interpolation filters from the candidate interpolation filters corresponding to the prediction mode and the decoding component of the area to be decoded, to form a plurality of candidate interpolation filters. Referring to the above example, assuming that the prediction mode of the area to be decoded is prediction mode 1 and the decoding component is the luminance component, wherein prediction mode 1 corresponds to m1 candidate interpolation filters and the luminance component corresponds to m5 candidate interpolation filters, a preset number of candidate interpolation filters can be selected from the m1 candidate interpolation filters and the m5 candidate interpolation filters, that is, the m1+m5 candidate interpolation filters, to form a plurality of candidate interpolation filters.

[0155] For another example, the decoding end removes the repeated interpolation filters from the selected prediction candidate interpolation filters to obtain multiple candidate interpolation filters. For example, assuming that the decoding information of the area to be decoded includes a prediction mode and a decoding component, where the prediction mode is a non-BIO mode and the decoding component is a brightness component, assuming that the candidate interpolation filters corresponding to the non-BIO mode include a 6-tap interpolation filter and an 8-tap interpolation filter, and the brightness component corresponds to an 8-tap interpolation filter and a 12-tap interpolation filter. An 8-tap interpolation filter is removed from these 4 interpolation filters, and the remaining 3 interpolation filters are a 6-tap interpolation filter, an 8-tap interpolation filter, and a 12-tap interpolation filter, respectively. Then, the 6-tap interpolation filter, the 8-tap interpolation filter, and the 12-tap interpolation filter are determined as multiple candidate interpolation filters.

[0156] As can be seen from the above description, in the first approach, the decoding end selects multiple candidate interpolation filters from preset candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, decoding component and interpolation target of the to-be-decoded region.

[0157] In some embodiments, the decoding end may also determine multiple candidate interpolation filters in the manner shown in the following second manner.

[0158] In the second method, the decoding end determines multiple candidate interpolation filters through the following steps S102-A1-b:

[0159] S102-A1-b, obtain the selection strategy corresponding to at least one of the prediction mode, interpolation direction, decoding component and interpolation target, and based on the selection strategy, select multiple candidate interpolation filters from the preset M interpolation filters, where M is a positive integer greater than 1.

[0160] In this implementation, M interpolation filters are preset. The decoder can select multiple interpolation filters from these M preset interpolation filters to form multiple candidate interpolation filters based on at least one of the prediction mode, interpolation direction, decoding component, and interpolation target of the to-be-decoded region. Optionally, the M interpolation filters have different numbers of taps.

[0161] Specifically, the decoder obtains a selection strategy corresponding to at least one of the prediction mode, interpolation direction, decoding component, and interpolation target for the region to be decoded. This selection strategy is either a default strategy set by the codec or transmitted from the codec to the decoder via the bitstream. Based on this strategy, the decoder can select multiple candidate interpolation filters from the M pre-set interpolation filters.

[0162] In this implementation, the selection strategies corresponding to at least one of the prediction mode, interpolation direction, decoding component and interpolation target of the area to be decoded may be the same or different, or partially the same and partially different.

[0163] Exemplarily, it is assumed that the decoding information of the area to be decoded includes a prediction mode and a decoding component, the prediction mode corresponds to selection strategy 1, and the decoding component corresponds to selection strategy 2. In this way, the decoding end can select one or more interpolation filters from M interpolation filters based on selection strategy 1. At the same time, the decoding end can select one or more interpolation filters from M interpolation filters based on selection strategy 2. Then, the decoding end obtains multiple candidate interpolation filters based on the one or more interpolation filters selected based on strategy 1 and the one or more interpolation filters selected based on strategy 2. For example, the decoding end determines the interpolation filters with different tap numbers among the one or more interpolation filters selected based on strategy 1 and the one or more interpolation filters selected based on strategy 2 as multiple candidate interpolation filters.

[0164] In some embodiments, the decoding end may also determine multiple candidate interpolation filters in the manner shown in the following third manner.

[0165] In a third approach, the decoder determines multiple candidate interpolation filters through the following steps S102-A1-c:

[0166] S102-A1-c, obtain N candidate combinations consisting of M interpolation filters, and decode the code stream to obtain second indication information, and based on the second indication information, select a target candidate combination from the N candidate combinations, and then determine the interpolation filters included in the target candidate combination as multiple candidate interpolation filters, wherein each candidate combination includes two or more interpolation filters from the M interpolation filters, the second indication information is used to indicate the target candidate combination, and N is a positive integer greater than 1.

[0167] In this implementation, M interpolation filters are preset, and the decoding end obtains N candidate combinations composed of these M interpolation filters. Optionally, the number of taps of these M interpolation filters is different.

[0168] In one example, the N candidate combinations are default at both the encoder and decoder. That is, the decoder itself stores the N candidate combinations, or the decoder can use the same rules as the encoder to combine the M interpolation filters to obtain N candidate combinations.

[0169] In one example, the N candidate combinations are indicated by the encoder to the decoder. For example, the encoder writes the determined N candidate combinations into the bitstream. The decoder obtains the N candidate combinations by decoding the bitstream.

[0170] The embodiment of the present application does not limit the specific combination form of the N candidate combinations.

[0171] In one possible implementation, the interpolation filters included in different candidate combinations among the above-mentioned N candidate combinations are randomly distributed. For example, it is assumed that the M interpolation filters include interpolation filters with 5 types of taps, namely 4 taps, 6 taps, 8 taps, 12 taps and 16 taps. The N candidate combinations composed of the interpolation filters with these 5 types of taps are: {6 taps, 8 taps, 12 taps}, {2 taps, 4 taps, 6 taps} and {4 taps, 6 taps, 8 taps, 12 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 1:

[0172] Table 1

[0173] In one possible implementation, the number of interpolation filters included in each candidate combination of the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination increases in sequence. For example, it is assumed that the M interpolation filters include interpolation filters with 6 types of taps, namely 2 taps, 4 taps, 6 taps, 8 taps, 12 taps and 16 taps. The N candidate combinations composed of the interpolation filters with these 6 types of taps include: {2 taps, 4 taps}, {2 taps, 4 taps, 6 taps}, {2 taps, 4 taps, 6 taps, 8 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 2:

[0174] Table 2

[0175] In one possible implementation, the number of interpolation filters included in each candidate combination of the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination decreases in sequence. For example, it is assumed that the M interpolation filters include interpolation filters with 6 types of taps, namely 2 taps, 4 taps, 6 taps, 8 taps, 12 taps and 16 taps. The N candidate combinations composed of the interpolation filters with these 6 types of taps include: {12 taps, 16 taps}, {8 taps, 12 taps, 16 taps}, {6 taps, 8 taps, 12 taps, 16 taps}, {4 taps, 6 taps, 8 taps, 12 taps, 16 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps, 16 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 3:

[0176] Table 3

[0177] In this method three, when determining multiple candidate interpolation filters, the encoding end obtains N candidate combinations and selects a target candidate combination from the N candidate combinations, for example, based on the rate-distortion cost, selects a target candidate combination from the N candidate combinations, and then determines the interpolation filters included in the target candidate combination as multiple candidate interpolation filters. At the same time, the encoding end writes second indication information in the bitstream, and the second indication information is used to indicate the target candidate combination. Exemplarily, the second indication information includes the index of the target candidate combination in the N candidate combinations, for example, the second indication information includes index 100. In this way, the decoding end can obtain the second indication information by decoding the bitstream, and then select the target candidate combination from the N candidate combinations obtained above based on the second indication information. For example, the second indication information includes the index 100 of the target candidate combination in the N candidate combinations, so that the decoding end can select the target candidate combination from the N candidate combinations obtained above based on index 100. For example, the N candidate combinations are shown in Table 3. The candidate combination corresponding to index 100 includes interpolation filters with three tap numbers of 8, 12, and 16. These three interpolation filters are then used as multiple candidate interpolation filters.

[0178] As described above, the decoding end can determine a plurality of candidate interpolation filters based on the decoding information of the to-be-decoded area in the above manner.

[0179] The specific implementation process of the above S102-A2 and S102-A3 is introduced below.

[0180] It should be noted that the embodiment of the present application does not limit the specific implementation order of the above S102-A2 and the above S102-A1. For example, the above S102-A2 can be executed before the above S102-A1, or after the above S102-A1, or simultaneously with the above S102-A1.

[0181] In an embodiment of the present application, when determining the target interpolation filter, the encoding end determines multiple candidate interpolation filters based on the encoding information of the area to be encoded, and then selects a candidate interpolation filter from these multiple candidate interpolation filters as the target interpolation filter. For example, the encoding end selects the target interpolation filter from these multiple candidate interpolation filters based on the rate-distortion cost of each of the multiple candidate interpolation filters. Then, the encoding end writes first indication information into the bitstream, and the first indication information is used to indicate the index information of the target interpolation filter among the multiple candidate interpolation filters. In this way, when determining the target interpolation filter, the decoding end determines multiple candidate interpolation filters through the above steps, and decodes the bitstream at the same time to obtain the first indication information, and then selects the target interpolation filter from the multiple candidate interpolation filters based on the first indication information.

[0182] In some embodiments, the first indication information includes a flag array or multiple flags, and the flag array or multiple flags are used to indicate the decoding information and the target interpolation filter. In this case, the decoding end can obtain the first indication information by decoding the bitstream, and then obtain the decoding information of the to-be-decoded area based on the first indication information.

[0183] In one example, the first indication information indicates the prediction mode, the decoding component, and the target interpolation filter. In this case, the first indication information includes a parsing flag array Flag1[A][I] (A and I are both positive integers), where A represents the decoding component, for example, A=0 represents the luminance component, and A=1 represents the chrominance component. I represents the prediction mode, for example, I=0 represents mode 1, and I=1 represents mode 2. Flag1[A][I] represents the index information of the target interpolation filter of the A component under prediction mode I. For example, the first indication information and its meaning are shown in Table 4:

[0184] Table 4

[0185] As shown in Table 4, assuming that the decoding information of the area to be decoded includes a prediction mode and a decoding component, and the prediction mode is mode 1, and the decoding component is a chroma component, the multiple candidate interpolation filters determined by the decoding end based on the prediction mode 1 and the chroma component include a 4-tap interpolation filter and a 6-tap interpolation filter. In this way, if the decoding end decodes the bit stream and the first indication information obtained includes a flag array Flag1[1][0]=1, then the target interpolation filter is determined to be a 6-tap interpolation filter. If the flag array Flag1[1][0] included in the first indication information is 0, then the target interpolation filter is determined to be a 4-tap interpolation filter.

[0186] In one example, the first indication information indicates the interpolation direction, the decoded component, and the target interpolation filter. In this case, the first indication information includes a parsing flag array Flag2[A][B] (A and B are both positive integers), where A represents the decoded component, for example, A=0 represents the luminance component, and A=1 represents the chrominance component. B represents the interpolation direction, for example, B=0 represents the x direction, and B=1 represents the y direction. Flag2[A][B] represents the index information of the target interpolation filter of the A component in the interpolation direction B. For example, the first indication information and its meaning are shown in Table 5:

[0187] Table 5

[0188] As shown in Table 5, assuming that the decoding information of the area to be decoded includes an interpolation direction and a decoding component, and the interpolation direction is the y direction, and the decoding component is the chrominance component, the multiple candidate interpolation filters determined by the decoding end based on the interpolation direction y direction and the chrominance component of the area to be decoded include a 6-tap interpolation filter and a 10-tap interpolation filter. In this way, if the decoding end decodes the bit stream and the first indication information obtained includes a flag array Flag2[1][1]=0, then the target interpolation filter is determined to be a 6-tap interpolation filter. If the first indication information includes a flag array Flag2[1][1]=1, then the target interpolation filter is determined to be a 10-tap interpolation filter.

[0189] In one example, the first indication information indicates the interpolation target, the decoding component, and the target interpolation filter. In this case, the first indication information includes the parsing flag array Flag3[A][D] (A and D are both positive integers), where A represents the decoding component, for example, A=0 represents the luminance component, and A=1 represents the chrominance component. D represents the interpolation target, for example, D=0 represents the interpolation target to determine the predicted value, and D=1 represents the interpolation target to determine the gradient. Flag3[A][D] represents the index information of the target interpolation filter for the A component of the interpolation target D. Exemplarily, the first indication information and its meaning are shown in Table 6:

[0190] Table 6

[0191] As shown in Table 6, assuming that the decoding information of the area to be decoded includes an interpolation target and a decoding component, and the interpolation target is a determined prediction value, and the decoding component is a chrominance component, the multiple candidate interpolation filters determined by the decoding end based on the interpolation target (i.e., the determined prediction value) and the chrominance component of the area to be decoded include a 4-tap interpolation filter and a 6-tap interpolation filter. In this way, if the decoding end decodes the bit stream and the first indication information obtained includes a flag array Flag3[0][1]=0, then the target interpolation filter is determined to be a 4-tap interpolation filter. If the first indication information includes a flag array Flag3[0][1]=1, then the target interpolation filter is determined to be a 6-tap interpolation filter.

[0192] In this implementation, the decoding end determines multiple candidate interpolation filters with different numbers of taps based on at least one of the prediction mode, interpolation direction, decoding component, and interpolation target corresponding to the area to be decoded, and then selects the target interpolation filter from these multiple candidate interpolation filters with different numbers of taps.

[0193] In some embodiments, the decoding end may determine a target interpolation filter corresponding to the area to be decoded based on a default interpolation filter corresponding to at least one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target corresponding to the area to be decoded.

[0194] In Example 1, the decoding end determines a target interpolation filter based on the prediction mode, decoding components, and interpolation target of the area to be decoded.

[0195] For example, both the encoding end and the decoding end default to using a specific N1-tap interpolation filter for luminance and / or chrominance for j1 interpolation targets in a certain i prediction modes, and using a specific N2-tap interpolation filter for luminance and / or chrominance for j2 interpolation contents in the remaining Ii prediction modes, where i, j1, j2, N1 and N2 are all integers, and N1≠N2.

[0196] For example, if the prediction mode of the area to be decoded is the bidirectional optical flow prediction mode, the decoding component is the luminance component, and the interpolation target is to determine the gradients of the forward prediction block and the backward prediction block in the x-direction and the y-direction of the area to be decoded, then the n1-tap interpolation filter is determined as the target interpolation filter;

[0197] If the prediction mode of the area to be decoded is the bidirectional optical flow prediction mode, the decoding component is the luminance component, and the interpolation target is to determine the prediction value of the area to be decoded in the x direction and the y direction, the n2-tap interpolation filter is determined as the target interpolation filter;

[0198] If the prediction mode of the area to be decoded is the bidirectional optical flow prediction mode, the decoding component is the chrominance component, and the interpolation target is to determine the prediction value of the area to be decoded in the x direction and the y direction, the interpolation filter with n3 taps is determined as the target interpolation filter;

[0199] If the prediction mode of the area to be decoded is a non-bidirectional optical flow prediction mode, the decoding component is a luminance component, and the interpolation target is to determine the prediction value of the area to be decoded in the x direction and the y direction, the n4-tap interpolation filter is determined as the target interpolation filter;

[0200] If the prediction mode of the area to be decoded is a non-bidirectional optical flow prediction mode, the decoding component is a chrominance component, and the interpolation target is to determine the prediction value of the area to be decoded in the x direction and the y direction, the interpolation filter with n5 taps is determined as the target interpolation filter;

[0201] Among them, n1, n2, n3, n4 and n5 are all positive integers.

[0202] The embodiment of the present application does not limit the specific values ​​of the above n1, n2, n3, n4 and n5.

[0203] In a possible implementation, the above n1 is not equal to the above n2.

[0204] In one example, n1=8, n2=12, n3=6, n4=12, and n5=6.

[0205] A specific implementation example is shown in Table 7:

[0206] Table 7

[0207] As can be seen from Table 7 above, in the BIO mode of the present embodiment, the number of taps of the interpolation filter used when calculating the gradient is different from the number of taps of the interpolation filter used when calculating the pixel prediction value. For example, an 8-tap interpolation filter is used to calculate the gradient, and a 12-tap interpolation filter is used for pixel prediction value interpolation calculation.

[0208] Example 2: The decoding end determines a target interpolation filter based on the prediction mode and the decoded components of the area to be decoded.

[0209] For example, both the encoding end and the decoding end default to using only a specific N1-tap interpolation filter for luminance and / or chrominance in a certain i prediction modes, and using a certain N2-tap interpolation filter for luminance and / or chrominance in the remaining Ii prediction modes (i, N1 and N2 are all integers, N1≠N2).

[0210] For example, if the prediction mode of the area to be decoded is the bidirectional optical flow prediction mode and the decoding component is the luminance component, the interpolation filter with the m1 tap is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with the m1 tap to perform gradient calculation and prediction value interpolation;

[0211] If the prediction mode of the area to be decoded is the bidirectional optical flow prediction mode and the decoding component is the chrominance component, the interpolation filter with m2 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with m2 taps to perform interpolation calculations on the predicted values ​​in the x and y directions.

[0212] If the prediction mode of the area to be decoded is a non-bidirectional optical flow prediction mode and the decoding component is a luminance component, the interpolation filter with the m3 tap is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with the m3 tap to perform interpolation calculations on the predicted values ​​in the x and y directions.

[0213] If the prediction mode of the area to be decoded is a non-bidirectional optical flow prediction mode and the decoding component is a chrominance component, the interpolation filter with m4 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with m4 taps to perform interpolation calculations on the predicted values ​​in the x and y directions.

[0214] Among them, m1, m2, m3 and m4 are all positive integers.

[0215] The embodiment of the present application does not limit the specific values ​​of the above m1, m2, m3 and m4.

[0216] In one example, m1=8, m2=6, m3=12, and m4=6.

[0217] A specific implementation example is shown in Table 8:

[0218] Table 8

[0219] Example 3: The decoding end determines the target interpolation filter based on the interpolation direction and the decoding component.

[0220] For example, the encoding end and the decoding end default to using an N1-tap interpolation filter for interpolation of the predicted values, gradients, and other contents of luminance and / or chrominance in the x-direction, and an N2-tap interpolation filter for interpolation of the predicted values, gradients, and other contents of luminance and / or chrominance in the y-direction (N1 and N2 are both positive integers, and N1≠N2).

[0221] For example, if the interpolation direction of the area to be decoded is the x direction and the decoding component is the luminance component, the interpolation filter with k1 tap is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with k1 tap to perform gradient calculation and prediction value interpolation;

[0222] If the interpolation direction of the area to be decoded is the x direction and the decoded component is the chrominance component, the k2-tap interpolation filter is determined as the target interpolation filter. At this time, the decoding end uses the k2-tap interpolation filter to perform gradient calculation and prediction value interpolation;

[0223] If the interpolation direction of the area to be decoded is the y direction and the decoded component is the luminance component, the k3-tap interpolation filter is determined as the target interpolation filter. At this time, the decoding end uses the k3-tap interpolation filter for gradient calculation and prediction value interpolation;

[0224] If the interpolation direction of the area to be decoded is the y direction and the decoded component is the chrominance component, the k4-tap interpolation filter is determined as the target interpolation filter. At this time, the decoding end uses the k4-tap interpolation filter for gradient calculation and prediction value interpolation;

[0225] Wherein, k1, k2, k3 and k4 are all positive integers.

[0226] The embodiment of the present application does not limit the specific values ​​of the above k1, k2, k3 and k4.

[0227] In one example, k1=12, k2=6, k3=8, and k4=4.

[0228] A specific implementation example is shown in Table 9:

[0229] Table 9

[0230] Exemplarily, the interpolation filters used in the x-direction and the y-direction may be swapped, that is, an interpolation filter with a shorter tap number (e.g., 8 taps and 4 taps) may be used in the x-direction, and an interpolation filter with a longer tap number (e.g., 12 taps and 6 taps) may be used in the y-direction.

[0231] Example 4: The decoding end determines a target interpolation filter based on the decoded components and the size of the area to be decoded.

[0232] For example, both the encoding end and the decoding end use an N1-tap interpolation filter by default within a certain block size range, and use an N2-tap interpolation filter within other block size ranges (N1 and N2 are both positive integers, and N1≠N2).

[0233] For example, if the decoding component of the area to be decoded is a luminance component, and the scale of the area to be decoded is w<=W1 or h<=H1, the interpolation filter with tap a1 is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with tap a1 to perform gradient calculation and prediction value interpolation;

[0234] If the decoding component of the area to be decoded is the luminance component, and the scale of the area to be decoded is w>W1 and h>H1, the interpolation filter with a2 tap is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with a2 tap to perform gradient calculation and prediction value interpolation;

[0235] If the decoding component of the area to be decoded is a chrominance component and the scale of the area to be decoded is w<=W2 or h<=H2, the interpolation filter with a3 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with a3 taps to perform gradient calculation and prediction value interpolation;

[0236] If the decoding component of the area to be decoded is a chrominance component, and the scale of the area to be decoded is w>W2 and h>H2, the a4-tap interpolation filter is determined as the target interpolation filter. At this time, the decoding end uses the a4-tap interpolation filter for gradient calculation and prediction value interpolation;

[0237] Wherein, a1, a2, a3 and a4 are all positive integers.

[0238] The embodiment of the present application does not limit the specific values ​​of the above a1, a2, a3 and a4.

[0239] In one example, a1=8, a2=12, a3=4, and a4=6.

[0240] A specific implementation example is shown in Table 10:

[0241] Table 10

[0242] Wherein, w and h represent the block size of the to-be-decoded area in the current component, and W1, H1, W2, and H2 are all positive integers.

[0243] After the decoding end determines the target interpolation filter based on the above steps, it executes the following step S103.

[0244] S103 : Determine a reference area of ​​the area to be decoded, and interpolate the reference area based on a target interpolation filter to determine a prediction value of the area to be decoded.

[0245] Based on the above steps, the decoding end determines the target interpolation filter corresponding to the area to be decoded, and then uses the target interpolation filter to interpolate the reference area of ​​the area to be decoded to determine the predicted value of the area to be decoded.

[0246] The embodiments of the present application do not restrict the order in which the decoder determines the reference area of ​​the area to be decoded and determines the target interpolation filter corresponding to the area to be decoded. That is, the decoder may first determine the target interpolation filter corresponding to the area to be decoded, and then determine the reference area of ​​the area to be decoded. Alternatively, the decoder may first determine the reference area of ​​the area to be decoded, and then determine the target interpolation filter corresponding to the area to be decoded. Alternatively, the decoder may determine the target interpolation filter corresponding to the area to be decoded at the same time as determining the reference area of ​​the area to be decoded.

[0247] In the embodiment of the present application, the reference area of ​​the area to be decoded can be a reference block or a reference frame, which is not limited in the embodiment of the present application. The specific method for the decoding end to determine the reference area of ​​the area to be decoded can refer to the description of the relevant technology and will not be repeated here.

[0248] Next, the decoding end interpolates the reference area of ​​the area to be decoded based on the target interpolation filter determined above, and further determines the predicted value of the area to be decoded.

[0249] In one example, if the interpolation target is to interpolate pixel prediction values ​​(i.e., determine the prediction values), the decoder uses the target interpolation filter to interpolate the reference region of the to-be-decoded region to obtain an interpolated reference region. Furthermore, based on the motion vector corresponding to the to-be-decoded region, the prediction block of the to-be-decoded region is obtained in the interpolated reference region.

[0250] In one example, if the prediction mode corresponding to the area to be decoded in the embodiment of the present application is BIO, the decoding end first uses the bidirectional prediction mode to determine a forward reference block and a backward reference block for the area to be decoded. Next, the target interpolation filter is sampled, and the forward and backward reference blocks are interpolated to determine the gradients of the forward and backward reference blocks. The motion vectors corresponding to the forward and backward reference blocks are then corrected based on the gradients to obtain the predicted value of the area to be decoded.

[0251] In some embodiments, the bit widths of the filter coefficients of the interpolation filters with different numbers of taps in the embodiments of the present application may be different, and / or the bit widths of the filter coefficients of the interpolation filters with the same number of taps may also be different.

[0252] In some embodiments, the present application also provides interpolation filter coefficients for some different numbers of taps. Specifically, MV position can be understood as the position of the pixel point, coefficients is the filter coefficient, and in some embodiments, the interpolation filter that determines the gradient is also referred to as a gradient filter or a gradient interpolation filter.

[0253] Table 11: 8-tap gradient filter coefficients used in BIO mode

[0254] Table 12: 12-tap filter coefficients used in BIO and non-BIO modes except AFFINE mode

[0255] Table 13: 12-tap filter system used in AFFINE mode

[0256] Table 14: 6-tap filter coefficients used in BIO and non-BIO modes except AFFINE mode

[0257] Table 15: 6-tap filter coefficients used in AFFINE mode

[0258] The video decoding method provided in the embodiment of the present application obtains the decoding information of the area to be decoded by decoding the code stream, and the decoding information includes at least one of the prediction mode, interpolation direction, decoding component, and interpolation target; based on the decoding information, the target interpolation filter corresponding to the area to be decoded is determined; the reference area of ​​the area to be decoded is determined, and based on the target interpolation filter, the reference area is interpolated to determine the prediction value of the area to be decoded. That is to say, in the embodiment of the present application, the decoding end adaptively selects the target interpolation filter from the interpolation filters with different tap numbers based on the encoding and decoding information of the area to be encoded and decoded, such as the prediction mode, interpolation direction, decoding component, interpolation target and other information, thereby improving the selection accuracy of the target interpolation filter. In this way, based on the accurately selected target interpolation filter, the reference area of ​​the area to be encoded and decoded is interpolated to determine the prediction value of the area to be encoded and decoded, which can improve the prediction effect of the area to be encoded and decoded, thereby improving the performance of video encoding and decoding.

[0259] The above describes the video decoding method of the present application using the decoding end as an example, and the following describes it using the encoding end as an example.

[0260] FIG6 is a flow chart of a video encoding method according to an embodiment of the present application, which is applied to the video encoders shown in FIG1 and FIG2 or executed by a processor. As shown in FIG6 , the method according to the embodiment of the present application includes:

[0261] Step S201: Obtain coding information of the area to be coded and a reference area corresponding to the area to be coded.

[0262] The coding information includes at least one of a prediction mode, an interpolation direction, a coding component, and an interpolation target.

[0263] The embodiment of the present application does not limit the specific size and shape of the coding area.

[0264] In some embodiments, the region to be encoded may be one or more image blocks to be encoded in the current frame to be encoded, for example, one or more CTUs, one or more CUs, or one or more PUs. Of course, the region to be encoded may also include an incomplete CTU, an incomplete CU, or an incomplete PU. In one example, the image block to be encoded may be referred to as a current block, an image block currently to be encoded, or the like.

[0265] In some embodiments, the area to be encoded may be the current frame to be encoded, that is, the area to be encoded is a whole frame of image.

[0266] During the video encoding process, the encoder determines the prediction mode for the region to be encoded, and based on the prediction mode, determines the predicted value for the region to be encoded. Furthermore, based on the predicted value for the region to be encoded, it determines the residual value for the region to be encoded. The residual value is then transformed to obtain transform coefficients, which are then quantized to obtain quantized coefficients. Finally, the quantized coefficients are encoded to obtain a bitstream.

[0267] The embodiments of the present application mainly relate to the prediction process of the area to be encoded.

[0268] In the embodiment of the present application, in order to improve the prediction accuracy of the to-be-encoded region, an interpolation filter is used to interpolate the reference region of the to-be-encoded region. However, currently, the encoding information of the to-be-encoded region is not considered when selecting the interpolation filter, resulting in inaccurate selection of the interpolation filter.

[0269] In order to solve the above technical problems, in an embodiment of the present application, a target interpolation filter of the area to be encoded is determined based on encoding information of the area to be encoded, so as to improve the accuracy of selecting the interpolation filter.

[0270] The embodiment of the present application does not limit the specific content of the encoding information of the encoding area to be encoded, which can be understood as all encoding information related to encoding the area to be encoded, such as prediction mode, transformation method, quantization method, interpolation direction, encoding component and interpolation target, etc.

[0271] In some embodiments, the encoding information of the region to be encoded in the embodiments of the present application includes at least one of a prediction mode of the region to be encoded, an interpolation direction to be interpolated, an encoding component to be encoded, and an interpolation target.

[0272] Exemplarily, the prediction mode of the area to be encoded may include various modes in the intra-frame prediction mode, various modes in the inter-frame prediction mode, various modes in the hybrid prediction mode, and the like.

[0273] Exemplarily, the interpolation direction includes an x ​​direction and a y direction.

[0274] Exemplarily, the coding components include a luminance component and a chrominance component.

[0275] Exemplarily, the interpolation targets include determining a prediction value (e.g., performing pixel interpolation filtering on a reference area of ​​the area to be encoded to determine a prediction value of the area to be encoded), and determining a gradient (e.g., determining the gradient of the forward reference block and the backward reference block in the x-direction or y-direction in BIO).

[0276] In one example, the encoding information of the to-be-encoded region may include only one of a prediction mode, an interpolation direction, an encoding component, and an interpolation target. For example, if the encoding information of the to-be-encoded region includes a prediction mode, and the prediction mode is prediction mode 1, the encoder may determine the target interpolation filter from the interpolation filter corresponding to the prediction mode.

[0277] In one example, the encoding information of the to-be-encoded region may include any two or three of a prediction mode, an interpolation direction, an encoding component, and an interpolation target. For example, if the encoding information of the to-be-encoded region includes a prediction mode and an interpolation direction, and the prediction mode is prediction mode 1 and the interpolation direction is the x-direction, the encoder may determine a target interpolation filter based on the interpolation filters corresponding to prediction mode 1 and the x-direction.

[0278] In one example, the encoding information of the area to be encoded may include a prediction mode, an interpolation direction, an encoding component, and an interpolation target. This embodiment of the present application is not limited to this. For example, the encoding information of the area to be encoded includes the prediction mode being prediction mode 1, the interpolation direction being the x direction, the encoding component being the luminance component, and the interpolation target being the prediction value for determining the area to be encoded. In this way, the encoding end can determine the target interpolation filter based on the interpolation filters corresponding to the prediction mode, interpolation direction, encoding component, and interpolation target.

[0279] In some embodiments, the encoding end may also obtain at least one of the prediction mode, interpolation direction, encoding component, and interpolation target corresponding to the to-be-encoded region in other ways.

[0280] In the embodiment of the present application, the reference area of ​​the area to be encoded can be a reference block or a reference frame, which is not limited in the embodiment of the present application. The specific method for the encoder to determine the reference area of ​​the area to be encoded can refer to the description of the relevant technology and will not be repeated here.

[0281] After the encoding end determines the encoding information of the area to be encoded, the following step S202 is performed.

[0282] Step S202: Determine a target interpolation filter corresponding to the area to be encoded based on the encoding information.

[0283] In an embodiment of the present application, the encoding end determines the target interpolation filter corresponding to the area to be encoded from interpolation filters with different numbers of taps based on the encoding information of the area to be encoded, fully considering the performance differences of interpolation filters with different numbers of taps in different encoding modes, different decoding components, different interpolation directions, and different interpolation targets, thereby improving the accuracy of selecting the target interpolation filter.

[0284] The embodiment of the present application does not limit the specific manner in which the encoder determines the target interpolation filter corresponding to the to-be-encoded area based on at least one of the prediction mode, the interpolation direction, the encoding component, and the interpolation target.

[0285] For example, in the embodiments of the present application, interpolation filters corresponding to different prediction modes, interpolation filters corresponding to different interpolation directions, interpolation filters corresponding to different coding components, and interpolation filters corresponding to different interpolation targets are preset. In this way, the encoder can determine the target interpolation filter corresponding to the area to be encoded based on the lookup filters corresponding to the prediction mode, interpolation direction, coding component, and interpolation target of the area to be encoded.

[0286] In some embodiments, the above step S202 includes the following steps S202-A1 to S202-A2:

[0287] S202-A1, determining multiple candidate interpolation filters based on the encoding information;

[0288] S202-A2: Select a target interpolation filter from multiple candidate interpolation filters.

[0289] In this implementation, when the encoding end determines the target interpolation filter corresponding to the to-be-encoded area based on the encoding information of the to-be-encoded area determined above, it first determines multiple candidate interpolation filters based on the encoding information, and then selects an interpolation filter from these candidate interpolation filters as the target interpolation filter. In other words, the embodiment of the present application determines multiple candidate interpolation filters with different numbers of taps based on the encoding information, and then selects the target interpolation filter from these multiple candidate interpolation filters with different numbers of taps, fully considering the performance differences of interpolation filters with different numbers of taps in different encoding modes, different decoding components, different interpolation directions, and different interpolation targets, thereby improving the accuracy of determining the target interpolation filter.

[0290] The following describes a specific process of determining multiple candidate interpolation filters at the encoding end based on encoding information.

[0291] In the embodiment of the present application, the implementation of determining multiple candidate interpolation filters based on the encoding information in S202-A1 includes at least the following:

[0292] In the first approach, the encoder determines multiple candidate interpolation filters through the following steps S202-A1-a1 and S202-A1-a2:

[0293] S202-A1-a1, obtaining at least one of a prediction mode, an interpolation direction, a coding component, and an interpolation target, and corresponding preset candidate interpolation filters;

[0294] S202-A1-a2: Determine multiple candidate interpolation filters based on preset candidate interpolation filters.

[0295] In this implementation, at least one of the prediction mode, interpolation direction, coding component, and interpolation target in the coding information corresponds to a preset candidate interpolation filter.

[0296] In one example, different prediction modes may correspond to different candidate interpolation filters. For example, prediction mode 1 corresponds to m1 candidate interpolation filters, and prediction mode 2 corresponds to m2 candidate interpolation filters. The number of taps of the m1 candidate interpolation filters is not completely consistent with the number of taps of the m2 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the BIO mode include an 8-tap interpolation filter and a 12-tap interpolation filter. The candidate interpolation filters corresponding to the non-BIO mode include a 6-tap interpolation filter and an 8-tap interpolation filter.

[0297] In one example, different interpolation directions may correspond to different candidate interpolation filters. For example, the x-direction corresponds to m3 candidate interpolation filters, and the y-direction corresponds to m4 candidate interpolation filters. The number of taps of the m3 candidate interpolation filters is not completely consistent with the number of taps of the m4 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the x (horizontal) direction include an 8-tap interpolation filter and a 12-tap interpolation filter, and the candidate interpolation filters corresponding to the y (vertical) direction include a 6-tap interpolation filter and an 8-tap interpolation filter.

[0298] In one example, different coding components may correspond to different candidate interpolation filters. For example, the luminance component corresponds to m5 candidate interpolation filters, and the chrominance component corresponds to m6 candidate interpolation filters. The number of taps of the m5 candidate interpolation filters is not completely consistent with the number of taps of the m6 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the luminance component include an 8-tap interpolation filter and a 12-tap interpolation filter, and the candidate interpolation filters corresponding to the chrominance component include a 4-tap interpolation filter and a 6-tap interpolation filter.

[0299] For example, different interpolation targets may correspond to different candidate interpolation filters. For example, m7 candidate interpolation filters may be used when determining the predicted value, while m8 candidate interpolation filters may be used when determining the gradient. The number of taps in the m7 candidate interpolation filters may not be exactly the same as the number of taps in the m8 candidate interpolation filters.

[0300] In this way, the encoding end can determine multiple candidate interpolation filters based on the candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, coding component, and interpolation target of the to-be-encoded region.

[0301] For example, the encoding end selects a preset number of candidate interpolation filters from the candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, coding component, and interpolation target of the to-be-encoded region, as a plurality of candidate interpolation filters. For example, assuming that the coding information of the to-be-encoded region includes a prediction mode and a coding component, the encoding end selects a preset number (e.g., n) of the candidate interpolation filters corresponding to the prediction mode and coding component of the to-be-encoded region, forming a plurality of candidate interpolation filters. Referring to the above example, assuming that the prediction mode of the to-be-encoded region is prediction mode 1 and the coding component is the luminance component, wherein prediction mode 1 corresponds to m1 candidate interpolation filters and the luminance component corresponds to m5 candidate interpolation filters, a preset number of candidate interpolation filters can be selected from the m1 candidate interpolation filters and the m5 candidate interpolation filters, i.e., the m1+m5 candidate interpolation filters, to form a plurality of candidate interpolation filters.

[0302] For another example, the encoding end removes duplicate interpolation filters from the selected prediction candidate interpolation filters to obtain multiple candidate interpolation filters. For example, assuming that the coding information of the area to be encoded includes a prediction mode and a coding component, where the prediction mode is a non-BIO mode and the coding component is a luminance component, assuming that the candidate interpolation filters corresponding to the non-BIO mode include a 6-tap interpolation filter and an 8-tap interpolation filter, and the luminance component corresponds to an 8-tap interpolation filter and a 12-tap interpolation filter. An 8-tap interpolation filter is removed from these four interpolation filters, and the remaining three interpolation filters are a 6-tap interpolation filter, an 8-tap interpolation filter, and a 12-tap interpolation filter, respectively. Then, the three interpolation filters of the 6-tap interpolation filter, the 8-tap interpolation filter, and the 12-tap interpolation filter are determined as multiple candidate interpolation filters.

[0303] As can be seen from the above description, in the first approach, the encoder selects multiple candidate interpolation filters from preset candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, coding component and interpolation target of the to-be-encoded region.

[0304] In some embodiments, the encoding end may also determine multiple candidate interpolation filters in the manner shown in the following second manner.

[0305] In the second method, the encoder selects multiple candidate interpolation filters from the preset M interpolation filters.

[0306] For example, as shown in Figure 7, the M interpolation filters include six interpolation filters with 2 taps, 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The encoder can select multiple interpolation filters from these six interpolation filters with different numbers of taps as multiple candidate interpolation filters.

[0307] In an example of the second approach, the encoder determines multiple candidate interpolation filters through the following steps S202-A1-b1 and S202-A1-b2:

[0308] S202-A1-b1, obtaining a selection strategy corresponding to at least one of a prediction mode, an interpolation direction, a coding component, and an interpolation target;

[0309] S202-A1-b2. Based on the selection strategy, select multiple candidate interpolation filters from the preset M interpolation filters, where M is a positive integer greater than 1.

[0310] In this implementation, M interpolation filters are preset. The encoder can select multiple candidate interpolation filters from these M preset interpolation filters based on at least one of the prediction mode, interpolation direction, coding component, and interpolation target of the to-be-encoded region. Optionally, the M interpolation filters have different numbers of taps.

[0311] Specifically, the encoder obtains a selection strategy corresponding to at least one of the prediction mode, interpolation direction, coding component, and interpolation target for the region to be encoded. This selection strategy is a default strategy for both the encoder and decoder. Based on this strategy, the encoder selects multiple candidate interpolation filters from the M preset interpolation filters.

[0312] In this implementation, the selection strategies corresponding to at least one of the prediction mode, interpolation direction, coding component and interpolation target of the to-be-encoded region may be the same or different, or partially the same and partially different.

[0313] Exemplarily, assuming that the coding information of the area to be encoded includes a prediction mode and a coding component, the prediction mode corresponds to selection strategy 1, and the coding component corresponds to selection strategy 2. In this way, the encoding end can select one or more interpolation filters from M interpolation filters based on selection strategy 1. At the same time, the encoding end can select one or more interpolation filters from M interpolation filters based on selection strategy 2. Then, the encoding end obtains multiple candidate interpolation filters based on the one or more interpolation filters selected based on strategy 1 and the one or more interpolation filters selected based on strategy 2. For example, the encoding end determines the interpolation filters with different tap numbers among the one or more interpolation filters selected based on strategy 1 and the one or more interpolation filters selected based on strategy 2 as multiple candidate interpolation filters.

[0314] In some embodiments, the above S202-A1-b2 includes the following steps S202-A1-b21 and S202-A1-b22:

[0315] S202-A1-b21. Obtain N candidate combinations consisting of M interpolation filters, each candidate combination including two or more interpolation filters from the M interpolation filters, where N is a positive integer greater than 1;

[0316] S202-A1-b22: Based on the selection strategy, select a target candidate combination from the N candidate combinations, and determine the interpolation filters included in the target candidate combination as a plurality of candidate interpolation filters.

[0317] In this implementation, M interpolation filters are preset, and the encoding end obtains N candidate combinations composed of these M interpolation filters. Optionally, the number of taps of these M interpolation filters is different.

[0318] In one example, the N candidate combinations are default at both the encoder and decoder. That is, the encoder itself stores the N candidate combinations, or the encoder can use the same rules as the decoder to combine M interpolation filters to obtain N candidate combinations.

[0319] The embodiment of the present application does not limit the specific combination form of the N candidate combinations.

[0320] In one possible implementation, the interpolation filters included in different candidate combinations among the above-mentioned N candidate combinations are randomly distributed. For example, assume that the M interpolation filters include interpolation filters with 5 types of taps: 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The N candidate combinations composed of interpolation filters with these 5 types of taps are: {6 taps, 8 taps, 12 taps}, {2 taps, 4 taps, 6 taps}, and {4 taps, 6 taps, 8 taps, 12 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 1.

[0321] In one possible implementation, the number of interpolation filters included in each candidate combination of the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination increases in sequence. For example, assume that the M interpolation filters include interpolation filters with 6 types of taps: 2 taps, 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The N candidate combinations composed of the interpolation filters with these 6 types of taps include: {2 taps, 4 taps}, {2 taps, 4 taps, 6 taps}, {2 taps, 4 taps, 6 taps, 8 taps}, {2 taps, 4 taps, 6 taps, 8 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 2.

[0322] In one possible implementation, the number of interpolation filters included in each of the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination decreases in sequence. For example, assume that the M interpolation filters include interpolation filters with 6 types of taps: 2 taps, 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The N candidate combinations composed of these 6 types of taps include: {12 taps, 16 taps}, {8 taps, 12 taps, 16 taps}, {6 taps, 8 taps, 12 taps, 16 taps}, {4 taps, 6 taps, 8 taps, 12 taps, 16 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps, 16 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 3.

[0323] Next, the encoder selects a target candidate combination from the N candidate combinations based on a selection strategy corresponding to at least one of the prediction mode, interpolation direction, coding component, and interpolation target, and further determines the interpolation filters included in the target candidate combination as multiple candidate interpolation filters. The embodiments of the present application do not limit the specific content of the selection strategy.

[0324] In some embodiments, after the encoder selects a target candidate combination from the N candidate combinations based on the above steps, it writes second indication information into the bitstream. The second indication information is used to indicate the target candidate combination. Exemplarily, the second indication information includes an index of the target candidate combination in the N candidate combinations, for example, the second indication information includes index 100. In this way, the decoder can obtain the second indication information by decoding the bitstream, and then select the target candidate combination from the N candidate combinations obtained above based on the second indication information.

[0325] As described above, the encoder can determine multiple candidate interpolation filters based on the coding information of the to-be-coded region in the above manner. Next, the encoder performs the above step S202-A2 to select a target interpolation filter from the multiple candidate interpolation filters.

[0326] The embodiment of the present application does not limit the specific manner in which the encoder selects a target interpolation filter from multiple candidate interpolation filters.

[0327] In one possible implementation, the encoder randomly selects a target interpolation filter from multiple candidate interpolation filters, or selects a target interpolation filter from multiple candidate interpolation filters based on some feature information (such as scale) of the region to be encoded.

[0328] In a possible implementation, the above S202-A2 includes the following steps S202-A21 to S202-A23:

[0329] S202-A21, for each candidate interpolation filter among the plurality of candidate interpolation filters, interpolate the candidate interpolation filter based on the reference region to obtain an interpolated prediction region;

[0330] S202-A22, determining a rate-distortion cost corresponding to a candidate interpolation filter based on the interpolated prediction region;

[0331] S202-A23: Select a target interpolation filter from the multiple candidate interpolation filters based on the rate-distortion cost corresponding to each candidate interpolation filter.

[0332] In this implementation, for each candidate interpolation filter among multiple candidate interpolation filters, for example, the i-th candidate interpolation filter, the encoding end uses the i-th candidate interpolation filter to interpolate the reference area to obtain the interpolated prediction area, and then, based on the interpolated prediction area, determines the rate-distortion cost corresponding to the i-th candidate interpolation filter. The calculation process of the rate-distortion cost can refer to the description of the relevant technology and will not be repeated here. In this way, the encoding end can determine the rate-distortion cost corresponding to each candidate interpolation filter among multiple candidate interpolation filters, and then select the target interpolation filter from multiple candidate interpolation filters based on the rate-distortion cost corresponding to each candidate interpolation filter among multiple candidate interpolation filters. For example, from multiple candidate interpolation filters, the candidate interpolation filter with the smallest rate-distortion cost is determined as the target interpolation filter.

[0333] For example, as shown in FIG8 , it is assumed that the multiple candidate interpolation filters include three candidate interpolation filters with 6 taps, 8 taps, and 12 taps. The encoder uses these three candidate interpolation filters to interpolate the reference area of ​​the area to be encoded, respectively, to obtain the prediction areas corresponding to the three candidate interpolation filters, and then calculates the rate-distortion cost based on the prediction areas corresponding to the three candidate interpolation filters, to obtain the rate-distortion cost corresponding to the three candidate interpolation filters, and then determines the candidate interpolation filter with the smallest rate-distortion cost as the target interpolation filter.

[0334] In some embodiments, the encoder selects a target interpolation filter from multiple candidate interpolation filters based on their respective rate-distortion costs, and then writes first indication information into the bitstream. This first indication information is used to indicate the index information of the target interpolation filter among the multiple candidate interpolation filters. Thus, when determining the target interpolation filter, the decoder uses the above steps to determine multiple candidate interpolation filters, decodes the bitstream simultaneously, obtains the first indication information, and then selects the target interpolation filter from the multiple candidate interpolation filters based on this first indication information.

[0335] In some embodiments, the first indication information includes a flag array or multiple flags, and the flag array or multiple flags are used to indicate the encoding information and the target interpolation filter.

[0336] In one example, the first indication information indicates the prediction mode, the coding component, and the target interpolation filter. In this case, the first indication information includes a parsing flag array Flag1[A][I] (A and I are both positive integers), where A represents the decoding component, for example, A=0 represents the luma component, and A=1 represents the chroma component. I represents the prediction mode, for example, I=0 represents mode 1, and I=1 represents mode 2. Flag1[A][I] represents the index information of the target interpolation filter of the A component under prediction mode I. Exemplarily, the first indication information and its meaning are shown in Table 4.

[0337] In one example, the first indication information indicates the interpolation direction, the coding component, and the target interpolation filter. In this case, the first indication information includes a parsing flag array Flag2[A][B] (A and B are both positive integers), where A represents the decoding component, for example, A=0 represents the luma component, and A=1 represents the chroma component. B represents the interpolation direction, for example, B=0 represents the x-direction, and B=1 represents the y-direction. Flag2[A][B] represents the index information of the target interpolation filter for the A component in the interpolation direction B. For example, the first indication information and its meaning are shown in Table 5.

[0338] In one example, the first indication information indicates the interpolation target, the coding component, and the target interpolation filter. In this case, the first indication information includes the parsing flag array Flag3[A][D] (A and D are both positive integers), where A represents the decoding component, for example, A=0 represents the luma component, and A=1 represents the chroma component. D represents the interpolation target, for example, D=0 represents the interpolation target to determine the predicted value, and D=1 represents the interpolation target to determine the gradient. Flag3[A][D] represents the index information of the target interpolation filter for the A component of the interpolation target D. Exemplarily, the first indication information and its meaning are shown in Table 6.

[0339] In this implementation, the encoding end determines multiple candidate interpolation filters with different numbers of taps based on at least one of the prediction mode, interpolation direction, encoding component, and interpolation target corresponding to the area to be encoded, and then selects the target interpolation filter from these multiple candidate interpolation filters with different numbers of taps.

[0340] In some embodiments, the encoding end may determine a target interpolation filter corresponding to the to-be-encoded region based on a default interpolation filter corresponding to at least one of a prediction mode, an interpolation direction, an encoding component, and an interpolation target corresponding to the to-be-encoded region.

[0341] In Example 1, the encoding end determines a target interpolation filter based on a prediction mode, an encoding component, and an interpolation target.

[0342] For example, both the encoding end and the decoding end default to using a specific N1-tap interpolation filter for luminance and / or chrominance for j1 interpolation targets in a certain i prediction modes, and using a specific N2-tap interpolation filter for luminance and / or chrominance for j2 interpolation contents in the remaining Ii prediction modes, where i, j1, j2, N1 and N2 are all integers, and N1≠N2.

[0343] For example, if the prediction mode of the area to be coded is the bidirectional optical flow prediction mode, the coding component is the luminance component, and the interpolation target is to determine the gradients of the forward prediction block and the backward prediction block of the area to be coded in the x-direction and the y-direction, then the n1-tap interpolation filter is determined as the target interpolation filter;

[0344] If the prediction mode of the to-be-encoded area is the bidirectional optical flow prediction mode, the encoding component is the luminance component, and the interpolation target is to determine the prediction value of the to-be-encoded area in the x-direction and the y-direction, then the n2-tap interpolation filter is determined as the target interpolation filter;

[0345] If the prediction mode of the area to be coded is the bidirectional optical flow prediction mode, the coding component is the chrominance component, and the interpolation target is to determine the prediction value of the area to be coded in the x direction and the y direction, the interpolation filter with n3 taps is determined as the target interpolation filter;

[0346] If the prediction mode of the area to be encoded is a non-bidirectional optical flow prediction mode, the encoding component is a luminance component, and the interpolation target is to determine the prediction value of the area to be encoded in the x direction and the y direction, the n4-tap interpolation filter is determined as the target interpolation filter;

[0347] If the prediction mode of the area to be coded is a non-bidirectional optical flow prediction mode, the coding component is a chrominance component, and the interpolation target is to determine the prediction value of the area to be coded in the x direction and the y direction, the interpolation filter with n5 taps is determined as the target interpolation filter;

[0348] Among them, n1, n2, n3, n4 and n5 are all positive integers.

[0349] The embodiment of the present application does not limit the specific values ​​of the above n1, n2, n3, n4 and n5.

[0350] In a possible implementation, the above n1 is not equal to the above n2.

[0351] In one example, n1=8, n2=12, n3=6, n4=12, and n5=6.

[0352] As can be seen from Table 7 above, in the BIO mode of the present embodiment, the number of taps of the interpolation filter used when calculating the gradient is different from the number of taps of the interpolation filter used when calculating the pixel prediction value. For example, an 8-tap interpolation filter is used to calculate the gradient, and a 12-tap interpolation filter is used for pixel prediction value interpolation calculation.

[0353] Example 2: The encoding end determines a target interpolation filter based on the prediction mode and the encoding component.

[0354] For example, both the encoding end and the encoding end default to using only a specific N1-tap interpolation filter for luminance and / or chrominance in a certain i prediction modes, and using a certain N2-tap interpolation filter for luminance and / or chrominance in the remaining Ii prediction modes (i, N1 and N2 are all integers, N1≠N2).

[0355] For example, if the prediction mode of the area to be encoded is the bidirectional optical flow prediction mode and the encoding component is the luminance component, the interpolation filter with the m1 tap is determined as the target interpolation filter. At this time, the encoder uses the interpolation filter with the m1 tap to perform gradient calculation and prediction value interpolation.

[0356] If the prediction mode of the area to be encoded is the bidirectional optical flow prediction mode and the encoding component is the chrominance component, the m2-tap interpolation filter is determined as the target interpolation filter. At this time, the encoder uses the m2-tap interpolation filter to perform prediction value interpolation calculations in the x-direction and y-direction.

[0357] If the prediction mode of the area to be encoded is a non-bidirectional optical flow prediction mode and the encoding component is a luminance component, the interpolation filter with m3 taps is determined as the target interpolation filter. At this time, the encoder uses the interpolation filter with m3 taps to perform interpolation calculations on the predicted values ​​in the x and y directions.

[0358] If the prediction mode of the area to be encoded is a non-bidirectional optical flow prediction mode and the encoding component is a chrominance component, the interpolation filter with m4 taps is determined as the target interpolation filter. At this time, the encoder uses the interpolation filter with m4 taps to perform prediction value interpolation calculations in the x-direction and y-direction.

[0359] Among them, m1, m2, m3 and m4 are all positive integers.

[0360] The embodiment of the present application does not limit the specific values ​​of the above m1, m2, m3 and m4.

[0361] In one example, m1=8, m2=6, m3=12, and m4=6.

[0362] Example 3: The encoding end determines the target interpolation filter based on the interpolation direction and the encoding component.

[0363] For example, the encoding end and the encoding end default to using an interpolation filter with N1 taps for interpolation of the predicted values, gradients, and other contents of luminance and / or chrominance in the x direction, and an interpolation filter with N2 taps for interpolation of the predicted values, gradients, and other contents of luminance and / or chrominance in the y direction (both N1 and N2 are positive integers, and N1≠N2).

[0364] For example, if the interpolation direction of the area to be encoded is the x direction and the encoding component is the luminance component, the interpolation filter with k1 taps is determined as the target interpolation filter. At this time, the encoder uses the interpolation filter with k1 taps to perform gradient calculation and prediction value interpolation.

[0365] If the interpolation direction of the area to be encoded is the x direction and the encoding component is the chrominance component, the k2-tap interpolation filter is determined as the target interpolation filter. At this time, the encoder uses the k2-tap interpolation filter to perform gradient calculation and prediction value interpolation;

[0366] If the interpolation direction of the area to be encoded is the y direction and the encoding component is the luminance component, the k3-tap interpolation filter is determined as the target interpolation filter. At this time, the encoder uses the k3-tap interpolation filter to perform gradient calculation and prediction value interpolation;

[0367] If the interpolation direction of the area to be encoded is the y direction and the encoding component is the chrominance component, the k4-tap interpolation filter is determined as the target interpolation filter. At this time, the encoder uses the k4-tap interpolation filter to perform gradient calculation and prediction value interpolation;

[0368] Wherein, k1, k2, k3 and k4 are all positive integers.

[0369] The embodiment of the present application does not limit the specific values ​​of the above k1, k2, k3 and k4.

[0370] In one example, k1=12, k2=6, k3=8, and k4=4.

[0371] In Example 4, the encoding end determines a target interpolation filter based on the encoding components and the size of the area to be encoded.

[0372] For example, both the encoding end and the encoding end use an N1-tap interpolation filter by default within a certain block size range, and use an N2-tap interpolation filter within other block size ranges (N1 and N2 are both positive integers, and N1≠N2).

[0373] For example, if the coding component of the to-be-coded region is a luminance component, and the scale of the to-be-coded region is w<=W1 or h<=H1, the interpolation filter with tap a1 is determined as the target interpolation filter. At this time, the encoder uses the interpolation filter with tap a1 to perform gradient calculation and prediction value interpolation.

[0374] If the coding component of the to-be-coded region is the luminance component, and the scale of the to-be-coded region is w>W1 and h>H1, the interpolation filter with a2 tap is determined as the target interpolation filter. At this time, the encoder uses the interpolation filter with a2 tap to perform gradient calculation and prediction value interpolation.

[0375] If the coding component of the to-be-coded region is a chrominance component and the scale of the to-be-coded region is w<=W2 or h<=H2, the interpolation filter with a3 taps is determined as the target interpolation filter. At this time, the encoder uses the interpolation filter with a3 taps to perform gradient calculation and prediction value interpolation.

[0376] If the coding component of the to-be-coded region is a chrominance component, and the scale of the to-be-coded region is w>W2 and h>H2, the a4-tap interpolation filter is determined as the target interpolation filter. At this time, the encoder uses the a4-tap interpolation filter for gradient calculation and prediction value interpolation.

[0377] Wherein, a1, a2, a3 and a4 are all positive integers.

[0378] The embodiment of the present application does not limit the specific values ​​of the above a1, a2, a3 and a4.

[0379] In one example, a1=8, a2=12, a3=4, and a4=6.

[0380] After the encoder determines the target interpolation filter based on the above steps, it executes the following step S203.

[0381] S203 . Interpolate the reference area based on the target interpolation filter to determine a prediction value of the area to be encoded.

[0382] Based on the above steps, the encoder determines the target interpolation filter corresponding to the area to be encoded, and then uses the target interpolation filter to interpolate the reference area of ​​the area to be encoded to determine the predicted value of the area to be encoded.

[0383] In one example, if the interpolation target is to interpolate pixel prediction values, the encoding end uses the target interpolation filter to interpolate the reference area to obtain an interpolated prediction area, and then searches in the interpolated prediction area to obtain a prediction block of the area to be encoded.

[0384] In one example, if the prediction mode corresponding to the region to be encoded in the embodiment of the present application is BIO, the encoder first uses a bidirectional prediction mode to determine a forward reference block and a backward reference block for the region to be encoded. Next, the target interpolation filter is sampled, and the forward and backward reference blocks are interpolated to determine the gradients of the forward and backward reference blocks. The motion vectors corresponding to the forward and backward reference blocks are then corrected based on the gradients to obtain the predicted value of the region to be encoded.

[0385] The video decoding method provided in the embodiment of the present application obtains the coding information of the area to be encoded and the reference area corresponding to the area to be encoded, wherein the coding information includes at least one of a prediction mode, an interpolation direction, a coding component, and an interpolation target; based on the coding information, the target interpolation filter corresponding to the area to be encoded is determined, and then based on the target interpolation filter, the reference area is interpolated to determine the predicted value of the area to be encoded. That is to say, in the embodiment of the present application, the encoding end adaptively selects the target interpolation filter from the interpolation filters with different tap numbers based on the coding information of the area to be encoded and decoded, such as the prediction mode, interpolation direction, coding component, interpolation target and other information, thereby improving the accuracy of the selection of the target interpolation filter. In this way, based on the accurately selected target interpolation filter, the reference area of ​​the area to be encoded is interpolated to determine the predicted value of the area to be encoded, which can improve the prediction effect of the area to be encoded, thereby improving the performance of video encoding.

[0386] The above describes in detail an embodiment of the audio encoding and decoding method of the present application in conjunction with Figures 5 to 8 , and the following describes in detail an embodiment of the device of the present application in conjunction with Figures 9 to 10 .

[0387] Figure 9 is a schematic block diagram of a video decoding apparatus according to an embodiment of the present application. The apparatus 10 can be applied to a decoding device or a decoder.

[0388] As shown in FIG9 , the video decoding apparatus 10 includes:

[0389] A decoding unit 11 is configured to decode a bitstream to obtain decoding information of a to-be-decoded region, wherein the decoding information includes at least one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target;

[0390] A determination unit 12, configured to determine a target interpolation filter corresponding to the to-be-decoded area based on the decoded information;

[0391] The interpolation unit 13 is configured to determine a reference area of ​​the area to be decoded, and interpolate the reference area based on the target interpolation filter to determine a prediction value of the area to be decoded.

[0392] In some embodiments, the interpolation unit 13 is specifically used to determine multiple candidate interpolation filters based on the decoding information; decode the code stream to obtain first indication information of the target interpolation filter, where the first indication information is used to indicate the index information of the target interpolation filter in the multiple candidate interpolation filters; and select the target interpolation filter from the multiple candidate interpolation filters based on the first indication information.

[0393] In some embodiments, the interpolation unit 13 is specifically used to obtain preset candidate interpolation filters corresponding to at least one of the prediction mode, the interpolation direction, the decoding component and the interpolation target; and determine the multiple candidate interpolation filters based on the preset candidate interpolation filters.

[0394] In some embodiments, the interpolation unit 13 is specifically configured to remove duplicate interpolation filters from the preset candidate interpolation filters to obtain the multiple candidate interpolation filters.

[0395] In some embodiments, the interpolation unit 13 is specifically used to obtain a selection strategy corresponding to at least one of the prediction mode, the interpolation direction, the decoding component and the interpolation target, and based on the selection strategy, select the multiple candidate interpolation filters from the preset M interpolation filters, where M is a positive integer greater than 1; or, obtain N candidate combinations consisting of the M interpolation filters, and decode the code stream to obtain second indication information, and based on the second indication information, select the target candidate combination from the N candidate combinations, and then determine the interpolation filters included in the target candidate combination as the multiple candidate interpolation filters, wherein each candidate combination includes two or more interpolation filters from the M interpolation filters, and the second indication information is used to indicate the target candidate combination, where N is a positive integer greater than 1.

[0396] In some embodiments, the interpolation filters included in different candidate combinations among the N candidate combinations are randomly distributed, or the number of interpolation filters included in each candidate combination among the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination increases or decreases in sequence.

[0397] In some embodiments, if the first indication information includes a flag array or multiple flags, and the flag array or multiple flags are used to indicate the decoding information and the target interpolation filter, the decoding unit 11 is also used to decode the code stream to obtain the first indication information; based on the first indication information, obtain the decoding information of the area to be decoded.

[0398] In some embodiments, the interpolation unit 13 is specifically used to determine the target interpolation filter based on the prediction mode, the decoded component and the interpolation target; or, determine the target interpolation filter based on the prediction mode and the decoded component; or, determine the target interpolation filter based on the interpolation direction and the decoded component; or, determine the target interpolation filter according to the decoded component and the size of the area to be decoded.

[0399] In some embodiments, the interpolation unit 13 is specifically used to determine the interpolation filter with n1 taps as the target interpolation filter if the prediction mode is the bidirectional optical flow prediction mode, the decoding component is the luminance component, and the interpolation target is to determine the gradients of the forward prediction block and the backward prediction block of the area to be decoded in the x direction and the y direction; if the prediction mode is the bidirectional optical flow prediction mode, the decoding component is the luminance component, and the interpolation target is to determine the predicted value of the area to be decoded in the x direction and the y direction, then the interpolation filter with n2 taps is determined as the target interpolation filter; if the prediction mode is the bidirectional optical flow prediction mode, the decoding component is the chrominance component, and the interpolation target is to determine When the predicted values ​​of the area to be decoded in the x and y directions are determined, the interpolation filter with the n3 tap is determined as the target interpolation filter; if the prediction mode is a non-bidirectional optical flow prediction mode, the decoding component is a luminance component, and the interpolation target is to determine the predicted values ​​of the area to be decoded in the x and y directions, the interpolation filter with the n4 tap is determined as the target interpolation filter; if the prediction mode is a non-bidirectional optical flow prediction mode, the decoding component is a chrominance component, and the interpolation target is to determine the predicted values ​​of the area to be decoded in the x and y directions, the interpolation filter with the n5 tap is determined as the target interpolation filter; wherein, n1, n2, n3, n4 and n5 are all positive integers.

[0400] In some embodiments, n1 is not equal to n2.

[0401] In some embodiments, n1=8, n2=12, n3=6, n4=12, and n5=6.

[0402] It should be understood that the apparatus embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, no further description is given here. Specifically, the apparatus shown in FIG9 can perform the above-mentioned video decoding method embodiment, and the aforementioned and other operations and / or functions of each module in the apparatus are respectively for implementing the above-mentioned method embodiment. For the sake of brevity, no further description is given here.

[0403] Figure 10 is a schematic block diagram of a video encoding apparatus according to an embodiment of the present application. The apparatus 20 can be applied to an encoding device or an encoder.

[0404] As shown in FIG10 , the video encoding apparatus 20 includes:

[0405] An acquisition unit 21 is configured to acquire encoding information of a region to be encoded and a reference region corresponding to the region to be encoded, wherein the encoding information includes at least one of a prediction mode, an interpolation direction, an encoding component, and an interpolation target;

[0406] a determining unit 22, configured to determine a target interpolation filter corresponding to the to-be-encoded area based on the encoding information;

[0407] The interpolation unit 23 is configured to perform interpolation filtering on the reference area based on the target interpolation filter to determine a prediction value of the area to be encoded.

[0408] In some embodiments, the interpolation unit 23 is specifically configured to determine a plurality of candidate interpolation filters based on the encoding information; and select the target interpolation filter from the plurality of candidate interpolation filters.

[0409] In some embodiments, the interpolation unit 23 is specifically used to obtain the preset candidate interpolation filters corresponding to at least one of the prediction mode, the interpolation direction, the coding component and the interpolation target; and determine the multiple candidate interpolation filters based on the preset candidate interpolation filters.

[0410] In some embodiments, the interpolation unit 23 is specifically configured to remove duplicate interpolation filters from the prediction candidate interpolation filters to obtain the multiple candidate interpolation filters.

[0411] In some embodiments, the interpolation unit 23 is specifically used to obtain a selection strategy corresponding to at least one of the prediction mode, the interpolation direction, the coding component and the interpolation target; based on the selection strategy, the multiple candidate interpolation filters are selected from the preset M interpolation filters, where M is a positive integer greater than 1.

[0412] In some embodiments, the interpolation unit 23 is specifically used to obtain N candidate combinations consisting of the M interpolation filters, each candidate combination includes two or more interpolation filters from the M interpolation filters, and N is a positive integer greater than 1; based on the selection strategy, a target candidate combination is selected from the N candidate combinations, and the interpolation filters included in the target candidate combination are determined as the multiple candidate interpolation filters.

[0413] In some embodiments, the interpolation filters included in different candidate combinations among the N candidate combinations are randomly distributed, or the number of interpolation filters included in each candidate combination among the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination increases or decreases in sequence.

[0414] In some embodiments, the interpolation unit 23 is specifically used to interpolate each of the multiple candidate interpolation filters using the candidate interpolation filter based on the reference area to obtain an interpolated prediction area; determine the rate-distortion cost corresponding to the candidate interpolation filter based on the interpolated prediction area; and select the target interpolation filter from the multiple candidate interpolation filters based on the rate-distortion cost corresponding to each of the multiple candidate interpolation filters.

[0415] In some embodiments, the interpolation unit 23 is further configured to write first indication information into the bitstream, where the first indication information is used to indicate index information of the target interpolation filter among the multiple candidate interpolation filters.

[0416] In some embodiments, the first indication information includes a flag array or multiple flags, and the flag array or multiple flags are used to indicate the encoding information and the target interpolation filter.

[0417] In some embodiments, the interpolation unit 23 is further configured to write second indication information into the bitstream, where the second indication information is used to indicate the target candidate combination.

[0418] In some embodiments, the interpolation unit 23 is specifically used to determine the target interpolation filter based on the prediction mode, the coding component and the interpolation target; or, to determine the target interpolation filter based on the prediction mode and the coding component; or, to determine the target interpolation filter based on the interpolation direction and the coding component; or, to determine the target interpolation filter according to the coding component and the size of the area to be encoded.

[0419] In some embodiments, the interpolation unit 23 is further used to determine the interpolation filter with n1 taps as the target interpolation filter if the prediction mode is the bidirectional optical flow prediction mode, the coding component is the luminance component, and the interpolation target is to determine the gradients of the forward prediction block and the backward prediction block of the to-be-encoded area in the x-direction and the y-direction; if the prediction mode is the bidirectional optical flow prediction mode, the coding component is the luminance component, and the interpolation target is to determine the predicted values ​​of the to-be-encoded area in the x-direction and the y-direction, then determine the interpolation filter with n2 taps as the target interpolation filter; if the prediction mode is the bidirectional optical flow prediction mode, When the coding component is a chroma component and the interpolation target is to determine the predicted values ​​of the area to be encoded in the x and y directions, the interpolation filter with n3 taps is determined as the target interpolation filter; if the prediction mode is a non-bidirectional optical flow prediction mode, the coding component is a luminance component, and the interpolation target is to determine the predicted values ​​of the area to be encoded in the x and y directions, the interpolation filter with n4 taps is determined as the target interpolation filter; if the prediction mode is a non-bidirectional optical flow prediction mode, the coding component is a chroma component, and the interpolation target is to determine the predicted values ​​of the area to be encoded in the x and y directions, the interpolation filter with n5 taps is determined as the target interpolation filter; wherein, n1, n2, n3, n4 and n5 are all positive integers.

[0420] In some embodiments, n1 is not equal to n2.

[0421] In some embodiments, n1=8, n2=12, n3=6, n4=12, and n5=6.

[0422] It should be understood that the apparatus embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, no further description is given here. Specifically, the apparatus shown in FIG10 can perform the above-mentioned video encoding method embodiment, and the aforementioned and other operations and / or functions of each module in the apparatus are respectively for implementing the above-mentioned method embodiment. For the sake of brevity, no further description is given here.

[0423] The apparatus of the embodiment of the present application is described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.

[0424] FIG11 is a schematic block diagram of an electronic device provided in an embodiment of the present application. The electronic device in FIG11 may be the above-mentioned encoding device (or encoder) or a decoding device (or decoder).

[0425] As shown in FIG11 , the electronic device 30 may include:

[0426] The memory 31 and the processor 32 are configured to store a computer program 33 and transmit the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiment of the present application.

[0427] For example, the processor 32 may be configured to execute the steps of the method 200 according to the instructions in the computer program 33 .

[0428] In some embodiments of the present application, the processor 32 may include but is not limited to:

[0429] General-purpose processor, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0430] In some embodiments of the present application, the memory 31 includes but is not limited to:

[0431] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0432] In some embodiments of the present application, the computer program 33 may be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to implement the method for recording a page provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 33 in the electronic device.

[0433] As shown in FIG11 , the electronic device 30 may further include:

[0434] The transceiver 34 may be connected to the processor 32 or the memory 31 .

[0435] The processor 32 may control the transceiver 34 to communicate with other devices. Specifically, the processor 32 may send information or data to other devices or receive information or data sent by other devices. The transceiver 34 may include a transmitter and a receiver. The transceiver 34 may further include one or more antennas.

[0436] It should be understood that the various components in the computing device 30 are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0437] According to one aspect of the present application, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by a computer, the computer is enabled to perform the method of the above-mentioned method embodiment. Alternatively, the present application also provides a computer program product containing instructions. When the computer is executed by the instructions, the computer is enabled to perform the method of the above-mentioned method embodiment.

[0438] According to another aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of the above-described method embodiment.

[0439] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0440] Finally, it should be noted that the above content is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A video decoding method, executed by a processor, characterized in that: The method comprises: Decoding the bitstream to obtain decoding information of the area to be decoded, wherein the decoding information includes at least one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target; Based on the decoded information, determining a target interpolation filter corresponding to the area to be decoded; A reference area of ​​the area to be decoded is determined, and based on the target interpolation filter, the reference area is interpolated to determine a prediction value of the area to be decoded.

2. The method according to claim 1, characterized in that The step of determining a target interpolation filter corresponding to the to-be-decoded area based on the decoded information includes: Based on the decoded information, determining a plurality of candidate interpolation filters; Decoding the bitstream to obtain first indication information of the target interpolation filter, where the first indication information is used to indicate index information of the target interpolation filter among the multiple candidate interpolation filters; Based on the first indication information, the target interpolation filter is selected from the multiple candidate interpolation filters.

3. The method according to claim 2, characterized in that The step of determining a plurality of candidate interpolation filters based on the decoded information comprises: Obtaining preset candidate interpolation filters corresponding to at least one of the prediction mode, the interpolation direction, the decoded component and the interpolation target; Based on the preset candidate interpolation filters, the plurality of candidate interpolation filters are determined.

4. The method according to claim 3, characterized in that The determining the plurality of candidate interpolation filters based on the preset candidate interpolation filters comprises: The repeated interpolation filters in the preset candidate interpolation filters are eliminated to obtain the multiple candidate interpolation filters.

5. The method according to any one of claims 2 to 4, characterized in that: The step of determining a plurality of candidate interpolation filters based on the decoded information comprises: Obtaining selection strategies corresponding to at least one of the prediction mode, the interpolation direction, the decoding component, and the interpolation target, and selecting the plurality of candidate interpolation filters from preset M interpolation filters based on the selection strategies, where M is a positive integer greater than 1; or Obtain N candidate combinations consisting of the M interpolation filters, and decode the bitstream to obtain second indication information, and based on the second indication information, select the target candidate combination from the N candidate combinations, and then determine the interpolation filters included in the target candidate combination as the multiple candidate interpolation filters, wherein each candidate combination includes two or more interpolation filters from the M interpolation filters, the second indication information is used to indicate the target candidate combination, and N is a positive integer greater than 1.

6. The method according to claim 5, characterized in that The interpolation filters included in different candidate combinations among the N candidate combinations are randomly distributed, or the number of interpolation filters included in each candidate combination among the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination increases or decreases in sequence.

7. The method according to any one of claims 2 to 6, characterized in that: If the first indication information includes a flag bit array or multiple flag bits, and the flag bit array or multiple flag bits are used to indicate the decoding information and the target interpolation filter, the decoding bit stream is obtained to obtain the decoding information of the area to be decoded, including: Decoding the bit stream to obtain the first indication information; Based on the first indication information, decoding information of the area to be decoded is obtained.

8. The method according to any one of claims 1 to 7, characterized in that: The step of determining a target interpolation filter corresponding to the to-be-decoded area based on the decoded information includes: determining the target interpolation filter based on the prediction mode, the decoded component and the interpolation target; or, determining the target interpolation filter based on the prediction mode and the decoded component; or, determining the target interpolation filter based on the interpolation direction and the decoded component; or, The target interpolation filter is determined according to the decoded components and the size of the area to be decoded.

9. The method according to any one of claims 1 to 8, characterized in that: The determining the target interpolation filter based on the prediction mode, the decoded component and the interpolation target comprises: If the prediction mode is a bidirectional optical flow prediction mode, the decoded component is a brightness component, and the interpolation target is to determine the gradients of the forward prediction block and the backward prediction block of the area to be decoded in the x direction and the y direction, an n1-tap interpolation filter is determined as the target interpolation filter; If the prediction mode is the bidirectional optical flow prediction mode, the decoded component is the brightness component, and the interpolation target To determine the predicted values ​​of the to-be-decoded area in the x direction and the y direction, an n2-tap interpolation filter is determined as the target interpolation filter; If the prediction mode is the bidirectional optical flow prediction mode, the decoded component is the chrominance component, and the interpolation target is to determine the prediction values ​​of the to-be-decoded area in the x direction and the y direction, then an n3-tap interpolation filter is determined as the target interpolation filter; If the prediction mode is a non-bidirectional optical flow prediction mode, the decoded component is a brightness component, and the interpolation target is to determine the prediction values ​​of the to-be-decoded area in the x direction and the y direction, an n4-tap interpolation filter is determined as the target interpolation filter; If the prediction mode is a non-bidirectional optical flow prediction mode, the decoded component is a chrominance component, and the interpolation target is to determine the prediction value of the to-be-decoded area in the x direction and the y direction, then an interpolation filter with n5 taps is determined as the target interpolation filter; Wherein, n1, n2, n3, n4 and n5 are all positive integers.

10. The method according to claim 9, characterized in that The n1 is not equal to n2.

11. The method according to claim 9 or 10, characterized in that: The n1=8, the n2=12, the n3=6, the n4=12, and the n5=6.

12. A video encoding method, executed by a processor, characterized in that: The method comprises: Acquire encoding information of a to-be-encoded region and a reference region corresponding to the to-be-encoded region, wherein the encoding information includes at least one of a prediction mode, an interpolation direction, an encoding component, and an interpolation target; Based on the encoding information, determining a target interpolation filter corresponding to the to-be-encoded area; Based on the target interpolation filter, interpolation filtering is performed on the reference area to determine a prediction value of the area to be encoded.

13. The method according to claim 12, characterized in that The determining, based on the encoding information, a target interpolation filter corresponding to the to-be-encoded area includes: Based on the encoding information, determining a plurality of candidate interpolation filters; The target interpolation filter is selected from the plurality of candidate interpolation filters.

14. The method according to claim 13, characterized in that The step of determining a plurality of candidate interpolation filters based on the encoding information comprises: Obtaining preset candidate interpolation filters corresponding to at least one of the prediction mode, the interpolation direction, the coding component and the interpolation target; Based on the preset candidate interpolation filters, the plurality of candidate interpolation filters are determined.

15. The method according to claim 13 or 14, characterized in that The step of determining a plurality of candidate interpolation filters based on the encoding information comprises: Obtaining a selection strategy corresponding to at least one of the prediction mode, the interpolation direction, the coding component, and the interpolation target; Based on the selection strategy, the plurality of candidate interpolation filters are selected from M preset interpolation filters, where M is a positive integer greater than 1.

16. The method according to any one of claims 12 to 15, characterized in that The determining, based on the encoding information, a target interpolation filter corresponding to the to-be-encoded area includes: Determining the target interpolation filter based on the prediction mode, the coding component and the interpolation target; or, Determining the target interpolation filter based on the prediction mode and the coding component; or, Determining the target interpolation filter based on the interpolation direction and the coding component; or, The target interpolation filter is determined according to the encoding component and the size of the area to be encoded.

17. A video decoding device, characterized in that: The device comprises: A decoding unit, configured to decode the bitstream to obtain decoding information of the area to be decoded, wherein the decoding information includes at least one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target; A determination unit, configured to determine a target interpolation filter corresponding to the to-be-decoded area based on the decoding information; The interpolation unit is used to determine a reference area of ​​the area to be decoded, and interpolate the reference area based on the target interpolation filter to determine a prediction value of the area to be decoded.

18. A video encoding device, characterized in that: The device comprises: an acquisition unit, configured to acquire encoding information of a to-be-encoded region and a reference region corresponding to the to-be-encoded region, wherein the encoding information includes at least one of a prediction mode, an interpolation direction, an encoding component, and an interpolation target; A determination unit, configured to determine a target interpolation filter corresponding to the to-be-encoded area based on the encoding information; An interpolation unit is used to perform interpolation filtering on the reference area based on the target interpolation filter to determine a prediction value of the area to be encoded.

19. An electronic device comprising a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the computer program to implement the method according to any one of claims 1 to 11 or 12 to 16.

20. A computer-readable storage medium, characterized in that: For storing computer programs; The computer program enables a computer to execute the method according to any one of claims 1 to 11 or 12 to 16 above.

21. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 11 or 12 to 16.

Citation Information

Patent Citations

  • Method and apparatus for coding video data

    CN108377393A

  • Image compression method based on direction lifting wavelet and improved SPIHT under IoT (Internet of Things)

    CN108810534A

  • Bidirectional optical flow based video coding and decoding

    CN113661708A

  • Method and apparatus for filtering

    CN116320498A

  • Prediction image generation device, moving image decoding device, and moving image encoding device

    JP2020096279A