Video coding and decoding method and device, equipment and storage medium
By adaptively selecting the interpolation filter based on the decoding information in video encoding and decoding, the problem of inaccurate selection of interpolation filters in the prior art is solved, and the video prediction effect and codec performance are improved.
Patent Information
- Application Number
- CN202311694332.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-09
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art does not consider the encoding and decoding information when selecting the interpolation filter, resulting in inaccurate selection of the interpolation filter, which in turn affects the video prediction effect.
The decoding information of the area to be decoded is obtained by decoding the code stream, including prediction mode, interpolation direction, decoding component and interpolation target, and the target interpolation filter is adaptively selected based on these information.
Improve the accuracy of the selection of interpolation filters, improve the prediction effect of video, and thus improve the performance of video encoding and decoding.
Smart Images

Figure CN120128703A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to a video encoding and decoding method, apparatus, device, and storage medium. Background Art
[0002] Digital video technology can be incorporated into various video devices, such as digital televisions, smartphones, computers, e-readers, or video players. With the development of video technology, the amount of data included in video data is large. To facilitate the transmission of video data, video devices perform video compression technology to make the video data more effectively transmitted or stored.
[0003] Due to the temporal redundancy in video, the temporal redundancy between video frames is eliminated through inter-frame prediction to improve the compression efficiency. In inter-frame prediction, in some cases, an interpolation filter is required to interpolate the reference image. However, currently, when determining the interpolation filter, the encoding and decoding information is not considered, resulting in inaccurate selection of the interpolation filter and poor prediction effect. Summary of the Invention
[0004] The present application provides a video encoding and decoding method, apparatus, device, and storage medium for improving the accuracy of selecting an interpolation filter, thereby enhancing the prediction effect of the video.
[0005] In a first aspect, the present application provides a video decoding method applied to a decoder, including:
[0006] Decoding a bitstream to obtain decoding information of a region to be decoded, where the decoding information includes at least one of a prediction mode, an interpolation direction, a decoded component, and an interpolation target;
[0007] Based on the decoding information, determining a target interpolation filter corresponding to the region to be decoded;
[0008] Determining a reference region of the region to be decoded, and interpolating the reference region based on the target interpolation filter to determine a predicted value of the region to be decoded.
[0009] In a second aspect, embodiments of the present application provide a video encoding method applied to an encoder, including:
[0010] Obtaining encoding information of a region to be encoded, and a reference region corresponding to the region to be encoded, where the encoding information includes at least one of a prediction mode, an interpolation direction, an encoded component, and an interpolation target;
[0011] Based on the encoding information, determining a target interpolation filter corresponding to the region to be encoded;
[0012] Performing interpolation filtering on the reference region based on the target interpolation filter to determine the predicted value of the region to be encoded.
[0013] In a third aspect, the present application provides a video decoding device, including:
[0014] A decoding unit configured to decode a bitstream to obtain decoding information of a region to be decoded, where the decoding information includes at least one of a prediction mode, an interpolation direction, a decoded component, and an interpolation target;
[0015] A determination unit configured to determine a target interpolation filter corresponding to the region to be decoded based on the decoding information;
[0016] An interpolation unit configured to determine a reference region of the region to be decoded and perform interpolation on the reference region based on the target interpolation filter to determine the predicted value of the region to be decoded.
[0017] In a fourth aspect, the present application provides a video encoding device, including:
[0018] An acquisition unit configured to acquire encoding information of a region to be encoded and a reference region corresponding to the region to be encoded, where the encoding information includes at least one of a prediction mode, an interpolation direction, an encoded component, and an interpolation target;
[0019] A determination unit configured to determine a target interpolation filter corresponding to the region to be encoded based on the encoding information;
[0020] An interpolation unit configured to perform interpolation filtering on the reference region based on the target interpolation filter to determine the predicted value of the region to be encoded.
[0021] In a fifth aspect, a video decoder is provided, including a processor and a memory. The memory is configured to store a computer program, and the processor is configured to call and run the computer program stored in the memory to execute the method in the first aspect or its various implementation manners described above.
[0022] In a sixth aspect, a video encoder is provided, including a processor and a memory. The memory is configured to store a computer program, and the processor is configured to call and run the computer program stored in the memory to execute the method in the second aspect or its various implementation manners described above.
[0023] In a seventh aspect, a video coding and decoding system is provided, including a video encoder and a video decoder. The video decoder is configured to execute the method in the first aspect or its various implementation manners described above, and the video encoder is configured to execute the method in the second aspect or its various implementation manners described above.
[0024] In an eighth aspect, a chip is provided for implementing the method in any one of the first aspect to the second aspect or its various implementation manners. Specifically, the chip includes: a processor for calling and running a computer program from a memory, so that a device equipped with the chip executes the method in any one of the first aspect to the second aspect or its various implementation manners.
[0025] In a ninth aspect, a computer-readable storage medium is provided for storing a computer program, and the computer program causes a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners.
[0026] In a tenth aspect, a computer program product is provided, including computer program instructions, and the computer program instructions cause a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners.
[0027] In an eleventh aspect, a computer program is provided, which when running on a computer, causes the computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners.
[0028] Based on the above technical solutions, in this application, by decoding a bitstream, decoding information of a region to be decoded is obtained, and the decoding information includes at least one of a prediction mode, an interpolation direction, a decoded component, and an interpolation target; based on the decoding information, a target interpolation filter corresponding to the region to be decoded is determined; a reference region of the region to be decoded is determined, and based on the target interpolation filter, interpolation is performed on the reference region to determine a predicted value of the region to be decoded. That is to say, in the embodiments of this application, at the decoding end, based on the encoding and decoding information of the region to be encoded and decoded, such as information such as a prediction mode, an interpolation direction, a decoded component, and an interpolation target, a target interpolation filter is adaptively selected from interpolation filters with different numbers of taps, improving the selection accuracy of the target interpolation filter. In this way, based on the accurately selected target interpolation filter, when interpolating the reference region of the region to be encoded and decoded to determine the predicted value of the region to be encoded and decoded, the prediction effect of the region to be encoded and decoded can be improved, thereby enhancing the performance of video encoding and decoding. Description of the Drawings
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0030] Figure 1 It is a schematic block diagram of a video encoding and decoding system related to the embodiments of this application;
[0031] Figure 2 It is a schematic block diagram of a video encoder according to an embodiment of the present application;
[0032] Figure 3 It is a schematic block diagram of a video decoder according to an embodiment of the present application;
[0033] Figure 4A It is a schematic diagram of inter-frame prediction;
[0034] Figure 4B It is a schematic diagram of pixel interpolation;
[0035] Figure 5 It is a schematic flowchart of a video decoding method provided by an embodiment of the present application;
[0036] Figure 6 It is a schematic flowchart of a video encoding method provided by an embodiment of the present application;
[0037] Figure 7 It is a schematic diagram for determining a plurality of candidate interpolation filters;
[0038] Figure 8 It is a schematic diagram for determining a target interpolation filter based on rate distortion cost;
[0039] Figure 9 It is a schematic block diagram of a video decoding device provided by an embodiment of the present application;
[0040] Figure 10 It is a schematic block diagram of a video encoding device provided by an embodiment of the present application;
[0041] Figure 11 It is a schematic block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0042] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0043] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In the embodiments of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In the description of this application, unless otherwise specified, "a plurality of" means two or more than two.
[0044] This application can be applied to the fields of image coding and decoding, video coding and decoding, hardware video coding and decoding, dedicated circuit video coding and decoding, real-time video coding and decoding, etc. For example, the solution of this application can be combined with audio video coding standards (AVS for short), such as the H.264 / audio video coding (AVC for short) standard, the H.265 / high efficiency video coding (HEVC for short) standard, and the H.266 / versatile video coding (VVC for short) standard. Or, the solution of this application can be combined with other proprietary or industry standards for operation, and the standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), including scalable video coding (SVC) and multi-view video coding (MVC) extensions. It should be understood that the technology of this application is not limited to any specific coding and decoding standard or technology.
[0045] For the sake of easy understanding, first combine Figure 1 to introduce the video coding and decoding system involved in the embodiments of this application.
[0046] Figure 1Schematic block diagram of a video encoding and decoding system according to an embodiment of the present application. It should be noted that Figure 1 This is only an example. The video encoding and decoding system of the embodiments of the present application includes but is not limited to Figure 1 as shown. As Figure 1 shown, the video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compressing) video data to generate a bitstream and transmit the bitstream to the decoding device. The decoding device decodes the bitstream generated by the encoding device to obtain the decoded video data.
[0047] The encoding device 110 of the embodiments of the present application can be understood as a device with video encoding function, and the decoding device 120 can be understood as a device with video decoding function. That is, the embodiments of the present application include a wider range of devices for the encoding device 110 and the decoding device 120, such as smart phones, desktop computers, mobile computing devices, notebooks (e.g., laptops) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, etc.
[0048] In some embodiments, the encoding device 110 can transmit the encoded video data (such as a bitstream) to the decoding device 120 via a channel 130. The channel 130 can include one or more media and / or devices capable of transmitting the encoded video data from the encoding device 110 to the decoding device 120.
[0049] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to directly transmit the encoded video data to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media can also include wired communication media, such as one or more physical transmission lines.
[0050] In another example, the channel 130 includes a storage medium that can store the video data encoded by the encoding device 110. The storage medium includes various locally accessible data storage media, such as optical discs, DVDs, flash memories, etc. In this example, the decoding device 120 can obtain the encoded video data from the storage medium.
[0051] In another example, the channel 130 may include a storage server that can store the video data encoded by the encoding device 110. In this example, the decoding device 120 may download the stored encoded video data from the storage server. Optionally, the storage server can store the encoded video data and transmit the encoded video data to the decoding device 120, such as a web server (e.g., for a website), a File Transfer Protocol (FTP) server, etc.
[0052] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113. Among them, the output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.
[0053] In some embodiments, in addition to the video encoder 112 and the input interface 113, the encoding device 110 may further include a video source 111.
[0054] The video source 111 may include at least one of a video capture device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system. Among them, the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data.
[0055] The video encoder 112 encodes the video data from the video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream contains the encoding information of the pictures or the sequence of pictures in the form of a bitstream. The encoding information may include encoded image data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. The SPS may contain parameters applied to one or more sequences. The PPS may contain parameters applied to one or more pictures. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.
[0056] The video encoder 112 directly transmits the encoded video data to the decoding device 120 via the output interface 113. The encoded video data may also be stored on a storage medium or a storage server for subsequent reading by the decoding device 120.
[0057] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122.
[0058] In some embodiments, in addition to the input interface 121 and the video decoder 122, the decoding device 120 may further include a display device 123.
[0059] Among them, the input interface 121 includes a receiver and / or a modem. The input interface 121 can receive the encoded video data through the channel 130.
[0060] The video decoder 122 is used to decode the encoded video data to obtain the decoded video data, and transmit the decoded video data to the display device 123.
[0061] The display device 123 displays the decoded video data. The display device 123 can be integrated with the decoding device 120 or outside the decoding device 120. The display device 123 can include various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0062] In addition, Figure 1 For example only, the technical solutions of the embodiments of the present application are not limited to Figure 1 , for example, the technology of the present application can also be applied to unilateral video encoding or unilateral video decoding.
[0063] Next, the video encoding framework involved in the embodiments of the present application will be introduced.
[0064] Figure 2 It is a schematic block diagram of a video encoder involved in the embodiments of the present application. It should be understood that the video encoder 200 can be used for lossy compression of images or lossless compression of images. This lossless compression can be visually lossless compression or mathematically lossless compression.
[0065] The video encoder 200 can be applied to image data in the luminance-chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4. Y represents luminance, Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance used to describe color and saturation. For example, in the color format, 4:2:0 means that for every 4 pixels, there are 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that for every 4 pixels, there are 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 means full-pixel display (YYYYCbCrCbCrCbCrCbCr).
[0066] For example, the video encoder 200 reads video data. For each frame of the video data, a frame of the image is divided into a number of coding tree units (CTUs). In some examples, a CTB can be referred to as a "tree block", "Largest Coding Unit" (LCU for short), or "coding tree block" (CTB for short). Each CTU can be associated with a pixel block of equal size within the image. Each pixel can correspond to one luminance sample and two chrominance samples. Therefore, each CTU can be associated with a luminance sample block and two chrominance sample blocks. The size of a CTU is, for example, 128×128, 64×64, 32×32, etc. A CTU can be further divided into a number of coding units (CUs) for encoding. A CU can be a rectangular block or a square block. A CU can be further divided into a prediction unit (PU for short) and a transform unit (TU for short), so that encoding, prediction, and transformation are separated, making the processing more flexible. In one example, a CTU is divided into CUs in a quadtree manner, and a CU is divided into TUs and PUs in a quadtree manner.
[0067] Video encoders and video decoders can support various PU sizes. Assuming that the size of a specific CU is 2N×2N, video encoders and video decoders can support PU sizes of 2N×2N or N×N for intra prediction, and support symmetric PUs of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter prediction. Video encoders and video decoders can also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter prediction.
[0068] In some embodiments, such asFigure 2 As shown, the video encoder 200 may include: a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filter unit 260, a decoded image buffer 270, and an entropy encoding unit 280. It should be noted that the video encoder 200 may include more, fewer, or different functional components.
[0069] Optionally, in this application, a current block may be referred to as a current coding unit (CU) or a current prediction unit (PU), etc. A prediction block may also be referred to as a predicted image block or an image prediction block, and a reconstructed image block may also be referred to as a reconstruction block or an image reconstruction block.
[0070] In some embodiments, the prediction unit 210 includes an inter-frame prediction unit 211 and an intra-frame prediction unit 212. Since there is a strong correlation between adjacent pixels in a frame of video, the intra-frame prediction method is used in video coding and decoding technologies to eliminate the spatial redundancy between adjacent pixels. Since there is a strong similarity between adjacent frames in video, the inter-frame prediction method is used in video coding and decoding technologies to eliminate the temporal redundancy between adjacent frames, thereby improving the coding efficiency.
[0071] The inter-frame prediction unit 211 can be used for inter-frame prediction. Inter-frame prediction may include motion estimation and motion compensation, and may refer to image information of different frames. Inter-frame prediction uses motion information to find a reference block from a reference frame, and generates a prediction block based on the reference block to eliminate temporal redundancy. The frames used for inter-frame prediction may be P-frames and / or B-frames. A P-frame refers to a forward prediction frame, and a B-frame refers to a bi-directional prediction frame. Inter-frame prediction uses motion information to find a reference block from a reference frame and generates a prediction block based on the reference block. The motion information includes the reference frame list where the reference frame is located, the reference frame index, and the motion vector. The motion vector can be an integer pixel or a fractional pixel. If the motion vector is a fractional pixel, then an interpolation filter needs to be used in the reference frame to create the required fractional pixel block. Here, the integer pixel or fractional pixel block found in the reference frame according to the motion vector is called the reference block. Some technologies directly use the reference block as the prediction block, and some technologies further process the reference block to generate the prediction block. Further processing the reference block to generate the prediction block can also be understood as using the reference block as the prediction block and then further processing the prediction block to generate a new prediction block.
[0072] The intra-frame prediction unit 212 only refers to the information of the same frame image and predicts the pixel information within the current coded image block to eliminate spatial redundancy. The frame used for intra-frame prediction may be an I-frame.
[0073] There are multiple prediction modes for intra prediction. Taking the international digital video coding standard H series as an example, the H.264 / AVC standard has 8 angular prediction modes and 1 non-angular prediction mode, and H.265 / HEVC extends to 33 angular prediction modes and 2 non-angular prediction modes. The intra prediction modes used by HEVC include the Planar mode, DC, and 33 angular modes, for a total of 35 prediction modes. The intra modes used by VVC include Planar, DC, and 65 angular modes, for a total of 67 prediction modes.
[0074] It should be noted that with the increase in the number of angular modes, intra prediction will be more accurate and more in line with the requirements for the development of high-definition and ultra-high-definition digital videos.
[0075] The residual unit 220 can generate the residual block of the CU based on the pixel block of the CU and the prediction block of the PU of the CU. For example, the residual unit 220 can generate the residual block of the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel block of the CU and the corresponding sample in the prediction block of the PU of the CU.
[0076] The transform / quantization unit 230 can quantize the transform coefficients. The transform / quantization unit 230 can quantize the transform coefficients associated with the TU of the CU based on the quantization parameter (QP) value associated with the CU. The video encoder 200 can adjust the quantization degree applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.
[0077] The inverse transform / quantization unit 240 can apply inverse quantization and inverse transform to the quantized transform coefficients respectively to reconstruct the residual block from the quantized transform coefficients.
[0078] The reconstruction unit 250 can add the samples of the reconstructed residual block to the corresponding samples of one or more prediction blocks generated by the prediction unit 210 to generate the reconstructed image block associated with the TU. By reconstructing the sample blocks of each TU of the CU in this way, the video encoder 200 can reconstruct the pixel block of the CU.
[0079] The loop filter unit 260 is used to process the pixels after inverse transform and inverse quantization, compensate for the distorted information, and provide a better reference for subsequent encoded pixels. For example, it can perform a deblocking filter operation to reduce the blocking effect of the pixel block associated with the CU.
[0080] In some embodiments, the loop filter unit 260 includes a deblocking filter unit and a sample adaptive offset / adaptive loop filter (SAO / ALF) unit, where the deblocking filter unit is used to remove the blocking effect and the SAO / ALF unit is used to remove the ringing effect.
[0081] The decoded picture buffer 270 may store the reconstructed pixel blocks. The inter prediction unit 211 may perform inter prediction on PUs of other pictures by using the reference pictures containing the reconstructed pixel blocks. Additionally, the intra prediction unit 212 may perform intra prediction on other PUs in the same picture as the CU by using the reconstructed pixel blocks in the decoded picture buffer 270.
[0082] The entropy coding unit 280 may receive the quantized transform coefficients from the transform / quantization unit 230. The entropy coding unit 280 may perform one or more entropy coding operations on the quantized transform coefficients to generate the entropy-coded data.
[0083] Figure 3 is a schematic block diagram of a video decoder according to an embodiment of the present application.
[0084] As Figure 3 shown, the video decoder 300 includes: an entropy decoding unit 310, a prediction unit 320, an inverse quantization / transformation unit 330, a reconstruction unit 340, a loop filter unit 350, and a decoded picture buffer 360. It should be noted that the video decoder 300 may include more, fewer, or different functional components.
[0085] The video decoder 300 may receive a bitstream. The entropy decoding unit 310 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, the entropy decoding unit 310 may parse the entropy-coded syntax elements in the bitstream. The prediction unit 320, the inverse quantization / transformation unit 330, the reconstruction unit 340, and the loop filter unit 350 may decode the video data according to the syntax elements extracted from the bitstream, that is, generate the decoded video data.
[0086] In some embodiments, the prediction unit 320 includes an intra prediction unit 322 and an inter prediction unit 321.
[0087] The intra prediction unit 322 may perform intra prediction to generate a prediction block of the PU. The intra prediction unit 322 may use an intra prediction mode to generate a prediction block of the PU based on the pixel blocks of spatially adjacent PUs. The intra prediction unit 322 may also determine the intra prediction mode of the PU according to one or more syntax elements parsed from the bitstream.
[0088] The inter prediction unit 321 may construct a first reference picture list (list 0) and a second reference picture list (list 1) according to the syntax elements parsed from the bitstream. Additionally, if the PU is encoded using inter prediction, the entropy decoding unit 310 may parse the motion information of the PU. The inter prediction unit 321 may determine one or more reference blocks of the PU according to the motion information of the PU. The inter prediction unit 321 may generate a prediction block of the PU according to one or more reference blocks of the PU.
[0089] The inverse quantization / transformation unit 330 inversely quantizes (i.e., dequantizes) the transform coefficients associated with the TU. The inverse quantization / transformation unit 330 may use the QP value associated with the CU of the TU to determine the degree of quantization.
[0090] After inversely quantizing the transform coefficients, the inverse quantization / transformation unit 330 may apply one or more inverse transforms to the inversely quantized transform coefficients to generate a residual block associated with the TU.
[0091] The reconstruction unit 340 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, the reconstruction unit 340 may add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the pixel block of the CU, obtaining a reconstructed image block.
[0092] The loop filter unit 350 may perform a deblocking filter operation to reduce the blocking artifacts of the pixel block associated with the CU.
[0093] The video decoder 300 may store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 may use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.
[0094] The basic process of video coding and decoding is as follows: At the encoding end, a frame of image is divided into blocks. For the current block, the prediction unit 210 generates a prediction block of the current block using intra prediction or inter prediction. The residual unit 220 may calculate a residual block based on the prediction block and the original block of the current block, that is, the difference between the prediction block and the original block of the current block. This residual block may also be referred to as residual information. The residual block undergoes processes such as transformation and quantization by the transform / quantization unit 230, which can remove information insensitive to the human eye to eliminate visual redundancy. Optionally, the residual block before being transformed and quantized by the transform / quantization unit 230 may be referred to as a temporal residual block, and the temporal residual block after being transformed and quantized by the transform / quantization unit 230 may be referred to as a frequency residual block or a frequency-domain residual block. The entropy coding unit 280 receives the quantized transform coefficients output by the transform quantization unit 230 and may perform entropy coding on the quantized transform coefficients to output a bitstream. For example, the entropy coding unit 280 may eliminate character redundancy according to the target context model and the probability information of the binary bitstream.
[0095] At the decoding end, the entropy decoding unit 310 can parse the code stream to obtain prediction information, quantization coefficient matrix, etc. of the current block. The prediction unit 320 uses intra prediction or inter prediction for the current block based on the prediction information to generate a predicted block of the current block. The inverse quantization / transformation unit 330 uses the quantization coefficient matrix obtained from the code stream to perform inverse quantization and inverse transformation on the quantization coefficient matrix to obtain a residual block. The reconstruction unit 340 adds the predicted block and the residual block to obtain a reconstructed block. The reconstructed blocks form a reconstructed image, and the loop filter unit 350 performs loop filtering on the reconstructed image based on the image or based on blocks to obtain a decoded image. The encoding end also needs to perform similar operations as the decoding end to obtain a decoded image. This decoded image can also be referred to as a reconstructed image, and the reconstructed image can be used as a reference frame for inter prediction for subsequent frames.
[0096] It should be noted that the block partitioning information determined by the encoding end, as well as mode information or parameter information such as prediction, transformation, quantization, entropy encoding, loop filtering, etc. are carried in the code stream when necessary. The decoding end determines the same block partitioning information, mode information or parameter information such as prediction, transformation, quantization, entropy encoding, loop filtering, etc. as the encoding end by parsing the code stream and analyzing according to the existing information, so as to ensure that the decoded image obtained by the encoding end is the same as the decoded image obtained by the decoding end.
[0097] The above is the basic process of a video codec under a block-based hybrid coding framework. With the development of technology, some modules or steps of this framework or process may be optimized. This application is applicable to the basic process of a video codec under this block-based hybrid coding framework, but is not limited to this framework and process.
[0098] In the embodiments of this application, the current block can be the current coding unit (CU) or the current prediction unit (PU), etc. Due to the need for parallel processing, an image can be divided into slices, etc. Slices in the same image can be processed in parallel, that is, there is no data dependency between them. And "frame" is a common term, and generally it can be understood that one frame is an image. The frame mentioned in the application can also be replaced by an image or a slice, etc.
[0099] Inter prediction refers to the process of predicting a reference block for a block to be encoded (such as the current block) in the current image from a neighboring encoded image (such as a reference image). The purpose is to remove the temporal redundancy of the video signal. The schematic diagram of inter prediction is as Figure 4A shown. According to the block matching criterion, the best matching block of the current block is searched for in the reference frame.
[0100] Inter-frame prediction uses motion information to represent "motion". The basic motion information includes information about the reference frame (or reference picture) and information about the motion vector (MV). The commonly used bi-directional prediction currently uses two reference blocks to predict the current block. The two reference blocks can use one forward reference block and one backward reference block. Later, it is also allowed that both are forward or both are backward. Forward means that the time corresponding to the reference frame is before the current frame, and backward means that the time corresponding to the reference frame is after the current frame. Or, forward means that the position of the reference frame in the video is before the current frame, and backward means that the position of the reference frame in the video is after the current frame. Or, forward means that the POC (picture order count) of the reference frame is less than the POC of the current frame, and backward means that the POC of the reference frame is greater than the POC of the current frame. In order to use bi-directional prediction, naturally, two reference blocks need to be found, so two sets of information about the reference frames and motion vectors are required. Each of these sets can be understood as a uni-directional motion information, and combining these two sets together forms a bi-directional motion information. In specific implementation, the uni-directional motion information and the bi-directional motion information can use the same data structure, except that the information about the reference frames and motion vectors of both sets in the bi-directional motion information is valid, while the information about the reference frames and motion vectors of one of the sets in the uni-directional motion information is invalid.
[0101] VVC supports two reference picture lists, denoted as RPL0 and RPL1, where RPL is the abbreviation of Reference Picture List. In VVC, the P slice can only use RPL0, and the B slice can use RPL0 and RPL1. For a slice, there are several reference pictures in each reference picture list, and the codec finds a certain reference picture through the reference picture index. VVC represents the motion information using the reference picture index and the motion vector. For the above bi-directional motion information, VVC uses the reference picture index refIdxL0 corresponding to reference picture list 0, the motion vector mvL0 corresponding to reference picture list 0, the reference picture index refIdxL1 corresponding to reference picture list 1, and the motion vector mvL0 corresponding to reference picture list 1. The reference picture index corresponding to reference picture list 0 and the reference picture index corresponding to reference picture list 1 here can be understood as the information of the above reference pictures. VVC uses two flag bits to represent whether to use the motion information corresponding to reference picture list 0 and whether to use the motion information corresponding to reference picture list 0, denoted as predFlagL0 and predFlagL1 respectively. It can also be understood that predFlagL0 and predFlagL1 represent whether the above unidirectional motion information is "valid". Therefore, although there is no explicit mention of the data structure of motion information in VVC, it uses the reference picture index, motion vector, and the flag bit of "whether it is valid" corresponding to each reference picture list to represent the motion information together. In the VVC standard text, motion information does not appear, but motion vectors are used. It can also be considered that the reference picture index and the flag bit of whether to use the corresponding motion information are appendages of the motion vector. In this article, for the convenience of description, "motion information" is still used, but it should be understood that "motion vector" can also be used for description.
[0102] As can be seen from the above, inter-frame prediction needs to use the motion vector to obtain the pixel prediction value in the reference picture. In actual situations, the motion of objects between adjacent images does not necessarily take the whole pixel as the basic unit. Therefore, in order to improve the prediction accuracy, it is necessary to improve the accuracy of motion estimation to the sub-pixel level, interpolate the reference picture to improve the accuracy of motion compensation, and then improve the coding efficiency. As Figure 4B shown, the pixel points where the capital A is located are whole pixels, such as A 0,0 、A 1,0 etc., and the pixel points where the lowercase a, b, c, etc. are located are sub-pixels, such as a 0,0 is a 1 / 4 pixel point in the X direction, b 0,0For example, it is 1 / 2 pixel point in the X direction. Exemplarily, the common inter-frame prediction modes in AVS3 support five motion vector precisions of 1 / 4, 1 / 2, 1, 2, and 4. When the MV precision is 1 / 4 or 1 / 2, an 8-tap interpolation filter can be used to interpolate the luminance prediction value, and a 4-tap interpolation filter can be used to interpolate the chrominance prediction value.
[0103] Traditional bidirectional prediction obtains the prediction value of the current block by weighted averaging of two reconstructed blocks, where the two reconstructed blocks come from the forward reference frame and the backward reference frame respectively. Further, in order to improve the prediction effect, the bidirectional optical flow (BIO) technique can be used to compensate the motion after bidirectional prediction to reduce motion deviation and improve coding efficiency. Specifically, BIO in AVS3 first calculates the gradient values in the x direction and the y direction (the x direction represents the horizontal direction and the y direction represents the vertical direction throughout the text), then obtains the calculation factor of each pixel according to the pixel value and the gradient value, derives the motion vector, and finally calculates a more accurate prediction value. Exemplarily, when calculating the gradient value in the x direction, BIO first uses an 8-tap gradient filter for gradient calculation, and then uses an 8-tap interpolation filter for interpolation. When calculating the gradient value in the y direction, BIO first uses an 8-tap interpolation filter for interpolation, and then uses an 8-tap gradient filter for gradient calculation. In addition, the BIO technique is only applied to the luminance component.
[0104] As can be seen from the above, currently when determining the prediction value or gradient by interpolation, the selected interpolation filter is pre-determined. For example, when interpolating for prediction, an 8-tap interpolation filter is selected to interpolate the luminance prediction value, and a 4-tap interpolation filter is selected to interpolate the chrominance prediction value. When calculating the gradient in BIO, an 8-tap interpolation filter is selected. Currently, when determining the interpolation filter, the codec-related information is not considered, resulting in inaccurate selection of the interpolation filter, and thus poor video prediction effect.
[0105] To solve the above technical problems, the embodiments of the present application adaptively select an interpolation filter based on the codec information of the area to be coded / decoded, such as prediction mode, interpolation direction, decoded component, interpolation target, etc., improving the accuracy of the interpolation filter selection. In this way, based on the accurately selected target interpolation filter, interpolating the reference area of the area to be coded / decoded to determine the prediction value of the area to be coded / decoded can improve the prediction effect of the area to be coded / decoded, and thus improve the performance of video coding and decoding.
[0106] The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0107] First, taking the decoding end as an example, the video decoding method provided by the embodiments of the present application will be introduced.
[0108] Figure 5 It is a schematic flowchart of the video decoding method provided by an embodiment of the present application. The embodiments of the present application are applied to Figure 1 and Figure 3 the video decoders shown. As Figure 5 shown, the method of the embodiments of the present application includes:
[0109] S101. Decode the bitstream to obtain the decoding information of the area to be decoded.
[0110] Among them, the decoding information includes at least one of a prediction mode, an interpolation direction, a decoded component, and an interpolation target.
[0111] The embodiments of the present application do not limit the specific size and shape of the area to be decoded.
[0112] In some embodiments, the area to be decoded may be one or several image blocks to be decoded in the current frame to be decoded. For example, it may be one or several CTUs, or one or several CUs, or one or several PUs, etc. In one example, the image block to be decoded may be referred to as the current block, the current image block to be decoded, etc.
[0113] In some embodiments, the area to be decoded may be the current frame to be decoded, that is, the area to be decoded is a whole-frame image.
[0114] During the video decoding process, when the decoding end decodes the area to be decoded, it decodes the bitstream to obtain the quantization coefficients of the area to be decoded, inverse-quantizes the quantization coefficients to obtain the transform coefficients of the area to be decoded, inverse-transforms the transform coefficients to obtain the residual values of the area to be decoded. Then, it determines the prediction mode of the area to be decoded, determines the predicted value of the area to be decoded based on the prediction mode, and obtains the reconstructed value of the area to be decoded based on the predicted value and the residual value of the area to be decoded.
[0115] The embodiments of the present application mainly relate to the prediction process of the area to be decoded.
[0116] In the embodiments of the present application, in order to improve the prediction accuracy of the area to be decoded, an interpolation filter is used to interpolate the reference area of the area to be decoded. However, currently, the decoding information of the area to be decoded is not considered when selecting the interpolation filter, resulting in inaccurate selection of the interpolation filter.
[0117] To solve the above technical problems, in the embodiments of the present application, the decoding information of the area to be decoded is obtained by decoding the bitstream, and then based on the decoding information, the target interpolation filter of the area to be decoded is determined to improve the selection accuracy of the interpolation filter.
[0118] The specific content of the decoding information for the area to be decoded in the embodiments of the present application is not limited, and can be understood as all decoding information related to decoding the area to be decoded, such as prediction mode, inverse transformation method, inverse quantization method, interpolation direction, decoded component, and interpolation target, etc.
[0119] In some embodiments, the decoding information of the area to be decoded in the embodiments of the present application includes at least one of the prediction mode of the area to be decoded, the interpolation direction to be interpolated, the decoded component to be decoded, and the interpolation target.
[0120] Exemplarily, the prediction mode of the area to be decoded may include various modes in the intra prediction mode, various modes in the inter prediction mode, and various modes in the hybrid prediction mode, etc.
[0121] Exemplarily, the interpolation direction includes the x direction and the y direction.
[0122] Exemplarily, the decoded component includes a luminance component and a chrominance component.
[0123] Exemplarily, the interpolation target includes determining a predicted value (for example, performing interpolation filtering on the reference area of the area to be decoded to determine the predicted value of the area to be decoded), and determining a gradient (for example, determining the gradient of the forward reference block and the backward reference block in the x direction or the y direction in BIO).
[0124] In a possible implementation manner, the encoding end may write the index or indication information of the prediction mode of the area to be decoded into the bitstream. In this way, the decoding end can obtain the prediction mode of the area to be decoded by decoding the bitstream.
[0125] In a possible implementation manner, the encoding end may write the interpolation direction of the area to be decoded into the bitstream. For example, if only the x direction of the area to be decoded is interpolated and the y direction is not interpolated, the encoding end may write the indication information of the interpolation direction in the bitstream to indicate that only the x direction is interpolated and the y direction is not interpolated. On the contrary, if only the y direction of the area to be decoded is interpolated and the x direction is not interpolated, the encoding end may write the indication information of the interpolation direction in the bitstream to indicate that only the y direction is interpolated and the x direction is not interpolated. For another example, if both the x direction and the y direction of the area to be decoded need to be interpolated, the encoding end may write the indication information of the interpolation direction in the bitstream to indicate that both the x direction and the y direction are interpolated. Optionally, the encoding end and the decoding end may pre-specify whether to interpolate the x direction and / or the y direction.
[0126] In a possible implementation, the encoding end separately encodes the luminance component and the chrominance component. Based on this, the encoding end indicates in the bitstream whether the information included in the bitstream corresponds to the luminance component or the chrominance component. In this way, by decoding the bitstream, the decoding end can obtain whether the current region to be decoded is the luminance component or the chrominance component of the region to be decoded.
[0127] In a possible implementation, the interpolation target can be derived from the prediction mode of the region to be decoded. For example, when the decoding end decodes the bitstream, determines that the prediction mode of the region to be decoded is the inter prediction mode and not the BIO mode, and at the same time, when the decoding end decodes the bitstream and obtains that the motion vector accuracy corresponding to the region to be decoded is sub-pixel, the decoding end can determine that the interpolation target is to determine the predicted value. If the decoding end decodes the bitstream and determines that the prediction mode of the region to be decoded is the BIO mode, the decoding end can determine that the interpolation target is to determine the gradient.
[0128] In one example, the decoding information of the above-mentioned region to be decoded may only include any one of the prediction mode, interpolation direction, decoded component, and interpolation target. For example, the decoding information of the region to be decoded includes the prediction mode, and the prediction mode is prediction mode 1. In this way, the decoding end can determine the target interpolation filter from the interpolation filter corresponding to the prediction mode.
[0129] In one example, the decoding information of the above-mentioned region to be decoded may include any two or three of the prediction mode, interpolation direction, decoded component, and interpolation target. For example, the decoding information of the region to be decoded includes the prediction mode and the interpolation direction, the prediction mode is prediction mode 1, and the interpolation direction is the x direction. In this way, the decoding end can determine the target interpolation filter based on the interpolation filter corresponding to prediction mode 1 and the x direction.
[0130] In one example, the decoding information of the above-mentioned region to be decoded may include the prediction mode, interpolation direction, decoded component, and interpolation target. The embodiments of the present application do not limit this. For example, the decoding information of the region to be decoded includes that the prediction mode is prediction mode 1, the interpolation direction is the x direction, the decoded component is the luminance component, and the interpolation target is to determine the predicted value of the region to be decoded. In this way, the decoding end can determine the target interpolation filter based on the interpolation filter corresponding to the prediction mode, interpolation direction, decoded component, and interpolation target.
[0131] In some embodiments, the decoding end can also obtain at least one of the prediction mode, interpolation direction, decoded component, and interpolation target corresponding to the region to be decoded in other ways.
[0132] After the decoding end determines the decoding information of the region to be decoded, it performs the steps of S102 as follows.
[0133] S102. Determine a target interpolation filter corresponding to the area to be decoded based on the decoding information.
[0134] In the embodiments of the present application, the decoding end determines a target interpolation filter corresponding to the area to be decoded from interpolation filters with different numbers of taps based on the decoding information of the area to be decoded, fully considering the performance differences of interpolation filters with different numbers of taps in different coding modes, different decoded components, different interpolation directions, and different interpolation targets, thereby improving the selection accuracy of the target interpolation filter.
[0135] The embodiments of the present application do not limit the specific manner in which the decoding end determines a target interpolation filter corresponding to the area to be decoded based on at least one of the prediction mode, interpolation direction, decoded component, and interpolation target.
[0136] In some embodiments, the above S102 includes the following steps S102-A1 to S102-A3:
[0137] S102-A1. Determine a plurality of candidate interpolation filters based on the decoding information;
[0138] S102-A2. Decode the code stream to obtain first indication information of the target interpolation filter, where the first indication information is used to indicate the index information of the target interpolation filter among the plurality of candidate interpolation filters;
[0139] S102-A3. Select the target interpolation filter from the plurality of candidate interpolation filters based on the first indication information.
[0140] In this implementation manner, when the decoding end determines a target interpolation filter corresponding to the area to be decoded based on the decoding information of the area to be decoded determined above, it first determines a plurality of candidate interpolation filters based on the decoding information, and then selects an interpolation filter from these candidate interpolation filters as the target interpolation filter. That is to say, the embodiments of the present application determine a plurality of candidate interpolation filters with different numbers of taps based on the decoding information, and then select the target interpolation filter from these plurality of candidate interpolation filters with different numbers of taps, fully considering the performance differences of interpolation filters with different numbers of taps in different coding modes, different decoded components, different interpolation directions, and different interpolation targets, thereby improving the determination accuracy of the target interpolation filter.
[0141] The following introduces the specific process of the decoding end determining a plurality of candidate interpolation filters based on the decoding information.
[0142] In the embodiments of the present application, the implementation manners of determining a plurality of candidate interpolation filters based on the decoding information in the above S102-A1 include at least the following several types:
[0143] In Method 1, the decoding end determines multiple candidate interpolation filters through the following steps S102-A1-a1 and S102-A1-a2:
[0144] S102-A1-a1: Obtain at least one of the prediction mode, interpolation direction, decoded component, and interpolation target, and respectively obtain the corresponding preset candidate interpolation filters;
[0145] S102-A1-a2: Based on the preset candidate interpolation filters, determine multiple candidate interpolation filters.
[0146] In this implementation, at least one of the prediction mode, interpolation direction, decoded component, and interpolation target in the decoded information respectively corresponds to a preset candidate interpolation filter.
[0147] In one example, different prediction modes can correspond to different candidate interpolation filters. For example, prediction mode 1 corresponds to m1 candidate interpolation filters, and prediction mode 2 corresponds to m2 candidate interpolation filters. The number of taps of each of the m1 candidate interpolation filters is not exactly the same as that of each of the m2 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the BIO mode include an 8-tap interpolation filter and a 12-tap interpolation filter. The candidate interpolation filters corresponding to the non-BIO mode include a 6-tap interpolation filter and an 8-tap interpolation filter.
[0148] In one example, different interpolation directions can correspond to different candidate interpolation filters. For example, the x direction corresponds to m3 candidate interpolation filters, and the y direction corresponds to m4 candidate interpolation filters. The number of taps of each of the m3 candidate interpolation filters is not exactly the same as that of each of the m4 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the x (horizontal) direction include an 8-tap interpolation filter and a 12-tap interpolation filter, and the candidate interpolation filters corresponding to the y (vertical) direction include a 6-tap interpolation filter and an 8-tap interpolation filter.
[0149] In one example, different decoded components can correspond to different candidate interpolation filters. For example, the luminance component corresponds to m5 candidate interpolation filters, and the chrominance component corresponds to m6 candidate interpolation filters. The number of taps of each of the m5 candidate interpolation filters is not exactly the same as that of each of the m6 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the luminance component include an 8-tap interpolation filter and a 12-tap interpolation filter, and the candidate interpolation filters corresponding to the chrominance component include a 4-tap interpolation filter and a 6-tap interpolation filter.
[0150] Exemplarily, different interpolation targets may correspond to different candidate interpolation filters. For example, when determining the predicted value, there are m7 candidate interpolation filters, and when determining the gradient, there are m8 candidate interpolation filters. The number of taps of each of the m7 candidate interpolation filters is not exactly the same as the number of taps of each of the m8 candidate interpolation filters.
[0151] In this way, the decoding end can determine multiple candidate interpolation filters based on the candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, decoded component, and interpolation target.
[0152] For example, a preset number of candidate interpolation filters are selected from the candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, decoded component, and interpolation target as the multiple candidate interpolation filters. For example, when the above decoding information includes the prediction mode and the decoded component, the decoding end selects a preset number (for example, n) from the candidate interpolation filters corresponding to the prediction mode and the decoded component respectively to form multiple candidate interpolation filters. Referring to the above example, assume that the prediction mode of the area to be decoded is prediction mode 1 and the decoded component is the luminance component. Among them, prediction mode 1 corresponds to m1 candidate interpolation filters, and the luminance component corresponds to m5 candidate interpolation filters. In this way, a preset number of candidate interpolation filters can be selected from the m1 + m5 candidate interpolation filters of the m1 candidate interpolation filters and the m5 candidate interpolation filters to form multiple candidate interpolation filters.
[0153] For another example, the decoding end removes the duplicate interpolation filters from the selected prediction candidate interpolation filters to obtain multiple candidate interpolation filters. For example, assume that the decoding information includes the prediction mode and the decoded component, where the prediction mode is the non - BIO mode and the decoded component is the luminance component. Assume that the candidate interpolation filters corresponding to the non - BIO mode include a 6 - tap interpolation filter and an 8 - tap interpolation filter, and the 8 - tap interpolation filter and a 12 - tap interpolation filter corresponding to the luminance component. One 8 - tap interpolation filter is removed from these 4 interpolation filters, and the remaining 3 interpolation filters are the 6 - tap interpolation filter, the 8 - tap interpolation filter, and the 12 - tap interpolation filter. Then, the 3 interpolation filters of the 6 - tap interpolation filter, the 8 - tap interpolation filter, and the 12 - tap interpolation filter are determined as multiple candidate interpolation filters.
[0154] As can be seen from the above description, in the first method, multiple candidate interpolation filters are selected from the preset candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, decoded component, and interpolation target.
[0155] In some embodiments, the decoding end can also determine multiple candidate interpolation filters by the method shown in the following Method 2.
[0156] In the second method, the decoding end determines multiple candidate interpolation filters through the following steps of S102-A1-b:
[0157] S102-A1-b. Obtain the selection strategies respectively corresponding to at least one of the prediction mode, interpolation direction, decoded component, and interpolation target, and based on the selection strategies, select multiple candidate interpolation filters from the preset M interpolation filters, where M is a positive integer greater than 1.
[0158] In this implementation manner, there are M preset interpolation filters, and the decoding end can select multiple candidate interpolation filters from these M preset interpolation filters based on at least one of the prediction mode, interpolation direction, decoded component, and interpolation target. Optionally, the number of taps of these M interpolation filters is different.
[0159] Specifically, the decoding end obtains the selection strategies respectively corresponding to at least one of the prediction mode, interpolation direction, decoded component, and interpolation target. The selection strategies are default for both the encoding and decoding ends, or the encoding end transmits them to the decoding end through the code stream. In this way, the decoding end can select multiple candidate interpolation filters from these M preset interpolation filters based on the selection strategies.
[0160] In this implementation manner, the selection strategies respectively corresponding to at least one of the above prediction mode, interpolation direction, decoded component, and interpolation target can be the same, different, or partially the same and partially different.
[0161] Exemplarily, assume that the above decoding information includes the prediction mode and the decoded component, the selection strategy corresponding to the prediction mode is Strategy 1, and the selection strategy corresponding to the decoded component is Strategy 2. In this way, the decoding end can select one or more interpolation filters from the M interpolation filters based on Strategy 1. At the same time, the decoding end can select one or more interpolation filters from the M interpolation filters based on Strategy 2. Furthermore, the decoding end obtains multiple candidate interpolation filters according to one or more interpolation filters selected based on Strategy 1 and one or more interpolation filters selected based on Strategy 2. For example, the decoding end determines the interpolation filters with different numbers of taps among the one or more interpolation filters selected based on Strategy 1 and the one or more interpolation filters selected based on Strategy 2 as multiple candidate interpolation filters.
[0162] In some embodiments, the decoding end can also determine multiple candidate interpolation filters through the method shown in the following Method 3.
[0163] Method 3. The decoding end determines multiple candidate interpolation filters through the following steps of S102-A1-c:
[0164] S102 - A1 - c. Obtain N candidate combinations composed of M interpolation filters, decode the bitstream to obtain second indication information, and based on the second indication information, select a target candidate combination from the N candidate combinations, and then determine the interpolation filters included in the target candidate combination as multiple candidate interpolation filters, where each candidate combination includes two or more interpolation filters among the M interpolation filters, and the second indication information is used to indicate the target candidate combination, and N is a positive integer greater than 1.
[0165] In this implementation manner, M interpolation filters are preset, and the decoding end obtains N candidate combinations composed of these M interpolation filters. Optionally, the number of taps of these M interpolation filters is different.
[0166] In one example, the above N candidate combinations are default for both the encoding and decoding ends. That is, the decoding end itself stores N candidate combinations, or the decoding end can combine the M interpolation filters using the same rule as the encoding end to obtain N candidate combinations.
[0167] In one example, the above N candidate combinations are indicated by the encoding end to the decoding end. For example, the encoding end writes the determined N candidate combinations into the bitstream. The decoding end obtains the N candidate combinations by decoding the bitstream.
[0168] The embodiments of this application do not limit the specific combination form of the N candidate combinations.
[0169] In a possible implementation manner, the interpolation filters included in different candidate combinations among the above N candidate combinations are randomly distributed. For example, assume that the M interpolation filters include interpolation filters with 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The N candidate combinations composed of these 5 types of interpolation filters are: {6 taps, 8 taps, 12 taps}, {2 taps, 4 taps, 6 taps}, and {4 taps, 6 taps, 8 taps, 12 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 1:
[0170] Table 1
[0171] Index Bit Meaning 0 00 Candidate combinations: {6 - tap, 8 - tap, 12 - tap} 1 01 Candidate combinations: {2 - tap, 4 - tap, 6 - tap} 2 10 Candidate combinations: {4 - tap, 6 - tap, 8 - tap, 12 - tap} …… …… ……
[0172] In a possible implementation, the number of interpolation filters included in each of the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination increases sequentially. For example, assume that M interpolation filters include interpolation filters with 6 types of taps: 2 taps, 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The N candidate combinations formed by these 6 types of tap interpolation filters include: {2 taps, 4 taps}, {2 taps, 4 taps, 6 taps}, {2 taps, 4 taps, 6 taps, 8 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps, 16 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 2:
[0173] Table 2
[0174] Index Bit Meaning 0 000 Do not use any interpolation filter 1 010 Use the first 1 interpolation filter (2 - tap) 2 011 Use the first 2 interpolation filters (2, 4 - tap) 3 100 Use the first 3 interpolation filters (2, 4, 6 - tap) 4 101 Use the first 4 interpolation filters (2, 4, 6, 8 - tap) 5 110 Use the first 5 interpolation filters (2, 4, 6, 8, 12 - tap) 6 111 Use the first 6 interpolation filters (2, 4, 6, 8, 12, 16 - tap)
[0175] In a possible implementation, the number of interpolation filters included in each of the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination decreases sequentially. For example, assume that M interpolation filters include interpolation filters with 6 types of taps: 2 taps, 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The N candidate combinations formed by these 6 types of tap interpolation filters include: {12 taps, 16 taps}, {8 taps, 12 taps, 16 taps}, {6 taps, 8 taps, 12 taps, 16 taps}, {4 taps, 6 taps, 8 taps, 12 taps, 16 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps, 16 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 3:
[0176] Table 3
[0177]
[0178]
[0179] In the third method, when the encoding end determines multiple candidate interpolation filters, it obtains N candidate combinations and selects a target candidate combination from the N candidate combinations. For example, based on the rate-distortion cost, it selects a target candidate combination from the N candidate combinations, and then determines the interpolation filter included in the target candidate combination as the multiple candidate interpolation filters. At the same time, the encoding end writes second indication information into the bitstream. The second indication information is used to indicate the target candidate combination. Exemplarily, the second indication information includes the index of the target candidate combination among the N candidate combinations. For example, the second indication information includes index 100. In this way, the decoding end can obtain the second indication information by decoding the bitstream, and then based on the second indication information, select the target candidate combination from the above-obtained N candidate combinations. For example, the second indication information includes index 100 of the target candidate combination among the N candidate combinations. In this way, the decoding end can select the target candidate combination from the above-obtained N candidate combinations based on index 100. For example, the above N candidate combinations are shown in Table 3, and the candidate combination corresponding to index 100 includes interpolation filters with 8, 12, and 16 taps, and then these 3 interpolation filters are used as the multiple candidate interpolation filters.
[0180] As described above, the decoding end can determine multiple candidate interpolation filters based on the decoding information of the area to be decoded in the above manner.
[0181] The following introduces the specific implementation processes of the above S102-A2 and S102-A3.
[0182] It should be noted that the embodiments of the present application do not limit the specific implementation order of the above S102-A2 and the above S102-A1. For example, the above S102-A2 can be executed before the above S102-A1, or after the above S102-A1, or executed simultaneously with the above S102-A1.
[0183] In the embodiments of the present application, when the encoding end determines the target interpolation filter, it determines multiple candidate interpolation filters based on the decoding information of the area to be decoded, and then selects a candidate interpolation filter from these multiple candidate interpolation filters as the target interpolation filter. For example, based on the rate-distortion cost of each of the multiple candidate interpolation filters, it selects the target interpolation filter from these multiple candidate interpolation filters. Then, it writes first indication information into the bitstream. The first indication information is used to indicate the index information of the target interpolation filter among the multiple candidate interpolation filters. In this way, when the decoding end determines the target interpolation filter, it determines multiple candidate interpolation filters through the above steps, and at the same time decodes the bitstream to obtain the first indication information, and then based on the first indication information, selects the target interpolation filter from the multiple candidate interpolation filters.
[0184] In some embodiments, the above first indication information includes a flag bit array or a plurality of flag bits, and when the flag bit array or the plurality of flag bits are used to indicate decoding information and a target interpolation filter, the decoding end can obtain the first indication information by decoding the code stream, and then obtain the decoding information of the area to be decoded based on the first indication information.
[0185] In one example, the first indication information indicates a prediction mode, a decoded component, and a target interpolation filter. At this time, the first indication information includes a parsed flag bit array Flag1[A][I] (both A and I are positive integers), where A represents the decoded component. For example, A = 0 represents the luminance component, and A = 1 represents the chrominance component. I represents the prediction mode. For example, I = 0 represents mode 1, and I = 1 represents mode 2. Flag1[A][I] represents the index information of the target interpolation filter of component A in prediction mode I. Exemplarily, the first indication information and its meaning are shown in Table 4:
[0186] Table 4
[0187]
[0188]
[0189] As shown in Table 4, assume that the decoding information of the area to be decoded includes a prediction mode and a decoded component, the prediction mode is mode 1, and the decoded component is the chrominance component. The multiple candidate interpolation filters determined by the decoding end based on prediction mode 1 and the chrominance component include a 4-tap interpolation filter and a 6-tap interpolation filter. In this way, if the decoding end decodes the code stream and the obtained first indication information includes a flag bit array Flag1[1][0] = 1, then the target interpolation filter is determined to be a 6-tap interpolation filter; if the first indication information includes a flag bit array Flag1[1][0] = 0, then the target interpolation filter is determined to be a 4-tap interpolation filter.
[0190] In one example, the first indication information indicates an interpolation direction, a decoded component, and a target interpolation filter. At this time, the first indication information includes a parsed flag bit array Flag2[A][B] (both A and B are positive integers), where A represents the decoded component. For example, A = 0 represents the luminance component, and A = 1 represents the chrominance component. B represents the interpolation direction. For example, B = 0 represents the x direction, and B = 1 represents the y direction. Flag2[A][B] represents the index information of the target interpolation filter of component A in interpolation direction B. Exemplarily, the first indication information and its meaning are shown in Table 5:
[0191] Table 5
[0192]
[0193] As shown in Table 5, assume that the decoded information of the area to be decoded includes an interpolation direction and a decoded component, the interpolation direction is the y direction, and the decoded component is a chrominance component. The multiple candidate interpolation filters determined by the decoding end based on the interpolation direction in the y direction and the chrominance component include a 6-tap interpolation filter and a 10-tap interpolation filter. In this way, if the decoding end decodes the code stream and the flag bit array Flag2[1][1] included in the obtained first indication information is 0, the target interpolation filter is determined to be a 6-tap interpolation filter; if the flag bit array Flag2[1][1] included in the first indication information is 1, the target interpolation filter is determined to be a 10-tap interpolation filter.
[0194] In one example, the first indication information indicates an interpolation target, a decoded component, and a target interpolation filter. At this time, the first indication information includes a parsed flag bit array Flag3[A][D] (both A and D are positive integers), where A represents the decoded component. For example, A = 0 represents a luminance component, and A = 1 represents a chrominance component. D represents the interpolation target. For example, D = 0 represents that the interpolation target is to determine a predicted value, and D = 1 represents that the interpolation target is to determine a gradient. Flag3[A][D] represents index information of the target interpolation filter for the A component with respect to the interpolation target D. Exemplarily, the first indication information and its meaning are shown in Table 6:
[0195] Table 6
[0196]
[0197] As shown in Table 6, assume that the decoded information of the area to be decoded includes an interpolation target and a decoded component, the interpolation target is to determine a predicted value, and the decoded component is a chrominance component. The multiple candidate interpolation filters determined by the decoding end based on the interpolation target (i.e., determining a predicted value) and the chrominance component include a 4-tap interpolation filter and a 6-tap interpolation filter. In this way, if the decoding end decodes the code stream and the flag bit array Flag3[0][1] included in the obtained first indication information is 0, the target interpolation filter is determined to be a 4-tap interpolation filter; if the flag bit array Flag3[0][1] included in the first indication information is 1, the target interpolation filter is determined to be a 6-tap interpolation filter.
[0198] In this implementation manner, the decoding end determines multiple candidate interpolation filters with different tap numbers based on at least one of the prediction mode, interpolation direction, decoded component, and interpolation target corresponding to the area to be decoded, and then selects the target interpolation filter from these multiple candidate interpolation filters with different tap numbers.
[0199] In some embodiments, the encoding end and the decoding end may determine the target interpolation filter corresponding to the area to be decoded based on the default interpolation filter corresponding to at least one of the prediction mode, interpolation direction, decoded component, and interpolation target corresponding to the area to be decoded.
[0200] In Example 1, the decoding end determines the target interpolation filter based on the prediction mode, decoded component, and interpolation target.
[0201] For example, both the encoding end and the decoding end default to using a specific N1-tap interpolation filter for luminance and / or chrominance for a certain j1 interpolation targets in a certain i prediction modes, and using a certain N2-tap interpolation filter for luminance and / or chrominance for a certain j2 interpolation contents in the remaining I - i encoding modes, where i, j1, j2, N1, and N2 are all integers, and N1 ≠ N2.
[0202] Illustratively, if the prediction mode is the bidirectional optical flow prediction mode, the decoded component is the luminance component, and the interpolation target is to determine the gradients of the forward prediction block and the backward prediction block of the area to be decoded in the x direction and the y direction, then the n1-tap interpolation filter is determined as the target interpolation filter;
[0203] If the prediction mode is the bidirectional optical flow prediction mode, the decoded component is the luminance component, and the interpolation target is to determine the predicted values of the area to be decoded in the x direction and the y direction, then the n2-tap interpolation filter is determined as the target interpolation filter;
[0204] If the prediction mode is the bidirectional optical flow prediction mode, the decoded component is the chrominance component, and the interpolation target is to determine the predicted values of the area to be decoded in the x direction and the y direction, then the n3-tap interpolation filter is determined as the target interpolation filter;
[0205] If the prediction mode is the non-bidirectional optical flow prediction mode, the decoded component is the luminance component, and the interpolation target is to determine the predicted values of the area to be decoded in the x direction and the y direction, then the n4-tap interpolation filter is determined as the target interpolation filter;
[0206] If the prediction mode is the non-bidirectional optical flow prediction mode, the decoded component is the chrominance component, and the interpolation target is to determine the predicted values of the area to be decoded in the x direction and the y direction, then the n5-tap interpolation filter is determined as the target interpolation filter;
[0207] Among them, n1, n2, n3, n4, and n5 are all positive integers.
[0208] The embodiments of the present application do not limit the specific values of the above n1, n2, n3, n4, and n5.
[0209] In a possible implementation manner, the above n1 is not equal to the above n2.
[0210] In an example, n1 = 8, n2 = 12, n3 = 6, n4 = 12, n5 = 6.
[0211] A specific embodiment is shown in Table 7:
[0212] Table 7
[0213]
[0214] As can be seen from Table 7 above, in the BIO mode of this application embodiment, the number of taps of the interpolation filter used when calculating the gradient is different from the number of taps of the interpolation filter used when calculating the pixel prediction value. For example, an 8-tap interpolation filter is used to calculate the gradient, and a 12-tap interpolation filter is used for pixel prediction value interpolation calculation.
[0215] Example 2, the decoding end determines the target interpolation filter based on the prediction mode and the decoded component.
[0216] For example, both the encoding end and the decoding end default to using a specific interpolation filter with N1 taps only for luminance and / or chrominance in a certain i prediction modes, and using an interpolation filter with N2 taps for luminance and / or chrominance in the remaining I - i prediction modes (i, N1, and N2 are all integers, and N1 ≠ N2).
[0217] Illustratively, if the prediction mode is the bidirectional optical flow prediction mode and the decoded component is the luminance component, then the interpolation filter with m1 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with m1 taps for gradient calculation and prediction value interpolation;
[0218] If the prediction mode is the bidirectional optical flow prediction mode and the decoded component is the chrominance component, then the interpolation filter with m2 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with m2 taps for prediction value interpolation calculation in the x and y directions;
[0219] If the prediction mode is the non - bidirectional optical flow prediction mode and the decoded component is the luminance component, then the interpolation filter with m3 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with m3 taps for prediction value interpolation calculation in the x and y directions;
[0220] If the prediction mode is the non - bidirectional optical flow prediction mode and the decoded component is the chrominance component, then the interpolation filter with m4 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with m4 taps for prediction value interpolation calculation in the x and y directions;
[0221] Among them, m1, m2, m3, and m4 are all positive integers.
[0222] The application embodiment does not limit the specific values of m1, m2, m3, and m4 above.
[0223] In one example, m1 = 8, m2 = 6, m3 = 12, m4 = 6.
[0224] A specific implementation example is shown in Table 8:
[0225] Table 8
[0226]
[0227] In Example 3, the decoding end determines the target interpolation filter based on the interpolation direction and the decoded component.
[0228] For example, by default, the encoding end and the decoding end use an interpolation filter with N1 taps for interpolation of prediction values, gradients, etc. in the x direction, and use an interpolation filter with N2 taps for interpolation of prediction values, gradients, etc. in the y direction. (Both N1 and N2 are positive integers, and N1 ≠ N2).
[0229] For example, if the interpolation direction is the x direction and the decoded component is the luminance component, then the interpolation filter with k1 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with k1 taps for gradient calculation and prediction value interpolation;
[0230] If the interpolation direction is the x direction and the decoded component is the chrominance component, then the interpolation filter with k2 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with k2 taps for gradient calculation and prediction value interpolation;
[0231] If the interpolation direction is the y direction and the decoded component is the luminance component, then the interpolation filter with k3 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with k3 taps for gradient calculation and prediction value interpolation;
[0232] If the interpolation direction is the y direction and the decoded component is the chrominance component, then the interpolation filter with k4 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with k4 taps for gradient calculation and prediction value interpolation;
[0233] Among them, k1, k2, k3, and k4 are all positive integers.
[0234] The embodiments of the present application do not limit the specific values of k1, k2, k3, and k4 above.
[0235] In one example, k1 = 12, k2 = 6, k3 = 8, k4 = 4.
[0236] A specific implementation example is shown in Table 9:
[0237] Table 9
[0238]
[0239]
[0240] Exemplarily, an interpolation filter that can also be used by swapping the x-direction and the y-direction can be used, that is, an interpolation filter with a shorter number of taps (such as 8 taps and 4 taps) is used in the x-direction, and an interpolation filter with a longer number of taps (such as 12 taps and 6 taps) is used in the y-direction.
[0241] Example 4: The decoding end determines the target interpolation filter based on the decoded component and the size of the area to be decoded.
[0242] For example, both the encoding end and the decoding end default to using an interpolation filter with N1 taps within a certain block size range, and using an interpolation filter with N2 taps within other block size ranges (both N1 and N2 are positive integers, and N1 ≠ N2).
[0243] For illustration, if the decoded component is a luminance component, and the scale of the area to be decoded is w <= W1 or h <= H1, then the interpolation filter with a1 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with a1 taps for gradient calculation and prediction value interpolation;
[0244] If the decoded component is a luminance component, and the scale of the area to be decoded is w > W1 and h > H1, then the interpolation filter with a2 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with a2 taps for gradient calculation and prediction value interpolation;
[0245] If the decoded component is a chrominance component, and the scale of the area to be decoded is w <= W2 or h <= H2, then the interpolation filter with a3 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with a3 taps for gradient calculation and prediction value interpolation;
[0246] If the decoded component is a chrominance component, and the scale of the area to be decoded is w > W2 and h > H2, then the interpolation filter with a4 taps is determined as the target interpolation filter. At this time, the decoding end uses the interpolation filter with a4 taps for gradient calculation and prediction value interpolation;
[0247] Among them, a1, a2, a3, and a4 are all positive integers.
[0248] The embodiments of the present application do not limit the specific values of a1, a2, a3, and a4 described above.
[0249] In one example, a1 = 8, a2 = 12, a3 = 4, a4 = 6.
[0250] A specific implementation example is shown in Table 10:
[0251] Table 10
[0252]
[0253] Wherein, w and h represent the block size of the area to be decoded under the current component, and W1, H1, W2, and H2 are all positive integers.
[0254] After the decoding end determines the target interpolation filter based on the above steps, it executes the following steps of S103.
[0255] S103: Determine the reference area of the area to be decoded, and based on the target interpolation filter, interpolate the reference area to determine the predicted value of the area to be decoded.
[0256] After the decoding end determines the target interpolation filter corresponding to the area to be decoded based on the above steps, it uses the target interpolation filter to interpolate the reference area of the area to be decoded to determine the predicted value of the area to be decoded.
[0257] The embodiments of the present application do not limit the sequence of the decoding end to determine the reference area of the area to be decoded and the target interpolation filter corresponding to the area to be decoded. That is to say, the decoding end can first determine the target interpolation filter corresponding to the area to be decoded, and then determine the reference area of the area to be decoded. Or, the decoding end can first determine the reference area of the area to be decoded, and then determine the target interpolation filter corresponding to the area to be decoded. Or, the decoding end can determine the target interpolation filter corresponding to the area to be decoded while determining the reference area of the area to be decoded.
[0258] In the embodiments of the present application, the reference area of the area to be decoded can be a reference block or a reference frame, and the embodiments of the present application do not limit this. Among them, the specific manner for the decoding end to determine the reference area of the area to be decoded can refer to the description of related technologies and will not be elaborated here.
[0259] Then, the decoding end interpolates the reference area of the area to be decoded based on the determined target interpolation filter above, and further determines the predicted value of the area to be decoded.
[0260] In one example, if the above interpolation target is to interpolate the pixel predicted value, the decoding end uses the target interpolation filter for interpolation to obtain the interpolated reference area, and further obtains the predicted block of the area to be decoded in the interpolated reference area based on the motion vector corresponding to the area to be decoded.
[0261] In one example, if the prediction mode corresponding to the area to be decoded in the embodiments of the present application is BIO, the decoding end first adopts the bidirectional prediction mode to determine a forward reference block and a backward reference block of the area to be decoded. Then, it samples the target interpolation filter to interpolate the forward reference block and the backward reference block to determine the gradients of the forward reference block and the backward reference block, and further corrects the motion vectors corresponding to the forward reference block and the backward reference block based on the gradients to obtain the predicted value of the area to be decoded.
[0262] In some embodiments, the bit widths of the filter coefficients of the interpolation filters with different numbers of taps in the embodiments of the present application may be different. And / or the bit widths of the filter coefficients of the interpolation filters with the same number of taps may also be different.
[0263] In some embodiments, the embodiments of the present application also provide the coefficients of some interpolation filters with different numbers of taps. Specifically as follows, where MV position can be understood as the position of a pixel point, and coefficients are the filter coefficients. In some embodiments, the interpolation filter for determining the gradient is also referred to as a gradient filter or a gradient interpolation filter.
[0264] Table 11: Coefficients of the 8-tap gradient filter used in the BIO mode
[0265] MV position coefficients 0 –4,11,–39,–1,41,–14,8,–2 1 / 4 –2,6,–19,–31,53,–12,7,–2 1 / 2 0,–1,0,–50,50,0,1,0 3 / 4 2,–7,12,–53,31,19,–6,2
[0266] Table 12: Coefficients of the 12-tap filter used in the BIO mode and non-BIO modes other than the AFFINE mode
[0267]
[0268]
[0269] Table 13: Coefficients of the 12-tap filter used in the AFFINE mode
[0270] MV position coefficients 0 0,0,0,0,0,256,0,0,0,0,0,0 1 / 16 -1,2,-3,6,-14,254,16,-7,4,-2,1,0 2 / 16 -1,3,-7,12,-26,249,35,-15,8,-4,2,0 3 / 16 -2,5,-9,17,-36,241,54,-22,12,-6,3,-1 4 / 16 -2,5,-11,21,-43,230,75,-29,15,-8,4,-1 5 / 16 -2,6,-13,24,-48,216,97,-36,19,-10,4,-1 6 / 16 -2,7,-14,25,-51,200,119,-42,22,-12,5,-1 7 / 16 -2,7,-14,26,-51,181,140,-46,24,-13,6,-2 8 / 16 -2,6,-13,25,-50,162,162,-50,25,-13,6,-2 9 / 16 -2,6,-13,24,-46,140,181,-51,26,-14,7,-2 10 / 16 -1,5,-12,22,-42,119,200,-51,25,-14,7,-2 11 / 16 -1,4,-10,19,-36,97,216,-48,24,-13,6,-2 12 / 16 -1,4,-8,15,-29,75,230,-43,21,-11,5,-2 13 / 16 -1,3,-6,12,-22,54,241,-36,17,-9,5,-2 14 / 16 0,2,-4,8,-15,35,249,-26,12,-7,3,-1 15 / 16 0,1,-2,4,-7,16,254,-14,6,-3,2,-1
[0271] Table 14: Coefficients of the 6-tap filter used in the BIO mode and non-BIO modes other than the AFFINE mode
[0272] MV position coefficients 0 0,0,256,0,0,0 1 / 8 4,-21,248,33,-10,2 2 / 8 8,-35,227,73,-22,5 3 / 8 10,-42,196,117,-34,9 4 / 8 10,-40,158,158,-40,10 5 / 8 9,-34,117,196,-42,10 6 / 8 5,-22,73,227,-35,8 7 / 8 2,-10,33,248,-21,4
[0273] Table 15: Coefficients of the 6-tap filter used in the AFFINE mode
[0274]
[0275]
[0276] The video decoding method provided by the embodiments of the present application obtains decoding information of a region to be decoded by decoding a bitstream. The decoding information includes at least one of a prediction mode, an interpolation direction, a decoded component, and an interpolation target. Based on the decoding information, a target interpolation filter corresponding to the region to be decoded is determined. A reference region of the region to be decoded is determined, and interpolation is performed on the reference region based on the target interpolation filter to determine a predicted value of the region to be decoded. That is to say, in the embodiments of the present application, the decoding end adaptively selects a target interpolation filter from interpolation filters with different tap numbers based on the encoding and decoding information of the region to be encoded and decoded, such as prediction mode, interpolation direction, decoded component, interpolation target, etc., improving the selection accuracy of the target interpolation filter. Thus, when interpolating the reference region of the region to be encoded and decoded based on the accurately selected target interpolation filter to determine the predicted value of the region to be encoded and decoded, the prediction effect of the region to be encoded and decoded can be improved, thereby enhancing the performance of video encoding and decoding.
[0277] The video decoding method of the present application is introduced above with the decoding end as an example, and the following is an illustration with the encoding end as an example.
[0278] Figure 6 It is a schematic flowchart of a video encoding method provided by an embodiment of the present application. The embodiments of the present application are applied to Figure 1 and Figure 2 the video encoder shown. As Figure 6 shown, the method of the embodiments of the present application includes:
[0279] S201. Obtain encoding information of a region to be encoded and a reference region corresponding to the region to be encoded.
[0280] Among them, the encoding information includes at least one of a prediction mode, an interpolation direction, an encoding component, and an interpolation target.
[0281] The embodiments of the present application do not limit the specific size and shape of the region to be encoded.
[0282] In some embodiments, the region to be encoded may be one or several image blocks to be encoded in the current frame to be encoded. For example, it may be one or several CTUs, or one or several CUs, or one or several PUs, etc. In one example, the image block to be encoded may be referred to as a current block, the current image block to be encoded, and so on.
[0283] In some embodiments, the region to be encoded may be the current frame to be encoded. That is to say, the region to be encoded is a whole-frame image.
[0284] During the video encoding process, when the encoding end encodes the area to be encoded, it determines the prediction mode of the area to be encoded. Based on the prediction mode, it determines the prediction value of the area to be encoded, and then based on the prediction value of the area to be encoded, it determines the residual value of the area to be encoded. Then, the residual value is transformed to obtain transform coefficients, the transform coefficients are quantized to obtain quantization coefficients, and finally the quantization coefficients are encoded to obtain a bitstream.
[0285] The embodiments of the present application mainly relate to the prediction process of the area to be encoded.
[0286] In the embodiments of the present application, in order to improve the prediction accuracy of the area to be encoded, an interpolation filter is used to interpolate the reference area of the area to be encoded. However, currently, the encoding information of the area to be encoded is not considered when selecting the interpolation filter, resulting in inaccurate selection of the interpolation filter.
[0287] To solve the above technical problems, in the embodiments of the present application, based on the encoding information of the area to be encoded, the target interpolation filter of the area to be encoded is determined to improve the selection accuracy of the interpolation filter.
[0288] The embodiments of the present application do not limit the specific content of the encoding information of the area to be encoded, which can be understood as all encoding information related to encoding the area to be encoded, such as prediction mode, transformation method, quantization method, interpolation direction, encoding component, and interpolation target, etc.
[0289] In some embodiments, the encoding information of the area to be encoded in the embodiments of the present application includes at least one of the prediction mode of the area to be encoded, the interpolation direction to be interpolated, the encoding component to be encoded, and the interpolation target.
[0290] Exemplarily, the prediction mode of the area to be encoded may include various modes in the intra prediction mode, various modes in the inter prediction mode, and various modes in the hybrid prediction mode, etc.
[0291] Exemplarily, the interpolation direction includes the x direction and the y direction.
[0292] Exemplarily, the encoding component includes a luminance component and a chrominance component.
[0293] Exemplarily, the interpolation target includes determining a prediction value (for example, performing pixel interpolation filtering on the reference area of the area to be encoded to determine the prediction value of the area to be encoded), and determining a gradient (for example, determining the gradient of the forward reference block and the backward reference block in the x direction or y direction in BIO).
[0294] In one example, the coding information of the to-be-coded region described above may include only any one of a prediction mode, an interpolation direction, a coding component, and an interpolation target. For example, the coding information of the to-be-coded region includes a prediction mode, and the prediction mode is prediction mode 1. In this way, the coding end can determine the target interpolation filter from the interpolation filter corresponding to this prediction mode.
[0295] In one example, the coding information of the to-be-coded region described above may include any two or three of a prediction mode, an interpolation direction, a coding component, and an interpolation target. For example, the coding information of the to-be-coded region includes a prediction mode and an interpolation direction, the prediction mode is prediction mode 1, and the interpolation direction is the x direction. In this way, the coding end can determine the target interpolation filter based on prediction mode 1 and the interpolation filter corresponding to the x direction.
[0296] In one example, the coding information of the to-be-coded region described above may include a prediction mode, an interpolation direction, a coding component, and an interpolation target. The embodiments of the present application do not limit this. For example, the coding information of the to-be-coded region includes that the prediction mode is prediction mode 1, the interpolation direction is the x direction, the coding component is a luminance component, and the interpolation target is to determine the predicted value of the to-be-coded region. In this way, the coding end can determine the target interpolation filter based on the interpolation filter corresponding to the prediction mode, the interpolation direction, the coding component, and the interpolation target.
[0297] In some embodiments, the coding end may also obtain at least one of the prediction mode, the interpolation direction, the coding component, and the interpolation target corresponding to the to-be-coded region in other ways.
[0298] In the embodiments of the present application, the reference region of the to-be-coded region may be a reference block or a reference frame, and the embodiments of the present application do not limit this. Among them, the specific manner for the coding end to determine the reference region of the to-be-coded region may refer to the description of related technologies and will not be elaborated here.
[0299] After the coding end determines the coding information of the to-be-coded region, it performs the steps of S202 as follows.
[0300] S202: Determine the target interpolation filter corresponding to the to-be-coded region based on the coding information.
[0301] In the embodiments of the present application, the coding end determines the target interpolation filter corresponding to the to-be-coded region from interpolation filters with different tap numbers based on the coding information of the to-be-coded region, fully considering the performance differences of interpolation filters with different tap numbers in different coding modes, different decoding components, different interpolation directions, and different interpolation targets, thereby improving the selection accuracy of the target interpolation filter.
[0302] The embodiments of the present application do not limit the specific manner in which the encoding end determines the target interpolation filter corresponding to the area to be encoded based on at least one of the prediction mode, interpolation direction, encoding component, and interpolation target.
[0303] In some embodiments, the above S202 includes the following steps S202-A1 to S202-A2:
[0304] S202-A1. Determine a plurality of candidate interpolation filters based on the encoding information;
[0305] S202-A2. Select the target interpolation filter from the plurality of candidate interpolation filters.
[0306] In this implementation manner, when the encoding end determines the target interpolation filter corresponding to the area to be encoded based on the encoding information of the area to be encoded determined above, it first determines a plurality of candidate interpolation filters based on the encoding information, and then selects one interpolation filter from these candidate interpolation filters as the target interpolation filter. That is to say, the embodiments of the present application determine a plurality of candidate interpolation filters with different tap numbers based on the encoding information, and then select the target interpolation filter from these candidate interpolation filters with different tap numbers, fully considering the performance differences of the interpolation filters with different tap numbers in different encoding modes, different decoding components, different interpolation directions, and different interpolation targets, thereby improving the determination accuracy of the target interpolation filter.
[0307] The following introduces the specific process of the encoding end determining a plurality of candidate interpolation filters based on the encoding information.
[0308] In the embodiments of the present application, the implementation manners of determining a plurality of candidate interpolation filters based on the encoding information in the above S202-A1 at least include the following several types:
[0309] Method 1. The encoding end determines a plurality of candidate interpolation filters through the following steps S202-A1-a1 and S202-A1-a2:
[0310] S202-A1-a1. Obtain the preset candidate interpolation filters respectively corresponding to at least one of the prediction mode, interpolation direction, encoding component, and interpolation target;
[0311] S202-A1-a2. Determine a plurality of candidate interpolation filters based on the preset candidate interpolation filters.
[0312] In this implementation manner, at least one of the prediction mode, interpolation direction, encoding component, and interpolation target in the encoding information respectively corresponds to a preset candidate interpolation filter.
[0313] In one example, different prediction modes may correspond to different candidate interpolation filters. For example, prediction mode 1 corresponds to m1 candidate interpolation filters, and prediction mode 2 corresponds to m2 candidate interpolation filters. The number of taps of each of the m1 candidate interpolation filters is not exactly the same as that of each of the m2 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the BIO mode include an 8-tap interpolation filter and a 12-tap interpolation filter. The candidate interpolation filters corresponding to the non-BIO mode include a 6-tap interpolation filter and an 8-tap interpolation filter.
[0314] In one example, different interpolation directions may correspond to different candidate interpolation filters. For example, the m3 candidate interpolation filters corresponding to the x direction, and the m4 candidate interpolation filters corresponding to the y direction. The number of taps of each of the m3 candidate interpolation filters is not exactly the same as that of each of the m4 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the x (horizontal) direction include an 8-tap interpolation filter and a 12-tap interpolation filter, and the candidate interpolation filters corresponding to the y (vertical) direction include a 6-tap interpolation filter and an 8-tap interpolation filter.
[0315] In one example, different coding components may correspond to different candidate interpolation filters. For example, the luminance component corresponds to m5 candidate interpolation filters, and the chrominance component corresponds to m6 candidate interpolation filters. The number of taps of each of the m5 candidate interpolation filters is not exactly the same as that of each of the m6 candidate interpolation filters. Exemplarily, the candidate interpolation filters corresponding to the luminance component include an 8-tap interpolation filter and a 12-tap interpolation filter, and the candidate interpolation filters corresponding to the chrominance component include a 4-tap interpolation filter and a 6-tap interpolation filter.
[0316] Exemplarily, different interpolation targets may correspond to different candidate interpolation filters. For example, the m7 candidate interpolation filters corresponding to determining the predicted value, and the m8 candidate interpolation filters corresponding to determining the gradient. The number of taps of each of the m7 candidate interpolation filters is not exactly the same as that of each of the m8 candidate interpolation filters.
[0317] In this way, the encoding end can determine multiple candidate interpolation filters based on the candidate interpolation filters respectively corresponding to at least one of the prediction mode, interpolation direction, coding component, and interpolation target.
[0318] For example, a preset number of candidate interpolation filters are selected from the candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, coding component, and interpolation target respectively as multiple candidate interpolation filters. For example, when the above coding information includes a prediction mode and a coding component, the coding end selects a preset number (for example, n) from the candidate interpolation filters corresponding to the prediction mode and the coding component respectively to form multiple candidate interpolation filters. Referring to the above example, assume that the prediction mode of the area to be coded is prediction mode 1 and the coding component is the luminance component, where prediction mode 1 corresponds to m1 candidate interpolation filters and the luminance component corresponds to m5 candidate interpolation filters. In this way, a preset number of candidate interpolation filters can be selected from the m1 + m5 candidate interpolation filters of the m1 candidate interpolation filters and the m5 candidate interpolation filters to form multiple candidate interpolation filters.
[0319] For another example, the coding end removes the duplicate interpolation filters from the selected prediction candidate interpolation filters to obtain multiple candidate interpolation filters. For example, assume that the coding information includes a prediction mode and a coding component, where the prediction mode is a non-BIO mode and the coding component is the luminance component. Assume that the candidate interpolation filters corresponding to the non-BIO mode include a 6-tap interpolation filter and an 8-tap interpolation filter, and the 8-tap interpolation filter and a 12-tap interpolation filter corresponding to the luminance component. One 8-tap interpolation filter is removed from these 4 interpolation filters, and the remaining 3 interpolation filters are a 6-tap interpolation filter, an 8-tap interpolation filter, and a 12-tap interpolation filter. Then, the 3 interpolation filters of the 6-tap interpolation filter, the 8-tap interpolation filter, and the 12-tap interpolation filter are determined as multiple candidate interpolation filters.
[0320] As can be seen from the above description, in the first method, multiple candidate interpolation filters are selected from the preset candidate interpolation filters corresponding to at least one of the prediction mode, interpolation direction, coding component, and interpolation target.
[0321] In some embodiments, the coding end can also determine multiple candidate interpolation filters by the method shown in Method 2 below.
[0322] Method 2: The coding end selects multiple candidate interpolation filters from the preset M interpolation filters.
[0323] For example, as Figure 7 shown, the M interpolation filters include 6 interpolation filters with 2 taps, 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The coding end can select multiple interpolation filters from these 6 interpolation filters with different tap numbers as multiple candidate interpolation filters.
[0324] In an example of the second method, the encoding end determines multiple candidate interpolation filters through the following steps S202-A1-b1 and S202-A1-b2:
[0325] S202-A1-b1. Obtain the selection strategy corresponding to at least one of the prediction mode, interpolation direction, encoding component, and interpolation target respectively;
[0326] S202-A1-b2. Based on the selection strategy, select multiple candidate interpolation filters from the preset M interpolation filters, where M is a positive integer greater than 1.
[0327] In this implementation, there are M preset interpolation filters. The encoding end can select multiple candidate interpolation filters from these M preset interpolation filters based on at least one of the prediction mode, interpolation direction, encoding component, and interpolation target. Optionally, the number of taps of these M interpolation filters is different.
[0328] Specifically, the encoding end obtains the selection strategy corresponding to at least one of the prediction mode, interpolation direction, encoding component, and interpolation target respectively. This selection strategy is the default for both the encoding and decoding ends, so that the encoding end can select multiple candidate interpolation filters from these M preset interpolation filters based on this selection strategy.
[0329] In this implementation, the selection strategies corresponding to at least one of the above-mentioned prediction mode, interpolation direction, encoding component, and interpolation target can be the same, different, or partially the same and partially different.
[0330] Exemplarily, assume that the above encoding information includes a prediction mode and an encoding component. The selection strategy corresponding to the prediction mode is Selection Strategy 1, and the selection strategy corresponding to the encoding component is Selection Strategy 2. In this way, the encoding end can select one or more interpolation filters from the M interpolation filters based on Selection Strategy 1. At the same time, the encoding end can select one or more interpolation filters from the M interpolation filters based on Selection Strategy 2. Furthermore, the encoding end obtains multiple candidate interpolation filters according to the one or more interpolation filters selected based on Strategy 1 and the one or more interpolation filters selected based on Strategy 2. For example, the encoding end determines the interpolation filters with different numbers of taps among the one or more interpolation filters selected based on Strategy 1 and the one or more interpolation filters selected based on Strategy 2 as multiple candidate interpolation filters.
[0331] In some embodiments, the above S202-A1-b2 includes the following steps S202-A1-b21 and S202-A1-b22:
[0332] S202-A1-b21. Obtain N candidate combinations composed of M interpolation filters, where each candidate combination includes two or more of the M interpolation filters, and N is a positive integer greater than 1;
[0333] S202-A1-b22. Based on a selection strategy, select a target candidate combination from the N candidate combinations, and determine the interpolation filters included in the target candidate combination as multiple candidate interpolation filters.
[0334] In this implementation manner, there are M preset interpolation filters, and the encoding end obtains N candidate combinations composed of these M interpolation filters. Optionally, the number of taps of these M interpolation filters is different.
[0335] In one example, the above N candidate combinations are default for both the encoding and decoding ends. That is, the encoding end itself stores N candidate combinations, or the encoding end can combine the M interpolation filters using the same rules as the decoding end to obtain N candidate combinations.
[0336] The embodiments of the present application do not limit the specific combination form of the N candidate combinations.
[0337] In one possible implementation manner, the interpolation filters included in different candidate combinations among the above N candidate combinations are randomly distributed. For example, assume that the M interpolation filters include interpolation filters with 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The N candidate combinations composed of these 5 types of tap interpolation filters are: {6 taps, 8 taps, 12 taps}, {2 taps, 4 taps, 6 taps}, and {4 taps, 6 taps, 8 taps, 12 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 1.
[0338] In one possible implementation manner, the number of interpolation filters included in each candidate combination among the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination increases in sequence. For example, assume that the M interpolation filters include interpolation filters with 2 taps, 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The N candidate combinations composed of these 6 types of tap interpolation filters include: {2 taps, 4 taps}, {2 taps, 4 taps, 6 taps}, {2 taps, 4 taps, 6 taps, 8 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps, 16 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 2.
[0339] In a possible implementation, the number of interpolation filters included in each candidate combination among the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination decreases in sequence. For example, assume that M interpolation filters include interpolation filters with 6 types of taps, namely 2 taps, 4 taps, 6 taps, 8 taps, 12 taps, and 16 taps. The N candidate combinations composed of these 6 types of tap interpolation filters include: {12 taps, 16 taps}, {8 taps, 12 taps, 16 taps}, {6 taps, 8 taps, 12 taps, 16 taps}, {4 taps, 6 taps, 8 taps, 12 taps, 16 taps}, {2 taps, 4 taps, 6 taps, 8 taps, 12 taps, 16 taps}, etc. Exemplarily, the N candidate combinations are shown in Table 3.
[0340] Next, the encoding end selects a target candidate combination from the N candidate combinations based on the selection strategies respectively corresponding to at least one of the prediction mode, interpolation direction, encoding component, and interpolation target, and then determines the interpolation filters included in the target candidate combination as multiple candidate interpolation filters. The embodiments of the present application do not limit the specific content of the selection strategy.
[0341] In some embodiments, after the encoding end selects a target candidate combination from the N candidate combinations based on the above steps, second indication information is written into the bitstream. The second indication information is used to indicate the target candidate combination. Exemplarily, the second indication information includes the index of the target candidate combination in the N candidate combinations. For example, the second indication information includes index 100. In this way, the decoding end can obtain the second indication information by decoding the bitstream, and then select the target candidate combination from the obtained N candidate combinations based on the second indication information.
[0342] As described above, the encoding end can determine multiple candidate interpolation filters based on the encoding information of the area to be encoded in the above manner. Next, the encoding end executes the steps of S202-A2 above to select a target interpolation filter from the multiple candidate interpolation filters.
[0343] The embodiments of the present application do not limit the specific manner in which the encoding end selects a target interpolation filter from the multiple candidate interpolation filters.
[0344] In a possible implementation, the encoding end randomly selects a target interpolation filter from the multiple candidate interpolation filters. Or based on some feature information (such as scale size and other information) of the area to be decoded, a target interpolation filter is selected from the multiple candidate interpolation filters.
[0345] In a possible implementation, the above S202-A2 includes the following steps of S202-A21 to S202-A23:
[0346] S202-A21. For each candidate interpolation filter among multiple candidate interpolation filters, use the candidate interpolation filter to perform interpolation on the reference region to obtain an interpolated prediction region;
[0347] S202-A22. Based on the interpolated prediction region, determine the rate-distortion cost corresponding to the candidate interpolation filter;
[0348] S202-A23. Based on the rate-distortion cost corresponding to each candidate interpolation filter among multiple candidate interpolation filters, select a target interpolation filter from the multiple candidate interpolation filters.
[0349] In this implementation manner, for each candidate interpolation filter among multiple candidate interpolation filters, for example, the i-th candidate interpolation filter, use the i-th candidate interpolation filter to perform interpolation on the reference region to obtain an interpolated prediction region. Then, based on the interpolated prediction region, determine the rate-distortion cost corresponding to the i-th candidate interpolation filter. The calculation process of the rate-distortion cost can refer to the description of related technologies and will not be elaborated here. In this way, the encoding end can determine the rate-distortion cost corresponding to each candidate interpolation filter among multiple candidate interpolation filters, and then based on the rate-distortion cost corresponding to each candidate interpolation filter among multiple candidate interpolation filters, select a target interpolation filter from the multiple candidate interpolation filters. For example, among multiple candidate interpolation filters, the candidate interpolation filter with the minimum rate-distortion cost is determined as the target interpolation filter.
[0350] Exemplarily, as Figure 8 shown, assume that the multiple candidate interpolation filters include 3 candidate interpolation filters: 6-tap, 8-tap, and 12-tap. Use these 3 candidate interpolation filters to perform interpolation on the reference region of the region to be encoded respectively to obtain the prediction regions corresponding to these 3 candidate interpolation filters respectively. Then, based on the prediction regions corresponding to these 3 candidate interpolation filters respectively, calculate the rate-distortion cost to obtain the rate-distortion costs corresponding to these 3 candidate interpolation filters respectively. Further, the candidate interpolation filter with the minimum rate-distortion cost is determined as the target interpolation filter.
[0351] In some embodiments, after the encoding end selects a target interpolation filter from the multiple candidate interpolation filters based on the rate-distortion cost of each of the multiple candidate interpolation filters, write first indication information in the bitstream. This first indication information is used to indicate the index information of the target interpolation filter among the multiple candidate interpolation filters. In this way, when determining the target interpolation filter, the decoding end determines multiple candidate interpolation filters through the above steps, and at the same time decodes the bitstream to obtain the first indication information. Then, based on this first indication information, select the target interpolation filter from the multiple candidate interpolation filters.
[0352] In some embodiments, the above first indication information includes a flag bit array or a plurality of flag bits, and the flag bit array or the plurality of flag bits are used to indicate coding information and a target interpolation filter.
[0353] In one example, the first indication information indicates a prediction mode, a coding component, and a target interpolation filter. At this time, the first indication information includes a parsed flag bit array Flag1[A][I] (both A and I are positive integers), where A represents a decoded component. For example, A = 0 represents a luminance component, and A = 1 represents a chrominance component. I represents a prediction mode. For example, I = 0 represents mode 1, and I = 1 represents mode 2. Flag1[A][I] represents index information of the target interpolation filter of component A in prediction mode I. Exemplarily, the first indication information and its meaning are shown in Table 4.
[0354] In one example, the first indication information indicates an interpolation direction, a coding component, and a target interpolation filter. At this time, the first indication information includes a parsed flag bit array Flag2[A][B] (both A and B are positive integers), where A represents a decoded component. For example, A = 0 represents a luminance component, and A = 1 represents a chrominance component. B represents an interpolation direction. For example, B = 0 represents the x direction, and B = 1 represents the y direction. Flag2[A][B] represents index information of the target interpolation filter of component A in interpolation direction B. Exemplarily, the first indication information and its meaning are shown in Table 5.
[0355] In one example, the first indication information indicates an interpolation target, a coding component, and a target interpolation filter. At this time, the first indication information includes a parsed flag bit array Flag3[A][D] (both A and D are positive integers), where A represents a decoded component. For example, A = 0 represents a luminance component, and A = 1 represents a chrominance component. D represents an interpolation target. For example, D = 0 represents that the interpolation target is to determine a predicted value, and D = 1 represents that the interpolation target is to determine a gradient. Flag3[A][D] represents index information of the target interpolation filter of component A for interpolation target D. Exemplarily, the first indication information and its meaning are shown in Table 6.
[0356] In this implementation manner, the encoding end determines a plurality of candidate interpolation filters with different tap numbers based on at least one of the prediction mode, interpolation direction, coding component, and interpolation target corresponding to the area to be encoded, and then selects a target interpolation filter from these plurality of candidate interpolation filters with different tap numbers.
[0357] In some embodiments, the encoding end and the decoding end can determine a target interpolation filter corresponding to the area to be encoded based on a default interpolation filter corresponding to at least one of the prediction mode, interpolation direction, coding component, and interpolation target corresponding to the area to be encoded.
[0358] Example 1: The encoding end determines a target interpolation filter based on a prediction mode, an encoding component, and an interpolation target.
[0359] For example, both the encoding end and the decoding end default to using a specific N1-tap interpolation filter for luminance and / or chrominance for a certain j1 interpolation targets in a certain i prediction mode, and using a certain N2-tap interpolation filter for luminance and / or chrominance for a certain j2 interpolation contents in the remaining I - i encoding modes, where i, j1, j2, N1, and N2 are all integers, and N1 ≠ N2.
[0360] Illustratively, if the prediction mode is a bidirectional optical flow prediction mode, the encoding component is a luminance component, and the interpolation target is to determine the gradients of the forward prediction block and the backward prediction block of the area to be encoded in the x direction and the y direction, then an n1-tap interpolation filter is determined as the target interpolation filter;
[0361] If the prediction mode is a bidirectional optical flow prediction mode, the encoding component is a luminance component, and the interpolation target is to determine the predicted values of the area to be encoded in the x direction and the y direction, then an n2-tap interpolation filter is determined as the target interpolation filter;
[0362] If the prediction mode is a bidirectional optical flow prediction mode, the encoding component is a chrominance component, and the interpolation target is to determine the predicted values of the area to be encoded in the x direction and the y direction, then an n3-tap interpolation filter is determined as the target interpolation filter;
[0363] If the prediction mode is a non-bidirectional optical flow prediction mode, the encoding component is a luminance component, and the interpolation target is to determine the predicted values of the area to be encoded in the x direction and the y direction, then an n4-tap interpolation filter is determined as the target interpolation filter;
[0364] If the prediction mode is a non-bidirectional optical flow prediction mode, the encoding component is a chrominance component, and the interpolation target is to determine the predicted values of the area to be encoded in the x direction and the y direction, then an n5-tap interpolation filter is determined as the target interpolation filter;
[0365] Among them, n1, n2, n3, n4, and n5 are all positive integers.
[0366] The embodiments of the present application do not limit the specific values of the above n1, n2, n3, n4, and n5.
[0367] In a possible implementation, the above n1 is not equal to the above n2.
[0368] In an example, n1 = 8, n2 = 12, n3 = 6, n4 = 12, n5 = 6.
[0369] As can be seen from Table 7 above, in the BIO mode of this application embodiment, the number of taps of the interpolation filter used when calculating the gradient is different from the number of taps of the interpolation filter used when calculating the pixel prediction value. For example, an 8-tap interpolation filter is used to calculate the gradient, and a 12-tap interpolation filter is used for pixel prediction value interpolation calculation.
[0370] Example 2: The encoding end determines the target interpolation filter based on the prediction mode and the encoded component.
[0371] For example, both the encoding end and the encoding end default to using a specific interpolation filter with N1 taps only for luminance and / or chrominance in a certain i prediction modes, and using an interpolation filter with N2 taps for luminance and / or chrominance in the remaining I - i prediction modes (i, N1, and N2 are all integers, and N1 ≠ N2).
[0372] For example, if the prediction mode is the bidirectional optical flow prediction mode and the encoded component is the luminance component, then the interpolation filter with m1 taps is determined as the target interpolation filter. At this time, the encoding end uses the interpolation filter with m1 taps for gradient calculation and prediction value interpolation;
[0373] If the prediction mode is the bidirectional optical flow prediction mode and the encoded component is the chrominance component, then the interpolation filter with m2 taps is determined as the target interpolation filter. At this time, the encoding end uses the interpolation filter with m2 taps for prediction value interpolation calculation in the x direction and the y direction;
[0374] If the prediction mode is the non - bidirectional optical flow prediction mode and the encoded component is the luminance component, then the interpolation filter with m3 taps is determined as the target interpolation filter. At this time, the encoding end uses the interpolation filter with m3 taps for prediction value interpolation calculation in the x direction and the y direction;
[0375] If the prediction mode is the non - bidirectional optical flow prediction mode and the encoded component is the chrominance component, then the interpolation filter with m4 taps is determined as the target interpolation filter. At this time, the encoding end uses the interpolation filter with m4 taps for prediction value interpolation calculation in the x direction and the y direction;
[0376] Among them, m1, m2, m3, and m4 are all positive integers.
[0377] This application embodiment does not limit the specific values of m1, m2, m3, and m4 above.
[0378] In one example, m1 = 8, m2 = 6, m3 = 12, m4 = 6.
[0379] Example 3: The encoding end determines the target interpolation filter based on the interpolation direction and the encoded component.
[0380] For example, by default, the encoding end uses an interpolation filter with N1 taps for interpolation of prediction values, gradients, etc. in the x direction, and an interpolation filter with N2 taps for interpolation of prediction values, gradients, etc. in the y direction. (Both N1 and N2 are positive integers, and N1 ≠ N2).
[0381] For example, if the interpolation direction is the x direction and the encoding component is the luminance component, the interpolation filter with k1 taps is determined as the target interpolation filter. At this time, the encoding end uses the interpolation filter with k1 taps for gradient calculation and prediction value interpolation;
[0382] If the interpolation direction is the x direction and the encoding component is the chrominance component, the interpolation filter with k2 taps is determined as the target interpolation filter. At this time, the encoding end uses the interpolation filter with k2 taps for gradient calculation and prediction value interpolation;
[0383] If the interpolation direction is the y direction and the encoding component is the luminance component, the interpolation filter with k3 taps is determined as the target interpolation filter. At this time, the encoding end uses the interpolation filter with k3 taps for gradient calculation and prediction value interpolation;
[0384] If the interpolation direction is the y direction and the encoding component is the chrominance component, the interpolation filter with k4 taps is determined as the target interpolation filter. At this time, the encoding end uses the interpolation filter with k4 taps for gradient calculation and prediction value interpolation;
[0385] Among them, k1, k2, k3, and k4 are all positive integers.
[0386] The embodiments of the present application do not limit the specific values of k1, k2, k3, and k4 described above.
[0387] In one example, k1 = 12, k2 = 6, k3 = 8, k4 = 4.
[0388] Example 4: The encoding end determines the target interpolation filter based on the encoding component and the size of the area to be encoded.
[0389] For example, by default, the encoding end uses an interpolation filter with N1 taps within a certain block size range and an interpolation filter with N2 taps within other block size ranges. (Both N1 and N2 are positive integers, and N1 ≠ N2).
[0390] For example, if the encoding component is the luminance component and the scale of the area to be encoded is w <= W1 or h <= H1, the interpolation filter with a1 taps is determined as the target interpolation filter. At this time, the encoding end uses the interpolation filter with a1 taps for gradient calculation and prediction value interpolation;
[0391] If the coded component is a luminance component and the scale of the area to be coded is such that w > W1 and h > H1, then the interpolation filter with a2 taps is determined as the target interpolation filter. At this time, the coding end uses the interpolation filter with a2 taps for gradient calculation and predicted value interpolation.
[0392] If the coded component is a chrominance component and the scale of the area to be coded is such that w <= W2 or h <= H2, then the interpolation filter with a3 taps is determined as the target interpolation filter. At this time, the coding end uses the interpolation filter with a3 taps for gradient calculation and predicted value interpolation.
[0393] If the coded component is a chrominance component and the scale of the area to be coded is such that w > W2 and h > H2, then the interpolation filter with a4 taps is determined as the target interpolation filter. At this time, the coding end uses the interpolation filter with a4 taps for gradient calculation and predicted value interpolation.
[0394] Among them, a1, a2, a3, and a4 are all positive integers.
[0395] The embodiments of the present application do not limit the specific values of the above a1, a2, a3, and a4.
[0396] In one example, a1 = 8, a2 = 12, a3 = 4, and a4 = 6.
[0397] After the coding end determines the target interpolation filter based on the above steps, it executes the following step S203.
[0398] S203: Interpolate the reference area based on the target interpolation filter to determine the predicted value of the area to be coded.
[0399] After the coding end determines the target interpolation filter corresponding to the area to be coded based on the above steps, it uses the target interpolation filter to interpolate the reference area of the area to be coded to determine the predicted value of the area to be coded.
[0400] In one example, if the above interpolation target is to interpolate the pixel predicted value, then the coding end uses the target interpolation filter to interpolate the reference area to obtain the interpolated predicted area, and then searches in the interpolated predicted area to obtain the predicted block of the area to be coded.
[0401] In one example, if the prediction mode corresponding to the area to be coded in the embodiments of the present application is BIO, then the coding end first uses the bidirectional prediction mode to determine a forward reference block and a backward reference block of the area to be coded. Then, it samples the target interpolation filter to interpolate the forward reference block and the backward reference block to determine the gradients of the forward reference block and the backward reference block, and then corrects the motion vectors corresponding to the forward reference block and the backward reference block based on the gradients to obtain the predicted value of the area to be coded.
[0402] The video decoding method provided by the embodiment of the present application obtains the encoding information of the area to be encoded and the reference area corresponding to the area to be encoded, where the encoding information includes at least one of a prediction mode, an interpolation direction, an encoding component, and an interpolation target; based on the encoding information, determines the target interpolation filter corresponding to the area to be encoded, and then based on the target interpolation filter, interpolates the reference area to determine the predicted value of the area to be encoded. That is to say, in the embodiment of the present application, the encoding end adaptively selects the target interpolation filter from interpolation filters with different numbers of taps based on the encoding information of the area to be encoded and decoded, such as prediction mode, interpolation direction, encoding component, interpolation target, etc., improving the selection accuracy of the target interpolation filter. Thus, when interpolating the reference area of the area to be encoded based on the accurately selected target interpolation filter to determine the predicted value of the area to be encoded, the prediction effect of the area to be encoded can be improved, thereby enhancing the performance of video encoding.
[0403] As described above in conjunction with Figures 5 to 8 , the embodiments of the audio encoding and decoding method of the present application are described in detail. Below in conjunction with Figures 9 to 10 , the embodiments of the device of the present application are described in detail.
[0404] Figure 9 FIG. is a schematic block diagram of a video decoding device provided by an embodiment of the present application. The device 10 can be applied to a decoding device.
[0405] As Figure 9 shown, the video decoding device 10 includes:
[0406] A decoding unit 11, configured to decode a bitstream to obtain decoding information of the area to be decoded, where the decoding information includes at least one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target;
[0407] A determination unit 12, configured to determine the target interpolation filter corresponding to the area to be decoded based on the decoding information;
[0408] An interpolation unit 13, configured to determine the reference area of the area to be decoded, and based on the target interpolation filter, interpolate the reference area to determine the predicted value of the area to be decoded.
[0409] In some embodiments, the interpolation unit 13 is specifically configured to determine a plurality of candidate interpolation filters based on the decoding information; decode the bitstream to obtain first indication information of the target interpolation filter, where the first indication information is used to indicate the index information of the target interpolation filter among the plurality of candidate interpolation filters; and select the target interpolation filter from the plurality of candidate interpolation filters based on the first indication information.
[0410] In some embodiments, the interpolation unit 13 is specifically configured to obtain preset candidate interpolation filters respectively corresponding to at least one of the prediction mode, the interpolation direction, the decoded component, and the interpolation target; and determine the plurality of candidate interpolation filters based on the preset candidate interpolation filters.
[0411] In some embodiments, the interpolation unit 13 is specifically configured to remove duplicate interpolation filters from the preset candidate interpolation filters to obtain the plurality of candidate interpolation filters.
[0412] In some embodiments, the interpolation unit 13 is specifically configured to obtain a selection strategy respectively corresponding to at least one of the prediction mode, the interpolation direction, the decoded component, and the interpolation target, and select the plurality of candidate interpolation filters from M preset interpolation filters based on the selection strategy, where M is a positive integer greater than 1; or obtain N candidate combinations formed by the M interpolation filters, decode the code stream to obtain second indication information, and select the target candidate combination from the N candidate combinations based on the second indication information, and then determine the interpolation filters included in the target candidate combination as the plurality of candidate interpolation filters, where each candidate combination includes two or more of the M interpolation filters, the second indication information is used to indicate the target candidate combination, and N is a positive integer greater than 1.
[0413] In some embodiments, the interpolation filters included in different candidate combinations among the N candidate combinations are randomly distributed, or the number of interpolation filters included in each candidate combination among the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination increases or decreases in sequence.
[0414] In some embodiments, if the first indication information includes a flag bit array or a plurality of flag bits, and the flag bit array or the plurality of flag bits are used to indicate the decoded information and the target interpolation filter, the decoding unit 11 is further configured to decode the code stream to obtain the first indication information; and obtain the decoded information of the area to be decoded based on the first indication information.
[0415] In some embodiments, the interpolation unit 13 is specifically configured to determine the target interpolation filter based on the prediction mode, the decoded component, and the interpolation target; or determine the target interpolation filter based on the prediction mode and the decoded component; or determine the target interpolation filter based on the interpolation direction and the decoded component; or determine the target interpolation filter according to the decoded component and the size of the area to be decoded.
[0416] In some embodiments, the interpolation unit 13 is specifically configured to determine the n1-tap interpolation filter as the target interpolation filter when the prediction mode is the bidirectional optical flow prediction mode, the decoded component is the luminance component, and the interpolation target is to determine the gradients of the forward prediction block and the backward prediction block of the to-be-decoded region in the x direction and the y direction; determine the n2-tap interpolation filter as the target interpolation filter when the prediction mode is the bidirectional optical flow prediction mode, the decoded component is the luminance component, and the interpolation target is to determine the prediction values of the to-be-decoded region in the x direction and the y direction; determine the n3-tap interpolation filter as the target interpolation filter when the prediction mode is the bidirectional optical flow prediction mode, the decoded component is the chrominance component, and the interpolation target is to determine the prediction values of the to-be-decoded region in the x direction and the y direction; determine the n4-tap interpolation filter as the target interpolation filter when the prediction mode is a non-bidirectional optical flow prediction mode, the decoded component is the luminance component, and the interpolation target is to determine the prediction values of the to-be-decoded region in the x direction and the y direction; determine the n5-tap interpolation filter as the target interpolation filter when the prediction mode is a non-bidirectional optical flow prediction mode, the decoded component is the chrominance component, and the interpolation target is to determine the prediction values of the to-be-decoded region in the x direction and the y direction; where n1, n2, n3, n4, and n5 are all positive integers.
[0417] In some embodiments, n1 is not equal to n2.
[0418] In some embodiments, n1 = 8, n2 = 12, n3 = 6, n4 = 12, and n5 = 6.
[0419] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they are not elaborated here. Specifically, Figure 9 The apparatus shown can execute the embodiments of the above video decoding method, and the foregoing and other operations and / or functions of each module in the apparatus are respectively for implementing the above method embodiments. For the sake of brevity, they are not elaborated here.
[0420] Figure 10 is a schematic block diagram of a video encoding apparatus provided by an embodiment of the present application. The apparatus 20 can be applied to an encoding device.
[0421] As Figure 10 shown, the video encoding apparatus 20 includes:
[0422] An acquisition unit 21, configured to acquire encoding information of a to-be-encoded region and a reference region corresponding to the to-be-encoded region, where the encoding information includes at least one of a prediction mode, an interpolation direction, an encoding component, and an interpolation target;
[0423] A determination unit 22, configured to determine a target interpolation filter corresponding to the area to be encoded based on the encoded information;
[0424] An interpolation unit 23, configured to perform interpolation filtering on the reference area based on the target interpolation filter to determine a predicted value of the area to be encoded.
[0425] In some embodiments, the interpolation unit 23 is specifically configured to determine a plurality of candidate interpolation filters based on the encoded information; and select the target interpolation filter from the plurality of candidate interpolation filters.
[0426] In some embodiments, the interpolation unit 23 is specifically configured to obtain preset candidate interpolation filters respectively corresponding to at least one of the prediction mode, the interpolation direction, the encoded component, and the interpolation target; and determine the plurality of candidate interpolation filters based on the preset candidate interpolation filters.
[0427] In some embodiments, the interpolation unit 23 is specifically configured to remove duplicate interpolation filters from the prediction candidate interpolation filters to obtain the plurality of candidate interpolation filters.
[0428] In some embodiments, the interpolation unit 23 is specifically configured to obtain a selection strategy respectively corresponding to at least one of the prediction mode, the interpolation direction, the encoded component, and the interpolation target; and select the plurality of candidate interpolation filters from M preset interpolation filters based on the selection strategy, where M is a positive integer greater than 1.
[0429] In some embodiments, the interpolation unit 23 is specifically configured to obtain N candidate combinations formed by the M interpolation filters, each candidate combination includes two or more interpolation filters among the M interpolation filters, and N is a positive integer greater than 1; select a target candidate combination from the N candidate combinations based on the selection strategy, and determine the interpolation filters included in the target candidate combination as the plurality of candidate interpolation filters.
[0430] In some embodiments, the interpolation filters included in different candidate combinations among the N candidate combinations are randomly distributed, or the number of interpolation filters included in each candidate combination among the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination increases or decreases in sequence.
[0431] In some embodiments, the interpolation unit 23 is specifically configured to, for each of the multiple candidate interpolation filters, perform interpolation on the reference region using the candidate interpolation filter to obtain an interpolated prediction region; determine the rate-distortion cost corresponding to the candidate interpolation filter based on the interpolated prediction region; and select the target interpolation filter from the multiple candidate interpolation filters based on the rate-distortion cost corresponding to each of the multiple candidate interpolation filters.
[0432] In some embodiments, the interpolation unit 23 is further configured to write first indication information in the bitstream, where the first indication information is used to indicate the index information of the target interpolation filter among the multiple candidate interpolation filters.
[0433] In some embodiments, the first indication information includes a flag bit array or multiple flag bits, and the flag bit array or multiple flag bits are used to indicate the coding information and the target interpolation filter.
[0434] In some embodiments, the interpolation unit 23 is further configured to write second indication information in the bitstream, where the second indication information is used to indicate the target candidate combination.
[0435] In some embodiments, the interpolation unit 23 is specifically configured to determine the target interpolation filter based on the prediction mode, the coding component, and the interpolation target; or determine the target interpolation filter based on the prediction mode and the coding component; or determine the target interpolation filter based on the interpolation direction and the coding component; or determine the target interpolation filter according to the coding component and the size of the region to be coded.
[0436] In some embodiments, the interpolation unit 23 is further configured to, when the prediction mode is the bidirectional optical flow prediction mode, the encoded component is the luminance component, and the interpolation target is to determine the gradients of the forward prediction block and the backward prediction block of the to-be-encoded region in the x direction and the y direction, determine the interpolation filter with n1 taps as the target interpolation filter; when the prediction mode is the bidirectional optical flow prediction mode, the encoded component is the luminance component, and the interpolation target is to determine the predicted values of the to-be-encoded region in the x direction and the y direction, determine the interpolation filter with n2 taps as the target interpolation filter; when the prediction mode is the bidirectional optical flow prediction mode, the encoded component is the chrominance component, and the interpolation target is to determine the predicted values of the to-be-encoded region in the x direction and the y direction, determine the interpolation filter with n3 taps as the target interpolation filter; when the prediction mode is a non-bidirectional optical flow prediction mode, the encoded component is the luminance component, and the interpolation target is to determine the predicted values of the to-be-encoded region in the x direction and the y direction, determine the interpolation filter with n4 taps as the target interpolation filter; when the prediction mode is a non-bidirectional optical flow prediction mode, the encoded component is the chrominance component, and the interpolation target is to determine the predicted values of the to-be-encoded region in the x direction and the y direction, determine the interpolation filter with n5 taps as the target interpolation filter; where n1, n2, n3, n4, and n5 are all positive integers.
[0437] In some embodiments, n1 is not equal to n2.
[0438] In some embodiments, n1 = 8, n2 = 12, n3 = 6, n4 = 12, and n5 = 6.
[0439] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they are not elaborated here. Specifically, Figure 10 The apparatus shown can execute the embodiments of the above video encoding method, and the foregoing and other operations and / or functions of each module in the apparatus are respectively for implementing the above method embodiments. For the sake of brevity, they are not elaborated here.
[0440] The device according to the embodiments of the present application has been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.
[0441] Figure 11 is a schematic block diagram of an electronic device provided by an embodiment of the present application, Figure 11 The electronic device can be the above-mentioned encoding device or decoding device.
[0442] Such as Figure 11 As shown, the electronic device 30 may include:
[0443] A memory 31 and a processor 32. The memory 31 is used to store a computer program 33 and transmit the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiments of the present application.
[0444] For example, the processor 32 can be used to execute the steps in the above method 200 according to the instructions in the computer program 33.
[0445] In some embodiments of the present application, the processor 32 may include, but is not limited to:
[0446] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.
[0447] In some embodiments of the present application, the memory 31 includes, but is not limited to:
[0448] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double DataRate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0449] In some embodiments of the present application, the computer program 33 can be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to complete the method for recording a page provided by the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 33 in the electronic device.
[0450] As Figure 11 shown, the electronic device 30 may further include:
[0451] A transceiver 34, which can be connected to the processor 32 or the memory 31.
[0452] Among them, the processor 32 can control the transceiver 34 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 34 can include a transmitter and a receiver. The transceiver 34 may further include an antenna, and the number of antennas can be one or more.
[0453] It should be understood that the various components in the computing device 30 are connected through a bus system, where the bus system includes, in addition to a data bus, a power bus, a control bus, and a status signal bus.
[0454] According to one aspect of the present application, there is provided a computer storage medium having a computer program stored thereon, and when the computer program is executed by a computer, the computer is enabled to execute the method of the above method embodiment. Or rather, the embodiment of the present application further provides a computer program product including instructions, and when the instructions are executed by a computer, the computer is enabled to execute the method of the above method embodiment.
[0455] According to another aspect of the present application, there is provided a computer program product or a computer program, and the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the method of the above method embodiment.
[0456] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server, a data center, etc. that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0457] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0458] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.
[0459] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network elements. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the functional modules can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0460] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A video decoding method, characterized in that, it includes: decoding a bitstream to obtain decoding information of a region to be decoded, where the decoding information includes at least one of a prediction mode, an interpolation direction, a decoded component, and an interpolation target; determining a target interpolation filter corresponding to the region to be decoded based on the decoding information; determining a reference region of the region to be decoded, and performing interpolation on the reference region based on the target interpolation filter to determine a predicted value of the region to be decoded.
2. The method according to claim 1, characterized in that, the determining the target interpolation filter corresponding to the region to be decoded based on the decoding information includes: determining a plurality of candidate interpolation filters based on the decoding information; decoding the bitstream to obtain first indication information of the target interpolation filter, where the first indication information is used to indicate index information of the target interpolation filter among the plurality of candidate interpolation filters; selecting the target interpolation filter from the plurality of candidate interpolation filters based on the first indication information.
3. The method according to claim 2, characterized in that, the determining a plurality of candidate interpolation filters based on the decoding information includes: obtaining preset candidate interpolation filters respectively corresponding to at least one of the prediction mode, the interpolation direction, the decoded component, and the interpolation target; determining the plurality of candidate interpolation filters based on the preset candidate interpolation filters.
4. The method according to claim 3, characterized in that, the determining the plurality of candidate interpolation filters based on the preset candidate interpolation filters includes: removing duplicate interpolation filters from the preset candidate interpolation filters to obtain the plurality of candidate interpolation filters.
5. The method according to claim 2, characterized in that, the determining a plurality of candidate interpolation filters based on the decoding information includes: obtaining a selection strategy respectively corresponding to at least one of the prediction mode, the interpolation direction, the decoded component, and the interpolation target, and selecting the plurality of candidate interpolation filters from M preset interpolation filters based on the selection strategy, where M is a positive integer greater than 1; or, obtaining N candidate combinations formed by the M interpolation filters, decoding the bitstream to obtain second indication information, and selecting a target candidate combination from the N candidate combinations based on the second indication information, and then determining the interpolation filters included in the target candidate combination as the plurality of candidate interpolation filters, where each candidate combination includes two or more of the M interpolation filters, the second indication information is used to indicate the target candidate combination, and N is a positive integer greater than 1.
6. The method according to claim 5, characterized in that, the interpolation filters included in different candidate combinations among the N candidate combinations are randomly distributed, or the number of interpolation filters included in each candidate combination among the N candidate combinations gradually increases, and the number of taps of the interpolation filters included in each candidate combination increases or decreases in sequence.
7. The method according to any one of claims 2-6, It is characterized in that when the first indication information includes a flag bit array or multiple flag bits, and the flag bit array or multiple flag bits are used to indicate the decoded information and the target interpolation filter, decoding the coded stream to obtain the decoded information of the area to be decoded includes: decoding the coded stream to obtain the first indication information; obtaining the decoded information of the area to be decoded based on the first indication information.
8. The method according to claim 1, it is characterized in that determining the target interpolation filter corresponding to the area to be decoded based on the decoded information includes: determining the target interpolation filter based on the prediction mode, the decoded component, and the interpolation target; or, determining the target interpolation filter based on the prediction mode and the decoded component; or, determining the target interpolation filter based on the interpolation direction and the decoded component; or, determining the target interpolation filter according to the decoded component and the size of the area to be decoded.
9. The method according to claim 8, it is characterized in that determining the target interpolation filter based on the prediction mode, the decoded component, and the interpolation target includes: when the prediction mode is a bidirectional optical flow prediction mode, the decoded component is a luminance component, and the interpolation target is to determine the gradients of the forward prediction block and the backward prediction block of the area to be decoded in the x direction and the y direction, determining the interpolation filter with n1 taps as the target interpolation filter; when the prediction mode is the bidirectional optical flow prediction mode, the decoded component is a luminance component, and the interpolation target is to determine the predicted values of the area to be decoded in the x direction and the y direction, determining the interpolation filter with n2 taps as the target interpolation filter; when the prediction mode is the bidirectional optical flow prediction mode, the decoded component is a chrominance component, and the interpolation target is to determine the predicted values of the area to be decoded in the x direction and the y direction, determining the interpolation filter with n3 taps as the target interpolation filter; when the prediction mode is a non-bidirectional optical flow prediction mode, the decoded component is a luminance component, and the interpolation target is to determine the predicted values of the area to be decoded in the x direction and the y direction, determining the interpolation filter with n4 taps as the target interpolation filter; when the prediction mode is a non-bidirectional optical flow prediction mode, the decoded component is a chrominance component, and the interpolation target is to determine the predicted values of the area to be decoded in the x direction and the y direction, determining the interpolation filter with n5 taps as the target interpolation filter; wherein, the n1, n2, n3, n4, and n5 are all positive integers.
10. The method according to claim 9, it is characterized in that the n1 is not equal to n2.
11. The method according to claim 10, it is characterized in that the n1 = 8, the n2 = 12, the n3 = 6, the n4 = 12, and the n5 = 6.
12. A video coding method, it is characterized in that including: Obtain the coding information of the area to be coded, and the reference area corresponding to the area to be coded, where the coding information includes at least one of a prediction mode, an interpolation direction, a coding component, and an interpolation target; Based on the coding information, determine the target interpolation filter corresponding to the area to be coded; Based on the target interpolation filter, perform interpolation filtering on the reference area to determine the predicted value of the area to be coded.
13. The method according to claim 12, wherein, the determining the target interpolation filter corresponding to the area to be coded based on the coding information includes: Based on the coding information, determine a plurality of candidate interpolation filters; Select the target interpolation filter from the plurality of candidate interpolation filters.
14. The method according to claim 13, wherein, the determining a plurality of candidate interpolation filters based on the coding information includes: Obtain the preset candidate interpolation filters respectively corresponding to at least one of the prediction mode, the interpolation direction, the coding component, and the interpolation target; Based on the preset candidate interpolation filters, determine the plurality of candidate interpolation filters.
15. The method according to claim 13, wherein, the determining a plurality of candidate interpolation filters based on the coding information includes: Obtain the selection strategy respectively corresponding to at least one of the prediction mode, the interpolation direction, the coding component, and the interpolation target; Based on the selection strategy, select the plurality of candidate interpolation filters from M preset interpolation filters, where M is a positive integer greater than 1.
16. The method according to claim 12, wherein, the determining the target interpolation filter corresponding to the area to be coded based on the coding information includes: Based on the prediction mode, the coding component, and the interpolation target, determine the target interpolation filter; or, Based on the prediction mode and the coding component, determine the target interpolation filter; or, Based on the interpolation direction and the coding component, determine the target interpolation filter; or, Determine the target interpolation filter according to the coding component and the size of the area to be coded.
17. A video decoding device, wherein, it includes: A decoding unit, configured to decode a bitstream to obtain the decoding information of the area to be decoded, where the decoding information includes at least one of a prediction mode, an interpolation direction, a decoding component, and an interpolation target; A determining unit, configured to determine the target interpolation filter corresponding to the area to be decoded based on the decoding information; An interpolation unit, configured to determine the reference area of the area to be decoded, and based on the target interpolation filter, perform interpolation on the reference area to determine the predicted value of the area to be decoded.
18. A video encoding device, wherein, it includes: An obtaining unit, configured to obtain the coding information of the area to be coded, and the reference area corresponding to the area to be coded, where the coding information includes at least one of a prediction mode, an interpolation direction, a coding component, and an interpolation target; A determination unit, configured to determine a target interpolation filter corresponding to the region to be encoded based on the encoded information; An interpolation unit, configured to perform interpolation filtering on the reference region based on the target interpolation filter to determine a predicted value of the region to be encoded.
19. An electronic device, comprising a processor and a memory; The memory is configured to store a computer program; The processor is configured to execute the computer program to implement the method according to any one of claims 1 to 11 or 12 to 16 above.
20. A computer-readable storage medium, characterized in that it is configured to store a computer program; The computer program causes a computer to execute the method according to any one of claims 1 to 11 or 12 to 16 above.