Encoding method, decoding method, encoder, decoder and storage medium
By analyzing the pixel values of the reference block, it determines its corresponding traditional intra prediction mode and adding it to the candidate set, the problem of difficulty in accurately determining the prediction mode in the prior art is solved, and the prediction performance of the current block is improved.
Patent Information
- Application Number
- PCT/CN2023/129334
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-08
AI Technical Summary
The prior art is difficult to accurately determine the traditional intra prediction mode corresponding to the reference block around the current block, affecting the prediction performance of the current block.
By analyzing the pixel values of the reference block around the current block, its corresponding traditional intra prediction mode is determined and added to the candidate set of intra prediction modes to improve prediction performance.
By accurately determining the traditional intra prediction mode, the prediction performance of the current block is improved and the efficiency of video encoding and decoding is improved.
Smart Images

Figure CN2023129334_08052025_PF_FP_ABST
Abstract
Description
Coding and decoding method, codec and storage medium Technical Field
[0001] The present application relates to the field of video coding and decoding technology, and in particular to a coding and decoding method, a codec, and a storage medium. Background Art
[0002] The current block can be predicted by constructing a candidate set of intra-frame prediction modes based on one or more surrounding reference positions. If a reference block at a reference position around the current block does not correspond to a traditional intra-frame prediction mode (such as an angular prediction mode), how to accurately determine the traditional intra-frame prediction mode corresponding to the reference block to improve the prediction performance of the current block is a problem that needs to be solved.
[0003] Summary of the Invention
[0004] The embodiments of the present application provide a coding method, a codec, and a storage medium to improve prediction performance. The following describes various aspects of the present application.
[0005] In a first aspect, a decoding method is provided, which is applied to a decoder, and the method includes: determining at least one reference position around a current block, the at least one reference position including a first reference position, the first reference position corresponding to a first reference block, the first reference block being predicted based on a first prediction mode, and the first prediction mode does not belong to a target intra-frame prediction mode, and the target intra-frame prediction mode includes at least an angle prediction mode; determining a second prediction mode based on a pixel value of the first reference block, the second prediction mode belonging to the target intra-frame prediction mode; adding the second prediction mode to a candidate set of intra-frame prediction modes; and predicting the current block based on the candidate set.
[0006] In a second aspect, a coding method is provided, which is applied to an encoder, and the method includes: determining at least one reference position around a current block, the at least one reference position including a first reference position, the first reference position corresponding to a first reference block, the first reference block being predicted based on a first prediction mode, and the first prediction mode does not belong to a target intra-frame prediction mode, and the target intra-frame prediction mode includes at least an angle prediction mode; determining a second prediction mode based on a pixel value of the first reference block, the second prediction mode belonging to the target intra-frame prediction mode; adding the second prediction mode to a candidate set of intra-frame prediction modes; and predicting the current block based on the candidate set.
[0007] According to a third aspect, a decoder is provided, comprising: a first determination unit, configured to determine at least one reference position around a current block, the at least one reference position including a first reference position, the first reference position corresponding to a first reference block, the first reference block being predicted based on a first prediction mode, and the first prediction mode does not belong to a target intra-frame prediction mode, the target intra-frame prediction mode at least including an angle prediction mode; a second determination unit, configured to determine a second prediction mode based on a pixel value of the first reference block, the second prediction mode belonging to the target intra-frame prediction mode; an adding unit, configured to add the second prediction mode to a candidate set of intra-frame prediction modes; and a prediction unit, configured to predict the current block based on the candidate set.
[0008] In a fourth aspect, a decoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method described in the first aspect when running the computer program.
[0009] In a fifth aspect, an encoder is provided, comprising: a first determination unit, configured to determine at least one reference position around a current block, the at least one reference position including a first reference position, the first reference position corresponding to a first reference block, the first reference block being predicted based on a first prediction mode, and the first prediction mode does not belong to a target intra-frame prediction mode, the target intra-frame prediction mode at least including an angle prediction mode; a second determination unit, configured to determine a second prediction mode based on a pixel value of the first reference block, the second prediction mode belonging to the target intra-frame prediction mode; an adding unit, configured to add the second prediction mode to a candidate set of intra-frame prediction modes; a prediction unit, configured to predict the current block based on the candidate set.
[0010] In a sixth aspect, an encoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method described in the second aspect when running the computer program.
[0011] In a seventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the method as described in the first aspect or the second aspect is implemented.
[0012] In an eighth aspect, a computer program product is provided, comprising a computer program, which implements the method described in the first aspect or the second aspect when executed.
[0013] In a ninth aspect, a non-volatile computer-readable storage medium for storing a bit stream is provided, wherein the bit stream is generated by an encoding method of an encoder, or the bit stream is decoded by a decoding method of a decoder, wherein the decoding method is the method described in the first aspect, and the encoding method is the method described in the second aspect.
[0014] The target intra-frame prediction mode mentioned above can be understood as a traditional intra-frame prediction mode (such as an angular prediction mode). In the embodiment of the present application, if the reference block at the reference position around the current block (i.e., the first reference block mentioned above) does not correspond to a traditional intra-frame prediction mode, then the corresponding traditional intra-frame prediction mode can be determined based on the pixel values of the reference block itself. The intra-frame prediction mode determined based on the pixel values of the reference block itself is more accurate and helps improve the prediction performance of the current block. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] FIG1 is a structural diagram illustrating an example of a video encoder to which an embodiment of the present application may be applied.
[0016] FIG2 is a diagram showing an example structure of a video decoder to which an embodiment of the present application can be applied.
[0017] FIG. 3A is a diagram illustrating an example of an intra prediction method based on a decoder-side intra mode derivation (DIMD) mode.
[0018] FIG3B is an example diagram of a gradient histogram of a current block.
[0019] FIG4 is an example diagram of a method for determining gradient information of a current block.
[0020] FIG5 is an example diagram of an angle partitioning method in a geometric partition mode (GPM).
[0021] FIG6 is a diagram showing an example of a reference template in a template-based intra mode derivation (TIMD) mode.
[0022] FIG7 is a diagram showing an example of reference positions used to construct a candidate set.
[0023] FIG8 is another example diagram of reference positions for constructing a candidate set.
[0024] FIG9 is a flowchart of a decoding method provided in an embodiment of the present application.
[0025] FIG10 is a flow chart of the encoding method provided in an embodiment of the present application.
[0026] FIG11 is a schematic diagram of the structure of a decoder provided in one embodiment of the present application.
[0027] FIG12 is a schematic diagram of the structure of a decoder provided in another embodiment of the present application.
[0028] FIG13 is a schematic diagram of the structure of an encoder provided in one embodiment of the present application.
[0029] FIG14 is a schematic diagram of the structure of an encoder provided in another embodiment of the present application. DETAILED DESCRIPTION
[0030] The technical solution in this application will be described below with reference to the accompanying drawings.
[0031] FIG1 is a schematic block diagram of a video encoder according to an embodiment of the present application.
[0032] It should be understood that the video encoder 100 can be used to perform lossy compression or lossless compression on an image. The lossless compression can be visually lossless compression or mathematically lossless compression.
[0033] The video encoder 100 can be applied to image data in a luminance and chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents brightness (Luma), Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance (Chroma) for describing color and saturation. For example, in terms of color format, 4:2:0 means that every 4 pixels have 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that every 4 pixels have 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 represents full pixel display (YYYYCbCrCbCrCbCrCbCr).
[0034] For example, the video encoder 100 reads video data, and for each image in the video data, divides the image into a number of coding tree units (CTUs). In some examples, CTUs may be referred to as "tree blocks", "largest coding units" (LCUs) or "coding tree blocks" (CTBs). Each CTU may be associated with a pixel block of equal size within the image. Each pixel may correspond to a luminance (luminance or luma) sample and two chrominance (chroma) samples. Therefore, each CTU may be associated with a luminance sample block and two chroma sample blocks. The size of a CTU is, for example, 128×128, 64×64, 32×32, etc. A CTU may be further divided into a number of coding units (CUs) for encoding. A CU may be a rectangular block or a square block. A CU can be further divided into prediction units (PUs) and transform units (TUs), allowing for separation of coding, prediction, and transform, and greater flexibility in processing. In one example, a CTU is divided into CUs using a quadtree, and a CU is divided into TUs and PUs using a quadtree.
[0035] The video encoder and video decoder can support various PU sizes. Assuming that the size of a particular CU is 2N×2N, the video encoder and video decoder can support PU sizes of 2N×2N or N×N for intra-frame prediction, and support symmetric PUs of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-frame prediction. The video encoder and video decoder can also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-frame prediction.
[0036] In some embodiments, as shown in FIG1 , the video encoder 100 may include a prediction unit 110, a residual unit 120, a transform / quantization unit 130, an inverse transform / quantization unit 140, a reconstruction unit 150, a loop filter unit 160, a decoded image buffer 170, and an entropy coding unit 180. It should be noted that the video encoder 100 may include more, fewer, or different functional components.
[0037] Optionally, in this application, the current block may be referred to as the current coding unit (CU) or the current prediction unit (PU), etc. The prediction block may also be referred to as a predicted image block or an image prediction block, and the reconstructed image block may also be referred to as a reconstructed block or an image reconstruction block.
[0038] In some embodiments, the prediction unit 110 includes an inter-frame prediction unit 111 and an intra-frame prediction unit 112. Because there is a strong correlation between adjacent pixels in a video image, intra-frame prediction is used in video coding and decoding technologies to eliminate spatial redundancy between adjacent pixels. Because there is a strong similarity between adjacent images in a video, inter-frame prediction is used in video coding and decoding technologies to eliminate temporal redundancy between adjacent images, thereby improving coding efficiency.
[0039] The inter-frame prediction unit 111 can be used for inter-frame prediction. Inter-frame prediction can include motion estimation and motion compensation. It can refer to image information from different images. Inter-frame prediction uses motion information to find a reference block from the reference image and generate a prediction block based on the reference block to eliminate temporal redundancy. Inter-frame prediction uses motion information to find a reference block from the reference image and generate a prediction block based on the reference block. Motion information includes the reference image list in which the reference image is located, the reference image index, and the motion vector. The motion vector can be integer pixel or fractional pixel. If the motion vector is fractional pixel, interpolation filtering is required to generate the required fractional pixel block in the reference image. Here, the integer pixel or fractional pixel block in the reference image found based on the motion vector is called a reference block. Some technologies directly use the reference block as the prediction block, while others further process the reference block to generate a prediction block. Reprocessing the reference block to generate a prediction block can also be understood as using the reference block as the prediction block and then processing the prediction block to generate a new prediction block.
[0040] The intra-frame prediction unit 112 only refers to information of the same image to predict pixel information within the current code image block to eliminate spatial redundancy.
[0041] Intra-frame prediction has multiple prediction modes. For example, the H-series international digital video coding standard H.264 / AVC has eight angular prediction modes and one non-angular prediction mode. H.265 / HEVC expands this to 33 angular prediction modes and two non-angular prediction modes. High-efficiency video coding (HEVC) uses planar, direct current (DC), and 33 angular modes for a total of 35 intra-frame prediction modes. Versatile video coding (VVC) uses planar, DC, and 65 angular modes for a total of 67 intra-frame prediction modes.
[0042] It should be noted that with the increase of angle modes, intra-frame prediction will be more accurate and more in line with the needs of high-definition and ultra-high-definition digital video development.
[0043] Residual unit 120 may generate a residual block for a CU based on the pixel blocks of the CU and the prediction blocks of the PUs of the CU. For example, residual unit 120 may generate a residual block for the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel blocks of the CU and the corresponding sample in the prediction blocks of the PUs of the CU.
[0044] The transform / quantization unit 130 may quantize the transform coefficients. The transform / quantization unit 130 may quantize the transform coefficients associated with the TUs of the CU based on a quantization parameter (QP) value associated with the CU. The video encoder 100 may adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.
[0045] The inverse transform / quantization unit 140 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficients to reconstruct a residual block from the quantized transform coefficients.
[0046] Reconstruction unit 150 may add samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by prediction unit 110 to generate a reconstructed image block associated with the TU. By reconstructing the sample blocks of each TU of a CU in this manner, video encoder 100 can reconstruct the pixel blocks of the CU.
[0047] The loop filter unit 160 is used to process the inverse transformed and inverse quantized pixels to compensate for distortion information and provide a better reference for subsequent coded pixels. For example, it can perform a deblocking filtering operation to reduce the blocking effect of pixel blocks associated with the CU.
[0048] In some embodiments, the loop filtering unit 160 includes a deblocking filtering unit and a sample adaptive offset / adaptive loop filtering (SAO / ALF) unit, wherein the deblocking filtering unit is used to remove blocking effects, and the SAO / ALF unit is used to remove ringing effects.
[0049] The decoded image buffer 170 may store the reconstructed pixel blocks. The inter-prediction unit 111 may use a reference image containing the reconstructed pixel blocks to perform inter-prediction on PUs of other images. In addition, the intra-prediction unit 112 may use the reconstructed pixel blocks in the decoded image buffer 170 to perform intra-prediction on other PUs in the same image as the CU.
[0050] The entropy encoding unit 180 may receive the quantized transform coefficients from the transform / quantization unit 130. The entropy encoding unit 180 may perform one or more entropy encoding operations on the quantized transform coefficients to generate entropy-encoded data.
[0051] FIG2 is a schematic block diagram of a video decoder according to an embodiment of the present application.
[0052] 2 , video decoder 200 includes an entropy decoding unit 210, a prediction unit 220, an inverse quantization / transformation unit 230, a reconstruction unit 240, a loop filter unit 250, and a decoded picture buffer 260. It should be noted that video decoder 200 may include more, fewer, or different functional components.
[0053] Video decoder 200 may receive a bitstream. Entropy decoding unit 210 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, entropy decoding unit 210 may parse the entropy-encoded syntax elements in the bitstream. Prediction unit 220, inverse quantization / transform unit 230, reconstruction unit 240, and loop filter unit 250 may decode video data based on the syntax elements extracted from the bitstream, thereby generating decoded video data.
[0054] In some embodiments, the prediction unit 220 includes an intra-frame prediction unit 222 and an inter-frame prediction unit 221 .
[0055] The intra-frame prediction unit 222 may perform intra-frame prediction to generate a prediction block for the PU. The intra-frame prediction unit 222 may use an intra-frame prediction mode to generate a prediction block for the PU based on the pixel blocks of spatially neighboring PUs. The intra-frame prediction unit 222 may also determine the intra-frame prediction mode of the PU based on one or more syntax elements parsed from the codestream.
[0056] The inter-frame prediction unit 221 may construct a first reference picture list (List 0) and a second reference picture list (List 1) based on syntax elements parsed from the codestream. In addition, if a PU is encoded using inter-frame prediction, the entropy decoding unit 210 may parse the motion information of the PU. The inter-frame prediction unit 221 may determine one or more reference blocks of the PU based on the motion information of the PU. The inter-frame prediction unit 221 may generate a prediction block for the PU based on the one or more reference blocks of the PU.
[0057] The inverse quantization / transform unit 230 may inversely quantize (ie, dequantize) the transform coefficients associated with the TU. The inverse quantization / transform unit 230 may use the QP value associated with the CU of the TU to determine the degree of quantization.
[0058] After inverse quantizing the transform coefficients, the inverse quantization / transform unit 230 may apply one or more inverse transforms to the inverse quantized transform coefficients in order to generate a residual block associated with the TU.
[0059] Reconstruction unit 240 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, reconstruction unit 240 can add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU to obtain a reconstructed image block.
[0060] The loop filtering unit 250 may perform a deblocking filtering operation to reduce blocking artifacts of pixel blocks associated with a CU.
[0061] The video decoder 200 may store the reconstructed image of the CU in the decoded image buffer 260. The video decoder 200 may use the reconstructed image in the decoded image buffer 260 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.
[0062] The basic process of video encoding and decoding is as follows: At the encoder end, an image is divided into blocks. For the current block, the prediction unit 110 uses intra-frame prediction or inter-frame prediction to generate a prediction block for the current block. The residual unit 120 calculates a residual block based on the predicted block and the original block of the current block. This residual block is the difference between the predicted block and the original block of the current block. This residual block can also be referred to as residual information. This residual block undergoes transformation and quantization by the transform / quantization unit 130, removing information that is insensitive to the human eye and eliminating visual redundancy. Optionally, the residual block before transformation and quantization by the transform / quantization unit 130 can be referred to as a time-domain residual block, and the time-domain residual block after transformation and quantization by the transform / quantization unit 130 can be referred to as a frequency residual block or a frequency-domain residual block. The entropy coding unit 180 receives the quantized change coefficients output by the transform and quantization unit 130, performs entropy coding on these quantized change coefficients, and outputs a bitstream. For example, the entropy coding unit 180 can eliminate character redundancy based on the target context model and probability information of the binary bitstream.
[0063] At the decoding end, the entropy decoding unit 210 can parse the code stream to obtain the prediction information, quantization coefficient matrix, etc. of the current block. The prediction unit 220 uses intra-frame prediction or inter-frame prediction on the current block based on the prediction information to generate a prediction block for the current block. The inverse quantization / transformation unit 230 uses the quantization coefficient matrix obtained from the code stream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block. The reconstruction unit 240 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed blocks constitute a reconstructed image, and the loop filtering unit 250 performs loop filtering on the reconstructed image based on the image or block to obtain a decoded image. The encoding end also requires similar operations as the decoding end to obtain a decoded image. The decoded image can also be called a reconstructed image, and the reconstructed image can be used as a reference image for inter-frame prediction of subsequent images.
[0064] It should be noted that the block division information determined by the encoder, as well as mode information or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering, etc., are carried in the bitstream when necessary. The decoder parses the bitstream and analyzes the existing information to determine the same block division information, prediction, transform, quantization, entropy coding, loop filtering, etc. mode information or parameter information as the encoder, thereby ensuring that the decoded image obtained by the encoder and the decoder are identical.
[0065] The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. This application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.
[0066] The preceding text describes in detail the codec framework provided by the embodiments of the present application. The embodiments of the present application mainly relate to the intra-frame prediction process, which can be implemented in the intra-frame prediction unit in the codec framework mentioned above. The following text describes in detail the relevant concepts involved in the embodiments of the present application.
[0067] DIMD mode
[0068] DIMD is an intra-frame prediction tool. DIMD technology can derive the intra-frame prediction mode (or prediction direction) of the current block based on the gradient information of the reconstructed area (the area where the reconstructed pixel values are located) surrounding the current block. DIMD technology can then obtain a prediction value based on the derived intra-frame prediction mode. The gradient information can include horizontal gradient information and vertical gradient information, and each set of horizontal gradient information and vertical gradient information can correspond to a traditional prediction angle. The following, in conjunction with Figures 3A and 3B, provides an example of how gradient information is matched with traditional prediction angles.
[0069] As shown in FIG3A , a 3×3 sliding window can be used to slide on the reconstruction area of 3 rows and 3 columns around the current block with a step size of 1 pixel to calculate the horizontal and vertical gradient values of each 3×3 window on the reconstruction area: G x and G y .G x and G y It can be obtained by multiplying the predicted value in the window position with the Sobel operator. The Sobel operator is shown in formula (1), where M x Used to calculate the horizontal gradient, M y Used to calculate vertical gradient. M x and M y The values of are as follows:
[0070] Then, according to G at each position x and G y , calculate the traditional angle direction O corresponding to each position according to the following formulas (2) and (3), and calculate the amplitude value G of the gradient of the angle corresponding to each position: G=|G x |+|G y | (2)
[0071] In some implementations, the calculation process of atan can be simplified. For example, the calculation process of atan can be simplified by looking up a table or using a modified formula.
[0072] Next, the magnitude value G of the gradient at each position may be accumulated in the derived traditional angle categories to obtain a gradient magnitude value histogram (see FIG3B ).
[0073] Finally, the traditional angle with the largest accumulated gradient magnitude value can be selected as the angle corresponding to the prediction block based on the DIMD mode. When the magnitude values derived from all traditional angles are zero, the prediction block can be matched to the traditional planar mode.
[0074] In addition to determining the prediction mode, DIMD can also be used to select the transform kernel group for the base transform and the non-separable secondary transform. Specifically, different traditional angular prediction modes can correspond to different transform kernels. The transform kernels mentioned here can include multiple transform selection (MTS), non-separable primary transform (NSPT), and low-frequency non-separable secondary transform (LFNST). In related art, transform kernels can be divided into multiple groups, each of which can contain multiple transform kernels. When the prediction mode is the traditional transform mode, the selected traditional transform mode determines the MTS, NSPT, or LFNST group, and the specific transform kernel selected is determined by the identifier parsed from the bitstream. If the prediction mode selected for the current block does not use the traditional prediction angle for prediction, then when predicting the current block, the transform kernel group corresponding to the current block cannot be determined, and thus the transform mode of the residual block of the current block cannot be determined. Therefore, in some implementations, the traditional prediction angle corresponding to the current block can be derived based on the DIMD mode. Then, the group of transform kernels corresponding to the current block can be determined based on the prediction angle. In conjunction with FIG4 , an example of determining a traditional prediction angle of the current block based on gradient information is given below.
[0075] As shown in FIG4 , a 3×3 sliding window can be used to slide in the current block with a step size of 1 pixel to calculate the horizontal and vertical gradient values of each 3×3 window in the current block: G x and G y .G x and G y It can be represented by the 3×3 horizontal gradient operator M x and the vertical gradient operator M y It is obtained by multiplying the predicted value within the window position.x and M y The value of is shown in formula (1) above.
[0076] Assuming that the current block is a block with a width and height of (w, h), the sliding 3×3 window can calculate the G of the (w-2)×(h-2) positions of the current block center. x and G y .
[0077] Then, according to G at each position x and G y , calculate the traditional angle direction O corresponding to each position according to the above formulas (2) and (3), and calculate the amplitude value G of the gradient of the angle corresponding to each position.
[0078] Next, the magnitude value G of the gradient at each position may be accumulated in the derived traditional angle categories to obtain a gradient magnitude value histogram (see FIG3B ).
[0079] Finally, the traditional angle with the largest cumulative gradient magnitude value can be selected as the traditional prediction angle corresponding to the current block. When the magnitude values derived from all traditional angles are zero, the current block can be matched to the traditional planar mode.
[0080] For the above-mentioned DIMD method, some variations may occur in actual implementation, and the present embodiment does not specifically limit this. For example, the operator used to solve the horizontal gradient and vertical gradient of each position may be different from the M in the present embodiment. x and M y For example, when calculating atan, integers and shifts or table lookups can be used to simplify floating-point number and division operations. For another example, when mapping angles to existing angle patterns, table lookups can also be used.
[0081] Most probable mode list (MPM)
[0082] Adjacent blocks within the same frame are typically highly correlated, so the intra-frame prediction modes of adjacent blocks are likely to be the same or similar. The MPM mode constructs an MPM list for the current block using the intra-frame prediction modes of adjacent blocks to predict the current block. For example, in HEVC, the length of the MPM list is 3, and in VVC, the length of the MPM list is 6.
[0083] Taking the VVC test model (VTM), a reference software in VVC, as an example, six MPMs can be generated by the prediction modes of two adjacent blocks. The modes in the MPM list can include the following three categories:
[0084] (1) Default mode;
[0085] (2) adjacent block mode;
[0086] (3) Computational derivation patterns.
[0087] Spatial geometric partition mode (SGPM) and GPM
[0088] Both SGPM and GPM are intra-frame partition prediction modes. SGPM is derived from GPM. Based on the partition prediction mode, the current block is divided into two parts, each of which is predicted using a different intra-frame prediction mode.
[0089] As shown in Figure 5, in the reference software ECM, GPM supports 64 partition modes. Compared to GPM, SGPM supports 26 different directional or positional partition modes out of a total of 64. The 64 partition modes supported by GPM include 32 partition angles. The correspondence between these 32 partition angles and traditional prediction angles is shown in Table 1. In Table 1, angleIdx represents the 32 partition angles, and intraMode represents the index of the traditional 67 intra-frame prediction modes (including planar mode, DC mode, and 65 angular prediction modes).
[0090] Table 1
[0091] TIMD mode
[0092] The intra-frame prediction technology based on the TIMD mode can derive information about the intra-frame prediction mode according to several rows of pixel values reconstructed around the current block, and further derive one or more traditional intra-frame prediction modes.
[0093] In the reference software ECM-7.0, the TIMD mode can derive four intra-frame prediction modes, namely: TIMD mode (timd mode), second TIMD mode (timd Secondary mode), TIMD horizontal mode (timd ver) and TIMD vertical mode (timd hor).
[0094] The TIMD mode can further divide the 65 traditional prediction angles into 129 prediction angles. In other words, a finer angle can be added between every two adjacent traditional prediction angles.
[0095] The TIMD mode derives the intra-frame prediction mode based on the cost value at the template position. For example, the sum of absolute transformed difference (SATD) can be used as the cost value of the template position in the reference software ECM. Referring to Figure 6, the L2 row above the current block can be used as the upper template, the L1 row on the left can be used as the left template, and a row of reconstructed pixel values (reference of the template) outside the template area is used as a reference pixel. Make a prediction on the template area under a given intra-frame prediction mode. The prediction process can be: based on the reference pixel value, use each traditional prediction mode in the MPM list to generate a prediction value for the template area. Then, calculate the SATD of the predicted value and the reconstructed value of the template area. Select the traditional prediction mode with the smallest SATD as the TIMD mode and use it for the prediction of the current block.
[0096] In the above template area, the SATD value between the predicted value and the reconstructed value obtained based on the prediction mode is the cost value. The TIMD mode can distinguish the above four TIMD prediction modes based on the SATD value, as follows:
[0097] (1) TIMD mode (timd mode) is the mode with the smallest total SATD value on the upper template and the left template;
[0098] (2) The second TIMD mode (timd secondary mode) is the mode with the second smallest total SATD value on the upper template and the left template;
[0099] (3) TIMD vertical mode (timd ver) is the mode with the smallest SATD value on the upper template;
[0100] (4) The TIMD horizontal mode (timd hor) is the mode with the smallest SATD value on the left template.
[0101] In the reference software ECM, the TIMD mode and the second TIMD mode can perform adaptive weighted prediction. Specifically, the predicted value of the current block in the TIMD mode is weighted with the predicted value of the current block in the second TIMD mode. Whether weighting is performed and the weight of the weighting are related to the SATD values of the two modes.
[0102] In some implementations, whether the TIMD mode and the second TIMD mode perform adaptive weighted prediction may be determined as follows:
[0103] (1) The TIMD mode is the mode with the smallest SATD value, and its SATD value can be cost0. The second TIMD mode is the mode with the second smallest SATD value, and its SATD value can be cost1.
[0104] (2) When cost0*2>cost1, the prediction values of the above two prediction modes on the current block will be weighted, otherwise the prediction value of the TIMD mode is directly used;
[0105] (3) When weighting, the weight ratio is shown in formulas (4) and (5), where Pred timdMode is the prediction result of the TIMD mode in the current block, and PredtimdSecondaryMode is the prediction value of the second TIMD mode in the current block. Pred=Pred timdMode ×w0+PredtimdSecondaryMode×w1 (5)
[0106] In some implementations, in order to avoid floating point calculations and trigger calculations, the value of w0+w1 is enlarged to 64, and the division calculation of w0 and w1 is also implemented using a table lookup method. The final weighting process is shown in formula (6). Pred = (Pred timdMode ×w0+PredtimdSecondaryMode×w1)>>6 (6)
[0107] The prediction mode derived by TIMD mode comes from the candidate list of intra prediction modes.
[0108] Template-based multiple reference line intra prediction (TMRL) mode
[0109] Reference software ECM-7.0 includes the TMRL mode. TMRL is an intra-frame prediction mode. TMRL technology is based on template matching and MRL technology. The process of predicting the current block based on the TMRL mode is as follows:
[0110] (1) In the bitstream parsing stage: parse the bitstream, confirm that the current block uses the TMRL mode, and parse the TMRL list index.
[0111] (2) During the prediction and reconstruction phase of the current block: First, a candidate list of TMRL modes is constructed. The syntax elements in the selected TMRL list are determined based on the candidate list of TMRL modes and the list index obtained by decoding. Each syntax element consists of an intra-frame prediction mode and a reference row index. The current block is predicted and reconstructed using the confirmed intra-frame prediction mode and the corresponding reference row.
[0112] TMRL list construction is an operation required by both the encoder and decoder. The TMRL mode selects a reference row from a candidate list by its index in the candidate list through codecs. This index is used to determine the selected reference row within the sorted candidate list. This reference row is then used for prediction with the selected intra prediction mode (traditional prediction mode).
[0113] The TMRL mode uses the sum of absolute differences (SAD) between the predicted value and the reconstructed value in the template area for up to 5 predefined extended reference lines (for example, the extended reference lines can be 1, 3, 5, 7, and 12. The specific extended reference lines used depend on the position of the current block in the current CTU. Of course, fewer than 5 extended reference lines can also be used). The 5 extended reference lines and 10 predefined prediction modes can form up to 5×10 combinations for sorting.
[0114] (x, -1) and (-1, y) are the coordinates relative to the upper left corner (0, 0) of the current block, and the template for calculating the SAD is an area of 1 row and 1 column.
[0115] From the above introduction, it can be seen that if the above-mentioned mode is used to predict the current block, it is necessary to establish a candidate set (or candidate list) constructed by traditional intra-frame prediction modes. In the related art, a candidate set can be constructed as follows: obtain the traditional intra-frame prediction modes of the blocks at the surrounding adjacent and non-adjacent positions, and add the above-mentioned traditional intra-frame prediction modes to the candidate set without duplication. In some implementations, when the obtained traditional intra-frame prediction modes cannot fill the candidate set to be constructed, it is further expanded based on the added traditional intra-frame prediction modes. For example, if the candidate set includes an angle prediction mode, then other angle prediction modes similar to the angle prediction mode can be added to the candidate set. For example, if angle prediction mode 50 is added to the candidate set, angle prediction mode 49 and angle prediction mode 51 can be added to the candidate set.
[0116] As an example, in SGPM and GPM, to construct the above candidate set, as shown in FIG7 , it is necessary to refer to reference blocks in five adjacent positions, namely: left, upper left, lower left, and upper right areas.
[0117] As another example, in MPM, TIMD, and TMRL, as shown in FIG8 , in addition to referring to reference blocks at 5 adjacent positions, 18 reference blocks at non-adjacent positions also need to be considered.
[0118] Of course, in addition to the prediction modes introduced above, there are other modes that also require the construction of a candidate set for intra-frame prediction, which will not be described here.
[0119] Based on the above introduction, it can be known that the current block can construct a candidate set of intra-frame prediction modes based on one or more surrounding reference positions to predict the current block. When constructing the candidate set, it is necessary to obtain the traditional intra-frame prediction mode of the reference block at the reference position and add these traditional intra-frame prediction modes to the candidate set. The way to obtain the traditional intra-frame prediction mode is different for reference blocks predicted based on different prediction modes. For example, the reference block can be predicted based on the following modes: intra-frame prediction mode, intra-frame block coding mode and inter-frame prediction mode. The following will be divided into two cases to introduce the method of obtaining the traditional intra-frame prediction mode of the reference block in the related art.
[0120] The first case: the reference block at the reference position corresponds to a traditional intra-frame prediction mode. For example, if the current block uses MIP, the planar mode in the traditional intra-frame mode of the reference block can be obtained. For another example, if the current block uses SGPM, the traditional intra-frame prediction angle corresponding to the reference block division angle can be obtained. For another example, if the current block uses TIMD or DIMD, the traditional intra-frame prediction mode with the highest weight derived from the reference block can be obtained. For another example, if the current block uses GPM, the traditional intra-frame prediction angle corresponding to the reference block division angle can be obtained.
[0121] The second case: the reference block at the reference position does not correspond to the traditional intra prediction mode. For example, if the reference block is predicted based on the intra block coding mode or the inter prediction mode, then there may be no corresponding traditional intra prediction mode. In the related art, in order to construct a candidate set, if the reference block does not correspond to the traditional intra prediction mode, the reference block of the reference block can be found through the motion information of the reference block, and the traditional intra prediction mode of the reference block of the reference block is added to the candidate set. There are many ways to find the reference block of the reference block based on motion information. For example, if the reference block is predicted based on the inter prediction mode, then the reference block of the reference block can be found through the motion vector (motion vector) obtained by the inter prediction mode. For another example, if the reference block is predicted based on the intra block coding mode, then the reference block of the reference block can be found through the block vector (block vector) obtained by the intra block coding mode.
[0122] As mentioned above, if the reference block at a certain reference position around the current block does not correspond to the traditional intra-frame prediction mode (such as the angle prediction mode), how to accurately determine the traditional intra-frame prediction mode corresponding to the reference block to improve the prediction performance of the current block is a problem that needs to be solved.
[0123] To address the above issues, in an embodiment of the present application, if a reference block at a reference position surrounding the current block does not correspond to a traditional intra prediction mode or is not predicted using a traditional intra prediction mode, the corresponding traditional intra prediction mode can be determined based on the pixel values of the reference block itself. The intra prediction mode determined based on the pixel values of the reference block itself is more accurate and helps improve the prediction performance of the current block.
[0124] The decoding method of the embodiment of the present application is described in detail below with reference to FIG9 .
[0125] FIG9 is a flowchart of a decoding method provided by an embodiment of the present application. The method of FIG9 can also be referred to as an intra-frame prediction method. The method of FIG9 can be applied to a decoder.
[0126] 9 , in step S910 , at least one reference position around a current block is determined.
[0127] The current block may refer to a current block to be predicted or a current block to be decoded. In some implementations, the current block is a luminance block. In other implementations, the current block may also be a chrominance block.
[0128] The reference positions around the current block may include, for example, the reference positions shown in Figure 7. Alternatively, the reference positions around the current block may include the reference positions shown in Figure 8.
[0129] The at least one reference position mentioned above may include a first reference position. The first reference position corresponds to a first reference block (or, the first reference block includes the first reference position). The first reference block is predicted based on a first prediction mode. The first prediction mode mentioned here does not belong to the target prediction mode.
[0130] The target intra-frame prediction mode mentioned in the embodiments of the present application may refer to a traditional intra-frame prediction mode. For example, the target intra-frame prediction mode may include at least a traditional angle prediction mode. In some implementations, the target intra-frame prediction mode may also include one or more of a planar mode and a DC mode. Therefore, "the first reference block is predicted based on the first prediction mode, and the first prediction mode does not belong to the target prediction mode" can also be understood as the first prediction mode corresponding to the first reference block is not a traditional intra-frame prediction mode.
[0131] In step S920, a second prediction mode is determined according to the pixel values of the first reference block. The second prediction mode belongs to the target intra prediction mode, or in other words, the second prediction mode belongs to the traditional intra prediction mode.
[0132] The pixel values of the first reference block mentioned above may be pixel values of the prediction block, or the pixel values of the first reference block may also be pixel values of the reconstructed block.
[0133] In some implementations, gradient information corresponding to at least one angle can be determined based on pixel values of the first reference block, and the second prediction mode can be determined based on the gradient information corresponding to the at least one angle. The gradient information corresponding to the at least one angle can also be referred to as gradient information corresponding to at least one angular prediction mode. The gradient information mentioned here can be, for example, a gradient histogram. Determining the second prediction mode based on the gradient information corresponding to the at least one angle mentioned above can include: determining the second prediction mode based on the gradient amplitude value corresponding to the at least one angle. For example, the angular prediction mode corresponding to the gradient with the largest amplitude value can be used as the second prediction mode.
[0134] There are many ways to determine the gradient information.
[0135] For example, the gradient information can be determined based on the DIMD mode. For example, taking the first reference block as the block in FIG4 as an example, a 3×3 sliding window can be used to slide within the first reference block to calculate the gradient information of each 3×3 window on the first reference block.
[0136] For another example, gradient information can be determined based on some variant modes of the DIMD mode. Exemplarily, other operators besides the sobel operator can also be used when obtaining gradient information. Exemplarily, when calculating the angle, operations such as shifting, multiplication, or table lookup can be used instead of calculating atan or triggering operations. Exemplarily, when calculating the angle corresponding to the traditional intra-frame prediction mode angle, table lookup can also be used. The gradient information may include, for example, the amplitude value of the gradient. The above "determine the second prediction mode based on the gradient information corresponding to at least one angle." can also be understood as determining the second prediction mode based on the amplitude value corresponding to at least one angle. For example, the angle value corresponding to the maximum amplitude value is found in the gradient histogram of the first reference block, and the angle prediction mode corresponding to the angle value can be used as the second prediction mode.
[0137] 9 , in step S930 , the second prediction mode is added to the candidate set of intra prediction modes. This embodiment of the present application does not specifically limit the candidate set, and the candidate set can be constructed based on any prediction mode.
[0138] In step S940 , the current block is predicted based on the candidate set.
[0139] In related art, if the prediction mode corresponding to the current block (i.e., the first prediction mode mentioned above) is a prediction mode based on motion information, the reference block of the first reference block is determined based on the motion information, and then the intra-frame prediction mode corresponding to the reference block of the first reference block is added to the intra-frame prediction mode candidate set. However, since the reference block of the first reference block is located far away from the current block, constructing the intra-frame prediction mode candidate set based on its corresponding intra-frame prediction mode will reduce prediction performance.
[0140] Therefore, in some implementations, if the first reference block is predicted based on motion information, the second prediction mode can be determined based on the pixel values of the first reference block in the manner described in step S920. In other words, if the first reference block is predicted based on motion information, the corresponding intra-frame prediction mode can be determined directly based on the pixel values of the first reference block itself. The intra-frame prediction mode determined based on the pixel values of the reference block itself is more accurate and helps improve the prediction performance of the current block.
[0141] It should be understood that the aforementioned operation information may be, for example, a motion vector used in inter-frame prediction or a block vector used in intra-frame prediction. For example, the first prediction mode may include an intra-frame block coding mode. The intra-frame block coding mode may also be referred to as an intra-frame block copy (IBC) mode. For another example, the first prediction mode may include an inter-frame prediction mode.
[0142] In some implementations, the first prediction mode mentioned above may include GPM. That is, if the first reference block is predicted based on GPM, then according to step S920, a corresponding intra-frame prediction mode may be determined directly based on the pixel values of the first reference block itself. The intra-frame prediction mode determined based on the pixel values of the reference block itself is more accurate and helps improve the prediction performance of the current block.
[0143] In some implementations, the first prediction mode mentioned above may not include GPM. This is because, if the prediction mode corresponding to the current block is GPM, the traditional prediction mode corresponding to the first reference block can be determined based on the correspondence between GPM and traditional intra-frame prediction modes, and this traditional prediction mode can be added to the intra-frame prediction mode candidate list for the current block. This implementation eliminates the need to determine the traditional prediction mode corresponding to the current block based on the pixel values of the current block, thereby reducing the complexity of the encoding end.
[0144] If the amplitude values corresponding to each angle in the gradient histogram of the first reference block are relatively average, adding the second prediction mode determined according to the gradient information to the intra-frame prediction mode candidate set may reduce the effectiveness of the intra-frame prediction mode in the candidate set.
[0145] Therefore, in some implementations, if the amplitude value corresponding to at least one angle satisfies a first preset condition, the second prediction mode is added to the candidate set of intra prediction modes.
[0146] For example, the first preset condition can be associated with the maximum amplitude value among the amplitude values corresponding to at least one angle. For example, if the maximum amplitude value is greater than or equal to a target value, the second prediction mode is added to the candidate set of intra-frame prediction modes. The target value is determined by the sum of the amplitude values remaining in the amplitude values corresponding to at least one angle, excluding the maximum amplitude value. By setting the first preset condition for the amplitude value corresponding to at least one angle, some less accurate second prediction modes can be excluded from the candidate set, helping to further improve the accuracy of the candidate set predictions.
[0147] For example, let us continue to take the gradient histogram of the first reference block as an example. In this histogram, it is assumed that the maximum amplitude value is set to Amp0, and the sum of the amplitude values of all angles except the maximum amplitude value Amp0 is set to Amp sum-1 The derived traditional prediction mode is used as the second prediction mode of the first reference block only when the relationship of the following formula (8) is satisfied. sum-1 (8)
[0148] N is a positive integer greater than or equal to 1, for example, N is equal to 1, 2, 3, etc.
[0149] For another example, if the maximum amplitude value is greater than or equal to a target value, the second prediction mode is added to the candidate set of intra prediction modes, where the target value is determined according to the sum of all amplitude values corresponding to at least one angle.
[0150] For example, let's continue to take the gradient histogram of the first reference block as an example. In this histogram, it is assumed that the maximum amplitude value is Amp0, and the sum of the amplitude values of all angles is Amp sum The derived traditional prediction mode is used as the second prediction mode of the first reference block only when the relationship of the following formula (9) is satisfied. sum (9)
[0151] N is a positive integer greater than or equal to 2, for example, N is equal to 2, 3, 4, etc.
[0152] The embodiments of the present application do not specifically limit the size of the first reference block. In some implementations, the first reference block can be a reference block of any size. Alternatively, in other implementations, the first reference block can be a reference block whose size satisfies a second preset condition. That is, if the size of the first reference block satisfies the second preset condition, the second prediction mode is determined based on the pixel values of the first reference block. For example, the second preset condition can be the relationship between the size of the first reference block and the first size. For example, if the size of the first reference block is greater than or equal to the first size, the second prediction mode is allowed to be determined based on the pixel values of the first reference block. If the size of the first reference block is smaller than the first size, the second prediction mode is not allowed to be determined based on the pixel values of the first reference block. For reference blocks with larger sizes, the second prediction mode determined based on their pixel values may consume too much time. Therefore, by setting the above-mentioned second preset condition, such reference blocks can be excluded, thereby saving the prediction time of the current block.
[0153] As mentioned above, the current block can be predicted based on a candidate set. In some implementations, an intra-frame prediction mode for the current block can be determined based on the candidate set, and a prediction value for the current block can be determined based on the intra-frame prediction mode of the current block. For example, the bitstream can be parsed to determine first index information. The first index information is used to determine the intra-frame prediction mode for the current block from the candidate set. Then, the intra-frame prediction mode for the current block can be determined from the candidate set based on the first index information.
[0154] In some implementations, the method of FIG. 9 may further include parsing a bitstream and determining first identification information. The first identification information is used to indicate that the prediction mode of the current block is a target prediction mode. The target prediction mode is a mode for performing prediction based on a candidate set. For example, the target prediction mode may include one or more of MPM, TIMD, TMRL, SGPM, and GPM.
[0155] The first identification information may be represented by, for example, m_flag (m represents the type of prediction mode, for example, m may be MPM). Of course, the first identification information may also be represented by any other letters and / or numbers.
[0156] In some implementations, the method of FIG. 9 may further include: parsing the code stream to determine the quantized coefficients of the current block; performing inverse quantization on the quantized coefficients to determine the transform coefficients of the current block; and performing inverse transform on the transform coefficients to determine the residual information of the current block.
[0157] In some implementations, the method of FIG9 may further include determining reconstruction information for the current block based on the prediction value and residual information of the current block. For example, the prediction value and the residual value of the current block may be summed, and the summed result may be used as the reconstruction value of the current block.
[0158] The decoding method provided by the embodiment of the present application is described in detail above in conjunction with Figure 9. The encoding method provided by the embodiment of the present application is described in detail below in conjunction with Figure 10.
[0159] FIG10 is a flow chart of an encoding method provided by an embodiment of the present application. The method of FIG10 can also be referred to as an intra-frame prediction method. The method of FIG10 can be applied to an encoder.
[0160] 10 , in step S1010 , at least one reference position around a current block is determined.
[0161] The current block may refer to a current block to be predicted or a current block to be encoded. In some implementations, the current block is a luminance block. In other implementations, the current block may also be a chrominance block.
[0162] The reference positions around the current block may include, for example, the reference positions shown in Figure 7. Alternatively, the reference positions around the current block may include the reference positions shown in Figure 8.
[0163] The at least one reference position mentioned above may include a first reference position. The first reference position corresponds to a first reference block (or, the first reference block includes the first reference position). The first reference block is predicted based on a first prediction mode. The first prediction mode mentioned here does not belong to the target prediction mode.
[0164] The target intra-frame prediction mode mentioned in the embodiments of the present application may refer to a traditional intra-frame prediction mode. For example, the target intra-frame prediction mode may include at least a traditional angle prediction mode. In some implementations, the target intra-frame prediction mode may also include one or more of a planar mode and a DC mode. Therefore, "the first reference block is predicted based on the first prediction mode, and the first prediction mode does not belong to the target prediction mode" can also be understood as the first prediction mode corresponding to the first reference block is not a traditional intra-frame prediction mode.
[0165] In step S1020, a second prediction mode is determined according to the pixel values of the first reference block. The second prediction mode belongs to the target intra-frame prediction mode, or in other words, the second prediction mode belongs to the traditional intra-frame prediction mode.
[0166] The pixel values of the first reference block mentioned above may be pixel values of the prediction block, or the pixel values of the first reference block may also be pixel values of the reconstructed block.
[0167] In some implementations, gradient information corresponding to at least one angle can be determined based on pixel values of the first reference block, and the second prediction mode can be determined based on the gradient information corresponding to the at least one angle. The gradient information corresponding to the at least one angle can also be referred to as gradient information corresponding to at least one angular prediction mode. The gradient information mentioned here can be, for example, a gradient histogram. Determining the second prediction mode based on the gradient information corresponding to the at least one angle mentioned above can include: determining the second prediction mode based on the gradient amplitude value corresponding to the at least one angle. For example, the angular prediction mode corresponding to the gradient with the largest amplitude value can be used as the second prediction mode.
[0168] There are various ways to determine gradient information. For example, gradient information can be determined based on a DIMD mode. For example, taking the first reference block as the block in FIG. 4 as an example, a 3×3 sliding window can be used to slide within the first reference block to calculate gradient information for each 3×3 window on the first reference block.
[0169] For another example, gradient information can be determined based on some variant of the DIMD mode. For example, operators other than the Sobel operator can be used to obtain gradient information. For example, operations such as shifts, multiplications, or table lookups can be used to calculate angles instead of atan or trigger operations. For example, table lookups can also be used to calculate the angle corresponding to the traditional intra-frame prediction mode angle.
[0170] Gradient information may include, for example, gradient magnitude values. The phrase "determining the second prediction mode based on the gradient information corresponding to at least one angle" can also be understood as determining the second prediction mode based on the magnitude value corresponding to at least one angle. For example, the angle value corresponding to the maximum magnitude value in the gradient histogram of the first reference block is found, and the angle prediction mode corresponding to that angle value can be used as the second prediction mode.
[0171] 10 , in step S1030 , the second prediction mode is added to the candidate set of intra prediction modes. This embodiment of the present application does not specifically limit the candidate set, and the candidate set can be constructed based on any prediction mode.
[0172] In step S1040 , the current block is predicted based on the candidate set.
[0173] In related art, if the prediction mode corresponding to the current block (i.e., the first prediction mode mentioned above) is a prediction mode based on motion information, the reference block of the first reference block is determined based on the motion information, and then the intra-frame prediction mode corresponding to the reference block of the first reference block is added to the intra-frame prediction mode candidate set. However, since the reference block of the first reference block is located far away from the current block, constructing the intra-frame prediction mode candidate set based on its corresponding intra-frame prediction mode will reduce prediction performance.
[0174] Therefore, in some implementations, if the first reference block is predicted based on motion information, the second prediction mode can be determined based on the pixel values of the first reference block in the manner described in step S1020. In other words, if the first reference block is predicted based on motion information, the corresponding intra-frame prediction mode can be determined directly based on the pixel values of the first reference block itself. The intra-frame prediction mode determined based on the pixel values of the reference block itself is more accurate and helps improve the prediction performance of the current block.
[0175] It should be understood that the aforementioned operation information may be, for example, a motion vector used in inter-frame prediction or a block vector used in intra-frame prediction. For example, the first prediction mode may include an intra-frame block coding mode. The intra-frame block coding mode may also be referred to as an intra-frame block copy (IBC) mode. For another example, the first prediction mode may include an inter-frame prediction mode.
[0176] In some implementations, the first prediction mode mentioned above may include GPM. That is, if the first reference block is predicted based on GPM, then according to step S1020, a corresponding intra-frame prediction mode may be determined directly based on the pixel values of the first reference block itself. The intra-frame prediction mode determined based on the pixel values of the reference block itself is more accurate and helps improve the prediction performance of the current block.
[0177] In some implementations, the first prediction mode mentioned above may not include GPM. This is because, if the prediction mode corresponding to the current block is GPM, the traditional prediction mode corresponding to the first reference block can be determined based on the correspondence between GPM and traditional intra-frame prediction modes, and this traditional prediction mode can be added to the intra-frame prediction mode candidate list for the current block. This implementation eliminates the need to determine the traditional prediction mode corresponding to the current block based on the pixel values of the current block, thereby reducing the complexity of the encoding end.
[0178] If the amplitude values corresponding to each angle in the gradient histogram of the first reference block are relatively average, adding the second prediction mode determined according to the gradient information to the intra-frame prediction mode candidate set may reduce the effectiveness of the intra-frame prediction mode in the candidate set.
[0179] Therefore, in some implementations, if the amplitude value corresponding to at least one angle satisfies a first preset condition, the second prediction mode is added to the candidate set of intra prediction modes.
[0180] For example, the first preset condition can be associated with the maximum amplitude value among the amplitude values corresponding to at least one angle. For example, if the maximum amplitude value is greater than or equal to a target value, the second prediction mode is added to the candidate set of intra-frame prediction modes. The target value is determined by the sum of the amplitude values remaining in the amplitude values corresponding to at least one angle, excluding the maximum amplitude value. By setting the first preset condition for the amplitude value corresponding to at least one angle, some less accurate second prediction modes can be excluded from the candidate set, helping to further improve the accuracy of the candidate set predictions.
[0181] For example, let us continue to take the gradient histogram of the first reference block as an example. In this histogram, it is assumed that the maximum amplitude value is set to Amp0, and the sum of the amplitude values of all angles except the maximum amplitude value Amp0 is set to Amp sum-1 The derived traditional prediction mode is used as the second prediction mode of the first reference block only when the relationship of the following formula (8) is satisfied. sum-1 (8)
[0182] N is a positive integer greater than or equal to 1, for example, N is equal to 1, 2, 3, etc.
[0183] For another example, if the maximum amplitude value is greater than or equal to a target value, the second prediction mode is added to the candidate set of intra prediction modes, where the target value is determined according to the sum of all amplitude values corresponding to at least one angle.
[0184] For example, let's continue to take the gradient histogram of the first reference block as an example. In this histogram, it is assumed that the maximum amplitude value is Amp0, and the sum of the amplitude values of all angles is Amp sum The derived traditional prediction mode is used as the second prediction mode of the first reference block only when the relationship of the following formula (9) is satisfied. sum (9)
[0185] N is a positive integer greater than or equal to 2, for example, N is equal to 2, 3, 4, etc.
[0186] The embodiments of the present application do not specifically limit the size of the first reference block. In some implementations, the first reference block can be a reference block of any size. Alternatively, in other implementations, the first reference block can be a reference block whose size satisfies a second preset condition. That is, if the size of the first reference block satisfies the second preset condition, the second prediction mode is determined based on the pixel values of the first reference block. For example, the second preset condition can be the relationship between the size of the first reference block and the first size. For example, if the size of the first reference block is greater than or equal to the first size, the second prediction mode is allowed to be determined based on the pixel values of the first reference block. If the size of the first reference block is smaller than the first size, the second prediction mode is not allowed to be determined based on the pixel values of the first reference block. For reference blocks with larger sizes, the second prediction mode determined based on their pixel values may consume too much time. Therefore, by setting the above-mentioned second preset condition, such reference blocks can be excluded, thereby saving the prediction time of the current block.
[0187] As mentioned above, the current block may be predicted based on the candidate set. In some implementations, an intra prediction mode for the current block may be determined based on the candidate set, and a prediction value for the current block may be determined based on the intra prediction mode of the current block.
[0188] In some implementations, the method of FIG. 10 may further include writing first identification information into the bitstream. The first identification information is used to indicate that the prediction mode of the current block is a target prediction mode. The target prediction mode is a mode for performing prediction based on a candidate set. For example, the target prediction mode may include one or more of MPM, TIMD, TMRL, SGPM, and GPM.
[0189] The first identification information may be represented by, for example, m_flag (m represents the type of prediction mode, for example, m may be MPM). Of course, the first identification information may also be represented by any other letters and / or numbers.
[0190] In some implementations, the method of FIG10 may further include: writing first index information into the bitstream. The first index information is used to determine the intra prediction mode of the current block from the candidate set.
[0191] In some implementations, the index information may include an index value, based on which the intra prediction mode of the current block may be determined directly from the candidate set.
[0192] In some implementations, the method of FIG. 10 further includes: transforming the residual information to determine transform coefficients; quantizing the transform coefficients to obtain quantized coefficients; encoding the quantized coefficients and writing the encoded bits into the bitstream.
[0193] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 10 . The device embodiment of the present application is described in detail below in conjunction with Figures 11 to 14 . It should be understood that the description of the method embodiment corresponds to the description of the device embodiment. Therefore, for parts not described in detail, reference can be made to the above method embodiment.
[0194] FIG11 is a schematic diagram of the structure of a decoder provided by an embodiment of the present application. As shown in FIG11 , the decoder 1100 includes: a first determining unit 1120 , a second determining unit 1130 , an adding unit 1130 , and a predicting unit 1140 .
[0195] The first determination unit 1100 is configured to determine at least one reference position around the current block, the at least one reference position including a first reference position, the first reference position corresponding to a first reference block, the first reference block is predicted based on a first prediction mode, and the first prediction mode does not belong to a target intra-frame prediction mode, and the target intra-frame prediction mode at least includes an angle prediction mode.
[0196] The second determining unit 1120 is configured to determine a second prediction mode according to the pixel value of the first reference block, where the second prediction mode belongs to the target intra prediction mode.
[0197] The adding unit 1130 is configured to add the second prediction mode to the candidate set of intra prediction modes.
[0198] The prediction unit 1140 is configured to predict the current block according to the candidate set.
[0199] In some implementations, the first prediction mode is a prediction mode based on motion information.
[0200] In some implementations, the first prediction mode includes an intra-block coding mode and / or an inter-frame prediction mode.
[0201] In some implementations, the first prediction mode includes only intra-frame block coding mode; or, the first prediction mode includes only inter-frame prediction mode.
[0202] In some implementations, the first prediction mode does not include a geometric partitioning mode.
[0203] In some implementations, determining the second prediction mode based on the pixel value of the first reference block includes: determining gradient information corresponding to at least one angle based on the pixel value of the first reference block; and determining the second prediction mode based on the gradient information corresponding to the at least one angle.
[0204] In some implementations, the gradient information corresponding to the at least one angle is determined based on a DIMD mode derived from an intra-frame prediction mode at the decoding end.
[0205] In some implementations, the gradient information corresponding to the at least one angle includes an amplitude value corresponding to the at least one angle, and the method further includes: if the amplitude value corresponding to the at least one angle satisfies a first preset condition, adding the second prediction mode to the candidate set of intra-frame prediction modes.
[0206] In some implementations, the first preset condition is associated with a maximum amplitude value among the amplitude values corresponding to the at least one angle.
[0207] In some implementations, the first preset condition includes: the maximum amplitude value is greater than or equal to a target value, and the target value is determined based on the sum of remaining amplitude values excluding the maximum amplitude value among the amplitude values corresponding to the at least one angle.
[0208] In some implementations, determining the second prediction mode based on the pixel value of the first reference block includes: if the size of the first reference block meets a second preset condition, determining the second prediction mode based on the pixel value of the first reference block.
[0209] In some implementations, the first reference block is a prediction block or a reconstructed block.
[0210] In some implementations, the target intra-frame prediction mode further includes: a planar mode and a direct current mode.
[0211] In some implementations, predicting the current block based on the candidate set includes: determining an intra-frame prediction mode of the current block based on the candidate set; and determining a prediction value of the current block based on the intra-frame prediction mode of the current block.
[0212] In some implementations, the decoder 1100 further includes:
[0213] A decoding unit is configured to determine first identification information, where the first identification information is used to indicate that the prediction mode of the current block is a target prediction mode, and the target prediction mode is predicted based on the candidate set; determine first index information; and predict the current block based on the candidate set, including: determining the intra-frame prediction mode of the current block from the candidate set based on the first index information.
[0214] The third determining unit is configured to determine a residual value of the current block; and determine a reconstructed value of the current block according to the predicted value of the current block and the residual value of the current block.
[0215] The fourth determining unit is configured to perform an inverse transform on the transform coefficients to determine residual information of the current block.
[0216] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0217] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0218] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the decoder 1100. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the decoding method in the first embodiment.
[0219] Based on the composition of the above-mentioned decoder 1100 and the computer-readable storage medium, refer to Figure 12, which shows a specific hardware structure diagram of the decoder 1200 provided in an embodiment of the present application. As shown in Figure 12, the decoder 1200 may include: a communication interface 1210, a memory 1220 and a processor 1230; each component is coupled together through a bus system 1240. It can be understood that the bus system 1240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 1240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 1240 in Figure 12. Among them,
[0220] The communication interface 1210 is used to receive and send signals when sending and receiving information with other external network elements;
[0221] Memory 1220, used for storing computer programs;
[0222] The processor 1230 is configured to, when running the computer program, execute:
[0223] determining at least one reference position around the current block, the at least one reference position including a first reference position, the first reference position corresponding to a first reference block, the first reference block being predicted based on a first prediction mode, the first prediction mode not belonging to a target intra-frame prediction mode, the target intra-frame prediction mode including at least an angular prediction mode;
[0224] determining a second prediction mode according to the pixel value of the first reference block, where the second prediction mode belongs to the target intra prediction mode;
[0225] adding the second prediction mode to a candidate set of intra prediction modes;
[0226] The current block is predicted according to the candidate set.
[0227] It is understood that the memory 1220 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1220 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0228] The processor 1230 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 1230. The above-mentioned processor 1230 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1220 , and the processor 1230 reads the information in the memory 1220 and completes the steps of the above method in combination with its hardware.
[0229] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0230] Optionally, as another embodiment, the processor 1230 is further configured to execute the decoding method described in the above embodiment when running the computer program.
[0231] FIG13 is a schematic diagram of the structure of an encoder provided by an embodiment of the present application. As shown in FIG13 , the encoder 1300 includes: a first determining unit 1310 , a second determining unit 1320 , an adding unit 1330 , and a predicting unit 1340 .
[0232] The first determination unit 1310 is configured to determine at least one reference position around the current block, the at least one reference position including a first reference position, the first reference position corresponding to a first reference block, the first reference block is predicted based on a first prediction mode, and the first prediction mode does not belong to a target intra-frame prediction mode, and the target intra-frame prediction mode at least includes an angle prediction mode.
[0233] The second determining unit 1320 is configured to determine a second prediction mode according to the pixel value of the first reference block, where the second prediction mode belongs to the target intra prediction mode.
[0234] The adding unit 1330 is configured to add the second prediction mode to the candidate set of intra prediction modes.
[0235] The prediction unit 1340 is configured to predict the current block according to the candidate set.
[0236] In some implementations, the first prediction mode is a prediction mode based on motion information.
[0237] In some implementations, the first prediction mode includes an intra-block coding mode and / or an inter-frame prediction mode.
[0238] In some implementations, the first prediction mode includes only intra-frame block coding mode; or, the first prediction mode includes only inter-frame prediction mode.
[0239] In some implementations, the first prediction mode does not include a geometric partitioning mode.
[0240] In some implementations, determining the second prediction mode based on the pixel value of the first reference block includes: determining gradient information corresponding to at least one angle based on the pixel value of the first reference block; and determining the second prediction mode based on the gradient information corresponding to the at least one angle.
[0241] In some implementations, the gradient information corresponding to the at least one angle is determined based on a DIMD mode derived from an intra-frame prediction mode at the decoding end.
[0242] In some implementations, the gradient information corresponding to the at least one angle includes an amplitude value corresponding to the at least one angle, and the method further includes: if the amplitude value corresponding to the at least one angle satisfies a first preset condition, adding the second prediction mode to the candidate set of intra-frame prediction modes.
[0243] In some implementations, the first preset condition is associated with a maximum amplitude value among the amplitude values corresponding to the at least one angle.
[0244] In some implementations, the first preset condition includes: the maximum amplitude value is greater than or equal to a target value, and the target value is determined based on the sum of remaining amplitude values excluding the maximum amplitude value among the amplitude values corresponding to the at least one angle.
[0245] In some implementations, determining the second prediction mode based on the pixel value of the first reference block includes: if the size of the first reference block meets a second preset condition, determining the second prediction mode based on the pixel value of the first reference block.
[0246] In some implementations, the first reference block is a prediction block or a reconstructed block.
[0247] In some implementations, the target intra-frame prediction mode further includes: a planar mode and a direct current mode.
[0248] In some implementations, predicting the current block based on the candidate set includes: determining an intra-frame prediction mode of the current block based on the candidate set; and determining a prediction value of the current block based on the intra-frame prediction mode of the current block.
[0249] In some implementations, the encoder 1300 further includes:
[0250] The encoding unit is configured to write first identification information into a bitstream, where the first identification information is used to indicate that the prediction mode of the current block is a target prediction mode, and the target prediction mode is predicted based on the candidate set; write first index information into the bitstream, where the first index information is used to indicate a position of the intra-frame prediction mode of the current block in the candidate set; and encode a quantization coefficient of the current block.
[0251] The third determining unit is configured to determine the residual value of the current block according to the prediction value of the current block.
[0252] The fourth determining unit is configured to determine a quantization coefficient of the current block according to the residual value of the current block.
[0253] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0254] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0255] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 1300. The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the decoding method in the first embodiment.
[0256] Based on the composition of the above-mentioned encoder 1300 and the computer-readable storage medium, refer to Figure 14, which shows a specific hardware structure diagram of the encoder 1400 provided in an embodiment of the present application. As shown in Figure 14, the encoder 1400 may include: a communication interface 1410, a memory 1420 and a processor 1430; each component is coupled together through a bus system 1440. It can be understood that the bus system 1440 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 1440 in Figure 14. Among them,
[0257] The communication interface 1410 is used to receive and send signals when sending and receiving information with other external network elements;
[0258] Memory 1420, for storing computer programs;
[0259] The processor 1430 is configured to, when running the computer program, execute:
[0260] determining at least one reference position around the current block, the at least one reference position including a first reference position, the first reference position corresponding to a first reference block, the first reference block being predicted based on a first prediction mode, the first prediction mode not belonging to a target intra-frame prediction mode, the target intra-frame prediction mode including at least an angular prediction mode;
[0261] determining a second prediction mode according to the pixel value of the first reference block, where the second prediction mode belongs to the target intra prediction mode;
[0262] adding the second prediction mode to a candidate set of intra prediction modes;
[0263] The current block is predicted based on the candidate set. It can be understood that the memory 1420 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1420 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0264] The processor 1430 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 1430. The above-mentioned processor 1430 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1420 , and the processor 1430 reads the information in the memory 1420 and completes the steps of the above method in combination with its hardware.
[0265] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0266] Optionally, as another embodiment, the processor 1430 is further configured to execute the encoding method described in the above embodiment when running the computer program.
[0267] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0268] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0269] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0270] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0271] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0272] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A decoding method, applied to a decoder, the method comprising: Determine at least one reference position around the current block, the at least one reference position includes a first reference position, the first reference position corresponds to a first reference block, the first reference block is predicted based on a first prediction mode, and the first prediction mode does not belong to a target intra-frame prediction mode, and the target intra-frame prediction mode at least includes an angular prediction mode; Determining a second prediction mode according to the pixel value of the first reference block, where the second prediction mode belongs to the target intra prediction mode; Adding the second prediction mode to a candidate set of intra prediction modes; The current block is predicted according to the candidate set.
2. The method according to claim 1, wherein: The first prediction mode is a prediction mode based on motion information.
3. The method according to claim 1 or 2, wherein: The first prediction mode includes an intra-frame block coding mode and / or an inter-frame prediction mode.
4. The method according to any one of claims 1 to 3, wherein: The first prediction mode includes only the intra block coding mode; or, the first prediction mode includes only the inter prediction mode.
5. The method according to any one of claims 1 to 4, wherein: The first prediction mode does not include a geometric partitioning mode.
6. The method according to any one of claims 1 to 5, wherein: The determining the second prediction mode according to the pixel value of the first reference block includes: Determining, according to the pixel values of the first reference block, gradient information corresponding to at least one angle; The second prediction mode is determined according to gradient information corresponding to the at least one angle.
7. The method according to claim 6, wherein the gradient information corresponding to the at least one angle is determined based on a DIMD mode derived from an intra-frame prediction mode at the decoding end.
8. The method according to claim 6 or 7, wherein: The gradient information corresponding to the at least one angle includes an amplitude value corresponding to the at least one angle, and the method further includes: If the amplitude value corresponding to the at least one angle satisfies a first preset condition, the second prediction mode is added to a candidate set of intra prediction modes.
9. The method according to claim 8, wherein: The first preset condition is associated with a maximum amplitude value among the amplitude values corresponding to the at least one angle.
10. The method according to claim 9, wherein: The first preset condition includes: the maximum amplitude value is greater than or equal to a target value, and the target value is determined based on the sum of remaining amplitude values excluding the maximum amplitude value among the amplitude values corresponding to the at least one angle.
11. The method according to any one of claims 1 to 10, wherein: The determining the second prediction mode according to the pixel value of the first reference block includes: If the size of the first reference block meets a second preset condition, a second prediction mode is determined according to the pixel value of the first reference block.
12. The method according to any one of claims 1 to 11, wherein: The first reference block is a prediction block or a reconstructed block.
13. The method according to any one of claims 1 to 12, wherein: The target intra-frame prediction mode also includes: a planar mode and a direct current mode.
14. The method according to any one of claims 1 to 13, wherein: The predicting the current block according to the candidate set includes: Determining an intra prediction mode of the current block according to the candidate set; A prediction value of the current block is determined according to the intra prediction mode of the current block.
15. The method according to any one of claims 1 to 14, wherein: The method further comprises: Parsing a bitstream to determine a residual value of the current block; A reconstructed value of the current block is determined according to the predicted value of the current block and the residual value of the current block.
16. The method according to any one of claims 1 to 15, wherein: The method further comprises: Parse the bitstream and determine first identification information, where the first identification information is used to indicate that the prediction mode of the current block is a target prediction mode, and the target prediction mode is predicted based on the candidate set.
17. The method according to any one of claims 1 to 16, wherein: The method further comprises: Parse the code stream to determine the first index information; The predicting the current block according to the candidate set includes: According to the first index information, an intra prediction mode of the current block is determined from the candidate set.
18. A coding method, applied to an encoder, the method comprising: Determine at least one reference position around the current block, the at least one reference position includes a first reference position, the first reference position corresponds to a first reference block, the first reference block is predicted based on a first prediction mode, and the first prediction mode does not belong to a target intra-frame prediction mode, and the target intra-frame prediction mode at least includes an angular prediction mode; Determining a second prediction mode according to the pixel value of the first reference block, where the second prediction mode belongs to the target intra prediction mode; Adding the second prediction mode to a candidate set of intra prediction modes; The current block is predicted according to the candidate set.
19. The method according to claim 18, wherein: The first prediction mode is a prediction mode based on motion information.
20. The method according to claim 18 or 19, wherein: The first prediction mode includes an intra-frame block coding mode and / or an inter-frame prediction mode.
21. The method according to any one of claims 18 to 20, wherein: The first prediction mode includes only the intra block coding mode; or, the first prediction mode includes only the inter prediction mode.
22. The method according to any one of claims 18 to 21, wherein: The first prediction mode does not include a geometric partitioning mode.
23. The method according to any one of claims 18 to 22, wherein: The determining the second prediction mode according to the pixel value of the first reference block includes: Determining, according to the pixel values of the first reference block, gradient information corresponding to at least one angle; The second prediction mode is determined according to gradient information corresponding to the at least one angle.
24. The method according to claim 23, wherein the gradient information corresponding to the at least one angle is determined based on a DIMD mode derived from an intra-frame prediction mode at the decoding end.
25. The method according to claim 23 or 24, wherein: The gradient information corresponding to the at least one angle includes an amplitude value corresponding to the at least one angle, and the method further includes: If the amplitude value corresponding to the at least one angle satisfies a first preset condition, the second prediction mode is added to a candidate set of intra prediction modes.
26. The method according to claim 25, wherein: The first preset condition is associated with a maximum amplitude value among the amplitude values corresponding to the at least one angle.
27. The method according to claim 26, wherein: The first preset condition includes: the maximum amplitude value is greater than or equal to a target value, and the target value is determined based on the sum of remaining amplitude values excluding the maximum amplitude value among the amplitude values corresponding to the at least one angle.
28. The method according to any one of claims 18 to 27, wherein: The determining the second prediction mode according to the pixel value of the first reference block includes: If the size of the first reference block meets a second preset condition, a second prediction mode is determined according to the pixel value of the first reference block.
29. The method according to any one of claims 18 to 28, wherein: The first reference block is a prediction block or a reconstructed block.
30. The method according to any one of claims 18 to 29, wherein: The target intra-frame prediction mode also includes: a planar mode and a direct current mode.
31. The method according to any one of claims 18 to 30, wherein: The predicting the current block according to the candidate set includes: Determining an intra prediction mode of the current block according to the candidate set; A prediction value of the current block is determined according to the intra prediction mode of the current block.
32. The method according to any one of claims 18 to 31, wherein: The method further comprises: Determining a residual value of the current block according to the prediction value of the current block; Determining a quantization coefficient of the current block according to the residual value of the current block; The quantized coefficients of the current block are encoded.
33. The method according to any one of claims 18 to 32, wherein: The method further comprises: Writing first identification information into a bitstream, the first identification information is used to indicate that the prediction mode of the current block is a target prediction mode, and the target prediction mode is predicted based on the candidate set.
34. A method according to any one of claims 18 to 33, wherein: The method further comprises: The first index information is written into a bitstream, where the first index information is used to indicate a position of the intra-frame prediction mode of the current block in the candidate set.
35. A decoder comprising: A first determining unit is configured to determine at least one reference position around the current block, the at least one reference position includes a first reference position, the first reference position corresponds to a first reference block, the first reference block is predicted based on a first prediction mode, and the first prediction mode does not belong to a target intra-frame prediction mode, and the target intra-frame prediction mode at least includes an angle prediction mode; A second determining unit is configured to determine a second prediction mode according to a pixel value of the first reference block, where the second prediction mode belongs to the target intra prediction mode; an adding unit, configured to add the second prediction mode to a candidate set of intra prediction modes; The prediction unit is configured to predict the current block according to the candidate set.
36. A decoder comprising: Memory for storing computer programs; A processor, configured to execute the method according to any one of claims 1 to 17 when running the computer program.
37. An encoder comprising: A first determining unit is configured to determine at least one reference position around the current block, the at least one reference position includes a first reference position, the first reference position corresponds to a first reference block, the first reference block is predicted based on a first prediction mode, and the first prediction mode does not belong to a target intra-frame prediction mode, and the target intra-frame prediction mode at least includes an angle prediction mode; A second determining unit is configured to determine a second prediction mode according to a pixel value of the first reference block, where the second prediction mode belongs to the target intra prediction mode; an adding unit, configured to add the second prediction mode to a candidate set of intra prediction modes; The prediction unit is configured to predict the current block according to the candidate set.
38. An encoder comprising: Memory for storing computer programs; A processor, configured to execute the method according to any one of claims 18 to 34 when running the computer program.
39. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 17 or the method according to any one of claims 18 to 34 is implemented.
40. A non-volatile computer-readable storage medium storing a bit stream, wherein the bit stream is generated by an encoding method using an encoder, or the bit stream is decoded by a decoding method using a decoder, wherein: The decoding method is the method according to any one of claims 1 to 17, and the encoding method is the method according to any one of claims 18 to 34.
Citation Information
Patent Citations
Intra-frame prediction method, image coding method, image decoding method and apparatus
CN114938449A
Decoder end intra mode derivation
CN115462077A
Decoder-side intra mode derivation for most probable mode list construction in video coding
CN116636209A
Method, device, and medium for video processing
WO2022214028A1