Coding method, decoding method, bitstream, coder, decoder and storage medium
By utilizing the statistical results of the prediction mode set of decoded or encoded blocks in video encoding, multiple intra-frame prediction modes are selected for weighted prediction, which solves the problem of high computational complexity of intra-frame prediction modes and achieves more efficient encoding/decoding and video compression performance improvement.
Patent Information
- Application Number
- PCT/CN2024/099956
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2026-02-05
AI Technical Summary
In existing video coding standards, intra-frame prediction mode has high computational complexity, resulting in low encoding and decoding efficiency and difficulty in meeting the bandwidth requirements of high-resolution video.
By determining the statistical results in the set of prediction modes of decoded or encoded blocks, multiple intra-frame prediction modes are selected for weighted prediction, reducing computational complexity and improving encoding and decoding efficiency.
It reduces computational complexity, improves encoding and decoding efficiency, enhances video compression performance, and reduces prediction errors.
Smart Images

Figure CN2024099956_05022026_PF_FP_ABST
Abstract
Description
Coding method, code stream, encoder, decoder and storage medium TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of video coding technology, and in particular to a coding method, a code stream, an encoder, a decoder and a storage medium. BACKGROUND
[0002] With the increasing demand for video display quality, high-resolution videos such as high-definition and ultra-high-definition videos have emerged. However, high-resolution videos usually have more information, and thus require more bandwidth. To reduce the bandwidth requirement, video coding standards involving video compression have been introduced.
[0003] In the video coding standards, for relatively complex intra prediction modes, they all use the reconstructed area adjacent to the current block for processing analysis to derive the intra prediction mode used by the current block, and they all support the combination of multiple intra prediction modes. However, these relatively complex intra prediction modes all require a certain amount of calculation to derive the intra prediction mode used by the current block, thereby increasing the computational complexity and affecting the coding efficiency.
[0004] SUMMARY
[0005] Embodiments of the present application provide a coding method, a code stream, an encoder, a decoder and a storage medium, which can reduce the computational complexity, improve the coding efficiency, and further improve the compression performance.
[0006] The technical solutions of embodiments of the present application can be implemented as follows:
[0007] In a first aspect, the embodiments of the present application provide a decoding method applied to a decoder, and the method comprises:
[0008] determining one or more preset blocks that have been decoded;
[0009] determining statistical results of at least two candidate intra prediction modes in a first mode set according to the prediction modes used by the one or more preset blocks respectively;
[0010] determining at least two intra prediction modes of the current block in the first mode set according to the statistical results;
[0011] predicting the current block according to the at least two intra prediction modes to determine a prediction value of the current block.
[0012] In a second aspect, the embodiments of the present application provide an encoding method applied to an encoder, and the method comprises:
[0013] determining one or more preset blocks that have been encoded;
[0014] determine, according to the prediction modes used by the one or more preset blocks respectively, a statistical result of at least two candidate intra prediction modes in the first mode set;
[0015] determine, according to the statistical result, at least two intra prediction modes of the current block in the first mode set;
[0016] predict the current block according to the at least two intra prediction modes to determine a prediction value of the current block.
[0017] In a third aspect, an embodiment of the present application provides a code stream, which is generated by bit coding according to the encoding method in the second aspect; wherein the to-be-encoded information in the encoding method includes at least one of the following: a transform coefficient of the current block, a transform kernel group index of the current block, a transform kernel index of the current block, and a value of the first syntax element used to indicate whether the first prediction mode is used for the current block.
[0018] In a fourth aspect, an embodiment of the present application provides an encoder, which includes a first determination unit and a first prediction unit, wherein:
[0019] the first determination unit is configured to determine one or more preset blocks that have been encoded, and determine a statistical result of at least two candidate intra prediction modes in the first mode set according to the prediction modes used by the one or more preset blocks respectively;
[0020] the first determination unit is further configured to determine at least two intra prediction modes of the current block in the first mode set according to the statistical result;
[0021] the first prediction unit is configured to predict the current block according to the at least two intra prediction modes to determine a prediction value of the current block.
[0022] In a fifth aspect, an embodiment of the present application provides an encoder, which includes a first memory and a first processor, wherein:
[0023] the first memory is configured to store a computer program capable of running on the first processor;
[0024] the first processor is configured to execute the steps of the encoding method in the second aspect when the computer program is running.
[0025] In a sixth aspect, an embodiment of the present application provides a decoder, which includes a second determination unit and a second prediction unit, wherein:
[0026] the second determination unit is configured to determine one or more preset blocks that have been decoded, and determine a statistical result of at least two candidate intra prediction modes in the first mode set according to the prediction modes used by the one or more preset blocks respectively;
[0027] The second determining unit is further configured to determine at least two intra prediction modes of the current block from the first mode set according to the statistical result;
[0028] The second predicting unit is configured to predict the current block according to the at least two intra prediction modes, and determine a predicted value of the current block.
[0029] In a seventh aspect, an embodiment of the present application provides a decoder, comprising a second memory and a second processor, wherein:
[0030] The second memory is configured to store a computer program capable of running on the second processor;
[0031] The second processor is configured to execute the steps of the decoding method according to the first aspect when running the computer program.
[0032] In an eighth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the decoding method according to the first aspect, or implement the steps of the encoding method according to the second aspect.
[0033] In a ninth aspect, an embodiment of the present application provides a computer program product, which comprises a computer program or instructions, and the computer program or instructions are executed by a processor to implement the steps of the decoding method according to the first aspect, or implement the steps of the encoding method according to the second aspect.
[0034] In a tenth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a code stream, and the code stream is generated by executing the steps of the encoding method according to the second aspect.
[0035] The embodiment of the present application provides a coding method, a code stream, an encoder, a decoder and a storage medium. At the encoding end, one or more preset blocks are determined; statistical results of at least two candidate intra-frame prediction modes in a first mode set are determined according to prediction modes used by the one or more preset blocks respectively; at least two intra-frame prediction modes of a current block are determined in the first mode set according to the statistical results; and a prediction value of the current block is determined by predicting the current block according to the at least two intra-frame prediction modes. At the decoding end, one or more preset blocks are determined; statistical results of at least two candidate intra-frame prediction modes in a first mode set are determined according to prediction modes used by the one or more preset blocks respectively; at least two intra-frame prediction modes of a current block are determined in the first mode set according to the statistical results; and a prediction value of the current block is determined by predicting the current block according to the at least two intra-frame prediction modes. In this way, the statistical results of the related blocks are obtained by using the related blocks which have been coded and decoded, and the at least two intra-frame prediction modes used by the current block are derived according to the statistical results. Since it is not necessary to indicate which kind of intra-frame prediction mode is used, the overhead in the code stream can be saved. Moreover, the prediction accuracy can be improved by using the weighted prediction of the at least two intra-frame prediction modes, and the prediction error can be reduced to a certain extent. In addition, the at least two intra-frame prediction modes used by the current block are derived according to the statistical results, which can avoid the intra-frame prediction mode derivation by gradient calculation in the DIMD mode, and the calculation complexity is reduced, so that the coding efficiency can be improved, and the compression performance is improved. BRIEF DESCRIPTION OF DRAWINGS
[0036] Fig. 1 is a flow block diagram of a hybrid coding framework;
[0037] Fig. 2 is a template matching diagram of a current block;
[0038] Fig. 3 is a reference sample diagram of a current block;
[0039] Fig. 4 is a multi-reference line diagram of a current block;
[0040] Fig. 5 is a diagram of a plurality of prediction modes corresponding to intra-frame prediction;
[0041] Fig. 6 is a diagram of a plurality of prediction modes corresponding to intra-frame prediction;
[0042] Fig. 7 is a diagram of a plurality of prediction modes corresponding to intra-frame prediction;
[0043] Fig. 8 is a diagram of a plurality of prediction modes corresponding to intra-frame prediction;
[0044] Fig. 9 is an encoding diagram of screen content;
[0045] Fig. 10 is a prediction flow diagram of a MIP mode;
[0046] FIG. 11 is a diagram illustrating a template and a template reference region of a current block;
[0047] FIG. 12 is a diagram illustrating a histogram of gradients and intra prediction modes;
[0048] FIG. 13 is a diagram illustrating a weighted combination of three intra prediction modes;
[0049] FIG. 14 is a diagram illustrating a search range of an ITMP mode;
[0050] FIG. 15 is a diagram illustrating a plurality of mode weights in a GPM mode;
[0051] FIG. 16A is a diagram illustrating a typical filter 1;
[0052] FIG. 16B is a diagram illustrating a typical filter 2;
[0053] FIG. 16C is a diagram illustrating a typical filter 3;
[0054] FIG. 17A is a diagram illustrating a reconstructed sample region for training a filter 1;
[0055] FIG. 17B is a diagram illustrating a reconstructed sample region for training a filter 2;
[0056] FIG. 17C is a diagram illustrating a reconstructed sample region for training a filter 3;
[0057] FIG. 18 is a diagram illustrating a DCT transform;
[0058] FIG. 19 is a diagram illustrating a base image of a DCT transform;
[0059] FIG. 20 is a diagram illustrating a flowchart of a LFNST-free transform;
[0060] FIG. 21 is a diagram illustrating a flowchart of a LFNST transform;
[0061] FIG. 22 is a diagram illustrating a detailed flowchart of a LFNST transform;
[0062] FIG. 23 is a diagram illustrating a base image of a plurality of transform kernel groups;
[0063] FIG. 24 is a diagram illustrating a base image of an NSPT transform;
[0064] FIG. 25 is a diagram illustrating a network architecture of a video coding according to an embodiment of the present application;
[0065] FIG. 26 is a diagram illustrating a system composition block diagram of an encoder according to an embodiment of the present application;
[0066] FIG. 27 is a system block diagram of a decoder according to an embodiment of the present application;
[0067] FIG. 28 is a flowchart of a decoding method according to an embodiment of the present application;
[0068] FIG. 29 is a diagram of spatial relationships between a current block and neighboring blocks and non-neighboring blocks according to an embodiment of the present application;
[0069] FIG. 30 is a diagram of spatial relationships between a current block and a preset range according to an embodiment of the present application;
[0070] FIG. 31 is a flowchart of a decoding method according to an embodiment of the present application;
[0071] FIG. 32 is a diagram of a structure of a candidate sample according to an embodiment of the present application;
[0072] FIG. 33 is a flowchart of a decoding method according to an embodiment of the present application;
[0073] FIG. 34 is a flowchart of a decoding method according to an embodiment of the present application;
[0074] FIG. 35 is a flowchart of an encoding method according to an embodiment of the present application;
[0075] FIG. 36 is a flowchart of an encoding method according to an embodiment of the present application;
[0076] FIG. 37 is a diagram of a structure of an encoder according to an embodiment of the present application;
[0077] FIG. 38 is a diagram of a specific hardware structure of an encoder according to an embodiment of the present application;
[0078] FIG. 39 is a diagram of a structure of a decoder according to an embodiment of the present application;
[0079] FIG. 40 is a diagram of a specific hardware structure of a decoder according to an embodiment of the present application;
[0080] FIG. 41 is a diagram of a structure of a codec system according to an embodiment of the present application. DETAILED DESCRIPTION
[0081] In order to enable a person skilled in the art to more fully understand the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings, which are used only for reference and are not used to limit the embodiments of the present application.
[0082] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application.
[0083] In the following description, reference is made to the "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" can be the same subset or a different subset of all possible embodiments, and can be combined with each other without conflict.
[0084] It should also be noted that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects, and do not represent the specific order of the objects. It can be understood that "first\second\third" can be interchanged with specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0085] In a video image, a coding block (CB) is generally represented by a first color component, a second color component and a third color component. Among them, the three color components are a luminance component, a blue chroma component and a red chroma component, respectively. Specifically, the luminance component is usually represented by the symbol Y, the blue chroma component is usually represented by the symbol Cb or U, and the red chroma component is usually represented by the symbol Cr or V. In this way, the video image can be represented in YCbCr format or YUV format.
[0086] Before the embodiments of the present application are further described in detail, the terms and terms involved in the embodiments of the present application are explained, which are applicable to the following explanations:
[0087] H.265 / High Efficiency Video Coding (HEVC);
[0088] H.266 / Versatile Video Coding (VVC);
[0089] VVC Test Model (VTM) is the reference software test platform of VVC;
[0090] Enhanced Compression Model (ECM) is the platform for improving compression performance after VVC;
[0091] Joint Video Experts Team (JVET);
[0092] coding unit (CU);
[0093] coding tree unit (CTU);
[0094] largest coding unit (LCU);
[0095] motion vector (MV);
[0096] prediction unit (PU);
[0097] transform unit (TU);
[0098] merge technique (Merge);
[0099] skip technique (Skip);
[0100] quantization parameter (QP);
[0101] merge with motion vector difference (MMVD);
[0102] motion vector prediction (MVP);
[0103] temporal motion vector prediction (TMVP);
[0104] discrete cosine transform (DCT);
[0105] discrete sine transform (DST);
[0106] multiple transform selection (MTS);
[0107] low frequency non-separable transform (LFNST);
[0108] non-separable primary transform (NSPT);
[0109] Context-based Adaptive Binary Arithmetic Coding (CABAC).
[0110] At present, the general video coding standards all adopt a block-based hybrid coding framework. Each picture or sub-picture or frame in a video is divided into square maximum coding units or coding tree units of the same size (such as 256x256, 128x128, 64x64, etc.). Each maximum coding unit or coding tree unit can be divided into rectangular coding units according to a rule. The coding units can also be divided into prediction units, transform units, etc. Specifically, as shown in FIG. 1, the hybrid coding framework 10 can include a pre-encode filtering module 11, an intra prediction module 12, an inter prediction module 13, a motion estimation module 14, a transform module 15, a quantization module 16, an entropy coding module 17, an inverse quantization module 18, an inverse transform module 19, a loop filter module 20, a decoded picture buffer module 21, etc. Among them, the prediction here can include intra prediction and inter prediction, and the inter prediction can include motion estimation and motion compensation. Since there is a strong correlation between adjacent samples in a picture of a video, the intra prediction method is used in the video coding technology to eliminate the spatial redundancy between adjacent samples. In addition, since there is a strong similarity between adjacent pictures in a video, the inter picture prediction method is used in the video coding technology to eliminate the temporal redundancy between adjacent pictures, thereby improving the coding efficiency. It should be noted that the sample can also be referred to as a pixel, and the sample includes position information and value.
[0111] The basic process of a video codec is as follows: at the encoding end, an image is divided into blocks, a prediction block of a current block is generated using intra prediction or inter prediction, a residual block is obtained by subtracting the prediction block from the initial block of the current block, a quantized coefficient matrix is obtained by transforming and quantizing the residual block, and the quantized coefficient matrix is entropy encoded and output to a bitstream. At the decoding end, a prediction block of a current block is generated using intra prediction or inter prediction, and a quantized coefficient matrix is obtained by parsing the bitstream. The quantized coefficient matrix is dequantized and inverse transformed to obtain a residual block, and the prediction block and the residual block are added to obtain a reconstructed block. The reconstructed blocks form a reconstructed image, and the reconstructed image is loop filtered based on the image or based on the block to obtain a decoded image. The encoding end also needs to perform similar operations as the decoding end to obtain a decoded image. The decoded image can be used as a reference image for subsequent image inter prediction. If necessary, the block division information, prediction, transformation, quantization, entropy encoding, loop filtering, and other mode information or parameter information determined by the encoding end need to be output to the bitstream. The decoding end analyzes and determines the same block division information, prediction, transformation, quantization, entropy encoding, loop filtering, and other mode information or parameter information as the encoding end by parsing the bitstream, so as to ensure that the decoded image obtained by the encoding end is the same as the decoded image obtained by the decoding end. The decoded image obtained by the encoding end is usually also called a reconstructed image. The current block can be divided into a prediction unit during prediction, and the current block can be divided into a transform unit during transformation. The division of the prediction unit and the transform unit can be different. The above is the basic process of a video codec under a block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process can be optimized. The embodiments of the present application are applicable to the basic process of the video codec under the block-based hybrid coding framework, but are not limited to the framework and process.
[0112] In addition, in the embodiments of the present application, the current block (CB) can be a current coding unit, a current prediction unit, or a current transform unit, etc. Due to the need for parallel processing, an image can be divided into slices, etc., and the slices in the same image can be processed in parallel, that is, there is no data dependency between them. "Frame" is a commonly used term, which can generally be understood as one frame being one image. In the embodiments of the present application, the frame can also be replaced by an image or a slice, etc.
[0113] The related schemes of the prediction technology will be introduced in detail below.
[0114] (I) Template matching
[0115] The method of template matching (TM) is first used in inter prediction, which uses the correlation between adjacent samples to take some areas around the current block as templates. When the current block is coded, its left and top sides have been coded according to the coding order. Of course, in the existing hardware decoder implementation, it is not necessarily guaranteed that the left and top sides of the current block have been decoded when the current block starts to be decoded. Here, we refer to inter blocks, such as in HEVC, the blocks coded in inter mode generate prediction blocks without the need for the reconstructed samples around them, so the prediction process of inter blocks can be carried out in parallel. However, the blocks coded in intra mode definitely need the reconstructed samples on the left and top sides as reference samples. In theory, the left and top sides are available, that is, the hardware design can be adjusted accordingly. Relatively speaking, the right and bottom sides are not available in the coding order of the existing standards such as VVC.
[0116] As shown in Fig. 2, a rectangular area on the left and on the top of the current block is set as a template, the height of the template on the left is generally the same as the height of the current block, and the width of the template on the top is generally the same as the width of the current block, of course, it can also be different. The best matching position of the template is searched in the reference image to determine the motion information or motion vector of the current block. This process can be described as follows: in a certain reference image (Ref0), starting from a starting position, a search is performed within a certain range. The search rules such as search range and search step can be set in advance. The matching degree of the template corresponding to the position and the template on the periphery of the current block is calculated every time the position is moved. The matching degree can be measured by some distortion cost, such as sum of absolute difference (SAD), sum of absolute transformed difference (SATD), mean-square error (MSE), etc. The transformation used by SATD is Hadamard transformation, and the smaller the values of SAD, SATD and MSE are, the higher the matching degree is. The cost is calculated by the prediction block of the template corresponding to the position and the reconstructed block of the template on the periphery of the current block. In addition to the search of the integer sample position, the search of the fractional sample position can also be performed, and the motion information of the current block is determined according to the position with the highest matching degree. The correlation between adjacent samples is used, and the motion information suitable for the template can also be suitable for the current block. Of course, the template matching method can not be suitable for all blocks, and therefore some methods can be used to determine whether the template matching method is used for the current block, such as a control switch indicating whether the template matching method is used for the current block. The name of this template matching method is decoder side motion vector derivation (DMVD). The encoder and the decoder can use the template to search and derive the motion information or find better motion information on the basis of the original motion information. It does not need to transmit specific motion vector or motion vector difference, but the same rule search is performed by the encoder and the decoder to ensure the consistency of encoding and decoding. The template matching method can improve the compression performance, but it needs to perform "search" at the decoding end, thereby bringing certain decoding end complexity.
[0117] (ii) Intra prediction.
[0118] It can be understood that there is strong spatial correlation between adjacent parts or adjacent samples in an image, and the intra prediction is a prediction method using the spatial correlation of the samples in the periphery of the current block and the samples in the current block. For example, as shown in FIG. 3, the 4x4 white filled samples are the current block, the grid filled samples on the left side and the top side of the current block are the reference samples of the current block, and the intra prediction uses these reference samples to predict the current block. These reference samples can all be available, that is, all have been coded. Some of them can also be unavailable, for example, if the current block is the leftmost in the frame, the reference samples on the left side of the current block are unavailable. Or when the current block is coded, the part below the left side of the current block has not been coded, and the reference samples below the left side are also unavailable. For the case where the reference samples are unavailable, the available reference samples or some values or some methods can be used for filling, or no filling is performed.
[0119] The intra prediction method of multiple reference lines (MRL) can use more reference samples to improve the coding efficiency. As shown in FIG. 4, here is an example of using four reference lines / columns.
[0120] There are various prediction modes for intra prediction, as shown in FIG. 5, here are nine modes for intra prediction of a 4x4 block in H.264. Among them, mode 0 (vertical mode) is to copy the values of the samples above the current block to the current block as prediction values in the vertical direction, mode 1 (horizontal mode) is to copy the values of the reference samples on the left side to the current block as prediction values in the horizontal direction, mode 2 (DC mode) is to use the average value of points A-D and I-L as the prediction value of all points, and modes 3-8 are to copy the values of the reference samples to the corresponding positions of the current block at a certain angle, respectively, because some positions of the current block cannot be exactly matched to the reference samples, it can be necessary to use the weighted average value of the reference samples, or the sub-sample of the interpolated reference samples.
[0121] In addition, there are PLANE, PLANAR and other modes, and with the development of technology and the expansion of the block, the angle prediction mode is also increasing. For example, the intra prediction modes used by HEVC include PLANAR, DC and 33 angle modes, a total of 35 prediction modes, as shown in FIG. 6. The intra prediction modes used by VVC include PLANAR, DC and 65 angle modes, a total of 67 prediction modes, as shown in FIG. 7. Of course, in addition to the above 67 modes, VVC also provides wide angle modes for some rectangular blocks with large differences in length and width, such as the modes indicated by the dashed lines in FIG. 8, which are -14-1 and 67-80 intervals, which will replace some regular modes, as shown in FIG. 8.
[0122] (Three) Intra Block Copy (IBC).
[0123] IBC can significantly improve the compression efficiency of screen content coding (SCC), so IBC is used for screen content coding from HEVC to VVC. Screen content is different from camera captured content, which is computer generated, screen content has no noise, contains text, computer graphics, etc., and the boundary is clear. There is a lot of repeated content in the screen content, as shown in FIG. 9.
[0124] In the embodiments of the present application, IBC can be considered as using the method of inter prediction to intra prediction. Among them, inter prediction uses the reference block on the reference image to generate the prediction block of the current block, and the reference image is not the current image. And IBC is to find the reference block from the already coded or reconstructed part of the current image to generate the prediction block of the current block. IBC can also be called intra picture block compensation or current picture referencing (CPR).
[0125] IBC can use block vector (BV) to represent the position difference between the current block and the reference block, which is similar to the MV of inter prediction. The encoder determines the best matching block of the current block in the search range through block matching method, and encodes the BV. There are many methods for encoding BV, such as using merge mode, which is similar to inter prediction, which will not be described here.
[0126] IBC can be considered as a kind of intra prediction method, or as another kind of prediction method independent of intra prediction and inter prediction. IBC has high efficiency for screen content coding, and can also improve the compression efficiency in natural sequences captured by cameras.
[0127] (Four) Matrix-based Intra Prediction (MIP).
[0128] For MIP, it is a special intra prediction mode, which can also be called Matrix weighted Intra Prediction in some places.
[0129] As shown in FIG. 10, in order to predict a block with width W and height H, MIP needs H reconstructed samples on the left side of the current block and W reconstructed samples on the top side of the current block as input. MIP generates the predicted block in the following three steps: (a) Averaging, (b) Matrix Vector Multiplication, and (c) Interpolation. Here, the core of MIP is considered to be the Matrix Vector Multiplication. It can be considered as a process to generate the predicted block from the input samples (reference samples) in a way of matrix multiplication. MIP provides a variety of matrices, and the difference of the prediction modes is embodied in the difference of the matrices. Different matrices will result in different results using the same input samples. The processes of Averaging and Interpolation are a design of compromise between performance and complexity. For a block with a large size, Averaging can achieve an effect of approximate down-sampling, so that the input can be adapted to a smaller matrix, and Interpolation can achieve an effect of up-sampling. In this way, it is not necessary to provide matrices for MIP for every size of block, but only to provide one or several matrices for a specific size. With the increasing demand for compression performance and the increasing hardware capability, the next generation of standards may have MIP with higher complexity.
[0130] MIP is somewhat similar to PLANAR, but obviously MIP is more complex and flexible than PLANAR.
[0131] (five) Template-based Intra Mode Derivation (TIMD).
[0132] As shown in FIG. 11, for a current block, a region on the left side and on the top side of the current block is taken as a template. Except for the boundary condition, the left side and the top side of the current block can theoretically obtain reconstructed values when the current block is coded. This is the basis of many template adaptation methods. TIMD takes the diagonal filled region as shown in FIG. 11 as a template, and the reference of the template in FIG. 11 is the reference samples of the template (the grid filled region). The decoder can use a certain intra prediction mode to predict on the template, and compare the predicted value with the reconstructed value to obtain the cost of the intra prediction mode on the template, such as SAD, SATD, SSE, etc. Since the template and the current block are adjacent, they are related, so the performance of a prediction mode on the template can be used to estimate the performance of the prediction mode on the current block. TIMD predicts some candidate intra prediction modes on the template to obtain their costs on the template, and selects one or two intra prediction modes with the lowest cost as the intra prediction mode of the current block.
[0133] It is found that if the cost difference between two intra prediction modes is not large, the compression performance can be improved by weighting the prediction values of the two intra prediction modes. The weight of the prediction values of the two prediction modes is related to the cost mentioned above. In the current version, the weight is inversely proportional to the cost.
[0134] In summary, TIMD uses the prediction effect of the intra prediction modes on the template to screen the intra prediction modes, and can weight two intra prediction modes according to the cost on the template. The advantage of TIMD is that if the current block selects the TIMD mode, it does not need to indicate which intra prediction mode is used, but is derived by the decoder itself through the above process, which saves the overhead to some extent. Using two intra prediction modes for weighting can also bring compression performance improvement, and theoretically it can use more than two intra prediction modes for weighting.
[0135] (VI) Decoder-side Intra Mode Derivation (DIMD).
[0136] DIMD uses the reconstructed samples on the left and top of the current block to derive the prediction mode, but it does not predict on the template, but analyzes the gradient of the reconstructed samples.
[0137] As shown in FIG. 12, DIMD analyzes the gradient of the black point, such as the horizontal gradient and the vertical gradient, and adapts an intra prediction mode according to the gradient thereof.
[0138] In the embodiments of the present application, the gradient can be calculated using the Sobel operator. Exemplarily, for the Sobel operator, the following is specific:
[0139] The operator of the horizontal gradient is:
[0140] The operator of the vertical gradient is:
[0141] Thus, assuming that the sample value of the sample at position (x, y) in the prediction block is P x,y , the horizontal gradient grad x and the vertical gradient grad y are calculated as follows: grad x = P x+1,y-1 + 2*P x+1,y + P x+1,y+1 - P x-1,y-1 - 2*P x-1,y - P x-1,y+1 (1) grad y = P x-1,y+1+2*P x,y+1 +P x+1,y+1 -P x-1,y-1 -2*P x,y-1 -P x+1,y -1 (2)
[0142] Here, the gradient strength is denoted as amp, and an example is amp = abs(grad x ) + abs(grad y ).
[0143] The virtual intra prediction mode derived from grad x and grad y can be implemented by a lookup table. For example, if abs(grad x ) equals 0 and abs(grad y ) does not equal 0, which indicates there is horizontal texture, then the corresponding intra prediction mode is 18 in VVC. If abs(grad y ) equals 0 and abs(grad x ) does not equal 0, which indicates there is vertical texture, then the corresponding intra prediction mode is 50 in VVC. In the case that abs(grad x ) and abs(grad y ) do not equal 0, if abs(grad x ) equals abs(grad y ) and grad x and grad y have the same sign, then the corresponding intra prediction mode is 34 in VVC; if abs(grad x ) equals 2 times abs(grad y ) and grad x and grad y have the same sign, then the corresponding intra prediction mode is 40 in VVC. In addition, other cases can be determined by the same principle of lookup table.
[0144] The analysis of all the points to be checked can result in a histogram-like result as shown below. That is, the statistics of the number of points matched by each intra prediction mode. Of course, the histogram is just for understanding and in implementation it can be realized in many simple forms, for example, by using an array. Now the DIMD selects the top 2 intra prediction modes in the histogram, plus the PLANAR mode, a total of 3 intra prediction modes, and the prediction values of the 3 intra prediction modes are weighted, and the weights are related to the analysis results. Exemplarily, as shown in FIG. 13, the 3 intra prediction modes include an M1 mode, an M2 mode, and a PLANAR mode. The prediction values obtained for the 3 intra prediction modes are set as Pred1, Pred2, and Pred3 respectively, and the weight values of the 3 intra prediction modes are set as w1, w2, and w3 respectively, and the specific calculation formula is as follows:
[0145] The final prediction block can be as shown below:
[0146] In summary, DIMD uses gradient analysis of reconstructed samples to screen intra prediction modes, and can weight 2 intra prediction modes plus planar according to the analysis results. The advantage of DIMD is that if the current block selects the DIMD mode, it does not need to indicate which intra prediction mode is actually used, but is derived by the decoder itself through the above process, which saves the overhead to some extent. Specifically, using multiple intra prediction modes for weighting can also bring compression performance improvement, and in theory it can use more intra prediction modes for weighting. An example is to use at most 5 intra angular prediction modes and planar for weighting, or to use at most 5 intra angular prediction modes and IBC or ITMP for weighting.
[0147] (Seven) Intra Template Matching Prediction (ITMP).
[0148] ITMP can be considered as a technology that combines IBC and TM. It has been mentioned above that TM applied to inter can reduce the overhead of coding MVs, similarly, TM used on IBC can reduce the overhead of coding BVs. An example is that there is no need to code BVs, and the matching block found by TM is directly used as the prediction block of the ITMP mode of the current block.
[0149] An example of ITMP is shown in FIG. 14, where the inverse L-shaped region at the top-left corner of the current block is used as the template to search in the search range filled with dots, and the search range is the reconstructed region. The dot-filled region shown in FIG. 14 includes the current CTU R1, the CTU at the top-left side R2, the CTU at the top side R3, and the CTU at the left side R4. This is an example, and the search range can be different in actual applications. The best matching block is found in R2 in this example.
[0150] (eight) Spatial Geometric Partitioning Mode (SGPM).
[0151] In the VVC video coding standard, there is an inter prediction mode called Geometric Partitioning Mode (GPM). In the AVS3 video coding standard, there is an inter prediction mode called Angular Weighted Prediction (AWP). Although the names are different, the specific implementation forms are different, but the principles are the same.
[0152] Traditional uni-prediction only finds one reference block with the same size as the current block. Traditional bi-prediction uses two reference blocks with the same size as the current block, and the sample value of each point of the prediction block is the average of the corresponding positions of the two reference blocks, i.e. each point of each reference block accounts for 50% of the proportion. Bi-directional weighted prediction makes the proportions of the two reference blocks different, such as all points in the first reference block accounting for 75% of the proportion and all points in the second reference block accounting for 25% of the proportion. But the proportions of all points in the same reference block are the same. Other optimization methods such as decoder side motion vector refinement (DMVR) and bi-directional optical flow (BIO) will make some changes to the reference samples or prediction samples, but are not related to the above principles. BIO can also be abbreviated as BDOF. GPM or AWP also uses two reference blocks with the same size as the current block, but some sample positions use 100% of the sample value of the corresponding position of the first reference block, some sample positions use 100% of the sample value of the corresponding position of the second reference block, and in the interface area or transition area, the sample values of the corresponding positions of the two reference blocks are used in a certain proportion. The weights of the interface area are also gradually transitioned. How these weights are distributed is determined by the mode of GPM or AWP. The weight of each sample position is determined according to the mode of GPM or AWP. Of course, in some cases, such as when the block size is very small, it may not be possible to guarantee that there are certain sample positions that use 100% of the sample value of the corresponding position of the first reference block and certain sample positions that use 100% of the sample value of the corresponding position of the second reference block in some modes of GPM or AWP. It can also be considered that GPM or AWP uses two reference blocks with different sizes than the current block, i.e. taking only a part as the reference block. That is, the part with a weight of 0 is excluded.
[0153] As shown in FIG. 15, it is a weight map of 64 modes of GPM in VVC on a square block. Black represents that the weight value of the corresponding position of the first reference block is 0%, white represents that the weight value of the corresponding position of the first reference block is 100%, and the gray area represents that the weight value of the corresponding position of the first reference block is greater than 0% and less than 100% according to the color depth. The weight value of the corresponding position of the second reference block is 100% minus the weight value of the corresponding position of the first reference block.
[0154] The weight derivation methods of GPM and AWP are different. GPM determines the angle and offset according to each mode, and then calculates the weight matrix of each mode. AWP first makes a one-dimensional weight line, and then uses a method similar to intra angular prediction to cover the entire matrix with the one-dimensional weight line.
[0155] It should be noted that in the early coding standards, there is only a rectangular partitioning mode for CU, PU and TU. GPM and AWP achieve the effect of non-rectangular partitioning of prediction without partitioning. GPM and AWP use a mask of weights of two reference blocks, that is, the weight map or weight matrix mentioned above. This mask determines the weights of the two reference blocks when generating the prediction block, or can be simply understood as that part of the prediction block comes from the first reference block and part of the prediction block comes from the second reference block, and the transition area is obtained by weighting the corresponding positions of the two reference blocks, so that the transition is smoother. GPM and AWP do not divide the current block into two CUs or PUs according to the partitioning line, so the residual after prediction, such as transformation, quantization, inverse transformation, inverse quantization, etc. are all processed as a whole.
[0156] It should also be noted that GPM is an inter-frame technology in VVC, but it can also use intra-frame prediction. The two prediction modes of GPM can both be inter-frame prediction modes, one can be inter-frame prediction mode and the other can be intra-frame prediction mode, or both can be intra-frame prediction modes.
[0157] The SGPM mode in ECM uses this weight mask or weight matrix in intra-frame prediction. It uses the weight matrix to combine the prediction values of two intra-frame prediction modes into the prediction value of SGPM, which can produce more complex textures than a single prediction mode.
[0158] Because 1 "partitioning" mode and 2 intra-frame prediction modes are used, the general logic is to write syntax elements indicating the 3 modes in the code stream as shown in Table 1, that is, partition_mode_idx, intra_pred_mode0_idx, intra_pred_mode1_idx as shown in Table 1. However, in order to reduce the overhead of this information, SGPM uses the template around the current block to sort the combination of the 3 modes to obtain a candidate list of the combined mode, and only the candidate index of SGPM needs to be written in the code stream. At the decoding end, SGPM can construct a candidate list of the combined mode, and according to the candidate index of SGPM, sgpm_cand_idx in Table 2, derive 1 "partitioning" mode and 2 intra-frame prediction modes partition_mode_idx, intra_pred_mode0_idx, intra_pred_mode1_idx.
[0159] Table 1
[0160] Table 2
[0161] Here, for sgpm_cand_idx,
[0162] (ix) Extrapolation filter based Intra Prediction (EIP).
[0163] There is an EIP mode in ECM, and FIG. 16A, FIG. 16B and FIG. 16C are several typical filters of the EIP mode. Among them, the sample positions of the grid filled part represent the input of EIP, and the sample positions of the white filled part represent the output of EIP. Taking the filter of the first square as an example, usually the prediction value of the sample in the lower right corner is required, the values of the 15 samples on the left side, the upper left side and the upper side of the sample are known, and each grid filled position of the filter has a preset coefficient. Multiply the sample value of each position of the input by the corresponding coefficient, add all the multiplication results and normalize, and the prediction value of the to-be-predicted position can be obtained.
[0164] The EIP mode can provide a filter that has been pre-trained, or a filter that is trained according to the surrounding reconstructed sample region of the current block. FIG. 17A, FIG. 17B and FIG. 17C represent three kinds of reconstructed sample regions that can be selected to train the filter. It is assumed that the current block and the reconstructed sample region used to train the filter coefficients follow similar rules, so the filter trained by the reconstructed sample region can be applied to the current block.
[0165] Further, the following is related to the introduction of the transformation technology.
[0166] In encoding, the hybrid coding framework commonly used nowadays first makes a prediction, which uses the correlation in space or time to get an image that is the same or similar to the current block. For a block, it is possible that the predicted block is exactly the same as the current block, but it is difficult to guarantee that all blocks in a video are like this, especially for natural videos, or videos shot by a camera. Irregular motion, distortion, occlusion, brightness changes, etc. in the video are difficult to be completely predicted. Therefore, the hybrid coding framework subtracts the predicted image from the original image of the current block to get a residual image, or subtracts the predicted block from the current block to get a residual block. The residual block is usually much simpler than the original image, so prediction can significantly improve compression efficiency. Instead of directly encoding the residual block, a transform is usually performed first. The transform is to transform the residual image from the spatial domain to the frequency domain to remove the correlation of the residual image. After the residual image is transformed into the frequency domain, the non-zero coefficients are mostly concentrated in the upper left corner because the energy is mostly concentrated in the low frequency area. Next, quantization is used to further compress. Moreover, since the human eye is not sensitive to high frequencies, a larger quantization step size can be used for the high frequency area.
[0167] FIG. 18 is a schematic diagram of a DCT transform. As shown in FIG. 18, after the original image is subjected to the DCT transform, only the upper left corner area has non-zero coefficients. Of course, this example is for the DCT transform of the entire image, while in video coding, the image is divided into blocks for processing, and thus the transform is also performed based on the blocks.
[0168] Transforms are very useful in the compression of common videos, but not all blocks must be transformed. In some cases, the compression effect is better without transform than with transform, so in some standards such as VVC, the encoder can choose whether to use transform for the current block. Among them, DCT2 type (DCT-II) is the most commonly used transform in video compression standards, and the basic image of its transform is shown in FIG. 19.
[0169] In addition, DCT8 type (DCT-VIII) and DST7 type (DST-VII) can also be used in VVC. The basic formulas of these transforms are shown in Table 3, which shows the basic transform formulas of DCT2, DCT8 and DST7 for N-point input.
[0170] Table 3
[0171] Since images are two-dimensional, the amount of computation and memory overhead of direct two-dimensional transform is unacceptable under the hardware conditions at that time, so the above-mentioned DCT2, DCT8 and DST7 transforms used in the standard are all split into horizontal and vertical one-dimensional transforms and performed in two steps. For example, first perform the horizontal transform and then perform the vertical transform, or first perform the vertical transform and then perform the horizontal transform.
[0172] (I) Multi-transform selection (MTS).
[0173] VVC supports DCT2, DCT8, DST7, etc. transform kernel, the DCT2, DCT8, DST7, etc. transform kernel used in VVC is a horizontal and numerical separable transform kernel, which can be applied to horizontal or vertical direction separately. For a block, the encoder can select the appropriate transform kernel and transmit the index to the bitstream, and the decoder determines the inverse transform kernel according to the index. The horizontal and vertical directions can select different transform kernels, such as DCT8 for horizontal direction and DST7 for vertical direction. This technique is generally referred to as MTS.
[0174] VVC uses a syntax element mts_idx to determine the transform kernel of the basic transform. As shown in Table 4 below, where trTypeHor represents the transform kernel of the horizontal transform, trTypeVer represents the transform kernel of the vertical transform, and 0 of trTypeHor and trTypeVer represents DCT2 type transform, 1 represents DST7 type transform, and 2 represents DCT8 type transform. If mts_idx is not present, the value of mts_idx is inferred to be 0.
[0175] Table 4
[0176] (II) Low frequency non-separable transform (LFNST).
[0177] The above transform method is relatively effective for horizontal and vertical textures, but the effect on diagonal textures is not as good. Indeed, horizontal and vertical textures are the most common, so the above transform method is very useful for improving compression efficiency. With the increasing demand for compression efficiency, if diagonal textures can be processed more effectively, the compression efficiency can be further improved.
[0178] In order to more effectively process the residuals of diagonal texture, LFNST transform is used in VVC. The above-mentioned transforms such as DCT2, DCT8, DST7 are referred to as primary transform. At the encoding end of VVC, LFNST is used after DCT2 transform and before quantization. At the decoding end of VVC, LFNST is used after dequantization and before inverse DCT2 transform. Because it is a transform on the basis of DCT2 (primary transform), LFNST is a secondary transform. FIG. 20 is a schematic diagram of the encoding and decoding process without LFNST (secondary transform). As shown in FIG. 20, the process at the encoding end is in turn: transform 211 - quantization 212 - entropy encoding 213, and the obtained encoded bits are written into the bitstream; and the process at the decoding end is in turn: entropy decoding 214 - dequantization 215 - inverse transform 216. FIG. 21 is a schematic diagram of the encoding and decoding process with LFNST (secondary transform). As shown in FIG. 21, the process at the encoding end is in turn: DCT2 transform (primary transform) 221 - LFNST transform (secondary transform) 222 - quantization 223 - entropy encoding 224, and the obtained encoded bits are written into the bitstream; and the process at the decoding end is in turn: entropy decoding 225 - dequantization 226 - inverse LFNST transform (secondary transform) 227 - inverse DCT2 transform (primary transform) 228. It should be noted that at the encoding end, dequantization can also be directly performed on the saved quantization coefficients without entropy decoding, because entropy encoding is lossless.
[0179] FIG. 22 is a schematic diagram of a detailed encoding and decoding process with LFNST (secondary transform). At the encoding end, LFNST performs secondary transform on the low-frequency coefficients in the upper left corner after primary transform. Primary transform concentrates energy in the upper left corner by decorrelating the image. Secondary transform decorrelates the low-frequency coefficients of primary transform, and the result is intuitively shown in FIG. 22. At the encoding end, 16 coefficients are input to a 4x4 LFNST, and 8 coefficients are output; 48 coefficients are input to an 8x8 LFNST, and 8 coefficients are output for an 8x8 block, and 16 coefficients are output for other blocks. At the decoding end, 8 coefficients are input to a 4x4 inverse LFNST, and 16 coefficients are output; 8 coefficients are input to an 8x8 block, and 16 coefficients are input to other blocks, and 48 coefficients are input to an 8x8 inverse LFNST, and 48 coefficients are output.
[0180] FIG. 23 is some base images of LFNST in VVC. In FIG. 23, only 2 base images of the lowest frequency of each transform kernel in each transform kernel group are shown. It can be seen that there are some obvious diagonal textures. LFNST has transform kernels optimized for certain diagonal textures, and also has transform kernels optimized for flat gradient textures, such as transform kernel group 0 of LFNST in VVC.
[0181] Angle prediction tiles the values of reference samples to the current block as prediction values at a specified angle, which means that the prediction block will have a clear directional texture, and the residual of the current block after angle prediction will also have a clear angle characteristic in statistics. Thus, the transform kernel selected by LFNST can be bound to the intra prediction mode, that is, after the intra prediction mode is determined, LFNST can only use the set of transform kernels corresponding to the intra prediction mode.
[0182] Specifically, LFNST in VVC has a total of 4 sets of transform kernels, and each set can select 2 transform kernels. Table 5 shows the correspondence between the intra prediction mode and the transform kernel set. Note that the cross-component prediction modes used for chroma intra prediction are 81 to 83, and there are no such modes for luma intra prediction. The transform kernel of LFNST can be transposed to correspond to more angles with one transform kernel set. For example, modes 13 to 23 and 45 to 55 correspond to transform kernel set 2, but 13 to 23 are obviously close to horizontal modes and 45 to 55 are obviously close to vertical modes.
[0183] Table 5
[0184] LFNST of VVC has a total of 4 sets of transform kernels, and according to the intra prediction mode, it is specified which set of LFNST is used. This makes use of the correlation between the intra prediction mode and the transform kernel of LFNST, thereby reducing the transmission of the selected transform kernel of LFNST in the code stream. Whether the current block will use LFNST, and if it uses LFNST, whether to use the first or the second in a set, needs to be determined through the code stream and some conditions.
[0185] In the subsequent evolution of ECM technology, LFNST is further expanded. LFNST has more transform kernel sets, 35 sets in ECM, and the correspondence between the transform kernel set index (LFNST set index) and the intra prediction mode (Intra pred. mode) is as shown in Table 6. Each transform kernel set is more efficient for the texture of the corresponding angle. Here, each transform kernel set can select 3 transform kernels.
[0186] Table 6
[0187] (Three) Non-separable basic transform NSPT.
[0188] LFNST is a horizontal-vertical non-separable transform. Since there is a secondary transform, DCT2 can be called a base transform. So, first applying DCT2 and then applying LFNST is a kind of performance and complexity compromise solution. Since a non-separable base transform has higher efficiency but higher complexity, such as higher computation and higher transform kernel storage space.
[0189] In ECM10, some small blocks can use NSPT, while large blocks still use DCT2+LFNST. The size of the small blocks is, for example, 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, 16x8, 8x32, 32x8. In ECM10, NSPT also matches transform kernel groups according to the intra prediction mode. The matching method can refer to the method of LFNST. Each transform kernel group has 3 transform kernels to choose from. Exemplarily, an 8x8 base image of one NSPT in ECM10 is shown in FIG. 24, which corresponds to inter angular prediction mode 7. It can be seen that it is better at processing textures corresponding to the angle.
[0190] In summary, for relatively complex intra prediction modes (such as TIMD and DIMD), TIMD and DIMD have some similarities. They both use the reconstructed regions adjacent to the current block to analyze and derive the intra prediction mode to be used by the current block. The adjacent regions have strong correlation with the current block, so in many cases, a reasonable intra prediction mode can be derived. In addition, they both support the combination of multiple intra prediction modes. The combination of multiple intra prediction modes can reduce the prediction error. Since they do not need to use syntax elements in the code stream to indicate which intra prediction mode is used, that is, even if multiple intra prediction modes are used for weighting, there is no additional indication cost.
[0191] However, in the related art, TIMD and DIMD both need a certain amount of computation to derive the intra prediction mode to be used by the current block. For example, TIMD needs to generate prediction values on the template and calculate the matching cost, and DIMD needs to calculate the gradient, thereby increasing the computational complexity and affecting the coding efficiency.
[0192] Based on this, the embodiment of the present application provides an encoding method, determining one or more preset blocks which have been encoded; determining statistical results of at least two candidate intra prediction modes in a first mode set according to prediction modes used by the one or more preset blocks respectively; determining at least two intra prediction modes of a current block in the first mode set according to the statistical results; and predicting the current block according to the at least two intra prediction modes to determine a prediction value of the current block. The embodiment of the present application also provides a decoding method, determining one or more preset blocks which have been decoded; determining statistical results of at least two candidate intra prediction modes in a first mode set according to prediction modes used by the one or more preset blocks respectively; determining at least two intra prediction modes of a current block in the first mode set according to the statistical results; and predicting the current block according to the at least two intra prediction modes to determine a prediction value of the current block.
[0193] In this way, the statistical results of the related blocks are obtained by using the related blocks which have been encoded or decoded, and the at least two intra prediction modes used by the current block are derived according to the statistical results. Since it is not necessary to explicitly indicate which kind of intra prediction mode is used, the overhead in the code stream can be saved. Moreover, the prediction accuracy can be improved by using the weighted prediction of the at least two intra prediction modes, and the prediction error is reduced to a certain extent. In addition, the at least two intra prediction modes used by the current block are derived according to the statistical results, which can avoid the gradient calculation for deriving the intra prediction mode in the DIMD mode, and the calculation complexity is reduced, so that the encoding and decoding efficiency can be improved, and the compression performance is improved.
[0194] The embodiments of the present application will be described in detail below with reference to the drawings.
[0195] FIG. 25 is a schematic diagram of a network architecture of a video encoding and decoding provided by an embodiment of the present application. As shown in FIG. 25, the network architecture includes one or more electronic devices 31 to 3N and a communication network 01, wherein the electronic devices 31 to 3N can perform video interaction through the communication network 01. The electronic devices in the implementation can be various types of devices with video encoding and decoding functions, for example, the electronic devices can include a mobile phone, a tablet computer, a personal computer, a personal digital assistant, a navigation instrument, a digital telephone, a video telephone, a television, a sensing device, a server, etc., and the embodiments of the present application are not limited thereto.
[0196] In the embodiments of the present application, a network architecture of a video encoding and decoding system including a decoding method and an encoding method is provided. The decoder or the encoder in the embodiments of the present application can be the electronic devices described above. That is, the electronic devices in the embodiments of the present application have video encoding and decoding functions, and generally include a video encoder (i.e., an encoder) and a video decoder (i.e., a decoder).
[0197] FIG. 26 is a schematic diagram of a system composition of an encoder according to an embodiment of the present application. As shown in FIG. 26, the encoder 100 can include a partitioning unit 101, a prediction unit 102, a first adder 107, a transform unit 108, a quantization unit 109, a dequantization unit 110, an inverse transform unit 111, a second adder 112, a filter unit 113, a decoded picture buffer (DPB) unit 114, and an entropy encoding unit 115. Here, the input of the encoder 100 can be a video composed of a series of pictures or a still picture, and the output of the encoder 100 can be a bitstream (also referred to as a "code stream") representing a compressed version of the input video.
[0198] The partitioning unit 101 partitions a picture in the input video into one or more coding tree units (CTUs). The partitioning unit 101 partitions a picture into a plurality of tiles (or tiles), and can further partition a tile into one or more bricks. Here, a tile or a brick can include one or more complete and / or partial CTUs. In addition, the partitioning unit 101 can form one or more slices, where a slice can include one or more tiles arranged in raster order in a picture, or one or more tiles covering a rectangular region in a picture. The partitioning unit 101 can also form one or more sub-pictures, where a sub-picture can include one or more slices, tiles, or bricks.
[0199] In the encoding process of the encoder 100, the partition unit 101 delivers a CTU to the prediction unit 102. Generally, the prediction unit 102 can be composed of a block partition unit 103, a motion estimation (ME) unit 104, a motion compensation (MC) unit 105, and an intra prediction unit 106. Specifically, the block partition unit 103 iteratively partitions the input CTU into smaller coding units (CUs) using quad-tree partitioning, binary-tree partitioning, and ternary-tree partitioning. The prediction unit 102 can obtain an inter prediction block for a CU using the ME unit 104 and the MC unit 105. The intra prediction unit 106 can obtain an intra prediction block for a CU using various intra prediction modes including the MIP mode. In an example, a rate-distortion optimized motion estimation approach can be invoked by the ME unit 104 and the MC unit 105 to obtain the inter prediction block, and a rate-distortion optimized mode determination approach can be invoked by the intra prediction unit 106 to obtain the intra prediction block. The prediction unit 102 outputs the prediction block for the CU, and the first adder 107 calculates the difference between the CU in the output of the partition unit 101 and the prediction block for the CU, i.e., a residual CU. The transform unit 108 reads the residual CU and performs one or more transform operations on the residual CU to obtain coefficients. The quantization unit 109 quantizes the coefficients and outputs quantized coefficients (i.e., levels). The inverse quantization unit 110 performs a scaling operation on the quantized coefficients to output reconstructed coefficients. The inverse transform unit 111 performs one or more inverse transforms corresponding to the transforms in the transform unit 108 and outputs a reconstructed residual. The second adder 112 calculates a reconstructed CU by adding the reconstructed residual and the prediction block for the CU from the prediction unit 102. The second adder 112 also sends its output to the prediction unit 102 to be used as an intra prediction reference. After all CUs in a picture or subpicture are reconstructed, the filter unit 113 performs loop filtering on the reconstructed picture or subpicture. Here, the filter unit 113 contains one or more filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), luma mapping with chroma scaling (LMCS) filter, and a neural network-based filter, etc. Alternatively, when the filter unit 113 determines that a CU is not to be used as a reference for encoding of other CUs, the filter unit 113 performs loop filtering on one or more target samples in the CU. The output of the filter unit 113 is a decoded picture or subpicture, which is buffered to the DPB unit 114. The DPB unit 114 outputs the decoded picture or subpicture according to timing and control information.Here, the pictures stored in the DPB unit 114 can also be used as references for the inter prediction or intra prediction performed by the prediction unit 102. The final entropy encoding unit 115 converts the parameters necessary for decoding the pictures from the encoder 100 (such as control parameters and supplemental information, etc.) into binary form and writes such binary form into the bitstream according to the syntax structure of each data unit, i.e., the bitstream that the encoder 100 finally outputs.
[0200] Further, the encoder 100 can be a computing device having a first processor and a first memory storing a computer program. When the first processor reads and runs the computer program, the encoder 100 reads the input video and generates the corresponding bitstream. In addition, the encoder 100 can also be a computing device having one or more chips. These units implemented as integrated circuits on the chips have similar connection and data exchange functions as the corresponding units in FIG. 26.
[0201] FIG. 27 is a block diagram of a system composition of a decoder according to an embodiment of the present application. As shown in FIG. 27, the decoder 200 can include a parsing unit 201, a prediction unit 202, an inverse quantization unit 205, an inverse transformation unit 206, an adder 207, a filtering unit 208, and a decoded picture buffer unit 209. Here, the input of the decoder 200 is a bitstream representing a compressed version of a video or a still picture, and the output of the decoder 200 can be a decoded video composed of a series of pictures or a decoded still picture.
[0202] The input bitstream of the decoder 200 can be the bitstream generated by the encoder 100. The parsing unit 201 parses the input bitstream and obtains the values of the syntax elements from the input bitstream. The parsing unit 201 converts the binary representation of the syntax elements into digital values and sends the digital values to the units in the decoder 200 to obtain one or more decoded pictures. The parsing unit 201 can also parse one or more syntax elements from the input bitstream to display the decoded pictures.
[0203] In the decoding process of the decoder 200, the parsing unit 201 sends the values of the syntax elements and one or more variables for obtaining one or more decoded pictures, which are set or determined according to the values of the syntax elements, to the units in the decoder 200. The prediction unit 202 determines the prediction block of the current decoded block (e.g. CU). Here, the prediction unit 202 can include a motion compensation unit 203 and an intra prediction unit 204. Specifically, when the inter decoding mode is indicated for decoding the current decoded block, the prediction unit 202 passes the related parameters from the parsing unit 201 to the motion compensation unit 203 to obtain the inter prediction block; when the intra prediction mode (including the MIP mode indicated based on the MIP mode index value) is indicated for decoding the current decoded block, the prediction unit 202 passes the related parameters from the parsing unit 201 to the intra prediction unit 204 to obtain the intra prediction block. The inverse quantization unit 205 has the same function as the inverse quantization unit 110 in the encoder 100. The inverse quantization unit 205 performs a scaling operation on the quantized coefficients (i.e. levels) from the parsing unit 201 to obtain the reconstructed coefficients. The inverse transform unit 206 has the same function as the inverse transform unit 111 in the encoder 100. The inverse transform unit 206 performs one or more transform operations (i.e. the inverse operations of the one or more transform operations performed by the inverse transform unit 111 in the encoder 100) to obtain the reconstructed residual. The adder 207 performs an addition operation on its inputs (the prediction block from the prediction unit 202 and the reconstructed residual from the inverse transform unit 206) to obtain the reconstructed block of the current decoded block. The reconstructed block is also sent to the prediction unit 202 to be used as the reference for other blocks encoded in the intra prediction mode.
[0204] After all the CUs in a picture or sub-picture are reconstructed, the filter unit 208 performs loop filtering on the reconstructed picture or sub-picture. The filter unit 208 contains one or more filters, such as a deblocking filter, a sample adaptive offset filter, an adaptive loop filter, a luma mapping and chroma scaling filter, and a neural network based filter, etc. Alternatively, when the filter unit 208 determines that the reconstructed block is not used as the reference for decoding other blocks, the filter unit 208 performs loop filtering on one or more target samples in the reconstructed block. Here, the output of the filter unit 208 is the decoded picture or sub-picture, which is buffered to the DPB unit 209. The DPB unit 209 outputs the decoded picture or sub-picture according to the timing and control information. The pictures stored in the DPB unit 209 can also be used as the reference for performing inter prediction or intra prediction by the prediction unit 202.
[0205] Further, the decoder 200 can be a second processor and a second memory recording a computer program. When the first processor reads and runs the computer program, the decoder 200 reads the input code stream and generates the corresponding decoded video. In addition, the decoder 200 can also be a computing device with one or more chips. These units implemented as integrated circuits on the chip have similar connection and data exchange functions as the corresponding units in FIG. 27.
[0206] It should be further noted that when the embodiments of the present application are applied to the encoder 100, the "current block" specifically refers to a current to-be-encoded block (which can also be referred to as "encoding block") in a video image; when the embodiments of the present application are applied to the decoder 200, the "current block" specifically refers to a current to-be-decoded block (which can also be referred to as "decoding block") in a video image.
[0207] In an embodiment of the present application, FIG. 28 is a flow diagram of a decoding method provided by an embodiment of the present application. As shown in FIG. 28, the method can include:
[0208] S2801, determining one or more preset blocks that have been decoded.
[0209] In an embodiment of the present application, the method is applied to a decoder. Specifically, based on the composition structure of the decoder 200 shown in FIG. 27, the decoding method is mainly applied to the intra prediction unit 204 and the inverse transformation unit 206 in FIG. 27. Wherein, when the current block uses the occurrence-based intra prediction (OBIP) mode, the prediction mode used by the related blocks in the current block can be used to derive the intra prediction mode used by the current block. In this way, compared with the TIMD mode and the DIMD mode, which both need a certain amount of calculation to derive the used intra prediction mode, the OBIP mode can save the calculation amount and improve the compression performance.
[0210] It should be further noted that when the current block uses the OBIP mode, the one or more preset blocks that have been decoded need to be determined first, so as to facilitate the subsequent determination of the statistical results of the at least two candidate intra prediction modes in the first mode set.
[0211] In a possible implementation, determining the one or more preset blocks that have been decoded can include: determining at least one candidate block that has been decoded, and determining the one or more preset blocks according to the at least one candidate block.
[0212] In the embodiments of the present application, the candidate blocks can include the neighboring blocks of the current block and / or the non-neighboring blocks of the current block. That is, the one or more preset blocks can only include the neighboring blocks of the current block, or only include the non-neighboring blocks of the current block, or include both the neighboring blocks and the non-neighboring blocks of the current block.
[0213] It can be understood that when the OBIP mode is used for the current block, the neighboring blocks and the non-neighboring blocks of the current block can be counted. FIG. 29 is a schematic diagram of the spatial relationship between the neighboring blocks and the non-neighboring blocks of the current block according to an embodiment of the present application. As shown in FIG. 29, the black filled block represents the current block, the blocks labeled 1-7 are the neighboring blocks of the current block, and the blocks labeled 8-25 are the non-neighboring blocks of the current block. The neighboring blocks refer to the blocks spatially adjacent to the current block, and the non-neighboring blocks refer to the blocks not spatially adjacent to the current block.
[0214] For example, assuming that the top-left position of the current block is (x, y), the width of the current block is width, and the height of the current block is height. Then for the neighboring blocks of the current block, the neighboring block 1 is the block containing the coordinates (x-1, y-1), the neighboring block 2 is the block containing the coordinates (x, y-1), the neighboring block 2 is the block containing the coordinates (x, y-1), the neighboring block 3 is the block containing the coordinates (x-1, y), the neighboring block 4 is the block containing the coordinates (x+width-1, y-1), the neighboring block 5 is the block containing the coordinates (x+width, y-1), the neighboring block 6 is the block containing the coordinates (x-1, y+height-1), and the neighboring block 7 is the block containing the coordinates (x-1, y+height).
[0215] The non-neighboring blocks (i.e. the non-neighboring blocks) can also be positioned according to certain rules. For example, the positions of the non-neighboring blocks can be calculated according to (x, y), width and height. In FIG. 29, the distance of each grid in the vertical direction is height, and the distance of each grid in the horizontal direction is width. Then for the non-neighboring blocks of the current block, the non-neighboring block 8 is the block containing the coordinates (x-width-1, y-height-1), the non-neighboring block 9 is the block containing the coordinates (x+2*width-1, y-height-1), the non-neighboring block 10 is the block containing the coordinates (x-width-1, y+2*height-1), the non-neighboring block 11 is the block containing the coordinates (x-2*width-1, y-2*height-1), the non-neighboring block 12 is the block containing the coordinates (x+width / 2, y-2*height-1), and so on. It should be noted that the block containing the coordinates mentioned herein refers to the CU or PU containing the coordinates.
[0216] Thus, in the embodiments of the present application, for the decoded candidate block, only the neighboring blocks can be used, or only the non-neighboring blocks can be used, or both the neighboring blocks and the non-neighboring blocks can be used, or only specified blocks in the neighboring blocks and the non-neighboring blocks can be used, and no limitation is made herein.
[0217] In another possible implementation, determining the one or more preset blocks can include: determining a preset range based on the current block; and determining the one or more preset blocks based on all decoded blocks in the preset range.
[0218] In the embodiments of the present application, the preset range can be set arbitrarily in the decoded region. For example, assuming that the top-left position of the current block is (x, y), the width of the current block is width, and the height of the current block is height. Then, the preset range can be set according to the position of the current block and the width and the height.
[0219] It can also be understood that when the current block uses the OBIP mode, all blocks in the preset range can be counted. FIG. 30 is a schematic diagram of the spatial relationship between the current block and the preset range according to an embodiment of the present application. As shown in FIG. 30, the black-filled block represents the current block, and the dot-filled part represents the preset range. For example, one possible preset range is the rectangular region with the top-left corner at (x-2*width, y-2*height), the top-right corner at (x+3*width-1, y-2*height), and the bottom-left corner at (x-2*width, y+3*height-1). If the minimum size of the CU or PU is 4*4, all 4*4 blocks in the search range can be counted, for example, the first 4*4 block in the top-left corner is the block containing the coordinate (x-2*width, y-2*height), and the 4*4 blocks in the same row can be scanned by adding 4 horizontally, and the 4*4 blocks in the same column can be scanned by adding 4 vertically. It should be noted that the block containing the coordinate mentioned herein refers to the CU or PU containing the coordinate.
[0220] Thus, in the embodiments of the present application, after the decoded preset range is determined, all blocks found in the preset range can be determined as the one or more preset blocks.
[0221] S2802, determine a statistical result of at least two candidate intra prediction modes in the first mode set according to a prediction mode used by each of the one or more preset blocks.
[0222] In the embodiments of the present application, the first mode set can include at least two candidate intra prediction modes. Wherein, the intra prediction mode can include many modes, such as the PLANAR mode, the DC mode and the 65 angle prediction modes already existing in VVC, the MIP mode, etc. The DIMD mode, the TIMD mode, the SGPM mode, the ITMP mode and the OBIP mode newly added in the ECM, etc. Here, a first mode set can be defined, and in some embodiments, the first mode set includes at least one of the following: angle prediction mode, DC mode and PLANAR mode.
[0223] Exemplarily, the first mode set can only contain angle prediction modes, or can also contain angle prediction modes and PLANAR modes, or can also contain angle prediction modes, DC modes and PLANAR modes, etc., which are not limited here. That is, one possible implementation is that the first mode set only contains angle prediction modes, and the number of these modes can be 65, or can be more because of more fine-grained angle prediction modes. Another possible implementation is that the first mode set only contains angle prediction modes, DC modes and PLANAR modes. The following will be described in detail taking the first mode set containing angle prediction modes, DC modes and PLANAR modes as an example, and the first mode set at this time corresponds to the 0-66 intra prediction modes in VVC.
[0224] In the embodiments of the present application, for the OBIP mode, a data structure can be constructed here to record the occurrence number of each candidate intra prediction mode in the first mode set. Exemplarily, this data structure can be an array, denoted as occurrence
[0067] . It should be noted that because the first mode set is taken as an example of 67 modes, if it is other number, the length of the array is adjusted accordingly. In addition, each value in occurrence
[0067] is initialized to 0.
[0225] In the embodiments of the present application, after determining the one or more preset blocks that have been decoded, the occurrence of the modes in the first mode set in the one or more preset blocks can be counted, that is, the statistical result of the at least two candidate intra prediction modes in the first mode set is determined. In a specific embodiment, it can include: determining the prediction mode used by each of the one or more preset blocks; and according to the prediction mode used by each of the one or more preset blocks, the candidate intra prediction mode is counted to determine the statistical result of the at least two candidate intra prediction modes in the first mode set.
[0226] It should be noted that in the embodiments of the present application, after determining the preset block (for example, CU or PU) containing a certain coordinate in the decoded region, the prediction mode used by the preset block can be determined. For example, if the preset block is an intra prediction block, the intra prediction mode used by the preset block can be determined; if the preset block is an inter prediction block, the inter prediction mode used by the preset block can be determined. In this way, the statistical results of at least two candidate intra prediction modes in the first mode set can be determined according to the prediction modes used by the one or more preset blocks.
[0227] It should be further noted that in the embodiments of the present application, taking the first preset block as an example, for determining the statistical results of at least two candidate intra prediction modes in the first mode set, the method can include: determining the intra prediction mode used by the first preset block; when the intra prediction mode is a candidate intra prediction mode in the first mode set, performing accumulation operation on the intra prediction mode to determine the statistical result of the intra prediction mode in the first mode set.
[0228] It can be understood that in the embodiments of the present application, the first preset block can be any one of the one or more preset blocks. Taking the first preset block as an example, the first preset block can be an inter prediction block or an intra prediction block. The following will be described in detail for the processing case that the first preset block is an inter prediction block.
[0229] In a possible implementation, for determining the statistical results of at least two candidate intra prediction modes in the first mode set, the method can include: when the first preset block is an inter prediction block, not counting the first preset block.
[0230] In the embodiments of the present application, the first preset block can be any one of the one or more preset blocks. If the first preset block is an inter prediction block, it indicates that the first preset block does not use the mode in the first mode set for prediction, and then the first preset block can not be counted (or said to be skipped), that is, at this time, no accumulation operation is performed on any mode in the first mode set, and the statistical results of at least two candidate intra prediction modes in the first mode set can be determined according to the prediction modes used by other preset blocks other than the first preset block.
[0231] In another possible implementation, for determining the statistical results of at least two candidate intra prediction modes in the first mode set, referring to FIG. 31, the method can include:
[0232] S3101, when the first preset block is an inter prediction block, determining a first intra prediction mode derived based on candidate samples of the first preset block.
[0233] S3102, accumulate the first intra prediction mode to determine a statistical result of the first intra prediction mode in the first mode set.
[0234] In the embodiments of the present application, the first preset block can be any one of one or more preset blocks. If the first preset block is an inter prediction block, indicating that the first preset block is not predicted using a mode in the first mode set, the first intra prediction mode of the first preset block can also be derived according to the candidate samples of the first preset block, and then the statistical result of the first intra prediction mode in the first mode set is determined by accumulating the first intra prediction mode. Specifically, the current accumulated value of the first intra prediction mode can be determined by accumulating the accumulated value corresponding to the first intra prediction mode, and the accumulated value corresponding to the first intra prediction mode is updated based on the current accumulated value; the judgment of the next preset block is continued until the one or more preset blocks are traversed, and the final accumulated value corresponding to the first intra prediction mode is determined as the statistical result of the first intra prediction mode in the first mode set.
[0235] In some embodiments, the method can further include: when the first intra prediction mode is a candidate intra prediction mode in the first mode set, accumulating the first intra prediction mode to determine a statistical result of the first intra prediction mode in the first mode set. That is, after deriving the first intra prediction mode based on the candidate samples of the first preset block, if the first intra prediction mode is a candidate intra prediction mode in the first mode set, the accumulation of the first intra prediction mode can be performed to determine the statistical result of the first intra prediction mode in the first mode set.
[0236] In some embodiments, the method can further include: determining the candidate samples of the first preset block according to at least part of the samples in the reconstructed block of the first preset block; or determining the candidate samples of the first preset block according to at least part of the samples in the prediction block of the first preset block.
[0237] In the embodiments of the present application, when the first preset block is an inter prediction block, the first intra prediction mode of the first preset block can be determined by gradient statistics according to the candidate samples of the first preset block. The candidate samples here can be at least part of the samples in the reconstructed block or the prediction block of the block. That is, the first intra prediction mode can be derived by gradient statistics according to at least part of the samples in the reconstructed block or the prediction block of the block. FIG. 32 is a structure diagram of a candidate sample provided in an embodiment of the present application. As shown in FIG. 32, an 8x8 block is provided here, and the 6x6 candidate samples inside it can be used to derive the first intra prediction mode, that is, all points in the reconstructed block or the prediction block except the outermost row and column can be used as candidate samples (filled with grids).
[0238] Thus, if the first preset block is an inter-prediction block, since it is not predicted using a mode in the first mode set, the first preset block can be directly not counted, i.e., no accumulation operation is performed on any mode in the first mode set; or, the gradient counting manner can be used to derive the first intra-prediction mode for the reconstructed block or the predicted block of the first preset block, and the derived first intra-prediction mode is one of the candidate intra-prediction modes in the first mode set. Assuming that the mode index corresponding to the first intra-prediction mode is X, the accumulation operation is performed on occurrence[X].
[0239] It can also be understood that, in the embodiments of the present application, the processing of the case where the first preset block is an intra-prediction block is described in detail as follows.
[0240] In a possible implementation, for the statistical result of the at least two candidate intra-prediction modes in the first mode set, the method further includes: when the first preset block is an intra-prediction block, determining a second intra-prediction mode used by the first preset block; when the second intra-prediction mode is a candidate intra-prediction mode in the first mode set, performing accumulation operation on the second intra-prediction mode to determine the statistical result of the second intra-prediction mode in the first mode set.
[0241] In the embodiments of the present application, when the first preset block is an intra-prediction block, and the second intra-prediction mode used by the first preset block is a candidate intra-prediction mode in the first mode set, assuming that the mode index corresponding to the second intra-prediction mode is X, the accumulation operation is performed on occurrence[X]. The second intra-prediction mode is one of the candidate intra-prediction modes in the first mode set.
[0242] In another possible implementation, the second intra-prediction mode is not a candidate intra-prediction mode in the first mode set, and for the statistical result of the at least two candidate intra-prediction modes in the first mode set, the method further includes: when the second intra-prediction mode is a candidate intra-prediction mode in the second mode set, determining at least two third intra-prediction modes corresponding to the second intra-prediction mode; when the at least two third intra-prediction modes are at least two candidate intra-prediction modes in the first mode set, performing accumulation operation on the at least two third intra-prediction modes to determine the statistical result of the at least two third intra-prediction modes in the first mode set.
[0243] In the embodiments of the present application, the second mode set is different from the first mode set, and the second mode set includes modes using at least two intra-prediction modes for combined prediction.
[0244] In a specific embodiment, the mode for combined prediction using at least two intra prediction modes can be DIMD mode, TIMD mode, SGPM mode and OBIP mode, etc. The intra prediction modes include, but are not limited to, DC mode, PLANAR mode and angular prediction mode, etc.
[0245] That is, in a more specific embodiment, the second mode set can include at least one of the following: DIMD mode, TIMD mode, SGPM mode and OBIP mode.
[0246] In the embodiments of the present application, if the first preset block uses DIMD mode, TIMD mode, SGPM mode, OBIP mode, etc., it can use candidate intra prediction modes in the first mode set, and each candidate intra prediction mode is accumulated. For example, assuming that the mode uses candidate intra prediction modes X1 and X2 in the first mode set, occurrence[X1] and occurrence[X2] are accumulated.
[0247] In another possible implementation, the second intra prediction mode is not a candidate intra prediction mode in the first mode set, and for determining the statistical result of at least two candidate intra prediction modes in the first mode set, the method further includes: when the second intra prediction mode is a candidate intra prediction mode in the third mode set, not counting the first preset block.
[0248] In another possible implementation, the second intra prediction mode is not a candidate intra prediction mode in the first mode set, and for determining the statistical result of at least two candidate intra prediction modes in the first mode set, the method further includes: when the second intra prediction mode is a candidate intra prediction mode in the third mode set, determining a fourth intra prediction mode derived based on the candidate samples of the first preset block; and accumulating the fourth intra prediction mode to determine the statistical result of the fourth intra prediction mode in the first mode set. Alternatively, the method further includes: when the second intra prediction mode is a candidate intra prediction mode in the third mode set, determining a fourth intra prediction mode derived based on the candidate samples of the first preset block; and when the fourth intra prediction mode is a candidate intra prediction mode in the first mode set, accumulating the fourth intra prediction mode to determine the statistical result of the fourth intra prediction mode in the first mode set.
[0249] In the embodiments of the present application, the third mode set is different from the first mode set, and the third mode set includes at least one of the following: a mode for predicting by copying an intra block, a mode for predicting using an extrapolation filter, and a mode for predicting using matrix operation.
[0250] In a specific embodiment, the modes for predicting the intra-block copy are specifically ITMP mode and IBC mode.
[0251] In a specific embodiment, the modes for predicting using extrapolation filter are specifically EIP mode.
[0252] In a specific embodiment, the modes for predicting using matrix operation are specifically MIP mode.
[0253] That is, in a more specific embodiment, the third mode set can include at least one of the following: ITMP mode, IBC mode, EIP mode and MIP mode.
[0254] In the embodiments of the present application, if the first preset block is an intra-prediction block, and the first preset block uses MIP mode, ITMP mode, EIP mode, IBC mode, etc., since it does not use the candidate intra-prediction modes in the first mode set for prediction, the first preset block can be directly skipped, i.e., the first preset block is not counted, and at this time, no accumulation operation is performed on any mode in the first mode set; or, the reconstructed block or the prediction block of the block can be used to derive a fourth intra-prediction mode using the gradient counting method, and the derived fourth intra-prediction mode belongs to one of the candidate intra-prediction modes in the first mode set. Assuming that the mode index corresponding to the fourth intra-prediction mode is X, then the accumulation operation can be performed on occurrence[X].
[0255] Exemplarily, if the first preset block is an intra-prediction block, taking deriving a fourth intra-prediction mode from the first preset block as an example, the accumulation operation is performed on the fourth intra-prediction mode, which can specifically be that the accumulation value corresponding to the fourth intra-prediction mode is added to determine the current accumulation value of the fourth intra-prediction mode, and the accumulation value corresponding to the fourth intra-prediction mode is updated based on the current accumulation value; the judgment of the next preset block is continued, and the final accumulation value corresponding to the fourth intra-prediction mode is determined as the statistical result of the fourth intra-prediction mode in the first mode set after the one or more preset blocks are traversed.
[0256] It can also be understood that in the embodiments of the present application, each intra-prediction mode actually represents a texture feature. Therefore, the intra-prediction mode index is also a texture feature index. For example, the DC mode and the PLANAR mode correspond to gradual texture features, and a certain angle prediction mode corresponds to a texture feature of this angle. The texture feature index can avoid the appearance of intra-prediction mode in the "interframe" on the one hand, and it is also more conducive to possible expansion on the other hand, for example, one intra-prediction mode can correspond to multiple texture features, such as the DC mode can correspond to horizontal gradual texture, vertical gradual texture, and diagonal gradual texture, etc.
[0257] In the embodiments of the present application, when the first preset block-based candidate sample is determined to derive the intra prediction mode, the number of candidate samples used to derive the intra prediction mode (i.e., the "candidate texture feature index") can be at least one, for example, 1, 2, 3, or more.
[0258] In the embodiments of the present application, for the first preset block, the number of candidate samples can be determined according to the size parameter of the first preset block. That is, when one or more candidate texture feature indexes are derived according to the candidate samples, the number of candidate samples used can be determined by the size parameter of the first preset block. For example, if the size of the first preset block is small, all available samples can be counted; if the size of the first preset block is large, the first preset block can be down-sampled, such as one sample in every 2, or 4, or 8 samples in the horizontal direction and / or the vertical direction. Alternatively, if the size of the first preset block in one of the horizontal or vertical directions is less than or equal to 8, all available samples in that direction are counted; otherwise, if the size of the first preset block in one of the horizontal or vertical directions is less than or equal to 16, one sample in every 2 samples in that direction is counted; otherwise, one sample in every 4 samples in that direction is counted, which is not limited here.
[0259] In the embodiments of the present application, taking the derivation of the first intra prediction mode from the candidate samples of the first preset block as an example, accordingly, in some embodiments, the first intra prediction mode derived based on the candidate samples of the first preset block can be determined, and the method can include: determining the horizontal gradient value and the vertical gradient value of the candidate sample; determining the texture feature index and the gradient intensity value corresponding to the candidate sample according to the horizontal gradient value and the vertical gradient value of the candidate sample; determining the texture feature statistics table according to the texture feature index and the gradient intensity value corresponding to the candidate sample; and determining the first intra prediction mode derived by the first preset block according to the texture feature statistics table.
[0260] It should be noted that, in the embodiments of the present application, when the texture feature index and the gradient intensity value corresponding to the candidate sample are determined according to the horizontal gradient value and the vertical gradient value of the candidate sample, it can include: performing angle mapping according to the horizontal gradient value and the vertical gradient value of the candidate sample to determine the texture feature index corresponding to the candidate sample; and performing gradient intensity calculation according to the horizontal gradient value and the vertical gradient value of the candidate sample to determine the gradient intensity value corresponding to the candidate sample.
[0261] In a specific embodiment, performing angle mapping according to the horizontal gradient value and the vertical gradient value of the candidate sample to determine the texture feature index corresponding to the candidate sample can include: determining the texture feature index corresponding to the candidate sample by using a preset lookup table according to the horizontal gradient value and the vertical gradient value of the candidate sample.
[0262] In the embodiments of the present application, the horizontal gradient value of the candidate sample can be denoted as grad x , and the vertical gradient value of the candidate sample can be denoted as grad y . In this way, the texture feature index (or referred to as "virtual intra prediction mode") is derived from grad x and grad y , which can be implemented by a lookup table.
[0263] Exemplarily, if abs(grad x ) is equal to 0 and abs(grad y ) is not equal to 0, there is horizontal direction texture, corresponding to the intra prediction mode 18 in VVC. If abs(grad y ) is equal to 0 and abs(grad x ) is not equal to 0, there is vertical direction texture, corresponding to the intra prediction mode 50 in VVC. In the case that abs(grad x ) and abs(grad y ) are not equal to 0, if abs(grad x ) is equal to abs(grad y ) and grad x and grad y have the same sign, it corresponds to the intra prediction mode 34 in VVC. If abs(grad x ) is equal to 2 times abs(grad y ) and grad x and grad y have the same sign, it corresponds to the intra prediction mode 40 in VVC. In addition, other cases can be determined according to the same principle of lookup table.
[0264] In a specific embodiment, the gradient strength calculation according to the horizontal gradient value and the vertical gradient value of the candidate sample to determine the gradient strength value corresponding to the candidate sample can include: performing an addition operation on the absolute value of the horizontal gradient value and the absolute value of the vertical gradient value to determine the gradient strength value corresponding to the candidate sample.
[0265] Here, the gradient strength value corresponding to the candidate sample can be denoted as amp. Exemplarily, amp = abs(grad x ) + abs(grad y ).
[0266] It should be noted that in the embodiments of the present application, the horizontal gradient value and the vertical gradient value of the candidate sample can be calculated using the Sobel operator. Exemplarily, for the Sobel operator, the specific formula is as follows:
[0267] Operator of horizontal gradient value:
[0268] Operator of vertical gradient value:
[0269] Thus, assuming that the sample value of the sample with sample position (x, y) of the reconstructed block or the predicted block is P x,y , the horizontal gradient value grad x and the vertical gradient value grad y are calculated as shown below: grad x = P x+1,y-1 + 2*P x+1,y + P x+1,y+1 - P x-1,y-1 - 2*P x-1,y - P x-1,y+1 (7) grad y = P x-1,y+1 + 2*P x,y+1 + P x+1,y+1 - P x-1,y-1 - 2*P x,y-1 - P x+1,y-1 (8)
[0270] It should also be noted that in the embodiments of the present application, considering that the Sobel operator uses the samples above, below, left and right of the current sample, the embodiments of the present application can be configured to calculate the gradients of all the samples in the reconstructed block or the predicted block except the samples in the outermost row and column.
[0271] In some embodiments, the determining the texture feature statistics table according to the texture feature index corresponding to the candidate sample and the gradient strength value can include: when the number of candidate samples is at least one, determining at least one texture feature index and at least one gradient strength value corresponding thereto; determining at least one reference texture feature index with different characteristics according to the at least one texture feature index, and performing accumulation calculation on the gradient strength values belonging to the same reference texture feature index according to the at least one gradient strength value, to determine the gradient strength accumulation value corresponding to the at least one reference texture feature index; and determining the texture feature statistics table according to the at least one reference texture feature index and the gradient strength accumulation value corresponding to the at least one reference texture feature index.
[0272] That is, in the embodiments of the present application, taking at least part of the samples in the reconstructed block or the prediction block as the candidate samples as an example, the gradient of all or part of the samples in the reconstructed block or the prediction block is calculated, and generally, the horizontal gradient value and the vertical gradient value can be calculated, and here the Sobel operator can be used to calculate the gradient value. For a certain sample, the texture direction of the sample can be inferred according to the horizontal gradient value and the vertical gradient value, for example, if the horizontal gradient value is non-zero and the vertical gradient value is zero, the texture of the sample is in the vertical direction. Conversely, if the horizontal gradient value is zero and the vertical gradient value is non-zero, the texture of the sample is in the horizontal direction. For example, if the horizontal gradient value and the vertical gradient value are equal and non-zero, the texture of the sample is 45 degrees. Of course, in many other cases of the embodiments of the present application, the horizontal gradient value and the vertical gradient value are both non-zero, and the direction of the texture of the sample can be determined according to their ratio. In this way, the gradient strength value of each sample can be corresponded to the corresponding texture feature index according to the gradient strength value of each sample. Thus, a texture feature statistics table is constructed, and the gradient strength value of each calculated sample is accumulated in the corresponding texture feature index item in the statistics table to obtain the final texture feature statistics table, and then the first intra prediction mode derived by the first preset block can be determined according to the texture feature statistics table.
[0273] In some embodiments, according to the texture feature statistics table, the first intra prediction mode derived by the first preset block is determined, and the method can include: determining the reference texture feature index corresponding to the highest gradient strength accumulation value in the texture feature statistics table as the first intra prediction mode derived by the first preset block.
[0274] That is, in the embodiments of the present application, when the gradient statistics is performed according to the candidate samples, it is assumed that there are 67 intra prediction modes here, and then the texture feature statistics table can also be set as an array, for example, array A mpAccumulate
[0067] , and each item of A mpAccumulate
[0067] is initialized to 0. Taking the first preset block as an example, if the index (or referred to as “reference texture feature index”) of the intra prediction mode derived by a certain sample position is X, the amp calculated by the sample position is accumulated on A mpAccumulate[X], and the mode with the highest value in the array A mpAccumulate
[0067] is selected as the first intra prediction mode derived by the gradient statistics of the reconstructed block or the prediction block of the block. It should be noted that the first intra prediction mode derived here belongs to one of the candidate intra prediction modes in the first mode set.
[0275] It can also be understood that, in the embodiments of the present application, the process of deriving one of the candidate intra prediction modes in the first mode set by using gradient statistics for the reconstructed block or the prediction block of the block can be performed when the block is decoded, and the derived candidate intra prediction mode is saved, so that the derivation of gradient statistics does not need to be performed when the current block is used, thereby saving the amount of calculation.
[0276] It can also be understood that, in the embodiments of the present application, when it is determined that the intra prediction mode used or derived by the first preset block is a candidate intra prediction mode in the first mode set, the intra prediction mode can be subjected to accumulation operation at this time.
[0277] In a possible implementation, the accumulation operation on the first intra prediction mode to determine the statistical result of the first intra prediction mode in the first mode set can include: performing the accumulation operation on the first intra prediction mode according to the size parameter value of the first preset block to determine the statistical result of the first intra prediction mode in the first mode set.
[0278] In the embodiments of the present application, taking the first preset block as an example, assuming that the width of the first preset block is width and the height of the first preset block is height, the size parameter value of the first preset block is width*height, and if the accumulation operation at this time can consider the size of the preset block, i.e., occurrence[X]=occurrence[X]+width*height. The first mode set can be represented by an array occurrence[L], L is the number of modes in the first mode set, i.e., the length of the array; and X is the mode index of the first intra prediction mode in the first mode set.
[0279] In another possible implementation, the accumulation operation on the first intra prediction mode to determine the statistical result of the first intra prediction mode in the first mode set can include: performing the accumulation operation of adding one on the first intra prediction mode to determine the statistical result of the first intra prediction mode in the first mode set.
[0280] In the embodiments of the present application, the accumulation operation can also not consider the size of the preset block, and at this time, the accumulation operation can also be the accumulation operation of adding one, i.e., occurrence[X]=occurrence[X]+1, wherein X represents the mode index of the first intra prediction mode in the first mode set. That is, after the first intra prediction mode used or derived by the first preset block is determined, if the first intra prediction mode is a candidate intra prediction mode in the first mode set, the accumulation value on the first intra prediction mode can be subjected to the accumulation operation of adding one to obtain the statistical result of the first intra prediction mode in the first mode set.
[0281] Thus, after the above operation process, after traversing one or more preset blocks, occurrence
[0067] has been statistically completed, that is, the statistical results of at least two candidate intra-frame prediction modes in the first mode set can be obtained.
[0282] S2803, Based on the statistical results, determine at least two intra-prediction modes for the current block in the first mode set.
[0283] In this embodiment of the application, based on the obtained statistical results, at least two intra-frame prediction modes of the current block can be determined from the first mode set.
[0284] In some embodiments, determining at least two intra-prediction modes of the current block in the first mode set based on statistical results may include: sorting the candidate intra-prediction modes in the first mode set in descending order according to the statistical results, and determining the at least two candidate intra-prediction modes with the highest ranking as at least two intra-prediction modes of the current block.
[0285] In this embodiment of the application, based on the statistical results of occurrence
[0067] , at least two intra-prediction modes of the current block can be determined from the first mode set. For example, the top two candidate intra-prediction modes with the largest values in occurrence
[0067] can be selected as the at least two intra-prediction modes of the current block.
[0286] For example, if two intra-prediction modes for the current block are determined, then the candidate intra-prediction mode M1 with the largest value in occurrence
[0067] and the candidate intra-prediction mode M2 with the second largest value can be used as the two intra-prediction modes for the current block.
[0287] S2804, predict the current block based on at least two intra-frame prediction modes to determine the predicted value of the current block.
[0288] In some embodiments, predicting the current block based on at least two intra-frame prediction modes to determine the predicted value of the current block may include: determining the weight values of each of the at least two intra-frame prediction modes based on the statistical results corresponding to the at least two intra-frame prediction modes; and performing weighted prediction on the current block based on the at least two intra-frame prediction modes and their respective weight values to determine the predicted value of the current block.
[0289] In this embodiment of the application, based on the statistical results, not only can at least two intra-prediction modes of the current block be determined from the first mode set, but also the weight values of each of these at least two intra-prediction modes can be determined. Specifically, the weight values of each of the at least two intra-prediction modes can be determined based on the statistical results corresponding to these at least two intra-prediction modes.
[0290] For example, it is assumed that two intra prediction modes of the current block can be derived according to the statistical result, specifically, the intra prediction mode M1 and the intra prediction mode M2. When performing the weighted prediction on the current block, the weighted prediction can be performed according to the intra prediction mode M1, the intra prediction mode M2 and the PLANAR mode, and the respective weights of the intra prediction mode M1, the intra prediction mode M2 and the PLANAR mode are W1, W2 and W3 respectively. The specific calculation formula is as follows:
[0291] Wherein, occurrence[M1] represents the cumulative value corresponding to the intra prediction mode M1, and occurrence[M2] represents the cumulative value corresponding to the intra prediction mode M2.
[0292] In addition, when the current block uses the OBIP mode, it is assumed that the prediction value obtained by predicting the current block using the intra prediction mode M1 is Pred1, the prediction value obtained by predicting the current block using the intra prediction mode M2 is Pred2, and the prediction value obtained by predicting the current block using the PLANAR mode is Pred3; then the prediction value obtained by predicting using the OBIP mode is Pred OBIP , and the specific calculation formula is as follows:
[0293] In the embodiment of the present application, the intra prediction modes of two first mode sets and the PLANAR mode are used for weighted prediction. It can be understood that more intra prediction modes of the first mode set can be used here, for example, 3, 4 or 5. In addition, the PLANAR mode here can be replaced by other modes, for example, the ITMP mode, which is not limited here.
[0294] It can also be understood that in the embodiment of the present application, whether the current block uses the first prediction mode can be indicated by the first syntax element. Referring to FIG. 33, for step S2801, the method can further include:
[0295] S3301, decoding the code stream to determine the value of the first syntax element.
[0296] S3302, when the first syntax element indicates that the current block uses the first prediction mode, determining one or more preset blocks that have been decoded.
[0297] In the embodiments of the present application, after step S3302, step S2802-S2804 is executed. That is, when the first syntax element indicates that the current block uses the first prediction mode, one or more preset blocks are determined; then the steps of determining the statistical result of at least two candidate intra prediction modes in the first mode set according to the prediction mode used by each of the one or more preset blocks, determining at least two intra prediction modes of the current block in the first mode set according to the statistical result, and predicting the current block according to the at least two intra prediction modes to determine the prediction value of the current block are executed.
[0298] In the embodiments of the present application, the first syntax element is used to indicate whether the current block uses the first prediction mode. Wherein, the method can further include: if the first syntax element takes the first value, the first syntax element indicates that the current block uses the first prediction mode; if the first syntax element takes the second value, the first syntax element indicates that the current block does not use the first prediction mode.
[0299] In the embodiments of the present application, the first syntax element can be a block-level syntax element, and the first value is different from the second value. Wherein, the first value can be 1, and the second value can be 0; or the first value can be true, and the second value can be false; or the first value can be 0, and the second value can be 1; or the first value can be false, and the second value can be true, which is not limited here.
[0300] In addition, in the embodiments of the present application, the first prediction mode can be the OBIP mode, and the first syntax element can be represented by cu_obip_flag. That is, in the embodiments of the present application, a block-level (such as CU or PU level) syntax element can be set to indicate whether the current block uses the OBIP mode. Specifically, if the value of cu_obip_flag is 1, it indicates that the current block uses the OBIP mode for prediction; if the value of cu_obip_flag is 0, it indicates that the current block does not use the OBIP mode for prediction. If cu_obip_flag does not appear, it can be inferred that its value is 0.
[0301] That is, in the embodiments of the present application, if the value of cu_obip_flag obtained by decoding is 1, it indicates that the first syntax element indicates that the current block uses the first prediction mode, at this time one or more preset blocks can be determined to facilitate the determination of the statistical result of at least two candidate intra prediction modes in the first mode set, and then at least two intra prediction modes used by the current block are determined.
[0302] It can also be understood that the decoding method in the embodiments of the present application can also be applied to the Multiple Transform Set Selection (MTSS) mode. Among them, NSPT and LFNST are transforms for processing various angle textures, and they can have multiple transform cores, and a transform core can be optimized for a certain specific angle texture. Of course, in addition to the angle texture, NSPT and LFNST also include transform cores for processing gradient texture. In fact, these transform cores can also be said to be trained KL transform (Karhunen-Loeve Transform, KLT). That is, NSPT and LFNST have multiple transform cores, each of which is designed for a specific texture, including angle texture, gradient texture, and the like. In addition, gradient texture can be further extended, such as horizontal gradient texture, vertical gradient texture, and diagonal gradient texture. Further, the MTSS mode is not limited to being used for non-separable transforms such as NSPT and LFNST, and can also be applied to separable transforms optimized for specific textures.
[0303] It should be further noted that in the embodiments of the present application, the current block can have multiple transform core groups, and each intra prediction mode can correspond to a transform core group. In other words, each intra prediction mode actually represents a texture feature. Therefore, the intra prediction mode index is also a texture feature index. For example, DC and PLANAR correspond to gradient texture features, and a certain angle prediction mode corresponds to an angle texture feature. The texture feature index can avoid the appearance of intra prediction mode in "inter" on the one hand, and it is also more conducive to possible expansion on the other hand, for example, one intra prediction mode can correspond to multiple texture features, such as DC mode can correspond to horizontal gradient texture, vertical gradient texture, and diagonal gradient texture.
[0304] Specifically, in the embodiments of the present application, it is considered that some intra prediction modes are not simple texture features, but can contain two or more texture features. Therefore, the MTSS technology can be used here. In the MTSS technology, if the intra prediction mode of the current block is a certain special intra prediction mode, there is not only one optional transform core group, for example, the first transform core group corresponding to the intra prediction mode M1 and the second transform core group corresponding to the intra prediction mode M2.
[0305] In some embodiments, for the transform process of the current block, referring to FIG. 34, the method can include:
[0306] S3401, determining the transform core of the current block.
[0307] In a possible implementation, the determining the transform core of the current block can include: determining a transform core group index of the current block; determining a transform core group of the current block according to the transform core group index; and determining the transform core of the current block according to the transform core group.
[0308] It should be noted that in the embodiments of the present application, the transform core group index is used to represent the number of the transform core group of the current block in the at least two candidate transform core groups, and the transform core group index can be represented by lfnst_nspt_set_index. Wherein, the transform core group index of the current block can be an integer greater than or equal to zero, such as 0, 1, 2, etc. For example, if the value of lfnst_nspt_set_index is equal to 0, it indicates that the first candidate transform core group is selected as the transform core group of the current block; if the value of lfnst_nspt_set_index is equal to 1, it indicates that the second candidate transform core group is selected as the transform core group of the current block.
[0309] In some embodiments, for determining the transform core group index of the current block, the method can include: decoding the code stream to determine the transform core group index of the current block; or the method can also include: decoding the code stream to determine the value of the second syntax element; and determining the transform core group index of the current block according to the value of the second syntax element.
[0310] That is, in the embodiments of the present application, the transform core group index of the current block can be directly determined by decoding the code stream, or it can also be determined by decoding the value of the second syntax element.
[0311] It should be further noted that in the embodiments of the present application, the second syntax element can be represented by lfnst_nspt_set_index. The value of the second syntax element can be used to indicate the transform core group index of the current block, specifically the number of the transform core group of the current block in the at least two candidate transform core groups. Wherein, the value of the second syntax element can be an integer greater than or equal to zero, such as 0, 1, 2, etc.
[0312] Here, for the OBIP mode, because it uses multiple intra prediction modes weighting, for example, it uses two intra prediction modes (M1, M2) in the first mode set and the PLANAR mode for weighting. Then its prediction value has both the characteristics of M1 and M2, and its residual may have the characteristics of M1 or M2, of course, it may also have other characteristics, but the residual and M1 and M2 have correlation. Therefore, for the OBIP mode, when determining the transform kernel determined according to the prediction mode such as LFNST or NSPT, an option can be added. Exemplarily, a syntax element is used to indicate whether the transform kernel (group) is determined according to M1 or M2. For example, a second syntax element lfnst_nspt_set_index is set. If the value of lfnst_nspt_set_index is 0, the first candidate transform kernel group is used, and if the value of lfnst_nspt_set_index is 1, the second candidate transform kernel group is used. Wherein, the first candidate transform kernel group is determined by the first intra prediction mode (M1), and the second candidate transform kernel group is determined by the second intra prediction mode (M2). Optionally, for special cases, for example, the transform kernel group determined by the first intra prediction mode (M1) and the second intra prediction mode (M2) is exactly the same, then other intra prediction modes can be determined, for example, the third intra prediction mode participating in prediction or the PLANAR mode, etc.
[0313] It can also be understood that in the embodiments of the present application, the decoding end can also construct the first candidate list. In some embodiments, the method can further include: determining the first candidate list of the current block, the first candidate list indicating at least two candidate transform kernel groups; determining the transform kernel group of the current block according to the first candidate list and the transform kernel group index.
[0314] It should be noted that in the embodiments of the present application, when the current block uses the OBIP mode, the first candidate list can include at least two candidate texture feature indexes, or the first candidate list can include at least two candidate transform kernel groups. Here, each candidate texture feature index corresponds to a candidate transform kernel group. Therefore, it can be said that: the first candidate list indicates at least two candidate transform kernel groups.
[0315] Exemplarily, the transform core group determined by the first intra prediction mode (M1) can be taken as the first candidate transform core group of the first candidate list. If the transform core group determined by the second intra prediction mode (M2) is different from the first candidate transform core group, the transform core group determined by the second intra prediction mode (M2) is taken as the second candidate transform core group of the first candidate list; otherwise, if the transform core group determined by the second intra prediction mode (M2) is the same as the first candidate transform core group, the transform core groups determined by the third intra prediction mode, the fourth intra prediction mode and the like used by the current block are checked in sequence to see whether they are the same as the first candidate transform core group of the first candidate list. If they are different, the transform core group is taken as the second candidate transform core group of the first candidate list. If all the intra prediction modes used by the current block are checked and the second candidate transform core group is still not determined, a default transform core group can be taken as the second candidate transform core group, for example, the transform core group determined by the PLANAR mode is taken as the second candidate transform core group.
[0316] In another possible implementation, after the transform core of the current block is determined, the determining the transform core of the current block can include: determining a transform core index of the current block; and determining the transform core of the current block according to the transform core group and the transform core index.
[0317] In another possible implementation, after the transform core of the current block is determined, the determining the transform core of the current block can include: determining a transform core index of the current block; and determining the transform core of the current block according to the transform core group and the transform core index.
[0318] In the embodiments of the present application, for the first candidate list, assuming that the first candidate list indicates the transform cores included in two candidate transform core groups, if the first candidate transform core group includes 3 candidate transform cores and the second candidate transform core group includes 2 candidate transform cores, the first candidate list can indicate 5 candidate transform cores; if the first candidate transform core group includes 3 candidate transform cores and the second candidate transform core group includes 3 candidate transform cores, the first candidate list can indicate 6 candidate transform cores. In this case, after the transform core index of the current block is determined, the transform core of the current block can be determined in the first candidate list according to the transform core index.
[0319] It also needs to be explained that in the embodiments of the present application, the transform core index of the current block can be a positive integer, for example, 1, 2, 3, 4, 5, 6 and the like. The transform core index of the current block can be determined directly by decoding the code stream, or can be determined by decoding the value of the third syntax element.
[0320] Exemplarily, one possible implementation manner is: decoding the bitstream, determining the transform core index of the current block. Or, another possible implementation manner is: decoding the bitstream, determining the value of the third syntax element; when the third syntax element indicates that the first transform mode is used for the current block, determining the transform core index of the current block according to the value of the third syntax element.
[0321] It also needs to be explained that, in the embodiment of the present application, the first transform mode can be LFNST / NSPT, and the third syntax element can be represented by lfnst_nspt_index. Wherein, the third syntax element can be used to indicate whether the first transform mode is used for the current block, and the corresponding transform core index when the first transform mode is used for the current block.
[0322] It also needs to be explained that, in the embodiment of the present application, if the value of the third syntax element is the third value, it is determined that the first transform mode is not used for the current block; if the value of the third syntax element is the fourth value, it is determined that the first transform mode is used for the current block and the corresponding transform core index. Wherein, the third value can be set to 0, and the fourth value can be set to non-0, such as 1, 2, 3, 4, 5, 6 and so on.
[0323] That is to say, in the embodiment of the present application, for LFNST / NSPT, the transform core index of the current block can also be represented by lfnst_nspt_index. Wherein, lfnst_nspt_index=0 indicates that the current block does not use LFNST / NSPT, and each transform core group of LFNST / NSPT in the ECM has 3 transform cores, so the value of lfnst_nspt_index is 1 or 2 or 3, which indicates that the first transform core or the second transform core or the third transform core of the selected transform core group of LFNST / NSPT is used for the current block.
[0324] In the embodiment of the present application, the current block uses the OBIP mode, at this time it has more than one optional transform core group. Exemplarily, it has two optional transform core groups, so the possible values of lfnst_nspt_index are 0, 1, 2, 3, 4, 5, 6. Wherein 1, 2, 3 correspond to 3 transform cores of the first candidate transform core group, and 4, 5, 6 correspond to 3 transform cores of the second candidate transform core group.
[0325] Exemplarily, at this time, a syntax element can not be added, but a possible value of an original syntax element can be added. For example, lfnst_nspt_index is used to indicate a transform kernel of LFNST or NSPT. It is known that there are 3 transform kernels in each transform kernel group of LFNST / NSPT. For a general prediction mode, possible values of lfnst_nspt_index are 0, 1, 2, and 3. Among them, if the value of lfnst_nspt_index is 0, it indicates that LFNST or NSPT is not used; if the value of lfnst_nspt_index is 1, it indicates that the first transform kernel in the selected transform kernel group is used; if the value of lfnst_nspt_index is 2, it indicates that the second transform kernel in the selected transform kernel group is used; and if the value of lfnst_nspt_index is 3, it indicates that the third transform kernel in the selected transform kernel group is used. For the OBIP mode, possible values of lfnst_nspt_index are 0, 1, 2, 3, 4, 5, and 6. Among them, if the value of lfnst_nspt_index is 0, it indicates that LFNST or NSPT is not used; if the value of lfnst_nspt_index is 1, it indicates that the first transform kernel in the first candidate transform kernel group is used; if the value of lfnst_nspt_index is 2, it indicates that the second transform kernel in the first candidate transform kernel group is used; if the value of lfnst_nspt_index is 3, it indicates that the third transform kernel in the first candidate transform kernel group is used; if the value of lfnst_nspt_index is 4, it indicates that the first transform kernel in the second candidate transform kernel group is used; if the value of lfnst_nspt_index is 5, it indicates that the second transform kernel in the second candidate transform kernel group is used; and if the value of lfnst_nspt_index is 6, it indicates that the third transform kernel in the second candidate transform kernel group is used. In this way, the transform kernel of the current block can be determined.
[0326] In S3402, the transform coefficient of the current block is transformed according to the transform kernel, and a residual block of the current block is determined.
[0327] It should be noted that in the embodiments of the present application, the method can further include: decoding the code stream to determine the quantized coefficient of the current block; and dequantizing the quantized coefficient of the current block to determine the transform coefficient of the current block.
[0328] It should be noted that, in the embodiments of the present application, when the transform coefficients of the current block are transformed according to the transform kernel, and the residual block of the current block is determined, it can include: performing non-separable basis transform on the transform coefficients of the current block according to the transform kernel, and determining the residual block of the current block; or performing low-frequency non-separable transform on the transform coefficients of the current block according to the transform kernel, determining the transform block of the current block, and performing discrete cosine transform on the transform block of the current block, and determining the residual block of the current block.
[0329] In a specific embodiment, if the size parameter of the current block satisfies the first condition, the non-separable basis transform is performed on the transform coefficients of the current block according to the transform kernel, and the residual block of the current block is determined; if the size parameter of the current block satisfies the second condition, the low-frequency non-separable transform is performed on the transform coefficients of the current block according to the transform kernel, the transform block of the current block is determined, and the discrete cosine transform is performed on the transform block of the current block, and the residual block of the current block is determined.
[0330] Here, the size parameter of the current block satisfying the first condition includes that the size parameter of the current block is small, for example, the size parameter of the current block is less than a certain threshold. That is, for a block with a smaller size, the transform kernel of NSPT is used here, that is, the inverse NSPT transform is performed on the transform coefficients of the current block according to the transform kernel, and the residual block of the current block is determined.
[0331] Here, the size parameter of the current block satisfying the second condition includes that the size parameter of the current block is large, for example, the size parameter of the current block is greater than a certain threshold. That is, for a block with a larger size, the transform kernel of LFNST is used here, that is, the inverse LFNST transform is performed on the transform coefficients of the current block according to the transform kernel, and the transform block of the current block is determined; and the inverse DCT2 transform is performed on the transform block of the current block, and the residual block of the current block is determined.
[0332] It should be noted that, in the embodiments of the present application, the "inverse transform" of the transform coefficients at the decoding end can also be referred to as "transform" in the standard text. The "transform" and "inverse transform" in this paper correspond to two opposite processes. For example, "transform" converts the values in the spatial domain to the coefficients in the frequency domain, and "inverse transform" converts the coefficients in the frequency domain to the values in the spatial domain. "Inverse" is relative to "positive", and they are essentially both transforms. It should be noted that if the standard only specifies decoding, the "transform" in the standard text is the part of decoding, which specifically refers to the "inverse transform" in this paper. The "inverse transform" of the transform coefficients at the decoding end can also be referred to as "transform" in the standard text.
[0333] S3403, determining the reconstructed value of the current block according to the residual value of the current block and the prediction value of the current block.
[0334] In the embodiments of the present application, after determining the residual value of the current block, the prediction value of the current block and the residual value of the current block can be added to determine the reconstructed value of the current block.
[0335] For example, after the decoding end obtains the quantization coefficients from the code stream by entropy decoding, if it is NSPT transform, the quantization coefficients are dequantized to obtain decoded transform coefficients, the decoded transform coefficients are subjected to inverse NSPT transform to obtain a decoded residual block, and finally a reconstructed block is obtained according to the decoded residual block and a prediction block. Conversely, if it is LFNST transform, the quantization coefficients are dequantized to obtain decoded transform coefficients, the decoded transform coefficients are subjected to inverse LFNST transform, and then inverse DCT2 transform is performed to obtain a decoded residual block, and finally a reconstructed block is obtained according to the decoded residual block and a prediction block.
[0336] The embodiments of the present application provide a decoding method, specifically, an intra prediction mode and a transform scheme based on occurrence statistics. First, one or more preset blocks that have been decoded are determined; according to the prediction modes used by the one or more preset blocks respectively, statistical results of at least two candidate intra prediction modes in a first mode set are determined; then, according to the statistical results, at least two intra prediction modes of a current block are determined in the first mode set; and finally, the current block is predicted according to the at least two intra prediction modes to determine a prediction value of the current block. In this way, the statistical results of the related blocks that have been decoded are obtained by using the related blocks that have been decoded, and the at least two intra prediction modes used by the current block are derived according to the statistical results. Since it is not necessary to explicitly indicate which kind of intra prediction mode is used, the overhead in the code stream can be saved. Moreover, the accuracy of prediction can be improved by using the weighted prediction of the at least two intra prediction modes, and the prediction error can be reduced to a certain extent. In addition, the at least two intra prediction modes used by the current block are derived according to the statistical results, which can avoid the intra prediction mode derivation by gradient calculation required by the DIMD mode, thereby reducing the calculation complexity, improving the coding efficiency, and further improving the compression performance.
[0337] In another embodiment of the present application, FIG. 35 is a flowchart of an encoding method according to an embodiment of the present application. As shown in FIG. 35, the method can include the following steps.
[0338] S3501, determining one or more preset blocks that have been encoded.
[0339] In the embodiments of the present application, the method is applied to an encoder. Specifically, based on the constituent structure of the encoder 100 shown in FIG. 26, the encoding method is mainly applied to the intra prediction unit 106, the transformation unit 108 and the inverse transformation unit 111 in FIG. 26. Wherein, when the current block uses the OBIP mode, the intra prediction mode used by the current block can be derived by using the prediction mode used by the related block in the encoded area. In this way, compared with the TIMD mode and the DIMD mode, the OBIP mode can save the calculation amount and improve the compression performance.
[0340] In the embodiments of the present application, it is first needed to determine whether the current block uses the first prediction mode. In some embodiments, the method can include: determining a plurality of candidate prediction modes, wherein the plurality of candidate prediction modes at least include the first prediction mode; performing encoding cost calculation on the current block according to the plurality of candidate prediction modes to determine a cost result corresponding to each of the plurality of candidate prediction modes; determining a minimum cost result from the cost results corresponding to the plurality of candidate prediction modes; and determining whether the current block uses the first prediction mode according to the candidate prediction mode corresponding to the minimum cost result.
[0341] In the embodiments of the present application, the cost calculation here can be determined according to the cost result of the rate distortion optimization (RDO), or can be determined according to the cost result of the sum of absolute difference (SAD), or can be determined according to the cost result of the sum of absolute transformed difference (SATD), but here is not limited in any way.
[0342] In the embodiments of the present application, for determining whether the current block uses the first prediction mode, it can include: when the candidate prediction mode corresponding to the minimum cost result is the first prediction mode, determining that the current block uses the first prediction mode; and when the candidate prediction mode corresponding to the minimum cost result is a non-first prediction mode, determining that the current block does not use the first prediction mode.
[0343] Further, in some embodiments, the method can further include: determining a value of the first syntax element; and performing encoding processing on the value of the first syntax element and writing the obtained encoding bits into a bitstream.
[0344] It should be noted that in the embodiments of the present application, the first syntax element is used to indicate whether the current block uses the first prediction mode. Wherein, the value of the first syntax element can be determined as follows: when the candidate prediction mode corresponding to the minimum cost result is the first prediction mode, the value of the first syntax element is determined as the first value; when the candidate prediction mode corresponding to the minimum cost result is not the first prediction mode, the value of the first syntax element is determined as the second value. That is, if the current block uses the first prediction mode, the value of the first syntax element is determined as the first value; if the current block does not use the first prediction mode, the value of the first syntax element is determined as the second value.
[0345] It should be further noted that in the embodiments of the present application, the first syntax element can be a block-level syntax element, and the first value is different from the second value. Wherein, the first value can be 1, and the second value can be 0; or the first value can be true, and the second value can be false; or the first value can be 0, and the second value can be 1; or the first value can be false, and the second value can be true, which is not limited here.
[0346] In addition, in the embodiments of the present application, the first prediction mode can be the OBIP mode, and the first syntax element can be represented by cu_obip_flag. That is, in the embodiments of the present application, a block-level (such as CU or PU level) syntax element can be set to indicate whether the current block uses the OBIP mode. Specifically, if the current block uses the OBIP mode for prediction, the value of cu_obip_flag can be determined as 1; if the current block does not use the OBIP mode for prediction, the value of cu_obip_flag can be determined as 0. If cu_obip_flag does not appear in the code stream, it can also be inferred that the value of cu_obip_flag is 0.
[0347] Further, in some embodiments, the method can further include: when the current block uses the first prediction mode, performing the step of determining the coded first or more preset blocks.
[0348] That is, in the embodiments of the present application, if the current block uses the first prediction mode, the coded one or more preset blocks can be determined at this time, so as to facilitate the statistics of the occurrence of the at least two candidate intra-frame prediction modes in the first mode set in the one or more preset blocks.
[0349] In a possible implementation, determining the coded one or more preset blocks can include: determining at least one candidate block that has been coded, and determining the one or more preset blocks according to the at least one candidate block.
[0350] In the embodiments of the present application, the candidate blocks can include the neighboring blocks of the current block and / or the non-neighboring blocks of the current block. That is, the one or more preset blocks can include only the neighboring blocks of the current block, or can include only the non-neighboring blocks of the current block, or can include both the neighboring blocks and the non-neighboring blocks of the current block.
[0351] It can be understood that when the OBIP mode is used for the current block, the neighboring blocks and the non-neighboring blocks of the current block can be counted. As shown in FIG. 29, the black filled block is the current block, the blocks labeled 1 to 7 are the neighboring blocks of the current block, and the blocks labeled 8 to 25 are the non-neighboring blocks of the current block. The neighboring blocks refer to the blocks spatially adjacent to the current block, and the non-neighboring blocks refer to the blocks spatially non-adjacent to the current block.
[0352] For example, it is assumed that the top-left position of the current block is (x, y), the width of the current block is width, and the height of the current block is height. Then for the neighboring blocks of the current block, the neighboring block 1 is the block containing the coordinate (x-1, y-1), the neighboring block 2 is the block containing the coordinate (x, y-1), the neighboring block 2 is the block containing the coordinate (x, y-1), the neighboring block 3 is the block containing the coordinate (x-1, y), the neighboring block 4 is the block containing the coordinate (x+width-1, y-1), the neighboring block 5 is the block containing the coordinate (x+width, y-1), the neighboring block 6 is the block containing the coordinate (x-1, y+height-1), and the neighboring block 7 is the block containing the coordinate (x-1, y+height).
[0353] The non-neighboring blocks (i.e., the non-neighboring blocks) can also be positioned according to certain rules. For example, the positions of the non-neighboring blocks can be calculated according to (x, y), width and height. In FIG. 29, the distance of each grid in the vertical direction is height, and the distance of each grid in the horizontal direction is width. Then for the non-neighboring blocks of the current block, the non-neighboring block 8 is the block containing the coordinate (x-width-1, y-height-1), the non-neighboring block 9 is the block containing the coordinate (x+2*width-1, y-height-1), the non-neighboring block 10 is the block containing the coordinate (x-width-1, y+2*height-1), the non-neighboring block 11 is the block containing the coordinate (x-2*width-1, y-2*height-1), the non-neighboring block 12 is the block containing the coordinate (x+width / 2, y-2*height-1), and so on. It should be noted that the block containing a certain coordinate herein refers to the CU or PU containing the coordinate.
[0354] Thus, in the embodiments of the present application, for the coded candidate block, only the neighboring blocks can be used, or only the non-neighboring blocks can be used, or both the neighboring blocks and the non-neighboring blocks can be used, or only specified blocks in the neighboring blocks and the non-neighboring blocks can be used, and no limitation is made herein.
[0355] In another possible implementation, determining the one or more preset blocks coded can include: determining a preset range based on the current block; and determining the one or more preset blocks coded based on all the blocks in the preset range.
[0356] In the embodiments of the present application, the preset range can be any set range coded. For example, assuming that the top-left position of the current block is (x, y), the width of the current block is width, and the height of the current block is height. Then, the preset range can be set according to the position of the current block and the width and the height.
[0357] It can also be understood that when the current block uses the OBIP mode, all the blocks in the preset range can be counted. As shown in FIG. 30, the black filled block represents the current block, and the dot filled part represents the preset range. For example, one possible preset range is all the coded blocks in the rectangular region with the top-left corner at (x-2*width, y-2*height) and the right-top corner at (x+3*width-1, y-2*height) and the left-bottom corner at (x-2*width, y+3*height-1). If the minimum size of the CU or PU is 4*4, all the 4*4 blocks in the search range can be counted, for example, the first 4*4 block in the top-left corner is the block containing the coordinate (x-2*width, y-2*height), and the 4*4 blocks in the same row can be scanned by adding 4 horizontally, and the 4*4 blocks in the same column can be scanned by adding 4 vertically. It should be noted that the block containing the coordinate mentioned herein refers to the CU or PU containing the coordinate.
[0358] Thus, in the embodiments of the present application, after the preset range coded is determined, all the blocks found in the preset range can be determined as the one or more preset blocks.
[0359] S3502, determining a statistical result of at least two candidate intra prediction modes in the first mode set according to the prediction modes used by the one or more preset blocks respectively.
[0360] In the embodiments of the present application, the first mode set can include at least two candidate intra prediction modes. The intra prediction mode can include many modes, such as the PLANAR mode, the DC mode and the 65 angle prediction modes, the MIP mode, etc. which already exist in VVC. The DIMD mode, the TIMD mode, the SGPM mode, the ITMP mode and the OBIP mode, etc. which are newly added in the ECM. Here, a first mode set can be defined, and in some embodiments, the first mode set includes at least one of the following: angle prediction mode, DC mode and PLANAR mode.
[0361] In the embodiments of the present application, the first mode set includes at least two candidate intra prediction modes. Exemplarily, the first mode set can only contain angle prediction modes, or can contain angle prediction modes and PLANAR modes, or can contain angle prediction modes, DC modes and PLANAR modes, etc., which are not limited here. That is, one possible implementation is that the first mode set only contains angle prediction modes, and the number of these modes can be 65, or can be more because of more fine-grained angle prediction modes. Another possible implementation is that the first mode set only contains angle prediction modes, DC modes and PLANAR modes. The following will be described in detail taking the first mode set containing angle prediction modes, DC modes and PLANAR modes as an example, and the first mode set at this time corresponds to the 67 intra prediction modes 0-66 in VVC.
[0362] In the embodiments of the present application, for the OBIP mode, a data structure can be constructed here to record the occurrence number of each candidate intra prediction mode in the first mode set. Exemplarily, this data structure can be an array, denoted as occurrence
[0067] . It should be noted that because the first mode set is taken as an example with 67 modes, if it is other number, the length of the array is adjusted accordingly. In addition, each value in occurrence
[0067] is initialized to 0.
[0363] In the embodiments of the present application, after determining the one or more preset blocks that have been coded, the occurrence of the modes in the first mode set in the one or more preset blocks can be counted, that is, the statistical result of the at least two candidate intra prediction modes in the first mode set is determined. In a specific embodiment, it can include: determining the prediction mode used by each of the one or more preset blocks; and performing candidate intra prediction mode counting according to the prediction mode used by each of the one or more preset blocks to determine the statistical result of the at least two candidate intra prediction modes in the first mode set.
[0364] It should be noted that in the embodiments of the present application, after the preset block containing a certain coordinate is determined in the coded area, the prediction mode used by the preset block can be determined. For example, if the preset block is an intra-prediction block, the intra-prediction mode used by the preset block can be determined; if the preset block is an inter-prediction block, the inter-prediction mode used by the preset block can be determined. In this way, the statistical results of the at least two candidate intra-prediction modes in the first mode set can be determined according to the prediction modes used by the one or more preset blocks.
[0365] It should also be noted that in the embodiments of the present application, taking the first preset block as an example, for determining the statistical results of the at least two candidate intra-prediction modes in the first mode set, the method can include: determining the intra-prediction mode used by the first preset block; when the intra-prediction mode is a candidate intra-prediction mode in the first mode set, performing accumulation operation on the intra-prediction mode to determine the statistical result of the intra-prediction mode in the first mode set.
[0366] It can be understood that in the embodiments of the present application, the first preset block can be any one of the one or more preset blocks. That is, taking the first preset block as an example, the first preset block can be an inter-prediction block or an intra-prediction block. The following will first describe the processing case when the first preset block is an inter-prediction block.
[0367] In a possible implementation, for determining the statistical results of the at least two candidate intra-prediction modes in the first mode set, the method can include: when the first preset block is an inter-prediction block, not counting the first preset block.
[0368] In the embodiments of the present application, the first preset block can be any one of the one or more preset blocks. If the first preset block is an inter-prediction block, it indicates that the first preset block does not use the mode in the first mode set for prediction, and then the first preset block can not be counted (or said to be skipped), that is, at this time, no accumulation operation is performed on any mode in the first mode set, and the statistical results of the at least two candidate intra-prediction modes in the first mode set can be determined according to the prediction modes used by other preset blocks other than the first preset block.
[0369] In another possible implementation, for determining the statistical results of the at least two candidate intra-prediction modes in the first mode set, the method can include: when the first preset block is an inter-prediction block, determining a first intra-prediction mode derived based on the candidate samples of the first preset block; performing accumulation operation on the first intra-prediction mode to determine the statistical result of the first intra-prediction mode in the first mode set.
[0370] In the embodiments of the present application, the first preset block can be any one of the one or more preset blocks. If the first preset block is an inter-prediction block, indicating that the first preset block is not predicted using the modes in the first mode set, the first intra-prediction mode of the first preset block can also be derived according to the candidate samples of the first preset block, and then the statistical result of the first intra-prediction mode in the first mode set is determined by accumulating the first intra-prediction mode. Specifically, the current accumulated value of the first intra-prediction mode can be determined by accumulating the accumulated value corresponding to the first intra-prediction mode, and the accumulated value corresponding to the first intra-prediction mode is updated based on the current accumulated value; the judgment of the next preset block is continued until the one or more preset blocks are traversed, and the final accumulated value corresponding to the first intra-prediction mode is determined as the statistical result of the first intra-prediction mode in the first mode set.
[0371] In some embodiments, the method can further include: when the first intra-prediction mode is a candidate intra-prediction mode in the first mode set, accumulating the first intra-prediction mode to determine the statistical result of the first intra-prediction mode in the first mode set. That is, after deriving the first intra-prediction mode based on the candidate samples of the first preset block, if the first intra-prediction mode is a candidate intra-prediction mode in the first mode set, the accumulation operation of the first intra-prediction mode can be performed to determine the statistical result of the first intra-prediction mode in the first mode set.
[0372] In some embodiments, the method can further include: determining the candidate samples of the first preset block according to at least part of the samples in the reconstructed block of the first preset block; or determining the candidate samples of the first preset block according to at least part of the samples in the prediction block of the first preset block.
[0373] In the embodiments of the present application, when the first preset block is an inter-prediction block, the first intra-prediction mode of the first preset block can be determined by gradient statistics according to the candidate samples of the first preset block. The candidate samples here can be at least part of the samples in the reconstructed block or the prediction block of the block. That is, the first intra-prediction mode can be derived by gradient statistics according to at least part of the samples in the reconstructed block or the prediction block of the block. As shown in FIG. 32, an 8x8 block is provided here, and the 6x6 candidate samples inside it can be used to derive the first intra-prediction mode, that is, all points in the reconstructed block or the prediction block except the outermost row and column can be used as candidate samples.
[0374] Thus, if the first preset block is an inter-prediction block, and the first preset block is not predicted by using the modes in the first mode set, the first preset block can be directly not counted, i.e., no accumulation operation is performed on any mode in the first mode set; or, the gradient counting manner can be used to derive the first intra-prediction mode for the reconstructed block or the predicted block of the first preset block, and the derived first intra-prediction mode is one of the candidate intra-prediction modes in the first mode set. Assuming that the mode index corresponding to the first intra-prediction mode is X, the accumulation operation is performed on occurrence[X].
[0375] It can also be understood that, in the embodiments of the present application, the processing of the case that the first preset block is an intra-prediction block is described in detail below.
[0376] In a possible implementation, for determining the statistical result of the at least two candidate intra-prediction modes in the first mode set, the method further includes: when the first preset block is an intra-prediction block, determining a second intra-prediction mode used by the first preset block; and when the second intra-prediction mode is a candidate intra-prediction mode in the first mode set, performing accumulation operation on the second intra-prediction mode to determine the statistical result of the second intra-prediction mode in the first mode set.
[0377] In the embodiments of the present application, when the first preset block is an intra-prediction block, and the second intra-prediction mode used by the first preset block is a candidate intra-prediction mode in the first mode set, assuming that the mode index corresponding to the second intra-prediction mode is X, the accumulation operation is performed on occurrence[X]. The second intra-prediction mode is one of the candidate intra-prediction modes in the first mode set.
[0378] In another possible implementation, the second intra-prediction mode is not a candidate intra-prediction mode in the first mode set, and for determining the statistical result of the at least two candidate intra-prediction modes in the first mode set, the method further includes: when the second intra-prediction mode is a candidate intra-prediction mode in a second mode set, determining at least two third intra-prediction modes corresponding to the second intra-prediction mode; and when the at least two third intra-prediction modes are at least two candidate intra-prediction modes in the first mode set, performing accumulation operation on the at least two third intra-prediction modes to determine the statistical result of the at least two third intra-prediction modes in the first mode set.
[0379] In the embodiments of the present application, the second mode set is different from the first mode set, and the second mode set includes modes using at least two intra-prediction modes for combined prediction.
[0380] In a specific embodiment, the mode for combined prediction using at least two intra prediction modes can be DIMD mode, TIMD mode, SGPM mode and OBIP mode, etc. The intra prediction modes include, but are not limited to, DC mode, PLANAR mode and angular prediction mode, etc.
[0381] That is, in a more specific embodiment, the second mode set can include at least one of the following: DIMD mode, TIMD mode, SGPM mode and OBIP mode.
[0382] In the embodiments of the present application, if the first preset block uses DIMD mode, TIMD mode, SGPM mode, OBIP mode, etc., it can use candidate intra prediction modes in the first mode set, and each candidate intra prediction mode is accumulated. For example, if the mode uses candidate intra prediction modes X1 and X2 in the first mode set, occurrence[X1] and occurrence[X2] are accumulated.
[0383] In another possible implementation, the second intra prediction mode is not a candidate intra prediction mode in the first mode set, and for determining the statistical result of at least two candidate intra prediction modes in the first mode set, the method further includes: when the second intra prediction mode is a candidate intra prediction mode in the third mode set, not counting the first preset block.
[0384] In another possible implementation, the second intra prediction mode is not a candidate intra prediction mode in the first mode set, and for determining the statistical result of at least two candidate intra prediction modes in the first mode set, the method further includes: when the second intra prediction mode is a candidate intra prediction mode in the third mode set, determining a fourth intra prediction mode derived based on the candidate samples of the first preset block; and accumulating the fourth intra prediction mode to determine the statistical result of the fourth intra prediction mode in the first mode set. Alternatively, the method further includes: when the second intra prediction mode is a candidate intra prediction mode in the third mode set, determining a fourth intra prediction mode derived based on the candidate samples of the first preset block; and when the fourth intra prediction mode is a candidate intra prediction mode in the first mode set, accumulating the fourth intra prediction mode to determine the statistical result of the fourth intra prediction mode in the first mode set.
[0385] In the embodiments of the present application, the third mode set is different from the first mode set, and the third mode set includes at least one of the following: a mode for predicting by copying an intra block, a mode for predicting using an extrapolation filter, and a mode for predicting using matrix operation.
[0386] In a specific embodiment, the modes for predicting the intra block copy are specifically ITMP mode and IBC mode.
[0387] In a specific embodiment, the modes for predicting using extrapolation filter are specifically EIP mode.
[0388] In a specific embodiment, the modes for predicting using matrix operation are specifically MIP mode.
[0389] That is, in a more specific embodiment, the third mode set can include at least one of the following: ITMP mode, IBC mode, EIP mode and MIP mode.
[0390] In the embodiments of the present application, if the first preset block is an intra prediction block, and the first preset block uses MIP mode, ITMP mode, EIP mode, IBC mode, etc., since it does not use the candidate intra prediction mode in the first mode set for prediction, the first preset block can be directly skipped, i.e. the first preset block is not counted, at this time, no accumulation operation is performed on any mode in the first mode set; or, the reconstructed block or the prediction block of the block can be used to derive the fourth intra prediction mode using the gradient counting method, and the derived fourth intra prediction mode belongs to one of the candidate intra prediction modes in the first mode set. Assuming that the mode index corresponding to the fourth intra prediction mode is X, then the accumulation operation can be performed on occurrence[X].
[0391] Exemplarily, if the first preset block is an intra prediction block, taking deriving the fourth intra prediction mode from the first preset block as an example, the accumulation operation is performed on the fourth intra prediction mode, specifically, the accumulation value corresponding to the fourth intra prediction mode is accumulated to determine the current accumulation value of the fourth intra prediction mode, and the accumulation value corresponding to the fourth intra prediction mode is updated based on the current accumulation value; the judgment of the next preset block is continued until all the one or more preset blocks are traversed, and the final accumulation value corresponding to the fourth intra prediction mode is determined as the statistical result of the fourth intra prediction mode in the first mode set.
[0392] It can also be understood that in the embodiments of the present application, each intra prediction mode actually represents a texture feature. Therefore, the intra prediction mode index is also a texture feature index. For example, the DC mode and the PLANAR mode correspond to gradual texture features, and a certain angle prediction mode corresponds to a texture feature of this angle. The texture feature index can avoid the appearance of intra prediction mode in "inter" on the one hand, and it is also more conducive to possible expansion, for example, one intra prediction mode can correspond to multiple texture features, such as the DC mode can correspond to horizontal gradual texture, vertical gradual texture, diagonal gradual texture, etc.
[0393] In the embodiments of the present application, when the first preset block-based candidate sample is determined to derive the intra prediction mode, the number of candidate samples used to derive the intra prediction mode (i.e., "candidate texture feature index") can be at least one, for example, 1, 2, 3 or more.
[0394] In the embodiments of the present application, for the first preset block, the number of candidate samples can be determined according to the size parameter of the first preset block. That is, when one or more candidate texture feature indexes are derived according to the candidate samples, the number of candidate samples used can be determined by the size parameter of the first preset block. For example, if the size of the first preset block is small, all available samples can be counted; if the size of the first preset block is large, the first preset block can be down-sampled to count, for example, one sample in every 2, or 4, or 8 samples in the horizontal direction and / or the vertical direction. Alternatively, if the size of the first preset block in one of the horizontal or vertical directions is less than or equal to 8, all available samples in that direction are counted; otherwise, if the size of the first preset block in one of the horizontal or vertical directions is less than or equal to 16, one sample in every 2 samples in that direction is counted; otherwise, one sample in every 4 samples in that direction is counted, which is not limited here.
[0395] In the embodiments of the present application, taking the derivation of the first intra prediction mode according to the candidate samples of the first preset block as an example, accordingly, in some embodiments, the first intra prediction mode derived based on the candidate samples of the first preset block is determined, which can include: determining the horizontal gradient value and the vertical gradient value of the candidate sample; determining the texture feature index and the gradient intensity value corresponding to the candidate sample according to the horizontal gradient value and the vertical gradient value of the candidate sample; determining the texture feature statistics table according to the texture feature index and the gradient intensity value corresponding to the candidate sample; and determining the first intra prediction mode derived by the first preset block according to the texture feature statistics table.
[0396] It should be noted that, in the embodiments of the present application, when the horizontal gradient value and the vertical gradient value of the candidate sample are determined to determine the texture feature index and the gradient intensity value corresponding to the candidate sample, it can include: performing angle mapping on the horizontal gradient value and the vertical gradient value of the candidate sample to determine the texture feature index corresponding to the candidate sample; and performing gradient intensity calculation on the horizontal gradient value and the vertical gradient value of the candidate sample to determine the gradient intensity value corresponding to the candidate sample.
[0397] In a specific embodiment, the angle mapping according to the horizontal gradient value and the vertical gradient value of the candidate sample, and determining the texture feature index corresponding to the candidate sample, can comprise: determining the texture feature index corresponding to the candidate sample according to the horizontal gradient value and the vertical gradient value of the candidate sample by using a preset lookup table.
[0398] In the embodiments of the present application, the horizontal gradient value of the candidate sample can be represented by grad x , and the vertical gradient value of the candidate sample can be represented by grad y . In this way, the texture feature index (or also referred to as "virtual intra prediction mode") derived according to grad x and grad y can be realized by table lookup.
[0399] Exemplarily, if abs(grad x ) is equal to 0 and abs(grad y ) is not equal to 0, there is horizontal direction texture, corresponding to the intra prediction mode 18 in VVC. If abs(grad y ) is equal to 0 and abs(grad x ) is not equal to 0, there is vertical direction texture, corresponding to the intra prediction mode 50 in VVC. In the case that abs(grad x ) and abs(grad y ) are not equal to 0, if abs(grad x ) is equal to abs(grad y ) and grad x and grad y have the same sign, it corresponds to the intra prediction mode 34 in VVC. If abs(grad x ) is equal to 2 times abs(grad y ) and grad x and grad y have the same sign, it corresponds to the intra prediction mode 40 in VVC. In addition, other cases can be determined by table lookup according to the same principle.
[0400] In a specific embodiment, the gradient strength calculation according to the horizontal gradient value and the vertical gradient value of the candidate sample, and determining the gradient strength value corresponding to the candidate sample, can comprise: performing addition operation on the absolute value of the horizontal gradient value and the absolute value of the vertical gradient value to determine the gradient strength value corresponding to the candidate sample.
[0401] Here, the gradient strength value corresponding to the candidate sample can be denoted as amp. Exemplarily, amp = abs(grad x ) + abs(grad y ).
[0402] It should be noted that in the embodiments of the present application, the horizontal gradient value and the vertical gradient value of the candidate sample can be calculated using the Sobel operator. For example, taking the Sobel operator as an example, assuming that the sample value of the sample at the sample position (x, y) of the reconstructed block or the prediction block is P x,y , the horizontal gradient value grad x , and the vertical gradient value grad y are calculated according to the formula (7) and the formula (8) as described above.
[0403] It should also be noted that in the embodiments of the present application, considering that the Sobel operator uses the samples in the upper, lower, left and right rows and columns of the current sample, the embodiments of the present application can be configured to calculate the gradients of all the samples in the reconstructed block or the prediction block except the samples in the outermost row and column.
[0404] In some embodiments, according to the texture feature index corresponding to the candidate sample and the gradient strength value, determining the texture feature statistical table can include: when the number of candidate samples is at least one, determining at least one texture feature index and at least one gradient strength value corresponding thereto; according to the at least one texture feature index, determining at least one reference texture feature index having a different characteristic, and according to the at least one gradient strength value, performing accumulation calculation on the gradient strength values belonging to the same reference texture feature index to determine a gradient strength accumulation value corresponding to the at least one reference texture feature index; and according to the at least one reference texture feature index and the gradient strength accumulation value corresponding to the at least one reference texture feature index, determining the texture feature statistical table.
[0405] That is, in the embodiments of the present application, taking at least part of the samples in the reconstructed block or the prediction block as candidate samples as an example, the gradients of all or part of the samples in the reconstructed block or the prediction block are calculated, and generally the horizontal gradient value and the vertical gradient value can be calculated, and here the Sobel operator can be used to calculate the gradient value. For a certain sample, the texture direction of the sample can be inferred according to the horizontal gradient value and the vertical gradient value, such as the horizontal gradient value being non-zero and the vertical gradient value being zero, which means that the texture of the sample is in the vertical direction. Conversely, the horizontal gradient value is zero and the vertical gradient value is non-zero, which means that the texture of the sample is in the horizontal direction. For example, the horizontal gradient value and the vertical gradient value are equal and non-zero, which means that the texture of the sample is at 45 degrees. Of course, there are many other cases in the embodiments of the present application where the horizontal gradient value and the vertical gradient value are both non-zero, and the direction of the texture of the sample can be determined according to the ratio of the two values. In this way, the gradient strength value of each sample can correspond to a corresponding texture feature index, so as to construct a texture feature statistical table, and the gradient strength value of each calculated sample is accumulated in the corresponding texture feature index item in the statistical table to obtain a final texture feature statistical table, and then the first intra prediction mode derived by the first preset block can be determined according to the texture feature statistical table.
[0406] In some embodiments, according to the texture feature statistics table, the first intra prediction mode derived by the first preset block is determined, and the method can include: determining the reference texture feature index corresponding to the highest gradient intensity accumulation value in the texture feature statistics table as the first intra prediction mode derived by the first preset block.
[0407] That is, in the embodiments of the present application, when gradient statistics is performed according to candidate samples, assuming that there are 67 intra prediction modes, the texture feature statistics table can also be set as an array, for example, array A mpAccumulate
[0067] , and each item of A mpAccumulate
[0067] is initialized to 0. Taking the first preset block as an example, if the index (or called "reference texture feature index") of the intra prediction mode derived by a sample position is X, the amp calculated by the sample position is accumulated on A mpAccumulate[X], and the mode with the highest value in the array A mpAccumulate
[0067] is selected as the first intra prediction mode derived by the gradient statistics of the reconstruction block or the prediction block of the block. It should be noted that the first intra prediction mode derived here belongs to one of the candidate intra prediction modes in the first mode set.
[0408] It can also be understood that, in the embodiments of the present application, the process of deriving one of the candidate intra prediction modes in the first mode set for the reconstruction block or the prediction block of the block by gradient statistics can be performed when the block is encoded, and the derived candidate intra prediction mode is saved, so that gradient statistics derivation is not needed when the block is used, thereby saving the amount of calculation.
[0409] It can also be understood that, in the embodiments of the present application, when the intra prediction mode used or derived by the first preset block is determined to be the candidate intra prediction mode in the first mode set, the intra prediction mode can be accumulated at this time.
[0410] In a possible implementation, the first intra prediction mode is accumulated to determine the statistical result of the first intra prediction mode in the first mode set, which can include: accumulating the first intra prediction mode according to the size parameter value of the first preset block to determine the statistical result of the first intra prediction mode in the first mode set.
[0411] In the embodiments of the present application, taking the first preset block as an example, assuming that the width of the first preset block is width and the height of the first preset block is height, then the size parameter value of the first preset block is width*height, if the cumulative operation at this time can consider the size of the preset block, occurrence[X] = occurrence[X] + width*height. Wherein, the first mode set can be represented by an array occurrence[L], L is the number of modes in the first mode set, that is, the length of the array; X is the mode index of the first intra prediction mode in the first mode set.
[0412] In another possible implementation, the cumulative operation on the first intra prediction mode to determine the statistical result of the first intra prediction mode in the first mode set can include: one plus cumulative operation on the first intra prediction mode to determine the statistical result of the first intra prediction mode in the first mode set.
[0413] In the embodiments of the present application, the cumulative operation can also not consider the size of the preset block, at this time, it can also be one plus cumulative operation, that is, occurrence[X] = occurrence[X] + 1, wherein X represents the mode index of the first intra prediction mode in the first mode set. That is, after determining the first intra prediction mode used or derived by the first preset block, if the first intra prediction mode is a candidate intra prediction mode in the first mode set, then the cumulative value on the first intra prediction mode can be operated by one to obtain the statistical result of the first intra prediction mode in the first mode set.
[0414] In this way, after the above operation process, occurrence
[0067] has been statistically completed, that is, the statistical results of at least two candidate intra prediction modes in the first mode set can be obtained.
[0415] S3503, according to the statistical result, determining at least two intra prediction modes of the current block in the first mode set.
[0416] In the embodiments of the present application, according to the obtained statistical result, at least two intra prediction modes of the current block can be determined from the first mode set.
[0417] In some embodiments, according to the statistical result, determining at least two intra prediction modes of the current block in the first mode set can include: according to the statistical result, sorting the candidate intra prediction modes in the first mode set in descending order, and determining the at least two candidate intra prediction modes at the front of the sorting as the at least two intra prediction modes of the current block.
[0418] In the embodiments of the present application, according to the statistical result of occurrence
[0067] , at least two intra prediction modes of the current block can be determined from the first mode set, for example, the first at least two candidate intra prediction modes with the maximum values in occurrence
[0067] can be selected as the at least two intra prediction modes of the current block.
[0419] For example, if it is determined that two intra prediction modes of the current block, the candidate intra prediction mode M1 with the maximum value and the candidate intra prediction mode M2 with the second maximum value in occurrence
[0067] can be selected as the two intra prediction modes of the current block.
[0420] In S3504, the prediction value of the current block is determined according to the at least two intra prediction modes.
[0421] In some embodiments, determining the prediction value of the current block according to the at least two intra prediction modes can include: determining a weight value of each of the at least two intra prediction modes according to the corresponding cumulative value; and determining the prediction value of the current block by performing weighted prediction on the current block according to the at least two intra prediction modes and the respective weight values.
[0422] In the embodiments of the present application, according to the statistical result, not only the at least two intra prediction modes of the current block can be determined from the first mode set, but also the weight value of each of the at least two intra prediction modes can be determined. Specifically, the weight value of each of the at least two intra prediction modes can be determined according to the corresponding cumulative value.
[0423] For example, it is assumed that according to the statistical result, two intra prediction modes of the current block can be derived, specifically, the intra prediction mode M1 and the intra prediction mode M2. When performing weighted prediction on the current block, weighted prediction can be performed according to the intra prediction mode M1, the intra prediction mode M2 and the PLANAR mode, and the respective weights of the intra prediction mode M1, the intra prediction mode M2 and the PLANAR mode are W1, W2 and W3 respectively. The specific calculation formulas are shown in the foregoing formulas (9), (10) and (11). Wherein, occurrence[M1] represents the cumulative value corresponding to the intra prediction mode M1, and occurrence[M2] represents the cumulative value corresponding to the intra prediction mode M2.
[0424] In addition, when the current block uses the OBIP mode, assume that a prediction value obtained by using an intra prediction mode M1 to predict the current block is denoted as Pred1, a prediction value obtained by using an intra prediction mode M2 to predict the current block is denoted as Pred2, and a prediction value obtained by using a PLANAR mode to predict the current block is denoted as Pred3; then a prediction value obtained by using the OBIP mode to predict the current block is denoted as Pred OBIP The specific calculation formula is shown in the aforementioned formula (12).
[0425] In the embodiments of the present application, the intra prediction modes of the two first mode sets and the PLANAR mode are used for weighted prediction. It can be understood that more intra prediction modes of the first mode set can be used, for example, 3, 4, or 5. In addition, the PLANAR mode can be replaced by other modes, for example, the ITMP mode, which is not limited herein.
[0426] It can also be understood that in the embodiments of the present application, the encoding method can also be applied to a multiple transform set selection (MTSS) mode. The NSPT and the LFNST are both transforms for processing various angle textures, and they can have multiple transform kernels, and a transform kernel can be specifically optimized for a certain angle texture. Of course, in addition to the angle texture, the NSPT and the LFNST also include transform kernels for processing gradient textures. In fact, these transform kernels can also be referred to as trained KL transforms (Karhunen-Loeve Transform, KLT). That is, the NSPT and the LFNST have multiple transform kernels, each of which is designed for a specific texture, including angle texture, gradient texture, etc. In addition, gradient texture can be further extended, such as horizontal gradient texture, vertical gradient texture, and diagonal gradient texture. Further, the MTSS mode is not limited to being used for non-separable transforms such as the NSPT and the LFNST, and can also be applied to separable transforms optimized for specific textures.
[0427] It should be further noted that in the embodiments of the present application, the current block can have multiple transform kernel sets, and each intra prediction mode can correspond to a transform kernel set. In other words, each intra prediction mode actually represents a texture feature. Therefore, the intra prediction mode index is also a texture feature index. For example, DC and PLANAR correspond to gradient texture features, and a certain angle prediction mode corresponds to an angle texture feature. The texture feature index can avoid the appearance of intra prediction modes in the "interframe" on the one hand, and it is also more conducive to possible expansion on the other hand, for example, one intra prediction mode can correspond to multiple texture features, such as the DC mode can correspond to horizontal gradient texture, vertical gradient texture, and diagonal gradient texture.
[0428] Specifically, in the embodiments of the present application, it is considered that some intra prediction modes are not simple texture features, but are likely to contain two or more texture features. Therefore, the MTSS technology can be used here. In the MTSS technology, if the intra prediction mode of the current block is a certain special intra prediction mode, there are more than one optional transform kernel group, for example, the first transform kernel group corresponding to the intra prediction mode M1 and the second transform kernel group corresponding to the intra prediction mode M2.
[0429] In some embodiments, for the transform process of the current block, referring to FIG. 36, the method can include:
[0430] S3601, determining the transform kernel of the current block.
[0431] In a possible implementation, the determination of the transform kernel of the current block can include: determining the transform kernel group of the current block; and determining the transform kernel of the current block according to the transform kernel group.
[0432] In some embodiments, the determination of the transform kernel group of the current block can include: performing encoding cost calculation on the current block according to at least two candidate transform kernel groups, determining the cost result corresponding to each of the at least two candidate transform kernel groups, and determining the minimum cost result from the cost results corresponding to the at least two candidate transform kernel groups, and determining the candidate transform kernel group corresponding to the minimum cost result as the transform kernel group of the current block.
[0433] Further, in some embodiments, the method further includes: determining the transform kernel group index of the current block; and performing encoding processing on the transform kernel group index of the current block, and writing the obtained encoding bits into the bitstream.
[0434] It should be noted that in the embodiments of the present application, the transform kernel group index can be used to indicate the number of the transform kernel group of the current block in the at least two candidate transform kernel groups, and the transform kernel group index can be represented by lfnst_nspt_set_index. Wherein, the transform kernel group index of the current block can be an integer greater than or equal to zero, for example, 0, 1, 2, etc. For example, if the value of lfnst_nspt_set_index is equal to 0, it indicates that the first candidate transform kernel group is selected as the transform kernel group of the current block; if the value of lfnst_nspt_set_index is equal to 1, it indicates that the second candidate transform kernel group is selected as the transform kernel group of the current block.
[0435] It should be further explained that in the embodiments of the present application, the transform kernel group index of the current block can be directly written into the code stream, or can be represented by a second syntax element, and then the value of the second syntax element is written into the code stream. In a specific embodiment, the method further comprises: determining the transform kernel group index of the current block; determining the value of the second syntax element according to the transform kernel group index of the current block; and performing encoding processing on the value of the second syntax element, and writing the obtained coded bits into the code stream.
[0436] In the embodiments of the present application, the second syntax element can be represented by lfnst_nspt_set_index. The second syntax element can be used to indicate the transform kernel group index of the current block, specifically the number of the transform kernel group of the current block in the at least two candidate transform kernel groups. Wherein, the value of the second syntax element can be an integer greater than or equal to zero, for example 0, 1, 2, etc. Illustratively, if the value of the second syntax element is equal to 0, it indicates that the first candidate transform kernel group is selected as the transform kernel group of the current block; if the value of the second syntax element is equal to 1, it indicates that the second candidate transform kernel group is selected as the transform kernel group of the current block; if the value of the second syntax element is equal to 2, it indicates that the third candidate transform kernel group is selected as the transform kernel group of the current block.
[0437] Here, for the OBIP mode, because it uses multiple intra prediction modes for weighting, such as it uses two intra prediction modes (M1, M2) in the first mode set and the PLANAR mode for weighting. Then its prediction value has both M1 characteristics and M2 characteristics, and its residual may have M1 characteristics or M2 characteristics, of course, it may also have other characteristics, but the residual and M1, M2 have correlation. Therefore, for the OBIP mode, when determining the transform kernel (group) determined according to the prediction mode such as LFNST or NSPT, a selection can be added. Illustratively, a syntax element is used to indicate whether to determine the transform kernel (group) according to M1 or to determine the transform kernel (group) according to M2. For example, set the second syntax element lfnst_nspt_set_index. If the value of lfnst_nspt_set_index is 0, the first candidate transform kernel group is used, and if the value of lfnst_nspt_set_index is 1, the second candidate transform kernel group is used. Wherein, the first candidate transform kernel group is determined by the first intra prediction mode (M1), and the second candidate transform kernel group is determined by the second intra prediction mode (M2). Optionally, for special cases, for example, the transform kernel group determined by the first intra prediction mode (M1) and the second intra prediction mode (M2) is exactly the same, then other intra prediction modes can be determined, for example, the third intra prediction mode or the PLANAR mode participating in prediction, etc.
[0438] It is also understood that in the embodiments of the present application, the encoding end can also construct the first candidate list. In some embodiments, the method can further include: determining the first candidate list of the current block, the first candidate list indicating at least two candidate transform kernel groups. Here, the transform kernel group index of the current block is specifically the number of the transform kernel group of the current block in the first candidate list.
[0439] It should be noted that in the embodiments of the present application, when the current block uses the OBIP mode, the first candidate list can include at least two candidate texture feature indexes, or the first candidate list can include at least two candidate transform kernel groups. Here, each candidate texture feature index corresponds to a candidate transform kernel group. Therefore, it can be said that: the first candidate list indicates at least two candidate transform kernel groups.
[0440] Exemplarily, the transform kernel group determined by the first intra prediction mode (M1) can be taken as the first candidate transform kernel group of the first candidate list. If the transform kernel group determined by the second intra prediction mode (M2) is different from the first candidate transform kernel group, the transform kernel group determined by the second intra prediction mode (M2) is determined as the second candidate transform kernel group of the first candidate list; otherwise, if the transform kernel group determined by the second intra prediction mode (M2) is the same as the first candidate transform kernel group, the transform kernel groups determined by the third intra prediction mode, the fourth intra prediction mode and the like used by the current block are sequentially checked to see whether they are the same as the first candidate transform kernel group of the first candidate list. If they are different, the transform kernel group is determined as the second candidate transform kernel group of the first candidate list. If all the intra prediction modes used by the current block are checked and the second candidate transform kernel group is still not determined, a default transform kernel group can be used as the second candidate transform kernel group, for example, the transform kernel group determined by the PLANAR mode is determined as the second candidate transform kernel group.
[0441] In a possible implementation, after the transform kernel group of the current block is determined, the transform kernel of the current block can be further determined. The method can include: determining at least two candidate transform kernels included in the transform kernel group; calculating the encoding cost of the current block according to the at least two candidate transform kernels, determining the cost result corresponding to each of the at least two candidate transform kernels; determining the minimum cost result from the cost results corresponding to each of the at least two candidate transform kernels, and determining the candidate transform kernel corresponding to the minimum cost result as the transform kernel of the current block.
[0442] In another possible implementation, when the first candidate list indicates transform cores included in at least two transform core groups, the method can further include: determining at least two candidate transform cores indicated by the first candidate list; performing encoding cost calculation on the current block according to the at least two candidate transform cores, to determine respective cost results of the at least two candidate transform cores; determining a minimum cost result from the respective cost results of the at least two candidate transform cores, and determining a candidate transform core corresponding to the minimum cost result as the transform core of the current block.
[0443] In the embodiments of the present application, the cost calculation can be determined according to a cost result of rate distortion optimization (RDO), a cost result of sum of absolute difference (SAD), or a cost result of sum of absolute transformed difference (SATD), but is not limited herein.
[0444] In a specific embodiment, the performing of encoding cost calculation on the current block according to the at least two candidate transform cores to determine respective cost results of the at least two candidate transform cores can include: performing transform and quantization on a residual block of the current block based on a first candidate transform core, to determine first candidate quantization coefficients of the current block, and performing entropy coding processing on the first candidate quantization coefficients to determine a first cost value of the first candidate transform core; performing inverse quantization and inverse transform on the first candidate quantization coefficients to determine a first candidate residual block of the current block, and determining a first candidate prediction block of the current block according to the first candidate residual block; performing cost calculation according to the first candidate prediction block and an original image of the current block to determine a second cost value of the first candidate transform core; and determining the cost result corresponding to the first candidate transform core according to the first cost value and the second cost value of the first candidate transform core, wherein the first candidate transform core is any one of the at least two candidate transform cores.
[0445] In the embodiments of the present application, for the first candidate list, if the first candidate list indicates transform cores included in two candidate transform core groups, and the first candidate transform core group includes three candidate transform cores and the second candidate transform core group includes two candidate transform cores, the first candidate list can indicate five candidate transform cores; or if the first candidate transform core group includes three candidate transform cores and the second candidate transform core group includes three candidate transform cores, the first candidate list can indicate six candidate transform cores. In this case, after the transform core index of the current block is determined, the transform core of the current block can be determined in the first candidate list according to the transform core index.
[0446] That is, in the embodiments of the present application, for the transform core index of the current block, the transform core index can be used to indicate the number of the transform core of the current block in the transform core group of the current block or the first candidate list. Wherein, the transform core index of the current block can be a positive integer, such as 1, 2, 3, 4, 5, 6, etc. The transform core index sequence number can be directly written into the code stream, or can also be written into the code stream through the value of the third syntax element.
[0447] In a possible implementation, the transform core index of the current block is determined; the transform core index of the current block is encoded, and the obtained encoding bits are written into the code stream.
[0448] In another possible implementation, the value of the third syntax element is determined; wherein the third syntax element is used to indicate whether the first transform mode is used for the current block and the corresponding transform core index used; the value of the third syntax element is encoded, and the obtained encoding bits are written into the code stream.
[0449] It should be further pointed out that in the embodiments of the present application, the first transform mode can be LFNST / NSPT, and the third syntax element can be represented by lfnst_nspt_index. Wherein, the third syntax element can be used to indicate whether the first transform mode is used for the current block, and the corresponding transform core index when the first transform mode is used for the current block. In this case, the value of the third syntax element can be 0, 1, 2, 3, 4, 5, 6, etc.
[0450] In a specific embodiment, if the value of the third syntax element is a third value, it is determined that the first transform mode is not used for the current block; if the value of the third syntax element is a fourth value, it is determined that the first transform mode is used for the current block and the corresponding transform core index. Wherein, the third value can be set to 0, and the fourth value can be set to non-0, such as 1, 2, 3, 4, 5, 6, etc.
[0451] That is, in the embodiments of the present application, for LFNST / NSPT, the transform core index of the current block can also be represented by lfnst_nspt_index. Wherein, lfnst_nspt_index is 0, which means that the current block does not use LFNST / NSPT, and each transform core group of LFNST / NSPT in the ECM has 3 transform cores, so the value of lfnst_nspt_index is 1 or 2 or 3, which means that the current block uses the first transform core or the second transform core or the third transform core of the selected transform core group of LFNST / NSPT.
[0452] In the embodiments of the present application, the current block uses the OBIP mode, at this time, it has more than one selectable transform kernel group. Exemplarily, it has two selectable transform kernel groups, and the possible values of lfnst_nspt_index are 0, 1, 2, 3, 4, 5, and 6. Among them, 1, 2, and 3 correspond to the three transform kernels of the first candidate transform kernel group, and 4, 5, and 6 correspond to the three transform kernels of the second candidate transform kernel group.
[0453] Exemplarily, at this time, the encoding end can also not add a syntax element, but add possible values of the original syntax element. For example, lfnst_nspt_index is used to indicate the transform kernel of LFNST or NSPT. It is known that LFNST / NSPT has three transform kernels in each transform kernel group. For general prediction modes, the possible values of lfnst_nspt_index are 0, 1, 2, and 3. Among them, if the value of lfnst_nspt_index is 0, it means that LFNST or NSPT is not used; if the value of lfnst_nspt_index is 1, it means that the first transform kernel in the selected transform kernel group is used; if the value of lfnst_nspt_index is 2, it means that the second transform kernel in the selected transform kernel group is used; and if the value of lfnst_nspt_index is 3, it means that the third transform kernel in the selected transform kernel group is used. For the OBIP mode, the possible values of lfnst_nspt_index are 0, 1, 2, 3, 4, 5, and 6. Among them, if the value of lfnst_nspt_index is 0, it means that LFNST or NSPT is not used; if the value of lfnst_nspt_index is 1, it means that the first transform kernel in the first candidate transform kernel group is used; if the value of lfnst_nspt_index is 2, it means that the second transform kernel in the first candidate transform kernel group is used; if the value of lfnst_nspt_index is 3, it means that the third transform kernel in the first candidate transform kernel group is used; if the value of lfnst_nspt_index is 4, it means that the first transform kernel in the second candidate transform kernel group is used; if the value of lfnst_nspt_index is 5, it means that the second transform kernel in the second candidate transform kernel group is used; and if the value of lfnst_nspt_index is 6, it means that the third transform kernel in the second candidate transform kernel group is used. In this way, the transform kernel of the current block can be determined.
[0454] S3602, determining the residual value of the current block according to the prediction value of the current block, and transforming the residual value of the current block according to the transform kernel to determine the transform coefficient of the current block.
[0455] It should be noted that in the embodiments of the present application, after the prediction value of the current block is determined, the method can further include: performing subtraction operation on the initial value of the current block and the prediction value of the current block to determine the residual value of the current block.
[0456] It should be further noted that in the embodiments of the present application, when the residual value of the current block is transformed according to the transform kernel to determine the transform coefficient of the current block, it can include: performing non-separable basis transform on the residual value of the current block according to the transform kernel to determine the transform coefficient of the current block; or performing discrete cosine transform on the residual value of the current block to determine the transform block of the current block; and performing low-frequency non-separable transform on the transform block of the current block according to the transform kernel to determine the transform coefficient of the current block.
[0457] In a specific embodiment, if the size parameter of the current block satisfies the first condition, the non-separable basis transform is performed on the residual value of the current block according to the transform kernel to determine the transform coefficient of the current block; if the size parameter of the current block satisfies the second condition, the discrete cosine transform is performed on the residual value of the current block to determine the transform block of the current block; and the low-frequency non-separable transform is performed on the transform block of the current block according to the transform kernel to determine the transform coefficient of the current block.
[0458] Here, the size parameter of the current block satisfying the first condition includes that the size parameter of the current block is small, for example, the size parameter of the current block is less than a certain threshold. That is, for a block with a smaller size, the transform kernel of NSPT is used here, that is, the NSPT transform is performed on the residual block of the current block according to the transform kernel to determine the transform coefficient of the current block.
[0459] Here, the size parameter of the current block satisfying the second condition includes that the size parameter of the current block is large, for example, the size parameter of the current block is greater than a certain threshold. That is, for a block with a larger size, the transform kernel of LFNST is used here, that is, the basis transform of DCT2 is first performed on the residual value of the current block, and then the LFNST transform is performed on the transform block of the current block according to the transform kernel to determine the transform coefficient of the current block.
[0460] S3603, encode the transform coefficient of the current block, and write the obtained encoding bits into the code stream.
[0461] It should be noted that in the embodiments of the present application, when the transform coefficient of the current block is encoded, the method can include: quantizing the transform coefficient of the current block to determine the quantized coefficient of the current block; and encoding the quantized coefficient of the current block to write the obtained encoding bits into the code stream.
[0462] In brief, the encoder determines the prediction block, derives the candidate texture feature index according to the prediction block, and determines the transform core set of NSPT / LFNST according to the candidate texture feature index. If there are multiple transform cores in the transform core set, the encoder tries each transform core in the transform core set.
[0463] If it is NSPT transform, the residual block is forward transformed using NSPT to obtain transform coefficients, the transform coefficients are quantized to obtain quantized coefficients, and the quantized coefficients are entropy coded. The cost of the overhead in the bitstream under the transform core can be obtained through entropy coding. The decoded transform coefficients are obtained by inverse quantization of the quantized coefficients, and the decoded residual block is obtained by inverse NSPT transform of the decoded transform coefficients. The decoded transform coefficients and the original transform coefficients can be different because quantization is lossy, and similarly, the decoded residual block and the original residual block can also be different. The reconstructed block is obtained according to the decoded residual block and the prediction block. The cost of distortion can be obtained according to the reconstructed block and the original image of the current block. The cost of encoding using the transform core of the current NSPT is the cost of the overhead plus the cost of distortion. The costs of several transform cores are compared, and the minimum one is selected as the best choice of NSPT for the current block.
[0464] If it is LFNST transform, the residual block is forward transformed using DCT2 and then forward transformed using LFNST to obtain transform coefficients, the transform coefficients are quantized to obtain quantized coefficients, and the quantized coefficients are entropy coded. The cost of the overhead in the bitstream under the transform core can be obtained through entropy coding. The decoded transform coefficients are obtained by inverse quantization of the quantized coefficients, and the decoded residual block is obtained by inverse LFNST transform of the decoded transform coefficients and inverse DCT2 transform. The decoded transform coefficients and the original transform coefficients can be different because quantization is lossy, and similarly, the decoded residual block and the original residual block can also be different. The reconstructed block is obtained according to the decoded residual block and the prediction block. The cost of distortion can be obtained according to the reconstructed block and the original image of the current block. The cost of encoding using the transform core of the current NSPT is the cost of the overhead plus the cost of distortion. The costs of several transform cores are compared, and the minimum one is selected as the best choice of NSPT for the current block.
[0465] It should be noted that in the embodiments of the present application, the residual block of the current block can include the residual value of at least one pixel in the current block, and therefore can also be referred to as the "residual value of the current block" herein; the prediction block of the current block can include the prediction value of at least one pixel in the current block, and therefore can also be referred to as the "prediction value of the current block" herein; similarly, the reconstructed block of the current block can include the reconstructed value of at least one pixel in the current block, and therefore can also be referred to as the "reconstructed value of the current block" herein.
[0466] It should be further noted that in the embodiments of the present application, the "transformation" of the encoding end on the residual block can also be referred to as "forward transformation", which specifically refers to the transformation from the spatial domain to the frequency domain to remove the correlation of the residual. It should be noted that if the standard only specifies decoding, the "transformation" in the standard text is the part of decoding, which specifically refers to the "inverse transformation" in this paper.
[0467] In yet another embodiment of the present application, the present embodiment also provides a code stream, wherein the code stream is generated by bit coding according to the encoding method described in the foregoing embodiments. Wherein the to-be-encoded information in the encoding method can include at least one of the following: the transform coefficients of the current block, the transform core group index of the current block, the transform core index of the current block, the value of the first syntax element, the value of the second syntax element and the value of the third syntax element. Here, these to-be-encoded information are encoded and processed to be written into the code stream.
[0468] In the embodiments of the present application, the first syntax element is used to indicate whether the current block uses the first prediction mode, the value of the second syntax element is used to indicate the transform core group index of the current block, and the value of the third syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform core index when the current block uses the first transform mode. For example, taking the third syntax element as an example, if the value of the third syntax element is 0, it is determined that the current block does not use the first transform mode; if the value of the second syntax element is non-zero, it is determined that the current block uses the first transform mode and the corresponding transform core index. For example, if the value of the second syntax element is 1, it is determined that the current block uses the first transform core; if the value of the second syntax element is 2, it is determined that the current block uses the second transform core, and so on, which is not limited here.
[0469] The embodiment of the present application provides a coding method, in particular, an intra prediction mode based on occurrence situation statistics and a transformation scheme thereof. One or more preset blocks are determined to be coded; at least two candidate intra prediction modes in a first mode set are determined according to the prediction modes used by the one or more preset blocks respectively; then, at least two intra prediction modes of a current block are determined in the first mode set according to the statistical results; and the current block is predicted according to the at least two intra prediction modes, and a prediction value of the current block is determined. In this way, the statistical results of the related blocks are obtained by using the coded related blocks, and the at least two intra prediction modes used by the current block are derived according to the statistical results. Since it is not necessary to explicitly indicate which intra prediction mode is used, the overhead in the code stream can be saved. Moreover, the prediction accuracy can be improved by using the weighted prediction of the at least two intra prediction modes, and the prediction error is reduced to a certain extent. In addition, the at least two intra prediction modes used by the current block are derived according to the statistical results, which can avoid the intra prediction mode derivation through gradient calculation required by the DIMD mode, reduce the calculation complexity, improve the coding and decoding efficiency, and further improve the compression performance.
[0470] In another embodiment of the present application, based on the coding method of the foregoing embodiment, the OBIP mode counts the intra prediction modes used by some related blocks, and derives the intra prediction mode used by the current block according to the intra prediction modes of the related blocks. Since each intra predicted block saves the intra prediction mode information used by it, compared with the DIMD mode which needs to derive through gradient calculation, the OBIP mode can save this link.
[0471] Briefly, the intra prediction mode includes many modes, such as the PLANAR mode, the DC mode and the 65 angle prediction modes, the MIP mode and the like which already exist in the VVC. The DIMD mode, the TIMD mode, the SGPM mode, the ITMP mode and the OBIP mode and the like are newly added in the ECM. A first mode set is defined here. A possible implementation manner is that the first mode set only contains the angle prediction modes, and the number of these modes can be 65, or the number can be more because of more fine-grained angle prediction modes. Another possible implementation manner is that the first mode set only contains the angle prediction modes and the DC mode and the PLANAR mode. Here, the second method is taken as an example for introduction, and the first mode set at this time corresponds to the 67 intra prediction modes numbered 0-66 in the VVC.
[0472] The OBIP mode can derive candidate intra prediction modes in a first mode set, and use them to generate a prediction value by weighting. The PLANAR mode or the ITMP mode can also be used to generate the prediction value by weighting. The following first introduces how to derive candidate intra prediction modes in the first mode set for the current block.
[0473] Here, a data structure can be constructed to record the occurrence times of various candidate intra prediction modes in the first mode set. The data structure can be an array, denoted as occurrence
[0067] . It should be noted that this is because there are 67 intra prediction modes in the first mode set, and if there are other numbers, the length of the array is set accordingly. Each number in occurrence
[0067] is initialized to 0.
[0474] (1) Data source.
[0475] Method one: using adjacent blocks and non-adjacent blocks for statistics.
[0476] The OBIP mode can count adjacent blocks and non-adjacent blocks. As shown in FIG. 29, adjacent blocks are labeled 1-7, and non-adjacent blocks are labeled 8-25. One possible implementation is to use only adjacent blocks. Another possible implementation is to use only non-adjacent blocks. Yet another possible implementation is to use both adjacent blocks and non-adjacent blocks.
[0477] Suppose the position of the top-left corner of the current block is (x, y), the width of the current block is width, and the height of the current block is height. Then, for the adjacent blocks of the current block, adjacent block 1 is the block containing the coordinates (x-1, y-1), adjacent block 2 is the block containing the coordinates (x, y-1), adjacent block 2 is the block containing the coordinates (x, y-1), adjacent block 3 is the block containing the coordinates (x-1, y), adjacent block 4 is the block containing the coordinates (x+width-1, y-1), adjacent block 5 is the block containing the coordinates (x+width, y-1), adjacent block 6 is the block containing the coordinates (x-1, y+height-1), and adjacent block 7 is the block containing the coordinates (x-1, y+height).
[0478] For non-adjacent blocks (i.e. non-adjacent blocks), they can also be located according to certain rules. For example, the position of non-adjacent blocks can be calculated according to (x, y), width and height. In FIG. 29, the distance of each grid in vertical direction is height, and the distance of each grid in horizontal direction is width. Then for the non-adjacent blocks of the current block, the non-adjacent block 8 is the block containing the coordinate (x-width-1, y-height-1), the non-adjacent block 9 is the block containing the coordinate (x+2*width-1, y-height-1), the non-adjacent block 10 is the block containing the coordinate (x-width-1, y+2*height-1), the non-adjacent block 11 is the block containing the coordinate (x-2*width-1, y-2*height-1), the non-adjacent block 12 is the block containing the coordinate (x+width / 2, y-2*height-1), and so on.
[0479] It should be noted that the block containing a coordinate mentioned herein refers to the CU or PU containing the coordinate. After the CU or PU containing the coordinate is determined, the intra prediction mode of the CU or PU can be obtained.
[0480] Method two: all blocks in a preset range.
[0481] A range can be set here. It is assumed that the top-left position of the current block is (x, y), the width of the current block is width, and the height of the current block is height. The preset range can be set according to the position of the current block and width and height.
[0482] As shown in FIG. 30, one possible preset range is all the blocks in the rectangular region with the top-left corner (x-2*width, y-2*height) and the top-right corner (x+3*width-1, y-2*height) and the bottom-left corner (x-2*width, y+3*height-1). If the minimum size of the CU or PU is 4*4, all the 4*4 blocks in the search range can be counted, for example, the first 4*4 block in the top-left corner is the block containing the coordinate (x-2*width, y-2*height), and the 4*4 blocks in the same row can be scanned by adding 4 horizontally, and the 4*4 blocks in the same column can be scanned by adding 4 vertically. The block containing a coordinate mentioned herein refers to the CU or PU containing the coordinate. After the CU or PU containing the coordinate is determined, the intra prediction mode of the CU or PU can be obtained.
[0483] (2) Statistics.
[0484] All blocks found in the data source.
[0485] (a) If this block is an intra coded block:
[0486] If this block uses a candidate intra prediction mode within the first mode set, say mode number X, then occurrence[X] is incremented.
[0487] If this block uses DIMD mode, TIMD mode, SGPM mode, OBIP mode, etc., it can use multiple candidate intra prediction modes within the first mode set, say mode numbers X1 and X2, then occurrence[X1] and occurrence[X2] are incremented respectively.
[0488] If this block uses MIP mode, ITMP mode, EIP mode, IBC mode, etc., it does not use any candidate intra prediction mode within the first mode set for prediction, one possible implementation is not to increment any mode within the first mode set. Another possible implementation is to derive a candidate intra prediction mode within the first mode set, say mode number X, from gradient statistics of the reconstructed or predicted block of this block, then occurrence[X] is incremented.
[0489] (b) If this block is an inter coded block:
[0490] This block does not use any mode within the first mode set for prediction, one possible implementation is not to increment any mode within the first mode set. Another possible implementation is to derive a candidate intra prediction mode within the first mode set, say mode number X, from gradient statistics of the reconstructed or predicted block of this block, then occurrence[X] is incremented. The process of deriving a candidate intra prediction mode within the first mode set from gradient statistics of the reconstructed or predicted block of this block can be done when this block is coded, and the derived candidate intra prediction mode within the first mode set is saved, so that it does not need to be derived again when this block is used, which can save computation.
[0491] (3) The method of incrementing the occurrence statistics.
[0492] For mode X, one possible accumulation method is occurrence[X] = occurrence[X] + 1. This method can be used for both the above-mentioned data source method one (using neighboring blocks and non-neighboring blocks for statistics) and data source method two (all blocks within a preset range). Another possible accumulation method is to consider the size of this block, i.e., occurrence[X] = occurrence[X] + width*height.
[0493] (4) Deriving a candidate intra prediction mode within a first mode set for a reconstructed block or a predicted block of a block using gradient statistics.
[0494] Exemplarily, the gradient can be calculated using a Sobel operator. An example of a Sobel operator is as follows
[0495] Operator for horizontal gradient value:
[0496] Operator for vertical gradient value:
[0497] Thus, assuming the sample value of the sample at position (x, y) of the reconstructed block or the predicted block is P x,y , the horizontal gradient value grad x and the vertical gradient value grad y are calculated as shown in the above-mentioned formula (7) and formula (8).
[0498] An example of the gradient strength denoted as amp is amp = abs(grad x ) + abs(grad y ).
[0499] Deriving a virtual intra prediction mode according to grad x and grad y can be achieved by looking up a table. Exemplarily, if abs(grad x ) is equal to 0 and abs(grad y ) is not equal to 0, there is horizontal direction texture, corresponding to the intra prediction mode 18 in VVC. If abs(grad y ) is equal to 0 and abs(grad x ) is not equal to 0, there is vertical direction texture, corresponding to the intra prediction mode 50 in VVC. In the case where abs(grad x ) and abs(grad y ) are not equal to 0, if abs(grad x ) is equal to abs(grad y ), grad x and grady The same as the intra prediction mode 34 in VVC. If abs(grad x ) is equal to 2 times abs(grad y ), and grad x and grad y The same as the intra prediction mode 40 in VVC. In addition, other cases can be determined according to the same principle.
[0500] Exemplarily, the above derivation is performed on all points in the reconstructed block or the prediction block except the outermost one row and one column. As shown in FIG. 32, if it is an 8x8 block, the above derivation is performed on the 6x6 samples in the interior of the block. Here, an array AmpAccumulate
[0067] is recorded, each item of AmpAccumulate
[0067] is initialized to 0. If the intra prediction mode derived for a sample position is X, the amp calculated for the sample position is added to AmpAccumulate[X], and the mode with the highest value in AmpAccumulate
[0067] is taken as the candidate intra prediction mode of the reconstructed block or the prediction block derived by gradient statistics.
[0501] (5) Determine the intra prediction mode and weight used by OBIP.
[0502] After the above operation, occurrence
[0067] is statistically completed. According to occurrence
[0067] , the intra prediction mode and weight used by OBIP can be determined.
[0503] One possible implementation is similar to the weighted prediction method of the DIMD mode in the related art, the first two modes with the highest values in occurrence
[0067] are selected and recorded as M1 and M2, and M1, M2 and PLANAR are weighted. The weights of M1, M2 and PLANAR are W1, W2 and W3 respectively, which are specifically shown in the above formula (9), formula (10) and formula (11).
[0504] In addition, the prediction values of M1, M2 and PLANAR are recorded as Pred1, Pred2 and Pred3 respectively, and the prediction value of the OBIP mode is Pred OBIP , which is specifically shown in the above formula (12).
[0505] In the above example, two intra prediction modes in the first mode set and the PLANAR mode are used for weighting. It can be understood that more intra prediction modes in the first mode set can also be used here, such as 3, 4, 5; in addition, the PLANAR mode here can be replaced by other modes, for example, the ITMP mode.
[0506] (6) Control of OBIP.
[0507] Here, a syntax element of a block level such as CU or PU can be set to indicate whether the current block uses the OBIP mode, for example, cu_obip_flag. If the value of cu_obip_flag is 1, it indicates that the current CU uses the OBIP mode for prediction. If the value of cu_obip_flag is 0, it indicates that the current CU does not use the OBIP mode for prediction. If cu_obip_flag does not exist, it is inferred that the value of cu_obip_flag is 0.
[0508] In a specific embodiment, the workflow of the OBIP mode prediction includes: counting the occurrence of candidate intra prediction modes in the first mode set in a specified block or region; determining the intra prediction mode and its corresponding weight according to the counting result; and generating a prediction value according to the determined intra prediction mode and its weight.
[0509] It should be noted that the data source for the OBIP mode is the same for the encoder and the decoder, so the encoder and the decoder will obtain the same intra prediction mode and weight, and the same prediction value will be generated using the OBIP mode.
[0510] At the encoding end, the encoder will attempt to use the OBIP mode for encoding and determine the encoding cost, such as the rate-distortion cost, of using the OBIP mode. The encoder will also attempt to use other modes for encoding and determine the encoding cost thereof. The encoder selects the mode with the lowest encoding cost as the prediction mode of the current block. Exemplarily, the encoder can write the value of cu_obip_flag in the bitstream.
[0511] At the decoding end, the decoder will parse the value of cu_obip_flag. If the value of cu_obip_flag is 1, the OBIP mode can be used for prediction.
[0512] (6) Adaptation of transform mode.
[0513] In VVC and ECM, the transform kernel set of LFNST or NSPT is derived from the current block's intra prediction mode. For the case of using only a single intra prediction mode, such as PLANAR mode, DC mode or angular prediction mode, this setting is reasonable. But for the case of OBIP mode, because it uses multiple intra prediction modes for weighting, such as it uses two intra prediction modes in the first mode set and PLANAR mode for weighting. Then its prediction value has both M1 and M2 characteristics, and its residual may have both M1 and M2 characteristics, of course, it may also have other characteristics, but the residual and M1, M2 have correlation. Thus for the OBIP mode, in determining the transform according to the prediction mode, such as LFNST or NSPT, a selection can be added. A syntax element is used to indicate whether to determine the transform kernel according to M1 or M2. For example, set a syntax element lfnst_nspt_set_index. If the value of lfnst_nspt_set_index is 0, the first candidate transform kernel set is used; if the value of lfnst_nspt_set_index is 1, the second candidate transform kernel set is used. The first candidate transform kernel set is determined by the first intra prediction mode (M1), and the first candidate transform kernel set is determined by the second intra prediction mode (M2). Optionally, for special cases, for example, the first intra prediction mode (M1) and the second intra prediction mode (M2) determine the transform kernel set exactly the same, then other intra prediction modes can be determined, such as the third intra prediction mode participating in prediction, or PLANAR mode, etc.
[0514] Of course, new syntax elements can not be added, but the possible values of the original syntax elements can be added. For example, lfnst_nspt_index is used to indicate the transform kernel of LFNST or NSPT. It is known that each transform kernel group of LFNST / NSPT has 3 transform kernels. For general modes, the possible values of lfnst_nspt_index are 0, 1, 2, and 3. When lfnst_nspt_index is 0, it means that LFNST or NSPT is not used. When lfnst_nspt_index is 1, it means that the first transform kernel in the selected transform kernel group is used. When lfnst_nspt_index is 2, it means that the second transform kernel in the selected transform kernel group is used. When lfnst_nspt_index is 3, it means that the third transform kernel in the selected transform kernel group is used. For OBIP mode, the possible values of lfnst_nspt_index are 0, 1, 2, 3, 4, 5, and 6. When lfnst_nspt_index is 0, it means that LFNST or NSPT is not used. When lfnst_nspt_index is 1, it means that the first transform kernel in the first candidate transform kernel group is used. When lfnst_nspt_index is 2, it means that the second transform kernel in the first candidate transform kernel group is used. When lfnst_nspt_index is 3, it means that the third transform kernel in the first candidate transform kernel group is used. When lfnst_nspt_index is 4, it means that the first transform kernel in the second candidate transform kernel group is used. When lfnst_nspt_index is 5, it means that the second transform kernel in the second candidate transform kernel group is used. When lfnst_nspt_index is 6, it means that the third transform kernel in the second candidate transform kernel group is used.
[0515] In a specific embodiment, taking the decoding end as an example, the working process of OBIP mode prediction plus transform includes: parsing the code stream to determine whether the current block uses the OBIP mode; if the current block uses the OBIP mode, counting the occurrence of the candidate intra prediction modes in the first mode set in the specified block or region; determining the intra prediction mode and its corresponding weight according to the counting result; generating a prediction value according to the determined intra prediction mode and its weight; if the current block uses LFNST / NSPT for transform, parsing the code stream to determine the transform kernel group index; determining the transform kernel group according to the transform kernel group index, and using the transform kernel in the transform kernel group to determine the residual value; and generating the reconstructed value according to the prediction value and the residual value.
[0516] At the encoding end, the encoder attempts to encode using the OBIP mode. After generating the prediction value, if there is a prediction residual, the encoder attempts to use various transform methods, such as the above-mentioned LFNST or NSPT, which can have 2 transform kernel groups selectable, and each transform kernel group has 3 transform kernels selectable. The encoder attempts to use the 6 possible LFNST or NSPT transforms, and selects the transform kernel with the minimum rate-distortion cost. If the encoder determines that the current block uses the OBIP mode, and the encoder determines that the current block uses LFNST / NSPT for transformation, the transform kernel group index needs to be written into the code stream.
[0517] In the embodiments of the present application, the specific implementation of the foregoing embodiments is described in detail through the foregoing embodiments. It can be seen from the foregoing embodiments that, according to the technical solutions of the foregoing embodiments, the intra prediction mode used by the current block is derived by using the intra prediction modes of the related blocks in the coded region to obtain relevant information. Since it is not necessary to explicitly indicate which intra prediction mode is used, the overhead in the code stream can be saved. Moreover, the prediction accuracy can be improved by using multiple intra prediction modes for weighted prediction, and the prediction error can be reduced to a certain extent. Overall, the compression performance can be improved.
[0518] In still another embodiment of the present application, based on the same inventive concept as the foregoing embodiments, FIG. 37 is a schematic diagram of the composition structure of an encoder provided in an embodiment of the present application. As shown in FIG. 37, the encoder 370 can include a first determination unit 3701 and a first prediction unit 3702, wherein:
[0519] The first determination unit 3701 is configured to determine one or more preset blocks that have been coded, and determine the statistical results of at least two candidate intra prediction modes in the first mode set according to the prediction modes used by the one or more preset blocks respectively;
[0520] The first determination unit 3701 is further configured to determine at least two intra prediction modes of the current block in the first mode set according to the statistical results;
[0521] The first prediction unit 3702 is configured to predict the current block according to the at least two intra prediction modes, and determine the prediction value of the current block.
[0522] In some embodiments, the first determination unit 3701 is further configured to, when the current block uses the first prediction mode, perform the steps of determining one or more preset blocks that have been coded, determining the statistical results of at least two candidate intra prediction modes in the first mode set according to the prediction modes used by the one or more preset blocks respectively, determining at least two intra prediction modes of the current block in the first mode set according to the statistical results, and predicting the current block according to the at least two intra prediction modes to determine the prediction value of the current block.
[0523] In some embodiments, the first determining unit 3701 is further configured to determine a plurality of candidate prediction modes, wherein the plurality of candidate prediction modes at least include the first prediction mode; calculate coding cost of the current block according to the plurality of candidate prediction modes, determine a cost result corresponding to each of the plurality of candidate prediction modes; determine a minimum cost result from the cost results corresponding to the plurality of candidate prediction modes; and determine whether the current block uses the first prediction mode according to a candidate prediction mode corresponding to the minimum cost result.
[0524] In some embodiments, the first determining unit 3701 is further configured to determine that the current block uses the first prediction mode when the candidate prediction mode corresponding to the minimum cost result is the first prediction mode; and determine that the current block does not use the first prediction mode when the candidate prediction mode corresponding to the minimum cost result is not the first prediction mode.
[0525] In some embodiments, referring to FIG. 37, the encoder 370 can further include an encoding unit 3703, wherein: the first determining unit 3701 is further configured to determine a value of a first syntax element, the first syntax element being used to indicate whether the current block uses the first prediction mode; and the encoding unit 3703 is configured to encode the value of the first syntax element and write the obtained coded bits into a bitstream.
[0526] In some embodiments, the first determining unit 3701 is further configured to determine at least one coded candidate block, wherein the candidate block includes: a neighboring block of the current block, and / or a non-neighboring block of the current block; and determine one or more preset blocks according to the at least one candidate block.
[0527] In some embodiments, the first determining unit 3701 is further configured to determine a preset range based on the current block; and determine the one or more preset blocks according to all coded blocks in the preset range.
[0528] In some embodiments, the first mode set includes at least one of the following: an angular prediction mode, a DC mode, and a PLANAR mode.
[0529] In some embodiments, referring to FIG. 37, the encoder 370 can further include a first statistical unit 3704 configured to not count a first preset block when the first preset block is an inter-prediction block, wherein the first preset block is any one of the one or more preset blocks.
[0530] In some embodiments, the first statistics unit 3704 is further configured to, when the first preset block is an inter-prediction block, determine a first intra-prediction mode derived based on candidate samples of the first preset block; and accumulate the first intra-prediction mode to determine a statistical result of the first intra-prediction mode in the first mode set, wherein the first preset block is any one of one or more preset blocks.
[0531] In some embodiments, the first statistics unit 3704 is further configured to, when the first intra-prediction mode is a candidate intra-prediction mode in the first mode set, accumulate the first intra-prediction mode to determine the statistical result of the first intra-prediction mode in the first mode set.
[0532] In some embodiments, the first determination unit 3701 is further configured to determine the candidate samples of the first preset block according to at least part of samples in a reconstructed block of the first preset block, or determine the candidate samples of the first preset block according to at least part of samples in a prediction block of the first preset block.
[0533] In some embodiments, the first determination unit 3701 is further configured to accumulate the first intra-prediction mode according to a size parameter value of the first preset block to determine the statistical result of the first intra-prediction mode in the first mode set.
[0534] In some embodiments, the first determination unit 3701 is further configured to accumulate the first intra-prediction mode by one to determine the statistical result of the first intra-prediction mode in the first mode set.
[0535] In some embodiments, the first determination unit 3701 is further configured to sort the candidate intra-prediction modes in the first mode set in descending order according to the statistical results, and determine at least two intra-prediction modes with higher orders as the at least two intra-prediction modes of the current block.
[0536] In some embodiments, the first prediction unit 3702 is further configured to determine a weight value of each of the at least two intra-prediction modes according to the statistical result corresponding to the at least two intra-prediction modes, and perform weighted prediction on the current block according to the at least two intra-prediction modes and the weight values to determine a prediction value of the current block.
[0537] In some embodiments, referring to FIG. 37, the encoder 370 can further include a first transform unit 3705, wherein: the first determining unit 3701 is further configured to determine a transform core of the current block, and determine a residual value of the current block according to the prediction value of the current block; and the first transform unit 3705 is configured to determine a transform coefficient of the current block by transforming the residual value of the current block according to the transform core; and the encoding unit 3703 is further configured to perform encoding processing on the transform coefficient of the current block, and write the obtained encoding bits into the bitstream.
[0538] In some embodiments, the first determining unit 3701 is further configured to determine a transform core group of the current block, and determine the transform core of the current block according to the transform core group.
[0539] In some embodiments, the first determining unit 3701 is further configured to determine a cost result corresponding to each of the at least two candidate transform core groups according to the encoding cost calculation of the current block on the at least two candidate transform core groups, and determine a minimum cost result from the cost results corresponding to the at least two candidate transform core groups, and determine the candidate transform core group corresponding to the minimum cost result as the transform core group of the current block.
[0540] In some embodiments, the first determining unit 3701 is further configured to determine a transform core group index of the current block, wherein the transform core group index is used to indicate a number of the transform core group of the current block in the at least two candidate transform core groups; and the encoding unit 3703 is further configured to perform encoding processing on the transform core group index of the current block, and write the obtained encoding bits into the bitstream.
[0541] In some embodiments, the first determining unit 3701 is further configured to determine at least two candidate transform cores included in the transform core group, determine a cost result corresponding to each of the at least two candidate transform cores according to the encoding cost calculation of the current block on the at least two candidate transform cores, and determine a minimum cost result from the cost results corresponding to the at least two candidate transform cores, and determine the candidate transform core corresponding to the minimum cost result as the transform core of the current block.
[0542] In some embodiments, the first determining unit 3701 is further configured to determine a transform core index of the current block, wherein the transform core index is used to indicate a number of the transform core of the current block in the transform core group, and the first candidate list indicates transform cores included in the at least two candidate transform core groups; and the encoding unit 3703 is further configured to perform encoding processing on the transform core index of the current block, and write the obtained encoding bits into the bitstream.
[0543] It can be understood that, in the embodiments of the present application, the "unit" can be a part of circuit, a part of processor, a part of program or software, etc., and of course can be a module, and can also be non-modular. Moreover, each component in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware, or in the form of a software functional module.
[0544] In still another embodiment of the present application, Fig. 38 is a specific hardware structure diagram of an encoder provided by the embodiments of the present application. As shown in Fig. 38, the encoder 370 can include a first communication interface 3801, a first memory 3802 and a first processor 3803; each component is coupled together through a first bus system 3804. It can be understood that the first bus system 3804 is used to realize the connection communication between the components. The first bus system 3804 includes not only a data bus, but also a power supply bus, a control bus and a status signal bus. However, in order to clearly illustrate, all kinds of buses are marked as the first bus system 3804 in Fig. 38. Among them,
[0545] The first communication interface 3801 is used for receiving and sending signals in the process of transceiving information with other external network elements;
[0546] The first memory 3802 is used for storing a computer program capable of running on the first processor 3803;
[0547] The first processor 3803 is used for, when running the computer program, performing:
[0548] determining one or more preset blocks that have been encoded; determining statistical results of at least two candidate intra-prediction modes in the first mode set according to the prediction modes used by the one or more preset blocks respectively; determining at least two intra-prediction modes of the current block in the first mode set according to the statistical results; and predicting the current block according to the at least two intra-prediction modes to determine the prediction value of the current block.
[0549] It is to be appreciated that the first memory 3802 in the embodiments of the application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. In one example, a non-volatile memory can be a read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory. A volatile memory can be a random access memory (RAM), which is used as the external cache. By way of example, and not limitation, many forms of RAM are available, for example, static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The first memory 3802 of the system and method described herein are intended to include, without being limited to, these and any other suitable types of memory.
[0550] The first processor 3803 can be an integrated circuit chip having a processing capability of signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware or the instruction in the form of software in the first processor 3803. The first processor 3803 described above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the first storage 3802, and the first processor 3803 reads the information in the first storage 3802, and combines the hardware to complete the steps of the above method.
[0551] It can be understood that the embodiments described in the present application can be realized by hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be realized in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, other electronic units for executing functions described in the present application or a combination thereof. For software implementation, the technology described in the present application can be realized by modules (such as processes, functions, etc.) for executing functions described in the present application. The software code can be stored in the memory and executed by the processor. The memory can be implemented in the processor or outside the processor.
[0552] Optionally, as another embodiment, the first processor 3803 is further configured to, when executing the computer program, perform the method according to any one of the preceding embodiments.
[0553] The embodiment provides an encoder, in which statistics of related blocks are obtained by using the related blocks, statistical results of the related blocks are obtained, and at least two intra-prediction modes used by a current block are derived according to the statistical results; since it is not necessary to explicitly indicate which kind of intra-prediction mode is used, the overhead in a code stream can be saved; moreover, since weighted prediction is performed by using the at least two intra-prediction modes, the prediction accuracy can be improved, and the prediction error can be reduced to a certain extent; in addition, since the at least two intra-prediction modes used by the current block are derived according to the statistical results, it is not necessary to perform intra-prediction mode derivation by using gradient calculation as in the DIMD mode, the calculation complexity is reduced, the coding and decoding efficiency can be improved, and the compression performance is improved.
[0554] In still another embodiment of the present application, based on the same inventive concept as in the preceding embodiments, FIG. 39 is a schematic diagram of a composition structure of a decoder provided in the embodiment of the present application. As shown in FIG. 39, the decoder 390 can include a second determination unit 3901 and a second prediction unit 3902, wherein:
[0555] The second determination unit 3901 is configured to determine one or more preset blocks that have been decoded; and determine statistical results of at least two candidate intra-prediction modes in a first mode set according to prediction modes used by the one or more preset blocks respectively;
[0556] The second determination unit 3901 is further configured to determine, according to the statistical results, the at least two intra-prediction modes of the current block in the first mode set;
[0557] The second prediction unit 3902 is configured to predict the current block according to the at least two intra-prediction modes, and determine a prediction value of the current block.
[0558] In some embodiments, the second determination unit 3901 is further configured to determine at least one candidate block that has been decoded, wherein the candidate block includes a neighboring block of the current block, and / or a non-neighboring block of the current block; and determine the one or more preset blocks according to the at least one candidate block.
[0559] In some embodiments, the second determination unit 3901 is further configured to determine a preset range based on the current block; and determine the one or more preset blocks according to all the decoded blocks in the preset range.
[0560] In some embodiments, the first mode set includes at least one of the following: an angular prediction mode, a DC mode and a PLANAR mode.
[0561] In some embodiments, referring to FIG. 39, the decoder 390 can further include a second statistics unit 3903 configured to, when the first preset block is an inter-predicted block, not count the first preset block; wherein the first preset block is any one of the one or more preset blocks.
[0562] In some embodiments, the second statistics unit 3903 is further configured to, when the first preset block is an inter-predicted block, determine a first intra-prediction mode derived based on candidate samples of the first preset block; and accumulate the first intra-prediction mode to determine a statistical result of the first intra-prediction mode in the first mode set; wherein the first preset block is any one of the one or more preset blocks.
[0563] In some embodiments, the second statistics unit 3903 is further configured to, when the first intra-prediction mode is a candidate intra-prediction mode in the first mode set, accumulate the first intra-prediction mode to determine the statistical result of the first intra-prediction mode in the first mode set.
[0564] In some embodiments, the second determination unit 3901 is further configured to determine the candidate samples of the first preset block according to at least part of samples in a reconstructed block of the first preset block; or determine the candidate samples of the first preset block according to at least part of samples in a prediction block of the first preset block.
[0565] In some embodiments, the second determination unit 3901 is further configured to accumulate the first intra-prediction mode according to a size parameter value of the first preset block to determine the statistical result of the first intra-prediction mode in the first mode set.
[0566] In some embodiments, the second determination unit 3901 is further configured to accumulate the first intra-prediction mode by one to determine the statistical result of the first intra-prediction mode in the first mode set.
[0567] In some embodiments, the second statistics unit 3903 is further configured to, when the first preset block is an intra-predicted block, determine a second intra-prediction mode used by the first preset block; and when the second intra-prediction mode is a first candidate intra-prediction mode in the first mode set, accumulate the second intra-prediction mode to determine a statistical result of the second intra-prediction mode in the first mode set; wherein the first preset block is any one of the one or more preset blocks.
[0568] In some embodiments, the second statistical unit 3903 is further configured to, when the second intra prediction mode is a candidate intra prediction mode in a second mode set, determine at least two third intra prediction modes corresponding to the second intra prediction mode; when the at least two third intra prediction modes are at least two candidate intra prediction modes in the first mode set, respectively perform accumulation operation on the at least two third intra prediction modes to determine statistical results of the at least two third intra prediction modes in the first mode set; and wherein the second mode set is different from the first mode set, and the second mode set comprises a mode of performing combined prediction using at least two intra prediction modes.
[0569] In some embodiments, the second statistical unit 3903 is further configured to, when the second intra prediction mode is a candidate intra prediction mode in a third mode set, not count the first preset block; and wherein the third mode set is different from the first mode set, and the third mode set comprises at least one of the following: a mode of performing prediction on the copy of the intra block, a mode of performing prediction using an extrapolation filter, and a mode of performing prediction using matrix operation.
[0570] In some embodiments, the second statistical unit 3903 is further configured to, when the second intra prediction mode is a candidate intra prediction mode in a third mode set, determine a fourth intra prediction mode derived based on candidate samples of the first preset block; and perform accumulation operation on the fourth intra prediction mode to determine statistical results of the fourth intra prediction mode in the first mode set; and wherein the third mode set is different from the first mode set, and the third mode set comprises at least one of the following: a mode of performing prediction on the copy of the intra block, a mode of performing prediction using an extrapolation filter, and a mode of performing prediction using matrix operation.
[0571] In some embodiments, the second determination unit 3901 is further configured to, according to the statistical results, sort the candidate intra prediction modes in the first mode set in descending order, and determine at least two candidate intra prediction modes at the front of the sorting as the at least two intra prediction modes of the current block.
[0572] In some embodiments, the second prediction unit 3902 is further configured to, according to the statistical results corresponding to the at least two intra prediction modes, determine weight values of the at least two intra prediction modes respectively; and perform weighted prediction on the current block according to the at least two intra prediction modes and the weight values respectively to determine a prediction value of the current block.
[0573] In some embodiments, referring to FIG. 39, the decoder 390 can further include a decoding unit 3904 configured to decode the bitstream to determine a value of the first syntax element, determine the one or more preset blocks when the first syntax element indicates that the current block uses the first prediction mode, determine the statistical results of the at least two candidate intra prediction modes in the first mode set according to the prediction modes used by the one or more preset blocks respectively, determine the at least two intra prediction modes for the current block in the first mode set according to the statistical results, and predict the current block according to the at least two intra prediction modes to determine the prediction value of the current block.
[0574] In some embodiments, referring to FIG. 39, the decoder 390 can further include a second transform unit 3905, wherein: the second determining unit 3901 is further configured to determine a transform core for the current block; and the second transform unit 3905 is configured to transform the transform coefficients of the current block according to the transform core to determine a residual block of the current block; and the second determining unit 3901 is further configured to determine the reconstructed value of the current block according to the residual value of the current block and the prediction value of the current block.
[0575] In some embodiments, the decoding unit 3904 is further configured to decode the bitstream to determine a transform core group index of the current block; the second determining unit 3901 is further configured to determine a transform core group of the current block according to the transform core group index; and determine the transform core of the current block according to the transform core group.
[0576] In some embodiments, the decoding unit 3904 is further configured to decode the bitstream to determine a transform core index of the current block; and the second determining unit 3901 is further configured to determine the transform core of the current block according to the transform core group and the transform core index.
[0577] In some embodiments, the second determining unit 3901 is further configured to determine a first candidate list of the current block, the first candidate list indicating the at least two candidate transform core groups; and determine the transform core group of the current block according to the first candidate list and the transform core group index.
[0578] In some embodiments, when the first candidate list indicates that the transform cores included in the at least two candidate transform core groups, the decoding unit 3904 is further configured to decode the bitstream to determine a transform core index of the current block; and the second determining unit 3901 is further configured to determine the transform core of the current block according to the first candidate list and the transform core index.
[0579] It can be understood that, in the embodiments, the "unit" can be a partial circuit, a partial processor, a partial program or software, etc., and can also be a module, and can also be non-modular. Moreover, the components in the embodiments can be integrated in a processing unit, or can be physically present individually, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function module.
[0580] In still another embodiment of the present application, FIG. 40 is a specific hardware structure diagram of a decoder provided by the embodiments of the present application. As shown in FIG. 40, the decoder 390 can include a second communication interface 4001, a second memory 4002 and a second processor 4003; and the components are coupled together through a second bus system 4004. It can be understood that the second bus system 4004 is used to realize the connection communication between the components. The second bus system 4004 includes a data bus, a power supply bus, a control bus and a state signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the second bus system 4004 in FIG. 40. Among them,
[0581] The second communication interface 4001 is used for receiving and sending signals in the information receiving and sending process between other external network elements;
[0582] The second memory 4002 is used for storing a computer program capable of running on the second processor 4003;
[0583] The second processor 4003 is used for, when running the computer program, performing:
[0584] determining one or more decoded preset blocks; determining statistical results of at least two candidate intra-prediction modes in the first mode set according to the prediction modes used by the one or more preset blocks respectively; determining at least two intra-prediction modes of the current block in the first mode set according to the statistical results; and predicting the current block according to the at least two intra-prediction modes to determine the prediction value of the current block.
[0585] Optionally, as another embodiment, the second processor 4003 is further configured to, when running the computer program, perform the method in any one of the preceding embodiments.
[0586] It can be understood that the hardware function of the second memory 4002 is similar to that of the first memory 3802, and the hardware function of the second processor 4003 is similar to that of the first processor 3803; and details are not described here.
[0587] The embodiment provides a decoder, in which statistics of decoded related blocks are obtained, statistical results of the related blocks are obtained, and at least two intra prediction modes used by a current block are derived according to the statistical results; since it is not necessary to explicitly indicate which kind of intra prediction mode is used, the overhead in a code stream can be saved; moreover, weighted prediction is performed by using the at least two intra prediction modes, the prediction accuracy can be improved, and the prediction error can be reduced to a certain extent; in addition, the at least two intra prediction modes used by the current block are derived according to the statistical results, the intra prediction mode derivation by gradient calculation required by the DIMD mode can be avoided, the calculation complexity is reduced, the coding efficiency can be improved, and the compression performance is improved.
[0588] In still another embodiment of the present application, FIG. 41 is a schematic diagram of a composition structure of a coding system provided by the embodiment of the present application. As shown in FIG. 41, the coding system 410 can include an encoder 4101 and a decoder 4102.
[0589] In the embodiment of the present application, the encoder 4101 can be the encoder described in any one of the foregoing embodiments, and the decoder 4102 can be the decoder described in any one of the foregoing embodiments.
[0590] In some embodiments, the embodiment of the present application provides a computer readable storage medium having a computer program stored thereon. The computer program is executed by a processor (such as the first processor or the second processor) to implement the steps of the coding method as described in the foregoing embodiments.
[0591] In some embodiments, the embodiment of the present application provides a computer program product including a computer program or instructions. The computer program or instructions are executed by a processor (such as the first processor or the second processor) to implement the steps of the coding method as described in the foregoing embodiments.
[0592] In some embodiments, the embodiment of the present application provides a computer program which is executed by a processor (such as the first processor or the second processor) to implement the steps of the coding method as described in the foregoing embodiments.
[0593] In some embodiments, the embodiment of the present application further provides a computer readable storage medium having a code stream stored thereon. The code stream is generated by executing the steps of the encoding method as described in the foregoing embodiments.
[0594] In the embodiment of the present application, the to-be-encoded information in the encoding method can include at least one of the following: a transform coefficient of the current block, a transform kernel group index of the current block, a transform kernel index of the current block, a value of the first syntax element, a value of the second syntax element, and a value of the third syntax element. Here, the to-be-encoded information is encoded and written into the code stream.
[0595] In the embodiments of the present application, the first syntax element is used to indicate whether the current block uses the first prediction mode, the value of the second syntax element is used to indicate the transform core group index of the current block, the value of the third syntax element is used to indicate whether the current block uses the first transform mode, and the corresponding transform core index when the current block uses the first transform mode. For example, taking the third syntax element as an example, if the value of the third syntax element is 0, it is determined that the current block does not use the first transform mode; if the value of the second syntax element is non-zero, it is determined that the current block uses the first transform mode and the corresponding transform core index. For example, if the value of the second syntax element is 1, it is determined that the current block uses the first transform core; if the value of the second syntax element is 2, it is determined that the current block uses the second transform core, and so on, which is not limited here.
[0596] Those skilled in the art can clearly understand that the units and algorithm steps of the examples described in combination with the embodiments disclosed in the present application can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software mode depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to realize the described functions, but such implementation should not be considered beyond the scope of the present application.
[0597] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0598] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other form.
[0599] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0600] In addition, each of the functional units in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0601] The functions described above can be implemented in software, firmware, hardware, or any combination thereof. Moreover, the functions will be implemented in software as embodied in one or more computer programs that are executed over one or more processors that are included in one or more computing devices.
[0602] It should be noted that, in the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0603] The above-mentioned sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0604] The methods disclosed in the several method embodiments of the present application can be combined arbitrarily without conflict to obtain new method embodiments.
[0605] The features disclosed in the several product embodiments of the present application can be combined arbitrarily without conflict to obtain new product embodiments.
[0606] The features disclosed in the several method or device embodiments of the present application can be combined arbitrarily without conflict to obtain new method or device embodiments.
[0607] The above merely describes specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims. Industrial applicability
[0608] In the embodiments of the present application, at the encoding end or the decoding end, after one or more preset blocks are encoded / decoded, statistical results of at least two candidate intra-prediction modes in the first mode set are determined according to the prediction modes used by the one or more preset blocks; then at least two intra-prediction modes of the current block are determined in the first mode set according to the statistical results; and finally, the prediction value of the current block is determined by predicting the current block according to the at least two intra-prediction modes. In this way, the statistical results of the relevant blocks are obtained by using the relevant blocks for statistics, and the at least two intra-prediction modes used by the current block are derived according to the statistical results. Since it is not necessary to explicitly indicate which intra-prediction mode is used, the overhead in the code stream can be saved. Moreover, the accuracy of prediction can be improved by using the at least two intra-prediction modes for weighted prediction, and the prediction error can be reduced to a certain extent. In addition, the at least two intra-prediction modes used by the current block are derived according to the statistical results, which can avoid the intra-prediction mode derivation by gradient calculation required by the DIMD mode, and the computational complexity is reduced, thereby improving the coding efficiency and further improving the compression performance.
Claims
1. A decoding method applied to a decoder, the method comprising: Identify one or more pre-defined blocks that have been decoded; Based on the prediction modes used by each of the one or more preset blocks, determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set; Based on the statistical results, at least two intra-frame prediction modes for the current block are determined from the first mode set; The current block is predicted based on the at least two intra-frame prediction modes to determine the predicted value of the current block.
2. The method according to claim 1, wherein, The determination of one or more preset blocks that have been decoded includes: Determine at least one candidate block that has been decoded, wherein the candidate block includes: neighboring blocks of the current block, and / or, non-neighboring blocks of the current block; Based on the at least one candidate block, determine the one or more preset blocks.
3. The method according to claim 1, wherein, The determination of one or more preset blocks that have been decoded includes: A preset range is determined based on the current block; The one or more preset blocks are determined based on all decoded blocks within the preset range.
4. The method according to claim 1, wherein, The first set of modes includes at least one of the following: angle prediction mode, DC mode, and PLANA mode.
5. The method according to claim 1, wherein, The step of determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks includes: When the first preset block is an inter-frame prediction block, the first preset block is not counted. Wherein, the first preset block is any one of the one or more preset blocks.
6. The method according to claim 1, wherein, The step of determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks includes: When the first preset block is an inter-frame prediction block, a first intra-frame prediction mode derived from candidate samples based on the first preset block is determined. The statistical results of the first intra-frame prediction modes in the first mode set are determined by accumulating the first intra-frame prediction modes. Wherein, the first preset block is any one of the one or more preset blocks.
7. The method according to claim 6, wherein, The method further includes: When the first intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, the first intra-frame prediction mode is accumulated to determine the statistical result of the first intra-frame prediction mode in the first mode set.
8. The method according to claim 6, wherein, The method further includes: Based on at least a portion of the samples within the reconstructed block of the first preset block, candidate samples for the first preset block are determined; or, Candidate samples for the first preset block are determined based on at least a portion of the samples within the prediction block of the first preset block.
9. The method according to claim 6 or 7, wherein, The step of performing cumulative calculations on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set includes: The first intra-frame prediction mode is cumulatively calculated based on the size parameter value of the first preset block to determine the statistical result of the first intra-frame prediction mode in the first mode set.
10. The method according to claim 6 or 7, wherein, The step of performing cumulative calculations on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set includes: The first intra-frame prediction mode is incremented by one to accumulate the results and determine the statistical results of the first intra-frame prediction mode in the first mode set.
11. The method according to claim 1, wherein, The step of determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks includes: When the first preset block is an intra-prediction block, determine the second intra-prediction mode used by the first preset block; When the second intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, the second intra-frame prediction mode is accumulated to determine the statistical result of the second intra-frame prediction mode in the first mode set. Wherein, the first preset block is any one of the one or more preset blocks.
12. The method according to claim 11, wherein, The method further includes: When the second intra-frame prediction mode is a candidate intra-frame prediction mode in the second mode set, at least two third intra-frame prediction modes corresponding to the second intra-frame prediction mode are determined. When at least two third intra-frame prediction modes are at least two candidate intra-frame prediction modes in the first mode set, cumulative calculations are performed on the at least two third intra-frame prediction modes respectively to determine the at least two third intra-frame prediction modes in the first mode set. Statistical results of the pattern; The second mode set is different from the first mode set, and the second mode set includes modes that use at least two intra-frame prediction modes for combined prediction.
13. The method according to claim 11, wherein, The method further includes: When the second intra-frame prediction mode is a candidate intra-frame prediction mode in the third mode set, the first preset block is not counted; The third mode set is different from the first mode set, and the third mode set includes at least one of the following: a mode that performs prediction by copying intra-frame blocks, a mode that performs prediction by using an extrapolation filter, and a mode that performs prediction by using matrix operations.
14. The method according to claim 11, wherein, The method further includes: When the second intra-frame prediction mode is a candidate intra-frame prediction mode in the third mode set, a fourth intra-frame prediction mode derived from the candidate samples of the first preset block is determined. The statistical results of the fourth intra-frame prediction mode in the first mode set are determined by performing cumulative calculation on the fourth intra-frame prediction mode. The third mode set is different from the first mode set, and the third mode set includes at least one of the following: a mode that performs prediction by copying intra-frame blocks, a mode that performs prediction by using an extrapolation filter, and a mode that performs prediction by using matrix operations.
15. The method according to claim 1, wherein, The step of determining at least two intra-frame prediction modes for the current block in the first mode set based on the statistical results includes: Based on the statistical results, the candidate intra-prediction modes in the first mode set are sorted in descending order, and the top two candidate intra-prediction modes are determined as the at least two intra-prediction modes of the current block.
16. The method according to claim 15, wherein, The step of predicting the current block based on the at least two intra-frame prediction modes to determine the predicted value of the current block includes: Based on the statistical results corresponding to the at least two intra-frame prediction modes, determine the weight values of each of the at least two intra-frame prediction modes; The predicted value of the current block is determined by performing a weighted prediction on the current block based on the at least two intra-frame prediction modes and their respective weight values.
17. The method according to any one of claims 1 to 16, wherein, The method further includes: Decode the bitstream to determine the value of the first syntax element; When the first syntax element indicates that the current block uses a first prediction mode, the following steps are performed: determining one or more preset blocks that have been decoded; determining statistical results of at least two candidate intra-prediction modes in a first mode set based on the prediction modes used by each of the one or more preset blocks; determining at least two intra-prediction modes of the current block in the first mode set based on the statistical results; and predicting the current block based on the at least two intra-prediction modes to determine the predicted value of the current block.
18. The method according to any one of claims 1 to 16, wherein, The method further includes: Determine the transform kernel of the current block; The transformation coefficients of the current block are transformed according to the transformation kernel to determine the residual block of the current block; The reconstruction value of the current block is determined based on the residual value of the current block and the predicted value of the current block.
19. The method according to claim 18, wherein, Determining the transform kernel of the current block includes: Decode the bitstream and determine the transform kernel group index of the current block; The transform kernel group of the current block is determined based on the transform kernel group index; The transformation kernel of the current block is determined based on the transformation kernel set.
20. The method according to claim 19, wherein, Determining the transform kernel of the current block based on the transform kernel set includes: Decode the bitstream and determine the transform kernel index of the current block; The transform kernel of the current block is determined based on the transform kernel group and the transform kernel index.
21. The method according to claim 19, wherein, Determining the transform kernel group of the current block based on the transform kernel group index includes: Determine a first candidate list for the current block, the first candidate list indicating at least two candidate transform kernel groups; The transform kernel group of the current block is determined based on the first candidate list and the transform kernel group index.
22. The method according to claim 21, wherein, When the first candidate list indicates the transform kernels included in the at least two candidate transform kernel groups, determining the transform kernel of the current block includes: Decode the bitstream and determine the transform kernel index of the current block; The transform kernel of the current block is determined based on the first candidate list and the transform kernel index.
23. An encoding method applied to an encoder, the method comprising: Identify one or more pre-defined blocks that have already been encoded; Based on the prediction modes used by each of the one or more preset blocks, at least two candidate intra-frame prediction modes in the first mode set are determined. Statistical results of the formula; Based on the statistical results, at least two intra-frame prediction modes for the current block are determined from the first mode set; The current block is predicted based on the at least two intra-frame prediction modes to determine the predicted value of the current block.
24. The method according to claim 23, wherein, The method further includes: When the current block uses a first prediction mode, the steps of determining one or more pre-coded blocks; determining statistical results of at least two candidate intra-prediction modes in a first mode set based on the prediction modes used by each of the one or more pre-coded blocks; determining at least two intra-prediction modes of the current block in the first mode set based on the statistical results; and predicting the current block based on the at least two intra-prediction modes to determine the predicted value of the current block.
25. The method according to claim 24, wherein, The method further includes: Multiple candidate prediction modes are determined, wherein the multiple candidate prediction modes include at least the first prediction mode; The encoding cost of the current block is calculated based on the multiple candidate prediction modes, and the cost result corresponding to each of the multiple candidate prediction modes is determined. Determine the minimum cost result among the cost results corresponding to each of the multiple candidate prediction modes; Based on the candidate prediction pattern corresponding to the minimum cost result, determine whether the current block uses the first prediction pattern.
26. The method of claim 25, wherein, Determining whether the current block uses the first prediction mode based on the candidate prediction mode corresponding to the minimum cost result includes: When the candidate prediction mode corresponding to the minimum cost result is the first prediction mode, it is determined that the current block uses the first prediction mode; When the candidate prediction mode corresponding to the minimum cost result is not the first prediction mode, it is determined that the current block does not use the first prediction mode.
27. The method according to claim 25, wherein, The method further includes: Determine the value of the first syntax element, which is used to indicate whether the current block uses the first prediction mode; The value of the first syntax element is encoded, and the resulting encoded bits are written into the bitstream.
28. The method according to claim 23, wherein, The determination of one or more pre-defined encoded blocks includes: Determine at least one candidate block that has been encoded, wherein the candidate block includes: neighboring blocks of the current block, and / or, non-neighboring blocks of the current block; Based on the at least one candidate block, determine the one or more preset blocks.
29. The method according to claim 23, wherein, The determination of one or more pre-defined encoded blocks includes: A preset range is determined based on the current block; The one or more preset blocks are determined based on all encoded blocks within the preset range.
30. The method according to claim 23, wherein, The first set of modes includes at least one of the following: angle prediction mode, DC mode, and PLANA mode.
31. The method according to claim 23, wherein, The step of determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks includes: When the first preset block is an inter-frame prediction block, the first preset block is not counted. Wherein, the first preset block is any one of the one or more preset blocks.
32. The method according to claim 23, wherein, The step of determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks includes: When the first preset block is an inter-frame prediction block, a first intra-frame prediction mode derived from candidate samples based on the first preset block is determined. The first intra-frame prediction mode is accumulated to determine the statistical results of the first intra-frame prediction mode in the first mode set; Wherein, the first preset block is any one of the one or more preset blocks.
33. The method according to claim 32, wherein, The method further includes: When the first intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, the first intra-frame prediction mode is accumulated to determine the statistical result of the first intra-frame prediction mode in the first mode set.
34. The method according to claim 32, wherein, The method further includes: Based on at least a portion of the samples within the reconstructed block of the first preset block, candidate samples for the first preset block are determined; or, Candidate samples for the first preset block are determined based on at least a portion of the samples within the prediction block of the first preset block.
35. The method according to claim 32 or 33, wherein, The step of performing cumulative calculations on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set includes: The first intra-frame prediction mode is cumulatively calculated based on the size parameter value of the first preset block to determine the statistical result of the first intra-frame prediction mode in the first mode set.
36. The method according to claim 32 or 33, wherein, The step of performing cumulative calculations on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set includes: The first intra-frame prediction mode is incremented by one to accumulate the results and determine the statistical results of the first intra-frame prediction mode in the first mode set.
37. The method according to any one of claims 23 to 36, wherein, The step of determining at least two intra-frame prediction modes for the current block in the first mode set based on the statistical results includes: Based on the statistical results, the candidate intra-prediction modes in the first mode set are sorted in descending order, and the top two candidate intra-prediction modes are determined as the at least two intra-prediction modes of the current block.
38. The method according to claim 37, wherein, The step of predicting the current block based on the at least two intra-frame prediction modes to determine the predicted value of the current block includes: Based on the statistical results corresponding to the at least two intra-frame prediction modes, determine the weight values of each of the at least two intra-frame prediction modes; The predicted value of the current block is determined by performing a weighted prediction on the current block based on the at least two intra-frame prediction modes and their respective weight values.
39. The method according to any one of claims 23 to 38, wherein, The method further includes: Determine the transform kernel of the current block; The residual value of the current block is determined based on the predicted value of the current block, and the residual value of the current block is transformed according to the transformation kernel to determine the transformation coefficient of the current block; The transform coefficients of the current block are encoded, and the resulting encoded bits are written into the bitstream.
40. The method according to claim 39, wherein, Determining the transform kernel of the current block includes: Determine the transform kernel set of the current block; The transformation kernel of the current block is determined based on the transformation kernel set.
41. The method according to claim 40, wherein, Determining the transform kernel set of the current block includes: The encoding cost of the current block is calculated based on at least two candidate transform kernel groups, and the cost result corresponding to each of the at least two candidate transform kernel groups is determined. The minimum cost result is determined from the cost results corresponding to each of the at least two candidate transformation kernel groups, and the candidate transformation kernel group corresponding to the minimum cost result is determined as the transformation kernel group of the current block.
42. The method according to claim 41, wherein, The method further includes: Determine the transform kernel group index of the current block; wherein the transform kernel group index is used to indicate the number of the transform kernel group of the current block in the at least two candidate transform kernel groups; The transform kernel group index of the current block is encoded, and the resulting encoded bits are written into the bit stream.
43. The method according to claim 40, wherein, Determining the transform kernel of the current block based on the transform kernel set includes: Determine at least two candidate transform kernels included in the transform kernel group; The encoding cost of the current block is calculated based on the at least two candidate transform kernels, and the cost result corresponding to each of the at least two candidate transform kernels is determined. The minimum cost result is determined from the cost results corresponding to each of the at least two candidate transformation kernels, and the candidate transformation kernel corresponding to the minimum cost result is determined as the transformation kernel of the current block.
44. The method according to claim 43, wherein, The method further includes: Determine the transform kernel index of the current block, wherein the transform kernel index is used to indicate the number of the transform kernel of the current block in the transform kernel group; The transform kernel index of the current block is encoded, and the resulting encoded bits are written into the bitstream.
45. A bitstream, wherein, The bitstream is generated by bit encoding according to the encoding method described in any one of claims 23 to 44; wherein the information to be encoded in the encoding method includes at least one of the following: The transform coefficients of the current block, the transform kernel group index of the current block, the transform kernel index of the current block, and the value of the first syntax element, wherein the first syntax element is used to indicate whether the current block uses the first prediction mode.
46. An encoder, the encoder comprising a first determining unit and a first predicting unit, wherein: The first determining unit is configured to determine one or more pre-coded blocks; and to determine statistical results of at least two candidate intra-frame prediction modes in a first mode set based on the prediction modes used by each of the one or more pre-coded blocks. The first determining unit is further configured to determine at least two intra-frame prediction modes for the current block in the first mode set based on the statistical results. The first prediction unit is configured to predict the current block based on the at least two intra-frame prediction modes and determine the predicted value of the current block.
47. An encoder, the encoder comprising a first memory and a first processor, wherein: The first memory is used to store computer programs that can run on the first processor; The first processor is configured to, when running the computer program, execute the steps of the encoding method as described in any one of claims 23 to 44.
48. A decoder, the decoder comprising a second determining unit and a second predicting unit, wherein: The second determining unit is configured to determine one or more preset blocks that have been decoded; And based on the prediction modes used by each of the one or more preset blocks, determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set; The second determining unit is further configured to determine at least two intra-frame prediction modes for the current block in the first mode set based on the statistical results. The second prediction unit is configured to predict the current block based on the at least two intra-frame prediction modes and determine the predicted value of the current block.
49. A decoder, the decoder comprising a second memory and a second processor, wherein: The second memory is used to store computer programs that can run on the second processor; The second processor is configured to execute the steps of the decoding method as described in any one of claims 1 to 22 when running the computer program.
50. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the decoding method as described in any one of claims 1 to 22, or the steps of the encoding method as described in any one of claims 23 to 44.
51. A computer-readable storage medium having a bitstream stored thereon, wherein, The bitstream is generated by performing the steps of the encoding method as described in any one of claims 23 to 44.