Coding method, decoding method, code stream, coder, decoder and storage medium

ZA202608025APending Publication Date: 2026-08-26GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
ZA202608025
Authority / Receiving Office
ZA · ZA
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-08-06
Publication Date
2026-08-26

AI Technical Summary

Technical Problem

The existing video encoding standards are incompletely considered when dealing with complex intra prediction modes, resulting in low compression efficiency.

Method used

When determining that the prediction mode of the current block meets specific conditions, at least two candidate transformation core groups are used to determine the transformation cores to improve the accuracy of the transformation.

Benefits of technology

By using multiple candidate transformation core groups, the compression efficiency of video encoding and decoding is improved and the encoding and decoding performance is improved.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

NOT VISIBLE DUE TO STATUS OF PATENT
Need to check novelty before this filing date? Find Prior Art

Description

Coding and decoding method, code stream, encoder, decoder and storage medium Technical Field

[0001] The embodiments of the present application relate to the field of video coding and decoding technology, and in particular to a coding and decoding method, a bit stream, an encoder, a decoder, and a storage medium. Background Art

[0002] As demand for video display quality increases, high-resolution video, such as HD and UHD, has emerged. However, high-resolution video typically contains more information and therefore requires more bandwidth. To reduce bandwidth requirements, video coding standards involving video compression have been introduced.

[0003] In video coding standards, each intra-frame prediction mode corresponds to a transform kernel group, so a texture feature index used to determine the transform kernel group can be derived for each block. However, for more complex intra-frame prediction modes, these modes can perform weighted calculations on the prediction values ​​of two or more intra-frame prediction modes. For blocks that require weighted prediction values ​​from two or more intra-frame prediction modes, the current transformation process does not fully consider these factors, resulting in low compression efficiency.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a coding and decoding method, a code stream, an encoder, a decoder, and a storage medium, which can improve compression efficiency and thereby enhance coding and decoding performance.

[0006] The technical solution of the embodiment of the present application can be implemented as follows:

[0007] In a first aspect, an embodiment of the present application provides a decoding method, applied to a decoder, the method comprising:

[0008] Determine the prediction mode of the current block;

[0009] When the prediction mode of the current block satisfies a first condition, determining a first candidate list of the current block; wherein the first candidate list indicates at least two candidate transform core groups;

[0010] Determine a transformation kernel for the current block according to the first candidate list;

[0011] A transform coefficient of the current block is determined, and the transform coefficient of the current block is inversely transformed according to the transform kernel to determine a residual block of the current block.

[0012] In a second aspect, an embodiment of the present application provides an encoding method, applied to an encoder, the method comprising:

[0013] Determine the prediction mode of the current block;

[0014] When the prediction mode of the current block satisfies a first condition, determining a first candidate list of the current block; wherein the first candidate list indicates at least two candidate transform core groups;

[0015] Determine a transformation kernel for the current block according to the first candidate list;

[0016] Determine a residual block of the current block, and transform the residual block of the current block according to the transformation kernel to determine a transformation coefficient of the current block;

[0017] The transform coefficients of the current block are coded and the resulting coded bits are written into the bitstream.

[0018] In a third aspect, an embodiment of the present application provides a code stream, which is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following:

[0019] The quantization coefficient of the current block, the transform core index number of the current block, the feature index number of the current block, the prediction mode of the current block and the value of the first syntax element; wherein the first syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform core index number used.

[0020] In a fourth aspect, an embodiment of the present application provides an encoder, comprising a first determining unit, a first transforming unit, and an encoding unit, wherein:

[0021] A first determining unit is configured to determine a prediction mode of a current block; when the prediction mode of the current block satisfies a first condition, determine a first candidate list for the current block; wherein the first candidate list indicates at least two candidate transform core groups;

[0022] The first determining unit is further configured to determine a transform kernel of the current block according to the first candidate list;

[0023] a first transform unit configured to determine a residual block of a current block, and transform the residual block of the current block according to a transform kernel to determine a transform coefficient of the current block;

[0024] The encoding unit is configured to perform encoding processing on the transformation coefficients of the current block and write the obtained encoding bits into the bit stream.

[0025] In a fifth aspect, an embodiment of the present application provides an encoder, comprising a first memory and a first processor, wherein:

[0026] a first memory for storing a computer program capable of running on the first processor;

[0027] The first processor is configured to execute the method according to the second aspect when running the computer program.

[0028] In a sixth aspect, an embodiment of the present application provides a decoder, comprising a second determination unit and a second transformation unit, wherein:

[0029] A second determining unit is configured to determine a prediction mode of the current block; when the prediction mode of the current block satisfies a first condition, determine a first candidate list for the current block; wherein the first candidate list indicates at least two candidate transform core groups;

[0030] The second determining unit is further configured to determine a transform kernel of the current block according to the first candidate list;

[0031] The second transform unit is configured to determine a transform coefficient of the current block, and perform an inverse transform on the transform coefficient of the current block according to the transform kernel to determine a residual block of the current block.

[0032] In a seventh aspect, an embodiment of the present application provides a decoder, comprising a second memory and a second processor, wherein:

[0033] a second memory for storing a computer program capable of running on the second processor;

[0034] The second processor is configured to execute the method according to the first aspect when running the computer program.

[0035] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by at least one processor, implements the method described in the first aspect or the method described in the second aspect.

[0036] The embodiments of the present application provide a coding and decoding method, a code stream, an encoder, a decoder, and a storage medium. At the encoding end, a prediction mode of the current block is determined; when the prediction mode of the current block meets a first condition, a first candidate list of the current block is determined; wherein the first candidate list indicates at least two candidate transform core groups; based on the first candidate list, a transform core of the current block is determined; a residual block of the current block is determined, and the residual block of the current block is transformed according to the transform core to determine the transform coefficient of the current block; the transform coefficient of the current block is encoded and the obtained coded bits are written into the code stream. At the decoding end, a prediction mode of the current block is determined; when the prediction mode of the current block meets a first condition, a first candidate list of the current block is determined; wherein the first candidate list indicates at least two candidate transform core groups; based on the first candidate list, a transform core of the current block is determined; the transform coefficient of the current block is determined, and the transform coefficient of the current block is inversely transformed according to the transform core to determine the residual block of the current block. In this way, both the encoding and decoding ends first determine the prediction mode for the current block. When the prediction mode for the current block meets the first condition, a first candidate list indicating at least two candidate transform core groups is determined, and then the transform core for the current block is determined based on these at least two candidate transform core groups. In other words, for a current block predicted using certain intra-frame prediction modes, the transform core for the current block is no longer determined based on just one transform core group, but rather based on at least two candidate transform core groups. This improves the accuracy of the transform, thereby increasing compression efficiency and, in turn, enhancing encoding and decoding performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] FIG1 is a flow chart diagram of a hybrid coding framework;

[0038] FIG2 is a schematic diagram of template matching of a current block;

[0039] FIG3 is a schematic diagram of reference pixels of a current block;

[0040] FIG4 is a schematic diagram of multiple reference rows of a current block;

[0041] FIG5 is a schematic diagram of multiple prediction modes corresponding to an intra-frame prediction;

[0042] FIG6 is a second schematic diagram of multiple prediction modes corresponding to an intra-frame prediction;

[0043] FIG7 is a third schematic diagram of multiple prediction modes corresponding to an intra-frame prediction;

[0044] FIG8 is a fourth schematic diagram of multiple prediction modes corresponding to an intra-frame prediction;

[0045] FIG9 is a schematic diagram of encoding of screen content;

[0046] FIG10 is a schematic diagram of a prediction process of a MIP mode;

[0047] FIG11 is a schematic diagram of a template of a current block and a template reference area;

[0048] FIG12 is a schematic diagram of a histogram of gradient and intra prediction mode;

[0049] FIG13 is a schematic diagram of weighted fusion of three intra-frame prediction modes;

[0050] FIG14 is a schematic diagram of the search range of the ITMP mode;

[0051] FIG15 is a schematic diagram of weights of various modes under a GPM mode;

[0052] FIG16 is a schematic diagram of a DCT transformation;

[0053] FIG17 is a schematic diagram of a base image of a DCT transformation;

[0054] FIG18 is a schematic diagram of a process flow without LFNST transformation;

[0055] FIG19 is a schematic diagram of a process flow with LFNST transformation;

[0056] FIG20 is a detailed flowchart of a LFNST transformation;

[0057] FIG21 is a schematic diagram of a base image of multiple transformation kernel groups;

[0058] FIG22 is a schematic diagram of a base image of NSPT transformation;

[0059] FIG23 is a schematic diagram of a network architecture of a video codec provided in an embodiment of the present application;

[0060] FIG24 is a schematic block diagram of a system composition of an encoder provided in an embodiment of the present application;

[0061] FIG25 is a schematic block diagram of a system composition of a decoder provided in an embodiment of the present application;

[0062] FIG26 is a flowchart diagram 1 of a decoding method provided in an embodiment of the present application;

[0063] FIG27 is a second flow chart of a decoding method provided in an embodiment of the present application;

[0064] FIG28 is a third flow chart of a decoding method provided in an embodiment of the present application;

[0065] FIG29 is a fourth flow chart of a decoding method provided in an embodiment of the present application;

[0066] FIG30 is a fifth flow chart of a decoding method provided in an embodiment of the present application;

[0067] FIG31 is a flowchart diagram 1 of an encoding method provided in an embodiment of the present application;

[0068] FIG32 is a second flow chart of an encoding method provided in an embodiment of the present application;

[0069] FIG33 is a schematic diagram of the composition structure of an encoder provided in an embodiment of the present application;

[0070] FIG34 is a schematic diagram of the hardware structure of an encoder provided in an embodiment of the present application;

[0071] FIG35 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application;

[0072] FIG36 is a schematic diagram of the hardware structure of a decoder provided in an embodiment of the present application;

[0073] Figure 37 is a schematic diagram of the composition structure of a coding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0074] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.

[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0076] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0077] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0078] In video images, a coding block (CB) is generally represented by a first color component, a second color component, and a third color component. These three color components are a luminance component, a blue chrominance component, and a red chrominance component. Specifically, the luminance component is typically represented by the symbol Y, the blue chrominance component is typically represented by the symbols Cb or U, and the red chrominance component is typically represented by the symbols Cr or V. Thus, video images can be represented in either the YCbCr or YUV format.

[0079] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained first. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations:

[0080] H.265 / High Efficiency Video Coding (HEVC);

[0081] H.266 / Versatile Video Coding (VVC);

[0082] VVC Test Model (VTM), a reference software testing platform for VVC;

[0083] A platform that improves compression performance after VVC (Enhanced Compression Model, ECM);

[0084] Joint Video Experts Team (JVET);

[0085] Coding Unit (CU);

[0086] Coding Tree Unit (CTU);

[0087] Largest Coding Unit (LCU);

[0088] Motion Vector (MV);

[0089] Prediction Unit (PU);

[0090] Transform Unit (TU);

[0091] Fusion technology (Merge);

[0092] Skip technology (Skip);

[0093] Quantization Parameter (QP);

[0094] Merge with Motion Vector Difference (MMVD)

[0095] Motion Vector Prediction (MVP);

[0096] Temporal Motion Vector Prediction (TMVP);

[0097] Subblock-based Temporal Motion Vector Prediction (SbTMVP);

[0098] Discrete Cosine Transform (DCT);

[0099] Discrete Sine Transform (DST);

[0100] Multiple Transform Selection (MTS);

[0101] Low Frequency Non-Separable Transform (LFNST);

[0102] Non-Separable Primary Transform (NSPT);

[0103] Context-based Adaptive Binary Arithmetic Coding (CABAC).

[0104] Currently, common video codec standards all adopt a block-based hybrid coding framework. Each image, sub-image, or frame in a video is divided into square maximum coding units (LCUs) or coding tree units (CTUs) of the same size (e.g., 256×256, 128×128, 64×64, etc.). Each LCU or CTU can be divided into rectangular CUs according to a set of rules. Coding units may also be divided into prediction units (PUs) and transform units (TUs). Specifically, as shown in Figure 1, the hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, inverse quantization, inverse transform, and in-loop filtering. The prediction module can include intra-frame prediction and inter-frame prediction, and inter-frame prediction can include motion estimation and motion compensation. Because adjacent pixels within a video image are highly correlated, intra-frame prediction is used in video coding and decoding to eliminate spatial redundancy between adjacent pixels. Furthermore, because adjacent images in a video image have strong similarities, inter-image prediction is used in video coding and decoding to eliminate temporal redundancy between adjacent images, thereby improving coding efficiency.

[0105] The basic process of a video codec is as follows: On the encoder side, an image is divided into blocks. Intra-frame prediction or inter-frame prediction is used on the current block to generate a prediction block for the current block. The prediction block is subtracted from the initial block of the current block to obtain a residual block. The residual block is transformed and quantized to obtain a quantization coefficient matrix. This quantization coefficient matrix is ​​entropy coded and output to the bitstream. On the decoder side, intra-frame prediction or inter-frame prediction is used on the current block to generate a prediction block for the current block. The bitstream is then parsed to obtain a quantization coefficient matrix. This quantization coefficient matrix is ​​inversely quantized and inversely transformed to obtain a residual block. The prediction block and residual block are added together to obtain a reconstructed block. The reconstructed blocks form a reconstructed image, which is then subjected to image-based or block-based loop filtering to obtain a decoded image. The encoder side also performs similar operations to the decoder side to obtain a decoded image. The decoded image can serve as a reference image for inter-frame prediction of subsequent images. Block division information, prediction, transform, quantization, entropy coding, loop filtering, and other mode or parameter information determined by the encoder are output to the bitstream if necessary. The decoding end determines the same block division information as the encoding end by parsing the bit stream and analyzing the existing information, as well as the mode information or parameter information such as prediction, transformation, quantization, entropy coding, and loop filtering, thereby ensuring that the decoded image obtained by the encoding end is the same as the decoded image obtained by the decoding end. The decoded image obtained by the encoding end is also usually called a reconstructed image. The current block can be divided into prediction units during prediction, and the current block can be divided into transformation units during transformation. The division of prediction units and transformation units can be different. The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. The embodiment of the present application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.

[0106] In addition, in the embodiments of the present application, the current block (CB) can be the current coding unit, the current prediction unit, or the current transform unit. Due to the need for parallel processing, the image can be divided into slices, etc. Slices in the same image can be processed in parallel, that is, there is no data dependency between them. "Frame" is a commonly used term, and it can generally be understood that a frame is an image. The frame described in the embodiments of the present application can also be replaced by an image or a slice, etc.

[0107] The following is a detailed introduction to the relevant solutions of prediction technology.

[0108] (1) Inter-frame prediction.

[0109] The Template Matching (TM) method was first used in inter-frame prediction. It exploits the correlation between adjacent pixels and uses areas surrounding the current block as templates. When the current block is encoded or decoded, its left and upper sides have already been encoded and decoded according to the coding order. Of course, existing hardware decoder implementations cannot guarantee that the left and upper sides have already been decoded when decoding of the current block begins. This refers to inter-frame blocks. For example, in HEVC, the prediction process for inter-frame coded blocks does not require surrounding reconstructed pixels, allowing the prediction process for inter-frame blocks to proceed in parallel. However, intra-frame coded blocks must use reconstructed pixels on the left and upper sides as reference pixels. Theoretically, the left and upper sides are available, meaning that they can be implemented with appropriate hardware design adjustments. In contrast, the right and lower sides are not available under the coding order of current standards such as VVC.

[0110] As shown in Figure 2, the rectangular areas to the left and above the current block are used as templates. The height of the left template portion is generally the same as the height of the current block, and the width of the upper template portion is generally the same as the width of the current block, but they can also be different. The best matching position of the template is found in the reference image to determine the motion information, or motion vector, of the current block. This process can be roughly described as starting from a starting position in a reference image (Ref0) and searching within a certain range around it. Search rules, such as the search range and search step size, can be predefined. At each position, the degree of match between the template corresponding to that position and the templates surrounding the current block is calculated. The degree of match can be measured using distortion metrics such as the sum of absolute differences (SAD), the sum of absolute transformed differences (SATD), and the mean-square error (MSE). SATD generally uses the Hadamard transform. Smaller values ​​of SAD, SATD, and MSE indicate a higher degree of match. The cost is calculated using the predicted block of the template corresponding to the position and the reconstructed blocks of the template surrounding the current block. In addition to searching at whole-pixel locations, sub-pixel locations can also be searched, with the motion information of the current block determined based on the location with the highest degree of match. By leveraging the correlation between adjacent pixels, the appropriate motion information for the template may also be appropriate for the current block. Of course, template matching may not be applicable to all blocks, so methods can be used to determine whether template matching should be used for the current block, such as using a control switch to indicate whether template matching should be used for the current block. This template matching method is called decoder-side motion vector derivation (DMVD). Both the encoder and decoder can use the template to search to derive motion information or find better motion information based on the existing motion information. This method does not require the transmission of specific motion vectors or motion vector differences. Instead, both the encoder and decoder perform the same search rules to ensure consistent encoding and decoding. Template matching can improve compression performance, but it requires a "search" at the decoder, which introduces a certain degree of decoding complexity.

[0111] (2) Intra-frame prediction.

[0112] It can be understood that there is a strong spatial correlation between adjacent parts or adjacent pixels within the image. Intra-frame prediction is a prediction method that uses the spatial correlation between the coded and decoded pixels around the current block and the pixels within the current block. For example, as shown in Figure 3, the 4×4 white filling pixels are the current block, and the grid filling pixels in the left column and the upper row of the current block are the reference pixels of the current block. Intra-frame prediction uses these reference pixels to predict the current block. These reference pixels may all be available, that is, all have been coded and decoded. Some may also be unavailable. For example, if the current block is the leftmost of the entire frame, then the reference pixels on the left side of the current block are unavailable. Or when encoding and decoding the current block, the lower left part of the current block has not been encoded and decoded, then the reference pixels on the lower left are also unavailable. In the case where reference pixels are unavailable, available reference pixels or certain values ​​or methods can be used for filling, or no filling can be performed.

[0113] The multiple reference line (MRL) intra prediction method can use more reference pixels to improve coding efficiency. As shown in Figure 4, there is a schematic diagram of using four reference rows / columns.

[0114] There are multiple prediction modes for intra-frame prediction, as shown in Figure 5. Here are the nine modes for intra-frame prediction for 4×4 blocks in H.264. Among them, mode 0 (vertical mode) copies the pixels above the current block vertically to the current block as the prediction value, mode 1 (horizontal mode) copies the reference pixels on the left side horizontally to the current block as the prediction value, mode 2 (DC mode) uses the average of the eight points A to D and I to L as the prediction value for all points, and modes 3 to 8 copy the reference pixels to the corresponding positions of the current block at a certain angle. Because some positions of the current block cannot correspond exactly to the reference pixels, it may be necessary to use the weighted average of the reference pixels, or the sub-pixels of the interpolated reference pixels.

[0115] In addition, there are modes such as PLANE and PLANAR. With technological advancements and the expansion of block sizes, the number of angular prediction modes is also increasing. For example, HEVC uses intra-frame prediction modes including PLANAR, DC, and 33 angular modes, for a total of 35 prediction modes (see Figure 6 for details). VVC uses intra-frame prediction modes including PLANAR, DC, and 65 angular modes, for a total of 67 prediction modes (see Figure 7 for details). Of course, in addition to the 67 modes mentioned above, VVC also provides wide-angle modes for rectangular blocks with a large difference between length and width. These modes, as indicated by the dashed lines in Figure 8, represent the ranges -14 to -1 and 67 to 80 degrees. These modes replace some conventional modes (see Figure 8 for details).

[0116] It can also be understood that intra block copy (IBC) can significantly improve the compression efficiency of screen content coding (SCC), and therefore IBC has been used for screen content coding from HEVC to VVC. Screen content is different from camera-captured content. It is computer-generated, noise-free, and contains text, computer graphics, and has clear boundaries. Screen content contains a large amount of repeated content, as shown in Figure 9.

[0117] In the embodiments of this application, IBC can be considered to apply the inter-frame prediction method to intra-frame prediction. Inter-frame prediction uses a reference block in a reference image, which is not the current image, to generate the prediction block for the current block. IBC, on the other hand, uses a reference block from the coded or reconstructed portion of the current image to generate the prediction block for the current block. IBC is also known as intra picture block compensation or current picture referencing (CPR).

[0118] IBC uses a block vector (BV) to represent the position difference between the current block and the reference block, similar to the MV used in inter-frame prediction. The encoder uses block matching within the search range to determine the best matching block for the current block and encodes the BV. There are various methods for encoding the BV, such as the merge mode, which is similar to inter-frame prediction and will not be discussed here.

[0119] IBC can be considered an intra-frame prediction method, or a different type of prediction method independent of intra-frame and inter-frame prediction. IBC is highly efficient for encoding screen content and can also improve compression efficiency in natural sequences captured by cameras.

[0120] It can also be understood that for matrix-based intra prediction (MIP), this is a special intra prediction mode, which may also be called Matrix weighted Intra Prediction in some places.

[0121] As shown in Figure 10, to predict a block of width W and height H, the MIP requires H reconstructed pixels in the column to the left of the current block and W reconstructed pixels in the row above the current block as input. MIP generates the prediction block using the following three steps: (a) reference pixel averaging, (b) matrix vector multiplication, and (c) interpolation. Here, the core of MIP is considered to be matrix multiplication. It can be thought of as a process of generating a prediction block from input pixels (reference pixels) using a matrix multiplication. MIP provides a variety of matrices, and different prediction methods are reflected in different matrices. Using different matrices for the same input pixels will produce different results. The reference pixel averaging and interpolation process is a design that compromises performance and complexity. For larger blocks, reference pixel averaging can achieve an effect similar to downsampling, allowing the input to fit into a smaller matrix, while interpolation achieves an upsampling effect. This eliminates the need to provide MIP matrices for every block size; instead, matrices of one or a few specific sizes are sufficient. As the demand for compression performance increases and hardware capabilities improve, more complex MIPs may appear in the next generation of standards.

[0122] MIP is somewhat similar to PLANAR, but it is obviously more complex and more flexible than PLANAR.

[0123] It can also be understood that for template-based intra mode derivation (TIMD), as shown in Figure 11, an area to the left and above the current block is used as a template. Except for edge cases, when encoding and decoding the current block, reconstructed values ​​can theoretically be obtained for the left and above areas of the current block. This is also the basis of many template adaptation methods. TIMD uses the diagonal filled area shown in Figure 11 as the template, and the template reference area in Figure 11 is the template's reference pixels (grid-filled area). The decoder can use a specific intra prediction mode to make predictions on the template and compare the predicted value with the reconstructed value to obtain the cost of the intra prediction mode on the template. Examples include SAD, SATD, and SSE. Since the template and the current block are adjacent and correlated, the performance of a prediction mode on the template can be used to estimate its performance on the current block. TIMD predicts some candidate intra-frame prediction modes on the template, obtains their costs on the template, and selects one or two intra-frame prediction modes with the lowest costs as the intra-frame prediction values ​​of the current block.

[0124] Research has found that if the cost difference between two intra-frame prediction modes on a template is small, taking a weighted average of the prediction values ​​of the two intra-frame prediction modes can improve compression performance. The weight of the prediction values ​​of the two prediction modes is related to the aforementioned cost, and in the current version, this weight is inversely proportional to the cost.

[0125] In summary, TIMD uses the prediction performance of intra-frame prediction modes on a template to select the appropriate intra-frame prediction mode, and can weight the two intra-frame prediction modes based on their cost on the template. The advantage of TIMD is that if the current block selects TIMD mode, the decoder does not need to indicate the specific intra-frame prediction mode to be used. Instead, the decoder derives the selected intra-frame prediction mode through the above process, which reduces the overhead to a certain extent.

[0126] It can also be understood that for decoder-side intra mode derivation (DIMD), DIMD uses the reconstructed pixels on the left and top sides of the current block to derive the prediction mode, but it does not predict on the template, but analyzes the gradient of the reconstructed pixels.

[0127] As shown in Figure 12, DIMD analyzes the gradient of black points, such as horizontal gradient and vertical gradient, and adapts an intra-frame prediction mode according to its gradient. The analysis of all points that need to be checked can obtain a result similar to the bar graph below. That is, the statistics of the number of points matched by each intra-frame prediction mode. Of course, the so-called bar graph is just to help understanding, and it can be implemented in a variety of simple forms. The current DIMD selects the two highest intra-frame prediction modes in the histogram, plus the PLANAR mode, and the prediction values ​​of a total of three intra-frame prediction modes are weighted. The weights are related to the results of the analysis. For example, as shown in Figure 13, the three intra-frame prediction modes include M1 mode, M2 mode and PLANAR mode. The prediction values ​​obtained for these three intra-frame prediction modes are set to Pred1, Pred2, and Pred3 respectively, and the weight values ​​of these three intra-frame prediction modes are set to w1, w2, and w3 respectively. The specific calculation formulas are as follows:

[0128] The final prediction block can be shown as follows:

[0129] In summary, DIMD uses gradient analysis of reconstructed pixels to select intra prediction modes, and can weight two intra prediction modes plus planar based on the analysis results. The advantage of DIMD is that if DIMD mode is selected for the current block, the decoder does not need to indicate the specific intra prediction mode to be used. Instead, the decoder can derive the selected mode through the above process, which saves a certain amount of overhead.

[0130] It's also understandable that Template Matching (TM) was first used in inter-frame prediction. It leverages the correlation between adjacent pixels and uses areas surrounding the current block as templates. When the current block is encoded or decoded, its left and upper sides are already encoded and decoded according to the coding order. Of course, existing hardware decoder implementations don't guarantee that the left and upper sides are already decoded when decoding begins. This refers to inter-frame blocks. For example, in HEVC, the prediction process for inter-frame coded blocks doesn't require surrounding reconstructed pixels, allowing the prediction process for inter-frame blocks to proceed in parallel. However, intra-frame coded blocks do require reconstructed pixels on the left and upper sides as reference pixels. Theoretically, the left and upper sides are available, meaning they can be implemented with appropriate hardware design adjustments. In contrast, the right and lower sides are unavailable under current coding standards like VVC.

[0131] As shown in Figure 2, the rectangular areas to the left and above the current block are used as templates. The height of the left template portion is generally the same as the height of the current block, and the width of the upper template portion is generally the same as the width of the current block, but can also be different. The best matching position of the template is found in the reference image to determine the motion information, or motion vector, of the current block. This process can be roughly described as starting from a starting position in a reference image (Ref0) and searching within a certain range around it. Search rules, such as the search range and search step size, can be predefined. At each position, the degree of match between the template corresponding to that position and the templates surrounding the current block is calculated. The degree of match can be measured using distortion costs, such as SAD or SATD. SATD generally uses transforms such as the Hadamard transform and MSE. Lower values ​​of SAD, SATD, and MSE indicate a higher degree of match. The cost is calculated using the predicted block of the template corresponding to that position and the reconstructed block of the template surrounding the current block. In addition to searching at integer pixel positions, sub-pixel positions can also be searched, with the motion information of the current block determined based on the position with the highest degree of match. By utilizing the correlation between adjacent pixels, the motion information that is appropriate for the template may also be appropriate for the current block. Of course, the template matching method may not necessarily be applicable to all blocks, so some methods can be used to determine whether the current block uses the above template matching method, such as using a control switch in the current block to indicate whether the template matching method is used. A classic template matching technology is called DMVD (Decoder side Motion Vector Derivation). Both the encoder and decoder can use the template to search to derive motion information or find better motion information based on the original motion information. It does not require the transmission of specific motion vectors or motion vector differences. Instead, both the encoder and decoder perform the same search rules to ensure consistency in encoding and decoding. The template matching method can improve compression performance, but it also requires "searching" in the decoder, which brings a certain degree of decoder complexity.

[0132] It's also understandable that intra-template matching prediction (ITMP) can be considered a combination of IBC and TM. As mentioned above, applying TM to inter-frames can reduce the overhead of encoding MVs. Similarly, using TM with IBC can reduce the overhead of encoding BVs. An example is to directly use the matching block found by TM as the ITMP prediction block for the current block, without encoding BVs.

[0133] Figure 14 shows an example of ITMP. The inverted L-shaped area in the upper left corner of the current block is used as a template. The search is performed within a point-filled search range, which covers the reconstructed area. The point-filled area shown in Figure 14 includes the current CTU (R1), the CTU to the upper left of R2, the CTU above R3, and the CTU to the left of R4. This is just an example; the search range may vary in actual applications. In this example, the best matching block is found in R2.

[0134] It can also be understood that for Spatial Geometric Partitioning Mode (SGPM), the VVC video codec standard has an inter-frame prediction mode called Geometric Partitioning Mode (GPM). The AVS3 video codec standard has an inter-frame prediction mode called Angular Weighted Prediction (AWP). Although these two modes have different names and specific implementations, they share the same principles.

[0135] Traditional unidirectional prediction only uses one reference block of the same size as the current block, while traditional bidirectional prediction uses two reference blocks of the same size. The pixel value of each point in the predicted block is the average of the corresponding positions in the two reference blocks, meaning that all points in each reference block account for 50% of the total pixel value. Bidirectional weighted prediction allows for different ratios in the two reference blocks, such as 75% for all points in the first reference block and 25% for all points in the second reference block. However, all points in the same reference block have the same ratio. Other optimization methods, such as decoder-side motion vector refinement (DMVR) and bidirectional optical flow (BIO), can cause some changes in reference or predicted pixels, but are unrelated to the principles described above. BIO can also be abbreviated as BDOF. GPM or AWP also uses two reference blocks of the same size as the current block, but some pixel locations use 100% of the pixel values ​​from the first reference block, while others use 100% of the pixel values ​​from the second reference block. In the boundary area, or transition zone, pixel values ​​from both reference blocks are used in a certain proportion. The weights in the boundary area also transition gradually. The specific distribution of these weights is determined by the GPM or AWP mode. The weight of each pixel position is determined based on the GPM or AWP mode. Of course, in some cases, such as very small block sizes, some GPM or AWP modes may not guarantee that some pixel locations will use 100% of the pixel values ​​from the first reference block, while others will use 100% of the pixel values ​​from the second reference block. Alternatively, GPM or AWP uses two reference blocks of different sizes, taking a desired portion of each as the reference block. Specifically, the portion with a non-zero weight is used as the reference block, while the portion with a zero weight is discarded.

[0136] Figure 15 shows the weights of the 64 GPM modes in VVC on square blocks. Black indicates a weight of 0% for the position corresponding to the first reference block, white indicates a weight of 100%, and gray areas, depending on the depth of the color, indicate a weight greater than 0% and less than 100% for the position corresponding to the first reference block. The weight for the position corresponding to the second reference block is 100% minus the weight for the position corresponding to the first reference block.

[0137] GPM and AWP use different weighting methods. GPM determines the angle and offset for each mode and then calculates a weight matrix for each mode. AWP first creates a one-dimensional weight line and then uses a method similar to intra-frame angle prediction to fill the entire matrix with this one-dimensional weight line.

[0138] It should be noted that earlier coding standards only had rectangular division methods, whether it was CU, PU or TU division. GPM and AWP achieved the predicted non-rectangular division effect without division. GPM and AWP use a mask of the weights of two reference blocks, that is, the weight map or weight matrix mentioned above. This mask determines the weights of the two reference blocks when generating the prediction block, or it can be simply understood that part of the position of the prediction block comes from the first reference block and part of the position comes from the second reference block, and the transition area (blending area) is weighted by the corresponding positions of the two reference blocks, so that the transition is smoother. GPM and AWP do not divide the current block into two CUs or PUs according to the dividing line, so the transformation, quantization, inverse transformation, inverse quantization, etc. of the residual after prediction also treat the current block as a whole.

[0139] It should also be noted that GPM is an inter-frame technology in VVC, but it can also use intra-frame prediction. The two prediction modes of GPM can both be inter-frame prediction modes, one can be inter-frame prediction mode and the other can be intra-frame prediction mode, or both can be intra-frame prediction modes.

[0140] The SGPM mode in ECM uses this weight mask or weight matrix in intra-frame prediction. It uses the weight matrix to combine the prediction values ​​of two intra-frame prediction modes into the prediction value of SGPM, which can produce more complex textures than a single prediction mode.

[0141] Because one partition mode and two intra-frame prediction modes are used, the general logic is to write syntax elements indicating these three modes into the bitstream, as shown in Table 1: partition_mode_idx, intra_pred_mode0_idx, and intra_pred_mode1_idx. However, to reduce the overhead of this information, SGPM uses a template around the current block to sort the combinations of the three modes, generating a candidate list of combined modes. Only the SGPM candidate indexes need to be written into the bitstream. At the decoder, SGPM constructs the candidate list of combined modes and, based on the SGPM candidate indexes (e.g., sgpm_cand_idx in Table 2), derives one partition mode and two intra-frame prediction modes: partition_mode_idx, intra_pred_mode0_idx, and intra_pred_mode1_idx.

[0142] Table 1

[0143] Table 2

[0144] Here, for sgpm_cand_idx,

[0145] Furthermore, the following introduces the transformation technology.

[0146] During encoding, commonly used hybrid coding frameworks first perform a prediction. This prediction leverages spatial or temporal correlation to produce an image identical or similar to the current block. While it's possible for the predicted block to be identical to the current block for a given block, it's difficult to guarantee this for all blocks in a video, especially in natural video or video captured by a camera. Irregular motion, distortion, occlusion, and brightness changes in video are difficult to fully predict. Therefore, hybrid coding frameworks subtract the predicted image from the original image of the current block to produce a residual image, or, in other words, subtract the predicted block from the current block to produce a residual block. This residual block is typically much simpler than the original image, so prediction can significantly improve compression efficiency. The residual block isn't encoded directly; instead, it's usually first transformed. This transform converts the residual image from the spatial domain to the frequency domain to remove correlation. After the residual image is transformed to the frequency domain, since most of the energy is concentrated in the low-frequency region, the non-zero coefficients are concentrated in the upper left corner. Quantization is then used to further compress the image. Furthermore, since the human eye is less sensitive to high frequencies, a larger quantization step size can be used in high-frequency regions.

[0147] Figure 16 illustrates a DCT transform. As shown in Figure 16, after the DCT transform, only the upper left corner of the original image has non-zero coefficients. Of course, this example applies the DCT transform to the entire image. In video codecs, images are processed by dividing them into blocks, so the transform is also performed on a block-by-block basis.

[0148] Transforms are very useful in typical video compression, but not all blocks require transforms. In some cases, transforms can even yield worse compression results than non-transforms. Therefore, in some standards, such as VVC, the encoder can choose whether to use transforms for the current block. DCT-II is the most commonly used transform in video compression standards, and its base image is shown in Figure 17.

[0149] In addition, VVC can also use DCT8 type (DCT-VIII) and DST7 type (DST-VII). The basic formulas of these transforms are shown in Table 3, which shows the basic transform formulas of DCT2, DCT8 and DST7 for N-point input.

[0150] Table 3

[0151] Because images are all two-dimensional, the computational complexity and memory overhead of performing a direct two-dimensional transform were prohibitive for the hardware available at the time. Therefore, the DCT2, DCT8, and DST7 transforms used in the standards were split into two steps: horizontal and vertical one-dimensional transforms. For example, the horizontal transform was performed first, followed by the vertical transform, or the vertical transform was performed first, followed by the horizontal transform.

[0152] (1) Multi-transformation selection MTS.

[0153] VVC supports transform kernels such as DCT2, DCT8, and DST7. The DCT2, DCT8, and DST7 kernels used in VVC are horizontally and numerically separable, allowing independent horizontal and vertical transforms. For a block, the encoder selects the appropriate transform kernel and transmits its index to the bitstream. The decoder then uses the index to determine the inverse transform kernel. Different transform kernels can be selected for the horizontal and vertical directions, such as using DCT8 horizontally and DST7 vertically. This technique is generally referred to as MTS.

[0154] VVC uses a syntax element, mts_idx, to determine the transform kernel of the base transform. As shown in Table 4 below, trTypeHor represents the transform kernel for the horizontal transform, and trTypeVer represents the transform kernel for the vertical transform. A value of 0 for trTypeHor or trTypeVer indicates a DCT2 transform, 1 indicates a DCT7 transform, and 2 indicates a DCT8 transform. If mts_idx is not present, the value of mts_idx is inferred to be 0.

[0155] Table 4

[0156] (2) Low-frequency non-separable transform LFNST.

[0157] The above transformation method is effective for horizontal and vertical textures, but less so for diagonal textures. Indeed, horizontal and vertical textures are the most common, making the above transformation method very useful for improving compression efficiency. As the demand for compression efficiency continues to increase, more efficient processing of diagonal textures could further improve compression efficiency.

[0158] In order to more effectively process the residual of oblique texture, LFNST transform is used in VVC. The above transforms such as DCT2, DCT8, and DST7 are called primary transforms. At the encoding end of VVC, LFNST is used after DCT2 transform and before quantization. At the decoding end of VVC, LFNST is used after inverse quantization and before inverse DCT2 transform. Because it is transformed on the basis of DCT2 (basic transform), LFNST is a secondary transform. Figure 18 is a schematic diagram of the encoding and decoding process without LFNST (secondary transform), and Figure 19 is a schematic diagram of the encoding and decoding process with LFNST (secondary transform). Of course, the encoding end can directly inverse quantize the saved quantization coefficients instead of entropy decoding, because entropy coding is lossless.

[0159] Figure 20 shows a detailed encoding and decoding process involving LFNST (secondary transform). On the encoder side, LFNST performs a secondary transform on the low-frequency coefficients in the upper left corner after the base transform. The base transform decorrelates the image, concentrating energy in the upper left corner. The secondary transform further decorrelates the low-frequency coefficients of the base transform. The result is intuitively shown in Figure 20. On the encoder side, 16 coefficients are input to the 4×4 LFNST, and the output is 8 coefficients. 48 coefficients are input to the 8×8 LFNST, and the output is 8 coefficients for the 8×8 block and 16 coefficients for other blocks. On the decoder side, 8 coefficients are input to the 4×4 inverse LFNST, and the output is 16 coefficients. 8 coefficients are input to the 8×8 block and 16 coefficients for other blocks. The coefficients are input to the 8×8 inverse LFNST, and the output is 48 coefficients.

[0160] Figure 21 shows some base images for LFNST in VVC. Only the two lowest-frequency base images for each kernel in each kernel group are shown. Some obvious diagonal textures can be seen. In addition to kernels optimized for certain diagonal textures, LFNST also has kernels optimized for flat, gradient textures, such as kernel group 0 in VVC.

[0161] LFNST is only applied to intra-coded blocks. Angular prediction tiles the reference pixels at a specified angle onto the current block as the prediction value. This means the predicted block will have a distinct directional texture, and the residual of the current block after angular prediction will also statistically exhibit significant angular characteristics. Therefore, the transform kernel selected by LFNST can be tied to the intra-prediction mode. That is, once the intra-prediction mode is determined, LFNST can only use the set of transform kernels corresponding to that intra-prediction mode.

[0162] Specifically, the LFNST in VVC has a total of 4 groups of transform kernels, and each group can select 2 transform kernels. Table 5 shows the correspondence between intra prediction modes and transform kernel groups. Note that the cross-component prediction modes used for chroma intra prediction are 81 to 83, and there are no such modes for luma intra prediction. The transform kernel of LFNST can be transposed to process more angles with one transform kernel group. For example, modes 13 to 23 and 45 to 55 both correspond to transform kernel group 2, but 13 to 23 is obviously close to the horizontal mode and 45 to 55 is obviously close to the vertical mode.

[0163] Table 5

[0164] VVC's LFNST uses four sets of transform kernels, with the intra-prediction mode specifying which set to use. This leverages the correlation between the intra-prediction mode and the LFNST transform kernel, reducing the transmission of the selected LFNST transform kernel in the bitstream. Whether the current block uses LFNST, and if so, whether to use the first or second set within a set, is determined by the bitstream and certain conditions.

[0165] In the subsequent evolution of ECM technology, LFNST was further expanded. LFNST has more transform kernel groups, 35 in ECM. The correspondence between the transform kernel group index (LFNST set index) and the intra prediction mode (Intra pred.mode) is shown in Table 6. Each transform kernel group is more efficient for textures at the corresponding angle. Here, each transform kernel group can select three transform kernels.

[0166] Table 6

[0167] (3) Non-separable basis transformation NSPT.

[0168] LFNST is a horizontally and vertically inseparable transform. Because it involves a secondary transform, DCT2 can be called the base transform. This approach of performing DCT2 before LFNST is a compromise between performance and complexity. While directly performing the inseparable base transform is more efficient, it also incurs higher complexity, such as increased computational effort and storage space required for the transform kernel.

[0169] In ECM10, some small blocks can use NSPT, while large blocks still use DCT2+LFNST. The sizes of small blocks are 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 8×32, and 32×8. In ECM10, NSPT also matches the transform kernel group according to the intra prediction mode. The matching method can refer to the LFNST method. Each transform kernel group has 3 transform kernels to choose from. For example, an 8×8 base image of an NSPT in ECM10 is shown in Figure 22. Figure 22 corresponds to inter-frame angle prediction mode 7. It can be seen that it handles the texture of the corresponding angle better. It should be noted that NSPT is only applied to intra-coded blocks.

[0170] In summary, both NSPT and LFNST process transforms for textures at various angles. They may have multiple transform kernels, and a transform kernel may be specifically optimized for a specific angular texture. Of course, in addition to angular textures, NSPT and LFNST also include transform kernels for gradient textures. In fact, these transform kernels can also be said to be trained KLTs (Karhunen-Loeve Transforms). It can also be summarized that both NSPT and LFNST have multiple transform kernels, each designed for a specific texture, including angular textures, gradient textures, and so on. Of course, gradient textures can be further expanded to include horizontal gradient textures, vertical gradient textures, and diagonal gradient textures. They have multiple transform kernel groups, and each intra-frame prediction mode can correspond to a transform kernel group. Each intra-frame prediction mode actually represents a texture feature. Therefore, the intra-frame prediction mode index is also a texture feature index.

[0171] In VVC and the current ECM, a block (CU or TU) can only derive a single texture feature index for deriving the LFNST / NSPT transform kernel set, thereby deriving a unique LFNST / NSPT transform kernel set. This texture feature can be said to be the link for using LFNST / NSPT. For blocks using standard intra prediction modes, namely DC, PLANAR, and various angle modes, this texture feature index is the intra prediction mode used by the current block. However, for special intra prediction modes such as DIMD, TIMD, MIP, SGPM, ITMP, and IBC, this is not so straightforward. DIMD and TIMD can both weight the prediction values ​​of two or more intra prediction modes, SGPM weights the prediction values ​​of two intra prediction modes using a weight matrix, MIP performs prediction based on a matrix operation, and ITMP and IBC perform predictions by replicating a reconstructed block. These are not simple texture features like standard intra prediction modes. For example, a block weighted by the prediction values ​​of two or more intra-frame prediction modes indicates that its texture may contain two or more texture features. Its residual after prediction may exhibit texture features of the first prediction mode or the second prediction mode. This shows that for blocks that require weighting of the prediction values ​​of two or more intra-frame prediction modes, the current transformation process does not fully consider the whole process, resulting in low compression efficiency.

[0172] Based on this, an embodiment of the present application provides an encoding method, which determines a prediction mode of a current block; when the prediction mode of the current block meets a first condition, determines a first candidate list of the current block; wherein the first candidate list indicates at least two candidate transform core groups; according to the first candidate list, determines a transform core of the current block; determines a residual block of the current block, and transforms the residual block of the current block according to the transform core to determine the transform coefficient of the current block; encodes the transform coefficient of the current block and writes the obtained coded bits into a bitstream. An embodiment of the present application also provides a decoding method, which determines a prediction mode of the current block; when the prediction mode of the current block meets a first condition, determines a first candidate list of the current block; wherein the first candidate list indicates at least two candidate transform core groups; according to the first candidate list, determines a transform core of the current block; determines the transform coefficient of the current block, and inversely transforms the transform coefficient of the current block according to the transform core to determine the residual block of the current block.

[0173] In this way, both the encoding and decoding ends first determine the prediction mode for the current block. When the prediction mode for the current block meets the first condition, a first candidate list indicating at least two candidate transform core groups is determined, and then the transform core for the current block is determined based on these at least two candidate transform core groups. In other words, for a current block predicted using certain intra-frame prediction modes, the transform core for the current block is no longer determined based on just one transform core group, but rather based on at least two candidate transform core groups. This improves the accuracy of the transform, thereby increasing compression efficiency and, in turn, enhancing encoding and decoding performance.

[0174] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0175] FIG23 is a schematic diagram of a network architecture for video encoding and decoding provided in an embodiment of the present application. As shown in FIG23 , the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction via the communication network 01. During implementation, the electronic devices can be various types of devices with video encoding and decoding capabilities. For example, the electronic devices can include mobile phones, tablet computers, personal computers, personal digital assistants, navigators, digital phones, video phones, televisions, sensor devices, servers, etc., and the embodiments of the present application do not limit this.

[0176] In an embodiment of the present application, a network architecture of a video encoding and decoding system including a decoding method and an encoding method is provided. The decoder or encoder in the embodiment of the present application can be the aforementioned electronic device. In other words, the electronic device in the embodiment of the present application has video encoding and decoding capabilities, and generally includes a video encoder (i.e., encoder) and a video decoder (i.e., decoder).

[0177] Figure 24 is a schematic block diagram of the system composition of an encoder provided in an embodiment of the present application. As shown in Figure 24, the encoder 100 may include: a segmentation unit 101, a prediction unit 102, a first adder 107, a transform unit 108, a quantization unit 109, an inverse quantization unit 110, an inverse transform unit 111, a second adder 112, a filtering unit 113, a decoded picture buffer (DPB) unit 114, and an entropy coding unit 115. Here, the input of the encoder 100 may be a video consisting of a series of pictures or a static picture, and the output of the encoder 100 may be a bitstream (also referred to as a "codestream") used to represent a compressed version of the input video.

[0178] Among them, the segmentation unit 101 segments the picture in the input video into one or more Coding Tree Units (CTUs). The segmentation unit 101 divides the picture into multiple tiles (or tiles), and can further divide a tile into one or more bricks. Here, a tile or a brick may include one or more complete and / or partial CTUs. In addition, the segmentation unit 101 can form one or more slices, where a slice can include one or more tiles arranged in a grid order in the picture, or one or more tiles covering a rectangular area in the picture. The segmentation unit 101 can also form one or more sub-pictures, where a sub-picture can include one or more slices, tiles or bricks.

[0179] During the encoding process of encoder 100, segmentation unit 101 transmits the CTU to prediction unit 102. Generally, prediction unit 102 may be composed of block segmentation unit 103, motion estimation (ME) unit 104, motion compensation (MC) unit 105, and intra prediction unit 106. Specifically, block segmentation unit 103 iteratively uses quadtree segmentation, binary tree segmentation, and ternary tree segmentation to further divide the input CTU into smaller coding units (CUs). Prediction unit 102 may use ME unit 104 and MC unit 105 to obtain inter-frame prediction blocks for the CU. Intra-frame prediction unit 106 may use various intra-frame prediction modes, including MIP mode, to obtain intra-frame prediction blocks for the CU. In an example, a rate-distortion optimized motion estimation method may be used by ME unit 104 and MC unit 105 to obtain inter-frame prediction blocks, and a rate-distortion optimized mode determination method may be used by intra-frame prediction unit 106 to obtain intra-frame prediction blocks. The prediction unit 102 outputs the prediction block of the CU, and the first adder 107 calculates the difference between the CU in the output of the segmentation unit 101 and the prediction block of the CU, i.e., the residual CU. The transform unit 108 reads the residual CU and performs one or more transform operations on the residual CU to obtain coefficients. The quantization unit 109 quantizes the coefficients and outputs the quantized coefficients (i.e., levels). The inverse quantization unit 110 performs a scaling operation on the quantized coefficients to output reconstructed coefficients. The inverse transform unit 111 performs one or more inverse transforms corresponding to the transform in the transform unit 108 and outputs the reconstructed residual. The second adder 112 calculates the reconstructed CU by adding the reconstructed residual and the prediction block of the CU from the prediction unit 102. The second adder 112 also sends its output to the prediction unit 102 for use as an intra-frame prediction reference. After all CUs in the picture or sub-picture are reconstructed, the filtering unit 113 performs loop filtering on the reconstructed picture or sub-picture. Here, the filtering unit 113 includes one or more filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), a luminance mapping and chroma scaling (LMCS) filter, and a neural network-based filter. Alternatively, when the filtering unit 113 determines that the CU is not used as a reference for encoding other CUs, the filtering unit 113 performs loop filtering on one or more target pixels in the CU. The output of the filtering unit 113 is a decoded picture or sub-picture, which is cached to the DPB unit 114. The DPB unit 114 outputs the decoded picture or sub-picture according to the timing and control information.Here, the picture stored in the DPB unit 114 can also be used as a reference for the prediction unit 102 to perform inter-frame prediction or intra-frame prediction. Finally, the entropy coding unit 115 converts the parameters necessary for decoding the picture from the encoder 100 (such as control parameters and supplementary information, etc.) into binary form, and writes such binary form into the code stream according to the syntax structure of each data unit. That is, the encoder 100 finally outputs the code stream.

[0180] Furthermore, the encoder 100 can be a device having a first processor and a first memory for recording a computer program. When the first processor reads and executes the computer program, the encoder 100 reads the input video and generates a corresponding bitstream. Alternatively, the encoder 100 can be a computing device having one or more chips. These units implemented as integrated circuits on the chip have similar connection and data exchange functions as the corresponding units in FIG. 24 .

[0181] Figure 25 is a block diagram of the system components of a decoder provided in an embodiment of the present application. As shown in Figure 25, the decoder 200 may include: a parsing unit 201, a prediction unit 202, an inverse quantization unit 205, an inverse transform unit 206, an adder 207, a filtering unit 208, and a decoded image buffer unit 209. Here, the input of the decoder 200 is a bitstream representing a compressed version of a video or a still image, and the output of the decoder 200 may be a decoded video consisting of a series of images or a decoded still image.

[0182] The input codestream to decoder 200 may be the codestream generated by encoder 100. Parsing unit 201 parses the input codestream and obtains syntax element values ​​from the input codestream. Parsing unit 201 converts the binary representation of the syntax elements into digital values ​​and sends the digital values ​​to units within decoder 200 to obtain one or more decoded pictures. Parsing unit 201 may also parse one or more syntax elements from the input codestream to display decoded pictures.

[0183] During the decoding process in decoder 200, parsing unit 201 transmits the values ​​of syntax elements and one or more variables set or determined based on the values ​​of the syntax elements, used to obtain one or more decoded pictures, to units within decoder 200. Prediction unit 202 determines a prediction block for the current decoding block (e.g., a CU). Prediction unit 202 may include motion compensation unit 203 and intra prediction unit 204. Specifically, when an inter decoding mode is indicated for decoding the current decoding block, prediction unit 202 passes relevant parameters from parsing unit 201 to motion compensation unit 203 to obtain an inter prediction block. When an intra prediction mode (including a MIP mode indicated based on a MIP mode index value) is indicated for decoding the current decoding block, prediction unit 202 passes relevant parameters from parsing unit 201 to intra prediction unit 204 to obtain an intra prediction block. Dequantization unit 205 has the same functionality as dequantization unit 110 in encoder 100. Dequantization unit 205 performs a scaling operation on the quantization coefficients (i.e., levels) from parsing unit 201 to obtain reconstructed coefficients. The inverse transform unit 206 has the same function as the inverse transform unit 111 in the encoder 100. The inverse transform unit 206 performs one or more transform operations (i.e., the inverse of the one or more transform operations performed by the inverse transform unit 111 in the encoder 100) to obtain a reconstructed residual. The adder 207 performs an addition operation on its input (the prediction block from the prediction unit 202 and the reconstructed residual from the inverse transform unit 206) to obtain a reconstructed block of the current decoded block. The reconstructed block is also sent to the prediction unit 202 to be used as a reference for other blocks encoded in the intra prediction mode.

[0184] After all CUs in the picture or sub-picture are reconstructed, the filtering unit 208 performs loop filtering on the reconstructed picture or sub-picture. The filtering unit 208 includes one or more filters, such as a deblocking filter, a sample adaptive offset filter, an adaptive loop filter, a luminance mapping and chroma scaling filter, and a neural network-based filter. Alternatively, when the filtering unit 208 determines that the reconstructed block is not used as a reference for decoding other blocks, the filtering unit 208 performs loop filtering on one or more target pixels in the reconstructed block. Here, the output of the filtering unit 208 is a decoded picture or sub-picture, which is cached to the DPB unit 209. The DPB unit 209 outputs the decoded picture or sub-picture based on timing and control information. The picture stored in the DPB unit 209 can also be used as a reference for performing inter-frame prediction or intra-frame prediction by the prediction unit 202.

[0185] Furthermore, the decoder 200 can be a second memory having a second processor and a computer program. When the first processor reads and runs the computer program, the decoder 200 reads the input code stream and generates a corresponding decoded video. In addition, the decoder 200 can also be a computing device having one or more chips. These units implemented as integrated circuits on the chip have similar connection and data exchange functions as the corresponding units in Figure 25.

[0186] It should also be noted that when the embodiment of the present application is applied to the encoder 100, the "current block" specifically refers to the current block to be encoded in the video image (which can also be simply referred to as the "encoding block"); when the embodiment of the present application is applied to the decoder 200, the "current block" specifically refers to the current block to be decoded in the video image (which can also be simply referred to as the "decoding block").

[0187] In one embodiment of the present application, FIG26 is a flowchart of a decoding method provided by the embodiment of the present application. As shown in FIG26 , the method may include:

[0188] S2601, determine the prediction mode of the current block.

[0189] It should be noted that in the embodiments of the present application, the method is applied to a decoder. Specifically, based on the structure of decoder 200 shown in FIG25 , the decoding method of the embodiments of the present application is primarily applied to intra-predicted blocks. Specifically, when the current block uses intra-prediction mode, the optimization scheme proposed here is mainly for the NSPT and LFNST transforms in intra-prediction mode to improve compression efficiency.

[0190] It should also be noted that in the embodiments of the present application, NSPT and LFNST are both transformations for processing textures at various angles. They may have multiple transformation kernels, and one transformation kernel may be specifically optimized for a certain specific angle texture. Of course, in addition to angle textures, NSPT and LFNST also include transformation kernels for processing gradient textures. In fact, these transformation kernels can also be said to be trained KL transforms (Karhunen-Loeve Transform, KLT). That is to say, NSPT and LFNST both have multiple transformation kernels, each of which is designed for a specific texture, and specific textures include angle textures, gradient textures, etc. In addition, gradient textures can be further extended to include horizontal gradient textures, vertical gradient textures, oblique gradient textures, etc. Furthermore, this technical solution is not limited to being used only for inseparable transformations such as NSPT and LFNST, and this technical solution can also be applied to separable transformations optimized for specific textures.

[0191] It should also be noted that in the intra-frame prediction blocks, they have multiple transform kernel groups, and each intra-frame prediction mode can correspond to a transform kernel group. In other words, each intra-frame prediction mode actually represents a texture feature. Therefore, the intra-frame prediction mode index is also a texture feature index. For example, DC and PLANAR correspond to gradient texture features, and a certain angle prediction mode corresponds to the texture feature of this angle. On the one hand, the texture feature index can avoid the appearance of intra-frame prediction mode in "inter-frame", and on the other hand, it is also more conducive to possible expansion. For example, an intra-frame prediction mode can correspond to multiple texture features, such as the DC mode can correspond to horizontal gradient texture, vertical gradient texture, oblique gradient texture, etc.

[0192] In embodiments of the present application, some intra-frame prediction modes are not based on simple texture features but may include two or more texture features. Therefore, it is necessary to first determine the prediction mode of the current block. In some embodiments, the method may include: decoding the bitstream and determining the prediction mode of the current block.

[0193] In the embodiment of the present application, some mode indication information (or mode flag) in the form of syntax elements can be written into the bitstream. In this way, by parsing the value of the mode indication information (or mode flag) in the bitstream, the prediction mode of the current block can be determined.

[0194] S2602 : When the prediction mode of the current block satisfies a first condition, determine a first candidate list for the current block; wherein the first candidate list indicates at least two candidate transform core groups.

[0195] It should be noted that, in the embodiment of the present application, the prediction mode of the current block satisfies the first condition, which may include: the prediction mode of the current block is one of the items in the first prediction mode set.

[0196] It should be noted that, in the embodiment of the present application, the first prediction mode set includes at least: DIMD mode, TIMD mode, SGPM mode, MIP mode, ITMP mode and IBC mode.

[0197] It should also be noted that, in an embodiment of the present application, the modes in the first prediction mode set are relatively complex prediction modes, and the corresponding textures may contain two or more texture features. For example, both the DIMD mode and the TIMD mode can weight the prediction values ​​of two or more intra-frame prediction modes, the SGPM mode weights the prediction values ​​of the two intra-frame prediction modes with a weight matrix, the MIP mode predicts based on a matrix operation, and the ITMP mode and the IBC mode predict based on copying a reconstructed reference block. Their texture features are not as simple as those of the DC mode, the PLANAR mode, etc., so it is necessary to determine the first candidate list here. Among them, the first candidate list may include at least two candidate texture feature indexes, or the first candidate list may include at least two transform core groups. Here, each candidate texture feature index corresponds to a transform core group. Therefore, it can be said that the first candidate list indicates at least two candidate transform core groups.

[0198] In some embodiments, determining a first candidate list for a current block may include: determining candidate pixels for deriving texture feature indexes; determining one or more candidate transform kernel groups for the current block based on the candidate pixels; and adding the one or more candidate transform kernel groups to the first candidate list. Here, when determining the one or more candidate transform kernel groups for the current block based on the candidate pixels, one or more candidate texture feature indexes for the current block may be first determined based on the candidate pixels, and then the one or more candidate transform kernel groups for the current block are determined based on the one or more candidate texture feature indexes.

[0199] It should also be noted that, generally, one candidate texture feature index corresponds to one candidate transform core group; however, in some cases, multiple similar candidate texture feature indexes may correspond to the same transform core group. For example, multiple intra-frame prediction modes with similar angles correspond to the same transform core group. In an embodiment of the present application, if multiple adjacent intra-frame prediction modes (or candidate texture feature indexes) correspond to one transform core group, then when determining the candidate transform core group, it is necessary to ensure that the candidate texture feature indexes are not determined to be the same transform core group.

[0200] In a possible implementation, for candidate pixels, a prediction block of the current block may be determined; and at least part of the pixels in the prediction block are used as candidate pixels.

[0201] In the embodiment of the present application, if a certain texture exists in the prediction block, it can be considered that the residual block has a texture with the same characteristics. In this way, the candidate pixels used for deducing the candidate texture feature index between frames can be all pixels in the prediction block or part of the pixels in the prediction block.

[0202] In another possible implementation, for candidate pixels, adjacent pixels of a reconstructed area of ​​the current block may be determined; and the adjacent pixels of the reconstructed area may be used as candidate pixels.

[0203] In the embodiment of the present application, the candidate pixels used to derive the candidate texture feature index between frames can be pixels adjacent to the reconstructed area of ​​the current block, such as the reconstructed areas to the left and right of the current block. Because the reconstructed areas to the left and above are not the current block but are adjacent to the current block, for example, if the textures are connected, they can be used to estimate the texture of the current block to a certain extent.

[0204] In yet another possible implementation, more pixels are considered. For candidate pixels, adjacent pixels in the reconstructed area and at least part of the pixels in the prediction block may be used as candidate pixels.

[0205] In an embodiment of the present application, the candidate pixels used to derive the candidate texture feature index between frames can also use the predicted block of the current block and the reconstructed areas on the left and above the current block at the same time. In this way, more pixels are used to derive the candidate texture feature index, making the derived candidate texture feature index more accurate.

[0206] It should also be noted that, in the embodiment of the present application, the number of candidate pixels used to derive the candidate texture feature index may be at least one, for example, 1, 2, 3 or more.

[0207] In some embodiments, the method may further include: determining the number of candidate pixels according to a size parameter of the current block.

[0208] That is to say, when deriving one or more candidate texture feature indexes based on candidate pixels, the number of candidate pixels used can be determined by the size parameter of the current block. For example, if the size of the current block is small, then all available pixels can be counted; if the size of the current block is large, then the current block can be downsampled and counted, such as counting one pixel out of every 2, or 4, or 8 pixels in the horizontal and / or vertical directions. Alternatively, if the size of the current block in one of the horizontal or vertical directions is less than or equal to 8, then all available pixels in that direction are counted; otherwise, if the size of the current block in one of the horizontal or vertical directions is less than or equal to 16, then one pixel out of every 2 pixels in that direction is counted; otherwise, one pixel out of every 4 pixels in that direction is counted, and no specific limitation is given here.

[0209] In some embodiments, determining a candidate texture feature list for a current block based on candidate pixels may include: determining a horizontal gradient value and a vertical gradient value of the candidate pixel; determining a texture feature index and a gradient intensity value corresponding to the candidate pixel based on the horizontal gradient value and the vertical gradient value of the candidate pixel; constructing a texture feature statistics table based on the texture feature index and the gradient intensity value corresponding to the candidate pixel; and determining a candidate texture feature list for the current block based on the texture feature statistics table.

[0210] It should be noted that in an embodiment of the present application, when determining the texture feature index and gradient intensity value corresponding to the candidate pixel based on the horizontal gradient value and vertical gradient value of the candidate pixel, it can include: performing angle mapping based on the horizontal gradient value and vertical gradient value of the candidate pixel to determine the texture feature index corresponding to the candidate pixel; and performing gradient intensity calculation based on the horizontal gradient value and vertical gradient value of the candidate pixel to determine the gradient intensity value corresponding to the candidate pixel.

[0211] In a specific embodiment, performing angle mapping based on the horizontal gradient value and the vertical gradient value of the candidate pixel to determine the texture feature index corresponding to the candidate pixel may include: determining the texture feature index corresponding to the candidate pixel using a preset lookup table based on the horizontal gradient value and the vertical gradient value of the candidate pixel.

[0212] In the embodiment of the present application, the horizontal gradient value of the candidate pixel can be expressed as grad x Indicates that the vertical gradient value of the candidate pixel can be expressed as grad y In this way, according to grad x and grad y Deriving the texture feature index (or referred to as a “virtual intra prediction mode”) can be achieved by looking up a table.

[0213] For example, if abs(grad x ) is equal to 0 and abs(grad y ) is not equal to 0, then there is horizontal texture, corresponding to intra prediction mode 18 in VVC. y ) is equal to 0 and abs(grad x ) is not equal to 0, then there is vertical texture, corresponding to intra prediction mode 50 in VVC. x ) and abs(grad y ) are not equal to 0, if abs(grad x ) is equal to abs(grad y ), and grad x and grad y The symbols are the same, corresponding to intra prediction mode 34 in VVC. If abs(gradx ) is equal to 2 times abs(grad y ), and grad x and grad y The symbols are the same, corresponding to intra prediction mode 40 in VVC. In addition, other situations can be determined by looking up the table according to the same principle.

[0214] In a specific embodiment, performing gradient strength calculation based on the horizontal gradient value and the vertical gradient value of the candidate pixel to determine the gradient strength value corresponding to the candidate pixel may include: performing an addition operation on the absolute value of the horizontal gradient value and the absolute value of the vertical gradient value to determine the gradient strength value corresponding to the candidate pixel.

[0215] Here, the gradient intensity value corresponding to the candidate pixel can be recorded as amp. For example, amp=abs(grad x )+abs(grad y ).

[0216] It should be noted that, in the embodiment of the present application, the horizontal gradient value and the vertical gradient value of the candidate pixel can be calculated using the Sobel operator. For example, the Sobel operator is as follows:

[0217] Operator for horizontal gradient value:

[0218] Operator for vertical gradient value:

[0219] So, suppose the pixel value at pixel position (x, y) is P x,y , then the horizontal gradient value grad x And the vertical gradient value grad y The calculation of grad is as follows: x =P x+1,y-1 +2*P x+1,y +P x+1,y+1 -P x-1,y-1 -2*P x-1,y -P x-1,y+1 (5) grad y =P x-1,y+1 +2*P x,y+1 +P x+1,y+1 -P x-1,y-1 -2*P x,y-1 -P x+1,y-1 (6)

[0220] It should also be noted that, in this embodiment of the present application, the candidate pixels may be set to exclude the pixels in the outermost row, column, and top, bottom, left, and right of the current block. Considering that the Sobel operator uses the pixels in the outermost row, column, and top, bottom, left, and right of the current pixel, the gradient of the pixels in the outermost row, column, and top, bottom, left, and right of the current pixel may not be calculated in this embodiment of the present application.

[0221] In some embodiments, constructing a texture feature statistics table based on the texture feature index and gradient intensity value corresponding to the candidate pixel may include: determining at least one texture feature index and at least one corresponding gradient intensity value when the number of candidate pixels is at least one; determining at least one reference texture feature index with mutually different characteristics based on the at least one texture feature index, and accumulating the gradient intensity values ​​belonging to the same reference texture feature index based on the at least one gradient intensity value to determine the gradient intensity cumulative value corresponding to the at least one reference texture feature index; constructing a texture feature statistics table based on the at least one reference texture feature index and the gradient intensity cumulative value corresponding to the at least one reference texture feature index.

[0222] That is to say, in an embodiment of the present application, taking at least part of the pixels in the prediction block as candidate pixels as an example, the gradients of all or part of the pixels in the prediction block are calculated. Generally speaking, the horizontal gradient value and the vertical gradient value can be calculated. Here, the Sobel operator can be used to calculate the gradient value. For a certain pixel, the texture direction of the pixel can be inferred based on its horizontal gradient value and vertical gradient value. For example, if the horizontal gradient value is non-zero and the vertical gradient value is zero, then the texture of the pixel is in the vertical direction. Conversely, if the horizontal gradient value is zero and the vertical gradient value is non-zero, then the texture of the pixel is in the horizontal direction. For example, if the horizontal gradient value and the vertical gradient value are equal and not zero, then the texture of the pixel is 45 degrees. Of course, there are many other cases in the embodiment of the present application where the horizontal gradient value and the vertical gradient value are not zero, and the texture direction of the pixel can be determined based on their ratio. In this way, the gradient intensity value of each pixel can be corresponded to the corresponding texture feature index. A texture feature statistics table is constructed, and the gradient intensity value of each calculated pixel is added to the corresponding texture feature index item in the statistics table to obtain the final texture feature statistics table. Then, the candidate texture feature list of the current block can be determined based on the texture feature statistics table.

[0223] In some embodiments, determining a candidate texture feature list for a current block based on a texture feature statistics table may include: sorting the texture feature statistics table from high to low according to gradient intensity accumulation values, determining reference texture feature indexes corresponding to top N gradient intensity accumulation values, and adding the N reference texture feature indexes to the candidate texture feature list for the current block, where N is a positive integer.

[0224] That is, in the embodiment of the present application, after the texture feature statistics table is constructed, in order to reduce complexity, the texture feature statistics table can also be sorted in descending order according to the accumulated gradient strength cumulative values, and then only the top N reference texture feature indexes are selected. Here, for complexity considerations, only a candidate texture feature list of length N can be maintained, and the candidate texture feature list is sorted from high to low according to the gradient strength cumulative values.

[0225] It should also be noted that in the embodiments of the present application, the value of N can be 2, 3, 4, 5, ..., 10, etc., and there is no specific limitation on the value of N here.

[0226] It can be understood that in the embodiment of the present application, considering that the angles of some adjacent intra-frame prediction modes are very close, in order to exclude intra-frame prediction modes that are too close, the method may also include: pruning the N reference texture feature indexes in the candidate texture feature list to determine one or more candidate transform core groups for the current block.

[0227] In some embodiments, pruning N reference texture feature indexes in a candidate texture feature list to determine one or more candidate transform core groups for a current block may include: determining a first candidate transform core group based on a reference texture feature index at a first position in the candidate texture feature list; when a second condition is satisfied between other reference texture feature indexes other than the first position in the candidate texture feature list and the i-th candidate transform core group, determining an i+1-th candidate transform core group based on the other reference texture feature indexes to determine one or more candidate transform core groups for the current block; wherein i is an integer greater than zero and less than N.

[0228] It should be noted that, in the embodiment of the present application, the second condition may include: the difference between the i-th candidate texture feature index and the i+1-th candidate texture feature index satisfies a preset threshold, i.e., there is a certain difference between two adjacent candidate texture feature indexes. Alternatively, when each candidate texture feature index corresponds to a candidate transform core group, it can also be said that the difference between the i-th candidate transform core group and the i+1-th candidate transform core group satisfies the preset threshold.

[0229] It should also be noted that, in the embodiment of the present application, the N reference texture feature indexes are sorted from high to low according to the accumulated gradient strength values. In this case, determining the first candidate transform kernel group based on the reference texture feature index that is first in the candidate texture feature list may include determining the first candidate transform kernel group based on the reference texture feature index with the largest accumulated gradient strength value in the candidate texture feature list.

[0230] It should also be noted that, in the embodiment of the present application, the preset threshold can be represented by THR, the i-th candidate texture feature index can be represented by candFeature(i), and the i+1-th candidate texture feature index can be represented by candFeature(i+1). In a specific embodiment, the difference between the i-th candidate texture feature index and the i+1-th candidate texture feature index meets the preset threshold, which can include: candFeature(i+1)+THR<candFeature(i)||candFeature(i+1)-THR> candFeature(i).

[0231] It should also be noted that, in the embodiment of the present application, the value of THR may be 3, 4, 5, 6, etc. Thus, when the cumulative values ​​of multiple adjacent angles are very high, this pruning method will give priority to candidate texture feature indexes with a certain degree of discrimination.

[0232] It is also understandable that in the embodiments of the present application, the above method does not take into account some intra-frame prediction modes derived from the DIMD, TIMD, SGPM, and other modes themselves. Therefore, in some embodiments, the method may further include: determining one or more intra-frame prediction modes derived from the prediction mode of the current block; determining one or more candidate transform core groups based on the one or more intra-frame prediction modes, and adding the one or more candidate transform core groups to the first candidate list.

[0233] In some embodiments, the method further includes: when the first candidate list is not full, determining a preset texture feature index of the current block; determining one or more candidate transform core groups according to the preset texture feature index, and adding the one or more candidate transform core groups to the first candidate list.

[0234] In some embodiments, the method further includes: when the first candidate list is not filled, determining candidate pixels for deriving texture feature indexes; determining a candidate texture feature list of the current block based on the candidate pixels; determining one or more candidate transform kernel groups based on the candidate texture feature list, and adding the one or more candidate transform kernel groups to the first candidate list.

[0235] That is to say, in the embodiment of the present application, some modes derived from the DIMD, TIMD, SGPM and other modes themselves are taken into consideration. For example, DIMD itself will derive one or several intra-frame prediction modes for weighting, and TIMD itself will also derive one or several intra-frame prediction modes for weighting. SGPM not only has two intra-frame prediction modes, but also has a "partitioning" mode that can also find the corresponding intra-frame prediction mode, and the residual often appears in the boundary area of ​​the "partition". Therefore, one possible implementation method is to determine the first candidate list based on the intra-frame prediction mode derived from the DIMD, TIMD, SGPM and other modes themselves and the candidate texture feature index derived by the above method; or another possible implementation method is to determine the first candidate list based on the intra-frame prediction mode derived from the DIMD, TIMD, SGPM and other modes themselves and the default texture feature index.

[0236] In addition, in the embodiment of the present application, since DIMD, TIMD, and SGPM can all derive more than one intra-frame prediction mode, these modes can also give priority to using multiple modes derived by each mode itself to determine the candidate texture feature index. When the modes derived by each mode itself cannot fill all the candidate texture feature indexes, one possible method is to add a default texture feature index. Another possible method is to use the above-mentioned sorted intra-frame prediction mode list of length N to determine the candidate texture feature index.

[0237] For example, when the prediction mode of the current block is the DIMD mode, if it uses the prediction values ​​of multiple intra-frame prediction modes for weighting during prediction, then the multiple intra-frame prediction modes used for weighting can be sequentially tried to be determined as candidate texture feature indexes.

[0238] For example, when the prediction mode of the current block is the TIMD mode, if it uses the prediction values ​​of multiple intra-frame prediction modes for weighting during prediction, then the multiple intra-frame prediction modes used for weighting can be sequentially tried to be determined as candidate texture feature indexes.

[0239] For example, when the prediction mode of the current block is the SGPM mode, the intra-frame prediction mode corresponding to the "partitioning" mode derived therefrom and the two intra-frame prediction modes used for prediction may be sequentially attempted to be determined as candidate texture feature indexes.

[0240] In another possible implementation, for determining the first candidate list of the current block, the method may include: determining a first candidate transformation core group based on the reference texture feature index at the first position in the candidate texture feature list; when the second condition is not satisfied between all reference texture feature indexes except the first position in the candidate texture feature list and the first candidate transformation core group, determining a second candidate transformation core group based on the reference texture feature index at the second position in the candidate texture feature list; and determining the first candidate list based on the first candidate transformation core group and the second candidate transformation core group.

[0241] It should be noted that in an embodiment of the present application, the first candidate transformation core group is determined based on the reference texture feature index in the first position in the candidate texture feature list. Specifically, it can be: the first candidate transformation core group is determined based on the reference texture feature index with the largest gradient intensity accumulated value in the candidate texture feature list.

[0242] It should also be noted that, in the embodiment of the present application, taking the first candidate list indicating two candidate transform core groups as an example, the first candidate transform core group can be represented by candFeature0, and the second candidate transform core group can be represented by candFeature1. For example, the intra-frame prediction mode with the largest cumulative gradient strength value is the first candidate transform core group candFeature0, then when selecting the second candidate transform core group candFeature1, it is required that candFeature1 and candFeature0 have a certain gap, such as candFeature1+THR<candFeature0||candFeature1-THR> If no matching candidate is found after checking all N-1 reference texture feature indices, a second candidate transform core group, candFeature1, can be determined based on the second-ranked reference texture feature index among the N reference texture feature indices. In other words, the candidate transform core groups corresponding to the first two reference texture feature indices with the largest cumulative gradient strength values ​​can be added to the first candidate list.

[0243] In another possible implementation, for determining the first candidate list of the current block, the method may include: determining a first candidate transformation core group based on the reference texture feature index at the first position in the candidate texture feature list; when the second condition is not satisfied between all reference texture feature indexes other than the first position in the candidate texture feature list and the first candidate transformation core group, determining a second candidate transformation core group based on a preset texture feature index; and determining the first candidate list based on the first candidate transformation core group and the second candidate transformation core group.

[0244] It should be noted that in an embodiment of the present application, if no one meets the requirements after checking all N-1 reference texture feature indexes, then the default texture feature index of the current block can also be determined, and then the candidate transform core group corresponding to the reference texture feature index with the largest gradient intensity accumulation value and the default texture feature index is added to the first candidate list.

[0245] In another possible implementation, for determining the first candidate list of the current block, the method may include: determining a first candidate transform core group based on a first intra-frame prediction mode derived from a prediction mode of the current block; determining a second candidate transform core group based on the first reference texture feature index when a second condition is satisfied between the first reference texture feature index in the candidate texture feature list and the first candidate transform core group; and determining the first candidate list based on the first candidate transform core group and the second candidate transform core group.

[0246] It should be noted that in this embodiment of the present application, the first candidate list not only considers the candidate texture feature indexes derived from the horizontal and vertical gradient values ​​of the candidate pixels, but also considers the intra-frame prediction mode derived from the prediction mode of the current block itself. This is explained in detail below using several examples.

[0247] Exemplarily, when the prediction mode of the current block is the DIMD mode, the first intra prediction mode derived from the DIMD mode is used as candFeature0.

[0248] Exemplarily, when the prediction mode of the current block is the TIMD mode, the first intra-frame prediction mode derived from the TIMD mode is used as candFeature0.

[0249] Exemplarily, when the prediction mode of the current block is the SGPM mode, the intra-frame prediction mode corresponding to the SGPM "partition" mode is used as candFeature0.

[0250] Then, candFeature1 is determined according to the above method. For example, if a reference texture feature index among the N reference texture feature indexes is tried, starting from the first position, and a reference texture feature index meets the THR restriction, it can be used as candFeature1.

[0251] It should also be noted that in the embodiment of the present application, for modes such as IBC and ITMP, a prediction block or a prediction block plus a template can be used to derive a candidate texture feature index (or intra-frame prediction mode).

[0252] It should also be noted that in an embodiment of the present application, if the candidate pixels used in the embodiment of the present application to derive N reference texture feature indexes are the same as the candidate pixels used in the DIMD mode, then the first intra-frame prediction mode derived by DIMD and the first reference texture feature index derived by the embodiment of the present application are the same.

[0253] S2603: Determine a transformation kernel for the current block according to the first candidate list.

[0254] It should be noted that after the first candidate list is constructed, the transformation kernel of the current block can be further determined.

[0255] In one possible implementation, for determining the transform core of the current block, the first candidate list indicates the transform cores included in at least two transform core groups. In this case, referring to FIG. 27 , the method may include:

[0256] S2701: Determine the transform core index number of the current block.

[0257] S2702 : Determine a transform core of the current block according to the first candidate list and the transform core index number of the current block.

[0258] In an embodiment of the present application, assuming that the first candidate list indicates the transformation cores included in two candidate transformation core groups, if the first candidate transformation core group includes 3 candidate transformation cores and the second candidate transformation core group includes 2 candidate transformation cores, the first candidate list can indicate 5 candidate transformation cores; if the first candidate transformation core group includes 3 candidate transformation cores and the second candidate transformation core group includes 3 candidate transformation cores, the first candidate list can indicate 6 candidate transformation cores.

[0259] In the embodiment of the present application, the transform core index number of the current block is first determined, and then the transform core of the current block is determined in the first candidate list according to the transform core index number.

[0260] It should be noted that, in the embodiment of the present application, the transform core index number of the current block can be a positive integer, such as 1, 2, 3, 4, 5, 6, etc. The transform core index number of the current block can be determined by directly decoding the code stream, or by decoding the value of the first syntax element.

[0261] For example, one possible implementation is to decode the bitstream and determine the transform core index number of the current block. Alternatively, another possible implementation is to decode the bitstream and determine the value of the first syntax element; when the first syntax element indicates that the current block uses the first transform mode, determine the transform core index number of the current block based on the value of the first syntax element.

[0262] It should also be noted that, in the embodiment of the present application, the first transform mode may be LFNST / NSPT, and the first syntax element may be represented by lfnst_idx. The first syntax element may be used to indicate whether the current block uses the first transform mode, and the corresponding transform core index number when the current block uses the first transform mode.

[0263] It should also be noted that in this embodiment of the present application, if the value of the first syntax element is the first value, it is determined that the current block does not use the first transform mode; if the value of the first syntax element is the second value, it is determined that the current block uses the first transform mode and the corresponding transform core index number. The first value can be set to 0, and the second value can be set to a non-zero value, such as 1, 2, 3, 4, 5, 6, etc.

[0264] That is, in the embodiment of the present application, for LFNST / NSPT, the transform core index number of the current block can also be represented by lfnst_idx. Among them, lfnst_idx is 0, which means that the current block does not use LFNST / NSPT. Each transform core group of LFNST / NSPT in ECM has 3 transform cores, so the value of lfnst_idx is 1, 2, or 3, which means that the current block uses the first transform core, the second transform core, or the third transform core of the selected transform core group of LFNST / NSPT.

[0265] In this embodiment of the present application, if the prediction mode of the current block is a special intra-frame prediction mode, then it has more than one selectable transform core group. For example, if it has two selectable transform core groups, then the possible values ​​of lfnst_idx are 0, 1, 2, 3, 4, 5, and 6. 1, 2, and 3 correspond to the three transform cores of the first transform core group, and 4, 5, and 6 correspond to the three transform cores of the second transform core group.

[0266] In a specific embodiment, the binary symbol correspondence table of lfnst_idx is shown in Table 7.

[0267] Table 7

[0268] The third binary symbol, that is, the binary symbol with BinIdx being 2, can also be understood as selecting the first candidate transformation core group or the second candidate transformation core group.

[0269] In another possible implementation, for determining the transformation kernel of the current block, referring to FIG. 28 , the method may include:

[0270] S2801, decoding the code stream and determining the feature index number of the current block.

[0271] S2802: Determine a transform core group for the current block according to the first candidate list and the feature index sequence number.

[0272] S2803 : Determine a transform core of the current block according to the transform core group and the transform core index number of the current block.

[0273] In an embodiment of the present application, the transform core group is one of the at least two candidate transform core groups indicated by the first candidate list, which may be a transform core group determined by a texture feature index of the current block. In some embodiments, determining the transform core group for the current block may include: determining a texture feature index for the current block based on the first candidate list and a feature index sequence number; and determining the transform core group for the current block based on the texture feature index.

[0274] It should be noted that, in an embodiment of the present application, the feature index number of the current block is used to indicate the number of the transform core group of the current block in the first candidate list, and the feature index number can be represented by lfnst_feature_idx. Among them, the feature index number of the current block can be an integer greater than or equal to zero, such as 0, 1, 2, 3, 4, and so on. Exemplarily, if the first candidate list indicates two candidate transform core groups, then lfnst_feature_idx is used to indicate which candidate transform core group is specifically selected. For example, if the value of lfnst_feature_idx is equal to 0, it indicates that the first candidate transform core group indicated by the first candidate list is selected; if the value of lfnst_feature_idx is equal to 1, it indicates that the second candidate transform core group indicated by the first candidate list is selected.

[0275] It should be noted that after determining the transform core group of the current block, the transform core of the current block can be determined based on the transform core group. Specifically, this can be done by first determining the transform core index number of the current block, and then determining the transform core of the current block based on the transform core group and the transform core index number.

[0276] In the embodiment of the present application, the transform core index number of the current block may be a positive integer, such as 1, 2, 3, 4, 5, 6, etc. The transform core index number of the current block may be determined by directly decoding the bitstream, or by decoding the value of the first syntax element.

[0277] For example, one possible implementation is to decode the bitstream and determine the transform core index number of the current block. Alternatively, another possible implementation is to decode the bitstream and determine the value of the first syntax element; when the first syntax element indicates that the current block uses the first transform mode, determine the transform core index number of the current block based on the value of the first syntax element.

[0278] It should also be noted that in an embodiment of the present application, the first transform mode may be LFNST / NSPT, and the first syntax element may be represented by lfnst_idx. The first syntax element may be used to indicate whether the current block uses the first transform mode, and the corresponding transform core index number when the current block uses the first transform mode. Exemplarily, in this case, for LFNST / NSPT, the transform core index number of the current block may also be represented by lfnst_idx. In VVC, lfnst_idx can have three values: 0, 1, and 2. A lfnst_idx of 0 indicates that the current block does not use LFNST. Each transform core group of LFNST in VVC has two transform cores, so a lfnst_idx value of 1 or 2 indicates that the current block uses the first or second transform core of the selected transform core group of LFNST. In existing ECMs, lfnst_idx can have four values: 0, 1, 2, and 3. Where lfnst_idx is 0, which means that the current block does not use LFNST / NSPT. Each transform core group of LFNST / NSPT in ECM has 3 transform cores, so the value of lfnst_idx is 1, 2, or 3, which means that the current block uses the first transform core, the second transform core, or the third transform core of the selected transform core group of LFNST / NSPT.

[0279] In a specific embodiment, the binary symbol correspondence table of lfnst_idx is shown in Table 8.

[0280] Table 8

[0281] It should also be noted that in this embodiment of the present application, when decoding the bitstream and determining the feature index number of the current block, the method may further include: when the current block uses the first transform mode, decoding the bitstream and determining the feature index number of the current block. In other words, when the prediction mode of the current block satisfies the first condition and lfnst_idx>0, the step of decoding the bitstream and determining the feature index number of the current block is performed.

[0282] S2604: Determine the transformation coefficients of the current block, and perform inverse transformation on the transformation coefficients of the current block according to the transformation kernel to determine a residual block of the current block.

[0283] It should be noted that, in the embodiment of the present application, determining the transformation coefficient of the current block may include: decoding the code stream to determine the quantization coefficient of the current block; and dequantizing the quantization coefficient of the current block to determine the transformation coefficient of the current block.

[0284] It should also be noted that, in an embodiment of the present application, when performing an inverse transform on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block, it can include: performing an inverse transform of an inseparable basic transform on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block; or, performing an inverse transform of a low-frequency inseparable transform on the transform coefficients of the current block according to the transform kernel to determine the transform block of the current block; and performing an inverse transform of a discrete cosine transform on the transform block of the current block to determine the residual block of the current block.

[0285] In a specific embodiment, if the size parameter of the current block meets the first condition, the inverse transform of the inseparable basic transform is performed on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block; if the size parameter of the current block meets the second condition, the inverse transform of the low-frequency inseparable transform is performed on the transform coefficients of the current block according to the transform kernel to determine the transform block of the current block; and the inverse transform of the discrete cosine transform is performed on the transform block of the current block to determine the residual block of the current block.

[0286] Here, the size parameter of the current block satisfies the first condition, including: the size parameter of the current block is relatively small, for example, the size parameter of the current block is less than a certain threshold. In other words, for relatively small blocks, an NSPT transform kernel is used, i.e., an inverse NSPT transform is performed on the transform coefficients of the current block according to the transform kernel to determine a residual block for the current block.

[0287] Here, the size parameter of the current block satisfies the second condition, including: the size parameter of the current block is relatively large, for example, the size parameter of the current block is greater than a certain threshold. In other words, for relatively large blocks, an LFNST transform kernel is used. Specifically, an inverse LFNST transform is performed on the transform coefficients of the current block based on the transform kernel to determine the transform block of the current block; and an inverse DCT2 transform is performed on the transform block of the current block to determine the residual block of the current block.

[0288] It should also be noted that in the embodiments of the present application, the "inverse transformation" of the transform coefficients at the decoding end may also be referred to as "transformation" in the standard text. The "transformation" and "inverse transformation" in this article correspond to two opposite processes. For example, if the "transformation" converts the numerical values ​​in the spatial domain to the coefficients in the frequency domain, then the "inverse transformation" converts the coefficients in the frequency domain to the numerical values ​​in the spatial domain. "Inverse" is relative to "positive", and they are essentially both transformations. It should be noted that if the standard only stipulates decoding, then the "transformation" in the standard text is the decoding part, specifically referring to the "inverse transformation" in this article. The "inverse transformation" of the transform coefficients at the decoding end may also be referred to as "transformation" in the standard text.

[0289] In some embodiments, referring to FIG. 29 , after step S2604, the method may further include:

[0290] S2901: Perform intra-frame prediction on the current block to determine a prediction block for the current block.

[0291] S2902 : Determine a reconstructed block of the current block according to the prediction block of the current block and the residual block of the current block.

[0292] It should be noted that in the embodiment of the present application, for step S2901, this step can be performed in parallel with steps S2601 to S2603, or it can be performed before steps S2601 to S2603. The order of the steps is not specifically limited here.

[0293] It should also be noted that, in the embodiment of the present application, after the prediction block of the current block is determined, an addition operation may be performed on the prediction block of the current block and the residual block of the current block to determine the reconstructed block of the current block.

[0294] Simply put, after determining the prediction block, the decoder derives candidate texture feature indices based on the prediction block. The decoder then uses these indices to determine the transform core group for NSPT / LFNST. If a transform core group has multiple selectable transform cores, the decoder determines the transform core by decoding the lfnst_nspt_idx field in the bitstream. The decoding of lfnst_nspt_idx in the bitstream is independent of the process of determining the transform core group. Quantized coefficients are obtained from the bitstream through entropy decoding.

[0295] If it is NSPT transformation, the quantized coefficients are inversely quantized to obtain decoded transform coefficients, the decoded transform coefficients are inversely NSPT transformed to obtain decoded residual blocks, and finally a reconstructed block is obtained based on the decoded residual block and the prediction block.

[0296] If it is LFNST transform, the quantized coefficients are inversely quantized to obtain decoded transform coefficients, the decoded transform coefficients are inversely LFNST transformed, and then inverse DCT2 transformed to obtain a decoded residual block, and finally a reconstructed block is obtained based on the decoded residual block and the prediction block.

[0297] In addition, in some embodiments, referring to FIG. 30 , after step S2601, the method may further include:

[0298] S3001: When the prediction mode of the current block does not satisfy the first condition, determine the texture feature index of the current block.

[0299] S3002: Determine a transformation kernel of the current block according to the texture feature index.

[0300] S3003 , determining a transform coefficient of the current block, and performing an inverse transform on the transform coefficient of the current block according to the transform kernel to determine a residual block of the current block.

[0301] It should be noted that, in an embodiment of the present application, determining the transform core of the current block according to the texture feature index may include: determining the transform core group of the current block according to the texture feature index; decoding the code stream to determine the transform core index number of the current block; and determining the transform core of the current block according to the transform core group and the transform core index number.

[0302] It should be noted that, in the embodiment of the present application, the prediction mode of the current block not meeting the first condition may include: the prediction mode of the current block is a prediction mode outside the first prediction mode set. Alternatively, the prediction mode of the current block not meeting the first condition may include: the prediction mode of the current block is one of the items in the second prediction mode set.

[0303] It should be noted that, in the embodiment of the present application, the second prediction mode set includes at least: DC mode, PLANAR mode, angular prediction mode and cross-component prediction mode.

[0304] It should also be noted that, in the embodiment of the present application, the modes in the second prediction mode set are relatively simple prediction modes. For example, a classification can be made, where the DC mode, PLANAR mode, and various angle prediction modes are classified as the first prediction mode set, and the DIMD mode, TIMD mode, MIP mode, SGPM mode, ITMP mode, and IBC mode are classified as the second prediction mode set.

[0305] For example, the second prediction mode set may include only one or more of the DIMD mode, TIMD mode, MIP mode, SGPM mode, ITMP mode, and IBC mode. For example, the second prediction mode set includes the DIMD mode, TIMD mode, and SGPM mode. The other modes all belong to the first prediction mode set. The encoding method of lfnst_idx of the first prediction mode set is the same as that of the related art, while the encoding method of lfnst_idx of the second prediction mode set is different from that of the related art.

[0306] For example, each transform core group in LFNST / NSPT has three transform cores. Blocks using the first prediction mode set can only select one transform core group, while blocks using the second prediction mode set can select two transform core groups, each with three transform cores. If the prediction mode of the current block belongs to the first prediction mode set, it can only select one texture feature index, and the possible values ​​of lfnst_idx are 0, 1, 2, and 3. If the prediction mode of the current block belongs to the second prediction mode set, the possible values ​​of lfnst_idx are 0, 1, 2, 3, 4, 5, and 6. 1, 2, and 3 correspond to the three transform cores of the transform core group corresponding to the first texture feature index, and 4, 5, and 6 correspond to the three transform cores of the transform core group corresponding to the second texture feature index. The binary symbol correspondence table for lfnst_idx for blocks using the first prediction mode set is shown in Table 8, and the binary symbol correspondence table for lfnst_idx for blocks using the second prediction mode set is shown in Table 7.

[0307] An embodiment of the present application provides a decoding method, specifically an intra-frame LFNST / NSPT multi-angle selection scheme. First, the prediction mode of the current block is determined; when the prediction mode of the current block meets the first condition, the first candidate list of the current block is determined; wherein the first candidate list indicates at least two candidate transform core groups; then, based on the first candidate list, the transform core of the current block is determined; the transform coefficient of the current block is determined, and the transform coefficient of the current block is inversely transformed according to the transform core to determine the residual block of the current block. In this way, when the prediction mode of the current block meets the first condition, it is necessary to determine the first candidate list indicating at least two candidate transform core groups, and then determine the transform core of the current block based on these at least two candidate transform core groups. That is to say, for the current block predicted using certain intra-frame prediction modes, when determining the transform core of the current block, it is no longer determined based on only one transform core group, but is determined using at least two candidate transform core groups, thereby improving the accuracy of the transformation, thereby improving the compression efficiency, and thus improving the encoding and decoding performance.

[0308] In another embodiment of the present application, FIG31 is a flow chart of a coding method provided in an embodiment of the present application. As shown in FIG31 , the method may include:

[0309] S3101, determine the prediction mode of the current block.

[0310] It should be noted that in the embodiments of the present application, the method is applied to an encoder. Specifically, based on the structure of encoder 200 shown in FIG24 , the encoding method of the embodiments of the present application is primarily applied to intra-predicted blocks. Specifically, when the current block uses intra-prediction mode, the optimization scheme proposed here is mainly for the NSPT and LFNST transforms in intra-prediction mode to improve compression efficiency.

[0311] It should also be noted that in the embodiments of the present application, NSPT and LFNST are both transformations for processing textures at various angles. They may have multiple transformation kernels, and one transformation kernel may be specifically optimized for a certain specific angle texture. Of course, in addition to angle textures, NSPT and LFNST also include transformation kernels for processing gradient textures. In fact, these transformation kernels can also be said to be trained KL transforms (Karhunen-Loeve Transform, KLT). That is to say, NSPT and LFNST both have multiple transformation kernels, each of which is designed for a specific texture, and specific textures include angle textures, gradient textures, etc. In addition, gradient textures can be further extended to include horizontal gradient textures, vertical gradient textures, oblique gradient textures, etc. Furthermore, this technical solution is not limited to being used only for inseparable transformations such as NSPT and LFNST, and this technical solution can also be applied to separable transformations optimized for specific textures.

[0312] It should also be noted that in the intra-frame prediction blocks, they have multiple transform kernel groups, and each intra-frame prediction mode can correspond to a transform kernel group. In other words, each intra-frame prediction mode actually represents a texture feature. Therefore, the intra-frame prediction mode index is also a texture feature index. For example, DC and PLANAR correspond to gradient texture features, and a certain angle prediction mode corresponds to the texture feature of this angle. On the one hand, the texture feature index can avoid the appearance of intra-frame prediction mode in "inter-frame", and on the other hand, it is also more conducive to possible expansion. For example, an intra-frame prediction mode can correspond to multiple texture features, such as the DC mode can correspond to horizontal gradient texture, vertical gradient texture, oblique gradient texture, etc.

[0313] In the embodiments of the present application, it is considered that some intra-frame prediction modes are not simple texture features, but may contain two or more texture features. Therefore, it is necessary to first determine the prediction mode of the current block. In some embodiments, the method may include: determining multiple candidate prediction modes; calculating the encoding cost of the current block based on the multiple candidate prediction modes, and determining the cost results corresponding to the multiple candidate prediction modes; determining the minimum cost result among the cost results corresponding to the multiple candidate prediction modes; and determining the candidate prediction mode corresponding to the minimum cost result as the prediction mode of the current block.

[0314] Furthermore, in some embodiments, the method may further include: performing encoding processing on the prediction mode of the current block, and writing the obtained encoding bits into the bitstream.

[0315] That is, in the embodiment of the present application, after determining the prediction mode of the current block, some mode indication information (or mode flag) in the form of syntax elements can be written into the bitstream. In this way, by parsing the value of the mode indication information (or mode flag) in the bitstream, the prediction mode of the current block can be determined.

[0316] S3102 : When the prediction mode of the current block satisfies a first condition, determine a first candidate list for the current block; wherein the first candidate list indicates at least two candidate transform core groups.

[0317] It should also be noted that, in an embodiment of the present application, the modes in the first prediction mode set are relatively complex prediction modes, and the corresponding textures may contain two or more texture features. For example, both the DIMD mode and the TIMD mode can weight the prediction values ​​of two or more intra-frame prediction modes, the SGPM mode weights the prediction values ​​of the two intra-frame prediction modes with a weight matrix, the MIP mode predicts based on a matrix operation, and the ITMP mode and the IBC mode predict based on copying a reconstructed reference block. Their texture features are not as simple as those of the DC mode, the PLANAR mode, etc., so it is necessary to determine the first candidate list here. Among them, the first candidate list may include at least two candidate texture feature indexes, or the first candidate list may include at least two transform core groups. Here, each candidate texture feature index corresponds to a transform core group. Therefore, it can be said that the first candidate list indicates at least two candidate transform core groups.

[0318] In some embodiments, determining a first candidate list for a current block may include: determining candidate pixels for deriving texture feature indexes; determining one or more candidate transform kernel groups for the current block based on the candidate pixels; and adding the one or more candidate transform kernel groups to the first candidate list. Here, when determining the one or more candidate transform kernel groups for the current block based on the candidate pixels, one or more candidate texture feature indexes for the current block may be first determined based on the candidate pixels, and then the one or more candidate transform kernel groups for the current block are determined based on the one or more candidate texture feature indexes.

[0319] It should also be noted that, generally, one candidate texture feature index corresponds to one candidate transform core group; however, in some cases, multiple similar candidate texture feature indexes may correspond to the same transform core group. For example, multiple intra-frame prediction modes with similar angles correspond to the same transform core group. In an embodiment of the present application, if multiple adjacent intra-frame prediction modes (or candidate texture feature indexes) correspond to one transform core group, then when determining the candidate transform core group, it is necessary to ensure that the candidate texture feature indexes are not determined to be the same transform core group.

[0320] In a possible implementation, for candidate pixels, a prediction block of the current block may be determined; and at least part of the pixels in the prediction block are used as candidate pixels.

[0321] In the embodiment of the present application, if a certain texture exists in the prediction block, it can be considered that the residual block has a texture with the same characteristics. In this way, the candidate pixels used for deducing the candidate texture feature index between frames can be all pixels in the prediction block or part of the pixels in the prediction block.

[0322] In another possible implementation, for candidate pixels, adjacent pixels of a reconstructed area of ​​the current block may be determined; and the adjacent pixels of the reconstructed area may be used as candidate pixels.

[0323] In the embodiment of the present application, the candidate pixels used to derive the candidate texture feature index between frames can be pixels adjacent to the reconstructed area of ​​the current block, such as the reconstructed areas to the left and right of the current block. Because the reconstructed areas to the left and above are not the current block but are adjacent to the current block, for example, if the textures are connected, they can be used to estimate the texture of the current block to a certain extent.

[0324] In yet another possible implementation, more pixels are considered. For candidate pixels, adjacent pixels in the reconstructed area and at least part of the pixels in the prediction block may be used as candidate pixels.

[0325] In an embodiment of the present application, the candidate pixels used to derive the candidate texture feature index between frames can also use the predicted block of the current block and the reconstructed areas on the left and above the current block at the same time. In this way, more pixels are used to derive the candidate texture feature index, making the derived candidate texture feature index more accurate.

[0326] It should also be noted that, in the embodiment of the present application, the number of candidate pixels used to derive the candidate texture feature index may be at least one, for example, 1, 2, 3 or more.

[0327] In some embodiments, the method may further include: determining the number of candidate pixels according to a size parameter of the current block.

[0328] That is to say, when deriving one or more candidate texture feature indexes based on candidate pixels, the number of candidate pixels used can be determined by the size parameter of the current block. For example, if the size of the current block is small, then all available pixels can be counted; if the size of the current block is large, then the current block can be downsampled and counted, such as counting one pixel out of every 2, or 4, or 8 pixels in the horizontal and / or vertical directions. Alternatively, if the size of the current block in one of the horizontal or vertical directions is less than or equal to 8, then all available pixels in that direction are counted; otherwise, if the size of the current block in one of the horizontal or vertical directions is less than or equal to 16, then one pixel out of every 2 pixels in that direction is counted; otherwise, one pixel out of every 4 pixels in that direction is counted, and no specific limitation is given here.

[0329] In some embodiments, determining a candidate texture feature list for a current block based on candidate pixels may include: determining a horizontal gradient value and a vertical gradient value of the candidate pixel; determining a texture feature index and a gradient intensity value corresponding to the candidate pixel based on the horizontal gradient value and the vertical gradient value of the candidate pixel; constructing a texture feature statistics table based on the texture feature index and the gradient intensity value corresponding to the candidate pixel; and determining a candidate texture feature list for the current block based on the texture feature statistics table.

[0330] It should be noted that in an embodiment of the present application, when determining the texture feature index and gradient intensity value corresponding to the candidate pixel based on the horizontal gradient value and vertical gradient value of the candidate pixel, it can include: performing angle mapping based on the horizontal gradient value and vertical gradient value of the candidate pixel to determine the texture feature index corresponding to the candidate pixel; and performing gradient intensity calculation based on the horizontal gradient value and vertical gradient value of the candidate pixel to determine the gradient intensity value corresponding to the candidate pixel.

[0331] In a specific embodiment, performing angle mapping based on the horizontal gradient value and the vertical gradient value of the candidate pixel to determine the texture feature index corresponding to the candidate pixel may include: determining the texture feature index corresponding to the candidate pixel using a preset lookup table based on the horizontal gradient value and the vertical gradient value of the candidate pixel.

[0332] In the embodiment of the present application, the horizontal gradient value of the candidate pixel can be expressed as grad x Indicates that the vertical gradient value of the candidate pixel can be expressed as grad y In this way, according to grad x and grad y Deriving the texture feature index (or referred to as a “virtual intra prediction mode”) can be achieved by looking up a table.

[0333] For example, if abs(grad x) is equal to 0 and abs(grad y ) is not equal to 0, then there is horizontal texture, corresponding to intra prediction mode 18 in VVC. y ) is equal to 0 and abs(grad x ) is not equal to 0, then there is vertical texture, corresponding to intra prediction mode 50 in VVC. x ) and abs(grad y ) are not equal to 0, if abs(grad x ) is equal to abs(grad y ), and grad x and grad y The symbols are the same, corresponding to intra prediction mode 34 in VVC. If abs(grad x ) is equal to 2 times abs(grad y ), and grad x and grad y The symbols are the same, corresponding to intra prediction mode 40 in VVC. In addition, other situations can be determined by looking up the table according to the same principle.

[0334] In a specific embodiment, performing gradient strength calculation based on the horizontal gradient value and the vertical gradient value of the candidate pixel to determine the gradient strength value corresponding to the candidate pixel may include: performing an addition operation on the absolute value of the horizontal gradient value and the absolute value of the vertical gradient value to determine the gradient strength value corresponding to the candidate pixel.

[0335] Here, the gradient intensity value corresponding to the candidate pixel can be recorded as amp. For example, amp=abs(grad x )+abs(grad y ).

[0336] It should be noted that, in the embodiment of the present application, the horizontal gradient value and the vertical gradient value of the candidate pixel can be calculated using the Sobel operator. For example, the Sobel operator is as follows:

[0337] Operator for horizontal gradient value:

[0338] Operator for vertical gradient value:

[0339] So, suppose the pixel value at pixel position (x, y) is P x,y , then the horizontal gradient value grad x And the vertical gradient value grad y The calculation of grad is as follows:x =P x+1,y-1 +2*P x+1,y +P x+1,y+1 -P x-1,y-1 -2*P x-1,y -P x-1,y+1 (7) grad y =P x-1,y+1 +2*P x,y+1 +P x+1,y+1 -P x-1,y-1 -2*P x,y-1 -P x+1,y-1 (8)

[0340] It should also be noted that, in this embodiment of the present application, the candidate pixels may be set to exclude the pixels in the outermost row, column, and top, bottom, left, and right of the current block. Considering that the Sobel operator uses the pixels in the outermost row, column, and top, bottom, left, and right of the current pixel, the gradient of the pixels in the outermost row, column, and top, bottom, left, and right of the current pixel may not be calculated in this embodiment of the present application.

[0341] In some embodiments, constructing a texture feature statistics table based on the texture feature index and gradient intensity value corresponding to the candidate pixel may include: determining at least one texture feature index and at least one corresponding gradient intensity value when the number of candidate pixels is at least one; determining at least one reference texture feature index with mutually different characteristics based on the at least one texture feature index, and accumulating the gradient intensity values ​​belonging to the same reference texture feature index based on the at least one gradient intensity value to determine the gradient intensity cumulative value corresponding to the at least one reference texture feature index; constructing a texture feature statistics table based on the at least one reference texture feature index and the gradient intensity cumulative value corresponding to the at least one reference texture feature index.

[0342] That is to say, in an embodiment of the present application, taking at least part of the pixels in the prediction block as candidate pixels as an example, the gradients of all or part of the pixels in the prediction block are calculated. Generally speaking, the horizontal gradient value and the vertical gradient value can be calculated. Here, the Sobel operator can be used to calculate the gradient value. For a certain pixel, the texture direction of the pixel can be inferred based on its horizontal gradient value and vertical gradient value. For example, if the horizontal gradient value is non-zero and the vertical gradient value is zero, then the texture of the pixel is in the vertical direction. Conversely, if the horizontal gradient value is zero and the vertical gradient value is non-zero, then the texture of the pixel is in the horizontal direction. For example, if the horizontal gradient value and the vertical gradient value are equal and not zero, then the texture of the pixel is 45 degrees. Of course, there are many other cases in the embodiment of the present application where the horizontal gradient value and the vertical gradient value are not zero, and the texture direction of the pixel can be determined based on their ratio. In this way, the gradient intensity value of each pixel can be corresponded to the corresponding texture feature index. A texture feature statistics table is constructed, and the gradient intensity value of each calculated pixel is added to the corresponding texture feature index item in the statistics table to obtain the final texture feature statistics table. Then, the candidate texture feature list of the current block can be determined based on the texture feature statistics table.

[0343] In some embodiments, determining a candidate texture feature list for a current block based on a texture feature statistics table may include: sorting the texture feature statistics table from high to low according to gradient intensity accumulation values, determining reference texture feature indexes corresponding to top N gradient intensity accumulation values, and adding the N reference texture feature indexes to the candidate texture feature list for the current block, where N is a positive integer.

[0344] That is, in the embodiment of the present application, after the texture feature statistics table is constructed, in order to reduce complexity, the texture feature statistics table can also be sorted in descending order according to the accumulated gradient strength cumulative values, and then only the top N reference texture feature indexes are selected. Here, for complexity considerations, only a candidate texture feature list of length N can be maintained, and the candidate texture feature list is sorted from high to low according to the gradient strength cumulative values.

[0345] It should also be noted that in the embodiments of the present application, the value of N can be 2, 3, 4, 5, ..., 10, etc., and there is no specific limitation on the value of N here.

[0346] It can be understood that in the embodiment of the present application, considering that the angles of some adjacent intra-frame prediction modes are very close, in order to exclude intra-frame prediction modes that are too close, the method may also include: pruning the N reference texture feature indexes in the candidate texture feature list to determine one or more candidate transform core groups for the current block.

[0347] In some embodiments, pruning N reference texture feature indexes in a candidate texture feature list to determine one or more candidate transform core groups for a current block may include: determining a first candidate transform core group based on a reference texture feature index at a first position in the candidate texture feature list; when a second condition is satisfied between other reference texture feature indexes other than the first position in the candidate texture feature list and the i-th candidate transform core group, determining an i+1-th candidate transform core group based on the other reference texture feature indexes to determine one or more candidate transform core groups for the current block; wherein i is an integer greater than zero and less than N.

[0348] It should be noted that, in the embodiment of the present application, the second condition may include: the difference between the i-th candidate texture feature index and the i+1-th candidate texture feature index satisfies a preset threshold, i.e., there is a certain difference between two adjacent candidate texture feature indexes. Alternatively, when each candidate texture feature index corresponds to a candidate transform core group, it can also be said that the difference between the i-th candidate transform core group and the i+1-th candidate transform core group satisfies the preset threshold.

[0349] It should also be noted that, in the embodiment of the present application, the N reference texture feature indexes are sorted from high to low according to the accumulated gradient strength values. In this case, determining the first candidate transform kernel group based on the reference texture feature index that is first in the candidate texture feature list may include determining the first candidate transform kernel group based on the reference texture feature index with the largest accumulated gradient strength value in the candidate texture feature list.

[0350] It should also be noted that, in the embodiment of the present application, the preset threshold can be represented by THR, the i-th candidate texture feature index can be represented by candFeature(i), and the i+1-th candidate texture feature index can be represented by candFeature(i+1). In a specific embodiment, the difference between the i-th candidate texture feature index and the i+1-th candidate texture feature index meets the preset threshold, which can include: candFeature(i+1)+THR<candFeature(i)||candFeature(i+1)-THR> candFeature(i).

[0351] It should also be noted that, in the embodiment of the present application, the value of THR may be 3, 4, 5, 6, etc. Thus, when the cumulative values ​​of multiple adjacent angles are very high, this pruning method will give priority to candidate texture feature indexes with a certain degree of discrimination.

[0352] It is also understandable that in the embodiments of the present application, the above method does not take into account some intra-frame prediction modes derived from the DIMD, TIMD, SGPM, and other modes themselves. Therefore, in some embodiments, the method may further include: determining one or more intra-frame prediction modes derived from the prediction mode of the current block; determining one or more candidate transform core groups based on the one or more intra-frame prediction modes, and adding the one or more candidate transform core groups to the first candidate list.

[0353] In some embodiments, the method further includes: when the first candidate list is not full, determining a preset texture feature index of the current block; determining one or more candidate transform core groups according to the preset texture feature index, and adding the one or more candidate transform core groups to the first candidate list.

[0354] In some embodiments, the method further includes: when the first candidate list is not filled, determining candidate pixels for deriving texture feature indexes; determining a candidate texture feature list of the current block based on the candidate pixels; determining one or more candidate transform kernel groups based on the candidate texture feature list, and adding the one or more candidate transform kernel groups to the first candidate list.

[0355] That is to say, in the embodiment of the present application, some modes derived from the DIMD, TIMD, SGPM and other modes themselves are taken into consideration. For example, DIMD itself will derive one or several intra-frame prediction modes for weighting, and TIMD itself will also derive one or several intra-frame prediction modes for weighting. SGPM not only has two intra-frame prediction modes, but also has a "partitioning" mode that can also find the corresponding intra-frame prediction mode, and the residual often appears in the boundary area of ​​the "partition". Therefore, one possible implementation method is to determine the first candidate list based on the intra-frame prediction mode derived from the DIMD, TIMD, SGPM and other modes themselves and the candidate texture feature index derived by the above method; or another possible implementation method is to determine the first candidate list based on the intra-frame prediction mode derived from the DIMD, TIMD, SGPM and other modes themselves and the default texture feature index.

[0356] In addition, in the embodiment of the present application, since DIMD, TIMD, and SGPM can all derive more than one intra-frame prediction mode, these modes can also give priority to using multiple modes derived by each mode itself to determine the candidate texture feature index. When the modes derived by each mode itself cannot fill all the candidate texture feature indexes, one possible method is to add a default texture feature index. Another possible method is to use the above-mentioned sorted intra-frame prediction mode list of length N to determine the candidate texture feature index.

[0357] For example, when the prediction mode of the current block is the DIMD mode, if it uses the prediction values ​​of multiple intra-frame prediction modes for weighting during prediction, then the multiple intra-frame prediction modes used for weighting can be sequentially tried to be determined as candidate texture feature indexes.

[0358] For example, when the prediction mode of the current block is the TIMD mode, if it uses the prediction values ​​of multiple intra-frame prediction modes for weighting during prediction, then the multiple intra-frame prediction modes used for weighting can be sequentially tried to be determined as candidate texture feature indexes.

[0359] For example, when the prediction mode of the current block is the SGPM mode, the intra-frame prediction mode corresponding to the "partitioning" mode derived therefrom and the two intra-frame prediction modes used for prediction may be sequentially attempted to be determined as candidate texture feature indexes.

[0360] In another possible implementation, for determining the first candidate list of the current block, the method may include: determining a first candidate transformation core group based on the reference texture feature index at the first position in the candidate texture feature list; when the second condition is not satisfied between all reference texture feature indexes except the first position in the candidate texture feature list and the first candidate transformation core group, determining a second candidate transformation core group based on the reference texture feature index at the second position in the candidate texture feature list; and determining the first candidate list based on the first candidate transformation core group and the second candidate transformation core group.

[0361] It should be noted that in an embodiment of the present application, the first candidate transformation core group is determined based on the reference texture feature index in the first position in the candidate texture feature list. Specifically, it can be: the first candidate transformation core group is determined based on the reference texture feature index with the largest gradient intensity accumulated value in the candidate texture feature list.

[0362] It should also be noted that, in the embodiment of the present application, taking the first candidate list indicating two candidate transform core groups as an example, the first candidate transform core group can be represented by candFeature0, and the second candidate transform core group can be represented by candFeature1. For example, the intra-frame prediction mode with the largest cumulative gradient strength value is the first candidate transform core group candFeature0, then when selecting the second candidate transform core group candFeature1, it is required that candFeature1 and candFeature0 have a certain gap, such as candFeature1+THR<candFeature0||candFeature1-THR> If no matching candidate is found after checking all N-1 reference texture feature indices, a second candidate transform core group, candFeature1, can be determined based on the second-ranked reference texture feature index among the N reference texture feature indices. In other words, the candidate transform core groups corresponding to the first two reference texture feature indices with the largest cumulative gradient strength values ​​can be added to the first candidate list.

[0363] In another possible implementation, for determining the first candidate list of the current block, the method may include: determining a first candidate transformation core group based on the reference texture feature index at the first position in the candidate texture feature list; when the second condition is not satisfied between all reference texture feature indexes other than the first position in the candidate texture feature list and the first candidate transformation core group, determining a second candidate transformation core group based on a preset texture feature index; and determining the first candidate list based on the first candidate transformation core group and the second candidate transformation core group.

[0364] It should be noted that in an embodiment of the present application, if no one meets the requirements after checking all N-1 reference texture feature indexes, then the default texture feature index of the current block can also be determined, and then the candidate transform core group corresponding to the reference texture feature index with the largest gradient intensity accumulation value and the default texture feature index is added to the first candidate list.

[0365] In another possible implementation, for determining the first candidate list of the current block, the method may include: determining a first candidate transform core group based on a first intra-frame prediction mode derived from a prediction mode of the current block; determining a second candidate transform core group based on the first reference texture feature index when a second condition is satisfied between the first reference texture feature index in the candidate texture feature list and the first candidate transform core group; and determining the first candidate list based on the first candidate transform core group and the second candidate transform core group.

[0366] It should be noted that in this embodiment of the present application, the first candidate list not only considers the candidate texture feature indexes derived from the horizontal and vertical gradient values ​​of the candidate pixels, but also considers the intra-frame prediction mode derived from the prediction mode of the current block itself. This is explained in detail below using several examples.

[0367] Exemplarily, when the prediction mode of the current block is the DIMD mode, the first intra prediction mode derived from the DIMD mode is used as candFeature0.

[0368] Exemplarily, when the prediction mode of the current block is the TIMD mode, the first intra-frame prediction mode derived from the TIMD mode is used as candFeature0.

[0369] Exemplarily, when the prediction mode of the current block is the SGPM mode, the intra-frame prediction mode corresponding to the SGPM "partition" mode is used as candFeature0.

[0370] Then, candFeature1 is determined according to the above method. For example, if a reference texture feature index among the N reference texture feature indexes is tried, starting from the first position, and a reference texture feature index meets the THR restriction, it can be used as candFeature1.

[0371] It should also be noted that in the embodiment of the present application, for modes such as IBC and ITMP, a prediction block or a prediction block plus a template can be used to derive a candidate texture feature index (or intra-frame prediction mode).

[0372] It should also be noted that in an embodiment of the present application, if the candidate pixels used in the embodiment of the present application to derive N reference texture feature indexes are the same as the candidate pixels used in the DIMD mode, then the first intra-frame prediction mode derived by DIMD and the first reference texture feature index derived by the embodiment of the present application are the same.

[0373] S3103: Determine a transformation kernel for the current block according to the first candidate list.

[0374] It should be noted that after the first candidate list is constructed, the transformation kernel of the current block can be further determined.

[0375] In one possible implementation, for determining the transform core of the current block, the first candidate list indicates transform cores included in at least two transform core groups. In this case, the transform core of the current block can be directly determined from the first candidate list based on cost calculation.

[0376] In an embodiment of the present application, assuming that the first candidate list indicates the transformation cores included in two candidate transformation core groups, if the first candidate transformation core group includes 3 candidate transformation cores and the second candidate transformation core group includes 2 candidate transformation cores, the first candidate list can indicate 5 candidate transformation cores; if the first candidate transformation core group includes 3 candidate transformation cores and the second candidate transformation core group includes 3 candidate transformation cores, the first candidate list can indicate 6 candidate transformation cores.

[0377] In some embodiments, determining the transformation core of the current block based on the first candidate list may include: determining at least two candidate transformation cores indicated by the first candidate list; performing encoding cost calculation on the current block based on the at least two candidate transformation cores, and determining the cost results corresponding to each of the at least two candidate transformation cores; determining the minimum cost result among the cost results corresponding to each of the at least two candidate transformation cores, and determining the candidate transformation core corresponding to the minimum cost result as the transformation core of the current block.

[0378] It should be noted that in the embodiment of the present application, the cost calculation here can be determined based on the cost result of Rate Distortion Optimization (RDO), or based on the cost result of Sum of Absolute Difference (SAD), or even based on the cost result of Sum of Absolute Transformed Difference (SATD), but no limitation is made here.

[0379] In a specific embodiment, encoding cost calculation is performed on the current block based on at least two candidate transform kernels to determine the cost results corresponding to each of the at least two candidate transform kernels, which may include: transforming and quantizing the residual block of the current block based on the first candidate transform kernel to determine the first candidate quantization coefficient of the current block, and performing entropy coding on the first candidate quantization coefficient to determine the first generation value of the first candidate transform kernel; inverse quantizing and inverse transforming the first candidate quantization coefficient to determine the first candidate residual block of the current block, and determining the first candidate prediction block of the current block based on the first candidate residual block; performing cost calculation based on the first candidate prediction block and the original image of the current block to determine the second generation value of the first candidate transform kernel; determining the cost result corresponding to the first candidate transform kernel based on the first generation value and the second generation value of the first candidate transform kernel; wherein the first candidate transform kernel is any one of the at least two candidate transform kernels.

[0380] In this embodiment of the present application, the transform core index number of the current block can be used to indicate the number of the transform core of the current block in the first candidate list. The transform core index number of the current block can be a positive integer, such as 1, 2, 3, 4, 5, 6, etc. The transform core index number can be written directly into the bitstream or written into the bitstream via the value of the first syntax element.

[0381] In a possible implementation, a transform core index number of a current block is determined; the transform core index number of the current block is coded, and the obtained coded bits are written into a bitstream.

[0382] In another possible implementation, a value of a first syntax element is determined; wherein the first syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform core index number used; the value of the first syntax element is encoded, and the obtained encoded bits are written into the bitstream.

[0383] It should also be noted that in this embodiment of the present application, the first transform mode can be LFNST / NSPT, and the first syntax element can be represented by lfnst_idx. The first syntax element can be used to indicate whether the current block uses the first transform mode, and the corresponding transform core index number when the current block uses the first transform mode. In this case, the value of the first syntax element can be 0, 1, 2, 3, 4, 5, 6, etc.

[0384] In a specific embodiment, if the value of the first syntax element is a first value, it is determined that the current block does not use the first transform mode; if the value of the first syntax element is a second value, it is determined that the current block uses the first transform mode and the corresponding transform core index number. The first value can be set to 0, and the second value can be set to a non-zero value, such as 1, 2, 3, 4, 5, 6, etc.

[0385] That is, in the embodiment of the present application, for LFNST / NSPT, the transform core index number of the current block can also be represented by lfnst_idx. Among them, lfnst_idx is 0, which means that the current block does not use LFNST / NSPT. Each transform core group of LFNST / NSPT in ECM has 3 transform cores, so the value of lfnst_idx is 1, 2, or 3, which means that the current block uses the first transform core, the second transform core, or the third transform core of the selected transform core group of LFNST / NSPT.

[0386] In this embodiment of the present application, if the prediction mode of the current block is a special intra-frame prediction mode, then it has more than one selectable transform core group. For example, if it has two selectable transform core groups, then the possible values ​​of lfnst_idx are 0, 1, 2, 3, 4, 5, and 6. 1, 2, and 3 correspond to the three transform cores of the first transform core group, and 4, 5, and 6 correspond to the three transform cores of the second transform core group.

[0387] In a specific embodiment, the binary symbol correspondence table of lfnst_idx is shown in the aforementioned Table 7. The third binary symbol, ie, the binary symbol with BinIdx being 2, can also be understood as selecting the first candidate transform core group or the second candidate transform core group.

[0388] In another possible implementation, the transform core of the current block is determined according to the first candidate list. The method may include: determining a transform core group of the current block according to the first candidate list; and determining the transform core of the current block according to the transform core group.

[0389] In an embodiment of the present application, the transform core group here is one of the at least two candidate transform core groups indicated by the first candidate list, which may be the transform core group determined by the texture feature index of the current block. In some embodiments, for determining the transform core group of the current block, the method may include: performing encoding cost calculation on the current block based on the at least two candidate transform core groups indicated by the first candidate list, determining cost results corresponding to each of the at least two candidate transform core groups; determining a minimum cost result among the cost results corresponding to each of the at least two candidate transform core groups, and determining the candidate transform core group corresponding to the minimum cost result as the transform core group of the current block.

[0390] Furthermore, in some embodiments, the method further includes: determining a feature index number of the current block; encoding the feature index number of the current block, and writing the obtained encoding bits into a bitstream.

[0391] It should be noted that, in an embodiment of the present application, the feature index number can be used to indicate the number of the transform core group of the current block in the first candidate list. The feature index number of the current block can be represented by lfnst_feature_idx. The feature index number of the current block can be an integer greater than or equal to zero, such as 0, 1, 2, 3, 4, and so on. Exemplarily, if the first candidate list indicates two candidate transform core groups, then lfnst_feature_idx is used to indicate which candidate transform core group is specifically selected. For example, if the value of lfnst_feature_idx is equal to 0, it indicates that the first candidate transform core group indicated by the first candidate list is selected; if the value of lfnst_feature_idx is equal to 1, it indicates that the second candidate transform core group indicated by the first candidate list is selected.

[0392] In some embodiments, after determining the transform core group of the current block, determining the transform core of the current block may include: determining at least two candidate transform cores included in the transform core group; performing encoding cost calculation on the current block based on the at least two candidate transform cores, and determining the cost results corresponding to each of the at least two candidate transform cores; determining the minimum cost result among the cost results corresponding to each of the at least two candidate transform cores, and determining the candidate transform core corresponding to the minimum cost result as the transform core of the current block.

[0393] In a specific embodiment, encoding cost calculation is performed on the current block based on at least two candidate transform kernels to determine the cost results corresponding to each of the at least two candidate transform kernels, which may include: transforming and quantizing the residual block of the current block based on the first candidate transform kernel to determine the first candidate quantization coefficient of the current block, and performing entropy coding on the first candidate quantization coefficient to determine the first generation value of the first candidate transform kernel; inverse quantizing and inverse transforming the first candidate quantization coefficient to determine the first candidate residual block of the current block, and determining the first candidate prediction block of the current block based on the first candidate residual block; performing cost calculation based on the first candidate prediction block and the original image of the current block to determine the second generation value of the first candidate transform kernel; determining the cost result corresponding to the first candidate transform kernel based on the first generation value and the second generation value of the first candidate transform kernel; wherein the first candidate transform kernel is any one of the at least two candidate transform kernels.

[0394] In this embodiment of the present application, the transform core index number of the current block can be used to indicate the number of the transform core of the current block in the transform core group of the current block. The transform core index number of the current block can be a positive integer, such as 1, 2, 3, etc. The transform core index number can be directly written into the bitstream, or can be written into the bitstream through the value of the first syntax element.

[0395] In a possible implementation, a transform core index number of a current block is determined; the transform core index number of the current block is coded, and the obtained coded bits are written into a bitstream.

[0396] In another possible implementation, a value of a first syntax element is determined; wherein the first syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform core index number used; the value of the first syntax element is encoded, and the obtained encoded bits are written into the bitstream.

[0397] It should also be noted that in this embodiment of the present application, the first transform mode can be LFNST / NSPT, and the first syntax element can be represented by lfnst_idx. The first syntax element can be used to indicate whether the current block uses the first transform mode, and the corresponding transform core index number when the current block uses the first transform mode. In this case, the value of the first syntax element can be 0, 1, 2, 3, etc.

[0398] Exemplarily, in this case, for LFNST / NSPT, the transform core index number of the current block can also be represented by lfnst_idx. In VVC, lfnst_idx can have three values: 0, 1, and 2. A lfnst_idx of 0 indicates that the current block does not use LFNST. Each transform core group of LFNST in VVC has two transform cores, so a lfnst_idx value of 1 or 2 indicates that the current block uses the first or second transform core of the selected transform core group of LFNST. In existing ECM, lfnst_idx can have four values: 0, 1, 2, and 3. A lfnst_idx of 0 indicates that the current block does not use LFNST / NSPT. Each transform core group of LFNST / NSPT in ECM has three transform cores, so a lfnst_idx value of 1, 2, or 3 indicates that the current block uses the first, second, or third transform core of the selected transform core group of LFNST / NSPT.

[0399] In a specific embodiment, the binary symbol correspondence table of lfnst_idx is shown in Table 8 above.

[0400] It should also be noted that, in this embodiment of the present application, when encoding the feature index number of the current block, the method may further include: when the current block uses the first transform mode, encoding the feature index number of the current block and writing the resulting coded bits into the bitstream. In other words, when the prediction mode of the current block satisfies the first condition and lfnst_idx>0, encoding the feature index number of the current block and writing the resulting coded bits into the bitstream is performed.

[0401] S3104: Determine a residual block of the current block, and transform the residual block of the current block according to the transformation kernel to determine a transformation coefficient of the current block.

[0402] S3105 , performing encoding processing on the transform coefficients of the current block, and writing the obtained encoding bits into the bitstream.

[0403] It should be noted that, in an embodiment of the present application, when encoding the transform coefficients of the current block, the method may include: quantizing the transform coefficients of the current block to determine the quantization coefficients of the current block; encoding the quantization coefficients of the current block and writing the obtained coded bits into the bitstream.

[0404] It should also be noted that in the embodiments of this application, the "transformation" of the residual block by the encoder can also be called a "forward transform," specifically referring to the transformation from the spatial domain to the frequency domain to remove residual correlation. It should be noted that if the standard only specifies decoding, then the "transformation" in the standard text refers to the decoding part, specifically referring to the "inverse transform" in this article.

[0405] It should also be noted that, in the embodiment of the present application, referring to FIG. 32 , for step S3104, the method may include:

[0406] S3201: Perform intra-frame prediction on the current block to determine a prediction block for the current block.

[0407] S3202 : Determine a residual block of the current block according to the initial block of the current block and the predicted block of the current block.

[0408] S3203 , transforming the residual block of the current block according to the transformation kernel to determine the transformation coefficient of the current block.

[0409] It should be noted that in the embodiment of the present application, steps S3201 to S3202 can be operated in parallel with steps S3101 to S3103, or can be executed before steps S3101 to S3103, or can be executed after steps S3101 to S3103. The order of the steps is not specifically limited here.

[0410] It should also be noted that, in the embodiment of the present application, after determining the prediction block of the current block, a subtraction operation may be performed on the initial block of the current block and the prediction block of the current block to determine the residual block of the current block.

[0411] It should also be noted that, in an embodiment of the present application, when transforming the residual block of the current block according to the transform kernel to determine the transform coefficient of the current block, it can include: performing an inseparable basic transform on the residual block of the current block according to the transform kernel to determine the transform coefficient of the current block; or performing a discrete cosine transform on the residual block of the current block to determine the transform block of the current block; and performing a low-frequency inseparable transform on the transform block of the current block according to the transform kernel to determine the transform coefficient of the current block.

[0412] In a specific embodiment, if the size parameter of the current block meets the first condition, an inseparable basic transform is performed on the residual block of the current block according to the transform kernel to determine the transform coefficient of the current block; if the size parameter of the current block meets the second condition, a discrete cosine transform is performed on the residual block of the current block to determine the transform block of the current block; and a low-frequency inseparable transform is performed on the transform block of the current block according to the transform kernel to determine the transform coefficient of the current block.

[0413] Here, the size parameter of the current block satisfies the first condition, including: the size parameter of the current block is relatively small, for example, the size parameter of the current block is less than a certain threshold. In other words, for relatively small blocks, an NSPT transform kernel is used. That is, an NSPT transform is performed on the residual block of the current block according to the transform kernel to determine the transform coefficients of the current block.

[0414] Here, the size parameter of the current block satisfies the second condition, including: the size parameter of the current block is relatively large, for example, the size parameter of the current block is greater than a certain threshold. In other words, for relatively large blocks, the LFNST transform kernel is used. Specifically, a basic DCT2 transform is first performed on the residual block of the current block, and then an LFNST transform is performed on the transform block of the current block based on the transform kernel to determine the transform coefficients of the current block.

[0415] In simple terms, after determining the prediction block, the encoder derives the candidate texture feature index based on the prediction block, and then determines the transform kernel group for NSPT / LFNST based on the candidate texture feature index. If a transform kernel group has multiple transform kernels to choose from, the encoder tries each transform kernel in the transform kernel group.

[0416] If the NSPT transform is used, the residual block can be forward transformed using NSPT to obtain transform coefficients, the transform coefficients are quantized to obtain quantized coefficients, and then the quantized coefficients are entropy encoded. Entropy encoding can be used to determine the overhead cost in the bitstream for this transform kernel. The quantized coefficients are inversely quantized to obtain decoded transform coefficients, and the decoded transform coefficients are inversely NSPT transformed to obtain the decoded residual block. The decoded transform coefficients may be different from the original transform coefficients because quantization is lossy. Similarly, the decoded residual block may also be different from the original residual block. A reconstructed block is obtained based on the decoded residual block and the prediction block. The distortion cost can be determined based on the reconstructed block and the original image of the current block. The cost of encoding using the current NSPT transform kernel is the sum of the overhead cost and the distortion cost. The costs of several transform kernels are compared, and the smallest one is selected as the optimal NSPT for the current block.

[0417] If the LFNST transform is used, the residual block can be forward transformed using DCT2, then forward transformed using LFNST to obtain transform coefficients. These transform coefficients are quantized to obtain quantized coefficients, and then entropy encoded. Entropy encoding can be used to determine the overhead cost in the bitstream for this transform kernel. The quantized coefficients are then inversely quantized to obtain decoded transform coefficients. These decoded transform coefficients are then subjected to an inverse LFNST transform and then an inverse DCT2 transform to obtain the decoded residual block. The decoded transform coefficients may differ from the original transform coefficients because quantization is lossy. Similarly, the decoded residual block may also differ from the original residual block. The decoded residual block and the prediction block are used to obtain a reconstructed block. The distortion cost can be determined based on the reconstructed block and the original image of the current block. The cost of encoding using the current NSPT transform kernel is the sum of the overhead cost and the distortion cost. The costs of several transform kernels are compared, and the one with the smallest cost is selected as the optimal NSPT for the current block.

[0418] In addition, in some embodiments, after step S3101, the method may further include: when the prediction mode of the current block does not meet the first condition, determining the texture feature index of the current block; determining the transformation kernel of the current block based on the texture feature index; determining the transformation coefficient of the current block, and performing an inverse transformation on the transformation coefficient of the current block based on the transformation kernel to determine the residual block of the current block.

[0419] It should be noted that, in the embodiment of the present application, determining the transformation kernel of the current block according to the texture feature index may include: determining the transformation kernel group of the current block according to the texture feature index; and determining the transformation kernel of the current block according to the transformation kernel group.

[0420] It should be noted that, in the embodiment of the present application, the prediction mode of the current block not meeting the first condition may include: the prediction mode of the current block is a prediction mode outside the first prediction mode set. Alternatively, the prediction mode of the current block not meeting the first condition may include: the prediction mode of the current block is one of the items in the second prediction mode set.

[0421] It should be noted that, in the embodiment of the present application, the second prediction mode set includes at least: DC mode, PLANAR mode, angular prediction mode and cross-component prediction mode.

[0422] It should also be noted that, in the embodiment of the present application, the modes in the second prediction mode set are relatively simple prediction modes. For example, a classification can be made, where the DC mode, PLANAR mode, and various angle prediction modes are classified as the first prediction mode set, and the DIMD mode, TIMD mode, MIP mode, SGPM mode, ITMP mode, and IBC mode are classified as the second prediction mode set.

[0423] For example, the second prediction mode set may include only one or more of the DIMD mode, TIMD mode, MIP mode, SGPM mode, ITMP mode, and IBC mode. For example, the second prediction mode set includes the DIMD mode, TIMD mode, and SGPM mode. The other modes all belong to the first prediction mode set. The encoding method of lfnst_idx of the first prediction mode set is the same as that of the related art, while the encoding method of lfnst_idx of the second prediction mode set is different from that of the related art.

[0424] For example, each transform core group in LFNST / NSPT has three transform cores. Blocks using the first prediction mode set can only select one transform core group, while blocks using the second prediction mode set can select two transform core groups, each with three transform cores. If the prediction mode of the current block belongs to the first prediction mode set, it can only select one texture feature index, and the possible values ​​of lfnst_idx are 0, 1, 2, and 3. If the prediction mode of the current block belongs to the second prediction mode set, the possible values ​​of lfnst_idx are 0, 1, 2, 3, 4, 5, and 6. 1, 2, and 3 correspond to the three transform cores of the transform core group corresponding to the first texture feature index, and 4, 5, and 6 correspond to the three transform cores of the transform core group corresponding to the second texture feature index. The binary symbol correspondence table for lfnst_idx for blocks using the first prediction mode set is shown in Table 8, and the binary symbol correspondence table for lfnst_idx for blocks using the second prediction mode set is shown in Table 7.

[0425] In another embodiment of the present application, the embodiment of the present application provides a code stream, wherein the code stream is generated by bit encoding based on the information to be encoded; wherein the information to be encoded may include at least one of the following: a quantization coefficient of the current block, a transform core index number of the current block, a feature index number of the current block, a prediction mode of the current block, and a value of the first syntax element.

[0426] In an embodiment of the present application, the first syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform core index number. That is to say, the transform core index number of the current block can also be determined by the value of the first syntax element. Among them, if the value of the first syntax element is 0, it is determined that the current block does not use the first transform mode; if the value of the first syntax element is non-zero, it is determined that the current block uses the first transform mode and the corresponding transform core index number. For example, if the value of the first syntax element is 1, it is determined that the current block uses the first transform core; if the value of the first syntax element is 2, it is determined that the current block uses the second transform core; if the value of the first syntax element is 4, it is determined that the current block uses the fourth transform core, etc., and no limitation is made here.

[0427] The embodiment of the present application provides a coding method, specifically a multi-angle selection scheme for intra-frame LFNST / NSPT. First, the prediction mode of the current block is determined; when the prediction mode of the current block meets the first condition, a first candidate list of the current block is determined; wherein the first candidate list indicates at least two candidate transform core groups; then, based on the first candidate list, the transform core of the current block is determined; the residual block of the current block is determined, and the residual block of the current block is transformed according to the transform core to determine the transform coefficient of the current block; the transform coefficient of the current block is encoded, and the obtained coded bits are written into the bitstream. In this way, when the prediction mode of the current block meets the first condition, a first candidate list indicating at least two candidate transform core groups is determined, and then the transform core of the current block is determined based on these at least two candidate transform core groups. That is, for the current block predicted using certain intra-frame prediction modes, when determining the transform core of the current block, it is no longer determined based on only one transform core group, but is determined using at least two candidate transform core groups, thereby improving the accuracy of the transformation, thereby improving the compression efficiency, and thus improving the encoding and decoding performance.

[0428] In another embodiment of the present application, based on the encoding and decoding method described in the aforementioned embodiment, in VVC and ECM, the LFNST / NSPT transform core index can be represented by the syntax element lfnst_idx. In VVC, lfnst_idx can have three values, namely 0, 1, and 2. Among them, lfnst_idx is 0, which means that the current block does not use LFNST. Each transform core group of LFNST in VVC has 2 transform cores, so the value of lfnst_idx is 1 or 2, which means that the current block uses the first transform core or the second transform core of the selected transform core group of LFNST. In the current ECM, lfnst_idx can have four values, namely 0, 1, 2, and 3. Where lfnst_idx is 0, which means that the current block does not use LFNST / NSPT. Each transform core group of LFNST / NSPT in ECM has 3 transform cores, so the value of lfnst_idx is 1, 2, or 3, which means that the current block uses the first transform core, the second transform core, or the third transform core of the selected transform core group of LFNST / NSPT.

[0429] In the embodiments of the present application, if the intra prediction mode of the current block is a special intra prediction mode, then it can select more than one transform kernel group. For example, it can select two transform kernel groups. These special intra prediction modes include, but are not limited to, DIMD, TIMD, MIP, SGPM, ITMP, and IBC.

[0430] In a specific embodiment, a classification is made here, DC, PLANAR, and various angle prediction modes are classified into a first prediction mode set, and DIMD, TIMD, MIP, SGPM, ITMP, IBC, etc. are classified into a second prediction mode set.

[0431] Each transform core group in LFNST / NSPT has three transform cores. Blocks using the first prediction mode set can only select one transform core group, while blocks using the second prediction mode set can select two transform core groups, each with three transform cores. If the prediction mode of the current block belongs to the first prediction mode set, it can only select one texture feature, so the possible values ​​of lfnst_idx are 0, 1, 2, and 3. If the prediction mode of the current block belongs to the second prediction mode set, the possible values ​​of lfnst_idx are 0, 1, 2, 3, 4, 5, and 6. Among them, 1, 2, and 3 correspond to the three transform cores of the transform core group corresponding to the first texture feature index, and 4, 5, and 6 correspond to the three transform cores of the transform core group corresponding to the second texture feature index.

[0432] In the embodiment of the present application, the binary symbol correspondence table of lfnst_idx of the block using the first prediction mode set is shown in the aforementioned Table 8, and the binary symbol correspondence table of lfnst_idx of the block using the second prediction mode set is shown in the aforementioned Table 7. Among them, the third binary symbol, that is, the binary symbol with BinIdx of 2, can also be understood as selecting the first candidate texture feature index or the second candidate texture feature index, or selecting the first candidate transform core group or the second candidate transform core group.

[0433] In another specific embodiment, the second prediction mode set may include only one or more of DIMD, TIMD, MIP, SGPM, ITMP, and IBC. For example, the second prediction mode set includes DIMD, TIMD, and SGPM. The other modes belong to the first prediction mode set. The lfnst_idx encoding method of the first prediction mode set is the same as that of the related art, while the lfnst_idx encoding method of the second prediction mode set is different from that of the related art.

[0434] It should also be noted that in the embodiment of the present application, a dedicated syntax element, such as lfnst_feature_idx, can also be set. The possible values ​​of lfnst_feature_idx are 0 or 1, indicating which candidate texture feature to select, or which candidate transform kernel group to select. If the intra prediction mode of the current block belongs to the second prediction mode set and lfnst_idx>0, lfnst_feature_idx is parsed, and the result is equivalent to the above embodiment.

[0435] It is understandable that for the derivation of candidate texture feature indexes, the embodiment of the present application can use the reconstructed areas on the left and above the current block to derive the candidate texture feature indexes, which is similar to the existing DIMD approach. Because the reconstructed areas on the left and above are not the current block but are adjacent to the current block, for example, when the textures are connected, they can be used to estimate the texture of the current block to a certain extent. Another possibility is to use the prediction block of the current block to derive the candidate texture feature indexes. Another possibility is to use the prediction block of the current block and the reconstructed areas on the left and above the current block at the same time, so that more pixels can be used to derive the candidate texture feature indexes.

[0436] In the embodiment of the present application, the candidate texture feature indexes (or candidate texture features) here can also be directly referred to as candidate transformation kernel groups.

[0437] One derivation method is to calculate the gradients of all or part of the pixels in the selected area. Generally, horizontal and vertical gradients can be calculated, and the Sobel operator can be used to calculate the gradients. For a particular pixel, the texture direction is inferred based on its horizontal and vertical gradients. For example, if the horizontal gradient is non-zero and the vertical gradient is zero, the texture at that point is vertical. Conversely, if the horizontal gradient is zero and the vertical gradient is non-zero, the texture at that point is horizontal. For example, if the horizontal and vertical gradients are equal and non-zero, the texture at that point is 45 degrees. Of course, there are many other cases where both the horizontal and vertical gradients are non-zero, and the texture direction of the pixel can be determined based on their ratio. This allows the gradient to be mapped to the corresponding intra-prediction mode. A statistical table of intra-prediction modes is constructed, and the gradient strength of each calculated pixel is accumulated into the corresponding intra-prediction mode entry in the statistical table. After the gradient statistics are completed, the intra-prediction modes are sorted in descending order of accumulated gradient strength. For complexity considerations, you can only maintain a list of length N, where the length of N can be 2, 3, 4, 5...10, etc.

[0438] In one possible implementation, an example of a Sobel operator is as follows.

[0439] Horizontal gradient operator:

[0440] Operator for vertical gradient:

[0441] Assume that the pixel value of the predicted block at pixel position (x, y) is P x,y , then the horizontal gradient grad x vertical gradient grad y The calculation is as follows: grad x =P x+1,y-1 +2*P x+1,y +Px+1,y+1 -P x-1,y-1 -2*P x-1,y -P x-1,y+1 (9) grad y =P x-1,y+1 +2*P x,y+1 +P x+1,y+1 -P x-1,y-1 -2*P x,y-1 -P x+1,y-1 (10)

[0442] An example of the gradient strength being denoted as amp is amp=abs(grad x )+abs(grad y ).

[0443] According to grad x and grad y The virtual intra prediction mode can be derived by looking up a table. For example, if abs(grad x ) is equal to 0 and abs(grad y ) is not equal to 0, there is horizontal texture, corresponding to intra prediction mode 18 in VVC. If abs(grad y ) is equal to 0 and abs(grad x ) is not equal to 0, there is vertical texture, corresponding to intra prediction mode 50 in VVC. If abs(grad x ) and abs(grad y ) are not equal to 0: If abs(grad x ) is equal to abs(grad y ), and grad x and grad y The same sign corresponds to intra prediction mode 34 in VVC; if abs(grad x ) is equal to 2 times abs(grad y ), and grad x and grad y The symbols are the same, corresponding to intra prediction mode 40 in VVC.

[0444] Other cases can be determined by looking up the table according to the same principle. DIMD also needs to perform similar statistical analysis to derive the intra-frame prediction mode. In terms of implementation, some logic can be reused with DIMD.

[0445] According to the above method, a list of intra-frame prediction modes with a length of N and sorted can be obtained. Then a list of candidate texture features is generated. An optional operation is to prune the N intra-frame prediction modes, or to exclude intra-frame prediction modes that are too close. Taking the intra-frame prediction mode in VVC as an example, the angles of some adjacent intra-frame prediction modes are very close. If two candidate texture features select two very close angles, their difference is not obvious. Therefore, a threshold can be set, which represents the gap between the two intra-frame prediction modes. For example, the intra-frame prediction mode with the largest cumulative gradient intensity value is the first candidate texture feature candFeature0, then when selecting the second candidate texture feature candFeature1, there needs to be a certain gap between candFeature1 and candFeature0, such as candFeature1+THR<candFeature0||candFeature1-THR> candFeature0. If no matching intra-frame prediction mode is found after checking all N-1 intra-frame prediction modes, one possible approach is to return to the second intra-frame prediction mode in the sorted intra-frame prediction mode list of length N and set it to candFeature1. Another possible approach is to add a default texture feature index. The THR may be 3, 4, 5, 6, etc. For situations where the cumulative values ​​of multiple adjacent angles are high, this method will prioritize angles with a certain degree of differentiation.

[0446] The above method does not take into account some of the modes derived from the DIMD, TIMD, SGPM and other modes themselves. For example, DIMD itself will derive one or several intra-frame prediction modes for weighting, and TIMD itself will also derive one or several intra-frame prediction modes for weighting. SGPM not only has two intra-frame prediction modes, but also a "partition" mode that can find the corresponding intra-frame prediction mode, and the residual often appears in the boundary area of ​​the "partition". One possible method is to determine the candidate texture feature index based on the modes derived from the DIMD, TIMD, SGPM and other modes themselves and the modes derived by the above method.

[0447] In a specific embodiment, as follows:

[0448] For DIMD mode, the first intra prediction mode derived by DIMD is used as candFeature0.

[0449] For TIMD mode, the first intra prediction mode derived by TIMD is used as candFeature0.

[0450] For SGPM mode, the intra prediction mode corresponding to the SGPM "split" mode is used as candFeature0.

[0451] Then, candFeature1 is determined according to the above method. That is, the intra-frame prediction modes in the sorted intra-frame prediction mode list of length N are tried, starting from the first intra-frame prediction mode. If an intra-frame prediction mode meets the THR restriction, it is used as candFeature1.

[0452] For IBC, ITMP, etc., the prediction block, or the prediction block plus the template can be used to derive the candidate texture feature index.

[0453] If the pixels used in deriving the list of N sorted intra-frame prediction modes in the embodiment of the present application are the same as the pixels used by DIMD, then the first intra-frame prediction mode derived by DIMD is the same as the first intra-frame prediction mode derived by the embodiment of the present application.

[0454] In another specific embodiment, as follows:

[0455] Since DIMD, TIMD, and SGPM can all derive more than one intra-frame prediction mode, these modes can also give priority to using multiple modes derived by each mode to determine the candidate texture feature index. When the modes derived by each mode itself cannot fill all the candidate texture feature indexes, one possible method is to add a default texture feature index. Another possible method is to use the above-mentioned sorted intra-frame prediction mode list of length N to determine the candidate texture feature index.

[0456] For example, for the DIMD mode, if it uses the prediction values ​​of multiple intra-frame prediction modes for weighting during prediction, then the multiple intra-frame prediction modes used for weighting can be sequentially tried to be determined as candidate texture feature indexes.

[0457] For example, for the TIMD mode, if it uses the prediction values ​​of multiple intra-frame prediction modes for weighting during prediction, then the multiple intra-frame prediction modes used for weighting can be sequentially tried to be determined as candidate texture feature indexes.

[0458] For example, for the SGPM mode, the intra-frame prediction mode corresponding to the "partition" mode derived therefrom and the two intra-frame prediction modes used for prediction may be sequentially attempted to be determined as candidate texture feature indexes.

[0459] It should also be noted that the existing ECM LFNST / NSPT has many transform core groups, so that adjacent intra-frame prediction modes correspond to different transform core groups. A compromise solution in subsequent standards is to appropriately reduce the number of transform core groups so that multiple intra-frame prediction modes correspond to one transform core group. This is similar to the use of 4 transform core groups in VVC, and multiple intra-frame prediction modes of similar angles correspond to the same transform core group. The number of transform core groups in subsequent standards may be 4 or other, such as 8, 12, etc. For the embodiments of the present application, if multiple adjacent intra-frame prediction modes correspond to one transform core group, then when determining the candidate texture features, it is necessary to ensure that each candidate texture feature is not determined to be the same transform core group.

[0460] That is to say, the texture feature index described in this article is used to determine the transformation kernel group. The texture feature index is used for ease of understanding, and the transformation kernel group can also be used directly to replace the texture feature index.

[0461] It should also be noted that the number of pixels whose gradients are calculated can be determined by the size of the current block. Because the Sobel operator uses the pixels in the upper, lower, left, and right rows and columns of the current pixel, it can be set not to calculate the gradients of the pixels in the outermost rows and columns of the current block. If the current block size is small, statistics can be calculated for all pixels for which gradients can be calculated. If the current block size is large, statistics can be calculated for downsampling, such as counting one pixel out of every two, four, or eight pixels in the horizontal and / or vertical directions.

[0462] For example, if the horizontal or vertical size of the current block is less than or equal to 8, gradient statistics are performed on all available pixels in that direction. Otherwise, if the horizontal or vertical size of the current block is less than or equal to 16, gradient statistics are performed on every two pixels in that direction. Otherwise, gradient statistics are performed on every four pixels in that direction.

[0463] In the embodiments of the present application, the specific implementation of the aforementioned embodiments is elaborated in detail through the above embodiments. It can be seen that according to the technical solution of the aforementioned embodiments, candidate texture feature indexes are provided for blocks predicted by certain special intra-frame prediction modes, and a "pruning" operation is also performed to exclude intra-frame prediction modes that are too close; in this way, for blocks predicted by certain special intra-frame prediction modes in intra-frame coded blocks, their predicted residual texture features are not as clear as those of ordinary intra-frame prediction modes. The use of this technical solution to derive multiple candidate texture features to guide the transformation can improve compression efficiency and thus improve encoding and decoding performance.

[0464] In yet another embodiment of the present application, based on the same inventive concept as the aforementioned embodiment, FIG33 is a schematic diagram of the composition structure of an encoder provided in an embodiment of the present application. As shown in FIG33 , the encoder 330 may include a first determination unit 3301, a first transformation unit 3302, and an encoding unit 3303, wherein:

[0465] A first determining unit 3301 is configured to determine a prediction mode of a current block; when the prediction mode of the current block satisfies a first condition, determine a first candidate list for the current block; wherein the first candidate list indicates at least two candidate transform core groups;

[0466] The first determining unit 3301 is further configured to determine a transform kernel of the current block according to the first candidate list;

[0467] A first transform unit 3302 is configured to determine a residual block of a current block, and transform the residual block of the current block according to a transform kernel to determine a transform coefficient of the current block;

[0468] The encoding unit 3303 is configured to perform encoding processing on the transformation coefficients of the current block and write the obtained encoding bits into the bitstream.

[0469] In some embodiments, the prediction mode of the current block satisfies a first condition, including: the prediction mode of the current block is one of the items in the first prediction mode set; the first prediction mode set includes at least: DIMD mode, TIMD mode, SGPM mode, MIP mode, ITMP mode and IBC mode.

[0470] In some embodiments, the first determination unit 3301 is further configured to determine multiple candidate prediction modes; perform encoding cost calculations on the current block based on the multiple candidate prediction modes to determine the cost results corresponding to each of the multiple candidate prediction modes; determine the minimum cost result among the cost results corresponding to the multiple candidate prediction modes; and determine the candidate prediction mode corresponding to the minimum cost result as the prediction mode of the current block.

[0471] In some embodiments, the first determination unit 3301 is further configured to determine one or more intra-frame prediction modes derived from the prediction mode of the current block; determine one or more candidate transform core groups based on the one or more intra-frame prediction modes, and add the one or more candidate transform core groups to the first candidate list.

[0472] In some embodiments, the first determination unit 3301 is further configured to determine a preset texture feature index of the current block when the first candidate list is not filled; determine one or more candidate transform core groups based on the preset texture feature index, and add the one or more candidate transform core groups to the first candidate list.

[0473] In some embodiments, the first determination unit 3301 is further configured to determine candidate pixels for deriving texture feature indexes when the first candidate list is not filled; determine a candidate texture feature list of the current block based on the candidate pixels; and determine one or more candidate transform core groups based on the candidate texture feature list, and add the one or more candidate transform core groups to the first candidate list.

[0474] In some embodiments, the first determination unit 3301 is further configured to determine the horizontal gradient value and the vertical gradient value of the candidate pixel; determine the texture feature index and the gradient intensity value corresponding to the candidate pixel based on the horizontal gradient value and the vertical gradient value of the candidate pixel; construct a texture feature statistics table based on the texture feature index and the gradient intensity value corresponding to the candidate pixel; and determine the candidate texture feature list of the current block based on the texture feature statistics table.

[0475] In some embodiments, the first determining unit 3301 is further configured to determine the number of candidate pixels according to a size parameter of the current block.

[0476] In some embodiments, the first determining unit 3301 is further configured to determine a prediction block of the current block; and use at least part of the pixels in the prediction block as candidate pixels.

[0477] In some embodiments, the first determining unit 3301 is further configured to determine adjacent pixels of a reconstructed area of ​​the current block; and use the adjacent pixels of the reconstructed area as candidate pixels.

[0478] In some embodiments, the first determining unit 3301 is further configured to use adjacent pixels in the reconstructed area and at least part of the pixels in the prediction block as candidate pixels.

[0479] In some embodiments, the first determination unit 3301 is further configured to determine at least one texture feature index and at least one corresponding gradient intensity value when the number of candidate pixels is at least one; determine at least one reference texture feature index with different characteristics based on the at least one texture feature index, and perform cumulative calculation on the gradient intensity values ​​belonging to the same reference texture feature index based on the at least one gradient intensity value to determine the gradient intensity cumulative value corresponding to the at least one reference texture feature index; and construct a texture feature statistical table based on the at least one reference texture feature index and the gradient intensity cumulative value corresponding to the at least one reference texture feature index.

[0480] In some embodiments, the first determination unit 3301 is further configured to sort the texture feature statistics table from high to low according to the gradient intensity accumulation value, determine the reference texture feature indexes corresponding to the top N gradient intensity accumulation values; where N is a positive integer; and add the N reference texture feature indexes to the candidate texture feature list of the current block.

[0481] In some embodiments, the first determining unit 3301 is further configured to prune the N reference texture feature indexes in the candidate texture feature list to determine one or more candidate transform groups for the current block.

[0482] In some embodiments, the first determination unit 3301 is further configured to determine the first candidate transform core group based on the reference texture feature index at the first position in the candidate texture feature list; when the second condition is met between other reference texture feature indexes other than the first position in the candidate texture feature list and the i-th candidate transform core group, determine the i+1-th candidate transform core group based on the other reference texture feature indexes to determine one or more candidate transform core groups for the current block; wherein i is an integer greater than zero and less than N.

[0483] In some embodiments, the first determination unit 3301 is further configured to determine a first candidate transformation core group based on the reference texture feature index at the first position in the candidate texture feature list; when the second condition is not satisfied between all reference texture feature indexes except the first position in the candidate texture feature list and the first candidate transformation core group, determine a second candidate transformation core group based on the reference texture feature index at the second position in the candidate texture feature list; determine the first candidate list based on the first candidate transformation core group and the second candidate transformation core group.

[0484] In some embodiments, the first determination unit 3301 is further configured to determine a first candidate transform core group based on a first intra-frame prediction mode derived from a prediction mode of the current block; determine a second candidate transform core group based on the first reference texture feature index when a second condition is satisfied between the first reference texture feature index in the candidate texture feature list and the first candidate transform core group; and determine a first candidate list based on the first candidate transform core group and the second candidate transform core group.

[0485] In some embodiments, the first candidate list indicates transform cores included in at least two transform core groups.

[0486] In some embodiments, the first determination unit 3301 is further configured to determine at least two candidate transformation cores indicated by the first candidate list; perform encoding cost calculation on the current block based on the at least two candidate transformation cores, and determine the cost results corresponding to each of the at least two candidate transformation cores; and determine the minimum cost result among the cost results corresponding to each of the at least two candidate transformation cores, and determine the candidate transformation core corresponding to the minimum cost result as the transformation core of the current block.

[0487] In some embodiments, the first determination unit 3301 is further configured to determine the transformation core group of the current block based on the first candidate list; wherein the transformation core group is one of the at least two candidate transformation core groups indicated by the first candidate list; and determine the transformation core of the current block based on the transformation core group.

[0488] In some embodiments, the first determination unit 3301 is further configured to determine at least two candidate transform cores included in the transform core group; perform encoding cost calculation on the current block based on the at least two candidate transform cores, and determine the cost results corresponding to each of the at least two candidate transform cores; and determine the minimum cost result among the cost results corresponding to each of the at least two candidate transform cores, and determine the candidate transform core corresponding to the minimum cost result as the transform core of the current block.

[0489] In some embodiments, the first determination unit 3301 is further configured to transform and quantize the residual block of the current block based on the first candidate transform kernel, determine the first candidate quantization coefficient of the current block, and perform entropy coding on the first candidate quantization coefficient to determine the first generation value of the first candidate transform kernel; dequantize and detransform the first candidate quantization coefficient to determine the first candidate residual block of the current block, and determine the first candidate prediction block of the current block based on the first candidate residual block; perform cost calculation based on the first candidate prediction block and the original image of the current block to determine the second generation value of the first candidate transform kernel; and determine the cost result corresponding to the first candidate transform kernel based on the first generation value and the second generation value of the first candidate transform kernel; wherein the first candidate transform kernel is any one of the at least two candidate transform kernels.

[0490] In some embodiments, the first determination unit 3301 is further configured to determine the transform core index number of the current block; wherein the transform core index number is used to indicate the number of the transform core of the current block in the first candidate list or the transform core group of the previous block; the encoding unit 3303 is further configured to encode the transform core index number of the current block and write the obtained encoded bits into the bitstream.

[0491] In some embodiments, the first determination unit 3301 is further configured to determine the value of the first syntax element; wherein the first syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform core index number used, and the transform core index number is used to indicate the number of the transform core of the current block in the first candidate list or the transform core group of the current block; the encoding unit 3303 is further configured to encode the value of the first syntax element and write the obtained encoded bits into the bitstream.

[0492] In some embodiments, the first determination unit 3301 is further configured to perform encoding cost calculation on the current block based on at least two candidate transform core groups in the first candidate list, determine the cost results corresponding to each of the at least two candidate transform core groups; and determine the minimum cost result among the cost results corresponding to each of the at least two candidate transform core groups, and determine the candidate transform core group corresponding to the minimum cost result as the texture feature index of the current block.

[0493] In some embodiments, the first determination unit 3301 is further configured to determine the feature index number of the current block; wherein the feature index number is used to indicate the number of the transform core group of the current block in the first candidate list; the encoding unit 3303 is further configured to encode the feature index number of the current block and write the obtained encoded bits into the bitstream.

[0494] In some embodiments, the first determination unit 3301 is further configured to determine the texture feature index of the current block when the prediction mode of the current block does not meet the first condition; and determine the transform kernel of the current block based on the texture feature index; the first transform unit 3302 is further configured to determine the transform coefficient of the current block, and perform an inverse transform on the transform coefficient of the current block based on the transform kernel to determine the residual block of the current block.

[0495] In some embodiments, the prediction mode of the current block does not satisfy the first condition, including: the prediction mode of the current block is a prediction mode outside the first prediction mode set.

[0496] In some embodiments, the prediction mode of the current block does not meet the first condition, including: the prediction mode of the current block is one of the second prediction mode set; the second prediction mode set includes at least: DC mode, PLANAR mode, angle prediction mode and cross-component prediction mode.

[0497] In some embodiments, the first determining unit 3301 is further configured to perform intra-frame prediction on the current block to determine a prediction block of the current block; and determine a residual block of the current block based on the initial block of the current block and the prediction block of the current block.

[0498] In some embodiments, the first determination unit 3301 is further configured to quantize the transform coefficients of the current block to determine the quantization coefficients of the current block; the encoding unit 3303 is further configured to encode the quantization coefficients of the current block and write the obtained encoding bits into the bitstream.

[0499] In some embodiments, the first transformation unit 3302 is further configured to perform an inseparable basic transform on the residual block of the current block according to the transform kernel to determine the transform coefficient of the current block; or, perform a discrete cosine transform on the residual block of the current block to determine the transform block of the current block; and perform a low-frequency inseparable transform on the transform block of the current block according to the transform kernel to determine the transform coefficient of the current block.

[0500] In some embodiments, the encoding unit 3303 is further configured to perform encoding processing on the prediction mode of the current block and write the obtained encoding bits into the bitstream.

[0501] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.

[0502] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0503] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 330. The computer-readable storage medium stores a computer program, and when the computer program is executed by the first processor, it implements the method described in any one of the aforementioned embodiments.

[0504] Based on the composition of the above-mentioned encoder 330 and the computer-readable storage medium, Figure 34 is a schematic diagram of the specific hardware structure of an encoder provided by an embodiment of the present application. As shown in Figure 34, the encoder 330 may include: a first communication interface 3401, a first memory 3402 and a first processor 3403; each component is coupled together through a first bus system 3404. It can be understood that the first bus system 3404 is used to achieve connection and communication between these components. In addition to the data bus, the first bus system 3404 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the first bus system 3404 in Figure 34. Among them,

[0505] The first communication interface 3401 is used to receive and send signals when sending and receiving information with other external network elements;

[0506] A first memory 3402 is used to store computer programs that can be run on the first processor 3403;

[0507] The first processor 3403 is configured to, when running the computer program, execute:

[0508] Determine a prediction mode for a current block; when the prediction mode of the current block satisfies a first condition, determine a first candidate list for the current block; wherein the first candidate list indicates at least two candidate transform core groups; determine a transform core for the current block based on the first candidate list; determine a residual block for the current block, and transform the residual block for the current block based on the transform core to determine a transform coefficient for the current block; encode the transform coefficient for the current block, and write the obtained coded bits into a bitstream.

[0509] It is understood that the first memory 3402 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The first memory 3402 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0510] The first processor 3403 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the first processor 3403. The above-mentioned first processor 3403 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the first memory 3402 , and the first processor 3403 reads the information in the first memory 3402 and completes the steps of the above method in combination with its hardware.

[0511] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0512] Optionally, as another embodiment, the first processor 3403 is further configured to execute any one of the methods described in the foregoing embodiments when running the computer program.

[0513] This embodiment provides an encoder that first determines a prediction mode for a current block. When the prediction mode for the current block satisfies a first condition, the encoder then determines a first candidate list indicating at least two candidate transform core groups. The encoder then determines a transform core for the current block based on these at least two candidate transform core groups. In other words, for a current block predicted using certain intra-frame prediction modes, the transform core for the current block is no longer determined based on a single transform core group, but rather based on at least two candidate transform core groups. This improves transform accuracy, thereby increasing compression efficiency and, consequently, encoding and decoding performance.

[0514] In yet another embodiment of the present application, based on the same inventive concept as the aforementioned embodiment, FIG35 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application. As shown in FIG35 , the decoder 350 may include a second determination unit 3501 and a second transformation unit 3502, wherein:

[0515] The second determining unit 3501 is configured to determine a prediction mode of the current block; when the prediction mode of the current block satisfies a first condition, determine a first candidate list for the current block; wherein the first candidate list indicates at least two candidate transform core groups;

[0516] The second determining unit 3501 is further configured to determine a transform kernel for the current block according to the first candidate list;

[0517] The second transform unit 3502 is configured to determine transform coefficients of the current block, and perform inverse transform on the transform coefficients of the current block according to the transform kernel to determine a residual block of the current block.

[0518] In some embodiments, the prediction mode of the current block satisfies a first condition, including: the prediction mode of the current block is one of the items in the first prediction mode set; the first prediction mode set includes at least: DIMD mode, TIMD mode, SGPM mode, MIP mode, ITMP mode and IBC mode.

[0519] In some embodiments, referring to FIG. 35 , the decoder 350 further includes a decoding unit 3503 configured to decode the code stream and determine a prediction mode of the current block.

[0520] In some embodiments, the second determination unit 3501 is further configured to determine one or more intra-frame prediction modes derived from the prediction mode of the current block; determine one or more candidate transform core groups based on the one or more intra-frame prediction modes, and add the one or more candidate transform core groups to the first candidate list.

[0521] In some embodiments, the second determination unit 3501 is further configured to determine a preset texture feature index of the current block when the first candidate list is not filled; determine one or more candidate transform core groups based on the preset texture feature index, and add the one or more candidate transform core groups to the first candidate list.

[0522] In some embodiments, the second determination unit 3501 is further configured to determine candidate pixels for deriving texture feature indexes when the first candidate list is not filled; determine a candidate texture feature list of the current block based on the candidate pixels; and determine one or more candidate transform core groups based on the candidate texture feature list, and add the one or more candidate transform core groups to the first candidate list.

[0523] In some embodiments, the second determination unit 3501 is further configured to determine the horizontal gradient value and the vertical gradient value of the candidate pixel; determine the texture feature index and the gradient intensity value corresponding to the candidate pixel based on the horizontal gradient value and the vertical gradient value of the candidate pixel; construct a texture feature statistics table based on the texture feature index and the gradient intensity value corresponding to the candidate pixel; and determine the candidate texture feature list of the current block based on the texture feature statistics table.

[0524] In some embodiments, the second determining unit 3501 is further configured to determine the number of candidate pixels according to a size parameter of the current block.

[0525] In some embodiments, the second determining unit 3501 is further configured to determine a prediction block of the current block; and use at least part of the pixels in the prediction block as candidate pixels.

[0526] In some embodiments, the second determining unit 3501 is further configured to determine adjacent pixels of a reconstructed area of ​​the current block; and use the adjacent pixels of the reconstructed area as candidate pixels.

[0527] In some embodiments, the second determining unit 3501 is further configured to use adjacent pixels in the reconstructed area and at least part of the pixels in the prediction block as candidate pixels.

[0528] In some embodiments, the second determination unit 3501 is further configured to determine at least one texture feature index and at least one corresponding gradient intensity value when the number of candidate pixels is at least one; determine at least one reference texture feature index with different characteristics based on the at least one texture feature index, and perform cumulative calculation on the gradient intensity values ​​belonging to the same reference texture feature index based on the at least one gradient intensity value to determine the gradient intensity cumulative value corresponding to the at least one reference texture feature index; and construct a texture feature statistical table based on the at least one reference texture feature index and the gradient intensity cumulative value corresponding to the at least one reference texture feature index.

[0529] In some embodiments, the second determination unit 3501 is further configured to sort the texture feature statistics table from high to low according to the gradient intensity accumulation value, determine the reference texture feature indexes corresponding to the top N gradient intensity accumulation values; where N is a positive integer; and add the N reference texture feature indexes to the candidate texture feature list of the current block.

[0530] In some embodiments, the second determining unit 3501 is further configured to prune the N reference texture feature indices in the candidate texture feature list to determine one or more candidate transform groups for the current block.

[0531] In some embodiments, the second determination unit 3501 is further configured to determine the first candidate transform core group based on the reference texture feature index at the first position in the candidate texture feature list; when the second condition is met between other reference texture feature indexes other than the first position in the candidate texture feature list and the i-th candidate transform core group, determine the i+1-th candidate transform core group based on the other reference texture feature indexes to determine one or more candidate transform core groups for the current block; wherein i is an integer greater than zero and less than N.

[0532] In some embodiments, the second determination unit 3501 is further configured to determine a first candidate transformation core group based on the reference texture feature index at the first position in the candidate texture feature list; when the second condition is not satisfied between all reference texture feature indexes except the first position in the candidate texture feature list and the first candidate transformation core group, determine a second candidate transformation core group based on the reference texture feature index at the second position in the candidate texture feature list; determine the first candidate list based on the first candidate transformation core group and the second candidate transformation core group.

[0533] In some embodiments, the second determination unit 3501 is further configured to determine a first candidate transform core group based on a first intra-frame prediction mode derived from a prediction mode of the current block; determine a second candidate transform core group based on the first reference texture feature index when a second condition is satisfied between the first reference texture feature index in the candidate texture feature list and the first candidate transform core group; and determine a first candidate list based on the first candidate transform core group and the second candidate transform core group.

[0534] In some embodiments, the first candidate list indicates transform cores included in at least two transform core groups.

[0535] In some embodiments, the second determining unit 3501 is further configured to determine a transform core index number of the current block; and determine the transform core of the current block according to the first candidate list and the transform core index number.

[0536] In some embodiments, the decoding unit 3503 is further configured to decode the code stream and determine the feature index number of the current block; the second determination unit 3501 is further configured to determine the transformation core group of the current block based on the first candidate list and the feature index number; wherein the transformation core group is one of the at least two candidate transformation core groups indicated by the first candidate list; and the transformation core of the current block is determined based on the transformation core group.

[0537] In some embodiments, the second determining unit 3501 is further configured to determine a transform core index number of the current block; and determine the transform core of the current block according to the transform core group and the transform core index number.

[0538] In some embodiments, the decoding unit 3503 is further configured to decode the code stream and determine the transform core index number of the current block.

[0539] In some embodiments, the decoding unit 3503 is further configured to decode the code stream and determine the value of the first syntax element; the second determination unit 3501 is further configured to determine the transform core index number of the current block according to the value of the first syntax element when the first syntax element indicates that the current block uses the first transform mode.

[0540] In some embodiments, the second determination unit 3501 is further configured to determine the texture feature index of the current block when the prediction mode of the current block does not meet the first condition; and determine the transformation kernel of the current block based on the texture feature index; the second transformation unit 3502 is further configured to determine the transformation coefficient of the current block, and perform inverse transformation on the transformation coefficient of the current block based on the transformation kernel to determine the residual block of the current block.

[0541] In some embodiments, the second determination unit 3501 is further configured to determine the transform core group of the current block based on the texture feature index; decode the code stream to determine the transform core index number of the current block; and determine the transform core of the current block based on the transform core group and the transform core index number.

[0542] In some embodiments, the prediction mode of the current block does not satisfy the first condition, including: the prediction mode of the current block is a prediction mode outside the first prediction mode set.

[0543] In some embodiments, the prediction mode of the current block does not meet the first condition, including: the prediction mode of the current block is one of the second prediction mode set; the second prediction mode set includes at least: DC mode, PLANAR mode, angle prediction mode and cross-component prediction mode.

[0544] In some embodiments, the decoding unit 3503 is further configured to decode the code stream to determine the quantization coefficient of the current block; the second determination unit 3501 is further configured to dequantize the quantization coefficient of the current block to determine the transformation coefficient of the current block.

[0545] In some embodiments, the second transformation unit 3502 is further configured to perform an inverse transform of an inseparable basic transform on the transformation coefficients of the current block according to the transformation kernel to determine the residual block of the current block; or, perform an inverse transform of a low-frequency inseparable transform on the transformation coefficients of the current block according to the transformation kernel to determine the transformation block of the current block; and perform an inverse transform of a discrete cosine transform on the transformation block of the current block to determine the residual block of the current block.

[0546] In some embodiments, the second determining unit 3501 is further configured to perform intra-frame prediction on the current block to determine a prediction block of the current block; and determine a reconstructed block of the current block based on the prediction block of the current block and the residual block of the current block.

[0547] It is understood that in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular system. Furthermore, the various components in this embodiment can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The aforementioned integrated units can be implemented in the form of hardware or software functional modules.

[0548] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, this embodiment provides a computer-readable storage medium, which is applied to the decoder 350 and stores a computer program. When the computer program is executed by the second processor, it implements any of the methods in the aforementioned embodiments.

[0549] Based on the composition of the above-mentioned decoder 350 and the computer-readable storage medium, Figure 36 is a schematic diagram of the specific hardware structure of a decoder provided in an embodiment of the present application. As shown in Figure 36, the decoder 350 may include: a second communication interface 3601, a second memory 3602 and a second processor 3603; each component is coupled together through a second bus system 3604. It can be understood that the second bus system 3604 is used to realize the connection and communication between these components. In addition to the data bus, the second bus system 3604 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the second bus system 3604 in Figure 36. Among them,

[0550] The second communication interface 3601 is used to receive and send signals when sending and receiving information with other external network elements;

[0551] The second memory 3602 is used to store computer programs that can be run on the second processor 3603;

[0552] The second processor 3603 is configured to, when running the computer program, execute:

[0553] Determine a prediction mode of the current block; when the prediction mode of the current block satisfies a first condition, determine a first candidate list for the current block; wherein the first candidate list indicates at least two candidate transform core groups; determine a transform kernel for the current block based on the first candidate list; determine a transform coefficient for the current block, and inversely transform the transform coefficient for the current block based on the transform kernel to determine a residual block for the current block.

[0554] Optionally, as another embodiment, the second processor 3603 is further configured to execute any one of the methods described in the foregoing embodiments when running the computer program.

[0555] It can be understood that the hardware functions of the second memory 3602 are similar to those of the first memory 3402, and the hardware functions of the second processor 3603 are similar to those of the first processor 3403; they will not be described in detail here.

[0556] This embodiment provides a decoder that first determines a prediction mode for a current block. When the prediction mode for the current block satisfies a first condition, the decoder then determines a first candidate list indicating at least two candidate transform core groups. The decoder then determines a transform core for the current block based on these at least two candidate transform core groups. In other words, for a current block predicted using certain intra-frame prediction modes, the transform core for the current block is no longer determined based on a single transform core group, but rather on at least two candidate transform core groups. This improves transform accuracy, thereby increasing compression efficiency and, consequently, encoding and decoding performance.

[0557] In yet another embodiment of the present application, FIG37 is a schematic diagram of the structure of a coding and decoding system provided by an embodiment of the present application. As shown in FIG37 , the coding and decoding system 370 may include an encoder 3701 and a decoder 3702 .

[0558] In the embodiment of the present application, the encoder 3701 may be the encoder described in any one of the aforementioned embodiments, and the decoder 3702 may be the decoder described in any one of the aforementioned embodiments.

[0559] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0560] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0561] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0562] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0563] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0564] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims. Industrial Applicability

[0565] In an embodiment of the present application, at the encoding end, a prediction mode of the current block is determined; when the prediction mode of the current block meets a first condition, a first candidate list of the current block is determined; wherein the first candidate list indicates at least two candidate transform core groups; based on the first candidate list, a transform core of the current block is determined; a residual block of the current block is determined, and the residual block of the current block is transformed according to the transform core to determine the transform coefficients of the current block; the transform coefficients of the current block are encoded, and the obtained coded bits are written into the bitstream. At the decoding end, a prediction mode of the current block is determined; when the prediction mode of the current block meets the first condition, a first candidate list of the current block is determined; wherein the first candidate list indicates at least two candidate transform core groups; based on the first candidate list, a transform core of the current block is determined; the transform coefficients of the current block are determined, and the transform coefficients of the current block are inversely transformed according to the transform core to determine the residual block of the current block. In this way, both the encoding and decoding ends first determine the prediction mode for the current block. When the prediction mode for the current block meets the first condition, a first candidate list indicating at least two candidate transform core groups is determined, and then the transform core for the current block is determined based on these at least two candidate transform core groups. In other words, for a current block predicted using certain intra-frame prediction modes, the transform core for the current block is no longer determined based on just one transform core group, but rather based on at least two candidate transform core groups. This improves the accuracy of the transform, thereby increasing compression efficiency and, in turn, enhancing encoding and decoding performance.

Claims

1. A decoding method, applied to a decoder, the method comprising: Determine the prediction mode of the current block; When the prediction mode of the current block satisfies a first condition, determine a first candidate list for the current block; wherein the first candidate list indicates at least two candidate transform kernel groups; Determine the transform kernel of the current block according to the first candidate list; Determine the transform coefficients of the current block, and perform an inverse transform on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block.

2. The method according to claim 1, wherein That the prediction mode of the current block satisfies the first condition includes: the prediction mode of the current block is one of a first set of prediction modes; The first set of prediction modes includes at least: DIMD mode, TIMD mode, SGPM mode, MIP mode, ITMP mode, and IBC mode.

3. The method according to claim 1, wherein, The determining the prediction mode of the current block includes: Decode the bitstream to determine the prediction mode of the current block.

4. The method according to claim 1, wherein, The determining the first candidate list for the current block includes: Determine one or more intra prediction modes derived from the prediction mode of the current block; Determine one or more candidate transform kernel groups according to the one or more intra prediction modes, and add the one or more candidate transform kernel groups to the first candidate list.

5. The method according to claim 4, wherein, The method further includes: When the first candidate list is not full, determine a preset texture feature index of the current block; Determine one or more candidate transform kernel groups according to the preset texture feature index, and add the one or more candidate transform kernel groups to the first candidate list.

6. The method according to claim 4, wherein The method further includes: When the first candidate list is not full, determine candidate pixels for deriving the texture feature index; Determine a candidate texture feature list of the current block according to the candidate pixels; Determine one or more candidate transform kernel groups according to the candidate texture feature list, and add the one or more candidate transform kernel groups to the first candidate list.

7. The method according to claim 6, wherein The determining the candidate texture feature list of the current block according to the candidate pixels includes: Determine the horizontal gradient value and the vertical gradient value of the candidate pixels; Determine the texture feature index and the gradient intensity value corresponding to the candidate pixels according to the horizontal gradient value and the vertical gradient value of the candidate pixels; Construct a texture feature statistical table according to the texture feature index and the gradient intensity value corresponding to the candidate pixels; Determine the candidate texture feature list of the current block according to the texture feature statistical table.

8. The method according to claim 6, wherein, The method further includes: Determine the number of candidate pixels according to the size parameter of the current block.

9. The method according to claim 6, wherein The method further includes: Determine the predicted block of the current block; Use at least some of the pixels in the predicted block as the candidate pixels.

10. The method according to claim 9, wherein The method further includes: Determine the adjacent pixels of the reconstructed area of the current block; Use the adjacent pixels of the reconstructed area as the candidate pixels.

11. The method according to claim 10, wherein, The method further includes: Use the adjacent pixels of the reconstructed area and at least some of the pixels in the predicted block as the candidate pixels.

12. The method according to claim 7, wherein The constructing a texture feature statistical table according to the texture feature index and the gradient intensity value corresponding to the candidate pixels includes: When the number of the candidate pixels is at least one, determine at least one texture feature index and corresponding at least one gradient intensity value; Determine at least one reference texture feature index with distinct characteristics according to the at least one texture feature index, and according to the at least one gradient intensity value, perform an accumulation calculation on the gradient intensity values belonging to the same reference texture feature index to determine the gradient intensity accumulation value corresponding to the at least one reference texture feature index; Construct the texture feature statistical table according to the at least one reference texture feature index and the gradient intensity accumulation value corresponding to the at least one reference texture feature index.

13. The method according to claim 12, wherein, The determining the candidate texture feature list of the current block according to the texture feature statistical table includes: Sort the texture feature statistical table in descending order according to the gradient intensity accumulation value, and determine the reference texture feature indices corresponding to the top N gradient intensity accumulation values; where N is a positive integer; Add the N reference texture feature indices to the candidate texture feature list of the current block.

14. The method according to claim 13, wherein, The determining one or more candidate transform kernel groups according to the candidate texture feature list includes: Prune the N reference texture feature indices in the candidate texture feature list to determine one or more candidate transform groups of the current block.

15. The method according to claim 14, wherein, The pruning the N reference texture feature indices in the candidate texture feature list to determine one or more candidate transform kernel groups of the current block includes: Determine the first candidate transform kernel group according to the reference texture feature index at the first position in the candidate texture feature list; When the other reference texture feature indices except the reference texture feature index at the first position in the candidate texture feature list satisfy a second condition with the i-th candidate transform kernel group, determine the (i + 1)-th candidate transform kernel group according to the other reference texture feature indices, so as to determine one or more candidate transform kernel groups of the current block; where i is an integer greater than zero and less than N.

16. The method according to claim 13, wherein, The determining the first candidate list of the current block includes: Determine the first candidate transform kernel group according to the reference texture feature index at the first position in the candidate texture feature list; When all the reference texture feature indices except the reference texture feature index at the first position in the candidate texture feature list do not satisfy the second condition with the first candidate transform kernel group, determine the second candidate transform kernel group according to the reference texture feature index at the second position in the candidate texture feature list; Determine the first candidate list according to the first candidate transform kernel group and the second candidate transform kernel group.

17. The method according to claim 13, wherein, The determining the first candidate list of the current block includes: Determine the first candidate transform kernel group according to the first intra prediction mode derived from the prediction mode of the current block; When the first reference texture feature index in the candidate texture feature list satisfies the second condition with the first candidate transform kernel group, determine the second candidate transform kernel group according to the first reference texture feature index; Determine the first candidate list according to the first candidate transform kernel group and the second candidate transform kernel group.

18. The method according to any one of claims 1 to 17, wherein The first candidate list indicates the transform kernels included in the at least two groups of transform kernels.

19. The method according to claim 18, wherein, Determining the transform kernel of the current block according to the first candidate list includes: Determining the transform kernel index number of the current block; Determining the transform kernel of the current block according to the first candidate list and the transform kernel index number.

20. The method according to any one of claims 1 to 17, wherein, Determining the transform kernel of the current block according to the first candidate list includes: Decoding the bitstream to determine the feature index number of the current block; Determining the group of transform kernels of the current block according to the first candidate list and the feature index number; wherein, the group of transform kernels is one of at least two candidate groups of transform kernels indicated by the first candidate list; Determining the transform kernel of the current block according to the group of transform kernels.

21. The method according to claim 20, wherein Determining the transform kernel of the current block according to the group of transform kernels includes: Determining the transform kernel index number of the current block; Determining the transform kernel of the current block according to the group of transform kernels and the transform kernel index number.

22. The method according to claim 19 or 21, wherein, Determining the transform kernel index number of the current block includes: Decoding the bitstream to determine the transform kernel index number of the current block.

23. The method according to claim 19 or 21, wherein, Determining the transform kernel index number of the current block includes: Decoding the bitstream to determine the value of a first syntax element; When the first syntax element indicates that the current block uses a first transform mode, determining the transform kernel index number of the current block according to the value of the first syntax element.

24. The method according to claim 1, wherein The method further includes: When the prediction mode of the current block does not meet a first condition, determining the texture feature index of the current block; Determining the transform kernel of the current block according to the texture feature index; Determining the transform coefficients of the current block, and performing an inverse transform on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block.

25. The method according to claim 24, wherein, Determining the transform kernel of the current block according to the texture feature index includes: Determining the group of transform kernels of the current block according to the texture feature index; Decoding the bitstream to determine the transform kernel index number of the current block; Determining the transform kernel of the current block according to the group of transform kernels and the transform kernel index number.

26. The method according to claim 24, wherein, The prediction mode of the current block not meeting the first condition includes: the prediction mode of the current block is a prediction mode outside a first set of prediction modes.

27. The method according to claim 24, wherein, The prediction mode of the current block not meeting the first condition includes: the prediction mode of the current block is one of a second set of prediction modes; The second set of prediction modes at least includes: DC mode, PLANAR mode, angular prediction mode, and cross-component prediction mode.

28. The method according to any one of claims 1 to 27, wherein Determining the transform coefficients of the current block includes: Decoding the bitstream to determine the quantization coefficients of the current block; Performing an inverse quantization on the quantization coefficients of the current block to determine the transform coefficients of the current block.

29. The method according to any one of claims 1 to 27, wherein, Performing an inverse transform on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block includes: Performing an inverse transform of an inseparable basis transform on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block; or, Inverse-transform the transform coefficients of the current block by the low-frequency non-separable transform according to the transformation, and determine the transform block of the current block; and inverse-transform the transform block of the current block by the discrete cosine transform to determine the residual block of the current block.

30. The method according to any one of claims 1 to 27, wherein, The method further includes: Perform intra prediction on the current block to determine the predicted block of the current block; Determine the reconstructed block of the current block according to the predicted block of the current block and the residual block of the current block.

31. An encoding method applied to an encoder, the method includes: Determine the prediction mode of the current block; When the prediction mode of the current block satisfies a first condition, determine the first candidate list of the current block; wherein, the first candidate list indicates at least two candidate transform kernel groups; Determine the transform kernel of the current block according to the first candidate list; Determine the residual block of the current block, and transform the residual block of the current block according to the transform kernel to determine the transform coefficients of the current block; Perform encoding processing on the transform coefficients of the current block, and write the obtained encoded bits into the code stream.

32. The method according to claim 31, wherein, The prediction mode of the current block satisfies the first condition, including: the prediction mode of the current block is one of the first prediction mode set; The first prediction mode set at least includes: DIMD mode, TIMD mode, SGPM mode, MIP mode, ITMP mode, and IBC mode.

33. The method according to claim 31, wherein, The determining the prediction mode of the current block includes: Determine multiple candidate prediction modes; Calculate the encoding cost of the current block according to the multiple candidate prediction modes, and determine the cost results corresponding to the multiple candidate prediction modes respectively; Determine the minimum cost result among the cost results corresponding to the multiple candidate prediction modes; Determine the candidate prediction mode corresponding to the minimum cost result as the prediction mode of the current block.

34. The method according to claim 31, wherein The determining the first candidate list of the current block includes: Determine one or more intra prediction modes derived from the prediction mode of the current block; Determine one or more candidate transform kernel groups according to the one or more intra prediction modes, and add the one or more candidate transform kernel groups to the first candidate list.

35. The method according to claim 34, wherein, The method further includes: When the first candidate list is not full, determine the preset texture feature index of the current block; Determine one or more candidate transform kernel groups according to the preset texture feature index, and add the one or more candidate transform kernel groups to the first candidate list.

36. The method according to claim 34, wherein, The method further includes: When the first candidate list is not full, determine the candidate pixels for deriving the texture feature index; Determine the candidate texture feature list of the current block according to the candidate pixels; Determine one or more candidate transform kernel groups according to the candidate texture feature list, and add the one or more candidate transform kernel groups to the first candidate list.

37. The method according to claim 36, wherein, The determining the candidate texture feature list of the current block according to the candidate pixels includes: Determine the horizontal gradient value and the vertical gradient value of the candidate pixels; Determine the texture feature index and the gradient intensity value corresponding to the candidate pixels according to the horizontal gradient value and the vertical gradient value of the candidate pixels; Construct a texture feature statistical table according to the texture feature index and gradient intensity value corresponding to the candidate pixels; Determine the candidate texture feature list of the current block according to the texture feature statistical table.

38. The method according to claim 36, wherein, The method further includes: Determine the number of candidate pixels according to the size parameter of the current block.

39. The method according to claim 36, wherein, The method further includes: Determine the prediction block of the current block; Use at least some of the pixels in the prediction block as the candidate pixels.

40. The method according to claim 39, wherein, The method further includes: Determine the adjacent pixels of the reconstructed area of the current block; Use the adjacent pixels of the reconstructed area as the candidate pixels.

41. The method according to claim 40, wherein, The method further includes: Use the adjacent pixels of the reconstructed area and at least some of the pixels in the prediction block as the candidate pixels.

42. The method according to claim 37, wherein The constructing a texture feature statistical table according to the texture feature index and gradient intensity value corresponding to the candidate pixels includes: When the number of candidate pixels is at least one, determine at least one texture feature index and the corresponding at least one gradient intensity value; Determine at least one reference texture feature index with distinct characteristics according to the at least one texture feature index, and according to the at least one gradient intensity value, perform an accumulative calculation on the gradient intensity values belonging to the same reference texture feature index to determine the gradient intensity accumulative value corresponding to the at least one reference texture feature index; Construct the texture feature statistical table according to the at least one reference texture feature index and the gradient intensity accumulative value corresponding to the at least one reference texture feature index.

43. The method according to claim 42, wherein, The determining the candidate texture feature list of the current block according to the texture feature statistical table includes: Sort the texture feature statistical table in descending order according to the gradient intensity accumulative value, and determine the reference texture feature indexes corresponding to the top N gradient intensity accumulative values; where N is a positive integer; Add the N reference texture feature indexes to the candidate texture feature list of the current block.

44. The method according to claim 43, wherein The determining one or more candidate transform kernel groups according to the candidate texture feature list includes: Prune the N reference texture feature indexes in the candidate texture feature list to determine one or more candidate transform groups of the current block.

45. The method according to claim 44, wherein, The pruning the N reference texture feature indexes in the candidate texture feature list to determine one or more candidate transform kernel groups of the current block includes: Determine the first candidate transform kernel group according to the reference texture feature index at the first position in the candidate texture feature list; When the other reference texture feature indexes other than the reference texture feature index at the first position in the candidate texture feature list satisfy a second condition with the i-th candidate transform kernel group, determine the (i + 1)-th candidate transform kernel group according to the other reference texture feature indexes to determine one or more candidate transform kernel groups of the current block; where i is an integer greater than zero and less than N.

46. The method according to claim 43, wherein, The determining the first candidate list of the current block includes: Determine the first candidate transform kernel group according to the reference texture feature index at the first position in the candidate texture feature list; When the second condition is not satisfied between all the reference texture feature indices except the first position in the candidate texture feature list and the first candidate transform kernel group, determine a second candidate transform kernel group according to the reference texture feature index at the second position in the candidate texture feature list; Determine the first candidate list according to the first candidate transform kernel group and the second candidate transform kernel group.

47. The method according to claim 43, wherein, The determining the first candidate list of the current block includes: Determine a first candidate transform kernel group according to the first intra prediction mode derived from the prediction mode of the current block; When the second condition is satisfied between the first reference texture feature index in the candidate texture feature list and the first candidate transform kernel group, determine a second candidate transform kernel group according to the first reference texture feature index; Determine the first candidate list according to the first candidate transform kernel group and the second candidate transform kernel group.

48. The method according to any one of claims 31 to 47, wherein, The first candidate list indicates the transform kernels included in the at least two transform kernel groups.

49. The method according to claim 48, wherein, The determining the transform kernel of the current block according to the first candidate list includes: Determine at least two candidate transform kernels indicated by the first candidate list; Perform coding cost calculation on the current block according to the at least two candidate transform kernels, and determine the cost results corresponding to the at least two candidate transform kernels respectively; Determine the minimum cost result among the cost results corresponding to the at least two candidate transform kernels respectively, and determine the candidate transform kernel corresponding to the minimum cost result as the transform kernel of the current block.

50. The method according to any one of claims 31 to 47, wherein, The determining the transform kernel of the current block according to the first candidate list includes: Determine the transform kernel group of the current block according to the first candidate list; wherein, the transform kernel group is one of the at least two candidate transform kernel groups indicated by the first candidate list; Determine the transform kernel of the current block according to the transform kernel group.

51. The method according to claim 50, wherein, The determining the transform kernel of the current block according to the transform kernel group includes: Determine at least two candidate transform kernels included in the transform kernel group; Perform coding cost calculation on the current block according to the at least two candidate transform kernels, and determine the cost results corresponding to the at least two candidate transform kernels respectively; Determine the minimum cost result among the cost results corresponding to the at least two candidate transform kernels respectively, and determine the candidate transform kernel corresponding to the minimum cost result as the transform kernel of the current block.

52. The method according to claim 49 or 51, wherein, The performing coding cost calculation on the current block according to the at least two candidate transform kernels, and determining the cost results corresponding to the at least two candidate transform kernels respectively includes: Perform transform and quantization on the residual block of the current block based on the first candidate transform kernel, determine the first candidate quantization coefficient of the current block, and perform entropy coding processing on the first candidate quantization coefficient to determine the first generation value of the first candidate transform kernel; Inverse quantize and inverse transform the first candidate quantization coefficients to determine the first candidate residual block of the current block, and determine the first candidate prediction block of the current block according to the first candidate residual block; calculate a cost based on the first candidate prediction block and the original image of the current block to determine the second-generation cost value of the first candidate transform kernel. Determine the cost result corresponding to the first candidate transform kernel according to the first-generation cost value and the second-generation cost value of the first candidate transform kernel; wherein, the first candidate transform kernel is any one of the at least two candidate transform kernels.

53. The method according to claim 49 or 51, wherein, The method further includes: Determine the transform kernel index number of the current block; wherein, the transform kernel index number is used to indicate the number of the transform kernel of the current block in the first candidate list or the transform kernel group of the current block. Perform encoding processing on the transform kernel index number of the current block, and write the obtained encoded bits into the bitstream.

54. The method according to claim 49 or 51, wherein, The method further includes: Determine the value of a first syntax element; wherein, the first syntax element is used to indicate whether the current block uses a first transform mode and the corresponding transform kernel index number used, and the transform kernel index number is used to indicate the number of the transform kernel of the current block in the first candidate list or the transform kernel group of the current block. Perform encoding processing on the value of the first syntax element, and write the obtained encoded bits into the bitstream.

55. The method according to claim 50, wherein, The determining the transform kernel group of the current block according to the first candidate list includes: Perform encoding cost calculation on the current block according to at least two candidate transform kernel groups indicated by the first candidate list, and determine the cost results corresponding to the at least two candidate transform kernel groups respectively. Determine the minimum cost result among the cost results corresponding to the at least two candidate transform kernel groups respectively, and determine the candidate transform kernel group corresponding to the minimum cost result as the transform kernel group of the current block.

56. The method according to claim 55, wherein, The method further includes: Determine the feature index number of the current block; wherein, the feature index number is used to indicate the number of the transform kernel group of the current block in the first candidate list. Perform encoding processing on the feature index number of the current block, and write the obtained encoded bits into the bitstream.

57. The method according to claim 31, wherein, The method further includes: When the prediction mode of the current block does not meet the first condition, determine the texture feature index of the current block. Determine the transform kernel of the current block according to the texture feature index. Determine the transform coefficients of the current block, and perform inverse transform on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block.

58. The method according to claim 57, wherein, The prediction mode of the current block not meeting the first condition includes: the prediction mode of the current block is a prediction mode outside the first prediction mode set.

59. The method according to claim 57, wherein, The prediction mode of the current block not meeting the first condition includes: the prediction mode of the current block is one of the second prediction mode set. The second prediction mode set at least includes: DC mode, PLANAR mode, angular prediction mode, and cross-component prediction mode.

60. The method according to any one of claims 31 to 59, wherein, The determining the residual block of the current block includes: Perform intra prediction on the current block to determine the prediction block of the current block. Determine the residual block of the current block based on the initial block of the current block and the predicted block of the current block.

61. The method according to claim 60, wherein, Encoding the transform coefficients of the current block and writing the obtained encoded bits into the bitstream includes: Quantize the transform coefficients of the current block to determine the quantization coefficients of the current block; Encode the quantization coefficients of the current block and write the obtained encoded bits into the bitstream.

62. The method according to any one of claims 31 to 59, wherein, The step of determining the transform coefficients of the current block by performing a transform on the residual block of the current block according to the transform kernel includes: Performing a non-separable base transform on the residual block of the current block according to the transform kernel to determine the transform coefficients of the current block; or, Performing a discrete cosine transform on the residual block of the current block to determine the transform block of the current block; and performing a low-frequency non-separable transform on the transform block of the current block according to the transform kernel to determine the transform coefficients of the current block.

63. The method according to any one of claims 31 to 59, wherein, The method further includes: Encoding the prediction mode of the current block and writing the obtained encoded bits into the bitstream.

64. A bitstream, wherein, The bitstream is generated by bit encoding according to the information to be encoded; wherein, the information to be encoded includes at least one of the following: The quantization coefficients of the current block, the transform kernel index number of the current block, the feature index number of the current block, the prediction mode of the current block, and the value of the first syntax element; wherein, the first syntax element is used to indicate whether the first transform mode is used for the current block and the corresponding transform kernel index number used.

65. An encoder, the encoder includes a first determination unit, a first transform unit, and an encoding unit, wherein: The first determination unit is configured to determine the prediction mode of the current block; when the prediction mode of the current block satisfies a first condition, determine the first candidate list of the current block; wherein, the first candidate list indicates at least two candidate transform kernel groups; The first determination unit is further configured to determine the transform kernel of the current block according to the first candidate list; The first transform unit is configured to determine the residual block of the current block and perform a transform on the residual block of the current block according to the transform kernel to determine the transform coefficients of the current block; The encoding unit is configured to encode the transform coefficients of the current block and write the obtained encoded bits into the bitstream.

66. An encoder, the encoder includes a first memory and a first processor, wherein: The first memory is used to store a computer program that can run on the first processor; The first processor is used to execute the method according to any one of claims 31 to 63 when running the computer program.

67. A decoder, the decoder includes a second determination unit and a second transform unit, wherein: The second determination unit is configured to determine the prediction mode of the current block; When the prediction mode of the current block satisfies a first condition, determine the first candidate list of the current block; wherein, the first candidate list indicates at least two candidate transform kernel groups; The second determination unit is further configured to determine the transform kernel of the current block according to the first candidate list; The second transformation unit is configured to determine the transformation coefficients of the current block, and perform an inverse transformation on the transformation coefficients of the current block according to the transformation to determine the residual block of the current block.

68. A decoder, the decoder includes a second memory and a second processor, wherein: The second memory is used to store a computer program that can run on the second processor; The second processor is used to execute the method according to any one of claims 1 to 30 when running the computer program.

69. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed by at least one processor, it implements the method according to any one of claims 1 to 30, or the method according to any one of claims 31 to 63.